Hello,
For the past couple of weeks, since the official start of LSST, I have been reviewing the results of the ALeRCE stamp classifier (Stamp Classifier Rubin Beta, 20260421). I noticed that some objects tagged as “VS” (variable stars) by the classifier are, in fact, QSOs/AGNs.
Are there ways that the stamp classifier could be improved? It does classify plenty of known variable stars correctly, but it isn’t the most reliable resource if I were searching for valid candidates for new variable stars.
On Day 1 of the RCW last week, there was a session, “Early Science: Data Preview 2 and Prompt Products,” organized by Melissa Graham, Eric Bellm, and Bob Blum, that touched on your question. Check out the recorded presentation on YouTube:
The talk touches on the issue of false positives. In the discussion at the end, at ~1h 4m 14s (the link should take you there), someone (it sounds like Michael Strauss—I attended remotely so I’m not sure) asked about the false positive issue. The respondent (Eric?) points out that variable star classification reliability is particularly affected by some template issues, and that “we’re continuing to improve the training bit by bit.” Listen to the discussion for a bit more detail. They also mention that the topic will be addressed in an upcoming paper. I don’t know how directly this explains what you’re seeing from ALeRCE in particular.
I’ll ping Eric to make sure he sees your question: @ebellm. Melissa @MelissaGraham has an eagle eye on the forum, so perhaps she will weigh in as a co-presenter.
Hi @kkruszynska and @TomLoredo, the ALeRCE stamp classifier is independent of the Rubin reliability scoring. @fforster can comment on the ALeRCE team’s plans for tuning it further.
I tend to work with Lasair, looking for variable stars, and find the opposite quite commonly - known VS flagged as AGN, though in Lasair this is done with contextual crossmatching with external catalogs.
I would agree with previous comments that this is likely early data teething problems. Templating, slightly askew photometric calibrations in the pre-LSST visits, and mismatching with different catalogs. We recently released a paper touching on some of this, looking at RR Lyrae (https://ui.adsabs.harvard.edu/abs/2026AJ....172...84N/abstract).
Hopefully, now that LSST has started, the stability of observations and improved completeness of templating will help with automated classification going forward.
Thanks for pointing this out. I inspected the predictions by ALeRCE Stamp Classifier Rubin Beta 20260421 against some catalogs of known objects for the currently available sky regions, and I noticed the mismatch that you mention. However, the fraction of catalog QSOs/AGNs incorrectly predicted as “VS” is quite small (~5% of objects catalogued as QSOs/AGNs). Meanwhile, the fraction of variable stars incorrectly predicted as “AGN” is larger (~40% of objects catalogued as variable stars), in agreement with what @steventgk mentions about Lasair. If I remove objects located in regions that have recently available templates, that fraction drops from ~40% to ~24%, where “recently” means after we released this stamp classifier version (April). That difference makes sense, because recently available templates include a big ecliptic field and regions closer to the Galactic plane, so variable star predictions that worked well in regions seen by our classifier in training (i.e., DDFs) later worked not so well in regions not seen by it.
In ALeRCE we’re currently working on updating both our training sets and stamp classifier, that may help alleviate these discrepancies. However, note that the main focus of ALeRCE stamp classifiers for Rubin is so far the same as for ZTF, in the sense that it is aimed to find transient candidates early after explosion/first light. A light curve classifier, better suited to do science for more classes besides transients, is also in the works, but the limited/changing LSST cadence so far has made the training process harder.
I’m marking this reply as the solution, but please feel free to unmark it if more info is needed.