Rob, thank you very much for these questions. They are genuinely important to me. In addition to identifying several reasons why an association may reasonably be left unmade, your comment highlighted aspects of the pipeline that should be made explicit in the public diagnostics rather than remaining implicit in the implementation.
A few clarifications about the current MVP may be useful.
The initial filtering model does not rely only on alert-level parameters. It works with the Science, Template, and Difference cutouts and was trained on more than 20,000 examples that I labeled in a largely manual process. The most difficult part of that work was distinguishing moving sources from transients and static astrophysical sources. Whenever even a weak signal was visible in the Template, I measured the Science and Template centroids. If their positions agreed, I treated the example as a stationary or transient source rather than as a moving-object detection.
I do not yet cross-match every candidate against external transient catalogs. At present, the first-line rejection is primarily image-based and learned from those labeled cutout triplets. This filtering reduces the current pool of approximately 3.85 million unassociated alerts to about 952,000 candidate moving-object detections before orbital association begins.
The cutouts remain available for inspection on every individual DiaSource page. For example:
https://sso.fullnode.pro/dia-source/170631103303385438
This also makes it possible to identify cases where a geometrically plausible match is actually a streak, a subtraction artifact, or another unsuitable source morphology.
Regarding repeated detections, 5,005 of the 6,579 orbits with pipeline-added observations—approximately 76%—have at least one night containing multiple detections. This is not yet an exact measurement of recovery in both members of the nominal visit pair, and your question made it clear that paired-visit support should be exposed as a separate diagnostic.
For comparison, 99.998% of the currently active X05-associated orbits in MPC have at least one multi-detection night. I consider that a useful benchmark, although not a property that the current recovery sample fully matches.
I also do not force an assignment when a detection remains genuinely ambiguous between multiple known orbits. During the current processing, four conflict groups were identified. One involved a streaked, fast-moving near-Earth source rather than a point-like detection suitable for this association branch, so it was excluded. Three groups remain explicitly marked as unresolved because the observations fit both competing orbits well:
2024 RX153 / 2000 KU53 — 5 observations;
2026 DT28 / 2026 DJ35 — 44 observations;
2026 DY28 / 2026 DX31 — 270 observations.
Those observations have not been assigned to either competing orbit. A more detailed photometric or image-level analysis may eventually resolve them, but the important point for the current pipeline is that the ambiguity is detected, retained, and not silently converted into an association.
I have intentionally not used agreement with the predicted magnitude as a hard filter during the initial recovery stage. The current purpose of this layer is high-recall recovery: removing as many detections of already known objects as possible from the unknown pool before beginning a search for genuinely new Solar System objects.
Stricter photometric checks, centroid remeasurement, aperture SNR measurements, and an ADES submission-quality policy belong to a later stage, where the optimization changes from completeness to precision. My current approach is to first collect observations that are strongly consistent with a known orbit, and then apply a much stricter policy before any observation could be considered suitable for submission. It is easier to remove a questionable point at that stage than to recover a real observation that was prematurely discarded.
The accepted observations nevertheless show strong internal orbital consistency. For a clean sample of 6,389 asteroid orbits and 263,156 pipeline-added observations, the median two-dimensional residual is 0.076 arcseconds. Of those observations, 95.2% are below 0.2 arcseconds and 99.87% are below 0.5 arcseconds. Only two used observations are slightly above 1 arcsecond.
I do not treat a small residual by itself as proof against source confusion or an alternative association, since the observations are selected and then participate in the fit. However, these results do show that the final catalog is not simply the result of a positional cross-match inside the initial search radius.
On the faint end, I agree that completeness must degrade near the single-visit detection limit. One member of a visit pair may pass the detection threshold while the other does not, causing pairs or individual observations to disappear before the association stage.
I think it is useful, however, to separate that loss of input completeness from the reliability of a DiaSource that has already been generated. When an alert has passed the approximately 5-sigma detection threshold, shows a clean signal in the cutouts, is classified consistently by the image model, and fits a multi-observation orbit with small residuals, proximity to the limiting magnitude alone does not necessarily make the association unreliable. In that situation, much of the loss may already have occurred upstream, when another weak observation failed to become an alert at all.
I fully accept your broader point: a useful independent pipeline should not only report the associations it can make, but should also help explain why Rubin or the official SSSC processing may reasonably have chosen not to make them.
The system already retains rejected fits, competing orbits, duplicate assignments, unstable seeds, and unresolved observations. After reading your comment, I see that these rejection reasons should become a much more visible and structured part of the public diagnostics.
Thank you again. This is exactly the kind of feedback that helps turn a working experiment into a more testable and useful system.