APOGEE Pipeline

No new APOGEE Data in DR20

From the beginning of APOGEE observations, many stars were observed at several epochs. An observation at a given epoch is called a ‘visit’. The separate visits from a given telescope in SDSS-V are combined into a single spectrum at the end of processing. In other words, the APO and LCO spectra of the same object are not combined. Spectra taken before SDSS-V, in APOGEE-1 and APOGEE-2, were processed by a previous version of the pipeline and released in DR17. These spectra have not been reprocessed with the current tag, but they have been analyzed by Astra for stars that were not re-observed in SDSS-V. Comparison of ASPCAP results from DR17 and DR19 for the same stars in Meszaros et al. 2025 show good agreement in derived parameters despite the differences between the surveys in the hardware (e.g., FPS vs plates) and in the APOGEE and the ASPCAP/Astra pipelines. Note that this comparison is still valid for DR20 because the APOGEE spectra are copies of the ones in DR19 and have not been reduced with an updated pipeline.

The original APOGEE data reduction pipeline (DRP) is described in Nidever et al. 2015. Below are the main processing steps, with notes on updates. Clicking on the names of files will take you to their datamodels.

Visit Processing Steps

1. Plan files

The pipeline uses human readable “plan files” which describe chunks of data. In the past these corresponded to data taken for a single plate. Now they correspond to exposures taken on a given field and robot configuration. The plan files also include the names of the calibration products to use for the reduction of the exposures.

2. Calibrations

There are master/super calibration products that are updated every year or so and daily calibration products that are produced daily. The daily products are PSF, Flux, Wave and WaveFPI. They are processed in that order but the jobs are parallelized on the Utah computer cluster.

3. APRED

The heart of the reduction pipeline is apred. It has three main steps:

  • AP3D: Collapse of the up-the-ramp data cube.
  • AP2D: Extraction of spectra.
  • AP1DVISIT: Sky and telluric correction, dither combination.

The processing is parallelized over plan files/fields.

AP3D takes the 3D up-the-ramp data cube for each exposure and detector and collapses it to a 2D image. It tried to remove cosmic rays and fix saturation (as much as possible). Dark current is removed, linearity corrections are applied, and the final 2D image is flat fielded. It outputs the ap2D file.

AP2D extracts the 300 fibers to 1D from the 2D images. The extraction currently uses empirical point spread function (EPSF) profiles that are generated from domeflats taken during the night. During the plate era, a domeflat was taken for each plate and it was used for extraction. To cut down on overhead during the FPS era, these between-exposure domeflats are no longer being taken, and we are instead using a “library” of historical flats (also regularly updated with flats taken at the start and end of each observing night) that serves as a more sophisticated model of the PSF. A relative flux calibration using a daily apFlux calibration file (generated from a domeflat) is applied to remove fiber-to-fiber throughput and relative spectral response (see the Flux Calibration discussion at readthedocs for more details). At the end of AP2D, a wavelength solution is “attached” to the output ap1D file. The wavelengths are correct/shifted using night sky emission lines if the exposure was taken on sky. See the FPI discussion at readthedocs for more information on how the new Fabry-Perot Interferometer is used to improve the wavelength solutions.

AP1DVISIT first performs corrections on each exposure and then combines multiple exposure at the end. Sky fibers are used to remove the sky continuum and line emission from each fiber. Hot star (“telluric”) spectra are used to fit a model of the telluric absorption from CH4, CO2 and H2O and how it changes across the field. This is then used to divide out the telluric absorption in each spectrum. Because APOGEE spectra are slightly undersampled, observations are taken at two different spectral dither shifts, offset by ~0.5 pixels in the spectral dimension, to recover good sampling. To combine these two sets of exposures, a precise dither shift is calculated between the exposures using the data themselves. The spectra are then combined on a fiber-by-fiber basis using the measured shifts and sinc interpolation. Finally, the spectra are flux calibrated using a 5th order polynomial fit to the telluric stars to get the correct spectral shape, and the 2MASS H-magnitudes to set the absolute flux scale (see Flux Calibration at readthedocs for more details). The final product is apVisit files for each star.

Visit Combination

Visit Combination Steps

  1. Determination of Doppler shift (i.e., radial velocity) for each visit spectrum.
  2. Resampling of each visit spectrum onto the same rest wavelength scale to remove the Doppler shift
  3. Removal of continuum from each visit spectrum
  4. Weighted combination of all the resampled visit spectra.
  5. Re-application of mean continuum shape.

Radial Velocity Determination

Beginning with DR17, radial velocity (RV) determination for APOGEE spectra has used the Doppler code (Nidever 2021). Doppler uses a model to provide a continuous set of template spectra for a simultaneous template+RV fit, as well as an effort to improve radial velocities for faint stars.

Resampling

For spectral combination, the individual Doppler shifts for each visit spectrum are removed, and then the visit spectra are resampled onto the same wavelength grid. Using the radial velocity determined as described above, the wavelengths are corrected to the rest wavelength values λrest=λobs1+RV/c\lambda_{rest}=\frac{\lambda_{obs}}{1+RV/c}, where c is the speed of light). Once corrected to the rest wavelength scale, the spectrum is resampled onto a final logarithmically-spaced wavelength scale using interpolation with a sinc function. As in previous data releases, the combined star wavelength scale, in vacuum wavelengths, is given by:
log10λi+1log10λ=6.0×106log_{10}\lambda_{i+1}-log_{10}\lambda=6.0\times10^6

with a starting wavelength of 15100.802 A˚\AA (log10λ0=4.719log_{10}\lambda_0=4.719). This provides roughly 3 samples per resolution element, and corresponds to a velocity sampling of about 4.145 km/s per pixel.

Weighted Combination

To create a final combined spectrum for each star, we do a weighted sum of the individual visit, rest-frame shifted, and resampled spectra. To avoid biasing the weighted sum in the case of variations in the flux calibration, we first remove the mean continuum (by dividing out a median-filtered spectrum) from each visit.

The combination is done in two different ways: (1) pixel-by-pixel weighting, where each pixel is weighted by its inverse variance, and (2) “global” weighting, where each visit spectrum is weighted by a median-filtered inverse variance. In both cases, bad pixels in individual visit spectra are discarded before the combination. Also, the uncertainties of pixels affected by persistence are inflated, resulting in reduced weight of these pixels to the combined spectrum. The two schemes give similar results in most cases, but both are saved in the apStar FITS file.

The apStar and apVisit files are used by Astra to determine stellar parameters for all the APOGEE data analysis pipelines.

Finally, the combined spectra are multiplied by the average (over the multiple visit spectra) of the continuum shape to return the continuum to the combined spectrum. These apStar FITS file are the final output of the APOGEE Data Reduction Pipeline.

Radial Velocity Determination

The Doppler code uses a “Cannon” model (Ness et al. 2015 and Casey et al. 2016) that allows one to quickly generate a template spectrum for arbitrary choices of effective temperature, surface gravity, and metallicity. This model is trained on a set of synthetic spectra generated using the Synspec spectral synthesis code (Hubeny et al. 2011) with Kurucz model atmospheres.

The Doppler code first resamples the observed visit spectra and cross correlates each visit spectrum against a small set of templates that span a range of stellar parameters. The cross correlation that yields the highest peak gives a set of starting guesses for the stellar parameters of the object and for individual visit radial velocities. These are passed to a non-linear least squares routine that then simultaneously solves for a set of stellar parameters (Teff, log g, [Fe/H]) and individual radial velocities. This fit is done on the “raw” (i.e., not resampled) visit spectra, and all spectra are fit simultaneously; the pipeline LSF characterization is used by the routine to convolve the model spectra so that they can be directly compared with the data.

After the best fit values are found, a template is generated and cross-correlated against resampled individual visit spectra. The resulting cross-correlation functions are saved in the apStar files (see below). While these are not used to provide radial velocities, the agreement between this cross-correlation RV and the best-fit RV are used as a diagnostic to flag suspicious or bad radial velocities. The cross-correlation functions can also be used to identify potential spectroscopic binaries.

For faint stars, this procedure can often fail, finding significantly discrepant radial velocities. To improve the success for these, a procedure has been developed in which an initial combined spectrum is made under the assumption of no radial velocity variations, i.e., the spectra are combined after shifting to the barycentric frame with no additional velocity shifts. This spectrum is then run through Doppler to provide an estimate of a systemic velocity. The full set of individual spectra are then passed to Doppler for the normal procedure, but the fits are constrained so that only velocities within 50 km/s of the initial systemic estimate are allowed. This generally provides significantly improved radial velocities, as confirmed, in particular, by assessing the radial velocities of stars in nearby dwarf galaxies.

Reference frame for radial velocities

In SDSS_V,the radial velocity relative to the solar system barycenter is reported as v_rad, not v_helio as it was in APOGEE-1 and APOGEE-2

Radial velocities in APOGEE are reported with respect to the center of mass of the Solar System – the barycenter. The individual exposures are corrected for the relative motion of the Earth along the line of sight to the star during each observation. This correction is called the “barycentric correction,” and it can be calculated very accurately (to m/s levels). When these corrections are applied to the absolute RVs, we determine the RV with respect to the barycenter or v_rad for short.

Spectroscopic Binary Identification

The cross-correlation functions can provide information about potential spectroscopic binaries, as such stars can yield double-peaked cross correlation functions if they are observed when the two stars have measurably different radial velocities. Following the work of Kounkel et al. (2021), we attempt to recognize these using autonomous Gaussian decomposition of the cross-correlation function. Our main goal is to identify the most-likely spectroscopic binaries, as these are the ones for which the stellar parameter and abundance determinations are likely to be most suspect, since we analyze spectra under the assumption that they can be modeled with the spectrum of a single star. We flag suspect systems with the MULTIPLE_SUSPECT bit (bit 21) in the STARFLAG bitmask in the allStar and allVisit files (which is also propagated into the spectrum_flags bitmask in the mwmAllStar and mwmAllVisit files); for a combined spectrum, if any of the individual visits have MULTIPLE_SUSPECT, then the combined spectrum has this bit set. Note that this determination is neither complete nor perfect: there may be spectroscopic binaries for which the flag is set, and there may be spectra for which the bit is set that are not spectroscopic binaries, although we have tuned things so that we expect more of the former (false negatives) than the latter (false positives).

For significantly enhanced detection and characterization of spectroscopic multiples that includes visual inspection, identification of individual components, and attempts to fit orbits, see this value added catalog (Kounkel et al. 2021)

Changes for SDSS-V

There have been many upgrades to the original APOGEE DRP for SDSS-V. In the last several years, nearly the entire pipeline has been translated from IDL to Python and “refactored” (improving the internal structure without changing the external behavior). A substantial amount of documentation was also added including a ReadTheDocs page.

In the past, arclamp exposure and wavelength solutions were created roughly every two weeks. In SDSS-V, the reduction pipeline was upgraded to process every arclamp taken on an average wavelength solution (averaging over arclamp exposures taken from ±3 days) on a daily basis. This required a substantial upgrade to the wavelength software to make it very robust and a culling of the arclamp linelists. In addition, the wavelength solutions were upgraded to make use of the new Fabry-Perot Interferometer (FPI) information. Full-frame (all 300 fibers) FPI exposures are taken during the afternoon and morning calibration sequences, while every science exposure has two fibers dedicated to FPI light which can be used to track small shifts throughout the night. The average daily wavelength solution is used to full-frame FPI exposures averaged over the 300 fibers to obtain very precise absolute wavelengths for each unique FPI line (there is no absolute wavelength for the FPI lines and they shift slightly over time). These calibrated FPI wavelengths are then used to refine the wavelength solution for each fiber by using the full-frame FPI information.

New quality assurance checks were written to evaluate the quality of every exposure taken. Only data passing these checks were processed through the rest of the pipeline. This helps with the complete automation of the entire data reduction process. A new APOGEE DRP database was created on the Utah SDSS-V servers. This tracked every aspect of the DRP processing including the status of the processing steps and the results. This makes it easier for SDSS-V operations staff and collaboration members with Utah accounts to access reduction results directly.

During the plate era, every plate visit required a separate dome flat exposure. This was used to: (1) trace out the positions of the spectra on the detector and create an empirical PSF, and (2) measure the fiber-to-fiber throughput (mostly used for sky subtraction). During the FPS era, the gang connector is left connected in the FPS throughout the entire night and dome flats are only taken at the beginning and end of the night (mainly for fiber-to-fiber throughput measurements). As a result, another procedure was needed to generate the empirical PSF for each exposure (as the traces move slightly throughout the night, especially at APO). Therefore, a “model PSF” was created which can be used to generate an empirical PSF for any exposure by measuring the trace positions on the data itself. The quicklook software on the mountain was rewritten from IDL to Python and was vastly simplified and sped up. A detailed description of the DRP software upgrades will be given in the APOGEE data reduction software paper (Nidever et al.,in prep.).

Output Files

The APOGEE DRP produces the following files, ordered from most processed to least processed. Clicking the name of the file takes you to the data model

NameDescription
allStarSummary information from the DRP (e.g. radial velocity scatter) for each star in the data release
allVisitSummary information from the DRP (e.g., individual radial velocities) for every visit of each star in the data release
apStarCombined data from multiple visits, and combine the data from the three chips. Spectra are resampled onto a simple logarithmic wavelength scale. Both combined spectra as well as the resampled individual visit spectra are output.
apVisitData from each visit for a star, along with an initial RV measurement; this is the level where files are split by fiber with an associated format change: the three chips are now included as separate rows in each HDU (since these contain data only from a single fiber). Native wavelength sampling is still preserved; wavelength solution coefficients are now moved to individual headers.

Browsing the list of data models with the Environment Variable = APOGEE_REDUX details the many other files that the APOGEE DRP produces, if you are interested in the intermediate data products.

Back to Top