TechOtsu Research

TechOtsu Research

100 problems. 7 domains. Real physics. Real products.

This document is a walk through one hundred real problems in astrophysics β€” the kind that sit on the desks of people who build spacecraft trajectories, hunt for exoplanets, map the Milky Way, and read the light of galaxies billions of years old. It is written for someone who is curious, comfortable with basic math, and wants to understand not just the physics but the software that physics becomes. After reading it, you will be able to explain why a three-body problem cannot be solved with pen and paper, why a neural network that ignores conservation of energy will quietly get worse the longer it runs, and how a handful of these ideas turn into products that research groups, universities, and mission-design teams actually pay for.

Mathematics is the language of physics. In this document, we will not hide the equations. We will explain every symbol. If you see something unfamiliar, keep reading β€” it will be explained.

References used throughout this document:

  • Classical Mechanics β€” Herbert Goldstein
  • Solar System Dynamics β€” Murray & Dermott
  • Galactic Dynamics β€” Binney & Merrifield
  • The Feynman Lectures on Physics β€” Feynman, Leighton & Sands
  • Machine Learning for Physics β€” Mehta et al.
  • Data Analysis in Astronomy β€” various
Track A Β· 15 Problems

Multi-Body Orbital Dynamics & N-Body Acceleration

How do we predict where everything in the solar system is β€” and why is it so hard?

Module A.0 β€” Before we begin: what is an orbit?

Start from the simplest possible version. An orbit is not a planet going in a circle. It is a planet that is falling β€” but falling sideways so fast that it keeps missing the Earth. Newton figured this out. If you fire a cannonball horizontally from a tall mountain, it will travel some distance before hitting the ground. Fire it faster, it goes farther. Fire it at exactly the right speed, it goes so far that the Earth curves away beneath it at the same rate it is falling. It never lands. That is an orbit.

F = GMm / rΒ²

Here F is the gravitational force between two masses, G is the gravitational constant (6.674 Γ— 10⁻¹¹ NΒ·mΒ²/kgΒ²), M and m are the two masses, and r is the distance between their centers. Double the mass of either object, and the force doubles β€” gravity is directly proportional to mass. But distance behaves differently: it appears as rΒ² in the denominator, so doubling the distance does not halve the force β€” it cuts it to a quarter. This is the inverse square law, and the intuition is the same as light spreading out from a lamp: at twice the distance, the same amount of light is spread over four times the area (because area scales with the square of distance), so the intensity you receive is one quarter. Gravity spreads out through space the same way.

What is a two-body problem? Two masses in space, each pulling on the other. This is solvable exactly, and the solution is an ellipse β€” this is Kepler's first law. An ellipse is described by its semi-major axis a (half the longest diameter β€” roughly the "average" distance from the central body) and its eccentricity e (how stretched the ellipse is: e = 0 is a perfect circle, e close to 1 is a long thin sliver). Earth's orbit has a = 1 AU (149.6 million km, the definition of one Astronomical Unit) and e = 0.017 β€” nearly circular. Compare that to Halley's Comet, with e = 0.967 β€” a needle-thin ellipse that swings from just inside Venus's orbit out past Neptune.

What is a three-body problem? Three masses, all pulling on each other simultaneously. In 1887, Henri PoincarΓ© proved this has no general closed-form solution β€” no equation you can write down that gives position as a function of time for arbitrary starting conditions. The differential equations exist. They are just not solvable with pen and paper. This is why every space agency on Earth runs numerical simulations rather than solving equations by hand.

The insight that connects to TechOtsu Research: we do not need to solve the equations analytically. We need to compute the answer fast and accurately. That is the N-body acceleration problem, and it is what Problems A-1 through A-15 address.

Problem 1: Fast trajectory prediction for the three-body problem

The Circular Restricted Three-Body Problem (CR3BP) is the most practical simplification of the three-body problem for space mission design: one very large mass (the Sun), one medium mass (Jupiter, say), and one tiny mass (a spacecraft) so small that its own gravity doesn't affect the other two. Every time you want to know where the spacecraft will be after time T, you have to integrate the differential equations step by step β€” a process called numerical integration. It works, but it is slow: a single high-fidelity trajectory can take hours of compute time to evaluate at the precision mission designers need.

A Hamiltonian is the total energy of a system, written as the sum of kinetic and potential energy:

H = (pΒ² / 2m) + V(q)

where p is momentum, q is position, m is mass, and V(q) is the potential energy at position q. For a conservative system β€” no friction, no energy loss β€” H stays constant over time. This is conservation of energy. Hamilton's equations describe how position and momentum evolve:

dq/dt = βˆ‚H/βˆ‚p dp/dt = -βˆ‚H/βˆ‚q

These two equations fully describe the motion of any conservative system, and they are mathematically equivalent to Newton's F = ma β€” just expressed in terms of energy instead of force. Their advantage is that they naturally conserve energy by construction.

A Hamiltonian Neural Network (HNN) is a neural network trained to output H β€” the Hamiltonian β€” given the current state (position and momentum). Once you have H, you use Hamilton's equations to evolve the state forward. The key property: because the network predicts the energy function itself rather than the trajectory directly, the resulting simulation automatically conserves energy. A classical neural network trained to predict position directly will drift β€” the simulated energy creeps upward or downward over time, and the orbit slowly spirals in a way that never happens in reality. HNNs do not drift, because the very structure of the network's output guarantees the conservation law. (Reference: Greydanus, Dzamba & Yosinski, 2019, "Hamiltonian Neural Networks" β€” a paper worth reading in full.)

Case Study

The Voyager missions (1977) are what precise orbital mechanics looks like in practice. NASA used classical numerical integration to plan Voyager 1 and 2's trajectories through the outer solar system, exploiting gravitational assists from Jupiter, Saturn, Uranus, and Neptune. Every maneuver was computed by hand and then verified by computer simulation. The trajectories were so precise that Voyager 2 hit Neptune's atmosphere within 100 km of its target after an 8-billion-kilometre journey β€” a targeting error smaller than the width of a small city, at the end of a twelve-year flight.

B2B Pitch

Candidate Project: solve_ivp-Compatible Orbit-Propagation Shim. Target client: any Python-based astrodynamics research team already using SciPy. The pain: their entire codebase calls scipy.solve_ivp for trajectory integration β€” replacing it means rewriting everything. The pitch: swap one import line, get a physics-preserving HNN underneath. Every existing script keeps working. Faster. More accurate over long horizons. Why they pay: it is not a new tool, it is a better version of the exact tool they already use.

Visual Output

A 3D plot showing the computed trajectory of a spacecraft through the Sun-Jupiter-Earth system. One trajectory in blue (classical RK45 integrator β€” notice the slow energy drift over 1,000 orbital periods). One trajectory in orange (HNN-based β€” flat energy, no drift). Below: an energy conservation plot of H vs. time. The classical solver's line slowly climbs; the HNN's line stays flat. This is the product demonstration.

Problem 2: Energy drift in long-horizon N-body simulation

Standard numerical integrators β€” Euler, Runge-Kutta β€” are not designed to preserve the geometric structure of Hamiltonian mechanics. Over millions of timesteps, tiny numerical errors accumulate. Energy drifts. Orbits slowly expand or contract. For a short simulation this doesn't matter. For a billion-year planetary evolution simulation, it destroys the result.

A symplectic integrator preserves the "symplectic structure" that Hamiltonian systems naturally possess β€” a geometric property of phase space that, when preserved, keeps the numerical solution bounded near the true energy surface instead of drifting away from it. The simplest example is the leapfrog (Verlet) integrator:

v(t + h/2) = v(t) + (h/2) Γ— a(t) x(t + h) = x(t) + h Γ— v(t + h/2) a(t + h) = F(x(t + h)) / m v(t + h) = v(t + h/2) + (h/2) Γ— a(t + h)

where h is the timestep, x is position, v is velocity, and a is acceleration. The leapfrog alternates between updating position and velocity β€” "leaping" over each other in half-steps β€” and this alternation is what preserves the symplectic structure. Energy is not exactly conserved but oscillates in a small band around its true value without drifting. Over a billion-year simulation, the difference between a drifting and a non-drifting integrator is the difference between a solar system that stays put and one that silently unravels in the simulation while doing nothing of the sort in reality. (Reference: Murray & Dermott, Solar System Dynamics, Chapter 8 β€” the definitive reference for symplectic integration in planetary science.)

Note

Symplectic integrators are the workhorse of every long-term solar-system-stability study ever published β€” including the famous 5-billion-year integrations that showed Mercury's orbit has a small but nonzero chance of eventually becoming unstable.

Problem 3: Approximate force calculation for large N β€” the Barnes-Hut algorithm

Computing gravity directly between N bodies requires N(Nβˆ’1)/2 pairwise force calculations β€” for N = 1 million particles, that's roughly 500 billion calculations per timestep. This is the direct-summation approach, and it does not scale. The Barnes-Hut algorithm (1986) fixes this by building an octree: recursively dividing space into cubes until each contains at most one particle. To compute the force on a given particle, distant clusters of particles are treated as a single point mass at their center of mass, while nearby particles are still summed individually.

accept cluster as single mass if: s / d < ΞΈ

where s is the size of the cluster's bounding box, d is the distance from the particle to the cluster's center of mass, and ΞΈ (theta) is an accuracy parameter, typically 0.5–1.0. This reduces the cost per timestep from O(NΒ²) to O(N log N) β€” for a million particles, a reduction from 500 billion operations to roughly 20 million. This single algorithmic change is why galaxy-scale N-body simulations with billions of particles are possible at all on modern hardware.

Problem 4: GPU-accelerated direct summation for small, dense clusters

Barnes-Hut wins for large, sparse systems, but for small dense systems β€” a globular star cluster with 100,000 tightly packed stars, or the near-field of a planetary encounter β€” direct O(NΒ²) summation is often simpler and, on the right hardware, competitive. GPUs are extremely good at exactly this kind of embarrassingly parallel arithmetic: computing the same force formula millions of times independently. A single modern GPU can evaluate on the order of 10ΒΉΒ² pairwise gravitational interactions per second, because every pair is independent and can be computed on a separate thread simultaneously.

a_i = Ξ£β±Ό GΒ·mβ±ΌΒ·(rβ±Ό βˆ’ rα΅’) / (|rβ±Ό βˆ’ rα΅’|Β² + Ρ²)^(3/2)

Here a_i is the acceleration of particle i summed over every other particle j, and Ξ΅ is a small "softening length" added to prevent the force from blowing up to infinity when two particles pass extremely close to each other β€” a numerical safety valve, not a physical effect.

Problem 5: Adaptive timestep control for close encounters

A fixed timestep is wasteful: most of the time, planets move slowly relative to the timescale of their orbits, and a large timestep is fine. But during a close encounter β€” two asteroids passing near each other, or a spacecraft swinging close to a planet β€” the local dynamics happen much faster, and the same large timestep produces garbage. Adaptive timestep control shrinks h automatically based on local conditions, typically using the free-fall time:

h_i = Ξ· Γ— √(Ξ΅ / |a_i|)

where Ξ· is a small tunable safety factor and |a_i| is the current acceleration magnitude of particle i. When acceleration spikes during a close pass, the timestep automatically shrinks to keep the integration accurate, then grows back once the encounter is over. Getting this wrong is one of the most common sources of silently wrong N-body results in published research.

Problem 6: Satellite conjunction screening β€” will two objects collide?

There are more than 30,000 tracked objects in Earth orbit β€” active satellites and debris β€” and space agencies must continuously check every close pass ("conjunction") between them. The core computation is a probability of collision, not a yes/no answer, because both objects' positions are known only to some uncertainty.

P_c = ∬_A (1 / (2Ο€ Οƒ_x Οƒ_y)) exp(βˆ’xΒ²/2Οƒ_xΒ² βˆ’ yΒ²/2Οƒ_yΒ²) dx dy

This integrates a 2D Gaussian probability density β€” representing the combined position uncertainty of both objects, projected onto the plane perpendicular to their relative velocity β€” over the area A defined by the combined physical size of the two objects (their "hard body radius"). Οƒ_x and Οƒ_y are the position uncertainties along each axis. If P_c exceeds a threshold (commonly 1 in 10,000), an operator maneuvers the satellite. Running this calculation for every pair of 30,000+ objects, repeatedly, as new tracking data arrives, is itself a nontrivial computational problem β€” a natural candidate for the same fast-propagation techniques described above.

Problem 7: Halo orbits and the Lagrange points

In the CR3BP, five special equilibrium points exist where the combined gravity of the two large bodies and the centrifugal effect of the rotating frame balance out β€” the Lagrange points, L1 through L5. L1 sits between the two bodies; L4 and L5 form equilateral triangles with them and are stable, which is why Jupiter's Trojan asteroids cluster there. Spacecraft don't sit exactly at these points β€” they fly in halo orbits, closed 3D loops around them. The James Webb Space Telescope orbits Sun-Earth L2, 1.5 million km from Earth, chosen because it keeps the Sun, Earth, and Moon all behind the telescope's sunshield at once. Halo orbits are computed by finding periodic solutions to the CR3BP equations of motion β€” a search problem, not a closed-form one, typically solved with differential correction (a targeting algorithm that nudges an initial guess until the orbit closes on itself).

Problem 8: Secular and resonant perturbation theory

Real planets don't just feel one other body's gravity β€” they feel small tugs from every other planet. Perturbation theory treats these as small corrections to the two-body ellipse rather than re-solving the whole N-body problem from scratch. Perturbations split into two flavors: secular perturbations build up steadily over time (slowly rotating an orbit's ellipse, for instance), while resonant perturbations occur when two orbital periods form a near-integer ratio and repeatedly reinforce each other at the same point in each cycle. Both are captured by expanding the disturbing function β€” the extra potential energy from the third body β€” as a Fourier series in the orbital angles, then studying which terms grow and which cancel out over many orbits.

Problem 9: Orbital resonance β€” the Kirkwood gaps

The asteroid belt between Mars and Jupiter is not evenly populated. At specific distances from the Sun β€” where an asteroid's orbital period would be a simple fraction of Jupiter's (1/3, 2/5, 1/2) β€” there are almost no asteroids at all. These are the Kirkwood gaps, discovered by Daniel Kirkwood in 1866. A mean-motion resonance occurs when:

p Γ— n_asteroid β‰ˆ q Γ— n_Jupiter (p, q small integers)

where n is mean orbital motion (2Ο€ divided by orbital period). At resonance, Jupiter's gravitational tug lines up repeatedly at the same orbital phase instead of averaging out, and the accumulated kicks pump up the asteroid's eccentricity until it either collides with something or gets ejected. The gaps are direct, visible evidence of a purely gravitational effect playing out over millions of years β€” one of the cleanest demonstrations of orbital dynamics in the solar system.

Problem 10: Tidal forces and the Roche limit

Gravity is not uniform across an extended body β€” the near side of a moon feels a stronger pull from its planet than the far side, and this difference stretches the moon along the line connecting the two. This differential pull is a tidal force. If a moon gets too close to its planet, the tidal stretching exceeds the moon's own self-gravity holding it together, and it is torn apart. The distance at which this happens is the Roche limit:

d = R Γ— (2 Γ— ρ_planet / ρ_moon)^(1/3)

where R is the planet's radius, and ρ_planet, ρ_moon are the two bodies' densities. Saturn's rings sit almost entirely inside Saturn's Roche limit β€” direct evidence they are debris from a body that either formed there and never coalesced, or wandered too close and was shredded.

Problem 11: Low-thrust trajectory optimization for electric propulsion

Chemical rocket engines burn fuel in seconds and produce a large, near-instantaneous velocity change (a "delta-v burn"). Electric (ion) propulsion produces a tiny continuous thrust over months or years β€” far more fuel-efficient, but the trajectory design problem changes completely: instead of picking a handful of discrete burns, you must optimize a continuous thrust direction over the entire flight. This is an optimal control problem, typically solved by applying Pontryagin's Minimum Principle to derive a "costate" system alongside the equations of motion, then solving a two-point boundary value problem to find the thrust profile that minimizes fuel use (or flight time) subject to the spacecraft's dynamics. NASA's Dawn mission, which orbited both Vesta and Ceres using ion propulsion, is only possible because this optimization can be solved reliably.

Problem 12: Gravitational assist design via patched conics

A gravitational assist (flyby) uses a planet's gravity and orbital motion to bend and accelerate a spacecraft's trajectory without spending fuel β€” how Voyager reached the outer solar system. The standard first-pass design tool is the patched conic approximation: treat the trajectory as a sequence of separate two-body problems (spacecraft-and-Sun, then spacecraft-and-planet during the flyby, then spacecraft-and-Sun again), "patching" them together at the boundary of the planet's sphere of influence. Inside the flyby, the spacecraft's speed relative to the planet is unchanged, but its direction rotates β€” and because the planet itself is moving relative to the Sun, that rotation translates into a real change in the spacecraft's Sun-relative speed. This approximation is what mission designers use to rapidly screen thousands of candidate flyby sequences before handing the best few to a full numerical integrator for refinement.

Problem 13: Uncertainty propagation in orbit determination

No tracking measurement is perfect β€” every estimate of a satellite's position and velocity carries uncertainty, usually represented as a covariance matrix. As that state is propagated forward in time through nonlinear gravitational dynamics, the uncertainty ellipse stretches and rotates in complicated ways. Two approaches dominate: linear covariance propagation, which propagates the covariance matrix directly using the dynamics' local linearization (fast, but breaks down over long time horizons or strongly nonlinear regimes like close flybys), and Monte Carlo propagation, which samples thousands of individual trajectories from the initial uncertainty distribution and propagates each one independently (accurate at any horizon, but computationally expensive β€” exactly the kind of expense a fast surrogate propagator is built to cut).

Problem 14: A learned surrogate for ephemeris generation

An ephemeris is a table (or a callable function) giving a body's position at any requested time. JPL's Horizons system is the authoritative source, built from high-precision numerical integration of the entire solar system's equations of motion, including relativistic corrections. For applications that need millions of ephemeris queries β€” Monte Carlo conjunction screening, large-scale mission trade studies β€” querying Horizons directly is often too slow. A neural surrogate trained to reproduce Horizons output to sub-kilometre accuracy over a bounded time window can answer the same query orders of magnitude faster, trading a small, well-characterized accuracy loss for throughput.

Problem 15: Real-time onboard trajectory correction

A spacecraft executing an autonomous maneuver β€” landing on an asteroid, docking without ground-in-the-loop control, or reacting to an unplanned deviation β€” cannot wait for a round-trip radio signal to Earth (which can take tens of minutes for deep-space missions). It needs to compute a corrected trajectory onboard, on flight hardware far less powerful than a ground workstation, within a strict time budget. This is where every idea in this track converges: fast, energy-consistent, low-memory trajectory propagation is not a nice-to-have for these missions, it is the only thing that makes autonomous operation possible at all.

Track B Β· 15 Problems

Exoplanet Detection & Characterization

We know of over 5,600 planets outside our solar system. We found them all with math.

Module B.0 β€” How do we find planets we cannot see?

No telescope on Earth or in orbit can directly resolve most exoplanets β€” they are drowned out by the light of their host star, often by a factor of a billion or more. So astronomers infer their existence indirectly, through four main methods. The transit method watches for the tiny dimming of starlight as a planet crosses in front of its star. Radial velocity watches for the Doppler shift in the star's spectrum as the planet's gravity tugs the star back and forth in a small orbit of its own. Direct imaging β€” physically photographing the planet β€” works only for very large planets very far from their star, where the light can be separated. Transit timing variations (TTV) catch planets that don't even transit, by noticing that another planet's transits arrive slightly early or late due to gravitational tugging.

The transit method is responsible for the majority of confirmed exoplanets, and its signal is simple geometry β€” how much of the star's disk does the planet block?

Ξ΄ = (R_p / R_*)Β²

where Ξ΄ is the fractional dimming, R_p is the planet's radius, and R_* is the star's radius. For Earth transiting the Sun: R_p = 6,371 km, R_* = 696,000 km, so Ξ΄ = (6371/696000)Β² = 0.0000838, or 0.0084%. That is the signal we are looking for β€” a dimming less than one part in ten thousand. NASA's Kepler telescope measured stellar brightness to about 20 parts per million precision to catch signals this small β€” the equivalent of standing far enough from a stadium floodlight to notice it flicker by the width of a gnat crossing the bulb.

Problem 16: Transit light-curve fitting

A raw transit light curve β€” brightness vs. time β€” is not a clean box shape. It has a smooth ingress and egress (the planet doesn't block the star instantaneously), and the dip is shallower toward the edges of the transit than the center because of limb darkening: a star appears dimmer near its visible edge than at its center, since we are seeing cooler, higher layers of the stellar atmosphere at the limb. The standard model for this is the Mandel-Agol transit model (2002), which analytically computes the light curve for a limb-darkened star given the planet's radius ratio, orbital inclination, semi-major axis, and a limb-darkening law such as:

I(ΞΌ) / I(1) = 1 βˆ’ u₁(1 βˆ’ ΞΌ) βˆ’ uβ‚‚(1 βˆ’ ΞΌ)Β²

where ΞΌ = cos(Ξ³) is the cosine of the angle between the line of sight and the local stellar surface normal (ΞΌ = 1 at disk center, ΞΌ β†’ 0 at the limb), and u₁, uβ‚‚ are empirical limb-darkening coefficients for the star's temperature and the observing band. Fitting a transit model to noisy real data means finding the combination of parameters (planet radius, orbital period, inclination, limb-darkening coefficients) that best reproduces the observed light curve β€” and quantifying the uncertainty on each. This is why MCMC (Markov Chain Monte Carlo) is the standard tool: rather than finding a single best-fit answer, MCMC samples thousands of parameter combinations weighted by how well each fits the data, building up a full probability distribution for every parameter. This is statistically rigorous but slow β€” a single well-converged MCMC fit can take hours of compute time, and a large survey may have thousands of candidate systems waiting to be fit. A trained surrogate model that approximates the same posterior distributions in a fraction of the time is a direct, measurable win for any team running this pipeline at scale.

Case Study

The discovery of the TRAPPIST-1 system (2016–2017): seven Earth-sized planets around a dim red dwarf star, three of them in the habitable zone. All seven were found via transit photometry β€” first by a small ground-based telescope in Chile (TRAPPIST, hence the name), then confirmed and refined with the Spitzer and Hubble space telescopes. The transit light curves are publicly available, and this remains one of the most important planetary system discoveries of the last decade, precisely because the transit method scales down to Earth-sized, Earth-temperature worlds around small stars.

B2B Pitch

Fast Transit Light-Curve Fitting Add-On. Target: university astronomy departments with a backlog of unanalyzed Kepler/TESS light curves. The pain: MCMC fitting takes hours per system, and they have thousands of candidate systems. The pitch: cut fitting time by 100Γ—, with the same statistical rigor. More papers, faster.

Problem 17: Radial velocity β€” detecting the star's wobble

A planet doesn't just orbit its star β€” the star also orbits the system's common center of mass, in a much smaller loop. As the star moves toward and away from us over that small orbit, its spectral lines shift blue then red β€” the Doppler effect. The size of that wobble, the radial velocity semi-amplitude K, is:

K = (2Ο€G / P)^(1/3) Γ— (m_p sin i) / (M_* + m_p)^(2/3) Γ— 1/√(1βˆ’eΒ²)

where P is the orbital period, m_p is the planet's mass, M_* is the star's mass, i is the orbital inclination, and e is eccentricity. This is where 51 Pegasi b (1995) β€” the first exoplanet confirmed around a Sun-like star, and a Nobel-Prize-winning discovery β€” came from: a periodic 59 m/s wobble in the host star's spectrum, corresponding to a Jupiter-mass planet orbiting absurdly close to its star, once every 4.2 days. Radial velocity is complementary to the transit method: transits give you the planet's radius and require the orbit to be edge-on as seen from Earth; radial velocity gives you the planet's mass (times sin i) and works for a much wider range of orbital orientations.

Problem 18: Transit timing variations β€” finding planets that don't transit

If a system has two or more planets, their mutual gravity tugs each other slightly off a perfectly periodic orbit. A transiting planet's transits will then arrive a little early or a little late, cycle to cycle, and the pattern of those deviations β€” the transit timing variations β€” encodes the mass and orbital period of the perturbing planet, even if that second planet never transits at all. TTV analysis is a signal-processing problem: extract a small periodic deviation (often seconds to minutes) from a series of transit-center measurements, each with its own timing uncertainty, and fit it to an N-body model of the full system to recover both planets' masses simultaneously β€” something the transit method alone cannot do (transits alone give radius, not mass).

Problem 19: False positive vetting

Not every dip in a light curve is a planet. An eclipsing binary star system in the background, blended into the same pixel as the target star, can produce a transit-shaped dip. So can starspots, instrumental noise, or a grazing eclipse. Vetting pipelines check for tell-tale signs: does the dip depth differ between odd- and even-numbered transits (a sign of an eclipsing binary with two similarly-sized stars)? Does the transit shape match a limb-darkened planetary transit or a V-shaped stellar eclipse? Is there a secondary eclipse at the expected phase, of the right depth for a planet's reflected/thermal light rather than a stellar companion? Statistical validation tools like vespa compute the relative probability of a "planet" hypothesis versus every plausible false-positive scenario, given the light curve shape, the star's brightness, and its position on the sky.

Problem 20: Atmospheric characterization via transmission spectroscopy

During a transit, some starlight passes through the thin ring of the planet's atmosphere at its limb before reaching us. Different atmospheric gases absorb different wavelengths, so the transit depth Ξ΄ = (R_p/R_*)Β² varies slightly with wavelength β€” deeper at wavelengths where the atmosphere is more opaque, because the planet's effective radius is larger there. Measuring Ξ΄ across many wavelengths produces a transmission spectrum that reveals atmospheric composition: JWST has used this technique to detect water vapor, carbon dioxide, and sulfur dioxide (a marker of photochemistry) in exoplanet atmospheres light-years away, from a signal often just tens of parts per million deep.

Problem 21: Direct imaging and coronagraphy

Photographing a planet directly means separating light a billion times fainter from a source less than an arcsecond away β€” roughly like trying to photograph a firefly next to a lighthouse from ten kilometres away. A coronagraph physically blocks the star's light inside the telescope, while adaptive optics correct for atmospheric blurring in real time. The minimum angular separation a telescope can resolve is set by the diffraction limit:

ΞΈ = 1.22 Γ— Ξ» / D

where Ξ» is the wavelength of light and D is the telescope's aperture diameter. Direct imaging currently only reaches young, hot, Jupiter-sized planets on wide orbits β€” but it is the only method that captures a planet's actual spectrum directly, without needing a transit or a favorable orbital alignment.

Problem 22: Stellar activity jitter in radial velocity data

Starspots β€” cooler, darker regions on a star's surface β€” rotate in and out of view as the star spins, subtly distorting its spectral lines in a way that can mimic the radial velocity signal of a planet. This stellar jitter is one of the leading sources of false planet detections in RV surveys. Distinguishing a real planet from stellar activity means checking whether the RV signal's period matches the star's known rotation period (a red flag), and whether the signal correlates with independent activity indicators like the star's chromospheric emission lines β€” a cross-validation problem as much as a signal-detection one.

Problem 23: Mass-radius relationships and interior structure

Once you have both a planet's mass (from radial velocity) and radius (from transits), you can compute its bulk density and start to infer what it's made of β€” rock, water, gas, or some mix. Plotting mass against radius across the known exoplanet population reveals distinct populations: small rocky worlds, a puzzling gap around 1.5–2 Earth radii (the "radius valley," thought to mark the boundary where a planet's primordial hydrogen atmosphere is stripped away by stellar radiation), and a range of gas-rich sub-Neptunes and giants. Interior structure models solve for the internal pressure-density-composition profile consistent with the observed mass and radius, the same way geophysicists model Earth's core from seismic data.

Problem 24: Calculating the habitable zone

The habitable zone is the range of distances from a star where a rocky planet with the right atmosphere could sustain liquid water on its surface. Its boundaries are set by radiative balance β€” how much starlight the planet absorbs versus how effectively an atmosphere can retain heat before runaway greenhouse or global freezing sets in. A simplified inner/outer boundary (Kasting et al., 1993) scales with stellar luminosity:

d_hz = √(L_* / L_β˜‰) (in AU, for the Sun-equivalent boundary distances)

where L_* is the star's luminosity relative to the Sun's. This is a first-order estimate β€” real habitable zone boundaries depend on atmospheric composition, cloud cover, and planetary rotation, which is why habitable-zone modeling remains an active research area rather than a solved formula.

Problem 25: Gravitational microlensing

When a foreground star (with a planet) passes almost exactly in front of a distant background star, the foreground star's gravity bends and magnifies the background star's light β€” a real-time demonstration of general relativity. A planet orbiting the foreground "lens" star adds a brief, characteristic spike to the magnification light curve. Microlensing is the only method sensitive to free-floating planets with no host star at all, and to planets on wide orbits that transits and RV struggle to reach β€” but each event is a one-time, non-repeating alignment, so it demands continuous, wide-field monitoring of millions of stars to catch enough events to matter statistically.

Problem 26: Automated transit search β€” the Box Least Squares algorithm

Before you can fit a transit model, you first have to find the transit in a light curve containing years of data, instrumental drift, stellar variability, and noise. The Box Least Squares (BLS) algorithm searches systematically over a grid of candidate periods, transit durations, and phases, testing how well a simple box-shaped dip fits the data at each combination, and returns the combination with the strongest signal. Before BLS can run cleanly, the light curve must be detrended β€” long-term stellar brightness variations (rotation, pulsation, instrumental drift) removed without accidentally removing the transit signal itself, typically with a smoothing filter tuned to timescales much longer than the expected transit duration.

Problem 27: Dynamical stability screening for multi-planet systems

A newly proposed multi-planet system, fit from transit or RV data, isn't automatically physical β€” some parameter combinations, though consistent with the data within its uncertainties, would cause the system to fly apart or collide within a few thousand orbits. Stability screening runs N-body integrations (the same machinery from Track A) across the range of parameter uncertainty to check which combinations remain dynamically stable over the star's age, ruling out unphysical solutions and often tightening the allowed parameter ranges beyond what the raw data alone would give.

Problem 28: Biosignature detection

A biosignature is a spectral fingerprint that is difficult to produce or maintain without biological activity β€” the classic example is atmospheric oxygen alongside methane, two gases that react with each other and would disappear within a geologically short time unless something is continuously replenishing both (on Earth: photosynthesis and methanogenic life). Detecting this kind of chemical disequilibrium in an exoplanet's transmission spectrum, at signal levels of tens of parts per million, while ruling out abiotic processes that can mimic the same features, is one of the most consequential and most statistically demanding problems in the entire field.

Problem 29: Occurrence rates and survey completeness corrections

"How common are Earth-like planets?" is a harder question than it sounds, because every survey is blind to some planets β€” small planets produce weaker signals that get missed, and short surveys miss long-period planets that transit only once or twice. Computing a true occurrence rate requires correcting the raw detection count for the survey's completeness: the probability that a planet of a given size and period, if it existed, would actually have been detected. Getting this correction wrong β€” a common and subtle statistical trap β€” silently biases every downstream estimate of how common habitable worlds really are.

Problem 30: Machine-learning vetting of transit candidates

At the scale of Kepler and TESS β€” tens of thousands of candidate signals β€” human-in-the-loop vetting doesn't scale. Convolutional neural networks trained directly on light-curve shapes (Google's AstroNet, 2018, is the widely cited example) learn to distinguish real transits from instrumental artifacts and eclipsing binaries directly from the raw flux time series, without hand-engineered features, and were used to confirm new planets in archival Kepler data that traditional pipelines had missed.

Track C Β· 15 Problems

Stellar Astrometry & Galactic Dynamics

The Gaia mission has measured the position and velocity of 1.8 billion stars. What do we do with all of it?

Module C.0 β€” What is Gaia?

The European Space Agency's Gaia spacecraft has been measuring the positions, parallaxes, proper motions, and radial velocities of stars in the Milky Way since 2013. Data Release 3 (2022) contains 1.8 billion sources β€” the largest and most precise map of our galaxy ever made.

Parallax is how we measure distance to nearby stars. As Earth moves from one side of its orbit to the other over six months, nearby stars appear to shift slightly against the far more distant background stars β€” the same effect as holding a finger up and closing one eye, then the other. The size of that shift is the parallax angle Ο€, and distance follows directly:

d = 1 / Ο€ (d in parsecs, Ο€ in arcseconds)

Proper motion is a star's actual movement across the sky, measured in milliarcseconds per year. Gaia measures this to about 20 microarcseconds precision for bright stars β€” comparable to measuring the width of a human hair from a thousand kilometres away. Combined with radial velocity (from spectroscopy) and parallax distance, this gives the full 6D phase space for each star: (x, y, z, vβ‚“, v_y, v_z) β€” complete 3D position and velocity. With this, we can integrate stellar orbits backward in time and reconstruct the dynamical history of the galaxy itself.

Problem 31: A fast galactic orbit integrator

To integrate a star's orbit through the Milky Way, you need a model of the galaxy's overall gravitational potential β€” the standard approach uses a multi-component model (Bovy 2015, and the widely used galpy Python library) combining a disk, a central bulge, and an extended dark matter halo, each contributing its own term to the total potential Ξ¦(x, y, z). Integrating a single stellar orbit in this potential takes milliseconds. Integrating 1.8 billion stars, each for millions of years into the past, takes a very long time indeed. The equation of motion is the same one from Track A, generalized to an arbitrary potential:

dΒ²r/dtΒ² = βˆ’βˆ‡Ξ¦(r)

In the cylindrical coordinates (R, Ο†, z) natural to a disk galaxy, this splits into:

dΒ²R/dtΒ² βˆ’ R(dΟ†/dt)Β² = βˆ’βˆ‚Ξ¦/βˆ‚R RΒ² (dΟ†/dt) = L_z (conserved) dΒ²z/dtΒ² = βˆ’βˆ‚Ξ¦/βˆ‚z

where L_z is the star's angular momentum about the galactic rotation axis β€” conserved for any potential that doesn't depend on Ο† (i.e. axisymmetric potentials), which gives a powerful built-in check on any integrator's correctness.

Case Study

Antoja et al. (2018) discovered that the Milky Way's disk shows phase-space spirals β€” a rippling pattern in the distribution of stellar positions and vertical velocities, visible only because Gaia DR2 provided full 6D phase-space data for millions of stars simultaneously. The pattern is best explained as the disk still "ringing" from a relatively recent (few-hundred-million-year-old) close passage of the Sagittarius dwarf galaxy β€” a piece of the Milky Way's violent history, hiding in plain sight in the data, invisible to any survey that measured position without velocity.

B2B Pitch

Gaia-Scale Galactic Orbit Integration Service. Target: galactic dynamics research groups working with Gaia data releases. The pain: back-integrating millions to billions of stellar orbits through a realistic Milky Way potential, repeatedly, as models and data releases are updated, is a standing computational cost. The pitch: a physics-consistent fast propagator that turns a multi-day batch job into an interactive query.

Problem 32: The Milky Way rotation curve and dark matter

Plotting the orbital speed of stars and gas against their distance from the galactic center produces the rotation curve. Newtonian gravity sourced only by visible matter predicts the curve should fall off at large radii, the way planetary orbital speed falls off with distance from the Sun. It doesn't β€” the curve stays roughly flat far beyond where the visible disk ends. This single, simply-stated observational fact, repeated across hundreds of galaxies, is one of the strongest pieces of evidence for dark matter: an unseen mass distribution, roughly spherical, that dominates the outer galaxy's gravity. Fitting a rotation curve means solving for the dark matter halo's density profile that, added to the visible disk and bulge, reproduces the observed orbital speeds at every radius.

Problem 33: Stellar streams β€” fossils of galactic cannibalism

When a small dwarf galaxy or globular cluster passes too close to the Milky Way, tidal forces (the same physics as Problem 10's Roche limit) stretch it into a long, thin ribbon of stars called a stellar stream, which then continues to orbit roughly along its progenitor's original path. The GD-1 and Sagittarius streams are the best-studied examples. Because a stream's stars all share a nearly identical orbital history, their positions and velocities trace out the galactic potential along the stream's path with unusual precision β€” small kinks or gaps in a stream can even reveal the gravitational fingerprint of a dark matter subhalo passing through it, too faint to see directly in any other way.

Problem 34: Chemical tagging

Stars born together in the same cluster share a nearly identical chemical fingerprint β€” the same ratios of iron, magnesium, and other elements, inherited from the same star-forming gas cloud. Chemical tagging uses high-resolution spectroscopy to measure dozens of these abundance ratios per star, then clusters stars in this high-dimensional chemical space to reconstruct groups that have long since dispersed dynamically β€” reuniting a stellar family that orbital dynamics alone can no longer identify, because random gravitational encounters over billions of years have scrambled their positions and velocities but not their chemistry.

Problem 35: Astrometric binary star orbits

Some binary star systems are too close, or too far, to resolve as separate points β€” but Gaia can still detect the tiny back-and-forth wobble in a star's position on the sky as it orbits an unseen companion (the astrometric analogue of radial velocity's Doppler wobble). Fitting the observed wobble to a Keplerian orbit yields the orbital period, eccentricity, and β€” combined with the star's own mass estimate β€” the companion's mass, even when that companion is a faint star, a white dwarf, or in some remarkable cases a stellar-mass black hole with no light of its own at all.

Problem 36: Finding wide binaries and comoving pairs at scale

A wide binary is a pair of stars gravitationally bound but separated by thousands of AU β€” too far apart to show any orbital motion over a human observing lifetime, but identifiable because both stars share the same distance (matching parallax) and move through space together (matching proper motion and radial velocity). Finding these pairs among 1.8 billion Gaia sources is a nearest-neighbor search problem in a 5-dimensional space (sky position, parallax, and 2D proper motion), at a scale where naive pairwise comparison is computationally impossible β€” exactly the kind of problem that spatial indexing (k-d trees, or the same octree structure from Track A's Barnes-Hut algorithm) is built to solve.

Problem 37: Disk heating and the age-velocity-dispersion relation

Older stars in the Milky Way's disk have, on average, more scattered (higher dispersion) velocities than younger stars β€” a trend called disk heating. Repeated gentle gravitational encounters with giant molecular clouds, spiral arms, and past minor mergers nudge stellar orbits over billions of years, and the accumulated effect grows roughly with the square root of a star's age. Measuring this age-velocity-dispersion relation precisely β€” which requires independently knowing both a star's age (hard) and its full 3D velocity (which Gaia provides) β€” constrains how violent or how quiet the Milky Way's dynamical history has actually been.

Problem 38: Globular cluster dynamics

A globular cluster is a dense, roughly spherical ball of up to a million stars, old enough and dynamically relaxed enough that its structure can be modeled with a small number of parameters β€” the King profile being the standard empirical model for how stellar density falls off with distance from the cluster center. Over billions of years, gravitational interactions between cluster stars drive heavier stars toward the center and lighter stars outward (mass segregation), and can eventually lead to core collapse β€” a runaway contraction of the cluster's dense center. Modeling this evolution requires either direct N-body simulation of every star (the Track A machinery, again) or a statistical Fokker-Planck treatment of the cluster as a smooth distribution.

Problem 39: Moving groups in the solar neighborhood

Plotting nearby stars' velocities in the U (toward galactic center), V (direction of galactic rotation), and W (out of the disk plane) coordinate system reveals clumps β€” moving groups β€” of stars sharing similar velocities despite being scattered across the sky and not necessarily born together. Some moving groups are genuine dispersed star clusters; others are a resonant effect of the galaxy's rotating central bar or spiral arms, which can herd unrelated stars into similar orbits. Telling the two apart requires combining the kinematic clump with independent age and chemical information β€” kinematics alone is ambiguous.

Problem 40: Astrometric microlensing

Beyond the brightness-magnification microlensing of Problem 25, a close alignment between a foreground and background star also causes a tiny, characteristic shift in the background star's apparent position during the event β€” astrometric microlensing. Because this positional shift depends on the lensing object's mass in a clean, well-understood way, and does not require the lens itself to be luminous, it is one of the few techniques capable of directly weighing isolated stellar remnants β€” including, in a landmark 2022 detection, an isolated stellar-mass black hole with no companion star at all.

Problem 41: The galactic bar and spiral resonances

The Milky Way's central bar and spiral arms are not static β€” they rotate as coherent patterns (though not necessarily at the same rate as individual stars orbit), and stars whose orbital period matches the pattern's rotation rate at a particular radius experience a resonance β€” the same physics as the Kirkwood gaps in Track A, but for an entire galaxy's disk. The Sun sits close to one such resonance (the "corotation" or near it), which is part of why local solar-neighborhood kinematics show clumpy, non-random structure. Mapping these resonances requires solving for periodic orbits in a rotating, non-axisymmetric galactic potential β€” a genuinely harder problem than the axisymmetric case in Problem 31.

Problem 42: Local Group dynamics and dwarf galaxy orbits

The Milky Way, Andromeda, and dozens of smaller dwarf galaxies form the gravitationally bound Local Group. Reconstructing a dwarf galaxy's orbital history β€” where it has been, and when it will next pass close to the Milky Way β€” means integrating its motion backward and forward in a potential that includes not just the Milky Way but Andromeda and the other significant members too, a genuine few-body problem (Track A's foundational subject) at galactic scale, complicated further by dynamical friction: a drag-like force a massive dwarf galaxy exerts on itself by gravitationally stirring up the Milky Way's dark matter halo as it passes through.

Problem 43: Classifying stellar populations from Gaia photometry

Gaia's color-magnitude diagram β€” brightness plotted against color for millions of stars β€” separates naturally into distinct populations: the main sequence, giant branches, white dwarfs. But separating, say, thin-disk from thick-disk stars, or identifying members of a specific stream or cluster, from photometry and astrometry alone is a classification problem well suited to unsupervised or semi-supervised machine learning, clustering stars in a combined space of color, magnitude, kinematics, and (where available) chemistry β€” at a scale where no human could inspect each star individually.

Problem 44: Calibrating the cosmic distance ladder

Gaia's parallax distances are the most precise ever measured for nearby stars β€” but parallax stops being measurable beyond a few thousand parsecs, the angle becomes too small to resolve. Beyond that, astronomers rely on standard candles: objects whose true brightness can be inferred from an observable property, like a Cepheid variable star's brightness-vs-pulsation-period relation. Gaia's precise parallax distances to nearby Cepheids are what calibrate that relation in the first place, which is then used to measure distances to far more distant galaxies β€” making Gaia, somewhat surprisingly, a foundational input to measuring the expansion rate of the universe itself.

Problem 45: Real-time cross-matching at billion-object scale

Gaia's 1.8 billion sources need to be matched against other catalogs β€” SDSS, 2MASS, WISE, ground-based surveys β€” to combine astrometry with spectra, infrared photometry, and other measurements of the same physical stars. At this scale, naive matching (comparing every source in one catalog against every source in another) is computationally infeasible; production cross-matching uses spatial indexing (HEALPix pixelization of the sky is the standard scheme) to reduce the search to nearby candidates only, plus a probabilistic matching step that accounts for each catalog's own positional uncertainty β€” a deceptively hard infrastructure problem underlying nearly every multi-survey astronomical study published today.

Track D Β· 15 Problems

Spectroscopy, Galaxies, Quasars & Cosmology

The Sloan Digital Sky Survey photographed one-third of the sky and took spectra of 3 million objects. This is what the universe looks like in data.

Module D.0 β€” What is a spectrum?

White light passed through a prism becomes a rainbow. Each color is a different wavelength of light, and atoms absorb and emit light at very specific wavelengths β€” spectral lines β€” determined by the discrete energy levels their electrons can occupy. Every element has a unique fingerprint of lines, which is how we know what a star or galaxy light-years away is made of without ever touching it.

The Balmer series gives hydrogen's visible-light emission lines:

1/Ξ» = R_H (1/n₁² βˆ’ 1/nβ‚‚Β²)

where R_H = 1.097 Γ— 10⁷ m⁻¹ (the Rydberg constant), n₁ = 2 is the lower energy level, and nβ‚‚ = 3, 4, 5, … are the upper levels an electron falls from. For nβ‚‚ = 3: Ξ» = 656.3 nm β€” red light, the HΞ± line, the single most recognizable spectral feature in astronomy, marking active star formation across entire galaxies. For nβ‚‚ = 4: Ξ» = 486.1 nm β€” blue-green, the HΞ² line.

When a galaxy moves away from us, every one of its spectral lines shifts to longer wavelengths β€” redshift:

z = (Ξ»_observed βˆ’ Ξ»_rest) / Ξ»_rest = v/c (for small velocities)

For distant galaxies, Hubble's law connects that recession velocity directly to distance: v = Hβ‚€ Γ— d, where Hβ‚€ is the Hubble constant, currently measured at roughly 67–73 km/s/Mpc depending on the method used β€” a persistent, unresolved discrepancy between different measurement techniques known as the Hubble tension, and one of the most consequential open problems in cosmology today.

Problem 46: A fast spectroscopic redshift estimator

Measuring a galaxy's redshift precisely requires a full spectrum β€” light dispersed across many wavelengths, with individual spectral lines identified and their shift measured. This is spectroscopic redshift, and it is expensive: it demands dedicated telescope time, one galaxy (or a limited number via multi-object spectroscopy) at a time. Photometric redshift ("photo-z") instead estimates redshift from just a handful of brightness measurements in different color filters β€” far faster and cheaper to obtain for millions of galaxies at once, at the cost of much larger uncertainty, since it is inferring a continuous spectral shift from a coarse, discretized color signal rather than measuring it directly.

The classical photo-z approach fits a library of template spectra β€” models of what different galaxy types look like at different redshifts β€” against the observed colors, searching for the best match. A neural approach instead learns the mapping from observed colors (and other features) directly to redshift from a training set of galaxies that have both photometry and a known spectroscopic redshift, implicitly learning the same physics the templates encode, but evaluated as a single fast forward pass rather than a template search across a large grid β€” a natural target for a fast, physics-consistent surrogate model, the same pattern as Problem 16's transit fitting and Problem 31's orbit integration.

Problem 47: Galaxy morphological classification

Galaxies broadly fall along the Hubble tuning fork: ellipticals (smooth, featureless light profiles), spirals (a disk with arms, often a central bar), and irregulars (everything that doesn't fit cleanly, often disturbed by a recent merger). Morphology correlates strongly with a galaxy's star-formation history and environment, which is why classifying it at scale matters. The Galaxy Zoo project first crowdsourced millions of visual classifications from volunteers; convolutional neural networks trained on those labels now reproduce human-level classification automatically across entire survey catalogs, freeing the crowdsourcing effort for the genuinely ambiguous, unusual cases a model flags with low confidence.

Problem 48: Quasar identification and black hole mass estimation

A quasar is the extremely luminous core of a galaxy, powered by gas spiraling into a supermassive black hole and heating to enormous temperatures before it crosses the event horizon. Quasar spectra show very broad emission lines β€” broadened by the high orbital velocities of gas deep in the black hole's gravitational well β€” and the width of those lines, combined with the source's luminosity, feeds a virial mass estimator for the black hole itself:

M_BH ∝ R_BLR Γ— Ξ”vΒ² / G

where R_BLR (the radius of the broad-line emitting region) is inferred from the source's luminosity via an empirical relation, and Ξ”v is the emission line width. This is how astronomers weigh black holes billions of light-years away, using nothing but a spectrum.

Problem 49: Type Ia supernova light-curve fitting

A Type Ia supernova β€” a white dwarf's thermonuclear explosion after it accretes too much mass from a companion star β€” reaches a strikingly consistent peak brightness across the whole population, after correcting for a well-characterized relationship between how bright it gets and how fast it fades. This makes them standard candles, usable to measure cosmological distances far beyond where Cepheids (Problem 44) reach. The correction is fit with a light-curve model such as SALT2, which parametrizes the supernova's brightness, color, and light-curve stretch (fast-fading vs. slow-fading) from photometric observations over the weeks surrounding peak brightness β€” this exact technique, applied to distant supernovae in the late 1990s, produced the surprise discovery that the universe's expansion is accelerating, evidence for dark energy, a Nobel-Prize-winning result.

B2B Pitch

Type Ia Supernova Light-Curve Fitting Accelerator. Target: cosmology survey teams (Rubin Observatory / LSST-scale programs) expecting to discover thousands of new supernovae per night. The pain: SALT2-style fitting at that volume, with rigorous uncertainty quantification, is a genuine computational bottleneck. The pitch: a fast, validated surrogate that keeps pace with the discovery rate instead of falling behind it.

Problem 50: Baryon Acoustic Oscillations as a standard ruler

In the early universe, sound waves rippled through the hot plasma of ordinary matter and radiation before recombination "froze" them in place β€” leaving a faint but statistically measurable preferred separation distance between galaxies today, roughly 490 million light-years. Because this scale is set by well-understood early-universe physics, it functions as a standard ruler: measuring the apparent size of this preferred separation at different redshifts (different cosmic distances) directly probes the universe's expansion history, complementary to the standard-candle approach of supernovae. Detecting it requires measuring the galaxy correlation function (Problem 57) with enough precision across a large enough survey volume to see a subtle statistical bump against a much larger smooth background.

Problem 51: Galaxy spectral energy distribution fitting

A galaxy's full spectral energy distribution (SED) β€” its brightness across ultraviolet, optical, and infrared wavelengths β€” encodes its star-formation history, stellar mass, dust content, and age, all superimposed. Stellar population synthesis models build a galaxy's predicted SED as a weighted combination of simple stellar populations of different ages and metallicities, plus a dust attenuation model; fitting this combination against an observed SED (again typically via MCMC, echoing Problem 16) recovers the underlying physical parameters β€” essentially reading a galaxy's life story out of its combined starlight.

Problem 52: AGN variability and reverberation mapping

An active galactic nucleus (AGN) flickers in brightness over days to years, as the accretion disk around its central black hole varies. Light from that variable disk reaches surrounding gas clouds (the broad-line region) after a light-travel-time delay, so the gas's emission-line brightness echoes the disk's continuum variability with a measurable lag. Reverberation mapping cross-correlates the continuum and emission-line light curves to measure that lag directly, which β€” combined with the emission-line width from Problem 48 β€” gives a geometric, physics-based black hole mass measurement, rather than relying on an empirical scaling relation.

Problem 53: Strong gravitational lensing mass reconstruction

A sufficiently massive foreground galaxy or cluster bends light from a background source enough to produce multiple distorted images, arcs, or even a complete "Einstein ring." Because the exact distortion pattern depends precisely on the foreground mass distribution β€” including its invisible dark matter β€” strong lensing mass reconstruction works backward from the observed image distortions to solve for the mass distribution that could have produced them: an inverse problem, mathematically related to medical imaging reconstruction, and one of the few direct probes of dark matter's spatial distribution independent of any assumption about how it interacts with light.

Problem 54: Weak lensing and cosmic shear

Far from any single massive lens, the combined gravity of all the matter along a line of sight subtly distorts the shapes of countless background galaxies by a percent-level effect too small to see in any single galaxy, but statistically measurable by averaging the coherent shape distortion β€” cosmic shear β€” across millions of galaxies. This maps the large-scale distribution of matter, including dark matter, across enormous volumes of the universe, and is one of the primary science goals of the Rubin Observatory's decade-long LSST survey β€” a problem defined entirely by its statistics, since the signal in any individual measurement is smaller than the noise.

Problem 55: Galaxy cluster mass estimation

A galaxy cluster's mass can be estimated three independent ways: from the velocity dispersion of its member galaxies (via the virial theorem), from the temperature and shape of the hot X-ray-emitting gas trapped in its gravitational well, or from its weak lensing signature (Problem 54) on background galaxies. The three methods rely on different physics and different systematic errors, so combining them β€” and understanding when and why they disagree β€” is a standard, high-value cross-validation problem in observational cosmology, often revealing that a cluster is dynamically disturbed (mid-merger) rather than the relaxed, simple system the simplest models assume.

Problem 56: Fitting the Cosmic Microwave Background power spectrum

The Cosmic Microwave Background (CMB) β€” the afterglow of the early universe, 380,000 years after the Big Bang β€” has a nearly uniform temperature across the sky, with tiny fluctuations (one part in 100,000) that encode the initial density variations from which every galaxy, star, and planet eventually formed. Decomposing the pattern of those fluctuations into a power spectrum β€” the amplitude of temperature variation as a function of angular scale β€” and fitting it against cosmological models pins down the universe's fundamental parameters (its age, geometry, composition of ordinary matter, dark matter, and dark energy) to remarkable precision, essentially reading the universe's birth certificate out of a faint microwave hiss.

Problem 57: The galaxy correlation function and large-scale structure

Galaxies are not scattered randomly through space β€” they cluster into a vast cosmic web of filaments, sheets, and voids. The two-point correlation function ΞΎ(r) quantifies this clustering: the excess probability, over a random distribution, of finding a galaxy pair separated by distance r. Computing it for a modern survey containing tens of millions of galaxies means counting pairwise separations across an enormous dataset β€” again an O(NΒ²) problem in its naive form, and again solvable with the spatial-indexing techniques introduced in Track A and Track C β€” and comparing the resulting clustering pattern against cosmological simulations to constrain how structure grows over cosmic time.

Problem 58: Automated emission-line deblending

Real galaxy spectra often show several emission lines crowded close together in wavelength β€” [NII] flanking HΞ± is the textbook case β€” overlapping enough that a naive single-line fit systematically misattributes flux between them. A robust automated pipeline must simultaneously fit multiple overlapping line profiles (typically Gaussian or Voigt profiles) with shared physical constraints, such as tying two lines from the same ion to a fixed flux ratio predicted by atomic physics β€” turning what looks like a simple curve-fitting problem into a constrained, physically-informed optimization run millions of times across a large spectroscopic survey.

Problem 59: Neural photometric redshift estimation

Extending Problem 46's theme: modern photo-z pipelines increasingly replace template-fitting entirely with neural networks trained end-to-end on the largest available spectroscopic samples, often outperforming template methods particularly in the regime where template libraries are incomplete or don't represent the true diversity of real galaxy spectra β€” but they inherit whatever biases and gaps exist in their training data, in a domain where the training data (spectroscopically confirmed galaxies) is itself not a representative, unbiased sample of the full underlying galaxy population β€” a genuine, unresolved methodological tension in the field.

Problem 60: Automated spectral classification at survey scale

SDSS alone has collected spectra for roughly 3 million objects, and needs to classify each one β€” star, galaxy, or quasar; and within each category, more detailed subtypes (stellar spectral type, galaxy type, quasar redshift and properties) β€” automatically, since no team could review that many spectra by hand. This is the endpoint of everything in this track: a production pipeline that ingests a raw spectrum and outputs a validated physical classification and measurement set, fast enough and reliable enough to run unattended, night after night, as the true final stage of a spectroscopic survey.

Track E Β· 15 Problems

How the Physics Becomes Code

Classical neural networks don't conserve energy. Physics-informed ones do. Here is why that matters.

Module E.0 β€” Why standard ML fails for physics

A standard neural network is trained to minimize a loss function on data. It has no knowledge of the underlying physics β€” it will fit the training data well, but it has no mechanism forcing it to respect conservation of energy, conservation of momentum, or any other fundamental law when making predictions outside that training data.

For orbital mechanics, this is not a minor inaccuracy β€” it is catastrophic. A network predicting "where will this satellite be in ten years" that quietly violates energy conservation will produce a slowly drifting answer: the simulated orbit spirals inward or outward over time, exactly the failure mode Track A's symplectic integrators exist to prevent. Real satellites do not do this. A model that predicts they do is not a slightly-wrong model β€” it is wrong in a way that compounds without bound.

Physics-informed neural networks (PINNs) fix this by adding physics as an explicit constraint in training, not just a hope baked into the training data. The loss function includes both the ordinary data-fitting term and a penalty for violating the governing physics equations:

L_total = L_data + Ξ» Γ— L_physics

where L_physics might be the residual of Hamilton's equations (Track A, Problem 1), or of any other governing differential equation for the system, and Ξ» is a weighting term balancing fidelity to data against fidelity to physics. This single idea β€” training on the physics, not just the data β€” is the throughline connecting every idea in this track.

Problem 61: A general-purpose PINN module

Building a physics-informed neural network for a specific problem means writing down the governing partial differential equation (PDE) β€” an equation involving a function and its derivatives with respect to space and/or time β€” and enforcing it during training. The key technical trick is automatic differentiation: the same backpropagation machinery used to compute gradients of a loss function with respect to a network's weights can also compute derivatives of the network's output with respect to its inputs β€” position, time, whatever the PDE is written in terms of. This is fundamentally different from ordinary training, where you only ever need gradients with respect to weights; here, you differentiate the network itself, symbolically and exactly (not via finite-difference approximation), to construct the PDE residual directly from the network's own predictions.

residual(x, t) = N[u(x, t; ΞΈ)] βˆ’ f(x, t)

where u(x, t; ΞΈ) is the network's output (parametrized by weights ΞΈ), N[Β·] is the differential operator from the governing PDE (built from automatic derivatives of u with respect to x and t), and f(x, t) is any known source term. Training pushes this residual toward zero everywhere in the domain, not just at the data points β€” so the network is forced to behave like a genuine solution of the underlying physics, not merely an interpolator between observations.

Visual Output

Two side-by-side heatmaps of a simulated field (temperature, pressure, or gravitational potential) over space and time: one from the network's prediction, one from a classical finite-difference solver used as ground truth. A difference plot beneath shows the residual error is small and β€” critically β€” does not grow at the domain's edges or at long time horizons, the tell-tale sign that the physics constraint is doing its job.

Problem 62: Neural ODEs for continuous-time dynamics

A standard recurrent neural network models a sequence in discrete steps. A Neural ODE (Chen et al., 2018) instead defines the hidden state's evolution as a continuous-time differential equation, dh/dt = f(h, t; ΞΈ), with f itself a neural network β€” and computes the state at any time by handing that equation to an off-the-shelf ODE solver. This is a natural fit for irregularly-sampled astronomical time series (observations don't arrive on a neat fixed grid), and it connects directly back to Track A: a Hamiltonian Neural Network is, structurally, a Neural ODE whose dynamics function is constrained to come from a learned Hamiltonian.

Problem 63: Symmetry-preserving and equivariant networks

Physical laws don't care which direction you're facing β€” rotate an entire N-body system, and the physics evolves identically, just rotated. A standard neural network has no reason to respect this rotational symmetry unless it happens to learn it from data, approximately, at the cost of needing far more training examples to cover every orientation. An equivariant network architecture builds the symmetry into its structure directly β€” guaranteeing that rotating the input rotates the output correspondingly, exactly, for every possible rotation, not just the ones seen during training β€” which typically means dramatically better data efficiency for physical systems with known symmetries (translation, rotation, and sometimes permutation symmetry between identical particles).

Problem 64: Uncertainty quantification in physics-informed ML

A single point prediction from a neural network says nothing about how confident that prediction is β€” and for a mission-critical application like conjunction screening (Track A, Problem 6) or trajectory correction (Problem 15), knowing the uncertainty is often as important as the prediction itself. Bayesian neural networks place a probability distribution over the network's weights rather than a single fixed value, and deep ensembles β€” training several independently-initialized networks and looking at how much they disagree β€” offer a simpler, often equally effective, alternative. Both approaches turn a black-box point estimate into something closer to the covariance-aware propagation already standard in classical orbit determination (Track A, Problem 13).

Problem 65: Surrogate modeling for expensive simulations

When a single high-fidelity simulation run β€” a full N-body integration, a detailed radiative transfer calculation β€” takes hours, and you need to explore how the outcome changes across thousands of different input parameters, running the full simulation for every parameter combination is infeasible. A surrogate model β€” trained on a limited number of full simulation runs β€” learns to approximate the simulation's input-output mapping directly, and can then be queried in milliseconds. Gaussian processes are the classical choice, prized for built-in, principled uncertainty estimates; neural network surrogates scale better to high-dimensional inputs and larger training sets, but need more care to get well-calibrated uncertainty out of them.

Problem 66: Differentiable physics simulators

If an entire simulation β€” not just a neural network β€” is written so that every step is differentiable, automatic differentiation can compute the gradient of the simulation's final output with respect to its initial conditions or parameters directly, end to end. This makes previously hard inverse problems tractable: given an observed outcome, what initial conditions or physical parameters would have produced it? Instead of a slow trial-and-error search, gradient descent directly through the simulator finds the answer β€” an approach increasingly used to infer, for instance, a galaxy cluster's mass distribution (Track D, Problem 53) directly from lensing observations, by differentiating through the full forward lensing model.

Problem 67: Transfer learning across physics domains

Labeled real astronomical data is often scarce β€” confirmed exoplanets, well-characterized supernovae, spectroscopically confirmed redshifts all take real telescope time and real human validation to produce. A common workaround: pretrain a model on abundant, cheaply-generated synthetic data from a physics simulator, where the ground truth is known exactly by construction, then fine-tune on the smaller set of real, labeled examples. The synthetic data needs to be realistic enough β€” noise, instrumental effects, and all β€” that what the model learns actually transfers, which is itself an active area of methodological research, not a solved problem.

Problem 68: Physics-constrained regression on sparse, noisy data

Real observational data is never complete or clean β€” missing time points, measurement noise, instrumental gaps. Data assimilation techniques, long used in weather forecasting, combine a physics model's predictions with sparse, noisy observations in a principled way, weighting each according to its known uncertainty, to produce a best estimate of the true underlying state that is more accurate than either the physics model or the raw data alone. Applying the same philosophy to sparse astronomical time series β€” a stellar light curve with large gaps, for instance β€” means the physics model fills in what the data can't see, rather than a naive interpolation with no physical basis at all.

Problem 69: Interpreting what a physics-informed network has learned

Even a network trained to respect known physical laws is still, internally, a large set of learned weights β€” and understanding why it makes a particular prediction, beyond "the physics-constrained loss was low," matters for scientific trust and for catching failure modes before they cause a bad prediction downstream. Techniques from the broader interpretability field β€” probing which input features most influence a given output, or checking whether a Hamiltonian Neural Network's learned H function actually resembles the true physical energy function it was implicitly meant to discover β€” turn a physics-informed model from a black box that happens to conserve energy into a genuine, inspectable scientific tool.

Problem 70: Multi-fidelity modeling

Not every simulation needs to be run at maximum precision. A multi-fidelity approach deliberately combines many cheap, lower-accuracy simulation runs with a small number of expensive, high-accuracy ones, learning a correction term that maps the cheap model's systematic errors onto the expensive model's more trustworthy results. This gets most of the accuracy of the expensive simulator at a small fraction of its total compute cost β€” directly useful anywhere a low-fidelity approximate model already exists and is fast (the patched-conic approximation of Track A, Problem 12, is a natural low-fidelity partner for a full numerical integrator as the high-fidelity source).

Problem 71: Conservation-law-constrained architectures beyond Hamiltonians

Hamiltonian Neural Networks (Track A, Problem 1) aren't the only way to bake a conservation law into a network's structure. Lagrangian Neural Networks instead learn the system's Lagrangian L = T βˆ’ V (kinetic minus potential energy, rather than the Hamiltonian's sum) and derive dynamics via the Euler-Lagrange equation β€” an approach that generalizes more naturally to systems described in arbitrary coordinates, including ones with physical constraints (a pendulum's fixed arm length, for instance) that are awkward to express directly in the Hamiltonian formulation. Choosing between the two is a genuine engineering decision, not a matter of one being strictly better than the other.

Problem 72: Real-time inference on edge hardware

Everything in this track assumes a network can be evaluated fast enough to matter β€” but "fast enough" looks very different on a research workstation than on the radiation-hardened flight computer of Track A, Problem 15's autonomous spacecraft, or the modest edge hardware of a remote observatory (Track F+G's Idea #12). Quantization (reducing numerical precision from 32-bit to 8-bit or lower) and distillation (training a much smaller "student" network to reproduce a larger "teacher" network's outputs) are the two standard techniques for shrinking a physics-informed model down to something that fits real deployment constraints β€” without giving up the conservation guarantees that made it worth building in the first place.

Problem 73: Benchmarking physics-ML against classical solvers

A claim of "10Γ— faster" or "100Γ— faster" is only as credible as the benchmark behind it. Rigorous benchmarking means comparing against a properly tuned classical solver (not a deliberately weak baseline), on the same hardware, across a representative range of problem sizes and difficulty, with accuracy reported alongside speed β€” since a faster but less accurate method is a different product than a faster and equally accurate one, and a credible technical buyer will ask exactly this question before paying for either.

B2B Pitch

Benchmark Report: Physics-ML vs. GMAT/STK/REBOUND. Target: any team currently deciding whether to trust a physics-ML approach at all. The pain: marketing claims about ML speedups are cheap; credible, reproducible, apples-to-apples benchmarks against the industry-standard tools (NASA's GMAT, AGI's STK, the open-source REBOUND N-body code) are not. The pitch: an independently reproducible benchmark suite becomes the trust artifact that unlocks every other sale in this document.

Problem 74: Active learning for simulation-heavy domains

When each new labeled training example costs a full expensive simulation run, which parameter combinations should you simulate next to improve your surrogate model the most? Active learning answers this by having the current surrogate model itself flag the regions of parameter space where its own uncertainty is highest (echoing Problem 64's uncertainty quantification), and prioritizing new simulation runs there β€” squeezing far more model accuracy out of a fixed simulation budget than randomly or exhaustively sampling the parameter space ever could.

Problem 75: From research prototype to production API

A validated, benchmarked, physics-conserving model sitting in a Jupyter notebook is not yet a product. The last mile β€” wrapping it in a stable API, handling edge cases and malformed inputs gracefully, versioning the model so results are reproducible months later, monitoring for silent performance drift once it's deployed β€” is unglamorous engineering work, and it is also where the great majority of research code dies before ever reaching a real user. Every idea in Tracks A through E ends here, at this problem, whether or not it says so explicitly.

Tracks F & G Β· 15 Candidate Products

From Research to Product

Here is what gets built from all of the above. Here is who buys it and why. The 75 problems in Tracks A–E are the research; these 15 are what a client actually signs a contract for β€” most of them bundle several research ideas into one deployable thing.

1. solve_ivp-Compatible Orbit-Propagation Shim M

Built on the Hamiltonian Neural Network approach from Track A, Problem 1 (Module A.1, above, has the full pitch and visual). A drop-in replacement for scipy.solve_ivp that any existing Python astrodynamics codebase can adopt by changing one import β€” faster, and energy-consistent over long horizons where the classical integrator drifts. Target client: Python-based astrodynamics research teams who cannot justify rewriting a working codebase just to get a faster solver.

2. Satellite Conjunction Screening Service L

Built on the collision-probability math of Track A, Problem 6, powered by the fast propagator of Idea #1 and #13. Technical implementation: a service that ingests two-line element sets (or higher-fidelity ephemerides) for tracked objects, screens every close-approach pair against a probability-of-collision threshold, and flags conjunctions requiring operator action β€” running continuously as new tracking data arrives, at a scale (tens of thousands of tracked objects, growing) where the pairwise math alone is a real computational load.

B2B Pitch

Target client: satellite operators and constellation operators who currently rely on manually-triggered conjunction reports from third-party providers. The pain: as constellations grow into the thousands of satellites, conjunction volume grows roughly quadratically, and screening latency becomes a real operational risk. The pitch: continuous, low-latency screening at a cost that scales with the fast propagator underneath it, not with the O(NΒ²) cost of the naive approach.

What you'd see on screen: a live dashboard of upcoming close approaches ranked by probability of collision, each with a mini trajectory plot of both objects' relative motion through the encounter. Defensible because: the propagator's speed and accuracy is the entire product β€” a client who has integrated it into their operational workflow has genuine switching costs.

3. Dockerized DeepInsight: Orbital Specimen S

A single, self-contained, containerized demo build of the DeepInsight product line (see also techotsu.com/products/deepinsight) focused narrowly on one worked orbital-dynamics example β€” the CR3BP trajectory comparison from Module A.1's visual output box, packaged as a docker run-able artifact a prospective client's engineering team can pull down and inspect on their own infrastructure in an afternoon, with no sales call required first.

B2B Pitch

Target client: technical evaluators at a research group or space agency who want to kick the tires before any conversation with sales. The pain: most vendor demos are either a slide deck or a hosted web demo neither of which lets a skeptical engineer verify the claims on their own trusted hardware, with their own data. The pitch: run it yourself, on your own machine, with your own trajectory β€” no cloud dependency, no trust required.

Minimum viable version: one containerized Jupyter notebook reproducing the A.1 energy-conservation comparison end to end. Defensible because: it isn't really the product β€” it's the trust-building step that gets a technical evaluator to champion a full engagement internally.

4. Benchmark Report: Physics-ML vs. GMAT/STK/REBOUND M

Covered in full in Track E, Problem 73, above β€” the credibility artifact underneath every other pitch in this section. Technical implementation: a reproducible benchmark harness comparing accuracy and wall-clock time against NASA's GMAT, AGI's STK, and the open-source REBOUND N-body code, across a representative spread of problem sizes, published with enough methodological detail that a skeptical reader can rerun it themselves.

5. Fast Transit Light-Curve Fitting Add-On M

Built on the Mandel-Agol transit model and MCMC fitting bottleneck described in Track B, Problem 16 (Module B.1, above, has the full pitch and case study). Target client: university astronomy departments and survey teams with a backlog of unanalyzed Kepler/TESS/Rubin light curves, for whom MCMC fitting time is a direct constraint on publication rate.

6. Gaia-Scale Galactic Orbit Integration Service L

Built on the galactic potential orbit integrator of Track C, Problem 31 (Module C.1, above, has the full pitch and case study). Target client: galactic dynamics research groups working with successive Gaia data releases, who re-run billions of orbit integrations every time the underlying potential model or the data itself is updated.

7. Unified Multi-Catalog Automation Dashboard L

Built on the cross-matching infrastructure problem of Track C, Problem 45. Technical implementation: a dashboard that automates the ingestion, cross-matching, and quality-flagging of new data releases across Gaia, SDSS, TESS, and other major public catalogs, so a research group's pipeline updates itself when a new data release lands rather than requiring a graduate student to manually re-run everything.

B2B Pitch

Target client: research groups whose published results depend on staying current with multiple, independently-updated public catalogs. The pain: cross-matching and re-validating a pipeline against a new data release is repetitive, error-prone manual work that happens every time any one of several upstream surveys ships an update. The pitch: automate the boring, error-prone part; keep the science part.

8. Multi-Planet Stability Screening API M

Built on the dynamical stability screening of Track B, Problem 27, using the fast N-body propagation of Track A. Technical implementation: an API that takes a proposed multi-planet system's orbital parameters (with uncertainties) and returns a stability verdict β€” computed by integrating across the parameter uncertainty range and checking for close encounters or ejections over the star's estimated age β€” turning a task that used to require setting up a custom N-body simulation into a single API call.

What you'd see on screen: a stability heatmap across two of the system's uncertain parameters (say, two planets' eccentricities), green where the system remains stable and red where it doesn't β€” letting a researcher immediately see which part of their fitted parameter space is physically plausible.

9. MATLAB/Simulink Adapter for the Orbital Toolkit S

The same drop-in philosophy as Idea #1, for the mission-design teams who work in MATLAB/Simulink rather than Python β€” a substantial share of the aerospace industry. Technical implementation: a MATLAB-callable interface (and a Simulink block, for teams that build trajectory logic visually) wrapping the same underlying fast propagator, so a MATLAB-native team gets the same speedup without learning a new language or rebuilding an existing Simulink model from scratch.

B2B Pitch

Target client: aerospace mission-design teams standardized on MATLAB/Simulink, a large share of the industry outside pure research groups. The pain: a Python-only tool, however good, is a non-starter for a team that has no reason to introduce a second language into an established toolchain. The pitch: the exact same physics-conserving speedup, inside the tool you already use every day.

10. Type Ia Supernova Light-Curve Fitting Accelerator M

Covered in full in Track D, Problem 49, above. Target client: cosmology survey teams expecting thousands of new supernova discoveries per night from Rubin Observatory-scale programs, for whom SALT2-style fitting at that volume, with proper uncertainty quantification, is a genuine bottleneck rather than a convenience.

11. Wide Binary / Co-Moving Pair Finder at Scale M

Built on the spatial-indexing search problem of Track C, Problem 36. Technical implementation: a service that runs the 5-dimensional nearest-neighbor search (sky position, parallax, proper motion) across the full 1.8-billion-source Gaia catalog and returns validated wide-binary and co-moving-pair candidates, ranked by the statistical confidence of the match β€” a search a single research group's own infrastructure often cannot run at full catalog scale without dedicated engineering effort.

Minimum viable version: a pre-computed, downloadable candidate catalog for a single well-studied sky region, with an API for on-demand queries against the full sky added once there's a paying client for it. Defensible because: the underlying spatial index, once built and validated, is expensive to reproduce independently β€” a genuine moat, not just a convenience.

12. Edge-Deployment Package for Remote Observatories M

Built on the quantization and distillation techniques of Track E, Problem 72. Technical implementation: a shrunk, edge-hardware-compatible build of the transit-vetting and light-curve pipelines (Track B, Problems 26 and 30), packaged to run on modest on-site computers at remote or bandwidth-limited observatories, so candidate signals can be flagged locally in near-real-time rather than waiting on a slow or intermittent connection to a central data center.

B2B Pitch

Target client: small and mid-size observatories, especially in locations with limited or expensive internet bandwidth. The pain: shipping raw data to a central facility for processing introduces both latency and real infrastructure cost. The pitch: the same validated pipeline, small enough to run on the hardware already sitting next to the telescope.

13. Orbit-Propagation Speed Layer (Component Licensing) S

Rather than a full product, this is the fast propagator from Ideas #1, #2, and #6 licensed as a standalone component β€” a speed layer another company can embed inside their own existing product, rather than a tool their end users interact with directly. Technical implementation: a well-documented, versioned library with a stable interface and a formal accuracy/performance specification, suitable for another vendor's due-diligence process before they build it into their own shipped product.

B2B Pitch

Target client: other software vendors in the space-tech or astrodynamics tooling space who need faster propagation inside their own product but don't want to build and maintain the physics-ML expertise in-house. The pain: building this in-house is a multi-quarter research investment with real technical risk. The pitch: license the component that took us months to validate, in weeks instead.

14. Exoplanet Archive Nightly Watch + Auto-Analysis Service L

Built on Track B's transit vetting (Problem 30), false-positive screening (Problem 19), and occurrence-rate correction (Problem 29). Technical implementation: a service that watches public exoplanet archives (NASA Exoplanet Archive, TESS releases) for new candidate signals every night, automatically runs vetting, transit fitting, and stability screening on anything new, and delivers a ranked, pre-analyzed shortlist each morning β€” turning "check the archive and analyze anything interesting" from a standing manual chore into something a research group wakes up to already done.

What you'd see on screen: a daily digest email or dashboard: new candidates overnight, each with a vetting verdict, a fitted transit model, and a stability screen already run β€” ready for a human to make the final call rather than to do the first-pass analysis from scratch.

15. Solver-Benchmarking-as-a-Service for Mission Design Teams M

The productized, ongoing version of Idea #4 β€” instead of a one-time benchmark report, a subscription service that continuously re-runs the benchmark suite as both the physics-ML solvers and the classical baselines (GMAT, STK, REBOUND's own version updates) evolve, so a mission-design team always has a current, trustworthy answer to "which solver should we actually use for this class of problem, today."

B2B Pitch

Target client: mission-design teams who must justify solver choices to their own stakeholders and cannot afford to have that justification go stale as tools update. The pain: a one-time benchmark ages out of relevance the moment either side of the comparison ships a new version. The pitch: a living benchmark, not a PDF that was accurate on the day it was written.

What Now

If you have read this far, you understand the problems. You understand the physics. You understand where the software products come from.

The Theoretical Astrophysicist role at TechOtsu Research is not about writing papers. It is about taking one of these problems, going deep, and helping the BD team understand whether a real customer will pay for the solution.

If you want to apply: techotsu.com/atwork

If you want to read more: