AI · BIOLOGY · SCIENTIFIC DISCOVERY

Signals into
Discovery

I’m Kuan-Ting (Woody) Lin. I find biological signals that existing tools and assumptions overlook—and turn unexpected observations into scientific understanding and opportunities for better therapies.

Existing tools shape what we think we can see.

My work begins where
that view stops.

Across my career, I have questioned apparent limits, investigated signals that were easy to dismiss, and built the methods needed to test them. The goal is to turn newly visible biology into something we can understand—and use.

THE DISCOVERY STORIES

Three examples of turning
data into discovery.

Selected examples from a broader body of work: discovering from limited material, developing a missing measurement, and revisiting accepted results. Each produced new biological insight and a question worth pursuing.

01 / Oncogene · 2014

New RNA discoveries from one patient pair.

Discovery from one patient: one tumour and one matched normal sample.

AN ASSUMPTION PUT TO THE TEST

The assumption I challenged

One patient pair seemed too small for discovery. Gene-expression arrays defined the established view of RNA, while early Solexa sequencing was too expensive to apply to a large patient cohort and lacked a mature analytical workflow.

What we showed

Sequencing beyond predefined probes revealed DUNQU1, an uncharacterized transcript with protein-coding potential that was preferentially expressed in HCC (liver-cancer) tumours compared with adjacent normal liver tissue. Different proportions of FGFR2 RNA versions were linked to tumour size and whether the cancer returned, as well as hepatitis infection and liver scarring.

My contribution

I worked out the early Solexa analysis and developed the qPCR approach to measure the relative proportions of FGFR2 RNA variants.

The evidence we established

The team validated preferential DUNQU1 expression in liver tumours across 55 tumour–normal pairs. Functional experiments showed that increasing DUNQU1 expression enhanced colony formation in liver-cancer cells, adding biological evidence to the computational discovery.

Why it matters

The study uncovered new liver-cancer markers with potential to improve prediction of disease outcomes and guide the search for new treatment targets. It turned previously unmeasured RNA signals into concrete candidates for biomarker and therapeutic research.

Discovery can begin with one pair. Validation must go further.

2014 · CHANGE THE MEASUREMENT

Unprecedented transcriptome resolution.

Look for sequences already on the array

A microarray detects RNA through predefined probes. What is outside those measurements can remain unseen.

Known RNAMatched by a probeUncharacterized RNAOutside this probe’s viewAnother RNA versionOutside this probe’s view

A useful measurement of known targets can still leave important biology outside the view.

Conceptual illustration, not experimental measurements. This view illustrates spliced RNA; RNA-seq can also detect unspliced or intron-retaining RNA. Microarray coverage depends on probe design; specialized arrays can examine splicing. RNA detection alone does not establish a functional protein, validated biomarker, or therapeutic target.Explore the published evidence · Oncogene 2014 ↗
Explore the evidence & discovery process
01A measurement beyond the array
Early Solexa generated approximately 250 million paired-end reads from the tumour and its adjacent normal tissue. The search included RNA sequences absent from existing microarray probe coverage.
02Turning sequencing candidates into cohort measurements
The qPCR analysis used ΔΔCt measurements to derive splicing inclusion ratios, so validation did not depend on sequencing every patient. The published method includes a corrected equation in the 2016 erratum.
03Connecting the signal to disease
Nested RT-PCR across 55 tumour–normal pairs detected DUNQU1 in most HCC tumours, with no detectable signal in most adjacent normal liver samples. Four male HBV-positive pairs showed signal in both tissues (Figure 2c), so the pattern was preferential rather than universally tumour-exclusive. The FGFR2 splicing pattern was associated with clinical features, including tumour size and recurrence.
04Testing cancer-cell behaviour
Collaborators increased DUNQU1 expression in Huh7 liver-cancer cells and observed enhanced colony formation in soft agar, providing functional evidence of altered cancer-cell behaviour (Figure 3). This cell-culture experiment did not establish a treatment effect in patients.
05The next question
Which of these RNA signals adds useful information in independent patients, and which reflects biology that can be therapeutically investigated?
Technical context & publication

HCC means hepatocellular carcinoma. RNA sequencing generated approximately 250 million paired-end reads from one tumour–adjacent-normal pair; empirical validation used 55 pairs. The qPCR method derives splicing inclusion measurements from ΔΔCt values; the corrected equation is provided in the 2016 erratum (DOI 10.1038/onc.2016.62). Protein-coding potential is not by itself proof of a functional protein or a validated therapeutic target. The account of skepticism reflects my experience, rather than a claim of universal agreement across the field.

Read the Oncogene paper ↗

02 / Genome Research · 2018

A hidden RNA switch connected to patient outcomes.

Discovery at scale: approximately 6,000 RNA-seq samples.

AN ASSUMPTION PUT TO THE TEST

The assumption I challenged

An almost unchanged AFMID total suggested little to investigate. Tumour heterogeneity made reproducible splicing patterns seem unlikely, while existing approaches struggled to capture complex changes involving several exons.

What we showed

AFMID’s overall expression barely changed, but the balance of its RNA versions did. The human-specific switch was already present in early liver cancer and associated with poorer outcomes.

My contribution

I developed an analysis for complex splicing events and applied it across approximately 6,000 RNA-seq samples, connecting the switch with clinical and molecular evidence.

The evidence we established

Its recurrence association had similar statistical significance to MKI67, the gene encoding the familiar proliferation marker Ki-67.

Why it matters

The integrative analysis connected an RNA splicing switch with the production of NAD+, a molecule important for metabolism and DNA repair, cancer-driving mutations, and liver-cancer recurrence. It uncovered a mechanistic link that gene totals alone had missed—and a biological pathway to investigate for biomarkers and treatment strategies.

An unchanged total can hide a biological switch.

CSHL news · Read the research story ↗

ONE GENE. DIFFERENT MESSAGES.
The same total can conceal a different compositionSample ASample BFull-length versionShortened version
Looking at RNA versions exposes a switch that a gene-level summary can miss.Illustrative proportions, not patient measurements. The AFMID study investigated actual isoform changes across biological datasets.
Explore the evidence & discovery process
01Testing a complex splicing event
The PSI analysis represented changes involving several skipped exons. The approximately 6,000 RNA-seq samples included liver tissues, different cell types, and cancer cell lines across datasets.
02Examining patient outcomes
Low versus high full-length AFMID splicing yielded reported hazard ratios of 1.7087 for death (P = 0.0035) and 1.8822 for recurrence (P = 3.60 × 10⁻⁵). Similar recurrence P-values for AFMID and MKI67 do not establish equal predictive accuracy; the analysis limitations are described below.
03Distinguishing tumour biology
Lower full-length AFMID was associated with TP53 mutations. CTNNB1—another common liver-cancer driver—showed no significant enrichment despite a similar mutation frequency. This raised the question of whether the switch helps distinguish biologically different forms of liver cancer.
04Choosing the relevant experimental model
The switch was human-specific. An unmodified mouse does not directly reproduce that splicing event; experiments must use a system that captures the human mechanism.
05Testing the biological effect
The team increased full-length AFMID in HepG2 liver-cancer cells. NAD⁺ levels rose and cell growth slowed, providing experimental support for a connection between the RNA switch and cell metabolism (Figure 4).
06The next question
Does the RNA-version balance improve prediction in independent patient cohorts, and do the effects seen in cell culture extend to other relevant models?
Technical context & publication

AFMID produces alternative RNA isoforms. The study analyzed approximately 6,000 RNA-sequencing samples across human and other biological datasets. In the reported qPCR comparison of 20 adjacent non-tumour livers and 19 HCCs, overall AFMID expression did not change significantly, while the full-length isoform decreased. A human-specific switch does not make mouse models universally uninformative; it limits their direct use for this mechanism without appropriate engineering. Figure 2C,D reports median-split PSI comparisons in TCGA LIHC: 130 patients for survival and 174 for recurrence, after excluding right-censored records with follow-up information only. That exclusion limits interpretation and generalization of the reported hazard ratios; they are not individual absolute risks or proof of independent predictive value. The recurrence log-rank P-value was similar to that for MKI67 (3.80 × 10⁻⁵), but similar P-values do not establish equivalent prediction accuracy. Across 369 HCCs, low-MKI67 samples had higher AFMID full-length PSI (P = 3.71 × 10⁻⁷). These associations motivate validation with appropriate censoring and adjustment for other prognostic factors. In the reported nonsilent-mutation comparison across 369 HCCs, 37 of 54 TP53-mutant samples fell in the low-AFMID full-length group (hypergeometric P = 0.0016). TTN and CTNNB1 were mutated in similar numbers of samples but showed no significant enrichment. CTNNB1 is a recognized HCC driver, not a passenger-gene comparator. These are reported enrichment tests, not proof that the associations differ statistically from one another or that the splicing switch causes TP53 mutations; absence of significance is not proof of no association. ARID1A enrichment was reported in the separate extreme-quartile analysis.

Read the Genome Research paper ↗

03 / Nature · 2019

An underestimated mutation. A consequential mechanism.

Revisit accepted results and follow the evidence to mechanism.

AN ASSUMPTION PUT TO THE TEST

The assumption I challenged

The flagship TCGA AML study reported SRSF2 mutations in fewer than 1% of patients. Taken at face value, that made the mutation appear uncommon and easy to give less attention.

What we showed

Reanalysis revealed SRSF2 mutations that the original report had largely missed. The collaborative study connected SRSF2 and IDH2 mutations to faulty INTS3 RNA processing.

My contribution

I recognized the mismatch in the RNA evidence, led the computational reanalysis, and checked the result across patient cohorts.

The evidence we established

Three cohorts converged on roughly one in nine patients. Restoring INTS3 slowed leukemia progression in an experimental mouse model.

Why it matters

The study corrected an underestimated mutation frequency and provided functional evidence that mutations in the machinery that splices RNA can drive myeloid blood cancers. Identifying the INTS3 mechanism—and showing that restoring INTS3 slowed leukemia in a mouse model—connected a missed mutation signal to how the disease develops and where future treatments might intervene.

A published frequency depends on the analysis behind it.

CSHL news · Read the research story ↗

2019 · REVEALING AN UNDERCOUNTED MUTATION

A rare finding became a consistent signal.

<1%Originally reported in TCGAOne case · NEJM 2013
≈11%Found in each of three patient groupsOur 2019 analysis
A different reading of existing data revealed a mutation affecting roughly one in nine patients across these independent cohorts.
Evidence & why this changed the research question

The 2019 paper states that the original TCGA AML publication reported one SRSF2-mutant case. Reanalysis identified 19 of 179 patients with RNA-sequencing data. The other two cohorts independently showed similar frequencies; these are cohort estimates, not a universal frequency for all AML patients.

The discovery led to a biological question: how does this mutation help drive leukemia? The collaborative study linked abnormal INTS3 RNA processing to loss of its protein. Restoring INTS3 slowed leukemia progression in a mouse model.

2019 study · Figure 1b–d ↗
Explore the evidence & discovery process
01Following a mismatch in the RNA
RNA-processing patterns suggested that the original mutation calls did not capture all affected patients. That discrepancy prompted a re-examination of the underlying data.
02Connecting mutations to mechanism
The collaborative work identified cooperation between SRSF2 and IDH2 mutations. Abnormal INTS3 splicing triggered RNA degradation and reduced protein expression, disrupting the Integrator complex.
03Testing biological consequence
Forced INTS3 expression promoted differentiation and slowed progression in IDH2/SRSF2 double-mutant HL-60 xenografts. This was experimental rescue evidence, not a treatment trial in patients.
04The next question
Which parts of this mechanism are therapeutically tractable, and which patient groups would be most informative for testing them?

Supporting analysis: age at diagnosis

Age context comes from a separate calculation using cBioPortal’s OHSU Nature 2018 study (aml_ohsu_2018), accessed 10 September 2026. Among mutation-profiled patients with age at diagnosis ≥70, 34/148 unique patients carried an SRSF2 mutation (23.0%). Counting samples instead gives 36/171 (21.1%), because some patients contributed multiple samples. Patients without recorded age were excluded. This portal subset differs from the 498-patient RNA analysis in our paper; it is not an additional independent cohort. The U.S. median age at AML diagnosis is 70 (SEER, 2019–2023).

Beat AML dataset · cBioPortal ↗Age at diagnosis · NCI SEER ↗

Why this matters today

Four studies connect SRSF2 alterations with rapid growth of mutant blood-cell populations, leukemia, future malignant disease, and earlier mortality. Population-scale UK Biobank analyses complement long-term tracking of individual clones.

Clonal expansion

>50% / year

SRSF2 P95H was associated with exceptionally rapid growth of mutant blood-cell populations, compared with about 5% for DNMT3A and TP53.

385 people · median 13-year follow-up

Fabre et al. · Nature · 2022

Disease association

82× odds

Myeloid leukemia odds: 82× for SRSF2, compared with 18× for ASXL1 in the same study.

454,787 UK Biobank exomes · population-scale study

Backman et al. · Nature · 2021

Human lifespan

Mortality HR 5.8

Mortality HR 5.8 for SRSF2, compared with 1.5 for both DNMT3A and TET2 in the same predicted-pathogenic missense analysis.

393,833 UK Biobank participants · discovery analysis

Park et al. · Nature Communications · 2025

HR = hazard ratio, a comparison of event rates over time. These relative measures describe different study populations and outcomes, not individual absolute risk.

Human studies connect SRSF2 alterations with rapid clonal expansion, myeloid malignancy, future progression, and earlier mortality. Our 2019 work identified INTS3 mis-splicing as one biologically consequential route to leukemia development—giving researchers a mechanism to investigate and potential vulnerabilities to test.

Study details, comparison groups & interpretation

Clonal expansion

Fabre et al. tracked 697 clones in 385 SardiNIA participants aged 55 or older. Estimated growth ranged from about 5% per year for DNMT3A and TP53 to more than 50% for SRSF2 P95H. This is relative clonal expansion, not a yearly probability of leukemia, and concerns this specific SRSF2 mutation.

Fabre et al. · Nature · 2022

Disease association

Backman et al. sequenced 454,787 participants. The SRSF2 rare-variant burden association with myeloid leukemia (ICD-10 C92) had an odds ratio of 82 (95% CI 44–150; P = 1.96 × 10⁻⁴³) in the European-ancestry analysis. This is a gene-burden association, not a prospective progression estimate for every SRSF2 mutation. For the same myeloid leukemia outcome, Supplementary Table 14 reports ASXL1 OR 18 (95% CI 11–28). These are each gene’s carrier-versus-non-carrier associations, not a direct comparison between carriers of the two genes. The reported burdens use different variant masks: SRSF2 M3.01 and ASXL1 M1.01. Supplementary Tables 6 and 14 report the results; the authors identify evidence consistent with somatic mutations in blood-derived DNA.

Backman et al. · Nature · 2021

Future progression

The study screened 438,890 participants; this estimate comes from 11,337 CHIP/CCUS participants in the derivation cohort. Among participants with CHIP/CCUS, SRSF2 mutation was associated with a hazard ratio of 17.24 (95% CI 12.44–23.87) for subsequent myeloid neoplasms, compared with CHIP/CCUS participants without that mutation. Figure 1C reports these univariate associations: TET2 HR 1.838 (95% CI 1.413–2.397), ASXL1 HR 1.869 (1.340–2.606), and SRSF2 HR 17.24. Each compares CHIP/CCUS participants with the named mutation against those without it; these are not direct comparisons between gene groups. SRSF2 was among the high-risk genes, but JAK2 had a higher estimate (HR 32.30). It is not an estimate for an isolated mutation independent of all other risk factors. CHIP/CCUS refers to acquired mutant blood-cell populations, without or with unexplained low blood counts.

Weeks et al. · NEJM Evidence · 2023

Human lifespan

The discovery analysis included 393,833 European-ancestry participants. Park et al. reported mortality HR 5.8 for carriers of AlphaMissense-predicted pathogenic SRSF2 variants versus non-carriers (Cox-analysis P = 3.3 × 10⁻⁶¹). The same AlphaMissense analysis reported DNMT3A HR 1.5 and TET2 HR 1.5, each against its own non-carrier group. These point estimates provide context rather than a formal test of differences between genes. SRSF2 had the largest HR among the significant AlphaMissense gene burdens reported in that analysis. The separate lifespan burden test gave P = 1.8 × 10⁻⁹⁴. These are distinct analyses, not a single HR/P-value pair; HR 5.8 does not mean a lifespan 5.8 times shorter.

Park et al. · Nature Communications · 2025

These studies examine different populations and outcomes; the UK Biobank analyses also draw on overlapping participants. Cross-gene comparisons within a study provide context; they are not head-to-head estimates or proof that the effects differ statistically. Their estimates cannot be multiplied or treated as one inevitable disease trajectory. Clonal growth, relative odds, and hazards do not by themselves establish an individual’s absolute risk.

Technical context & publication

AML means acute myeloid leukemia. Figure 1 reports SRSF2 mutations in 19/179 TCGA, 57/498 Beat AML, and 28/263 Leucegene patients. INTS3 mis-splicing triggered nonsense-mediated decay and reduced protein expression. Forced INTS3 expression promoted differentiation and slowed progression in IDH2/SRSF2 double-mutant HL-60 xenografts, not a clinical trial. The separate cBioPortal estimate uses age at diagnosis ≥70 and unique mutation-profiled patients; accessed 10 September 2026. Clinical frequency depends on the cohort, measurement, and analysis.

Read the Nature paper ↗

WHAT CONNECTS THE WORK

The signal was there.
The view had to change.

RNA sequencing

Look beyond predefined measurements.

Alternative splicing

Look beyond a gene’s total activity.

Mechanistic analysis

Look beyond a mutation’s label.

AI evaluation

Look beyond a model’s intended use.

(Coming soon after publication)

Extensive data can still leave the important biology undiscovered. My contribution is recognizing what the data could reveal, developing the analysis needed to test it, and helping turn the result into the next scientific decision.

THE SAME QUESTION, NOW APPLIED TO AI

What can this capability reveal
that our current methods cannot?

The shift from microarrays to RNA sequencing taught me to ask what new biology a technology makes accessible. I bring the same question to AI: what useful signal might a model contain beyond the task it was designed to perform?

The Discovery Loop describes how I approach that question. Revisit overlooked evidence. Test an unexpected capability. Let the result change the explanation—and reveal the next question worth asking.

BETTER QUESTIONSBetter evidence. New possibilities.12345

Question

Define what we need to understand. An unexpected observation, an unmet need, or a new technology can open a question that current methods have not answered.

One discovery, step by step
AFMID example: Could the choice of RNA version reveal liver-cancer biology that total gene activity obscures?

Each discovery opens the next question. Evidence changes what we ask and how we investigate. A result may support an explanation, rule it out, or reveal a new possibility. The next cycle starts from what we have learned.

MY VISION FOR AI DISCOVERY

New understanding is the measure.

What defines an AI innovator is not the ability to generate known hypotheses, but the ability to turn an unexpected insight into a new discovery—building and testing the evidence until it changes what we understand about biology.

AI IN PRACTICE · SUPPORTING EXAMPLEHow I worked with Astra to check this site’s evidence

My question: how can a non-specialist see why the SRSF2 discovery mattered, without losing the distinctions that make the evidence trustworthy?

My scientific direction

I identified the story worth telling, challenged vague explanations, pointed to the age-specific finding in cBioPortal, and asked that the claims be checked. I also set the boundaries around personal information and collaborative credit.

Astra’s assistance

Astra helped inspect the paper, calculate frequencies from public data, simplify the explanation, and implement the interactive website. My feedback shaped what to investigate and how the findings should be presented.

Try changing the question

Does the finding hold across independent groups?
TCGA10.6%19 of 179 patients
Beat AML11.4%57 of 498 patients
Leucegene10.6%28 of 263 patients

Similar frequencies across three cohorts support a recurring finding. They do not establish an identical frequency in every patient population.

Inspect Figure 1b–d · Nature 2019 ↗

What changed as we checked the evidence?

  • “0%” became “less than 1%”: the paper describes one originally reported TCGA case.
  • “About 20% in older patients” became a precise, reproducible comparison: 23.0% of unique patients versus 21.1% of samples.
  • The age-specific portal calculation was separated from the three published cohorts.

The value of the interaction was in turning a scientific insight into a clear explanation whose numbers, assumptions, and limits can be inspected.

Data and calculation notes

Fixed public-data snapshot, 10 September 2026; these controls select verified summaries, not a live AI response. cBioPortal study aml_ohsu_2018, mutation profile aml_ohsu_2018_mutations, SRSF2 gene 6427. Only mutation-profiled samples with recorded age at diagnosis were included. Patients were counted once and marked mutation-positive if any included sample carried an SRSF2 mutation. The age filter is ≥70. The all-age comparison includes all recorded ages, with no additional disease-stage filter.

These descriptive percentages do not establish a causal age effect, predict an individual’s outcome, or replace a formal clinical analysis. Astra helped communicate and inspect published research here; it did not make the original biological discoveries.

WHAT I BRING

Expanding what’s possible

A new technology creates possibilities. Scientific judgment determines which possibilities are real, what they mean, and which deserve investment.

The examples highlighted here emphasize work that can be shared publicly; much of my biotechnology experience has involved proprietary discovery programs and platform development.

New things to pursue

Previously unseen biology can suggest therapeutic targets, biomarkers, diagnostic approaches, or opportunities for new intellectual property.

Better ways to decide

Mechanistic understanding can sharpen patient-stratification hypotheses, improve experimental-model choices, and focus drug-development decisions.

Earlier evidence about what is real

Rigorous testing can establish a promising direction, expose a limitation, or show when a different question is more valuable.

THE CAPABILITIES BEHIND THE DISCOVERIES

Scientific judgment,
supported by technical depth.

01

AI & technology evaluation

Find capabilities beyond a model’s original objective. Test predictions, understand limitations, and define the evidence needed for adoption.

02

Computational biology

Connect RNA biology, functional genomics, and multi-omics to explain disease mechanisms and generate testable hypotheses.

03

Drug discovery strategy

Translate molecular and clinical evidence into target prioritization, patient-stratification hypotheses, and the next experiment.

04

Discovery from limited patient data

Extract useful biological signals from sparse or challenging clinical material, then strengthen the interpretation with independent evidence.

05

Scientific leadership & platforms

Build computational functions and reproducible workflows. Bring computational, experimental, and leadership teams to a shared scientific decision.

BUILDING WHAT THE QUESTION NEEDS

When the method is missing,
develop it.

PSI-SIGMA · BIOINFORMATICS 2019

Making RNA differences easier to detect

I designed and implemented PSI-Sigma to analyze alternative splicing using short- and long-read RNA sequencing. Developing a reusable method makes a way of seeing biology available to other researchers.

Value: a published analytical tool that supports further discovery.

Read the method ↗

SMN2 · NUCLEIC ACIDS RESEARCH 2022

Connecting sequence changes to therapeutic design

My analysis of minigene experiments mapped elements controlling SMN2 splicing, including previously uncharacterized elements. The work connected experimental measurements to a biological question relevant to antisense-oligonucleotide design.

Read the study ↗
Technical foundation

Python, Perl, C/C++, JavaScript, Linux/Unix, AWS, Nextflow, RNA-seq, NGS, and multi-omics analysis.

BACKGROUND

Scientific depth.
A broader perspective.

Across academic research and biotechnology, I have repeatedly worked at the boundary of what our tools can reveal.

I work across computational and experimental disciplines, helping teams turn complex findings into clear scientific recommendations. My PhD in Biomedical Informatics and MBA in Management Information Systems inform both the technical work and the decisions around it.

Education

PhD, Biomedical Informatics
National Yang-Ming University · 2009–2013

MBA, Management Information Systems
West Texas A&M University · 2004–2005

  1. 2023–2026

    Codify Therapeutics

    Associate Director, Bioinformatics

  2. 2020–2023

    Skyhawk Therapeutics

    Principal Scientist, Bioinformatics & Machine Learning

  3. 2015–2022

    Cold Spring Harbor Laboratory

    Adrian R. Krainer Lab
    Computational Postdoc, 2015–2020
    Visiting Scientist, 2020–2022

  4. 2006–2015

    Early research in Taiwan

    Research roles at Taipei Medical University, National Yang-Ming University, and Academia Sinica

INVENTION
Co-inventor, US10953059B2
Methods and compositions for treating non-small cell lung cancer

RECOGNITION
SOAR Award, Skyhawk Therapeutics
2021

SELECTED TALKS
National Academies, 2023
London Calling, Oxford Nanopore, 2019

THE PUBLISHED RECORD

A record of
collaborative discovery.

Selected collaborative research. Labels show my authorship position; co-first denotes a shared-first contribution. Some titles are abbreviated for readability.

Full publication bibliography ↗