Clinical Trial Site Selection and AI: Look-Alike Modeling Across Real-World Data and Care Delivery Signals

6 hours ago 6

Rommie Analytics

 Look-Alike Modeling Across Real-World Data and Care Delivery SignalsAaron Cohen, Vice President, Data Strategy at Definitive Healthcare

Site selection and enrollment failures are persistent challenges for life sciences organizations. A Tufts Center for the Study of Drug Development analysis of nearly 16,000 investigative sites across 151 Phase II and III global trials found that nearly half of all selected sites either failed to enroll a patient or under-enrolled.1 Site-selection decisions commonly draw on historical performance indicators like prior trial participation and past enrollment, typically favoring sites with an established track record over fresh site opportunities. This creates a sampling problem: if the pool of potential sites is largely composed of those selected in previous trials, analyses of that pool will naturally favor characteristics associated with prior selection. When sponsors draw candidates from a pool shaped by prior selection decisions, established sites are more likely to be selected again and trial speed, diversity, and overall quality are potentially impacted.   

AI can help organizations identify new clinical trial sites by evaluating a broader range of characteristics to inform site selection so that appropriate candidates are not overlooked. Using predictive algorithms to target new research characteristics including suitable patient populations, investigators, and clinical capabilities, among other data points, sponsors can potentially identify sites that are a unique fit, expanding their approach to site selection.

Basing site selection on trial history

Sponsors commonly evaluate sites across four broad categories: likely eligible-patient volume, clinical capability, investigator availability, and prior-trial performance. Problematically, these criteria are frequently measured using historical proxies, with past trial participation among the most readily available. As a result, these screens tend to return sites that already have a research history; they also fail to address the highly protocol-specific nature of enrollment.

By leveraging different data points for site selection, sponsors can identify the next generation of capable clinical trial sites. The key to tapping into new opportunities is broadening sponsors’ traditional searches to include specific characteristics of a particular trial. 

Further informing a sponsor’s site selection approach, data shows that trial participation varies sharply by care setting. For example, across Commission on Cancer programs from 2013 to 2017, the overall adult cancer treatment trial participation rate was 7.1%.2 

Participation rates varied significantly by program type:

21.6% at NCI-designated comprehensive cancer centers4.1% at community programs5.4% at academic, non-NCI-designated programs5.7% at integrated network programs

While these NCI findings reflect data from several years ago, little suggests that site selection practices have changed meaningfully in the interim. This noted variation is impactful because a site with a strong trial participation history is likely to appear more attractive in a selection process than those with little or none. Prior participation is a real signal. The problem is that it is also a filter, and using it as a primary screen ensures the pool never expands.

Selecting the next generation of sites with AI

By identifying specific characteristics related to clinical capability, patient population, investigator or site capability, and data visibility, AI-enabled tools can perform a look-alike analysis to find sites that resemble other ones that have successfully conducted a very similar trial, potentially surfacing new sites that are an appropriate match but lack a trial track record.

Using real-world patient claims data, procedure volumes, investigators’ expertise, and capabilities to perform specific services and procedures, AI can evaluate broader signals unrelated to trial history. This can be particularly important in the rare disease space, where diagnosis codes are often not specific enough to identify the patients who meet a study’s criteria. By incorporating broader clinical and longitudinal signals, AI can estimate which patients are likely affected by the disease of interest and, by extension, where those patients are receiving care. This can help sponsors identify qualified sites that may be overlooked by diagnosis-code- or trial-performance-based approaches alone. 

Although these models rely on prior data, they can use care-delivery and patient-level signals to estimate future suitability without requiring prior performance in a similar trial. AI is not a shortcut around the limitations of the underlying data, but rather a tool that can evaluate more data at scale. The opportunity comes from evaluating different signals, such as patient populations, care-delivery patterns, investigator expertise, and protocol-specific capabilities that can help identify qualified sites traditional approaches may overlook.  

The foundation of any analysis is and always will be the quality of the input data. AI makes it possible to evaluate larger and more complex datasets. Where data capture can be reliably estimated, it may also help differentiate genuinely low patient volume from limited data visibility.

It can be effective at extracting variables from clinical text where appropriately governed data are available to supplement site validation. AI can also be used for modeling across full feature sets at scale, beyond manual feasibility capabilities, to screen large numbers of sites that lack prior trial experience.

While AI cannot fully resolve trial eligibility from claims data, determine site selection, or directly recruit appropriate patients, it can quickly provide valuable estimates, helping sponsors better understand the patient population that is likely to meet treatment criteria. Of course, sponsors will still need to thoroughly assess the capabilities of sites and investigators, and consider the primary factors for managing and running the trial.

A practical way to evaluate these approaches is through a controlled exploration cohort. Sponsors could reserve a defined percentage of candidate sites for model-identified, trial-naïve organizations and subject them to the same readiness assessments and operational reviews as traditional candidates. Enrollment and study performance could then be measured against established benchmarks. This allows organizations to test whether expanding the candidate pool improves recruitment and site performance without abandoning proven site-selection practices.

Life sciences organizations don’t need to abandon their current site-selection approach entirely, but they should recognize that relying too heavily on past enrollment performance can create a self-reinforcing system in which the sites selected yesterday become the sites considered tomorrow. This approach comes at a cost: a smaller and less diverse enrollment pool, limited geographic and demographic representation, overlooked patient populations, slower recruitment, and higher costs. By incorporating AI-enabled approaches into site selection, sponsors can base decisions on capabilities and patient populations in addition to prior trial experience, broadening the candidate pool, expanding potential matches, and reaching more patients.


References

Getz, K. “Enrollment Performance: Weighing the ‘Facts’.” Applied Clinical Trials, May 2012, Vol. 21, Issue 5. Tufts CSDD study of nearly 16,000 investigative sites across 151 Phase II and III global trials conducted 2008–2010, with performance data supplied by 10 pharmaceutical companies and two CROs; typical trial ran in 11 countries among adult patients. Enrollment activation rate is defined as the percentage of sites ready to enroll that enroll at least one patient. Findings: 11 percent of ready-to-recruit sites enrolled no patient; 41 percent of activated sites did not achieve the target enrollment number, defined as a band from sites enrolling only one patient through those missing target by 10 percent or more; 40 percent hit target and 15 percent exceeded it; sites enrolling at least one patient reached 85 percent of target on average; and 48 percent of all selected sites enrolled nobody or under-enrolled.National Estimates of the Participation of Patients With Cancer in Clinical Research Studies Based on Commission on Cancer Accreditation Data,’ Journal of Clinical Oncology, 2024, doi:10.1200/JCO.23.01030 (covering 2013-2017).

About Aaron Cohenthe Author: 

Aaron Cohen is Director of Life Sciences Solutions Consulting at Definitive Healthcare. He specializes in healthcare claims data and real-world data, with deep expertise in claims analytics, provider reference data, patient tokenization and linkage, and cloud analytics. In his more than seven years at Definitive, Aaron has worked in Customer Success, Solutions Consulting, and leadership roles. He leads a team of subject matter experts supporting biopharma, diagnostics, and medical device customers, and is a co-author on a conference poster analyzing Medicare claims data.

Prior to joining Definitive, Aaron focused on patients’ access to novel medical devices at a reimbursement consulting firm, where he was a Definitive Healthcare customer. Aaron is completing an MS in Data Science at the University of Colorado Boulder.

Read Entire Article