Search
AI News ARCHIVE

How is Artificial Intelligence Changing How We do Science?

Since the late 1980s particle physicists have used AI even as the concept of a neural network was barely in the public’s consciousness. AI and particle physics go hand in hand as the experiments the physicists perform usually revolves around seeking out patterns in…

Daniel Okafor
Daniel OkaforSenior AI Reporter
12 min read
How is Artificial Intelligence Changing How We do Science?

Since the late 1980s particle physicists have used AI even as the concept of a neural network was barely in the public’s consciousness. AI and particle physics go hand in hand as the experiments the physicists perform usually revolves around seeking out patterns in the data from particle detectors and AI is excellent at pattern detection. Boaz Klima, a Physicists from the Fermi National Accelerator Laboratory, also called Fermilab, says “It took us several years to convince people that this is not just some magic, hocus-pocus, black box stuff.” He was amongst the first to adopt AI tools but today, it’s a part of standard particle physics practices.

Neural networks search for fingerprints of new particles in the debris of collisions at the LHC. © 2012 CERN, FOR THE BENEFIT OF THE ALICE COLLABORATION

Usually, particle physicists aim to comprehend the way the inner gears of the universe works, typically by colliding subatomic particles at hit speeds to break them down into even smaller and more unusual kinds of matter. For example the Higgs boson particle, discovered in 2012 by physicists using the Large Hadron Collider, the inscrutable particle is the key to a physics explanation of they way that fundamental particles gain their masses.

However, working with such fleeting particles is tricky and does not come with naturally occurring instructions. The Higgs boson only results in one out of one billion proton collisions, and it very quickly, in one-billionth of a picosecond (already one-trillionth of a second), decays into different particles like a couple of photons or four muons, a particle similar to an electron.  In order to study the Higgs boson particles, they must be ‘reconstructed’ from that other particulate matter and determine if they go together consistent with the parent particle, a difficult task due to the extra particles generated in the same collision.

According to Pushpalatha Bhat, a Fermilab physicist, algorithms, and neural networks do a superb job of sifting pertinent data from background data. Inside particle detectors, normally a massive cylindrical structure full of a variety of sensors, a photon forms a shower or spray of particles in an electrometer calorimeter, a type of subsystem. And while the particles they produce are different, electrons and hadrons produce similar showers as well. To tell the difference, machine-learning algorithms are used to detect correlations found in the variables that make up the showers. For example, an algorithm is able to differentiate between Higgs boson related photon pairs and other random photons. “This is the proverbial needle-in-the-haystack problem,” Bhat says. “That’s why it’s so important to extract the most information we can from the data.”

Despite advancements in machine learning, however, physicists remain dependent on their comprehension of the underlying science in order to determine how to sort through results for evidence of new phenomena and particles. But AI is becoming more and more significant tells computer scientist Paolo Calafiura from Lawrence Berkeley National Laboratory. Already plans are in action to upgrade the Large Hadron Collider at CERN in order to speed up the collision rate to ten times its current speed in 2024. Then, Calafiura explains, machine learning will be essential if scientists want to keep up the massive amount of data produced from particle collisions.

Algorithms for Social Media

Psychologist Martin Seligman and colleagues from the University of Pennsylvania’s Positive Psychology Center, have recognized the opportunity to utilize AI in the analysis of the massive amount of social data that can be collected from online social media. The World Well-Being Project combs through the hundreds of billions posts and tweets from billions of users using natural language processors and machine learning to track both physical and emotional health of the public.

Usually, this kind of data collection is achieved through surveys, but mining social media is “unobtrusive, it’s inexpensive, and the numbers you get are orders of magnitude greater,” explains Seligman. Although it is messier than traditional methods, AI is a great tool for cleaning it up and revealing important patterns.

Recently, in a study performed by Seligman and team, the Facebook posts of 29,000 users who had already depression self-assessment test. An algorithm correctly found correlations between words used in updates from 28,000 of those posts and levels of depression then due to its machine learning ability the AI was able to determine depression levels in users based solely on their update posts.

Additionally, the team could determine countywide mortality rates related to heart disease by using AI to analyze 148 million tweets. It turns out that words concerning negative relationships and anger could more accurately predict mortality rates than predictions relying on the 10 heart disease risk factors, including diabetes and smoking.

Furthermore, the researchers were able to develop a color-coded map of each U.S. county displaying levels of depression, well-being, trust, and 5 personality characteristics simply by examining Twitter. They’ve also been able to predict political beliefs, income, stereotypes, mystical experiences, evaluate hospital care, and individual personality solely on data gathered from social media.

“There’s a revolution going on in the analysis of language and its links to psychology,” explains James Pennebaker, a social psychologist at the University of Texas at Austin. By focusing on style rather than content he is able to examine the use of certain word types and predict the grade outcomes of potential students by what they write in college admission essays. For example, the use of prepositions and articles are indicators of analytical thinking and therefore translate to the likelihood of higher grades. Meanwhile, an abundance of adverbs and pronouns are indicators of narrative thinking, which translates to lower grades.

Also, through the analysis of language use, Pennebaker was able to support claims that a significant portion of the 1728 play Double Falsehoods was written by William Shakespeare. According to the machine learning algorithms, factors like rare words and cognitive complexity match up with other Shakespeare works.

“Now, we can analyze everything that you’ve ever posted, ever written, and increasingly how you and Alexa talk,” Pennebaker adds. As a result, they are developing  “richer and richer pictures of who people are.”

Finding the source of autism through examination of the human genome

Autism is a challenge for sufferers, caregivers, and the geneticists searching for causes of the disorder. It is known to have strong inheritance patterns and thus a powerful genetic component. However, only around 20 percent of cases can be explained by the currently known genetic variations. To find more variations that lead to autism researchers have search for indicators in the 25,000 other genes and their related DNA, which is a daunting task for geneticists.

As a result of Olga Troyanskaya, a computational biologist at Princeton University and the Simons Foundation has recruited AI tools to reveal thousands of genes that participate to the spectrum of autism disorders.

“We can only do so much as biologists to show what underlies diseases like autism,” explains Robert Darnell, founding director of the New York Genome Center and a physician-scientist at The Rockefeller University, who is working with Troyanskaya. “The power of machines to ask a trillion questions where a scientist can ask just 10 is a game-changer.”

Troyanskaya took multiple data sets on the way proteins interact, on what genes are functioning in certain cells, and in which area are located relevant transcription factor binding sites and other important genome features. After combining the information, the team built a guide of different gene interplays and contrasted those of the known gene indications for autism risk with the thousands of those that were unknown. When similarities were found, that added 2,500 more genes to the list of those implicated in cases of autism, as reported in the findings published last year in the journal Nature Neuroscience.

However, geneticists have recently found that the genes do not act alone but are in fact influenced by millions of local noncoding bases, which react with DNA-binding proteins and additional determinants. Suddenly their next task became to identify the noncoding variations which might affect adjacent genes responsible for autism, an even greater puzzle than the one they’d just seemingly solved and once again they are turning to AI for help.

This time Jian Zhou, a graduate student in Troyanskaya’s Princeton lab, was charged with teaching the deep-learning program. Zhou trained the system by exposing it to information gathered by Roadmap Epigenomics and the Encyclopedia of DNA Elements, two cataloging projects with a database of tens of thousands of noncoding DNA sites and their effects on neighboring genes. This taught the system, called DeepSEA, what to look for as it analyzed mysterious areas of noncoding DNA for possible activity.

After reading the efforts of Troyanskaya and Zhou in the journal Nature Methods last October, Xiaohui Xie, a computer scientist from the University of California at Irvine, called their program “milestone in applying deep learning to genomics.”

Currently, Troyanskaya’s Princeton team is examining autism patients’ genomes with DeepSEA to gain a deeper understanding of the effects of the noncoding bases.

Meanwhile, Xie is also using AI in genome work. He is also looking for mutations that lead disease and disorders, with a broader scope than autism. He is proceeding cautiously as deep learning systems can only be as reliable as their original provided data sets. “Right now I think people are skeptical”, he says. “But I think down the road more and more people will embrace deep learning.”

Clarifying the sky with AI

In April, Kevin Schawinski, an astrophysicist, took to Twitter looking for help. He had four fuzzy pictures of galaxies he couldn’t classify and asked if another astrophysicist might have more luck. Some answered that they appeared to be spirals and ellipticals, well-known galaxies types, while others, more familiar with Schawinski, asked if it was a trick: Were the pictures of real galaxies or only computer modeled simulations?

AI that “knows” what a galaxy should look like transforms a fuzzy image (left) into a crisp one (right). KIYOSHI TAKAHASE SEGUNDO/ALAMY STOCK PHOTO

The truth was a little more complex, he answered. Schawinski, Ce Zhang, a computer scientist, and other from ETH Zurich in Switzerland, had the pictures created by a neural network that didn’t have any knowledge of physics but did have a deep, innate understanding of how galaxies were supposed to look.

He posted the photos to Twitter to see how realistic the AI pictures were but with the technology, in general, he had a larger goal in mind. Schawinski wants to develop a film manipulation technology that emulates that magical surveillance function overused on TV and in movies: zoom and enhance. His true goal is to build a network that can take out of focusing long distance images of galaxies and correct and sharpen them to make them accurately appear as if the images were captured by a better telescope. Such an innovation could allow astronomers to glean for information and details from imaging data already gathered by older tools. “Hundreds of millions or maybe billions of dollars have been spent on sky surveys,” explains Schawinski. “With this technology, we can immediately extract somewhat more information.”

The infamous Twitter photos were the results of a type of machine-learning model, called generative adversarial network, that contains two contrasting neural networks designed to duel one another. One of the two systems is a generator that makes the images, while the second system is a discriminatory, designed to find flaws in the artificially created images, which in turn makes the generator better at the making the said images.

In order to teach the programs how to achieve his desired goals, Schawinski and his team degraded real pictures of galaxies and showed the generator how to refine the images until they fooled the discriminator. In the end, the network was capable of surpassing current techniques to “zoom and enhance,” and created more clear and smooth images of galaxies.

This use of AI in astronomy is rather unusual but it is not alone, according to Brian Nord, an astrophysicist from Fermi National Accelerator Laboratory. Nord presented his own machine-learning approach to seeking out gravitational lenses at the meeting of the American Astronomical Society in January. Gravitational lensing is the bending of light as it travels from its source to Earth due to spacetime and the galaxy and stars the light travels past. The effect is used in astronomy to calculate distances and discover invisible masses like black holes.

In this rare case, people are actually better at spotting this phenomenon than computers because the effect is difficult to describe in mathematical terms but visually unique to the human observer. Nord and colleagues wonder if AI trained by thousands of images of the lensing effect could learn to see the distinction. After a month, “there have been almost a dozen papers, actually, on searching for strong lenses using some kind of machine learning. It’s been a flurry,” Nord says.

And this case is yet another example of the nascent realization that AI and machine learning can be a powerful aid astronomy, especially when it comes to classifying petabytes of observational data. According to Schawinski “That’s one way I think in which real discovery is going to be made in this age of ‘Oh my God, we have too much data.’”

Teaching Computers Chemical Synthesis

Surprisingly, many organic chemists work backward. They start with a vision of the completed structure of a certain molecule and work from that to figure out how to construct it. “You need the right ingredients and a recipe for how to combine them,” said Marwin Segler, a graduate student from the University of Münster. Now he and others are adding Ai to the mix.

Their goal is use AI to concur one of the biggest challenges in making molecules: narrowing down the hundreds of options for molecule building blocks and the thousands of chemical guidelines ruling them. Over decades of intense and detailed work chemist have inserted known interactions into computers, desperately trying to build a program that could easily compute molecular recipes. Yet, according to Segler the field of chemistry “can be very subtle. It’s hard to write down all the rules in a binary way.”

Thus Segler, computer scientist Mike Preuss, and Mark Waller have come to AI for help. Rather than inserting hard line regulations for reactions, the team developed a deep learning neural network that learned the method by which chemical reactions occurred by being exposed millions of examples. “The more data you feed it the better it gets,” explained Segler. As the network improved it learned to expect the best possible outcome for a specific step in chemical procedure and even began making its own molecule making recipes.

By comparing and contrasting a typical molecular design program and the new AI program with 40 distinct molecule targets in a two-hour window, researchers found that the AI resulted in a recipe 95% or the time whereas the conventional computer only had a 22.5 percent success rate. Segler, who will be beginning work at a pharmaceutical company, wants to use this AI technique to improve the manufacturing of medicines.

Another organic chemist, Paul Wender from Stanford University is cautious regarding Segler’s approach to molecule synthesis but he’s optimistic about the possible future impact. He is also building his own AI synthesis program and thinks Segler’s “could have a profound impact” on the field.

In any case, Segler insists that his program and similar AI like it will not be replacing human organic chemists anytime soon since humans can do far more work than the programs. The AI may be able to find the route to a new molecule but it can’t design or complete the synthesis to make it.

Though it won’t be long before scientist turns their AI sights to those functionalities as well.

More News to Read

Related Coverage

Google · Preferred Sources

Don't miss new tech stories on Google

Add TrendinTech once in the Google app and our stories appear in your news suggestions.

Add Now
Daniel Okafor

Daniel Okafor

Senior AI Reporter

Daniel Okafor is the Senior AI Reporter at TrendinTech, where he covers large language models, machine learning research and the practical use of artificial intelligence across business and government. He previously reported on artificial intelligence for MIT Technology Review, covering the labs behind the current generation of frontier models and the policy debates in Washington and Brussels. Daniel holds a Master of Science in Machine Learning from Carnegie Mellon University and follows the research community closely, attending NeurIPS and ICML each year to speak with the people behind the papers. He has a particular interest in evaluation: how models are benchmarked, where those benchmarks fail and what that means for the companies betting on them.

All stories by Daniel Okafor (316)