This is the third post of an ongoing series on searching the academic literature. If you have not already done so, we invite you to check out part 1 on library catalogs, bibliographic databases, and repositories. It may be especially helpful to read part 2 on search interfaces for additional context for this post. Search engines are a type of search interface, and therefore the information provided here builds on that earlier post.
Search engines
A search engine (exp. Google, Bing, etc.) is a search interface for the full internet[1] rather than for curated collections of academic literature, like those covered in part 2 (EBSCO-host, ProQuest, etc.). Search engines are software programs that use proprietary algorithms to retrieve information[2] (more on search algorithms coming in Part 4). This is different from a web browser, which is the application that provides the graphical interface for using a search engine[3] (exp. Chrome, Safari, etc.). Search engines use bots to crawl the internet, collecting data and metadata from websites. The results that are returned when a user enters their search is based on both how the algorithm is designed and what the search engine’s bots have been able to collect. This process may be changing with the growth in the use of LLMs for information retrieval. Google is constantly changing its algorithm and is notoriously secretive about how it functions. Because of this, it is hard to tell if Google is still crawling and indexing the same way or if they are radically changing the ways in which Google functions. If you have noticed that you’re getting different results in your Google search than you expect or different results than you have seen for similar searches in the past, this is a sign that the algorithm has changed. The way Google’s algorithm selects the results it shows you can change from one day to the next.
Another important thing to note is that not everything on the internet is crawlable. Online information that exists behind a paywall may not be crawled or indexed adequately enough by a search engine to appear in search results. Since many academic journals and bibliographic databases are behind paywalls, they are not typically as well indexed in search engines as other types of online content. In these cases, search engines rely more heavily on the publicly available metadata. This means, unless an article is open access, it may not show up in a Google search. Additionally, due to the (often changing) ranking algorithm of a search engine that prioritizes things like paid and popular content for a general audience, you are far less likely to find relevant, credible academic literature through a search engine than you are through a tool designed for an academic audience.
Google Scholar
This brings us to Google Scholar (GS), which is also a search engine, but one designed for an academic audience. As with the larger Google search engine, GS uses bots to crawl the internet collecting information for indexing. Unlike regular Google bots, the scholar bots are programmed to identify items structured to look like academic literature. Their programming is designed to locate scholarly journal articles, conference papers, technical reports, dissertations, pre- and post-prints, and abstracts, and it does so based on how the information is structured on a website[4].
GS’s method of indexing materials is different than the methods used for curation by bibliographic databases. Databases target only materials from pre-identified, established academic sources (exp. journals, academic associations, and research centers) that fit within their collecting scope (see part 1 of this series for more about collecting scopes in databases). This gives GS the benefit of breadth (it may include relevant materials not indexed in a bibliographic database) but has no built-in means of assessing credibility or authority of the materials indexed. Instead, it uses factors like journal ranking, author impact, and citation count as part of the logic in its ranking system. These are certainly helpful in what filters to the top, but not foolproof. Often this leads to low-quality results that will ‘junk up’ your search results.
The search algorithm for GS more closely resembles that of Google than it does the search interfaces described in part 2 of this series. Rather than using keyword searching, GS uses semantic search, which looks not only for the words entered into the search, but any additional words or concepts the algorithm was designed to see in relationship to those keywords[5]. It is designed to return results based on natural human language patterns, more like mimicking a conversation with a human than with coding a query for a computer. This is part of why you usually end up with hundreds of thousands of results.
Google and Google Scholar also take your prior search history into account as well as your geographic location, what other people are searching, and other subjective elements to prioritize what you see first and what may get buried in page 12,456 of the results. Taken together, this way of designing the search algorithm takes away a large amount of your control over the search, and introduced biases into the process that you may not be aware of.
When to use it:
Broad search engines are valuable for searching things like gray literature, company and organizational reports or datasets, and blogs. Advanced Google Search functionality enables you to search specific domains (exp. site:.gov or even site:CDC.gov) or geographic regions. These types of materials can be useful as background information for a variety of academic purposes. However, search engines are generally not efficient for finding research studies or other types of rigorous academic literature.
GS is designed for certain types of academic purposes and is especially beneficial in identifying scholarly materials that bibliographic databases might miss, including conference papers, dissertations, other types of gray literature, and literature published in journals or geographic locations underrepresented by the bibliographic databases[6]. GS can also be helpful when you are unfamiliar with the relevant keywords for your field or topic and need to rely more on natural language mapping until you become more familiar with the academic language surrounding your topic.
There are downsides to using GS to search the academic literature. First, because the GS algorithm can’t differentiate between high-quality research and what has been structured to fool the bots, it is vulnerable to returning low-quality, noncredible sources including papermill and now increasingly AI-generated materials designed to look like research articles. This puts much more responsibility on you as the researcher to evaluate each article to determine what is a high-quality, credible source, and what is not. While it is still important to evaluate articles retrieved from the library catalog or bibliographic databases, you can trust that they have at least gone through an initial vetting process. Additionally, because of the breadth of materials GS indexes and its semantic search functionality, you are likely to be overwhelmed by tens of thousands of pages of results for your search rather than the more targeted and precise list of results returned through library-provided search tools. GS is focused on more results, and when it comes to finding the highest quality academic literature, more is not always better. When it comes to identifying high-quality research papers to support your literature review, GS may not be the most efficient search tool.
Finally, if part of your goal in searching the literature is to reduce bias in the results returned to you, as in comprehensive searching for systematic reviews, scoping reviews, and other advanced review research methods, GS poses some significant problems. By its nature, GS is designed to personalize results for you based on factors like those listed above (location, previous search history, etc.). Logging out and running GS in a private browser can partially mitigate this. However, the thing to note is that the algorithm takes control out of your hands to a large extent. By their nature, results from GS will come with some amount of bias programmed in. As with any tool, it is helpful to know how it works so that you can critically assess whether the tool meets your particular needs for your current project.
This post is part 3 in an ongoing series on searching the academic literature. In the next post, we will focus on search algorithms, mentioned in this post, and generative AI.
[1] To learn more about the history and functionality of search engines, see Tuhin, M. (2025, April 9). What is a Search Engine? Everything you Need to Know. Science News Today. Accessed April 28, 2026 from https://www.sciencenewstoday.org/what-is-a-search-engine-everything-you-need-to-know.
[2] What are Search Engines? (2025, July 23). Geeks for Geeks.Accessed April 27, 2026 from https://www.geeksforgeeks.org/computer-science-fundamentals/what-are-search-engines-and-how-do-they-work/. This page has a lot of great information on how search engines work.
[3] Search Engine vs Web Browser. (2025, Dec. 5). Geeks for Geeks. Accessed on July 20, 2026 from https://www.geeksforgeeks.org/techtips/difference-between-search-engine-and-web-browser/.
[4] Inclusion Guidelines for Webmasters. (n.d.) Google Scholar. Accessed April 27, 2026 from https://scholar.google.com/intl/en/scholar/inclusion.html#content. Note that this source is written for webmasters to understand how to make their content crawlable by Google Scholar’s bots. This implies that only properly structured data will be captured and makes it possible for webmasters to “game” the system with non-credible, non-scholarly work that is structured to look like scholarly work.
[5] What is semantic search? (2026, Jan. 14). Google Cloud. Accessed on April 27, 2026 from https://cloud.google.com/discover/what-is-semantic-search. See also What is Semantic Search? (2025, July 23). Geeks for Geeks. Accessed on April 27, 2026 from https://www.geeksforgeeks.org/nlp/what-is-semantic-search/.
[6] See How Does Google Scholar Work? (2025, Jan. 17). California Learning Resource Network. Accessed on April 28, 2026 from https://www.clrn.org/how-does-google-scholar-work/.
