In the last few years, web scraping has been one of my day to day and frequently needed tasks. I was wondering if I can make it smart and automatic to save lots of time. So I made AutoScraper! The project code is available on . Github This project is made for automatic web scraping to make scraping easy. It gets a url or the html content of a web page and a list of sample data which we want to scrape from that page. . It learns the scraping rules and returns the similar elements. Then you can use this learned object with new urls to get similar content or the exact same element of those new pages! This data can be text, url or any html tag value of that page Installation Install latest version from git repository using pip: $ pip install git+https: thub.com autoscraper.git //gi /alirezamika/ How to use Getting similar results Say we want to fetch all related post titles in a stackoverflow page: autoscraper AutoScraper

url = wanted_list = [ ]

scraper = AutoScraper()
result = scraper.build(url, wanted_list)
print(result) from import 'https://stackoverflow.com/questions/2081586/web-scraping-with-python' # We can add one or multiple candidates here. # You can also put urls here to retrieve urls. "How to call an external command?" Here's the output: [ , , , , , , , , ] 'How do I merge two dictionaries in a single expression in Python (taking union of dictionaries)?' 'How to call an external command?' 'What are metaclasses in Python?' 'Does Python have a ternary conditional operator?' 'How do you remove duplicates from a list whilst preserving order?' 'Convert bytes to a string' 'How to get line count of a large file cheaply in Python?' "Does Python have a string 'contains' substring method?" 'Why is “1000000000000000 in range(1000000000000001)” so fast in Python 3?' Now you can use the object to get related topics of any stackoverflow page: scraper scraper.get_result_similar( ) 'https://stackoverflow.com/questions/606191/convert-bytes-to-a-string' We can also append a topic from linked topics section to the wanted list to get all related & linked topics! Or if we want to get its urls, we can add one of the urls too. Getting exact result Say we want to scrape live stock prices from yahoo finance: autoscraper AutoScraper

url = wanted_list = [ ]

scraper = AutoScraper() result = scraper.build(url, wanted_list)
print(result) from import 'https://finance.yahoo.com/quote/AAPL/' "124.81" # Here we can also pass html content via the html parameter instead of the url (html=html_content) You can also pass any custom module parameter. for example you may want to use proxies or custom headers: requests proxies = { : , : ,
}

result = scraper.build(url, wanted_list, request_args=dict(proxies=proxies)) "http" 'http://127.0.0.1:8001' "https" 'https://127.0.0.1:8001' Now we can get the price of any symbol: scraper.get_result_exact( ) 'https://finance.yahoo.com/quote/MSFT/' You may want to get other info as well. For example if you want to get market cap too, you can just append it to the wanted list. By using the method, it will retrieve the data as the same exact order in the wanted list. get_result_exact Saving the model We can now save the built model to use it later. To save: # Give it a file path
scraper.save( ) 'yahoo-finance' And to load: scraper.load( ) 'yahoo-finance' Generating the scraper python code We can also generate a stand-alone code for the learned scraper to use it anywhere: code = scraper.generate_python_code()
print(code) It will print the generated code. There's a class named which has the methods and GeneratedAutoScraper get_result_similar which you can use. You can also use method to get both. get_result_exact get_result Thanks I hope this project is useful for you and save your time, too. I look forward to hear any feedback or suggestion.

AutoScraper Introduction: Fast and Light Automatic Web Scraper for Python

About Author

Comments

TOPICS

THIS ARTICLE WAS FEATURED IN

Related Stories

Untitled Story

AutoScraper and Flask: Create an API From Any Website in Less Than 5 Minutes

116 Stories To Learn About Web Scraping

3 Mejores Formas de Crawl Datos desde Website

5 Técnicas Anti-Scraping que Puedes Encontrar

53 Stories To Learn About Data Scraping

AutoScraper and Flask: Create an API From Any Website in Less Than 5 Minutes

116 Stories To Learn About Web Scraping

3 Mejores Formas de Crawl Datos desde Website

5 Técnicas Anti-Scraping que Puedes Encontrar

53 Stories To Learn About Data Scraping

Light-Mode

Classic

Newspaper

Dark-Mode

Neon Noir

Minty

HN StartUps