Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

39 results about "Web scraping" patented technology

Web scraping, web harvesting, or web data extraction is data scraping used for extracting data from websites. Web scraping software may access the World Wide Web directly using the Hypertext Transfer Protocol, or through a web browser. While web scraping can be done manually by a software user, the term typically refers to automated processes implemented using a bot or web crawler. It is a form of copying, in which specific data is gathered and copied from the web, typically into a central local database or spreadsheet, for later retrieval or analysis.

Advanced Cybersecurity System for Real-Time Phishing Detection, Account Takeover Fraud Prevention, and Software Repository Optimization Using Machine Learning Techniques

Systems and processes are disclosed for enhancing cybersecurity and optimizing software repositories through integration of web crawling, web scraping, feature engineering, and advanced machine learning algorithms to detect phishing attempts, prevent account takeover fraud, and identify unused code in repositories. The system collects and refines data from various sources, including transaction logs, customer databases, device details, external data sources, and historical fraud data, to build comprehensive datasets. Feature engineering creates new, meaningful features from the refined data, which are used to train and evaluate machine learning models. The best-performing models are deployed in production to monitor incoming communications and transactions in real-time, flagging suspicious activities and optimizing codebases. This processing ensures timely detection and prevention of security threats while maintaining efficient software development processes. Robust protection is provided against evolving cyber threats and enhances software performance and security through continuous learning and adaptation.
Owner:BANK OF AMERICA CORP

Ai-driven real-time publication process

A system for automated real-time publication processing may comprise a data collection module configured to automatically gather data from multiple external sources. A data analysis module may be configured to process the gathered data using artificial intelligence models. A data visualization module may be configured to generate interactive visual representations of the processed data. An update control module may be configured to automatically update published content and maintain version history with timestamps. The data collection module may utilize application programming interfaces and web scraping tools to gather data from government databases and real-time data feeds. The data analysis module may employ machine learning libraries to process and analyze the collected data. The data visualization module may use visualization tools to create interactive charts and graphs. The update control module may use Git-based version management and automated scheduling scripts to implement updates and record modification dates.
Owner:MCDONALD OLIVIA MARSHA

Vertical domain intelligence analysis method based on large language model

The invention discloses a vertical domain intelligence analysis method based on a large language model. The method comprises the following steps: receiving a natural language analysis demand of a user; based on pre-constructed vertical domain specific task planning knowledge, intelligently decomposing the demand into a subtask list corresponding to a preset dimension by using a large language model, and semantically matching a proper information acquisition tool for each subtask to generate a structured task plan; the system schedules and executes tool sets of network search, webpage capture, internal query and the like according to a plan, and performs cross verification and cleaning on collected fragmented information; and inputting the verified data into a large language model for integration analysis, and optionally carrying out multi-round iteration self-evaluation optimization to automatically generate a high-quality vertical domain intelligence analysis report. According to the method, full-process and high-intelligence automatic processing of information analysis in the vertical field from demand understanding to report generation is realized, and the method has excellent planning capability, process interpretability, result reliability and excellent modular expansibility.
Owner:FOCUS TECH

Webpage data extraction method based on large language model

The invention discloses a webpage data extraction method based on a large language model, which comprises the following steps of: generating an Xpath sequence webpage grabber by utilizing the large language model, and processing diversified and variable network environments through a two-stage framework: in the first stage, inversely checking and removing HTML (Hypertext Markup Language) noise by utilizing LLM (Language Language Model) information extraction capability, and in the second stage, inversely checking and removing HTML (Hypertext Markup Language Model) noise; self-adaptive generation of an Xpath action sequence is carried out according to the hierarchical structure of the HTML; according to the combination of an external evaluation mechanism and a local evaluation mechanism of LLMs, a plurality of Xpath action sequences generated on different webpages in one stage are integrated, and a general grabber specific to a website is generated. According to the method, a baseline method is always exceeded under zero sample setting, higher efficiency is shown in large-scale webpage information extraction tasks, the method can quickly adapt to different website and task requirements, dependence on LLMs is reduced when similar tasks are processed, and therefore the efficiency of processing a large number of webpage tasks is improved.
Owner:HANGZHOU DIANZI UNIV

Automated method for virtual technical assistance in the correction of computer vulnerabilities through combined usage of software automation and artificial intelligence technics

The invention relates to an automatic method for technical assistance in the correction of vulnerabilities of a computer system, where the aforementioned method comprises at least the following steps: a) receiving information relative to computer system vulnerabilities by text, voice, file input or through data exchange; b) launching tools based on artificial intelligence algorithms and machine learning to understand the information and requests entered relative to vulnerabilities; c) performing parsing and text mining of the information received and process it through intelligent data matching with information relative to the computer vulnerabilities present in a database or through interaction with external online vulnerability databases; d) in case of unsatisfactory results, starting an extended online search by web scraping using deep learning algorithms and unsupervised machine learning; e) updating an internal vulnerability database with the additional information found; and f) generating a vulnerability report accompanied by relative remediation.
Owner:CYLOCK SRL

Network anti-crawling method, system and computer device

This application relates to a method, system, and computer device for preventing web scraping. The method includes: when a client sends an access request to a target website, determining whether the target website has configured anti-scraping measures based on preset anti-scraping information; if the target website has configured anti-scraping measures, injecting an SDK into the access response, generating a first access response, and returning the first access response to the client; the access response is generated by the target website based on the access request; the client asynchronously submits its runtime information to a human-machine interface verification unit based on the first access response; when the client sends an access request to a subpage of the target website, determining whether the interface corresponding to the subpage is a protected interface; if the interface corresponding to the subpage is a protected interface, the human-machine interface verification unit performs human-machine interface verification based on the client's runtime information; if the human-machine interface verification result is successful, an access request is sent to the subpage. This method can prevent the target website from being accessed by malicious programs.
Owner:ZHONGAN INFORMATION TECH SERVICES CO LTD

Automated methods for generating labeled benchmark data set of geological thin-section images for machine learning and geospatial analysis

A method and a system for generating a labeled benchmark dataset are disclosed. The method includes obtaining a plurality of sources related to a thin section using web scraping and extracting a plurality of images from the plurality of sources related to the thin section, the plurality of images including a plurality of thin section images and a plurality of non-thin section images. Further, the method includes determining the plurality of thin section images from the plurality of extracted images and generating a classification of the plurality of thin section images based on a given classification criteria. The geological thin-section based machine learning models is trained based on the generated classification of the plurality of thin section images and a wellbore drilling plan is generated based on the geological thin-section based machine learning models.
Owner:ARAMCO SERVICES CO +1

Method and system for identification of product taxonomy from product abbreviation

The present invention generally relates to the field of taxonomy identification. Identifying product taxonomy from a product abbreviation is currently performed manually and consumes lot of time and effort. Hence, embodiments of present disclosure provide an automated method for identification of product taxonomy from product abbreviation. First, a brand name of product abbreviation is predicted using a Large Language Model (LLM) and a brand list. Then, possible expansions of the abbreviation are generated based on the brand name using acronym expansion dictionary. Among the generated possible expansions, a relevant one is identified using the LLM. Later, web scraping and web search techniques are applied on the relevant expansion to obtain associated top k matches of a supergroup, a product group and module of the predicted relevant expansion. Finally, product taxonomy is predicted based on the top k matches using a LLM augmented taxonomy classification by Retrieval Augmented Generation.
Owner:TATA CONSULTANCY SERVICES LTD

Price comparison method based on webpage capture and text similarity

The invention discloses a price comparison method based on webpage capture and text similarity. The method comprises the following steps: step 1, constructing a Scrioy framework tool to obtain information of each website, and constructing a database; 2, carrying out vectorization processing on the large-scale text data, and converting the text data into a digital form capable of being calculated and analyzed; step 3, constructing a cosine similarity algorithm; and 4, designing a user interaction interface. According to the invention, through the technical combination of Scrapy dynamic crawling, TF-ID vectorization, cosine similarity matching and interactive UI, three major pain points of data lag, inaccurate matching and result redundancy of a traditional price comparison tool are solved, and localization optimization is carried out for the Chinese market.
Owner:HANGZHOU DIANZI UNIV

Data processing tool for social media follower scrubbing

A computer-implemented data processing method of validating legitimacy of a plurality of social media followers of a selected social media account owner, comprising steps, carried out by a social media follower scrubber tool, of: receiving an uploaded spreadsheet from a user, the spreadsheet including results of a web scraping operation, where a web scraping tool has been used to scrape data regarding the plurality of social media followers of the social media account owner selected by the user, where the spreadsheet has a plurality of rows, with each row representing one of the plurality of followers and a plurality of columns, with each column representing a characteristic feature related to the plurality of followers; and presenting the user with a drag and drop graphical user interface functionality allowing the user to rearrange and rename the columns of the uploaded spreadsheet in accordance with a native spreadsheet format.
Owner:APRACITA INVENTIONS CO INC

Generating a path to a document element using machine learning

Disclosed herein are system, method, and computer program product embodiments for improving web scraping technology by using machine learning to generate parsing expressions. A system receives a request to identify an element in a first document at a target web page. The system downloads and modifies the first document by adding an index value as an attribute to a tag for the element. A query is submitted to a large language model (LLM), including the modified first document, a description of the element, and a request asking the LLM to identify the element based on the description. The system obtains, from the LLM, the index value assigned to the element. The system generates an expression defining a path to the element in the first document using the index returned by the large language model. The system downloads a second document, and parses data of a second element using the expression.
Owner:OXYLABS UAB

An artificial intelligence-based threat and attack

This invention introduces a server-based artificial intelligence-supported security and evacuation system designed to protect oil wells against threats such as warfare, terrorism, and chemical attacks The system comprises four main functions: detection of a warfare or terrorism atmosphere, rapid shutdown of the oil well, personnel evacuation, and detection of chemical weapon attacks. Within the server system, various technologies are utilised, including web scraping, Support Vector Machines (SVM), Convolutional Neural Networks (CNN), YOLO object detection, Text-to-Speech (TTS), Natural Language Processing (NLP) algorithms, as well as physical components such as Metal-Oxide Sensors (MOS) and Electrochemical Sensors. This integration ensures early threat detection, prompt response, and minimisation of casualties.
Owner:SONMEZ SELAHATTIN

Systems and methods for web scraping

A method includes receiving a request from a client web browser, generating a first instruction responsive to the request, modifying header information in the first instruction to produce a first header-modified instruction, forwarding the first header-modified instruction to a first target website, receiving a first response from the first target website, modifying, using a man in the middle (MITM) proxy, the first response to produce a modified first response, and sending a result to the client web browser responsive to the modified first response.
Owner:TOPMARQ INC

Proxy traffic optimization by caching media resources

Provided herein are system, apparatus, device, method and / or computer program product embodiments, and / or combinations and sub-combinations thereof, for caching media resources during a scraping operation. Web resources needed by webpage are stored in a cache that is used by multiple browsers that are scraping the webpage. When an unexpired entry for the web resource is present in the cache, a browser retrieves the web resource and cache instead of making a request from the webpage. This offers a technological improvement of reducing the traffic burden on proxy servers needed to forward the scraping requests and responses.
Owner:OXYLABS UAB

Webpage data scraping graphical user interface for electronic devices

1. The name of the design product: webpage data scraping graphical user interface for electronic devices. 2. The use of the design product: an electronic device. 3. The design points of the design product: in the graphical user interface. 4. The picture or photo that best indicates the design points: front view. 5. The use of the graphical user interface: for accessing applications that allow scraping or capturing webpage data, such as accessing webpage scraping APIs or using artificial intelligence assistants. 6. The human-computer interaction mode of the graphical user interface: it can be interacted through touching or clicking the graphical user interface shown on the display screen of the electronic device or operating the buttons of the electronic device.
Owner:CLARITEX PTE LTD

Web scraping through use of proxies, and applications thereof

PendingHK40134646AMechanical engineeringWeb scraping
The invention relates to a computer-implemented method for processing web scraping jobs, using a plurality of database servers (404A-404N) operating independently of one another and each being configured to manage data storage to at least a portion of a job database (314) that stores status of web scraping jobs while the web scraping jobs are being executed, the method comprising: - receiving a web scraping request from a client computing device (102); - when the web scraping request is received, selecting one of the plurality of database servers (404A-404N) that is identified as enabled in a table (1008); and sending a job description specified by the web scraping request to the selected database server (404A-404N) for storage in the job database (314) as a pending web scraping job; - repeatedly checking health of each of the plurality of database servers (404A-404N); and - based on the health checks, determine whether each of the plurality of database servers (404A-404N) are to be enabled or disabled in the table (1008).
Owner:OXYLABS UAB

Automated marketplace monitoring and alert system

A system is provided that alerts users to newly surfaced third-party digital listings matching predefined criteria, wherein said criteria are executed continuously or on-demand through distributed compute nodes. The system receives input for searching at least one online marketplace, the input received from a user device and providing description of a desired item of merchandise and criteria associated with the item. The system also launches web scraping of select online marketplaces in search of the item and continuously monitors the select online marketplaces using Python-based headless scrapers. The system also parses and normalizes data gathered during the web scraping with listings processed as they appear in near real time, identifies items of the data matching the received input, and formats and pushes alerts to the user device describing the identified items. The criteria initially comprise keywords, minimum and maximum prices, and selected online marketplaces to be searched.
Owner:DOVE JOHN

Automated device for cross-border e-commerce and method for controlling same

An automated device for cross-border e-commerce and a method for controlling same are disclosed. The automated device according to one embodiment of the present invention comprises: a database that stores data; and a processor that controls the automated device, wherein the processor may, when a preset word representing a popular search term is identified while scraping a web page, extract keywords searched at a preset frequency or higher and store the keywords in the database, search a shopping platform using the extracted and stored keywords, store a shopping mall URL address searched on the shopping platform in the database, and store product data of the shopping mall in the database using the stored shopping mall URL address.
Owner:GOPHERSOFT INC

Data processing tool for social media follower scrubbing

A computer-implemented data processing method of validating legitimacy of a plurality of social media followers of a selected social media account owner, comprising steps, carried out by a social media follower scrubber tool, of: receiving an uploaded spreadsheet from a user, the spreadsheet including results of a web scraping operation, where a web scraping tool has been used to scrape data regarding the plurality of social media followers of the social media account owner selected by the user, where the spreadsheet has a plurality of rows, with each row representing one of the plurality of followers and a plurality of columns, with each column representing a characteristic feature related to the plurality of followers; and presenting the user with a drag and drop graphical user interface functionality allowing the user to rearrange and rename the columns of the uploaded spreadsheet in accordance with a native spreadsheet format.
Owner:GILLESPIE RICHARD +1

Scrape time calculator

Disclosed herein are system, method, and computer program product embodiments for improving web scraping technology by dynamically updating scraping parameters. A scrape system may retrieve a webpage addressed at a target URL. The scrape system may compile an object list from the webpage. The scrape system may determine a number of objects in the object list. Based on the determined number of objects, the scrape system may determine a next time to retrieve the webpage addressed at the target URL such that, when the determined number of objects is greater, the next time is sooner. When the determined next time occurs, the scrape system may re-retrieve the webpage addressed at the target URL.
Owner:OXYLABS UAB

Big data enterprise environment public service system

The embodiment of the present application relates to a kind of big data enterprise environment public service system, comprising: data acquisition module, for by web page capture algorithm and multi-source API calling mechanism, non-structured data is handled, obtains enterprise environment information;Data processing module, by real-time monitoring the change of data source and user query record, dynamically adjusts the task priority and collection frequency of data acquisition;User interface module is used to receive the request information from multiple users and return corresponding result according to enterprise environment information;Request information includes through Web browser, mobile application, desktop client, API call request;Website function module is used to receive the request information sent by user interface module, and according to Web browser, mobile application, desktop client, API call request and enterprise environment information certificate query, enterprise evaluation and integral exchange.
Owner:CHINA ENVIRONMENTAL UNITED CERTIFICATION CENT CO LTD

Method and system for estimating duration and performance of a product over lifecycle of the same

Application, in a single software instrument, of AI and advanced statistics methods for monitoring and analyzing quality problems over the whole lifecycle of a product is provided. A method based on web scraping techniques and AI for monitoring the product out of the warranty period is also provided.
Owner:FRENI BREMBO SPA

Intelligent Technical Web-Based Approach Leveraging Web Scrapper and Random Forest Algorithm to Detect Phishing Emails and SMS

Systems and processes are disclosed for detecting phishing emails and text messages. The method involves accessing the internet to gather data from various online sources, executing multi-threaded downloaders to handle multiple data streams, and storing the downloaded data in a repository. A web scraping agent analyzes and extracts relevant features from the stored data, transforming unstructured data into a structured data model. Both are stored in a database. An after-processing dataset is generated, including testing and training datasets for machine learning analysis. Random Forest models are evaluated to determine accuracy in predicting phishing attempts, and optimal models are selected, which generate phishing predictions from new data, with feature extraction identifying attributes relevant for detection. An evaluation model assesses feature extraction accuracy and overall system performance. The machine learning algorithm adapts to new phishing techniques. The trained model is integrated into a security infrastructure, with real-time processing and continuous loop feedback.
Owner:BANK OF AMERICA CORP

Dynamic optimization of request parameters for proxy servers

PendingCN122372546AEngineeringWeb crawler
The present disclosure relates to dynamic optimization of request parameters for a proxy server. Systems and methods of task fulfillment are extended as provided herein and target the web scraping process through the step of a client submitting a request to a web crawler. The systems and methods allow for more complex requests to be defined for the web crawler in order to receive more specific data. In one aspect, a method for extracting and collecting data from a network by a service provider infrastructure includes the steps of inspecting parameters of a request received from a user's device, adjusting the request parameters according to pre-established scraping logic, selecting a proxy according to criteria of the pre-established scraping logic, sending the adjusted request to a target through the selected proxy, inspecting metadata received from the target, and forwarding the data to the user's device.
Owner:OKOSILA BOSE PTE LTD

Mass generation of content for educational applications

A system, apparatus, method, and instruction for generating content for an educational application, comprising: periodically web scraping a repository of learning content; saving the learning content and corresponding information; processing the learning content using a neural network architecture; summarizing the learning content using artificial intelligence or machine learning; and generating one or more questions and answers for an educational application based on the learning content using artificial intelligence or machine learning.
Owner:ACAPEDIA LLC

A webpage content intelligent crawling method and system based on data analysis

PendingCN122285979ADocumentationData science
This invention discloses a method and system for intelligent web page content crawling based on data analysis, belonging to the field of web page crawling technology. The method includes identifying candidate content blocks and calculating their Shannon entropy, generating an importance score by combining text density; constructing a global vocabulary based on the candidate content blocks and calculating inverse document frequencies of terms, defining a topic-specific factor, multiplying the importance score by the topic-specific factor to obtain a comprehensive priority, selecting candidate content blocks with a comprehensive priority higher than the average as target crawling blocks, and generating crawling rules for each target crawling block; and using the generated crawling rules to extract content from the initial HTML source code. This invention improves the content differentiation of web page crawling by combining text density analysis and Shannon entropy evaluation, and significantly enhances the accuracy of the crawling rules through in-depth analysis of web page structure and visual elements.
Owner:TIANJIN HONGCHENG TECHNOLOGY CO LTD

Web anti-crawler methods, devices, equipment, storage media and products

This application provides a method, apparatus, device, storage medium, and product for preventing web scraping, relating to the field of information security. The method includes: sending a first page request to a front-end server to obtain a first-screen HTML fragment and rendering it to obtain the first-screen page content; then sending a second page request to a back-end server to obtain a non-first-screen HTML fragment based on interaction with the back-end server; rendering the non-first-screen HTML fragment to obtain non-first-screen page content; and concatenating the first-screen page content and the non-first-screen page content. In the isomorphic rendering process between the front-end and back-end, this application ensures that the front-end server only returns the first-screen HTML fragment, not the complete HTML, while the browser requests other HTML fragments from the back-end server. This means the front-end server only retains the first-screen data, effectively preventing web crawlers from scraping data from the front-end server and fundamentally protecting website data security.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Quick information portal

In one aspect, a computer implemented method for providing rapid access to product information on a mobile computing device is disclosed. The method includes provisioning an online platform and associating a database repository. Next, the method generates a visual code and identifies the visual code to a domain name. Then the visual code is populated with information, such as product information, wherein the product information is associated to the visual code that is further linked to the data repository and online platform. The association allows for multiple file types and acquisition from web scraping. Further, the online platform allows for editing content behind the domains and generating further subdomains that further link to the original visual code. Thereby providing, in one aspect, a quick information portal through the use of visual codes and data structuring.
Owner:IMAGINE ONE SOLUTIONS LLC

Dynamic optimization of request parameters for proxy servers

The system and method of task fulfillment are extended as provided herein and target the web scraping process by the step of the client submitting a request to the web crawler. The system and method allow for defining more complex requests for the web crawler in order to receive more specific data. In one aspect, a method for extracting and collecting data from a network by a service provider infrastructure comprises the steps of: checking parameters of a request received from a user's device, adjusting the request parameters according to pre-established scraping logic, selecting a proxy according to criteria of the pre-established scraping logic, sending the adjusted request to a target through the selected proxy, checking metadata received from the target, and forwarding the data to the user's device.
Owner:OKOSILA BOSE PTE LTD

Systems and methods for automated assessment of media content for sincerity

A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause a system to perform a method for analyzing media content, the method comprising: ingesting, via one or more web scraping techniques, media content from one or more online sources; preprocessing, via one or more natural language processing tools, the media content comprising the steps of: tokenizing the media content into a plurality of components, removing stop words from the plurality of components, lemmatizing the plurality of components, and normalizing the plurality of components; analyzing, via the one or more natural language processing tools, the preprocessed media content for one or more sentiments; generating, via one or more scoring algorithms, a score for each of the one or more sentiments; compiling the score for each of the one or more sentiments to generate a final score; and displaying the final score on one or more client devices.
Owner:Y2 CONSULTING LLC