Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

45 results about "Web scraping" patented technology

Web scraping, web harvesting, or web data extraction is data scraping used for extracting data from websites. Web scraping software may access the World Wide Web directly using the Hypertext Transfer Protocol, or through a web browser. While web scraping can be done manually by a software user, the term typically refers to automated processes implemented using a bot or web crawler. It is a form of copying, in which specific data is gathered and copied from the web, typically into a central local database or spreadsheet, for later retrieval or analysis.

Webpage information processing method and device based on intelligent agent, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an agent-based webpage information processing method, device and equipment and a medium. And distributing the data to a capturing agent to obtain webpage original data. According to the method, model context information is generated by extracting webpage source information and release time information, page structure representation is constructed, key information nodes are positioned, structured data are generated, then multi-scene decision data are generated in combination with the structured data and the model context information in the analysis process, and a decision result is output through a decision reasoning module. By constructing a uniform data processing flow, webpage capture, context generation, structured analysis, analysis processing and decision reasoning are organically linked, efficient acquisition and intelligent processing of webpage information are realized, and timeliness and accuracy of decision are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Advanced Cybersecurity System for Real-Time Phishing Detection, Account Takeover Fraud Prevention, and Software Repository Optimization Using Machine Learning Techniques

Systems and processes are disclosed for enhancing cybersecurity and optimizing software repositories through integration of web crawling, web scraping, feature engineering, and advanced machine learning algorithms to detect phishing attempts, prevent account takeover fraud, and identify unused code in repositories. The system collects and refines data from various sources, including transaction logs, customer databases, device details, external data sources, and historical fraud data, to build comprehensive datasets. Feature engineering creates new, meaningful features from the refined data, which are used to train and evaluate machine learning models. The best-performing models are deployed in production to monitor incoming communications and transactions in real-time, flagging suspicious activities and optimizing codebases. This processing ensures timely detection and prevention of security threats while maintaining efficient software development processes. Robust protection is provided against evolving cyber threats and enhances software performance and security through continuous learning and adaptation.
Owner:BANK OF AMERICA CORP

AI-powered platform for recruitment, competency assessment, and onboarding with blockchain-based verification infrastructure and smart contract processing in HR

An AI-powered HR recruitment platform system that aggregates job postings in real time using API and web scraping technologies, classifies roles using NLP, and ranks candidates using trust-weighted AI models based on blockchain-verified qualifications, thereby achieving increased transparency, data protection, and verifiable fairness over traditional HR systems. The components can be used individually or in any combination without departing from the scope of the invention.
Owner:RIAZ AYSHA

AI-Powered Policy Evaluation and Ethical Compliance System

An artificial intelligence (AI)-powered system for policy evaluation, bias mitigation, and regulatory compliance. The system retrieves AI policies from multiple sources, including APIs, document parsing, and structured web scraping, ensuring real-time updates. It applies machine learning and natural language processing (NLP) to generate unbiased policy summaries and detect cybersecurity, ethics, and compliance gaps. A policy optimization engine analyzes governance trends and refines recommendations based on adoption feasibility. The system includes a sandbox testing framework that simulates regulatory and societal impacts to assess policy effectiveness. This AI-driven approach enhances fairness, transparency, and cybersecurity resilience in AI governance, enabling policymakers to proactively align with global regulatory frameworks and mitigate policy risks.
Owner:CYBER INSTITUTE

Ai-driven real-time publication process

A system for automated real-time publication processing may comprise a data collection module configured to automatically gather data from multiple external sources. A data analysis module may be configured to process the gathered data using artificial intelligence models. A data visualization module may be configured to generate interactive visual representations of the processed data. An update control module may be configured to automatically update published content and maintain version history with timestamps. The data collection module may utilize application programming interfaces and web scraping tools to gather data from government databases and real-time data feeds. The data analysis module may employ machine learning libraries to process and analyze the collected data. The data visualization module may use visualization tools to create interactive charts and graphs. The update control module may use Git-based version management and automated scheduling scripts to implement updates and record modification dates.
Owner:MCDONALD OLIVIA MARSHA

Vertical domain intelligence analysis method based on large language model

The invention discloses a vertical domain intelligence analysis method based on a large language model. The method comprises the following steps: receiving a natural language analysis demand of a user; based on pre-constructed vertical domain specific task planning knowledge, intelligently decomposing the demand into a subtask list corresponding to a preset dimension by using a large language model, and semantically matching a proper information acquisition tool for each subtask to generate a structured task plan; the system schedules and executes tool sets of network search, webpage capture, internal query and the like according to a plan, and performs cross verification and cleaning on collected fragmented information; and inputting the verified data into a large language model for integration analysis, and optionally carrying out multi-round iteration self-evaluation optimization to automatically generate a high-quality vertical domain intelligence analysis report. According to the method, full-process and high-intelligence automatic processing of information analysis in the vertical field from demand understanding to report generation is realized, and the method has excellent planning capability, process interpretability, result reliability and excellent modular expansibility.
Owner:FOCUS TECH

Webpage data extraction method based on large language model

The invention discloses a webpage data extraction method based on a large language model, which comprises the following steps of: generating an Xpath sequence webpage grabber by utilizing the large language model, and processing diversified and variable network environments through a two-stage framework: in the first stage, inversely checking and removing HTML (Hypertext Markup Language) noise by utilizing LLM (Language Language Model) information extraction capability, and in the second stage, inversely checking and removing HTML (Hypertext Markup Language Model) noise; self-adaptive generation of an Xpath action sequence is carried out according to the hierarchical structure of the HTML; according to the combination of an external evaluation mechanism and a local evaluation mechanism of LLMs, a plurality of Xpath action sequences generated on different webpages in one stage are integrated, and a general grabber specific to a website is generated. According to the method, a baseline method is always exceeded under zero sample setting, higher efficiency is shown in large-scale webpage information extraction tasks, the method can quickly adapt to different website and task requirements, dependence on LLMs is reduced when similar tasks are processed, and therefore the efficiency of processing a large number of webpage tasks is improved.
Owner:HANGZHOU DIANZI UNIV

Automated method for virtual technical assistance in the correction of computer vulnerabilities through combined usage of software automation and artificial intelligence technics

The invention relates to an automatic method for technical assistance in the correction of vulnerabilities of a computer system, where the aforementioned method comprises at least the following steps: a) receiving information relative to computer system vulnerabilities by text, voice, file input or through data exchange; b) launching tools based on artificial intelligence algorithms and machine learning to understand the information and requests entered relative to vulnerabilities; c) performing parsing and text mining of the information received and process it through intelligent data matching with information relative to the computer vulnerabilities present in a database or through interaction with external online vulnerability databases; d) in case of unsatisfactory results, starting an extended online search by web scraping using deep learning algorithms and unsupervised machine learning; e) updating an internal vulnerability database with the additional information found; and f) generating a vulnerability report accompanied by relative remediation.
Owner:CYLOCK SRL

Network anti-crawling method, system and computer device

This application relates to a method, system, and computer device for preventing web scraping. The method includes: when a client sends an access request to a target website, determining whether the target website has configured anti-scraping measures based on preset anti-scraping information; if the target website has configured anti-scraping measures, injecting an SDK into the access response, generating a first access response, and returning the first access response to the client; the access response is generated by the target website based on the access request; the client asynchronously submits its runtime information to a human-machine interface verification unit based on the first access response; when the client sends an access request to a subpage of the target website, determining whether the interface corresponding to the subpage is a protected interface; if the interface corresponding to the subpage is a protected interface, the human-machine interface verification unit performs human-machine interface verification based on the client's runtime information; if the human-machine interface verification result is successful, an access request is sent to the subpage. This method can prevent the target website from being accessed by malicious programs.
Owner:ZHONGAN INFORMATION TECH SERVICES CO LTD

Automated methods for generating labeled benchmark data set of geological thin-section images for machine learning and geospatial analysis

A method and a system for generating a labeled benchmark dataset are disclosed. The method includes obtaining a plurality of sources related to a thin section using web scraping and extracting a plurality of images from the plurality of sources related to the thin section, the plurality of images including a plurality of thin section images and a plurality of non-thin section images. Further, the method includes determining the plurality of thin section images from the plurality of extracted images and generating a classification of the plurality of thin section images based on a given classification criteria. The geological thin-section based machine learning models is trained based on the generated classification of the plurality of thin section images and a wellbore drilling plan is generated based on the geological thin-section based machine learning models.
Owner:ARAMCO SERVICES CO +1

Method and system for identification of product taxonomy from product abbreviation

The present invention generally relates to the field of taxonomy identification. Identifying product taxonomy from a product abbreviation is currently performed manually and consumes lot of time and effort. Hence, embodiments of present disclosure provide an automated method for identification of product taxonomy from product abbreviation. First, a brand name of product abbreviation is predicted using a Large Language Model (LLM) and a brand list. Then, possible expansions of the abbreviation are generated based on the brand name using acronym expansion dictionary. Among the generated possible expansions, a relevant one is identified using the LLM. Later, web scraping and web search techniques are applied on the relevant expansion to obtain associated top k matches of a supergroup, a product group and module of the predicted relevant expansion. Finally, product taxonomy is predicted based on the top k matches using a LLM augmented taxonomy classification by Retrieval Augmented Generation.
Owner:TATA CONSULTANCY SERVICES LTD

System and method for web scraping and countermeasure solver

A web scraping system configured with web scraping countermeasure resolution technology. The system comprises a configuration manager, a browser stack configured as a web browser client, a custom solver comprising a web scraping countermeasure solver; an application programming interface (API) gateway server operatively connected with the configuration manager and is configured to obtain a browser stack configuration and a session strategy from the configuration manager, and a session analysis server comprising a response analyzer configured to process a response from the target website to the target webpage request to solve a web scraping countermeasure challenge from the target website and provide an antibot solution to the custom solver.
Owner:ZYTE GRP LTD

Price comparison method based on webpage capture and text similarity

The invention discloses a price comparison method based on webpage capture and text similarity. The method comprises the following steps: step 1, constructing a Scrioy framework tool to obtain information of each website, and constructing a database; 2, carrying out vectorization processing on the large-scale text data, and converting the text data into a digital form capable of being calculated and analyzed; step 3, constructing a cosine similarity algorithm; and 4, designing a user interaction interface. According to the invention, through the technical combination of Scrapy dynamic crawling, TF-ID vectorization, cosine similarity matching and interactive UI, three major pain points of data lag, inaccurate matching and result redundancy of a traditional price comparison tool are solved, and localization optimization is carried out for the Chinese market.
Owner:HANGZHOU DIANZI UNIV

Data processing tool for social media follower scrubbing

A computer-implemented data processing method of validating legitimacy of a plurality of social media followers of a selected social media account owner, comprising steps, carried out by a social media follower scrubber tool, of: receiving an uploaded spreadsheet from a user, the spreadsheet including results of a web scraping operation, where a web scraping tool has been used to scrape data regarding the plurality of social media followers of the social media account owner selected by the user, where the spreadsheet has a plurality of rows, with each row representing one of the plurality of followers and a plurality of columns, with each column representing a characteristic feature related to the plurality of followers; and presenting the user with a drag and drop graphical user interface functionality allowing the user to rearrange and rename the columns of the uploaded spreadsheet in accordance with a native spreadsheet format.
Owner:APRACITA INVENTIONS CO INC

Generating a path to a document element using machine learning

Disclosed herein are system, method, and computer program product embodiments for improving web scraping technology by using machine learning to generate parsing expressions. A system receives a request to identify an element in a first document at a target web page. The system downloads and modifies the first document by adding an index value as an attribute to a tag for the element. A query is submitted to a large language model (LLM), including the modified first document, a description of the element, and a request asking the LLM to identify the element based on the description. The system obtains, from the LLM, the index value assigned to the element. The system generates an expression defining a path to the element in the first document using the index returned by the large language model. The system downloads a second document, and parses data of a second element using the expression.
Owner:OXYLABS UAB

An artificial intelligence-based threat and attack

This invention introduces a server-based artificial intelligence-supported security and evacuation system designed to protect oil wells against threats such as warfare, terrorism, and chemical attacks The system comprises four main functions: detection of a warfare or terrorism atmosphere, rapid shutdown of the oil well, personnel evacuation, and detection of chemical weapon attacks. Within the server system, various technologies are utilised, including web scraping, Support Vector Machines (SVM), Convolutional Neural Networks (CNN), YOLO object detection, Text-to-Speech (TTS), Natural Language Processing (NLP) algorithms, as well as physical components such as Metal-Oxide Sensors (MOS) and Electrochemical Sensors. This integration ensures early threat detection, prompt response, and minimisation of casualties.
Owner:SONMEZ SELAHATTIN

Systems and methods for web scraping

A method includes receiving a request from a client web browser, generating a first instruction responsive to the request, modifying header information in the first instruction to produce a first header-modified instruction, forwarding the first header-modified instruction to a first target website, receiving a first response from the first target website, modifying, using a man in the middle (MITM) proxy, the first response to produce a modified first response, and sending a result to the client web browser responsive to the modified first response.
Owner:TOPMARQ INC

Proxy traffic optimization by caching media resources

Provided herein are system, apparatus, device, method and / or computer program product embodiments, and / or combinations and sub-combinations thereof, for caching media resources during a scraping operation. Web resources needed by webpage are stored in a cache that is used by multiple browsers that are scraping the webpage. When an unexpired entry for the web resource is present in the cache, a browser retrieves the web resource and cache instead of making a request from the webpage. This offers a technological improvement of reducing the traffic burden on proxy servers needed to forward the scraping requests and responses.
Owner:OXYLABS UAB

Webpage data scraping graphical user interface for electronic devices

1. The name of the design product: webpage data scraping graphical user interface for electronic devices. 2. The use of the design product: an electronic device. 3. The design points of the design product: in the graphical user interface. 4. The picture or photo that best indicates the design points: front view. 5. The use of the graphical user interface: for accessing applications that allow scraping or capturing webpage data, such as accessing webpage scraping APIs or using artificial intelligence assistants. 6. The human-computer interaction mode of the graphical user interface: it can be interacted through touching or clicking the graphical user interface shown on the display screen of the electronic device or operating the buttons of the electronic device.
Owner:CLARITEX PTE LTD

Web scraping through use of proxies, and applications thereof

PendingHK40134646AMechanical engineeringWeb scraping
The invention relates to a computer-implemented method for processing web scraping jobs, using a plurality of database servers (404A-404N) operating independently of one another and each being configured to manage data storage to at least a portion of a job database (314) that stores status of web scraping jobs while the web scraping jobs are being executed, the method comprising: - receiving a web scraping request from a client computing device (102); - when the web scraping request is received, selecting one of the plurality of database servers (404A-404N) that is identified as enabled in a table (1008); and sending a job description specified by the web scraping request to the selected database server (404A-404N) for storage in the job database (314) as a pending web scraping job; - repeatedly checking health of each of the plurality of database servers (404A-404N); and - based on the health checks, determine whether each of the plurality of database servers (404A-404N) are to be enabled or disabled in the table (1008).
Owner:OXYLABS UAB

Automated marketplace monitoring and alert system

A system is provided that alerts users to newly surfaced third-party digital listings matching predefined criteria, wherein said criteria are executed continuously or on-demand through distributed compute nodes. The system receives input for searching at least one online marketplace, the input received from a user device and providing description of a desired item of merchandise and criteria associated with the item. The system also launches web scraping of select online marketplaces in search of the item and continuously monitors the select online marketplaces using Python-based headless scrapers. The system also parses and normalizes data gathered during the web scraping with listings processed as they appear in near real time, identifies items of the data matching the received input, and formats and pushes alerts to the user device describing the identified items. The criteria initially comprise keywords, minimum and maximum prices, and selected online marketplaces to be searched.
Owner:DOVE JOHN

Automated device for cross-border e-commerce and method for controlling same

An automated device for cross-border e-commerce and a method for controlling same are disclosed. The automated device according to one embodiment of the present invention comprises: a database that stores data; and a processor that controls the automated device, wherein the processor may, when a preset word representing a popular search term is identified while scraping a web page, extract keywords searched at a preset frequency or higher and store the keywords in the database, search a shopping platform using the extracted and stored keywords, store a shopping mall URL address searched on the shopping platform in the database, and store product data of the shopping mall in the database using the stored shopping mall URL address.
Owner:GOPHERSOFT INC

Data processing tool for social media follower scrubbing

A computer-implemented data processing method of validating legitimacy of a plurality of social media followers of a selected social media account owner, comprising steps, carried out by a social media follower scrubber tool, of: receiving an uploaded spreadsheet from a user, the spreadsheet including results of a web scraping operation, where a web scraping tool has been used to scrape data regarding the plurality of social media followers of the social media account owner selected by the user, where the spreadsheet has a plurality of rows, with each row representing one of the plurality of followers and a plurality of columns, with each column representing a characteristic feature related to the plurality of followers; and presenting the user with a drag and drop graphical user interface functionality allowing the user to rearrange and rename the columns of the uploaded spreadsheet in accordance with a native spreadsheet format.
Owner:GILLESPIE RICHARD +1

Scrape time calculator

Disclosed herein are system, method, and computer program product embodiments for improving web scraping technology by dynamically updating scraping parameters. A scrape system may retrieve a webpage addressed at a target URL. The scrape system may compile an object list from the webpage. The scrape system may determine a number of objects in the object list. Based on the determined number of objects, the scrape system may determine a next time to retrieve the webpage addressed at the target URL such that, when the determined number of objects is greater, the next time is sooner. When the determined next time occurs, the scrape system may re-retrieve the webpage addressed at the target URL.
Owner:OXYLABS UAB

Big data enterprise environment public service system

The embodiment of the present application relates to a kind of big data enterprise environment public service system, comprising: data acquisition module, for by web page capture algorithm and multi-source API calling mechanism, non-structured data is handled, obtains enterprise environment information;Data processing module, by real-time monitoring the change of data source and user query record, dynamically adjusts the task priority and collection frequency of data acquisition;User interface module is used to receive the request information from multiple users and return corresponding result according to enterprise environment information;Request information includes through Web browser, mobile application, desktop client, API call request;Website function module is used to receive the request information sent by user interface module, and according to Web browser, mobile application, desktop client, API call request and enterprise environment information certificate query, enterprise evaluation and integral exchange.
Owner:CHINA ENVIRONMENTAL UNITED CERTIFICATION CENT CO LTD

Method and system for estimating duration and performance of a product over lifecycle of the same

Application, in a single software instrument, of AI and advanced statistics methods for monitoring and analyzing quality problems over the whole lifecycle of a product is provided. A method based on web scraping techniques and AI for monitoring the product out of the warranty period is also provided.
Owner:FRENI BREMBO SPA

Intelligent Technical Web-Based Approach Leveraging Web Scrapper and Random Forest Algorithm to Detect Phishing Emails and SMS

Systems and processes are disclosed for detecting phishing emails and text messages. The method involves accessing the internet to gather data from various online sources, executing multi-threaded downloaders to handle multiple data streams, and storing the downloaded data in a repository. A web scraping agent analyzes and extracts relevant features from the stored data, transforming unstructured data into a structured data model. Both are stored in a database. An after-processing dataset is generated, including testing and training datasets for machine learning analysis. Random Forest models are evaluated to determine accuracy in predicting phishing attempts, and optimal models are selected, which generate phishing predictions from new data, with feature extraction identifying attributes relevant for detection. An evaluation model assesses feature extraction accuracy and overall system performance. The machine learning algorithm adapts to new phishing techniques. The trained model is integrated into a security infrastructure, with real-time processing and continuous loop feedback.
Owner:BANK OF AMERICA CORP

Quick information portal

In one aspect, a computer implemented method for providing rapid access to product information on a mobile computing device is disclosed. The method includes provisioning an online platform and associating a database repository. Next, the method generates a visual code and identifies the visual code to a domain name. Then the visual code is populated with information, such as product information, wherein the product information is associated to the visual code that is further linked to the data repository and online platform. The association allows for multiple file types and acquisition from web scraping. Further, the online platform allows for editing content behind the domains and generating further subdomains that further link to the original visual code. Thereby providing, in one aspect, a quick information portal through the use of visual codes and data structuring.
Owner:IMAGINE ONE SOLUTIONS LLC

Dynamic optimization of request parameters for proxy servers

PendingCN122372546AEngineeringWeb crawler
The present disclosure relates to dynamic optimization of request parameters for a proxy server. Systems and methods of task fulfillment are extended as provided herein and target the web scraping process through the step of a client submitting a request to a web crawler. The systems and methods allow for more complex requests to be defined for the web crawler in order to receive more specific data. In one aspect, a method for extracting and collecting data from a network by a service provider infrastructure includes the steps of inspecting parameters of a request received from a user's device, adjusting the request parameters according to pre-established scraping logic, selecting a proxy according to criteria of the pre-established scraping logic, sending the adjusted request to a target through the selected proxy, inspecting metadata received from the target, and forwarding the data to the user's device.
Owner:OKOSILA BOSE PTE LTD

Mass generation of content for educational applications

A system, apparatus, method, and instruction for generating content for an educational application, comprising: periodically web scraping a repository of learning content; saving the learning content and corresponding information; processing the learning content using a neural network architecture; summarizing the learning content using artificial intelligence or machine learning; and generating one or more questions and answers for an educational application based on the learning content using artificial intelligence or machine learning.
Owner:ACAPEDIA LLC