Website information extraction method and electronic equipment thereof

By obtaining the target website address and the content requirements entered by the user, combining the page layout data to identify and match link data, and using the AI ​​language large model to extract information, the problem of lack of user natural language understanding in existing technologies is solved, and intelligent and efficient information extraction is achieved.

CN120780930APending Publication Date: 2025-10-14ZHONGSHAN INNOLEBI PRECISION MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510920420.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing information extraction methods lack a deep understanding of users' natural language needs and the ability to extract targeted information, and are unable to automatically and accurately match target data based on user input.

Method used

By obtaining the address information of the target website and the target content language description data entered by the user, combining the page layout data to identify matching link data, and using the preset AI language model for extraction and processing, the information of the target content is determined.

Benefits of technology

It has improved the intelligence level and processing efficiency of information extraction, and can automatically identify target links and extract page field information based on user needs, significantly improving the accuracy and versatility of information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780930A_ABST
    Figure CN120780930A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a website information extraction method, device and system. The method comprises the steps that address information of a target website and corresponding target content language description data input by a user are obtained; based on the target content language description data and page layout data of a target website, identifying and matching corresponding link data in the at least one piece of address information, and determining at least one piece of target link data; and through a preset AI language large model, performing extraction processing on the target link data, and determining information of the target content. By acquiring the target website address and the target content requirement input by the user, identifying and matching the target link data in combination with the page layout data, and extracting the target content information by using the AI language large model, the target link can be automatically identified based on the user requirement, and the page field information can be extracted through the AI language large model; and the intelligent level and the processing efficiency of website information extraction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of wisdom recognition, and in particular to a website information extraction method, device, system, electronic equipment and storage medium thereof. BACKGROUND

[0002] With the rapid growth of Internet information, users often need to extract specific target content such as customer information, business data or contact information from a large number of web pages when accessing websites. However, existing information extraction methods mostly rely on fixed web page structure templates or rule-based crawler programs, which have poor adaptability, low universality and high manual maintenance costs when dealing with different website structures, complex page layouts and dynamically loaded content.

[0003] At the same time, the prior art lacks deep understanding of user natural language requirements and targeted information extraction capabilities, and cannot achieve the function of automatically and accurately matching target data based on user input.

[0004] Therefore, the existing information extraction method lacks deep understanding of user natural language requirements and targeted information extraction capabilities, and cannot achieve the function of automatically and accurately matching target data based on user input. SUMMARY

[0005] The embodiments of the present application provide a website information extraction method to solve the problem that the existing information extraction method lacks deep understanding of user natural language requirements and targeted information extraction capabilities, and cannot achieve the function of automatically and accurately matching target data based on user input, to improve the automation, accuracy and universality of information extraction.

[0006] In a first aspect, the embodiments of the present application provide a website information extraction method, which comprises the following steps: Obtaining address information of a target website and corresponding target content language description data input by a user; Based on the target content language description data and the page layout data of the target website, identifying and matching at least one corresponding link data in the address information to determine at least one target link data; Extracting and processing the target link data through a preset AI language model to determine the information of the target content.

[0007] Optionally, the obtaining of the address information of the target website and the corresponding target content language description data input by the user comprises: Determining the URL address of the target website; Converting and processing the target content through a colloquial description to determine the corresponding target content language description data, which is used to point to the content information corresponding to the user's requirements.

[0008] Optionally, before the determining at least one target link data by recognizing and matching the corresponding link data in the at least one address information based on the target content language description data and the page layout data of the target website, the method further comprises: determining a website type based on website data features, wherein the website type comprises a number type, a letter type, a loading type, and a map type; recognizing and analyzing the page layout of the target website by the number type, the letter type, the loading type, and the map type, to determine the website type possessed by the page layout data.

[0009] Optionally, the determining at least one target link data by recognizing and matching the corresponding link data in the at least one address information based on the target content language description data and the page layout data of the target website comprises: determining a target demand content of a user based on the target content language description data; analyzing the page layout data to determine position feature data and structure feature data corresponding to each link data in the page, wherein the structure data comprises a label type, a link attribute, and a nested hierarchical structure of a link element of the link data; matching the position feature data and the structure feature data corresponding to each link data with the target demand content to determine at least one target link data.

[0010] Optionally, the matching the position feature data and the structure feature data corresponding to each link data with the target demand content to determine at least one target link data comprises: determining a function feature of the link data in the current webpage based on the structure feature data corresponding to each link data, wherein the function feature comprises a click initiation feature and an automatic jump initiation feature; obtaining a corresponding correlation value by matching the corresponding function feature with the target demand content; comparing the correlation value with a preset correlation threshold to determine at least one target link data from the plurality of link data.

[0011] Optionally, the determining target content information by extracting and processing the target link data by a preset AI language large model comprises: determining content data of a page accessed in the target link data; performing semantic analysis and recognition on the content data by a preset AI language large model to determine at least one target content information corresponding to the target demand content.

[0012] In a second aspect, the embodiments of the present application also provide a website information extraction device, which comprises: The first obtaining module is configured to obtain address information of a target website and target content language description data input by a user; The first determining module is configured to identify and match corresponding link data in at least one of the address information based on the target content language description data and page layout data of the target website, and determine at least one target link data; The first processing module is configured to perform extraction processing on the target link data by using a preset AI language large model, and determine information of target content.

[0013] In a third aspect, the present application provides a multi-modal mechanical property test management system, which comprises a multi-modal mechanical property test management device, a server and a smart expansion device.

[0014] In a fourth aspect, the embodiments of the present application provide an electronic device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the website information extraction method provided by the embodiments of the present application when executing the computer program.

[0015] In a fifth aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the website information extraction method provided by the embodiments of the present application when executed by a processor.

[0016] In the embodiments of the present application, address information of a target website and target content language description data input by a user are obtained, corresponding link data in at least one of the address information is identified and matched based on the target content language description data and page layout data of the target website, at least one target link data is determined, and extraction processing is performed on the target link data by using a preset AI language large model to determine information of target content. By obtaining target website address and target content requirement input by a user, identifying and matching target link data in combination with page layout data, and extracting target content information by using an AI language large model, target link can be automatically identified based on user requirement, page field information can be extracted by using an AI language large model, and the intelligent level and processing efficiency of website information extraction are improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 This is a system architecture diagram of a website information extraction system provided by an embodiment of the present invention; Figure 2 This is a flow chart of a website information extraction method provided by an embodiment of the present invention; Figure 3 is a structural diagram of another website information extraction device provided in an embodiment of the present invention; Figure 4 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] like Figure 1 As shown, Figure 1 This is an architectural diagram of a website information extraction system 100 provided in an embodiment of the present invention. The website information extraction system includes: a website information extraction device 300, a server 101, and a smart device 102. Specifically, a first acquisition module can be used to obtain the address information of a target website and corresponding target content language description data input by a user; a first determination module can be used to identify and match corresponding link data within at least one address information based on the target content language description data and the page layout data of the target website, and determine at least one target link data; and a first processing module can be used to extract and process the target link data using a preset AI language model to determine information about the target content.

[0021] Specifically, the above-mentioned target website may refer to a web page platform from which the user wishes to extract specific information, which is usually a collection of web pages with a clear information structure and accessible online. It can be understood that in this embodiment, the above-mentioned target website may be the source data location for the AI ​​language large model to perform information extraction tasks. Therefore, its page structure, content arrangement, etc. will directly affect the extraction effect.

[0022] The address information can refer to a uniform resource locator (URL) in the target website for accessing web resources, which is a basic entry for the website information extraction system to load web pages, parse layouts, and extract subsequent information. Generally, when the user inputs a requirement, the address information is usually provided in the form of a URL, and the website information extraction system determines the range of web pages to be crawled based on the URL.

[0023] The target content language description data can refer to the user's description of the web content they want to extract in natural language. This data serves as a target-oriented parameter for the AI model, guiding the model to identify information fields that highly match the user's requirements in the web page. It can be understood that the target content language description data can be a complete sentence, a set of keywords, or a structured prompt, which is semantically analyzed by the website information extraction system and matched with the web page structure.

[0024] For example, the user inputs "I need the contact name and phone number of each company," which is the target content language description data. At this time, the website information extraction system converts this description into semantic tags, guiding the AI large model to preferentially extract fields such as "contact," "phone," "mobile," "contact name," and filter out irrelevant information, thereby accurately extracting the content of interest to the user.

[0025] In one possible embodiment, the website information extraction system prompts the user to input the website source and specific requirements of the information to be crawled. For example, the user wants to obtain the contact information of exhibitors from a certain industry exhibition website, so the user inputs the URL address of the website in the system and inputs the requirement content in a colloquial description manner, such as "I want to know the contact name, mobile number, and email of all exhibitors in the exhibition." The website information extraction system processes the above input and parses the following data: target website, address information, target content language description data, and based on the structure type of the target website (e.g., belongs to CrawlerMainLoadMore type, needs to scroll to load all customer information pages), further analyzes the link data and content structure characteristics in the page, and submits the link of each customer information page to the AI language large model for semantic extraction and organization, finally outputs the structured data in the form of an Excel table for the user to use.

[0026] The page layout data can refer to the structural arrangement information of various content elements in the page, mainly including the nested structure of HTML tags, the region to which the element belongs, the style class name (class), the ID identifier, the tag type, and the relative position relationship, etc. data structures used to reflect the visual structure and content organization method of the web page.

[0027] The link data can refer to all elements with jump capability in the webpage, usually in the form of a button with a label or dynamic JS binding jump, representing the address to which the user will jump after clicking. It can be understood that these links can jump to a detail page, an external website, or trigger a modal window, etc. The website information extraction system can filter out the part that may contain target information from these links. For example, the following links may exist in the webpage: html Company Details About Us Expand More These links are all "link data", but not every link contains useful information. For example, "About Us" may not contain contact information. Therefore, this step is only to collect all the jump paths as a candidate set for the next step of screening.

[0028] In another possible embodiment, the website information extraction system can perform semantic matching and structural comparison between the user's natural language input requirements (i.e., target content language description data) and the relevant content of the links in the webpage, to determine which links are more likely to contain target data. This process includes multiple steps such as text keyword extraction, context semantic analysis, and position feature comparison.

[0029] For example, assuming that the user's input requirement is "get all the contact information and contact numbers of the exhibitors", the website information extraction system will scan whether the surrounding or jump page of each link contains keywords such as "contact", "phone", "mobile phone", or whether it conforms to the common field structure. If a link jump page contains Contact: Zhang San and Phone: 13911112222 , the system will determine that the link has a high matching degree, and then mark it as a target candidate.

[0030] The target link data can refer to the set of webpage links that are most likely to contain the content required by the user after page structure analysis and semantic recognition matching by the website information extraction system, and these links will be used as input for the subsequent AI language large model to extract content.

[0031] The preset AI language large model can refer to a large model with natural language understanding and deep learning. Generally, the model can be trained based on the Transformer architecture and can handle multi-language and multi-task semantic information extraction tasks. In this embodiment, the preset AI language large model can perform semantic analysis on webpage content, automatically understand page text structure and field meaning, and extract content highly related to user requirements. For example, the following content exists in a webpage fragment: Company name: XYZ Robot Co., Ltd. Contact: Li Si Phone: 13812345678 Address: No. 88, Gaokeli, Chaoyang District, Beijing Therefore, after receiving the instruction "please extract the company name, contact person and phone number", the above preset AI language model will automatically extract the company name: XYZ Robot Co., Ltd., contact person: Li Si, and phone number: 13812345678.

[0032] In another possible embodiment, the above website information extraction system uses an AI language model to analyze unstructured or semi-structured text information in the webpage and extract field data that meets the user's specified target. It can be understood that this process may include but is not limited to text preprocessing, semantic recognition, field positioning, content extraction and result structuring.

[0033] The above target content information can be the data field set extracted by the website information extraction system for the user, which has a clear semantic correspondence. These information are highly matched with the user's original needs, usually in a structured form (such as JSON, table), which is convenient for users to use and subsequent processing.

[0034] As shown in Figure 2 , the flowchart of a website information extraction method provided by an embodiment of the present application includes the following steps: Figure 1 201. Obtain the address information of the target website and the user input corresponding target content language description data.

[0035] In the present embodiment, the target website can refer to a webpage platform from which the user wants to extract specific information, which is usually a set of webpages with clear information structure and online access. It can be understood that in the present embodiment, the target website can be the source data for the AI language model to perform information extraction tasks, so its page structure, content arrangement, etc. will directly affect the extraction effect.

[0036] The above address information can refer to the uniform resource locator (URL) used to access web resources in the target website, which is the basic entrance for website information extraction system to load webpages, parse layouts and extract information subsequently. Generally, when the user inputs the demand, the address information is usually provided in the form of URL, and the website information extraction system determines the range of webpages to be captured accordingly.

[0037] ​The target content language description data can be used to indicate the content information corresponding to the user demand, and the target content language description data can be a complete sentence, a keyword set or a structured prompt.

[0038] For example, the user inputs "I need the contact name and phone number of each company", which is the target content language description data. At this time, the website information extraction system converts the description into semantic tags to guide the AI large model to preferentially extract fields such as "contact", "phone", "mobile", "contact name" and the like when analyzing the webpage content, and to filter irrelevant information, thereby accurately extracting the content concerned by the user.

[0039] In a possible embodiment, the website information extraction system prompts the user to input the website source and specific demand of the information to be captured. For example, the user wants to obtain the contact information and contact method of the exhibitors from a certain industry exhibition website. The user inputs the URL address of the website in the system and inputs the demand content in a colloquial description manner, for example, "I want to know the contact name, mobile phone number and email address of all exhibitors in the exhibition". The website information extraction system processes the above input and analyzes the following data: target website, address information, target content language description data and the like. Based on the structure type of the target website (for example, it belongs to the CrawlerMainLoadMore type and needs to load all customer information pages), the website information extraction system further analyzes the link data and content structure characteristics in the page, and submits the link of each customer information page to the AI language large model for semantic extraction and arrangement, and finally outputs the structured data in the form of an Excel table for the user to use.

[0040] 202. Based on the target content language description data and the page layout data of the target website, at least one target link data corresponding to the link data in the at least one address information is identified and matched.

[0041] In the embodiment of the present application, the page layout data can be the structural arrangement information of each content element in the page, mainly including the nested structure of HTML tags, the region to which the element belongs, the style class name (class), the ID identifier, the tag type and the relative position relationship and the like data structures used to reflect the visual structure and content organization method of the webpage.

[0042] The link data can refer to all elements with jump capability in the webpage, usually in the form of a button with a label or dynamic JS binding jump, representing the address to which the user will jump after clicking. It can be understood that these links can jump to a detail page, an external website, or trigger a modal window, etc. The website information extraction system can filter out the part that may contain target information from these links. For example, there can be the following links in the webpage: html Company Details About Us Expand More These links are all "link data", but not every link contains useful information. For example, "About Us" may not contain contact information. Therefore, this step is only to collect all jump paths as a candidate set for the next step of screening.

[0043] In another possible embodiment, the website information extraction system can perform semantic matching and structural comparison between the user's natural language input requirements (i.e., target content language description data) and the relevant content of the links in the webpage, to determine which links are more likely to contain target data. This process includes multiple steps such as text keyword extraction, context semantic analysis, and position feature comparison.

[0044] For example, assuming that the user's input requirement is "get all the contact information and contact numbers of the exhibitors", the website information extraction system will scan whether the surrounding or jump page of each link contains keywords such as "contact", "phone", "mobile phone", or whether it conforms to the common field structure. If a link jump page contains Contact: Zhang San and Phone: 13911112222 , the system will determine that the link has a high matching degree, and then mark it as a target candidate.

[0045] The target link data can refer to the set of webpage links that are most likely to contain the content required by the user after page structure analysis and semantic recognition matching by the website information extraction system, and these links will be used as input for the subsequent AI language large model to extract content.

[0046] 203. Extracting and processing the target link data by a preset AI language large model to determine the information of the target content.

[0047] In the embodiment of the application, the above-mentioned preset AI language large model can refer to a large model with natural language understanding and deep learning. Generally, the model can be trained based on the Transformer architecture and can process multi-language and multi-task semantic information extraction tasks. In the embodiment, the preset AI language large model can perform semantic analysis on web page content, automatically understand page text structure and field meaning, and extract content information highly related to user demand, for example, the following content exists in a web page segment: Company name: XYZ robot Co., Ltd. Contact: Li Si Phone: 13812345678 Address: No. 88, Gaokeli, Chaoyang District, Beijing Therefore, after receiving the instruction "please extract the company name, contact and phone", the above-mentioned preset AI language large model will automatically extract the company name: XYZ robot Co., Ltd., the contact: Li Si, and the phone: 13812345678.

[0048] In another possible embodiment, the website information extraction system analyzes the unstructured or semi-structured text information in the web page through the AI language large model, extracts the field data meeting the user-specified target, and it can be understood that the process can include but is not limited to text preprocessing, semantic recognition, field positioning, content extraction and result structuring.

[0049] The above-mentioned target content information can be the data field set extracted by the website information extraction system for the user, which has a clear semantic correspondence relationship. These information is highly matched with the original demand of the user, and is usually presented in a structured form (such as JSON, table) for easy use and subsequent processing by the user.

[0050] In the embodiment of the application, the address information of the target website and the user input corresponding target content language description data are obtained; based on the target content language description data and the page layout data of the target website, at least one target link data corresponding to the link data in the at least one address information is identified and matched to determine at least one target link data; and the target link data is processed by a preset AI language large model to determine the target content information. By obtaining the target website address and the target content requirement input by the user, combining the page layout data to identify and match the target link data, and using the AI language large model to extract the target content information, the target link can be automatically identified based on the user demand, and the page field information can be extracted through the AI language large model, which improves the intelligent level and processing efficiency of website information extraction.

[0051] Optionally, in the step of acquiring the address information of the target website and the corresponding target content language description data input by the user, the URL address of the target website can also be determined; the target content is converted and processed through colloquial description to determine the corresponding target content language description data.

[0052] In the embodiment of the present application, the above-mentioned target content can refer to the specific information type that the user wants to extract from the webpage, which is usually represented as a data field with specific meaning and purpose. It can be understood that these contents represent the actual business demand orientation of the user, and the final output result of the above-mentioned website information extraction system can correspond to the target content one by one.

[0053] In a possible embodiment, the above-mentioned website information extraction system extracts the target field or keyword set that is helpful for subsequent processing by performing semantic analysis and structural standardization on the natural language (such as daily language, keywords, sentences) input by the user, that is, it converts "unstructured language expression" into "structured description that can be recognized by a program", which is used as a reference basis for AI matching and recognition. For example, the user inputs: "I want to see which companies are engaged in robot equipment, and do they have contact phone numbers?" After the above-mentioned website information extraction system performs semantic recognition, it is converted and processed as: "Target field": ["company name", "main business", "contact person", "contact phone number"], "Filtering condition": ["main business contains robot equipment"].

[0054] Through the above method steps, the user can express the information extraction demand in a colloquial and non-technical way, and the website information extraction system can automatically convert such expression into structured target content description data, thereby significantly reducing the use threshold, avoiding the user from having to understand the webpage structure or write extraction rules, and improving the intelligent interaction and adaptability of the website information extraction system. At the same time, combined with the recognition of the URL address and the clear indication of the target field, it is helpful to improve the accuracy and efficiency of subsequent link screening and information extraction, and enhance the automation and user friendliness of the overall information crawling process.

[0055] Optionally, in the step of identifying and matching the corresponding link data in at least one address information based on the target content language description data and the page layout data of the target website, it further includes determining the website type based on the website data characteristics; the page layout of the target website is identified and analyzed through digital type, alphabetical type, loading type and map type to determine the website type possessed in the page layout data.

[0056] In the embodiments of the present application, the website data features mentioned above can refer to identifiable patterns presented by the target webpage in terms of structure and content organization, including but not limited to pagination form, list display structure, jump mode (such as link jump or JavaScript loading), modular layout, etc. It can be understood that these features provide the basis for judging the website type and selecting the parsing strategy. For example: If the webpage bottom pagination is a numerical sequence (1, 2, 3…), it can be considered as a "numerical feature"; If the pagination uses A, B, C as the beginning (such as grouping companies by letters), it can be considered as an "alphabetical feature"; If the page initially loads only part of the data and the "Load More" button needs to be clicked to display the complete content, it indicates that it has a "loading feature"; If the enterprise information is popped up after clicking the map area, it indicates that it has a "map feature".

[0057] The website types mentioned above can include but are not limited to numerical type, alphabetical type, loading type, and map type, etc. which have specific structural data. Generally speaking, the website type can be used to guide the selection of the appropriate page parsing method. In this embodiment, the following four typical types are described: Numerical type (CrawlerMainNumber): The pagination button is a numerical index, such as 1, 2, 3…; Alphabetical type (CrawlerMainWord): Grouped display by letters or pinyin initials; Loading type (CrawlerMainLoadMore): The page needs to click the "Load More" button to display the complete content; Map type (CrawlerMainMap): The page uses visual map or region map interactive display content.

[0058] Since different website types correspond to different structural data triggers, the storage and content can also be different. Therefore, in one possible embodiment, the website information extraction system performs structured interpretation and type classification of the target website through website data features and HTML structure, so as to determine which structural type the website belongs to, and accordingly guide the subsequent layout parsing and link data extraction strategy.

[0059] By the above method steps, the structural features of the target website can be automatically judged before information extraction, and the page is classified into typical structural types such as digital type, alphabetical type, loading type or map type, so that a matching analysis strategy is adopted for different website layouts. The mechanism effectively improves the adaptability of the system to diversified web page structures, avoids the limitation of manually configuring rules in traditional crawlers, significantly enhances the universality, accuracy and automation level of information extraction, and is especially suitable for application scenarios such as exhibition official websites, enterprise directory platforms and the like with non-uniform structures and frequent dynamic loading.

[0060] Optionally, in the step of identifying and matching the corresponding link data in the at least one address information based on the target content language description data and the page layout data of the target website, and determining the at least one target link data, the step further comprises: determining the target demand content of the user based on the target content language description data; analyzing the page layout data to determine the position feature data and the structural feature data corresponding to each link data in the page; and matching the position feature data and the structural feature data corresponding to each link data with the target demand content to determine the at least one target link data.

[0061] In the embodiment of the application, the above-mentioned target demand content can be a specific information demand field extracted by the above-mentioned website information extraction system according to the natural language input of the user, which is used to guide the subsequent information identification and extraction operation, represents the core data element that the user really wants to obtain, and is usually a keyword such as “company name”, “contact person”, “telephone number” and “email address”. For example, when the user inputs: “I want to download the contact information and responsible person information of all suppliers”, the above-mentioned website information extraction system converts it into “contact person”, “contact telephone number” and “company name”, that is, a keyword or natural language sentence that can highlight the user's intention.

[0062] In a possible embodiment, the above-mentioned website information extraction system analyzes and processes the web page source code or DOM structure to identify the structural level, label distribution, module block and element attribute and the like feature data in the web page.

[0063] The above-mentioned position feature data can refer to the spatial or structural position of a link element in a web page, including the module (such as main content area, sidebar, footer) where it is located, the order level where it is located, the class name of the parent container and the like.

[0064] The above-mentioned structural feature data can refer to the HTML structure information of a link element itself and its internal nested content, including but not limited to the label type of the link data, the link attribute, the nested level structure of the link element.

[0065] In this embodiment, the structure and position feature data of the link can be compared with the target demand content of the user in terms of semantics or keywords, a similarity can be calculated by setting a weight rule or model, and link data with high matching degree can be screened out. The link data with a matching score exceeding a threshold value is determined as the target link.

[0066] For example: The system calculates the matching degree of a link with "contact + phone", finds that the link contains "contact: Li Gong" and "phone: 139…", and the matching score is 0.92, which is higher than the set threshold value 0.85. Therefore, the system submits the link as target link data for processing.

[0067] Optionally, in the step of matching the position feature data and the structure feature data corresponding to each link data with the target demand content to determine at least one target link data, the function characteristics of the link data in the current webpage are determined based on the structure feature data corresponding to each link data. The corresponding association value is obtained by matching the corresponding function characteristics with the target demand content. The association value is compared with a preset association threshold value, and at least one target link data is determined from the plurality of link data.

[0068] In the embodiment of the application, the function characteristics can be the behavior characteristics of the link data in the webpage, including but not limited to the click start characteristics and the automatic jump start characteristics, which are used to reflect the specific action mode in the user interaction or webpage loading process. The click start characteristics can be the behavior characteristics that require the user to click to trigger content loading or jump. The automatic jump start characteristics can be the behavior characteristics that are automatically triggered after the page is loaded.

[0069] The website information extraction system can analyze the logical correlation between the function characteristics of a link and the target content described by the user to determine whether the behavior of the link matches the required information. The specific matching process can be based on a rule engine or a semantic scoring mechanism. For example, if the user wants to extract "enterprise detailed information", and a link can open a page containing contact information after being clicked, the matching degree is high. Conversely, if the link only pops up an advertisement, the matching degree is low.

[0070] The association value can be a numerical score given by the website information extraction system to the matching degree between a link and the target demand. The association value is usually a decimal number between 0 and 1. The higher the association value, the more likely the link carries the data that the user wants to extract. Specifically, the association value can be scored in multiple dimensions, such as function characteristic matching degree, label structure complexity, content field hit rate, position weight, and attribute semantic credibility. The final association value of the link can be obtained by weighted calculation based on the scoring content.

[0071] The preset correlation threshold can be a standard boundary value set in the website information extraction system, which is used to determine whether the link has a high enough matching degree. Generally, in an enterprise directory website, the threshold can be set to 0.85 to ensure extraction accuracy; in a sparse data page, the threshold can be appropriately reduced to 0.75 to avoid omission.

[0072] Optionally, in the step of extracting the target link data by the preset AI language large model to determine the information of the target content, the method further includes determining content data of a page accessed in the target link data; and performing semantic analysis and recognition on the content data by the preset AI language large model to determine at least one target content information corresponding to the target demand content.

[0073] In the embodiment of the present application, after the website information extraction system identifies the target link, it automatically accesses the detail page corresponding to each link, obtains the page content data, and inputs it into the preset AI language large model for semantic analysis processing, extracts the field information corresponding to the user demand (such as company name, contact person, contact number), and finally outputs the structured result, completing intelligent information extraction of multiple records.

[0074] Specifically, when the website information extraction system identifies multiple target link data, it accesses the detail page pointed by each target link one by one, extracts the content data in the page body, such as the title, "person in charge", "contact number", etc. Then, after inputting the text field as content data into the preset AI language large model, the AI language large model performs semantic understanding and field positioning on the page content, automatically identifies the entity content corresponding to the target demand input by the user (such as "company name, person in charge, phone number"), and outputs the extraction result in a structured format (such as JSON or table field).

[0075] The embodiment introduces an AI language large model for semantic analysis processing after accessing the target link page, so that the system can automatically identify and extract the key information field corresponding to the user demand without presetting the web page structure rule, improve the adaptability to unstructured web page content, significantly enhance the intelligent level and accuracy of information extraction, and reduce the artificial configuration cost. It is especially suitable for website environments with complex page structures or diverse field expressions.

[0076] As shown in Figure 3 The website information extraction device 300 provided by the embodiment of the present application comprises: A first acquisition module 301 is configured to acquire address information of a target website and language description data of a target content input by a user. The first determination module 302 is configured to identify and match at least one corresponding link data in the address information based on the target content language description data and the page layout data of the target website, and determine at least one target link data. The first processing module 303 is configured to perform extraction processing on the target link data by using a preset AI language large model, and determine information of target content.

[0077] Optionally, the first acquisition module 301 includes: A first determination sub-module is configured to determine a URL address of the target website. A first processing sub-module is configured to perform conversion processing on the target content by using a spoken language description, and determine corresponding target content language description data, which is used to point to content information corresponding to a user demand.

[0078] Optionally, the apparatus further includes: A second determination module is configured to determine a website type based on website data features, wherein the website type includes a digital type, an alphabetical type, a loading type, and a map type. A third determination module is configured to identify and analyze a page layout of a target website by using the digital type, the alphabetical type, the loading type, and the map type, and determine a website type possessed by the page layout data.

[0079] Optionally, the first determination module 302 includes: A second determination sub-module is configured to determine a target demand content of a user based on the target content language description data. A third determination sub-module is configured to analyze the page layout data, and determine position feature data and structure feature data corresponding to each link data in the page, wherein the structure data includes a label type, a link attribute, and a nested hierarchical structure of a link element of the link data. A fourth determination sub-module is configured to perform matching processing on the position feature data and the structure feature data corresponding to each link data and the target demand content, and determine at least one target link data.

[0080] Optionally, the fourth determination sub-module includes: A first determination unit is configured to determine a function characteristic of link data in a current web page based on the structure feature data corresponding to each link data, wherein the function characteristic includes a click initiation characteristic and an automatic jump initiation characteristic. A second determination unit is configured to obtain a corresponding correlation value by performing matching processing on the corresponding function characteristic and the target demand content. The third determining unit is configured to compare the association value with a preset association threshold value, and determine at least one target link data in the plurality of link data.

[0081] Optionally, the first processing module 303 includes: The fifth determining sub-module is configured to determine content data of a page accessed in the target link data. The sixth determining sub-module is configured to perform semantic analysis and identification on the content data by using a preset AI language large model, and determine information of at least one target content corresponding to the target demand content.

[0082] As shown in Figure 4 The embodiment of the present application also provides an electronic device 400 including a processor, and the processor can execute any one of the website information extraction methods.

[0083] Specifically, the electronic device 400 includes a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401 to execute the website information extraction method, wherein: The processor 401 executes the computer program of the website information extraction method stored in the memory 402, and performs the following steps: Obtain address information of a target website and user input target content language description data; Based on the target content language description data and page layout data of the target website, identify and match at least one link data corresponding to the address information to determine at least one target link data; Extract and process the target link data by using a preset AI language large model to determine information of the target content.

[0084] Optionally, the processor 401 executes the obtaining of the address information of the target website and the user input target content language description data, including: Determine the URL address of the target website; Convert the target content by using a colloquial description to determine corresponding target content language description data, and the target content language description data is used to point to content information corresponding to user demand.

[0085] Optionally, before the processor 401 executes the identifying and matching of at least one link data corresponding to the address information based on the target content language description data and the page layout data of the target website to determine at least one target link data, the method further includes: Determine the website type based on the website data characteristics, and the website type includes a digital type, an alphabetical type, a loading type, and a map type. The page layout of the target website is identified and analyzed through the digital type, the letter type, the loading type and the map type, and the website type in the page layout data is determined.

[0086] Optionally, the processor 401 performs the identification matching of the corresponding link data in at least one of the address information based on the target content language description data and the page layout data of the target website, and determines at least one target link data, including: The target demand content of the user is determined based on the target content language description data. The page layout data is parsed to determine the position feature data and the structure feature data corresponding to each link data in the page, and the structure data includes the label type, the link attribute and the nested hierarchical structure of the link element of the link data. The position feature data and the structure feature data corresponding to each link data are matched with the target demand content to determine at least one target link data.

[0087] Optionally, the processor 401 further performs the matching of the position feature data and the structure feature data corresponding to each link data with the target demand content to determine at least one target link data, including: The function characteristics of the link data in the current webpage are determined based on the structure feature data corresponding to each link data, and the function characteristics include the click start characteristics and the automatic jump start characteristics. The corresponding correlation value is obtained by matching the corresponding function characteristics with the target demand content. The correlation value is compared with a preset correlation threshold value, and at least one target link data is determined from the plurality of link data.

[0088] Optionally, the processor 401 further performs the extraction of the target link data through the preset AI language large model to determine the information of the target content, including: The content data of the accessed page in the target link data is determined. The content data is subjected to semantic analysis and identification through the preset AI language large model, and at least one information of the target content corresponding to the target demand content is determined.

[0089] The embodiment of the application also provides a computer readable storage medium, and the computer readable storage medium stores a computer program. The computer program is executed by the processor to realize each process of the website information extraction method or the application end website information extraction method provided by the embodiment of the application, and the same technical effects can be achieved. To avoid repetition, details are not repeated here.

[0090] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and can be stored in a computer readable storage medium. The program can include the processes of the above-mentioned embodiment methods when executed. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM) or a random access memory (RAM).

[0091] The above disclosure is only the preferred embodiments of the present application, and of course cannot limit the scope of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope of the present application.

Claims

1. A website information extraction method, characterized in that: include: Obtain the address information of the target website and the corresponding target content language description data entered by the user; Based on the target content language description data and the page layout data of the target website, identifying and matching the corresponding link data in at least one of the address information to determine at least one target link data; By using a preset AI language model, the target link data is extracted and processed to determine the information of the target content.

2. The website information extraction method according to claim 1, wherein: The step of obtaining the address information of the target website and the corresponding target content language description data input by the user includes: Determine the URL address of the target website; The target content is converted through a spoken description to determine corresponding target content language description data, and the target content language description data is used to point to content information corresponding to the user's needs.

3. The website information extraction method according to claim 1, wherein: Before identifying and matching corresponding link data in at least one of the address information based on the target content language description data and the page layout data of the target website and determining at least one target link data, the method further includes: Determine the website type based on the website data characteristics, where the website type includes digital type, letter type, loading type, and map type; The page layout of the target website is identified and analyzed through the digital type, letter type, loading type and map type to determine the website type contained in the page layout data.

4. The website information extraction method according to claim 1, wherein: The step of identifying and matching corresponding link data in at least one of the address information based on the target content language description data and the page layout data of the target website to determine at least one target link data includes: Determining the user's target demand content based on the target content language description data; Parsing the page layout data to determine position feature data and structural feature data corresponding to each link data in the page, wherein the structural data includes a tag type of the link data, link attributes, and a nested hierarchical structure of link elements; The position feature data and the structure feature data corresponding to each link data are matched with the target requirement content to determine at least one target link data.

5. The website information extraction method according to claim 4, wherein: The matching of the position feature data and the structural feature data corresponding to each link data with the target requirement content to determine at least one target link data includes: Determining functional characteristics of the link data in the current webpage based on the structural characteristic data corresponding to each link data, wherein the functional characteristics include a click-to-start characteristic and an automatic jump-to-start characteristic; By matching the corresponding functional characteristics with the target demand content, a corresponding correlation value is obtained; The correlation value is compared with a preset correlation threshold value to determine at least one target link data from the plurality of link data.

6. The website information extraction method according to claim 1, wherein: The target link data is extracted and processed by a preset AI language model to determine the target content information, including: Determining content data of a page to be accessed in the target link data; By using a preset AI language model, semantic analysis and recognition are performed on the content data to determine at least one target content information corresponding to the target demand content.

7. A website information extraction device, characterized in that: include: A first acquisition module is used to acquire the address information of the target website and the corresponding target content language description data input by the user; A first determining module is configured to identify and match corresponding link data in at least one of the address information based on the target content language description data and the page layout data of the target website, and determine at least one target link data; The first processing module is used to extract and process the target link data through a preset AI language model to determine the information of the target content.

8. A website information extraction system, characterized in that: The website information extraction system includes: a website information extraction device; The website information extraction device implements the website information extraction method described in claim 1.

9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the website information extraction method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the website information extraction method according to any one of claims 1 to 6.