Enterprise information acquisition method and system, electronic device, and storage medium

WO2026200548A1PCT designated stage Publication Date: 2026-10-01HANGZHOU ALIBABA INT INTERNET IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/082944
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-11
Publication Date
2026-10-01

Smart Images

  • Figure CN2026082944_01102026_PF_FP_ABST
    Figure CN2026082944_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an enterprise information acquisition method and system, an electronic device, and a storage medium. The method comprises: in response to an enterprise information acquisition operation, a client acquires a page link of a current list page of a yellow pages website, a next-page link example and a detail-page link example, and sends an information acquisition request to a preset server by using the page link of the current list page, the next-page link example and the detail-page link example as request parameters, the information acquisition request being used for triggering the server to acquire, on the basis of the request parameters, enterprise information of an enterprise displayed on the list page of the yellow pages website; and the client displays execution information of an information acquisition task, wherein the execution information of the information acquisition task comprises one or more of: site information of the yellow pages website, acquisition progress of the enterprise information, and an execution status of the task. In the present method, enterprise information is automatically collected on the basis of a yellow pages website and subjected to structured processing, thereby effectively improving the efficiency of enterprise information acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Enterprise information acquisition methods, systems, electronic devices and storage media

[0001] This disclosure claims priority to Chinese Patent Application No. 202510378506.4, filed with the China National Intellectual Property Administration on March 27, 2025, entitled “Method, System, Electronic Device and Storage Medium for Obtaining Enterprise Information”, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of computer technology, and in particular to a method for acquiring enterprise information, a system for acquiring enterprise information, an electronic device, a storage medium, and a computer program product. Background Technology

[0003] Yellow Pages are internationally recognized business directories organized by company type and product category. In existing technology, standard yellow page websites display summary information of listed companies on the list page, which includes at least the company name. For companies with complete information, the list page may also include a company profile and address. Because the structure of different yellow page websites varies greatly, and the types of company information in yellow pages are limited, the utilization efficiency of existing yellow page websites is relatively low, primarily relying on manual searches to obtain company information.

[0004] Therefore, a method is needed to automatically obtain business information based on yellow pages. Summary of the Invention

[0005] This disclosure provides a method for obtaining enterprise information, which can automatically obtain enterprise information based on a business directory, thereby improving not only the efficiency of enterprise information acquisition but also the comprehensiveness of the information obtained.

[0006] Accordingly, this disclosure also provides an enterprise information acquisition system, an electronic device, a storage medium, and a computer program product to ensure the implementation and application of the above-mentioned enterprise information acquisition method.

[0007] To address the aforementioned problems, this disclosure provides a method for obtaining enterprise information, applied to a client-side application. The method includes:

[0008] In response to the enterprise information retrieval operation, retrieve the page link of the current list page, the next page link example, and the details page link example from the yellow pages website;

[0009] Using the page link of the current list page, the example link of the next page, and the example link of the details page as request parameters, an information retrieval request is sent to a preset server. The information retrieval request is used to trigger the server to perform the following information retrieval task: based on the request parameters, retrieve the enterprise information of the enterprises displayed on the list page of the yellow pages website;

[0010] The execution information of the information acquisition task is displayed, wherein the execution information includes one or more of the following: site information of the yellow pages website, the progress of acquiring the enterprise information, and the execution status of the information acquisition task.

[0011] This disclosure provides a method for obtaining enterprise information, applied to a server, the method comprising:

[0012] In response to an information retrieval request sent by the client, the request parameters of the information retrieval request are parsed to obtain the page link of the current list page, the next page link example, and the details page link example of the yellow pages website;

[0013] Create and execute the information retrieval task corresponding to the information retrieval request. The information retrieval task is used to: obtain the enterprise information of the enterprises displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website.

[0014] The execution information of the information acquisition task is sent to the client, so that the client displays the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, the progress of acquiring the enterprise information, and the execution status of the information acquisition task.

[0015] This disclosure provides an enterprise information acquisition system, including a client and a server, wherein...

[0016] The client is used to execute the enterprise information acquisition method as described above;

[0017] The server is used to execute the enterprise information acquisition method described above.

[0018] This disclosure also discloses a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the methods described in this disclosure.

[0019] This disclosure also discloses a computer program product, including a computer program / computer executable instructions, wherein the computer program / computer executable instructions, when executed by a processor in an electronic device, implement the method described in this disclosure.

[0020] Compared with the prior art, the embodiments of this disclosure have the following advantages:

[0021] The client responds to the enterprise information retrieval operation by obtaining the page link of the current list page, the next page link example, and the details page link example from the yellow pages website. Using these links as request parameters, it sends an information retrieval request to a pre-defined server. This request triggers the server to execute the following information retrieval task: based on the request parameters, it retrieves the enterprise information displayed on the yellow pages website's list page; then, it displays the execution information of the information retrieval task, which includes one or more of the following: the yellow pages website's site information, the retrieval progress of the enterprise information, and the task's execution status. This achieves automatic collection and structured processing of enterprise information based on the yellow pages website, effectively improving the efficiency of enterprise information retrieval. Attached Figure Description

[0022] Figure 1 is one of the flowcharts of the enterprise information acquisition method disclosed in this embodiment;

[0023] Figure 2 is a schematic diagram of the list page of a standard yellow pages website in the prior art;

[0024] Figure 3 is a schematic diagram of the implementation architecture of the enterprise information acquisition method disclosed in this embodiment;

[0025] Figure 4 is one of the schematic diagrams of the client interface in the enterprise information acquisition method disclosed in this embodiment;

[0026] Figure 5 is a second flowchart of the steps of the enterprise information acquisition method disclosed in this embodiment of the present disclosure;

[0027] Figure 6 is a flowchart of the third step of the enterprise information acquisition method disclosed in this embodiment;

[0028] Figure 7 is a schematic diagram of the interactive process of the enterprise information acquisition system disclosed in this embodiment.

[0029] Figure 8 is a schematic diagram of the structure of an exemplary device provided in an embodiment of this disclosure. Detailed Implementation

[0030] To make the above-mentioned objectives, features and advantages of this disclosure more apparent and understandable, the disclosure will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0031] In existing technologies, standard yellow pages websites display summary information of listed companies on their list pages, which also include pagination links, as shown in Figure 2. The company information list displays the information for each company. In this embodiment, by parsing the page files of the list pages and detail pages of the standard yellow pages website, the page files of all list pages of the yellow pages website are obtained. A large language model is then used to analyze the page content in the page files, automatically extracting the company information listed on each page of the yellow pages website. Furthermore, by combining this with supplementary information searched from preset data sources, the company information is further improved, resulting in more comprehensive company information.

[0032] In some optional embodiments, the enterprise information acquisition method disclosed in this disclosure can be implemented through an enterprise information acquisition system, the implementation architecture of which is shown in Figure 3. The enterprise information acquisition system provides external data sources to the system by accessing search engine services and database access services containing enterprise information; it extracts hyperlinks and text content from list pages by accessing hyperlink extraction tools and page text extraction tools; and it achieves the identification of next-page links and detail page links, and the analysis, reasoning, and structuring of page text content by accessing the information analysis capabilities of a large language model.

[0033] The enterprise information acquisition system provided in this embodiment acquires enterprise information on the server side through two intelligent agents: an information recognition intelligent agent and an information search intelligent agent. Optionally, the information recognition intelligent agent is used to identify next-page links and detail page links, and automatically acquires the page files of each list page based on the next-page links, identifies the detail page links in the page files, and then automatically acquires the page content of the detail page corresponding to each detail page link. Afterwards, it identifies the structured information of the corresponding enterprise from the page content. Optionally, the information search intelligent agent is used to search for enterprise information outside of websites (such as standard yellow pages websites) to improve the enterprise information extracted from websites. Furthermore, it is also used to infer and generate other enterprise information based on existing enterprise information, such as the enterprise's industry, whether it is an e-commerce company, or whether it is a foreign trade company.

[0034] By using the enterprise information acquisition method disclosed in this embodiment to analyze overseas business directory websites, it is possible to efficiently discover potential overseas customers.

[0035] As shown in Figure 1, the enterprise information acquisition method applied to the client of the enterprise information acquisition system includes steps 102 to 106.

[0036] Step 102: In response to the enterprise information retrieval operation, obtain the page link of the current list page, the next page link example, and the details page link example of the yellow pages website.

[0037] In some optional embodiments, an information input interface can be set on the client side for users to input the search start list page of the yellow pages website to be searched, the next page link example of the yellow pages website, and the details page link example. Taking the client interface shown in Figure 4 as an example, the user can input the hyperlink of the start list page (such as the homepage of the yellow pages website) of the current yellow pages website to be searched in edit box 402, input the hyperlink of the next page of any list page of the yellow pages website in edit box 404, and input the hyperlink of any details page of the yellow pages website in edit box 406. After the user triggers the enterprise information retrieval operation through the client interface, the client retrieves the hyperlinks entered in each edit box, uses the hyperlink in edit box 402 as the page link of the current list page, uses the hyperlink in edit box 404 as the next page link example, and uses the hyperlink in edit box 406 as the details page link example.

[0038] Step 104: Using the page link of the current list page, the example link of the next page, and the example link of the details page as request parameters, send an information retrieval request to a preset server. The information retrieval request is used to trigger the server to perform the following information retrieval task: based on the request parameters, retrieve the enterprise information of the companies displayed on the list page of the yellow pages website.

[0039] The client generates an information retrieval request using the current list page link, the next page link example, and the details page link example as request parameters, and sends the information retrieval request to the server of the enterprise information retrieval system.

[0040] After receiving an information retrieval request, the server parses the request, obtains the request parameters, and thus obtains the page link of the current list page, the next page link example, and the details page link example. Then, the server creates and executes a company information scraping task. During execution, the task retrieves company information from the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example.

[0041] Optionally, obtaining the enterprise information of the companies displayed on the list page of the yellow pages website based on the request parameters includes: sub-steps 1201, 1202, and 1203.

[0042] Sub-step 1201: Based on the page links of the current list page, the next page link example, and the details page link example, analyze the page files of each page of the yellow pages website layer by layer to obtain the structured information of the companies displayed on each page.

[0043] In existing technology, the page files of yellow pages websites, including list pages and detail pages, use HTML (HyperText Markup Language) instructions to describe the page structure. Taking a list page as an example, the page file uses... The `<link>` tag defines a hyperlink used to link from the current page to another HTML page, such as linking from the current list page to a company information detail page, the next list page, the previous list page, or a specified list page. Based on this, by parsing the page file of each list page, not only can we obtain the links to the detail pages of the companies listed on the current list page, but also the links to the next list page, as well as the text content of each page. Furthermore, based on the obtained text content of each company information detail page, the company information can be extracted.

[0044] Optionally, the step of progressively analyzing the page files of each page of the yellow pages website based on the page links of the current list page, the next page link example, and the details page link example to obtain the structured information of the companies displayed on each page includes: sub-steps S1 to S5.

[0045] Sub-step S1: Initialize the current page links with the page links of the current list page.

[0046] Sub-step S2: Based on the current page link, obtain the page file of the current list page.

[0047] For specific implementation methods of obtaining the page file of the current list page based on the current page link, please refer to the prior art, and will not be repeated in the embodiments of this disclosure.

[0048] Sub-step S3: Referring to the next page link example and the details page link example, identify the details page link and the next page link in the page file.

[0049] The current list page's page file uses HTML directives to describe one or more of the following information: a text description of the company displayed on the current page and a link to a details page displaying the company's information, a link to the next list page, a link to the previous list page, and a link to a specified list page.

[0050] Optionally, identifying the details page links and next page links in the page file by referring to the next page link example and the details page link example includes: obtaining a list of hyperlinks in the page file; using the next page link example, the details page link example, and the list of hyperlinks as input, triggering a preset large language model to refer to the next page link example and the details page link example, and identify the details page links and next page links in the list respectively.

[0051] In some optional embodiments, prompt words can be generated based on the next page link example and the content of the page file, triggering a preset large language model to refer to the next page link example and identify the next page link in the content of the page file; and prompt words can be generated based on the details page link example and the content of the page file, triggering a preset large language model to refer to the details page link example and identify the details page link in the content of the page file.

[0052] In some alternative embodiments, to improve the accuracy of identifying next page links and details page links, it can first be based on Tags are used to extract all hyperlinks from the page file, resulting in a list of hyperlinks in the current page file. Then, a large language model is triggered to identify next-page links and detail page links from this list. For example, based on the next-page link example and the list of hyperlinks, a prompt word is generated, triggering a preset large language model to identify the next-page links in the list by referring to the next-page link example; similarly, based on the detail page link example and the list of hyperlinks, a prompt word is generated, triggering a preset large language model to identify the detail page links in the list by referring to the detail page link example.

[0053] In some optional embodiments, the following prompt template can be used to generate prompt words that identify the next page link:

[0054] #Task

[0055] You will be provided with several links from a website.Your task is to identify and extra-ct the pagination URLs based on the given example.

[0056] #Requirements

[0057] **Only return the pagination URLs**,separated by commas.

[0058] **Do not include any explanation** in the output.

[0059] #Input

[0060] links: $! {links}

[0061] example:$!{example}”

[0062] In the prompt word template, "#Task" represents the task that the large language model is required to perform, "#Requirements" represents additional requirements for the large language model to perform the task, and "#Input" represents the source data for performing the task. Among them, "$!{links}" is the hyperlink variable to be identified, and "$!{example}" is the reference example variable. When generating prompt words based on the prompt word template, "$!{links}" uses the list of hyperlinks on the current page, and "$!{example}" is replaced with the next page link example entered by the user.

[0063] In some optional embodiments, the following prompt template can be used to generate prompts that identify links to detail pages:

[0064] #Task

[0065] You will be provided with several links from a website.Your task is to identify and extract the company detail's page URLs based on the given example.

[0066] #Requirements

[0067] **Only return the company detail's page URLs**,separated by commas.

[0068] **Do not include any explanation** in the output.

[0069] #Input

[0070] links: $! {links}

[0071] example:$!{example}”

[0072] In the prompt word template, "#Task" represents the task that the large language model is required to perform, "#Requirements" represents additional requirements for the large language model to perform the task, and "#Input" represents the source data for performing the task. Here, "$!{links}" represents the hyperlink variable to be identified, and "$!{example}" represents the reference example variable. When generating prompt words based on the prompt word template, "$!{links}" is replaced with the list of hyperlinks on the current page, and "$!{example}" is replaced with the example link from the details page entered by the user.

[0073] Directly calling the large language model to extract detail page links or next page links from all hyperlinks on the page has low accuracy. In the embodiments of this disclosure, by inputting a small number of sample prompts to provide the large language model with a few examples, the accuracy of extracting detail page links or next page links can be improved.

[0074] Sub-step S4 involves analyzing the text content in the enterprise information detail pages corresponding to each of the detail page links to obtain the structured information of each enterprise included in the current list page.

[0075] Optionally, the step of analyzing the text content in the enterprise information detail pages corresponding to each of the detail page links to obtain the structured information of each enterprise included in the current list page includes: for each of the detail page links, obtaining the page file corresponding to each detail page link; parsing the page file of the enterprise information detail page corresponding to each detail page link to obtain the text content in the enterprise information detail page; using a preset large language model to analyze and identify the text content to obtain the structured information of the enterprise corresponding to the enterprise information detail page; and integrating the structured information of the enterprises corresponding to each of the enterprise information detail pages to obtain the structured information of each enterprise included in the current list page.

[0076] The structured information includes, but is not limited to, the following: company name, mailing address, contact person, contact number, country, creation time, etc.

[0077] For details on obtaining the page file corresponding to the link to the details page, and parsing the page file of the enterprise information details page corresponding to each link to the details page, and obtaining the text content in the enterprise information details page, please refer to the prior art, and will not be repeated in the embodiments of this disclosure.

[0078] After obtaining the text content corresponding to each enterprise information detail page, prompt words can be generated based on the preset prompt word template and the text content corresponding to the current enterprise information detail page. Based on the generated prompt words, a preset large language model is called to trigger the preset large language model to analyze and identify the text content, extract the specified enterprise information, and output it according to the preset format, thereby obtaining the structured information of the enterprise corresponding to the enterprise information detail page.

[0079] Optionally, the prompts may include, but are not limited to, the following: task description text instructing the large language model to extract and structure detailed information about a specified company from the given webpage content; a description of the specified company details; output information format; information extraction rules; and the given webpage content. The specified company details include, but are not limited to, one or more of the following: company name, website, telephone number, mailing address, country code, annual budget, company type, contact person, email address, etc.

[0080] The specific content of the structured information is determined according to specific needs, and the specific content of the structured information is not limited in the embodiments of this disclosure.

[0081] In some alternative embodiments, business information obtained from analysis that cannot be analyzed from the page files of the yellow pages website can be output in a specific format, such as an empty string.

[0082] After executing sub-steps S2, S3, and S4, the structured information of each company displayed on the current list page can be obtained.

[0083] Sub-step S5: In response to successfully identifying the next page link, the next page link is used as the current page link, and the process jumps to the step of obtaining the page file of the current list page based on the current page link, and then to the step of analyzing the text content in the enterprise information detail pages corresponding to each of the detail page links to obtain the structured information of each enterprise included in the current list page.

[0084] As mentioned earlier, for lists that are not the last, the list page includes a link to the next list page; therefore, the large language model will output the next page link. However, for the last list page, the list page does not include a link to the next list page, and the next page link output by the large language model is either empty or invalid.

[0085] If the large language model outputs a valid next page link, it means that the current list is not the last list page, and we need to continue processing the next list page, obtain the page file of the next list page, and then execute sub-steps S3 and S4.

[0086] In some optional embodiments, after analyzing the text content in the enterprise information detail pages corresponding to each of the detail page links to obtain the structured information of each enterprise included in the current list page, the method further includes: in response to not identifying a next page link, ending the progressive analysis operation on the yellow pages website.

[0087] In the list pages of the yellow pages website, the last list page does not have a hyperlink pointing to the next page. Therefore, when the large language model analyzes and identifies the page file of the last list page, it cannot identify the hyperlink pointing to the next page. That is, when no next page link is identified in the page file, the hierarchical progressive analysis operation of the yellow pages website can be considered complete.

[0088] Sub-step 1202: Based on the structured information, perform a search operation on a preset data source to obtain the complete information of the corresponding enterprise.

[0089] By analyzing the page files of yellow pages websites, some company information displayed on the pages is obtained and structured information is generated. Further information not obtained from the yellow pages websites is then obtained through other channels to supplement the company information.

[0090] In some optional embodiments, the structured information includes: company name; the step of performing a search operation based on the structured information to obtain the corresponding company's complete information includes: obtaining a description of the missing information of each company based on the structured information; and searching the preset data source based on the company name and the description of the missing information to obtain the corresponding company's complete information. The preset data source includes one or more of the following: a preset search engine, a database of a preset customer management system.

[0091] The missing information can be: enterprise information in the structured information where the information value is an empty string; the description of the missing information can be: the field name or field description of the enterprise information field in the structured information where the information value is an empty string. Optionally, the missing enterprise information can be determined based on the format of the structured information output by the large language model.

[0092] For example, when the large language model extracts the company name of a company named "nameA" from the details page but fails to extract the contact information for company A, the company name field in the structured information will have the value "nameA", and the contact information field will have the value "" (i.e., an empty string). When retrieving the complete information, "nameA contact person" can be used as a keyword to search a preset search engine and retrieve the contact information for company nameA returned by the search engine. Then, the contact information can be extracted from the retrieved results returned by the preset large language model. Alternatively, when retrieving the complete information, "nameA" can be used as a parameter to call the customer information query interface of the preset customer management system database to retrieve the contact information for company nameA.

[0093] In some optional embodiments, the structured information includes: company names. The step of performing a search operation based on the structured information using a preset data source to obtain the corresponding company's complete information includes: obtaining a description of the missing information for each company based on the structured information; searching the preset data source based on the company name and the description of the missing information to obtain search results for the corresponding company, wherein the preset data source includes: a database of a preset search engine and / or a preset customer management system; and using a preset large language model to infer the search results to obtain the corresponding company's complete information.

[0094] Based on the structured information, the description of the missing information for each enterprise is obtained. The specific implementation method for searching a preset data source based on the enterprise name and the description of the missing information to obtain the search results for the corresponding enterprise is described above and will not be repeated here. After obtaining the search results, for enterprise information that cannot be obtained from the search results, such as the continent, city, industry, and type of the enterprise, it is necessary to infer it based on the search results. For example, after obtaining the enterprise introduction based on the target enterprise's name, the enterprise description can be further used as input data to generate prompts according to a preset template. The preset template includes candidate industries. Then, based on the prompts, a large language model is invoked, triggering the large language model to infer the candidate industries matching the enterprise based on the enterprise name and the enterprise introduction.

[0095] By searching for company information from preset data sources outside of yellow pages websites, supplementing the company information on yellow pages websites, and by inferring specific company information based on the company introductions obtained from the search, the channels for obtaining company information are expanded, which helps to obtain more comprehensive and complete company information.

[0096] Sub-step 1203: Based on the structured information and the completion information of each enterprise, obtain the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website.

[0097] By assigning values ​​to the unassigned information fields in the structured information using the company's completion information, the company's complete structured information can be obtained.

[0098] In some alternative embodiments, obtaining the enterprise information of the companies displayed on the list page of the yellow pages website based on the request parameters includes: initializing the current page link with the page link of the current list page; obtaining the page file of the current list page based on the current page link; identifying the detail page link and the next page link in the page file by referring to the next page link example and the detail page link example; analyzing the text content in the enterprise information detail page corresponding to each detail page link to obtain the structured information of each enterprise included in the current list page; performing a search operation on a preset data source based on the structured information to obtain the corresponding enterprise's complete information; obtaining the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website based on the structured information and the complete information of each enterprise; and, in response to successfully identifying the next page link, using the next page link as the current page link, jumping to the step of obtaining the page file of the current list page based on the current page link and the step of obtaining the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website based on the structured information and the complete information of each enterprise. The specific implementation methods for each step are described in the previous embodiments and will not be repeated here.

[0099] Those skilled in the art should understand that the steps of performing a search operation based on the structured information to obtain the corresponding enterprise's complete information, and the steps of obtaining the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website based on the structured information and the complete information of each enterprise, can be performed after sub-step S4, and are not limited to one execution order.

[0100] Afterwards, the server uses the obtained supplementary information to fill in the missing enterprise information in the structured information, obtains the final structured enterprise information for each enterprise, and sends the final enterprise information to the client.

[0101] Step 106: Display the execution information of the information acquisition task, wherein the execution information includes one or more of the following: site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task.

[0102] The progress of acquiring enterprise information includes: the total number of enterprises included in the yellow pages website and the total number of enterprises whose enterprise information has been acquired; the execution status of the information acquisition task is used to indicate whether the task has been completed.

[0103] After the client sends an information retrieval request to the preset server, the server creates an information retrieval task for the request and executes it. Then, the client receives progress information from the server regarding the execution of the information retrieval task and displays this progress information and the site information of the yellow pages website in real time, as shown in Figure 4.

[0104] Optionally, as shown in Figure 5, after displaying the execution information of the information acquisition task, the method further includes: step 108.

[0105] Step 108: In response to the execution record query operation of the information acquisition task, display a list of enterprise information of each enterprise acquired by the server.

[0106] In some optional embodiments, the server stores execution records of the information retrieval tasks performed for each information retrieval request. These execution records include: task execution time, corresponding yellow pages website, and execution results. The execution results include: company information displayed on the yellow pages website's list page. This company information includes: company information extracted from the yellow pages website's pages, and company information searched and inferred from a preset data source.

[0107] Optionally, in response to the execution record query operation of the information retrieval task, a list of enterprise information of each of the aforementioned enterprises obtained by the server is displayed, including: in response to the execution record query operation of the information retrieval task, sending an execution record query request to the server, the execution record query request carrying an execution record identifier, the execution record query request causing the server to send the execution result of the information retrieval task corresponding to the execution record identifier to the client, the execution result including: enterprise information of the enterprises displayed on the list page of the yellow pages website; and the client displaying a list of enterprise information of the aforementioned enterprises.

[0108] When a user clicks on the information displayed on the client to retrieve the execution record of the requested information retrieval task, the client further retrieves the execution result corresponding to the execution record clicked by the user, and displays a list of enterprise information on the client based on the execution result.

[0109] In summary, the enterprise information acquisition method disclosed in this embodiment, through a client responding to an enterprise information acquisition operation, obtains the page link of the current list page, the next page link example, and the details page link example of a yellow pages website. Using these page links as request parameters, it sends an information acquisition request to a preset server. This request triggers the server to execute the following information acquisition task: based on the request parameters, it acquires the enterprise information of the enterprises displayed on the yellow pages website's list page; then, it displays the execution information of the information acquisition task, which includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the task. This method achieves automatic collection and structured processing of enterprise information based on yellow pages websites, effectively improving the efficiency of enterprise information acquisition.

[0110] Furthermore, by searching preset data sources and reasoning based on search results, supplementary information is obtained, and this supplementary information is used to complete the missing enterprise information in the yellow pages website, effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0111] Based on the above embodiments, this embodiment also provides a method for obtaining enterprise information applied to the server, as shown in Figure 6. The method includes steps 602 to 606.

[0112] Step 602: In response to the information retrieval request sent by the client, parse the request parameters of the information retrieval request to obtain the page link of the current list page, the next page link example, and the details page link example of the yellow pages website.

[0113] For details on how the client sends an information retrieval request, please refer to the relevant descriptions in the previous embodiments, which will not be repeated here.

[0114] Subsequently, the server parses the request parameters of the information retrieval request according to the preset communication protocol between the client and the server, and obtains the page link of the current list page, the next page link example, and the details page link example of the yellow pages website.

[0115] Step 604: Create and execute the information retrieval task corresponding to the information retrieval request. The information retrieval task is used to: obtain the enterprise information of the enterprises displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website.

[0116] Optionally, the server can only create one information retrieval task for each information retrieval request and start executing that information retrieval task.

[0117] Optionally, obtaining the enterprise information of the companies displayed on the list page of the yellow pages website based on the page links of the current list page, the next page link example, and the details page link example of the yellow pages website includes: analyzing the page files of each page of the yellow pages website layer by layer based on the page links of the current list page, the next page link example, and the details page link example to obtain the structured information of the companies displayed on each page; performing a search operation on a preset data source based on the structured information to obtain the complete information of the corresponding companies; and obtaining the enterprise information of the corresponding companies displayed on the list page of the yellow pages website based on the structured information and the complete information of each company.

[0118] The specific implementation methods for each step of the server-side to obtain the enterprise information of the companies displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website are described in the relevant descriptions in the previous embodiments, and will not be repeated here.

[0119] Step 606: Send the execution information of the information acquisition task to the client, so that the client displays the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, the progress of acquiring the enterprise information, and the execution status of the information acquisition task.

[0120] During the execution of the information acquisition task, the server can push the execution information of the information acquisition task to the client in real time according to the progress of the information acquisition task in acquiring enterprise information. For example, the server can push the progress of acquiring enterprise information and the execution status of the information acquisition task for the client to display synchronously.

[0121] On the other hand, the server stores the execution records of the information retrieval tasks. These execution records include: task execution time, the corresponding yellow pages website, and execution results. The execution results include: enterprise information extracted from the yellow pages website, and enterprise information searched and inferred from preset data sources.

[0122] Optionally, after sending the execution information of the information retrieval task to the client, the method further includes: responding to the execution record query request sent by the client, obtaining the execution record identifier carried in the execution record query request; and sending the execution result of the information retrieval task corresponding to the execution record identifier to the client, wherein the execution result includes: enterprise information of the companies displayed on the list page of the yellow pages website obtained by executing the information retrieval task. The execution record identifier corresponds one-to-one with the information retrieval task, and the server can retrieve the execution record and execution result of the corresponding information retrieval task based on the execution record identifier.

[0123] The method for generating the execution record query request is described in the prior art and will not be repeated in this embodiment.

[0124] In summary, the enterprise information acquisition method disclosed in this embodiment involves the server responding to an information acquisition request sent by the client, parsing the request parameters of the information acquisition request to obtain the page link of the current list page, the next page link example, and the details page link example of the yellow pages website; then, an information acquisition task corresponding to the information acquisition request is created and executed. The information acquisition task is used to: acquire the enterprise information of the enterprises displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website; next, the server sends the execution information of the information acquisition task to the client, causing the client to display the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task, thereby realizing the automatic collection and structured processing of enterprise information based on the yellow pages website, effectively improving the efficiency of enterprise information acquisition.

[0125] Furthermore, by searching preset data sources and reasoning based on search results, supplementary information is obtained, and this supplementary information is used to complete the missing enterprise information in the yellow pages website, effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0126] Based on the above embodiments, this embodiment also provides an enterprise information acquisition system, which includes a client and a server. As shown in Figure 7, the specific implementation process of the enterprise information acquisition system includes steps 702 to 718.

[0127] Step 702: The client responds to the enterprise information retrieval operation by obtaining the page link of the current list page, the next page link example, and the details page link example from the yellow pages website.

[0128] Step 704: The client sends an information retrieval request to the preset server using the page link of the current list page, the next page link example, and the details page link example as request parameters.

[0129] The information retrieval request is used to trigger the server to perform the following information retrieval task: based on the request parameters, retrieve the enterprise information of the companies displayed on the list page of the yellow pages website.

[0130] Step 706: The server responds to the information retrieval request sent by the client, parses the request parameters of the information retrieval request, and obtains the page link of the current list page, the next page link example, and the details page link example of the yellow pages website.

[0131] Step 708: The server creates and executes the information retrieval task corresponding to the information retrieval request. The information retrieval task is used to: obtain the enterprise information of the enterprises displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website.

[0132] Step 710: The server sends the information to the client to obtain the task execution information.

[0133] The execution information includes one or more of the following: site information of the yellow pages website, the progress of acquiring the enterprise information, and the execution status of the information acquisition task.

[0134] Step 712: The client displays the execution information of the information acquisition task.

[0135] Step 714: In response to the execution record query operation of the information acquisition task, the client sends an execution record query request to the server.

[0136] The execution record query request carries an execution record identifier.

[0137] Step 716: In response to the execution record query request, the server sends the information corresponding to the execution record identifier to the client to obtain the execution result of the task. The execution result includes: enterprise information of the companies displayed on the list page of the yellow pages website.

[0138] Step 718: The client displays a list of the company's information.

[0139] For specific implementation details of the above steps performed by the client and server, please refer to the specific implementation details in the previous embodiments, which will not be repeated here.

[0140] In summary, the enterprise information retrieval system disclosed in this embodiment involves a client responding to an enterprise information retrieval operation by obtaining the page link of the current list page, a sample next page link, and a sample details page link from a yellow pages website. Then, using these links as request parameters, the client sends an information retrieval request to a pre-defined server. The server responds to the client's request by creating and executing an information retrieval task corresponding to the request. Based on the page link of the current list page, the sample next page link, and the sample details page link from the yellow pages website, the server retrieves the enterprise information displayed on the yellow pages website's list page. Next, the server sends execution information of the information retrieval task to the client, which then displays this information. This execution information includes one or more of the following: site information of the yellow pages website, the progress of enterprise information retrieval, and the execution status of the information retrieval task. This system achieves automatic collection and structured processing of enterprise information based on yellow pages websites, effectively improving the efficiency of enterprise information retrieval.

[0141] Furthermore, by searching preset data sources and reasoning based on search results, supplementary information is obtained, and this supplementary information is used to complete the missing enterprise information in the yellow pages website, effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0142] Based on the above embodiments, this embodiment also provides an enterprise information acquisition device, applied to a client, the device comprising:

[0143] The input link retrieval module is used to respond to enterprise information retrieval operations and retrieve the page link of the current list page, the next page link example, and the details page link example from the yellow pages website;

[0144] The enterprise information acquisition module is used to send an information acquisition request to a preset server using the page link of the current list page, the next page link example, and the details page link example as request parameters. The information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, acquire the enterprise information of the enterprises displayed on the list page of the yellow pages website.

[0145] The display module is used to display the execution information of the information acquisition task, wherein the execution information includes one or more of the following: site information of the yellow pages website, the progress of acquiring the enterprise information, and the execution status of the information acquisition task.

[0146] Optionally, obtaining the enterprise information of the companies displayed on the list page of the yellow pages website based on the request parameters includes:

[0147] Based on the page links of the current list page, the next page link example, and the details page link example, the page files of each page of the yellow pages website are analyzed layer by layer to obtain the structured information of the companies displayed on each page;

[0148] Based on the structured information, perform a search operation on a preset data source to obtain the complete information of the corresponding enterprise;

[0149] Based on the structured information and the completion information of each enterprise, obtain the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website.

[0150] Optionally, based on the page links of the current list page, the next page link example, and the details page link example, the page files of each page of the yellow pages website are analyzed layer by layer to obtain the structured information of the companies displayed on each page, including:

[0151] Initialize the current page link with the page link of the current list page;

[0152] Based on the current page link, obtain the page file of the current list page;

[0153] Referring to the examples of the next page link and the details page link, identify the details page link and the next page link in the page file;

[0154] By analyzing the text content in the enterprise information detail pages corresponding to each of the aforementioned detail page links, the structured information of each enterprise included in the current list page can be obtained;

[0155] In response to the successful identification of the next page link, the next page link is used as the current page link, and the process jumps to the step of obtaining the page file of the current list page based on the current page link, and then to the step of analyzing the text content in the enterprise information detail pages corresponding to each of the detail page links to obtain the structured information of each enterprise included in the current list page.

[0156] Optionally, the structured information includes: the company name, and the step of performing a search operation based on the structured information using a preset data source to obtain the corresponding company's complete information includes:

[0157] Based on the structured information, a description of the missing information for each of the enterprises is obtained;

[0158] Based on the company name and the description of the missing information, a preset data source is searched to obtain the corresponding company's complete information. The preset data source includes: a preset search engine and / or a database of a preset customer management system.

[0159] Optionally, the structured information includes: the company name, and the step of performing a search operation based on the structured information using a preset data source to obtain the corresponding company's complete information includes:

[0160] Based on the structured information, a description of the missing information for each of the enterprises is obtained;

[0161] Based on the description of the company name and the missing information, a preset data source is searched to obtain search results for the corresponding company. The preset data source includes: a database of a preset search engine and / or a preset customer management system.

[0162] The search results are inferred using a pre-defined large language model to obtain the corresponding company's complete information.

[0163] Optionally, after displaying the execution information of the information acquisition task, the device further includes:

[0164] In response to the execution record query operation of the information acquisition task, a list of enterprise information of each enterprise acquired by the server is displayed.

[0165] The enterprise information acquisition device disclosed in this embodiment is used to implement the above-described enterprise information acquisition method. For the specific implementation of each module of the device, please refer to the specific implementation of the corresponding steps in the foregoing method embodiment, which will not be repeated here.

[0166] In summary, the enterprise information acquisition device disclosed in this embodiment, through a client responding to an enterprise information acquisition operation, obtains the page link of the current list page, the next page link example, and the details page link example of a yellow pages website. Using these page links as request parameters, it sends an information acquisition request to a preset server. This request triggers the server to execute the following information acquisition task: based on the request parameters, it acquires the enterprise information of the enterprises displayed on the yellow pages website's list page; then, it displays the execution information of the information acquisition task, which includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the task. This achieves automatic collection and structured processing of enterprise information based on yellow pages websites, effectively improving the efficiency of enterprise information acquisition.

[0167] Furthermore, by searching preset data sources and reasoning based on search results, supplementary information is obtained, and this supplementary information is used to complete the missing enterprise information in the yellow pages website, effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0168] Based on the above embodiments, this embodiment also provides an enterprise information acquisition device, applied to a server, the device comprising:

[0169] The page link acquisition module is used to respond to the information acquisition request sent by the client, parse the request parameters of the information acquisition request, and obtain the page link of the current list page, the next page link example, and the details page link example of the yellow pages website;

[0170] The enterprise information acquisition module is used to create and execute the information acquisition task corresponding to the information acquisition request. The information acquisition task is used to: acquire the enterprise information of the enterprises displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website.

[0171] An execution information sending module is used to send execution information of the information acquisition task to the client, so that the client displays the execution information. The execution information includes one or more of the following: site information of the yellow pages website, the progress of acquiring the enterprise information, and the execution status of the information acquisition task.

[0172] Optionally, obtaining the enterprise information displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website includes:

[0173] Based on the page links of the current list page, the next page link example, and the details page link example, the page files of each page of the yellow pages website are analyzed layer by layer to obtain the structured information of the companies displayed on each page;

[0174] Based on the structured information, perform a search operation on a preset data source to obtain the complete information of the corresponding enterprise;

[0175] Based on the structured information and the completion information of each enterprise, obtain the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website.

[0176] Optionally, based on the page links of the current list page, the next page link example, and the details page link example, the page files of each page of the yellow pages website are analyzed layer by layer to obtain the structured information of the companies displayed on each page, including:

[0177] Initialize the current page link with the page link of the current list page;

[0178] Based on the current page link, obtain the page file of the current list page;

[0179] Referring to the examples of the next page link and the details page link, identify the details page link and the next page link in the page file;

[0180] By analyzing the text content in the enterprise information detail pages corresponding to each of the aforementioned detail page links, the structured information of each enterprise included in the current list page can be obtained;

[0181] In response to the successful identification of the next page link, the next page link is used as the current page link, and the process jumps to the step of obtaining the page file of the current list page based on the current page link, and then to the step of analyzing the text content in the enterprise information detail pages corresponding to each of the detail page links to obtain the structured information of each enterprise included in the current list page.

[0182] The enterprise information acquisition device disclosed in this embodiment is used to implement the above-described enterprise information acquisition method. For the specific implementation of each module of the device, please refer to the specific implementation of the corresponding steps in the foregoing method embodiment, which will not be repeated here.

[0183] In summary, the enterprise information acquisition device disclosed in this embodiment, through a server responding to an information acquisition request sent by a client, parses the request parameters of the information acquisition request to obtain the page link of the current list page, the next page link example, and the details page link example of a yellow pages website; then, it creates and executes an information acquisition task corresponding to the information acquisition request. The information acquisition task is used to: acquire the enterprise information of the enterprises displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website; next, the server sends the execution information of the information acquisition task to the client, causing the client to display the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task, thereby realizing the automatic collection and structured processing of enterprise information based on the yellow pages website, effectively improving the efficiency of enterprise information acquisition.

[0184] Furthermore, by searching preset data sources and reasoning based on search results, supplementary information is obtained, and this supplementary information is used to complete the missing enterprise information in the yellow pages website, effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0185] This disclosure also provides a non-volatile readable storage medium storing one or more modules (programs) that, when applied to a device, enable the device to execute instructions for the method steps in this disclosure.

[0186] This disclosure also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the methods described in this disclosure.

[0187] This disclosure also provides an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method described in this disclosure. In this disclosure, the electronic device includes devices such as servers and terminal devices.

[0188] This disclosure also discloses a computer program product, including a computer program / computer executable instructions, wherein the computer program / computer executable instructions, when executed by a processor in an electronic device, implement the method described in this disclosure.

[0189] Embodiments of this disclosure can be implemented as an apparatus configured using any suitable hardware, firmware, software, or any combination thereof, and may include electronic devices such as servers (clusters) and terminals. Figure 8 schematically illustrates an exemplary apparatus 800 that can be used to implement the various embodiments described in this disclosure.

[0190] For one embodiment, FIG8 illustrates an exemplary device 800 having one or more processors 802, a control module (chipset) 804 coupled to at least one of the processors 802, a memory 806 coupled to the control module 804, a non-volatile memory (NVM) / storage device 808 coupled to the control module 804, one or more input / output devices 810 coupled to the control module 804, and a network interface 812 coupled to the control module 804.

[0191] Processor 802 may include one or more single-core or multi-core processors, and processor 802 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 800 can serve as a server, terminal, or other device as described in the embodiments of this disclosure.

[0192] In some embodiments, apparatus 800 may include one or more computer-readable media (e.g., memory 806 or NVM / storage device 808) having instructions 814 and one or more processors 802 that are combined with the one or more computer-readable media and configured to execute the instructions 814 to implement the module and thus perform the actions described in this disclosure.

[0193] In one embodiment, the control module 804 may include any suitable interface controller to provide any suitable interface to at least one of the processors 802 and / or any suitable device or component communicating with the control module 804.

[0194] The control module 804 may include a memory controller module to provide an interface to the memory 806. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0195] Memory 806 may be used, for example, to load and store data and / or instructions 814 for device 800. In one embodiment, memory 806 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 806 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).

[0196] In one embodiment, the control module 804 may include one or more input / output controllers to provide an interface to the NVM / storage device 808 and (one or more) input / output devices 810.

[0197] For example, NVM / storage device 808 may be used to store data and / or instructions 814. NVM / storage device 808 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives (HDDs), one or more optical disc drives (CDs), and / or one or more digital universal optical disc (DVD) drives).

[0198] NVM / storage device 808 may include storage resources that are part of a device on which device 800 is mounted, or that are accessible to the device but do not necessarily have to be part of the device. For example, NVM / storage device 808 may be accessed via a network through one or more input / output devices 810.

[0199] One or more input / output devices 810 may provide an interface for device 800 to communicate with any other suitable device. Input / output devices 810 may include communication components, audio components, sensor components, etc. A network interface 812 may provide an interface for device 800 to communicate via one or more networks. Device 800 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing a wireless network based on communication standards, such as Bluetooth, WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.

[0200] In one embodiment, at least one of the processors 802 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 804. In one embodiment, at least one of the processors 802 may be logically packaged with one or more controllers of the control module 804 to form a system-in-package (SiP). In one embodiment, at least one of the processors 802 may be integrated with the logic of one or more controllers of the control module 804 on the same die. In one embodiment, at least one of the processors 802 may be integrated with the logic of one or more controllers of the control module 804 on the same die to form a system-on-a-chip (SoC).

[0201] In various embodiments, device 800 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop, handheld computing device, tablet, netbook, etc.). In various embodiments, device 800 may have more or fewer components and / or different architectures. For example, in some embodiments, device 800 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0202] The detection device can use a main control chip as a processor or control module, and sensor data, position information, etc. can be stored in a memory or NVM / storage device. The sensor group can be used as an input / output device, and the communication interface can include a network interface.

[0203] This disclosure also provides an electronic device, including: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform one or more methods as described in this disclosure. In this disclosure, the memory can store various types of data, such as target files, file-application association data, and user behavior data, thereby providing a data foundation for various processing methods.

[0204] This disclosure also provides one or more machine-readable media having executable code stored thereon, which, when executed, causes a processor to perform one or more of the methods described in this disclosure.

[0205] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0206] The various embodiments in this disclosure are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0207] This disclosure describes embodiments of methods, terminal devices (systems), and computer program products according to embodiments of this disclosure with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0208] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0209] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0210] While preferred embodiments of the present disclosure have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the present disclosure.

[0211] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0212] The foregoing has provided a detailed description of a method for obtaining enterprise information, a system for obtaining enterprise information, an electronic device, a storage medium, and a computer program product provided by this disclosure. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this disclosure. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this disclosure should not be construed as a limitation of this disclosure.

Claims

1. A method for acquiring enterprise information, applied to a client-side application, wherein, The method includes: In response to the enterprise information retrieval operation, retrieve the page link of the current list page, the next page link example, and the details page link example from the yellow pages website; Using the page link of the current list page, the example link of the next page, and the example link of the details page as request parameters, an information retrieval request is sent to a preset server. The information retrieval request is used to trigger the server to perform the following information retrieval task: based on the request parameters, retrieve the enterprise information of the enterprises displayed on the list page of the yellow pages website; The execution information of the information acquisition task is displayed, wherein the execution information includes one or more of the following: site information of the yellow pages website, the progress of acquiring the enterprise information, and the execution status of the information acquisition task.

2. The method according to claim 1, wherein, The step of obtaining the enterprise information of the companies displayed on the list page of the yellow pages website based on the request parameters includes: Based on the page links of the current list page, the next page link example, and the details page link example, the page files of each page of the yellow pages website are analyzed layer by layer to obtain the structured information of the companies displayed on each page; Based on the structured information, perform a search operation on a preset data source to obtain the complete information of the corresponding enterprise; Based on the structured information and the completion information of each enterprise, obtain the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website.

3. The method according to claim 2, wherein, The process involves progressively analyzing the page files of each page on the yellow pages website based on the page links of the current list page, the next page link example, and the details page link example, to obtain the structured information of the companies displayed on each page, including: Initialize the current page link with the page link of the current list page; Based on the current page link, obtain the page file of the current list page; Referring to the examples of the next page link and the details page link, identify the details page link and the next page link in the page file; By analyzing the text content in the enterprise information detail pages corresponding to each of the aforementioned detail page links, the structured information of each enterprise included in the current list page can be obtained; In response to the successful identification of the next page link, the next page link is used as the current page link, and the process jumps to the step of obtaining the page file of the current list page based on the current page link, and then to the step of analyzing the text content in the enterprise information detail pages corresponding to each of the detail page links to obtain the structured information of each enterprise included in the current list page.

4. The method according to claim 2 or 3, wherein, The structured information includes: company name; the step of performing a search operation based on the structured information using a preset data source to obtain the corresponding company's complete information includes: Based on the structured information, a description of the missing information for each of the enterprises is obtained; Based on the company name and the description of the missing information, a preset data source is searched to obtain the corresponding company's complete information. The preset data source includes: a preset search engine and / or a database of a preset customer management system.

5. The method according to claim 2 or 3, wherein, The structured information includes: company name; the step of performing a search operation based on the structured information using a preset data source to obtain the corresponding company's complete information includes: Based on the structured information, a description of the missing information for each of the enterprises is obtained; Based on the description of the company name and the missing information, a preset data source is searched to obtain search results for the corresponding company. The preset data source includes: a database of a preset search engine and / or a preset customer management system. The search results are inferred using a pre-defined large language model to obtain the corresponding company's complete information.

6. The method according to claim 3, wherein, The step of identifying the details page link and the next page link in the page file by referring to the next page link example and the details page link example includes: Retrieve the list of hyperlinks in the page file; Using the next page link example and the details page link example as input, a preset large language model is triggered to generate link classification rules; The hyperlink list is categorized based on the aforementioned link classification rules to identify details page links and next page links; The classification results are evaluated for confidence, and links with confidence scores higher than a preset threshold are retained as valid identification results.

7. The method according to claim 3, wherein, The step of analyzing the text content in the enterprise information detail pages corresponding to each of the aforementioned detail page links to obtain the structured information of each enterprise included in the current list page also includes: The text content is semantically segmented to extract semantic fragments related to enterprise information; Based on a pre-defined enterprise information ontology database, entity recognition and relation extraction are performed on the semantic fragments. The extracted results are mapped to a preset structured template to generate standardized information including company name, contact information, and industry classification. Cross-validate conflicting information and prioritize using frequently occurring valid information to populate structured fields.

8. The method according to claim 5, wherein, The description based on the company name and the missing information, searching a preset data source to obtain the corresponding company's complete information, also includes: The company name is semantically expanded to generate a list of synonyms and industry aliases; A multi-dimensional search query statement is constructed based on the description of the missing information, and the dimensions include keyword combination, semantic similarity, and contextual association. The search requests are executed in parallel within a preset data source, and the candidate result set is returned after merging and deduplication. The candidate results are sorted using a pre-defined scoring model, with priority given to results that have high information completeness.

9. The method according to claim 5, wherein, The step of using a preset large language model to reason about the search results and obtain the corresponding enterprise's complete information also includes: Align the search results with the structured information to identify the missing fields that need to be reasoned; Reasoning prompts are constructed based on enterprise knowledge graphs, and the prompts include contextual information, reasoning objectives, and constraints. A multi-turn conversational large language model is invoked for stepwise reasoning, with each reasoning session focusing on a missing field. Perform a logical consistency check on the reasoning results and correct any contradictory information items.

10. The method according to any one of claims 1-9, wherein, After displaying the execution information of the information acquisition task, the method further includes: In response to the execution record query operation of the information acquisition task, a list of enterprise information of each enterprise acquired by the server is displayed.

11. The method according to claim 10, wherein, The execution record query operation in response to the information acquisition task, which displays a list of enterprise information for each of the enterprises acquired by the server, further includes: Generate data access tokens based on user permissions; The system retrieves execution records using a paginated query method and supports filtering by time range, task status, and company name. The search results are anonymized to hide sensitive information fields; It provides visualization components that support dynamic switching between table view, card view, and chart view.

12. A method for acquiring enterprise information, applied on a server-side, wherein, The method includes: In response to an information retrieval request sent by the client, the request parameters of the information retrieval request are parsed to obtain the page link of the current list page, the next page link example, and the details page link example of the yellow pages website; Create and execute the information retrieval task corresponding to the information retrieval request. The information retrieval task is used to: obtain the enterprise information of the enterprises displayed on the list page of the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example of the yellow pages website. The execution information of the information acquisition task is sent to the client, so that the client displays the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, the progress of acquiring the enterprise information, and the execution status of the information acquisition task.

13. The method according to claim 12, wherein, The process of obtaining company information displayed on the list page of the yellow pages website based on the page links of the current list page, the next page link example, and the details page link example of the yellow pages website includes: Based on the page links of the current list page, the next page link example, and the details page link example, the page files of each page of the yellow pages website are analyzed layer by layer to obtain the structured information of the companies displayed on each page; Based on the structured information, perform a search operation on a preset data source to obtain the complete information of the corresponding enterprise; Based on the structured information and the completion information of each enterprise, obtain the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website.

14. The method according to claim 12 or 13, wherein, The process of creating and executing the information acquisition task corresponding to the information acquisition request further includes: Generate a unique task identifier based on the page links of the current list page; The task's unique identifier, request parameters, and execution status are stored in a distributed task queue. Tasks are executed using an asynchronous processing framework, supporting task interruption recovery and priority scheduling; Record task execution logs in real time, including the time spent on each step, error messages, and resource consumption.

15. The method according to any one of claims 12-14, wherein, Sending the information to the client to obtain the task execution information further includes: Establish a long-lived connection channel using the WebSocket protocol; The execution information is encapsulated into a standardized message format, which includes a timestamp, message type, and payload data. Supports incremental updates, transmitting only the changed execution information fields; When a connection interruption is detected, the system automatically reconnects and resends the execution information accumulated during the interruption.

16. An enterprise information acquisition system, comprising: Client and server, among which, The client is configured to perform the method as described in any one of claims 1-11; The server is used to execute the method as described in any one of claims 12-8.

17. An enterprise information acquisition application, comprising: Client and server, among which, The client is configured to perform the method as described in any one of claims 1-11; The server is used to execute the method as described in any one of claims 12-15.

18. An electronic device, wherein, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-8.

19. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.

20. A computer program product comprising a computer program / computer-executable instructions, wherein, When the computer program / computer executable instructions are executed by a processor in an electronic device, the method of any one of claims 1-8 is implemented.