Enterprise information acquisition method and system, electronic equipment and storage medium

The client obtains the page link sample of the Yellow Pages website and sends a request. The server automatically obtains and structures the enterprise information. Combining the large language model and data source supplementary information, it solves the problem of low-efficiency in obtaining information on the Yellow Pages website and realizes efficient and comprehensive enterprise information acquisition.

CN120492755APending Publication Date: 2025-08-15HANGZHOU ALIBABA INT INTERNET IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510378506.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the enterprise information acquisition efficiency of yellow page websites is low, mainly relying on manual query, and the information type is single, so the inefficiency caused by the inefficiency caused by the differences in website structure is ineffective.

Method used

The client obtains the page link examples of the yellow page website and sends an information acquisition request to the server. The server automatically obtains and structures the enterprise information based on these link examples, and combines the large language model and preset data source supplementary information to realize automated enterprise information acquisition.

Benefits of technology

It improves the efficiency and comprehensiveness of enterprise information acquisition, realizes automated enterprise information collection and structured processing, reduces manual intervention, and improves the speed and integrity of information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492755A_ABST
    Figure CN120492755A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an enterprise information acquisition method, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: in response to an enterprise information acquisition operation, a client acquires a page link of a current list page, a next page link sample and a detail page link sample of a yellow page website, and takes the page link of the current list page, the next page link sample and the detail page link sample as request parameters; sending an information acquisition request to a preset server, the information acquisition request being used for triggering the server to acquire enterprise information of an enterprise displayed on the list page of the yellow page website based on the request parameter; the client displays execution information of the information acquisition task, and the execution information of the information acquisition task comprises one or more of site information of the yellow page website, the acquisition progress of the enterprise information and the execution state of the task. According to the method, the enterprise information is automatically collected and subjected to structured processing based on the yellow page website, and the acquisition efficiency of the enterprise information is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method for acquiring enterprise information, a system for acquiring enterprise information, an electronic device, a storage medium, and a computer program product. Background Art

[0002] The Yellow Pages is an internationally recognized directory of business telephone numbers organized by business nature and product category. Prior art standard Yellow Pages websites display summary information for listed businesses on their listing pages, where this information includes at least the company name. For businesses with complete listing information, the listing page may also include a company profile and address. Because the structure of Yellow Pages websites can vary significantly, and the types of company information within them are limited, prior art utilizes Yellow Pages websites inefficiently, primarily relying on manual queries to obtain company information.

[0003] It can be seen that a method is needed to automatically obtain enterprise information based on the Yellow Pages. Summary of the Invention

[0004] The embodiment of the present application provides a method for obtaining enterprise information, which can automatically obtain enterprise information based on the enterprise yellow pages, which can not only improve the efficiency of obtaining enterprise information, but also improve the comprehensiveness of the obtained information.

[0005] Correspondingly, the embodiments of the present application also provide an enterprise information acquisition system, an electronic device, a storage medium and a computer program product to ensure the implementation and application of the above-mentioned enterprise information acquisition method.

[0006] In order to solve the above problems, the present application discloses a method for obtaining enterprise information, which is applied to a client. The method includes:

[0007] In response to the enterprise information acquisition operation, a page link of the current list page, a next page link sample, and a detail page link sample of the yellow page website are acquired;

[0008] An information acquisition request is sent to a preset server using the page link of the current list page, the next page link sample, and the details page link sample as request parameters. The information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website;

[0009] The execution information of the information acquisition task is displayed, wherein the execution information includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task.

[0010] The present application discloses a method for obtaining enterprise information, which is applied to a server. The method includes:

[0011] In response to an information acquisition request sent by a client, the request parameters of the information acquisition request are parsed to obtain a page link of a current list page of the yellow pages website, a next page link sample, and a details page link sample;

[0012] Creating and executing an information acquisition task corresponding to the information acquisition request, the information acquisition task is used to: acquire enterprise information displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample, and the details page link sample;

[0013] Sending the execution information of the information acquisition task to the client so that the client displays the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, acquisition progress of the enterprise information, and execution status of the information acquisition task.

[0014] The embodiment of the present application discloses an enterprise information acquisition system, including a client and a server, wherein:

[0015] The client is used to execute the enterprise information acquisition method as described above;

[0016] The server is used to execute the enterprise information acquisition method as described above.

[0017] The embodiment of the present application further discloses a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method described in the embodiment of the present application.

[0018] An embodiment of the present application further discloses a computer program product, including a computer program / computer executable instructions, characterized in that when the computer program / computer executable instructions are executed by a processor in an electronic device, the method described in the embodiment of the present application is implemented.

[0019] Compared with the prior art, the embodiments of the present application have the following advantages:

[0020] In response to the enterprise information acquisition operation, the client obtains the page link, next page link sample and details page link sample of the current list page of the yellow pages website, and uses the page link of the current list page, the next page link sample and the details page link sample as request parameters to send an information acquisition request to the preset server. The information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website; then, display the execution information of the information acquisition task, wherein the execution information of the information acquisition task includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the task, thereby realizing automatic collection and structured processing of enterprise information based on the yellow pages website, effectively improving the efficiency of obtaining enterprise information. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is one of the step flow charts of the enterprise information acquisition method disclosed in the embodiment of this application;

[0022] Figure 2 This is a schematic diagram of a list page of a standard yellow pages website in the prior art;

[0023] Figure 3 This is a schematic diagram of the implementation architecture of the enterprise information acquisition method disclosed in the embodiment of this application;

[0024] Figure 4 This is one of the client interface diagrams in the enterprise information acquisition method disclosed in the embodiment of this application;

[0025] Figure 5 This is the second step flow chart of the enterprise information acquisition method disclosed in the embodiment of this application;

[0026] Figure 6 This is the third step flow chart of the enterprise information acquisition method disclosed in the embodiment of this application;

[0027] Figure 7 This is a schematic diagram of the interaction process of the enterprise information acquisition system disclosed in the embodiment of this application;

[0028] Figure 8 It is a structural diagram of an exemplary device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0030] In the prior art, standard yellow pages websites display summary information of included companies on the list page, and the list page also includes page turning links. The list page display effect is as follows: Figure 2 As shown. The enterprise information list is used to display the included information of each enterprise. In the embodiment of the present application, by parsing the page files of the list page and the details page of the standard yellow pages website, the page files of all the list pages of the yellow pages website are obtained, and the page content in the page files is analyzed using a large language model to automatically extract the enterprise information included in each page of the yellow pages website. On the other hand, combined with searching for supplementary information from a preset data source, the enterprise information is further improved to obtain more comprehensive enterprise information.

[0031] In some optional embodiments, the enterprise information acquisition method disclosed in the embodiments of this application can be implemented through an enterprise information acquisition system. The implementation architecture of the enterprise information acquisition system is as follows: Figure 3 As shown in the figure, the enterprise information acquisition system provides off-site data sources by accessing search engine services and database access services containing enterprise information. It also extracts hyperlinks and text content from list pages by accessing hyperlink extraction tools and page text extraction tools. By accessing the information analysis capabilities of a large language model, it can identify links to the next page and details page, and analyze, infer, and structure page text content.

[0032] The enterprise information acquisition system provided by the embodiment of the present application realizes the acquisition of enterprise information through two agents on the server side, namely, an information recognition agent and an information search agent. Optionally, the information recognition agent is used to identify the next page link and the details page link, and automatically obtain the page file of each list page based on the next page link, and identify the details page link in the page file, and then automatically obtain the page content of the details page corresponding to each details page link, and then identify the structured information of the corresponding enterprise from the page content. Optionally, the information search agent is used to search for enterprise information of various enterprises outside the website (such as a standard yellow pages website) to improve the enterprise information extracted from the website. On the other hand, it is also used to infer and generate other information of the enterprise based on the existing enterprise information, such as the industry to which the enterprise belongs, whether it is an e-commerce company, whether it is a foreign trade company, etc.

[0033] By using the enterprise information acquisition method disclosed in the embodiment of this application to analyze overseas enterprise yellow pages websites, efficient mining of overseas potential customers can be achieved.

[0034] like Figure 1 As shown, the enterprise information acquisition method applied to the client of the enterprise information acquisition system includes: steps 102 to 106.

[0035] Step 102: In response to the enterprise information acquisition operation, a page link of the current list page, a next page link sample, and a detail page link sample of the yellow page website are acquired.

[0036] In some optional embodiments, an information input interface may be provided on the client for the user to input the search start list page of the yellow pages website to be searched, the next page link sample of the yellow pages website, and the detail page link sample. Figure 4 Taking the client interface shown as an example, a user can enter a hyperlink to the start list page (e.g., the homepage) of the yellow pages website currently being searched in edit box 402, enter a hyperlink to the next page of any list page of the yellow pages website in edit box 404, and enter a hyperlink to any next detail page of the yellow pages website in edit box 406. After the user triggers the enterprise information acquisition operation through the client interface, the client obtains the hyperlinks entered in each edit box, uses the hyperlink in edit box 402 as the page link of the current list page, uses the hyperlink in edit box 404 as a next page link example, and uses the hyperlink in edit box 406 as a detail page link example.

[0037] Step 104, using the page link of the current list page, the next page link sample and the details page link sample as request parameters, sends an information acquisition request to the preset server, and the information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website.

[0038] The client generates an information acquisition request using the page link of the current list page, the next page link example, and the details page link example as request parameters, and sends the information acquisition request to the server of the enterprise information acquisition system. After receiving the information acquisition request, the server parses the information acquisition request and obtains the request parameters, thereby obtaining the page link of the current list page, the next page link example, and the details page link example. The server then creates and executes an enterprise information crawling task. During execution, the enterprise information crawling task retrieves the enterprise information of the enterprises listed on the yellow pages website based on the page link of the current list page, the next page link example, and the details page link example.

[0039] Optionally, obtaining the enterprise information of the enterprise displayed on the list page of the yellow pages website based on the request parameters includes: sub-step 1201, sub-step 1202 and sub-step 1203.

[0040] Sub-step 1201, based on the page link of the current list page, the next page link sample and the detail page link sample, the page files of each page of the yellow pages website are analyzed layer by layer to obtain the structured information of the enterprise displayed on each page.

[0041] In the prior art, the page files of the yellow page website list page and detail page use HTML (HyperTextMarkup Language) instructions to describe the composition of the page. The tag defines a hyperlink, used to link from the current page to another HTML page. For example, linking from the current list page to the company details page, the next list page, the previous list page, or a specific list page. Based on this, by parsing the page files of each list page, we can not only obtain the company details page link for the company included in the current list page, but also the page link for the next list page, as well as the text content of each page. Based on the text content of each company details page, we can extract the company information.

[0042] Optionally, based on the page link of the current list page, the next page link sample and the details page link sample, the page files of each page of the yellow pages website are analyzed layer by layer to obtain structured information of the enterprises displayed on each page, including: sub-steps S1 to S5.

[0043] Sub-step S1, initializing the current page link with the page link of the current list page.

[0044] Sub-step S2: obtaining the page file of the current list page based on the current page link.

[0045] For the specific implementation method of obtaining the page file of the current list page based on the current page link, please refer to the prior art and will not be repeated in the embodiments of this application.

[0046] Sub-step S3, referring to the next page link sample and the details page link sample, identifying the details page link and the next page link in the page file.

[0047] The page file of the current list page uses HTML instructions to describe one or more of the following information: a text description of the company displayed on the current page and a link to the details page displaying the company information, a page link to the next list page, a page link to the previous list page, and a page link to a specified list page.

[0048] Optionally, referring to the next page link sample and the details page link sample to identify the details page link and the next page link in the page file includes: obtaining a list of hyperlinks in the page file; taking the next page link sample, the details page link sample and the list of hyperlinks as input, triggering a preset large language model to refer to the next page link sample and the details page link sample to respectively identify the details page link and the next page link in the list.

[0049] In some optional embodiments, a prompt word can be generated based on the next page link sample and the content of the page file, triggering a preset large language model to refer to the next page link sample and identify the next page link in the content of the page file; and, a prompt word can be generated based on the details page link sample and the content of the page file, triggering a preset large language model to refer to the details page link sample and identify the details page link in the content of the page file.

[0050] In some other optional embodiments, in order to improve the recognition accuracy of the next page link and the details page link, the The tag extracts all hyperlinks in the page file and obtains a list of hyperlinks in the current page file. Afterwards, the large model is triggered to identify the next page link and the details page link from the list of hyperlinks. For example, based on the next page link sample and the list of hyperlinks, a prompt word is generated, and the preset large language model is triggered to refer to the next page link sample to identify the next page link in the list, and based on the details page link sample and the list of hyperlinks, a prompt word is generated, and the preset large language model is triggered to refer to the details page link sample to identify the details page link in the list.

[0051] In some optional embodiments, the following prompt template may be used to generate prompt words for identifying the next page link:

[0052]

[0053] In the prompt word template, "#Task" represents the task that the large language model is required to perform, "#Requirements" represents the additional requirements for the large language model to perform the task, and "#Input" represents the source data for performing the task. Among them, "$!{links}" is the hyperlink variable to be identified, and "$!{example}" represents the reference sample variable. When generating prompt words based on the prompt word template, "$!{links}" uses the list of hyperlinks in the current page, and "$!{example}" is replaced with the next page link sample entered by the user.

[0054] In some optional embodiments, the following prompt template may be used to generate prompt words for identifying detail page links:

[0055]

[0056] In the prompt word template, "#Task" represents the task that the large language model is required to perform, "#Requirements" represents the additional requirements for the large language model to perform the task, and "#Input" represents the source data for performing the task. Among them, "$!{links}" represents the hyperlink variable to be identified, and "$!{example}" represents the reference example variable. When generating prompt words based on the prompt word template, "$!{links}" is replaced with the list of hyperlinks in the current page.

[0057] "$!{example}" is replaced with the sample link to the detail page entered by the user.

[0058] Directly calling the large language model to extract detail page links or next page links from all hyperlinks on the page has a low accuracy rate. In the embodiment of the present application, by inputting a small number of sample prompts and giving the large language model a few samples, the accuracy rate of extracting detail page links or next page links can be improved.

[0059] Sub-step S4: analyzing the text content in the enterprise information detail page corresponding to each detail page link to obtain the structured information of each enterprise included in the current list page.

[0060] Optionally, the analysis of the text content in the enterprise information details page corresponding to each of the details page links to obtain the structured information of each enterprise included in the current list page includes: for each of the details page links, respectively obtaining the page file corresponding to the details page link; parsing the page file of the enterprise information details page corresponding to each of the details page links to obtain the text content in the enterprise information details page; using a preset large language model to analyze and identify the text content to obtain the structured information of the enterprise corresponding to the enterprise information details page; integrating the structured information of the enterprise corresponding to each of the enterprise information details pages to obtain the structured information of each enterprise included in the current list page.

[0061] The structured information includes but is not limited to the following: company name, mailing address, contact person, contact number, country, creation time, etc.

[0062] The specific implementation method of obtaining the page file corresponding to the detail page link, parsing the page file of the enterprise information detail page corresponding to each detail page link, and obtaining the text content in the enterprise information detail page can be found in the prior art and will not be repeated in the embodiments of this application.

[0063] After obtaining the text content corresponding to each enterprise information details page, a prompt word can be generated based on the preset prompt word template and the text content corresponding to the current enterprise information details page, and the preset large language model can be called based on the generated prompt word to trigger the preset large language model to analyze and identify the text content, extract the specified enterprise information from it, and output it according to the preset format, so as to obtain the structured information of the enterprise corresponding to the enterprise information details page.

[0064] Optionally, the prompt includes, but is not limited to, the following: a task description text instructing the large language model to extract and structure detailed information about a specified company from the given webpage content, a description of the specified company details, an output information format, information extraction rules, and the given webpage content. The specified company details include, but are not limited to, one or more of the following: company name, website, phone number, mailing address, country code, annual budget, company type, contact person, email address, etc.

[0065] The specific content of the structured information is determined based on specific needs and is not limited in the embodiments of the present application. In some optional embodiments, the business information obtained from the yellow pages website that cannot be parsed can be output in a specific format, for example, as an empty string.

[0066] After executing sub-steps S2, S3 and S4, the structured information of each enterprise displayed in the current list page can be obtained.

[0067] Sub-step S5, in response to successfully identifying the next page link, uses the next page link as the current page link, jumps to the step of executing the step of obtaining the page file of the current list page based on the current page link, and then analyzes the text content in the enterprise information details page corresponding to each of the details page links to obtain the structured information of each enterprise included in the current list page.

[0068] As mentioned above, for non-final list pages, the page includes a link to the next list page, so the large language model outputs a next-page link. However, for the final list page, the page does not include a link to the next list page, so the next-page link output by the large language model is empty or invalid.

[0069] When the large language model outputs a valid next page link, it means that the current list is not the last list page, and it is necessary to continue processing the next list page and obtain the page file of the next list page. Then, sub-steps S3 and S4 are executed.

[0070] In some optional embodiments, after analyzing the text content in the enterprise information details page corresponding to each of the details page links to obtain the structured information of each enterprise included in the current list page, it also includes: in response to not identifying the next page link, ending the progressive analysis operation on the yellow pages website.

[0071] On a yellow pages website, the last listing page doesn't have a hyperlink pointing to the next page. Therefore, when the large language model analyzes and recognizes the page file for the last listing page, it can't identify the hyperlink pointing to the next page. In other words, if no next-page link is identified in the page file, the layer-by-layer analysis of the yellow pages website is considered complete.

[0072] Sub-step 1202: performing a search operation of a preset data source based on the structured information to obtain the completed information of the corresponding enterprise.

[0073] By analyzing the page files of the yellow pages website to obtain some corporate information of the companies displayed on the pages of the yellow pages website and generating structured information, we further obtain the company's information that is not obtained from the yellow pages website through other channels as a supplement to the company information.

[0074] In some optional embodiments, the structured information includes: company name, and the search operation of the preset data source is performed based on the structured information to obtain the completed information of the corresponding company, including: obtaining a description of the missing information of each company based on the structured information; searching the preset data source based on the company name and the description of the missing information to obtain the completed information of the corresponding company, wherein the preset data source includes one or more of the following: a preset search engine, a database of a preset customer management system.

[0075] The missing information may be enterprise information whose value in the structured information is an empty string; the description of the missing information may be the field name or description of an enterprise information field whose value in the structured information is an empty string. Alternatively, the missing information of an enterprise may be determined based on the format of the structured information output by the large language model.

[0076] For example, when the large language model extracts the company name of a company named nameA from the details page but fails to extract the contact information of company A, the value of the company name field in the structured information is "nameA", and the value of the contact field is "" (that is, an empty string). When obtaining the completed information, you can use "nameA contact" as a keyword to search the preset search engine to obtain the contact recall results of company nameA returned by the search engine. Then, you can use the preset large language model to extract the contact information from the recall results returned by the search engine. Alternatively, when obtaining the completed information, you can use "nameA" as a parameter to call the customer information query interface of the database of the preset customer management system to obtain the contact information of company nameA.

[0077] In some optional embodiments, the structured information includes: company names, and the search operation of a preset data source is performed based on the structured information to obtain the completed information of the corresponding company, including: obtaining a description of the missing information of each company based on the structured information; searching a preset data source based on the company name and the description of the missing information to obtain search results of the corresponding company, wherein the preset data source includes: a database of a preset search engine and / or a preset customer management system; and using a preset large language model to infer the search results to obtain the completed information of the corresponding company.

[0078] Based on the structured information, a description of the missing information of each of the enterprises is obtained, and based on the enterprise name and the description of the missing information, a preset data source is searched to obtain the search results of the corresponding enterprises. For the specific implementation method, please refer to the above description and will not be repeated here. After obtaining the search results, for some enterprise information that cannot be obtained in the search results, such as the continent and city to which the enterprise belongs, the industry described by the enterprise, the type of enterprise, etc., it is necessary to infer based on the search results. For example, after obtaining the enterprise introduction based on the enterprise name of the target enterprise, the enterprise description can be further used as input data to generate prompt words according to a preset template, wherein the preset template includes candidate industries. Afterwards, the large language model is called based on the prompt words to trigger the large language model to infer the candidate industry matching the enterprise based on the enterprise name and the enterprise introduction.

[0079] By searching for enterprise information in preset data sources outside the Yellow Pages website, the enterprise information in the Yellow Pages website is supplemented, and by inferring the specified enterprise information based on the enterprise introduction obtained by the search, the channels for obtaining enterprise information are expanded, which helps to obtain more comprehensive and complete enterprise information.

[0080] Sub-step 1203: Based on the structured information and the supplementary information of each enterprise, obtain the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website.

[0081] The unassigned information fields in the structured information are assigned values using the supplementary information of the enterprise to obtain the complete structured information of the enterprise.

[0082] In other optional embodiments, the step of obtaining the enterprise information of the enterprise displayed on the list page of the yellow pages website based on the request parameters includes: initializing the current page link with the page link of the current list page; obtaining the page file of the current list page based on the current page link; identifying the detail page link and the next page link in the page file with reference to the next page link sample and the detail page link sample; analyzing the text content in the enterprise information detail page corresponding to each detail page link to obtain the structured information of each enterprise included in the current list page; performing a search operation of a preset data source based on the structured information to obtain the supplementary information of the corresponding enterprise; obtaining the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website based on the structured information and the supplementary information of each enterprise; in response to successfully identifying the next page link, using the next page link as the current page link, jumping to the step of obtaining the page file of the current list page based on the current page link to the step of obtaining the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website based on the structured information and the supplementary information of each enterprise. The specific implementation of each step can be found in the relevant description in the previous embodiment, which will not be repeated here.

[0083] Those skilled in the art should understand that the step of performing a search operation of a preset data source based on the structured information to obtain the completed information of the corresponding enterprise, and the step of obtaining the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website based on the structured information and the completed information of each enterprise, can be performed after sub-step S4 and is not limited to one execution order.

[0084] Afterwards, the server uses the obtained supplementary information to fill in the enterprise information that is missing from the structured information, obtains the final structured enterprise information of each enterprise, and sends the final enterprise information to the client.

[0085] Step 106: Display the execution information of the information acquisition task, wherein the execution information includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task.

[0086] The progress of obtaining the enterprise information includes: the total number of enterprises included in the yellow pages website and the total number of enterprises whose enterprise information has been obtained; the execution status of the information acquisition task is used to indicate whether the task has been completed.

[0087] After the client sends an information acquisition request to the preset server, the server creates an information acquisition task for the information acquisition request and executes the information acquisition task. Afterwards, the client receives the progress information of the server executing the information acquisition task and displays the progress information and the site information of the yellow pages website in real time. The display effect is as follows: Figure 4 shown.

[0088] Optional, such as Figure 5 As shown, after displaying the execution information of the information acquisition task, the method further includes: step 108.

[0089] Step 108 : In response to the execution record query operation of the information acquisition task, a list of enterprise information of each enterprise acquired by the server is displayed.

[0090] In some optional embodiments, the server stores an execution record of the information acquisition task executed for each information acquisition request, wherein the execution record includes: the task execution time, the corresponding yellow pages website, the execution result, etc. The execution result includes: the enterprise information displayed on the list page of the yellow pages website, wherein the enterprise information includes: the enterprise information extracted from the page of the yellow pages website, and the enterprise information searched and inferred from the preset data source.

[0091] Optionally, in response to an execution record query operation of the information acquisition task, a list of enterprise information of each of the enterprises obtained by the server is displayed, including: in response to an execution record query operation of the information acquisition task, an execution record query request is sent to the server, the execution record query request carries an execution record identifier, the execution record query request enables the server to send the execution result of the information acquisition task corresponding to the execution record identifier to the client, and the execution result includes: the enterprise information of the enterprise displayed on the list page of the yellow pages website; the client displays a list of enterprise information of the enterprise.

[0092] When the user clicks on the execution record of the information acquisition task of the information acquisition request displayed on the client, the client further obtains the execution result corresponding to the execution record clicked by the user, and displays a list of enterprise information on the client based on the execution result.

[0093] In summary, the enterprise information acquisition method disclosed in the embodiment of the present application obtains the page link, next page link sample and details page link sample of the current list page of the yellow pages website in response to the enterprise information acquisition operation through the client, and uses the page link of the current list page, the next page link sample and the details page link sample as request parameters to send an information acquisition request to a preset server. The information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website; then, display the execution information of the information acquisition task, wherein the execution information of the information acquisition task includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the task, thereby realizing automatic collection and structured processing of enterprise information based on the yellow pages website, effectively improving the efficiency of acquiring enterprise information.

[0094] Furthermore, by searching preset data sources and making inferences based on search results, supplementary information is obtained, and the supplementary information is used to complete the missing enterprise information in the yellow pages website, thereby effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0095] Based on the above embodiment, this embodiment also provides a method for obtaining enterprise information applied to the server. Figure 6 As shown, the method includes: steps 602 to 606.

[0096] Step 602: In response to the information acquisition request sent by the client, the request parameters of the information acquisition request are parsed to obtain the page link of the current list page of the yellow page website, the next page link sample, and the details page link sample.

[0097] The specific implementation method of the client sending the information acquisition request can be found in the relevant description in the previous embodiment and will not be repeated here.

[0098] Afterwards, the server parses the request parameters of the information acquisition request according to the preset communication protocol between the client and the server, and obtains the page link of the current list page of the yellow page website, the next page link sample and the details page link sample.

[0099] Step 604, create and execute the information acquisition task corresponding to the information acquisition request, and the information acquisition task is used to: obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample.

[0100] Optionally, the server can only create one information acquisition task for each information acquisition request and start executing the information acquisition task.

[0101] Optionally, the obtaining of enterprise information of enterprises displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample includes: based on the page link of the current list page, the next page link sample and the details page link sample, progressively analyzing the page files of each page of the yellow pages website layer by layer to obtain structured information of the enterprises displayed on each said page; performing a search operation of a preset data source based on the structured information to obtain complete information of corresponding enterprises; and obtaining the enterprise information of corresponding enterprises displayed on the list page of the yellow pages website based on the structured information and the complete information of each said enterprise.

[0102] The specific implementation methods of the server's steps of obtaining the enterprise information displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample can be found in the relevant description in the previous embodiment and will not be repeated here.

[0103] Step 606, sending the execution information of the information acquisition task to the client, so that the client displays the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, acquisition progress of the enterprise information, and execution status of the information acquisition task.

[0104] During the execution of the information acquisition task, the server can push the execution information of the information acquisition task to the client in real time based on the progress of the information acquisition task in acquiring enterprise information. For example, it can push the acquisition progress of the enterprise information and the execution status of the information acquisition task for synchronous display by the client. On the other hand, the server will store the execution record of the information acquisition task. The execution record includes: the task execution time, the corresponding yellow pages website, the execution result, etc. Among them, the execution result includes: the enterprise information of each enterprise extracted from the page of the yellow pages website, and the enterprise information of each enterprise searched and inferred from the preset data source.

[0105] Optionally, after sending the execution information of the information acquisition task to the client, the method further includes: in response to an execution record query request sent by the client, obtaining an execution record identifier carried in the execution record query request; and sending the execution result of the information acquisition task corresponding to the execution record identifier to the client, the execution result including: enterprise information displayed on the list page of the yellow pages website obtained by executing the information acquisition task. The execution record identifier corresponds one-to-one to the information acquisition task, and the server can retrieve the execution record and execution result of the corresponding information acquisition task based on the execution record identifier.

[0106] The method for generating the execution record query request is described in the prior art and will not be described in detail in the embodiments of this application.

[0107] In summary, the enterprise information acquisition method disclosed in the embodiment of the present application is that the server responds to the information acquisition request sent by the client, parses the request parameters of the information acquisition request, and obtains the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample; then, creates and executes the information acquisition task corresponding to the information acquisition request, and the information acquisition task is used to: based on the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website; next, the server sends the execution information of the information acquisition task to the client, so that the client displays the execution information, wherein the execution information includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task, thereby realizing automatic collection of enterprise information and structured processing based on the yellow pages website, effectively improving the efficiency of acquiring enterprise information.

[0108] Furthermore, by searching preset data sources and making inferences based on search results, supplementary information is obtained, and the supplementary information is used to complete the missing enterprise information in the yellow pages website, thereby effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0109] Based on the above embodiment, this embodiment also provides a system for obtaining enterprise information. The system for obtaining enterprise information of yellow pages includes a client and a server. Figure 7 As shown, the specific implementation process of the enterprise information acquisition system includes: steps 702 to 718.

[0110] In step 702 , the client obtains a page link of the current list page, a next page link sample, and a details page link sample of the yellow page website in response to the enterprise information acquisition operation.

[0111] In step 704, the client sends an information acquisition request to a preset server using the page link of the current list page, the next page link sample, and the details page link sample as request parameters.

[0112] The information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website.

[0113] In step 706, the server responds to the information acquisition request sent by the client, parses the request parameters of the information acquisition request, and obtains the page link of the current list page of the yellow page website, the next page link sample, and the details page link sample.

[0114] Step 708, the server creates and executes the information acquisition task corresponding to the information acquisition request, and the information acquisition task is used to: obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample.

[0115] Step 710: The server sends the execution information of the information acquisition task to the client.

[0116] The execution information includes one or more of the following: site information of the yellow pages website, acquisition progress of the enterprise information, and execution status of the information acquisition task.

[0117] Step 712: The client displays the execution information of the information acquisition task.

[0118] Step 714: The client sends an execution record query request to the server in response to the execution record query operation of the information acquisition task.

[0119] The execution record query request carries an execution record identifier.

[0120] Step 716: In response to the execution record query request, the server sends the execution result of the information acquisition task corresponding to the execution record identifier to the client, wherein the execution result includes: the enterprise information displayed on the list page of the yellow pages website.

[0121] Step 718: The client displays a list of enterprise information of the enterprise.

[0122] The specific implementation methods for the client and server to perform the above steps can be found in the specific implementation methods in the previous embodiments, which will not be repeated here.

[0123] In summary, in the enterprise information acquisition system disclosed in the embodiment of the present application, the client responds to the enterprise information acquisition operation, obtains the page link, next page link sample and details page link sample of the current list page of the yellow pages website, and then sends an information acquisition request to the preset server with the page link of the current list page, the next page link sample and the details page link sample as request parameters; the server responds to the information acquisition request sent by the client, creates and executes the information acquisition task corresponding to the information acquisition request, and obtains the enterprise information of the enterprise displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample; next, the server sends the execution information of the information acquisition task to the client, and the client displays the execution information, wherein the execution information includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task, thereby realizing automatic collection and structured processing of enterprise information based on the yellow pages website, effectively improving the efficiency of acquiring enterprise information.

[0124] Furthermore, by searching preset data sources and making inferences based on search results, supplementary information is obtained, and the supplementary information is used to complete the missing enterprise information in the yellow pages website, thereby effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0125] Based on the above embodiment, this embodiment further provides an enterprise information acquisition device, which is applied to a client, and includes:

[0126] An input link acquisition module is used to obtain a page link of a current list page, a next page link sample, and a detail page link sample of a yellow page website in response to an enterprise information acquisition operation;

[0127] The enterprise information acquisition module is used to send an information acquisition request to a preset server using the page link of the current list page, the next page link sample, and the details page link sample as request parameters. The information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website;

[0128] The display module is used to display the execution information of the information acquisition task, wherein the execution information includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task.

[0129] Optionally, obtaining the enterprise information of the enterprise displayed on the list page of the yellow pages website based on the request parameter includes:

[0130] Based on the page link of the current list page, the next page link sample, and the detail page link sample, the page files of each page of the yellow pages website are progressively analyzed layer by layer to obtain structured information of the enterprise displayed on each page;

[0131] Performing a search operation of a preset data source based on the structured information to obtain the completed information of the corresponding enterprise;

[0132] Based on the structured information and the supplementary information of each enterprise, the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website is obtained.

[0133] Optionally, based on the page link of the current list page, the next page link sample, and the detail page link sample, the page files of each page of the yellow pages website are progressively analyzed layer by layer to obtain structured information of the enterprise displayed on each page, including:

[0134] Initialize the current page link with the page link of the current list page;

[0135] Based on the current page link, obtain the page file of the current list page;

[0136] Referring to the next page link example and the details page link example, identifying the details page link and the next page link in the page file;

[0137] Analyze the text content of the enterprise information details page corresponding to each of the details page links to obtain structured information of each enterprise included in the current list page;

[0138] In response to successfully identifying the next page link, the next page link is used as the current page link, and the process jumps to executing the step of obtaining the page file of the current list page based on the current page link, and then to the step of analyzing the text content in the enterprise information details page corresponding to each of the details page links to obtain the structured information of each enterprise included in the current list page.

[0139] Optionally, the structured information includes: a company name, and performing a search operation of a preset data source based on the structured information to obtain the corresponding company's complete information includes:

[0140] Based on the structured information, obtaining a description of the missing information of each of the enterprises;

[0141] Based on the enterprise name and the description of the missing information, a preset data source is searched to obtain the complementary information of the corresponding enterprise, wherein the preset data source includes: a database of a preset search engine and / or a preset customer management system.

[0142] Optionally, the structured information includes: a company name, and performing a search operation of a preset data source based on the structured information to obtain the corresponding company's complete information includes:

[0143] Based on the structured information, obtaining a description of the missing information of each of the enterprises;

[0144] Based on the enterprise name and the description of the missing information, searching a preset data source to obtain search results of the corresponding enterprise, wherein the preset data source includes: a database of a preset search engine and / or a preset customer management system;

[0145] A preset large language model is used to infer the search results to obtain the corresponding enterprise's complementary information.

[0146] Optionally, after displaying the execution information of the information acquisition task, the device further includes:

[0147] In response to the execution record query operation of the information acquisition task, a list of enterprise information of each of the enterprises acquired by the server is displayed.

[0148] The enterprise information acquisition device disclosed in the embodiment of the present application is used to implement the above-mentioned enterprise information acquisition method. The specific implementation methods of each module of the device can be found in the specific implementation methods of the corresponding steps in the above-mentioned method embodiment, which will not be repeated here.

[0149] In summary, the enterprise information acquisition device disclosed in the example of this application responds to the enterprise information acquisition operation through the client, obtains the page link, next page link sample and details page link sample of the current list page of the yellow pages website, and uses the page link of the current list page, the next page link sample and the details page link sample as request parameters to send an information acquisition request to the preset server. The information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website; then, display the execution information of the information acquisition task, wherein the execution information of the information acquisition task includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the task, thereby realizing automatic collection and structured processing of enterprise information based on the yellow pages website, effectively improving the efficiency of acquiring enterprise information.

[0150] Furthermore, by searching preset data sources and making inferences based on search results, supplementary information is obtained, and the supplementary information is used to complete the missing enterprise information in the yellow pages website, thereby effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0151] Based on the above embodiment, this embodiment further provides an enterprise information acquisition device, which is applied to a server, and includes:

[0152] A page link acquisition module is used to respond to an information acquisition request sent by a client, parse the request parameters of the information acquisition request, and obtain a page link of the current list page of the yellow pages website, a next page link sample, and a detail page link sample;

[0153] An enterprise information acquisition module is configured to create and execute an information acquisition task corresponding to the information acquisition request, wherein the information acquisition task is configured to: acquire enterprise information of enterprises displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample, and the details page link sample;

[0154] An execution information sending module is used to send the execution information of the information acquisition task to the client, so that the client displays the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, acquisition progress of the enterprise information, and execution status of the information acquisition task.

[0155] Optionally, the obtaining of enterprise information of enterprises displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample includes: based on the page link of the current list page, the next page link sample and the details page link sample, progressively analyzing the page files of each page of the yellow pages website layer by layer to obtain structured information of the enterprises displayed on each said page; performing a search operation of a preset data source based on the structured information to obtain complete information of corresponding enterprises; and obtaining the enterprise information of corresponding enterprises displayed on the list page of the yellow pages website based on the structured information and the complete information of each said enterprise.

[0156] Optionally, based on the page link of the current list page, the next page link sample and the details page link sample, the page files of each page of the yellow pages website are analyzed step by step to obtain the structured information of the enterprises displayed on each page, including: initializing the current page link with the page link of the current list page; obtaining the page file of the current list page based on the current page link; identifying the details page link and the next page link in the page file with reference to the next page link sample and the details page link sample; analyzing the text content in the enterprise information details page corresponding to each of the details page links to obtain the structured information of each enterprise included in the current list page; in response to successfully identifying the next page link, using the next page link as the current page link, jumping to executing the step of obtaining the page file of the current list page based on the current page link to the step of analyzing the text content in the enterprise information details page corresponding to each of the details page links to obtain the structured information of each enterprise included in the current list page.

[0157] The enterprise information acquisition device disclosed in the embodiment of the present application is used to implement the above-mentioned enterprise information acquisition method. The specific implementation methods of each module of the device can be found in the specific implementation methods of the corresponding steps in the above-mentioned method embodiment, which will not be repeated here.

[0158] In summary, the enterprise information acquisition device disclosed in the example of this application, by the server responding to the information acquisition request sent by the client, parsing the request parameters of the information acquisition request, and obtaining the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample; then, creating and executing the information acquisition task corresponding to the information acquisition request, the information acquisition task is used to: based on the page link of the current list page of the yellow pages website, the next page link sample and the details page link sample, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website; next, the server sends the execution information of the information acquisition task to the client, so that the client displays the execution information, wherein the execution information includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task, thereby realizing automatic collection of enterprise information and structured processing based on the yellow pages website, effectively improving the efficiency of acquiring enterprise information.

[0159] Furthermore, by searching preset data sources and making inferences based on search results, supplementary information is obtained, and the supplementary information is used to complete the missing enterprise information in the yellow pages website, thereby effectively improving the comprehensiveness and completeness of enterprise information acquisition.

[0160] An embodiment of the present application further provides a non-volatile readable storage medium, which stores one or more modules (programs). When the one or more modules are applied to a device, the device can execute instructions (instructions) of each method step in the embodiment of the present application.

[0161] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the embodiment of the present application.

[0162] The present application also provides an electronic device comprising: a processor and a memory communicatively connected to the processor; the memory storing computer-executable instructions; and the processor executing the computer-executable instructions stored in the memory to implement the method described in the present application. In the present application, the electronic device includes a server, a terminal device, and other devices.

[0163] An embodiment of the present application further discloses a computer program product, including a computer program / computer executable instructions, characterized in that when the computer program / computer executable instructions are executed by a processor in an electronic device, the method described in the embodiment of the present application is implemented.

[0164] The embodiments of the present disclosure may be implemented as a device configured as desired using any appropriate hardware, firmware, software, or any combination thereof, and the device may include electronic devices such as a server (cluster), a terminal, etc. Figure 8 An exemplary apparatus 800 that can be used to implement various embodiments described in this application is schematically illustrated.

[0165] For one embodiment, Figure 8 An exemplary apparatus 800 is shown having one or more processors 802, a control module (chip set) 804 coupled to at least one of the processor(s) 802, a memory 806 coupled to the control module 804, a non-volatile memory (NVM) / storage device 808 coupled to the control module 804, one or more input / output devices 810 coupled to the control module 804, and a network interface 812 coupled to the control module 804.

[0166] The processor 802 may include one or more single-core or multi-core processors, and the processor 802 may include any combination of general-purpose processors or dedicated processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, the apparatus 800 can serve as a server, terminal, or other device described in the embodiments of the present application.

[0167] In some embodiments, the apparatus 800 may include one or more computer-readable media (e.g., memory 806 or NVM / storage 808) having instructions 814 and one or more processors 802 configured in conjunction with the one or more computer-readable media to execute the instructions 814 to implement a module to perform the actions described in the present disclosure.

[0168] For one embodiment, the control module 804 may include any suitable interface controller to provide any suitable interface to at least one of the processor(s) 802 and / or any suitable device or component in communication with the control module 804 .

[0169] The control module 804 may include a memory controller module to provide an interface to the memory 806. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0170] The memory 806 can be used, for example, to load and store data and / or instructions 814 for the device 800. For one embodiment, the memory 806 can include any suitable volatile memory, such as a suitable DRAM. In some embodiments, the memory 806 can include double data rate type four synchronous dynamic random access memory (DDR4 SDRAM).

[0171] For one embodiment, the control module 804 may include one or more input / output controllers to provide interfaces to the NVM / storage device 808 and the input / output device(s) 810 .

[0172] For example, NVM / storage 808 may be used to store data and / or instructions 814. NVM / storage 808 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable non-volatile storage device(s) (e.g., one or more hard disk drives (HDDs), one or more compact disk (CD) drives, and / or one or more digital versatile disk (DVD) drives).

[0173] NVM / storage device 808 may include storage resources that are part of the device on which apparatus 800 is installed, or it may be accessible to the device without being part of the device. For example, NVM / storage device 808 may be accessed via input / output device(s) 810 over a network.

[0174] (One or more) input / output devices 810 may provide an interface for the apparatus 800 to communicate with any other appropriate device. The input / output device 810 may include a communication component, an audio component, a sensor component, etc. The network interface 812 may provide an interface for the apparatus 800 to communicate via one or more networks. The apparatus 800 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, for example, accessing a wireless network based on a communication standard such as Bluetooth, WiFi, 2G, 3G, 4G, 5G, etc., or a combination thereof for wireless communication.

[0175] For one embodiment, at least one of the processor(s) 802 may be packaged together with the logic of one or more controllers of the control module 804 (e.g., a memory controller module). For one embodiment, at least one of the processor(s) 802 may be packaged together with the logic of one or more controllers of the control module 804 to form a system-in-package (SiP). For one embodiment, at least one of the processor(s) 802 may be integrated on the same die with the logic of one or more controllers of the control module 804. For one embodiment, at least one of the processor(s) 802 may be integrated on the same die with the logic of one or more controllers of the control module 804 to form a system-on-chip (SoC).

[0176] In various embodiments, the apparatus 800 may be, but is not limited to, a terminal device such as a server, a desktop computing device, or a mobile computing device (e.g., a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.). In various embodiments, the apparatus 800 may have more or fewer components and / or a different architecture. For example, in some embodiments, the apparatus 800 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0177] Among them, the main control chip can be used as a processor or control module in the detection device, sensor data, location information, etc. are stored in the memory or NVM / storage device, the sensor group can be used as an input / output device, and the communication interface may include a network interface.

[0178] The present application also provides an electronic device comprising: a processor; and a memory storing executable code, wherein when the executable code is executed, the processor executes one or more methods described in the embodiments of the present application. The memory in the embodiments of the present application can store various data, such as target files, file-application association data, and other data, and can also include user behavior data, thereby providing a data foundation for various processing.

[0179] The embodiments of the present application further provide one or more machine-readable media on which executable codes are stored. When the executable codes are executed, the processor executes one or more methods described in the embodiments of the present application.

[0180] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0181] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0182] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0183] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0184] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0185] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0186] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0187] The above is a detailed introduction to an enterprise information acquisition method, an enterprise information acquisition system, an electronic device, a storage medium and a computer program product provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core ideas. At the same time, for general technical personnel in this field, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.

Claims

1. A method for obtaining enterprise information, applied to a client, characterized in that: The method comprises: In response to the enterprise information acquisition operation, a page link of the current list page, a next page link sample, and a detail page link sample of the yellow page website are acquired; An information acquisition request is sent to a preset server using the page link of the current list page, the next page link sample, and the details page link sample as request parameters. The information acquisition request is used to trigger the server to perform the following information acquisition task: based on the request parameters, obtain the enterprise information of the enterprise displayed on the list page of the yellow pages website; The execution information of the information acquisition task is displayed, wherein the execution information includes one or more of the following: the site information of the yellow pages website, the acquisition progress of the enterprise information, and the execution status of the information acquisition task.

2. The method according to claim 1, characterized in that The acquiring, based on the request parameters, enterprise information of the enterprise displayed on the list page of the yellow pages website includes: Based on the page link of the current list page, the next page link sample, and the detail page link sample, the page files of each page of the yellow pages website are progressively analyzed layer by layer to obtain structured information of the enterprise displayed on each page; Performing a search operation of a preset data source based on the structured information to obtain the completed information of the corresponding enterprise; Based on the structured information and the supplementary information of each enterprise, the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website is obtained.

3. The method according to claim 2, characterized in that The page files of each page of the yellow pages website are analyzed layer by layer based on the page link of the current list page, the next page link sample, and the details page link sample to obtain structured information of the enterprise displayed on each page, including: Initialize the current page link with the page link of the current list page; Based on the current page link, obtain the page file of the current list page; Referring to the next page link example and the details page link example, identifying the details page link and the next page link in the page file; Analyze the text content of the enterprise information details page corresponding to each of the details page links to obtain structured information of each enterprise included in the current list page; In response to successfully identifying the next page link, the next page link is used as the current page link, and the process jumps to executing the step of obtaining the page file of the current list page based on the current page link, and then to the step of analyzing the text content in the enterprise information details page corresponding to each of the details page links to obtain the structured information of each enterprise included in the current list page.

4. The method according to claim 2, characterized in that The structured information includes: a company name, and performing a search operation of a preset data source based on the structured information to obtain the corresponding company's supplementary information includes: Based on the structured information, obtaining a description of the missing information of each of the enterprises; Based on the enterprise name and the description of the missing information, a preset data source is searched to obtain the complementary information of the corresponding enterprise, wherein the preset data source includes: a database of a preset search engine and / or a preset customer management system.

5. The method according to claim 2, characterized in that The structured information includes: a company name, and performing a search operation of a preset data source based on the structured information to obtain the corresponding company's supplementary information includes: Based on the structured information, obtaining a description of the missing information of each of the enterprises; Based on the enterprise name and the description of the missing information, searching a preset data source to obtain search results of the corresponding enterprise, wherein the preset data source includes: a database of a preset search engine and / or a preset customer management system; A preset large language model is used to infer the search results to obtain the corresponding enterprise's complementary information.

6. The method according to claim 1, characterized in that After displaying the execution information of the information acquisition task, the method further includes: In response to the execution record query operation of the information acquisition task, a list of enterprise information of each of the enterprises acquired by the server is displayed.

7. A method for obtaining enterprise information, applied to a server, characterized in that: The method comprises: In response to an information acquisition request sent by a client, the request parameters of the information acquisition request are parsed to obtain a page link of a current list page of the yellow pages website, a next page link sample, and a details page link sample; Creating and executing an information acquisition task corresponding to the information acquisition request, the information acquisition task is used to: acquire enterprise information displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample, and the details page link sample; Sending the execution information of the information acquisition task to the client so that the client displays the execution information, wherein the execution information includes one or more of the following: site information of the yellow pages website, acquisition progress of the enterprise information, and execution status of the information acquisition task.

8. The method according to claim 7, characterized in that The acquiring of enterprise information displayed on the list page of the yellow pages website based on the page link of the current list page of the yellow pages website, the next page link sample, and the details page link sample includes: Based on the page link of the current list page, the next page link sample, and the detail page link sample, the page files of each page of the yellow pages website are progressively analyzed layer by layer to obtain structured information of the enterprise displayed on each page; Performing a search operation of a preset data source based on the structured information to obtain the completed information of the corresponding enterprise; Based on the structured information and the supplementary information of each enterprise, the enterprise information of the corresponding enterprise displayed on the list page of the yellow pages website is obtained.

9. An enterprise information acquisition system comprising: The client and server are characterized by: The client is configured to execute the method according to any one of claims 1 to 6; The server is used to execute the method according to any one of claims 7 to 8.

10. An enterprise information acquisition application, comprising: The client and server are characterized by: The client is configured to execute the method according to any one of claims 1 to 6; The server is used to execute the method according to any one of claims 7 to 8.

11. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.

13. A computer program product comprising a computer program / computer executable instructions, characterized in that When the computer program / computer executable instructions are executed by a processor in an electronic device, the method according to any one of claims 1 to 8 is implemented.