Commodity information retrieval method and device based on large language model and medium
Through the large language model and multimodal model, the content analysis and identification of web pages is solved, and the problems of low correlation and insufficient utilization of multimodal data in traditional product information retrieval methods are achieved, and efficient and intelligent product information extraction and storage are achieved.
Patent Information
- Application Number
- CN202510626972.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-02
AI Technical Summary
Traditional product information retrieval methods are difficult to understand the semantics and context of user queries, resulting in low correlation of search results and the inability to effectively extract and utilize product key information in multimodal data.
The content analysis and recognition of web pages is used to analyze and identify the content of the web pages through the large language model, and the relevance of web pages is initially judged through the large language model. The multi-modal model is used to extract key product information in pictures and videos, and convert it into a structured Markdown format to store it in the knowledge base.
It realizes efficient and intelligent product information retrieval, which can fully extract key product information in multimodal data, reduces information omissions and redundancy, and provides a solid analysis foundation.
Smart Images

Figure CN120578799A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and information retrieval technology, and in particular to a commodity information retrieval method, device and storage medium based on a large language model. Background Art
[0002] In today's digital age, the amount of product information on the Internet is exploding. Traditional product information retrieval methods rely primarily on keyword matching, which has many limitations:
[0003] On the one hand, it is difficult to understand the semantics and context of user queries, resulting in low relevance of retrieval results and users may need to spend a lot of time screening information;
[0004] On the other hand, traditional methods are often unable to effectively extract and utilize key product information contained in multimodal data such as pictures and videos.
[0005] With the development of artificial intelligence (AI), large language models (LLMs) have demonstrated powerful language understanding and generation capabilities, and multimodal models are also evolving, providing new approaches to solving these problems. However, there is currently no mature method that can fully utilize these technologies to achieve efficient and intelligent product information retrieval. Summary of the Invention
[0006] In view of the technical deficiencies mentioned in the background technology, the embodiment of the present invention aims to provide a product information retrieval method, device and storage medium based on a large language model.
[0007] To achieve the above objectives, in a first aspect, an embodiment of the present invention provides a product information retrieval method based on a large language model, comprising:
[0008] Receive a product query request input by a user; the product query request includes a product name or brand name;
[0009] Responding to the product query request to obtain search results, and obtaining top k web pages from the search results;
[0010] Analyze the content of the top k web pages to obtain title information and meta description information;
[0011] Filter the top k web pages based on the large language model, title information, meta description information and product query requests to obtain the pages to be parsed;
[0012] The page to be parsed is identified and analyzed using a multimodal model to obtain key product information; the key product information includes product selling points and product specifications.
[0013] As a specific implementation of this application, receiving a product query request input by a user is specifically as follows:
[0014] Configure the browser, install drivers and dependent libraries;
[0015] Writing program code to start the browser through a programming interface;
[0016] Utilize the browser to send a product query request to a network search engine.
[0017] As a specific implementation of this application, the title information and meta description information are specifically as follows:
[0018] Parse web pages and extract links from them;
[0019] The title information and meta description information of each web page are obtained according to the link.
[0020] As a specific implementation of this application, the page to be parsed is obtained as follows:
[0021] Based on the title information and meta description information, a large language model is used to preliminarily determine the relevance of any web page to the product name / brand name in the product query request;
[0022] Select pages with correlation higher than the threshold as pages to be parsed.
[0023] As a specific implementation of the present application, a large language model is used to preliminarily determine the relevance of any web page to the product name / brand name in the product query request, specifically:
[0024] The BGE Embedding model is used to convert the meta information and the search keywords in the product query request into vectors, and the similarity between the vectors is calculated to achieve a preliminary judgment of the relevance.
[0025] As a preferred implementation of the present application, after obtaining the page to be parsed, the method further includes:
[0026] Obtain the secondary links in the page to be parsed, obtain the webpage content corresponding to the secondary links, and perform preliminary screening on the webpage content.
[0027] As a specific implementation of this application, the key information of the product is obtained as follows:
[0028] Use the HTML parsing library to parse the text content of the page to be parsed and convert it into Markdown format;
[0029] A multimodal model is used to identify and analyze the images in the page to be parsed, and the product selling points and product specifications in the images are converted into text form, and integrated with the parsed text content; wherein, the converted Markdown format content is stored in the knowledge base as product brand knowledge.
[0030] In a second aspect, an embodiment of the present invention further provides a product information retrieval device based on a large language model, comprising:
[0031] A query input unit, configured to receive a product query request input by a user; the product query request includes a product name or a brand name;
[0032] A result response unit, configured to respond to the product query request to obtain search results, and obtain top k web pages from the search results;
[0033] A web page parsing unit is used to parse the content of the top k web pages to obtain title information and meta description information;
[0034] A webpage filtering unit is used to filter the top k webpages based on the large language model, title information, meta description information and product query request to obtain the pages to be parsed;
[0035] The multimodal analysis unit is used to use a multimodal model to identify and analyze the page to be analyzed to obtain key information of the product; the key information of the product includes the selling points and specifications of the product.
[0036] In a third aspect, an embodiment of the present invention further provides a product information retrieval device based on a large language model, comprising a processor, an input device, an output device, and a memory, wherein the processor, input device, output device, and memory are interconnected, wherein the memory is used to store a computer program, and the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method of the first aspect above.
[0037] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method of the first aspect.
[0038] By implementing the large language model-based product information retrieval method and device provided in the embodiments of the present invention, content analysis is performed on multiple web pages obtained based on a product query request, a preliminary correlation judgment is made on the multiple web pages using the large language model, and a multimodal analysis is performed using a multimodal model, ultimately obtaining key information such as product selling points and rules. This technical solution can effectively extract key product information contained in multimodal data such as images and videos, thereby realizing efficient and intelligent product information retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the specific implementation of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the specific implementation or the description of the prior art.
[0040] Figure 1 is a flowchart of a commodity information retrieval method based on a large language model provided by an embodiment of the present invention;
[0041] Figure 2 is another flow chart of the illustrated method;
[0042] Figure 3 1 is a structural diagram of a commodity information retrieval device based on a large language model provided by an embodiment of the present invention;
[0043] Figure 4 yes Figure 3 Another structural diagram of the device shown. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0045] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0046] Please refer to Figure 1 and Figure 2 The embodiment of the present invention provides a commodity information retrieval method based on a large language model, including:
[0047] S1, receiving a product query request input by a user.
[0048] In practice, first configure the Chrome browser's automated environment on the server or local device, installing the necessary drivers and dependent libraries. Then, write program code to launch the Chrome browser through the programming interface and send a search request containing the product name and brand information to the Baidu search engine. For example, if a user wants to search for information about "Suntory Oolong Tea," the program will send "Suntory Oolong Tea" as the keyword to the Baidu search engine.
[0049] S2, responding to the product query request to obtain search results, and obtaining top k web pages from the search results.
[0050] In specific implementation, after Baidu's search engine returns search results, the program uses web parsing tools such as BeautifulSoup to obtain the HTML page content of the top k (for example, k = 5) web pages, while removing inaccessible pages based on the response of the web pages. This process is achieved through network request and page parsing technology, ensuring that the obtained page content is complete and accurate.
[0051] S3, content analysis is performed on the top k web pages to obtain title information and meta description information.
[0052] In the specific implementation, the content of each web page is parsed and the links therein are extracted. For each link corresponding to the page, its title and meta description information are obtained.
[0053] S4, based on the large language model, title information, meta description information and product query request, the top k web pages are filtered to obtain the pages to be parsed.
[0054] In specific implementation, after obtaining the webpage content, a large language model is used to initially determine its relevance to the queried product brand information based on its title, meta description, and other information. If it is determined to be relevant, further content from that page is obtained; if the relevance is low, it is excluded. Secondary links on the page are also obtained. Some secondary links, such as "More Information," may contain more information about the product or brand. For each secondary link on the page, the corresponding webpage content is obtained, and the above process is repeated to obtain more comprehensive information. This process enables comprehensive and accurate collection of information related to product brands, greatly reducing information omissions and redundancies, and providing a solid foundation for subsequent in-depth analysis and utilization of product information.
[0055] For the preliminary determination of relevance, this embodiment adopts the following method:
[0056] The BGE Embedding model is used to convert page meta information and search keywords into embedding vectors respectively, and the similarity between the vectors is calculated.
[0057] Furthermore, the web pages are sorted from high to low by similarity, with pages with similarity above a certain threshold selected and those below it excluded. For pages with preliminary relevance, the presence of secondary links is checked. If so, the content of the secondary link pages is also obtained and preliminarily screened.
[0058] S5, use the multimodal model to identify and analyze the page to be parsed to obtain key product information.
[0059] In the specific implementation, it is necessary to perform multimodal analysis on the obtained web page content, specifically:
[0060] Markdown format is concise, easy to read, and easy to edit, and is convenient for storage and management. Therefore, in this embodiment, unstructured HTML text content is converted into structured Markdown format. Specifically, according to HTML tags and Markdown syntax rules, HTML elements such as titles, paragraphs, lists, and pictures are converted into corresponding Markdown formats. For example, <h1>Tags are converted to first-level headings represented by the # symbol in Markdown. Multimodal models are used to identify and analyze images on the page. Many key product information, such as selling points and specifications, is contained in images. Multimodal models can effectively extract this information and convert it into text. Image content is embedded in Markdown, and the converted Markdown content can be stored in the knowledge base as product brand knowledge, facilitating subsequent query, analysis, and utilization. Furthermore, by regularly updating product information in the knowledge base, the timeliness and accuracy of the information are ensured.
[0061] By implementing the large language model-based product information retrieval method provided in an embodiment of the present invention, content analysis is performed on multiple web pages obtained based on a product query request, a preliminary correlation judgment is made on the multiple web pages using the large language model, and a multimodal analysis is performed using a multimodal model, ultimately obtaining key information such as product selling points and rules. This technical solution can effectively extract key product information contained in multimodal data such as images and videos, thereby realizing efficient and intelligent product information retrieval.
[0062] Based on the same inventive concept, the present invention provides a commodity information retrieval device based on a large language model, such as Figure 3 Shown, including:
[0063] A query input unit, configured to receive a product query request input by a user; the product query request includes a product name or a brand name;
[0064] A result response unit, configured to respond to the product query request to obtain search results, and obtain top k web pages from the search results;
[0065] A web page parsing unit is used to parse the content of the top k web pages to obtain title information and meta description information;
[0066] A webpage filtering unit is used to filter the top k webpages based on the large language model, title information, meta description information and product query request to obtain the pages to be parsed;
[0067] The multimodal analysis unit is used to use a multimodal model to identify and analyze the page to be analyzed to obtain key information of the product; the key information of the product includes the selling points and specifications of the product.
[0068] The query input unit is specifically used to: configure the browser, install the driver and dependent libraries;
[0069] Writing program code to start the browser through a programming interface;
[0070] Utilize the browser to send a product query request to a network search engine.
[0071] Furthermore, the web page parsing unit is specifically used to:
[0072] Parse web pages and extract links from them;
[0073] The title information and meta description information of each web page are obtained according to the link.
[0074] Furthermore, the webpage filtering unit is specifically used to:
[0075] Based on the title information and meta description information, a large language model is used to preliminarily determine the relevance of any web page to the product name / brand name in the product query request;
[0076] Select pages with correlation higher than the threshold as pages to be parsed.
[0077] Furthermore, this embodiment uses the following method to determine the correlation:
[0078] The BGE Embedding model is used to convert the meta information and the search keywords in the product query request into vectors, and the similarity between the vectors is calculated to achieve a preliminary judgment of the relevance.
[0079] Furthermore, the webpage filtering unit is further configured to:
[0080] Obtain the secondary links in the page to be parsed, obtain the webpage content corresponding to the secondary links, and perform preliminary screening on the webpage content.
[0081] Furthermore, the multimodal parsing unit is specifically used to:
[0082] Use the HTML parsing library to parse the text content of the page to be parsed and convert it into Markdown format;
[0083] A multimodal model is used to identify and analyze the images in the page to be parsed, and the product selling points and product specifications in the images are converted into text form, and integrated with the parsed text content; wherein, the converted Markdown format content is stored in the knowledge base as product brand knowledge.
[0084] It should be noted that the specific workflow of this embodiment can be found in the aforementioned method embodiment section and will not be described in detail here.
[0085] Furthermore, another embodiment of the present invention provides a commodity information retrieval device based on a large language model. Figure 4 As shown, the apparatus may include: one or more processors 101, one or more input devices 102, one or more output devices 103, and a memory 104. The processors 101, input devices 102, output devices 103, and memory 104 are interconnected via a bus 105. The memory 104 is used to store a computer program, which includes program instructions. The processor 101 is configured to call the program instructions to execute the method of the above method embodiment.
[0086] It should be understood that in the embodiment of the present invention, the processor 101 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0087] The input device 102 may include a keyboard, etc., and the output device 103 may include a display (LCD, etc.), a speaker, etc.
[0088] The memory 104 may include a read-only memory and a random access memory, and provides instructions and data to the processor 101. A portion of the memory 104 may also include a non-volatile random access memory. For example, the memory 104 may also store device type information.
[0089] In a specific implementation, the processor 101, input device 102, and output device 103 described in the embodiment of the present invention can execute the implementation method described in the embodiment of the commodity information retrieval method based on a large language model provided by the embodiment of the present invention, which will not be repeated here.
[0090] Accordingly, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, a product information retrieval method based on a large language model is implemented.
[0091] The computer-readable storage medium may be an internal storage unit of the system described in any of the aforementioned embodiments, such as a hard disk or memory of the system. The computer-readable storage medium may also be an external storage device of the system, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the system. Furthermore, the computer-readable storage medium may also include both an internal storage unit of the system and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the system. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.
[0092] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0093] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or can be electrical, mechanical or other forms of connection.
[0094] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the embodiments of the present invention.
[0095] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0096] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0097] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.< / h1>
Claims
1. A commodity information retrieval method based on a large language model, characterized in that: include: Receive product query requests input by users; The product query request includes a product name or a brand name; Responding to the product query request to obtain search results, and obtaining top k web pages from the search results; Analyze the content of the top k web pages to obtain title information and meta description information; Filter the top k web pages based on the large language model, title information, meta description information and product query requests to obtain the pages to be parsed; The page to be parsed is identified and analyzed using a multimodal model to obtain key product information; the key product information includes product selling points and product specifications.
2. The product information retrieval method according to claim 1, wherein: The specific product query request received from the user is: Configure the browser, install drivers and dependent libraries; Writing program code to start the browser through a programming interface; Utilize the browser to send a product query request to a network search engine.
3. The product information retrieval method according to claim 1, wherein: Get the title information and meta description information specifically as follows: Parse web pages and extract links from them; The title information and meta description information of each web page are obtained according to the link.
4. The product information retrieval method according to claim 1, wherein: The specific page to be parsed is: Based on the title information and meta description information, a large language model is used to preliminarily determine the relevance of any web page to the product name / brand name in the product query request; Select pages with correlation higher than the threshold as pages to be parsed.
5. The commodity information retrieval method according to claim 4, wherein: Use the large language model to preliminarily determine the relevance of any web page to the product name / brand name in the product query request, specifically: The BGE Embedding model is used to convert the meta information and the search keywords in the product query request into vectors, and the similarity between the vectors is calculated to achieve a preliminary judgment of the relevance.
6. The commodity information retrieval method according to claim 4, wherein: After obtaining the page to be parsed, the method further includes: Obtain the secondary links in the page to be parsed, obtain the webpage content corresponding to the secondary links, and perform preliminary screening on the webpage content.
7. The commodity information retrieval method according to claim 1, wherein: The key information of the product is as follows: Use the HTML parsing library to parse the text content of the page to be parsed and convert it into Markdown format; A multimodal model is used to identify and analyze the images in the page to be parsed, and the product selling points and product specifications in the images are converted into text form, and integrated with the parsed text content; wherein, the converted Markdown format content is stored in the knowledge base as product brand knowledge.
8. A commodity information retrieval device based on a large language model, characterized in that: include: A query input unit, configured to receive a product query request input by a user; The product query request includes a product name or a brand name; A result response unit, configured to respond to the product query request to obtain search results, and obtain top k web pages from the search results; A web page parsing unit is used to parse the content of the top k web pages to obtain title information and meta description information; A webpage filtering unit is used to filter the top k webpages based on the large language model, title information, meta description information and product query request to obtain the pages to be parsed; The multimodal analysis unit is used to use a multimodal model to identify and analyze the page to be analyzed to obtain key information of the product; the key information of the product includes the selling points and specifications of the product.
9. A commodity information retrieval device based on a large language model, characterized in that: The method comprises a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 7.