Product information recommendation method and device, storage medium and electronic device
Patent Information
- Application Number
- CN202610774137.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-29
- Publication Date
- 2026-09-18
AI Technical Summary
然而,相关技术中仅支持对图像中单个显著主体进行识别与同款召回,导致推荐结果局限于孤立单品,无法生成符合用户潜在需求的场景化组合方案
[0023] In this embodiment, the following steps are adopted: obtaining image information input by the target object, wherein the image information includes information on multiple products; determining target scene information matching the multiple products based on the information on the multiple products in the image information; determining the target product to be recommended based on the target scene information, and pushing the information of the target product to the target object, thereby solving the technical problem of comparing the accuracy of product recommendations in related technologies.
Smart Images

Figure CN122779945A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a method and apparatus for recommending product information, a storage medium, and an electronic device. Background Technology
[0002] In the field of image search, users often upload complex scene images containing multiple product subjects, such as holiday outfits or camping gear combinations, hoping to obtain related product recommendations through the images. However, related technologies only support the identification and recall of single salient subjects in the image, resulting in recommendations limited to isolated single items and failing to generate scenario-based combination solutions that meet the user's potential needs.
[0003] Regarding the technical issues concerning the accuracy of product recommendations mentioned above, no effective solution has yet been proposed. Summary of the Invention
[0004] This application provides a method and apparatus for recommending product information, a storage medium, and an electronic device to at least solve the technical problem of comparing the accuracy of product recommendations in related technologies.
[0005] According to one aspect of the embodiments of this application, a method for recommending product information is provided, comprising: obtaining image information input by a target object, wherein the image information includes information on multiple products; determining target scene information matching the multiple products based on the information on the image information; determining target products to be recommended that match the target scene information based on the target scene information, and pushing the information of the target products to the target object.
[0006] Further, based on the target scene information, determining the target product to be recommended that matches the target scene information includes: calling a search term recommendation tool to determine at least one candidate search query term based on the target scene information; calling a relevance analysis tool to score the relevance of the at least one candidate search query term to obtain a target score value; if the target score value is greater than a preset threshold, then searching based on the at least one candidate search query term to obtain the target product to be recommended.
[0007] Furthermore, after calling the relevance analysis tool to score the relevance of the at least one candidate search query term and obtaining a target score, the method further includes: if the target score of a candidate search query term is less than or equal to the preset threshold, then calling the search term rewriting tool to rewrite the candidate search query term to obtain a rewritten candidate search query term; calling the search term recommendation tool to score the rewritten candidate search query term again, until the rewritten candidate search query term is greater than the preset threshold.
[0008] Furthermore, based on the information of multiple products in the image information, determining the target scene information that matches the multiple products includes: calling a search tool to search for the multiple products contained in the image information to obtain product description information corresponding to the multiple products; and calling a scene extraction skill tool to perform relationship reasoning on the product description information and the image information to obtain the target scene information.
[0009] Furthermore, the relevance analysis tool is invoked to score the relevance of the at least one candidate search query term to obtain a target score, including: scoring the relevance between the at least one candidate search query term and the target scene information to obtain a first score; scoring the relevance between the at least one candidate search query term and the image information to obtain a second score; and obtaining the target score based on the first score and the second score.
[0010] Furthermore, after obtaining the image information input by the target object, the method further includes: calling a recognition skill tool to identify the products contained in the image information, and obtaining the proportion area corresponding to the products in the image information; and determining the information of multiple products in the image information based on the proportion area.
[0011] Furthermore, the scene extraction tool is invoked to perform relationship reasoning on the product description information and the image information to obtain the target scene information, including: identifying the product description information to obtain corresponding product feature information, wherein the product feature information includes at least: product category information, product attribute information, and association information between products; identifying the image information to obtain global semantic information; and performing reasoning based on the product feature information and the global semantic information to obtain the target scene information.
[0012] According to another aspect of the embodiments of this application, a method for recommending product information is also provided, comprising: obtaining image information corresponding to a target object uploaded by a client, wherein the image information includes information of multiple products; determining target scene information matching the multiple products based on the information of the multiple products in the image information in a cloud server; determining a target product to be recommended based on the target scene information; and pushing the information of the target product to the client.
[0013] According to another aspect of the embodiments of this application, a product information recommendation device is also provided, comprising: an acquisition unit, configured to acquire image information input by a target object, wherein the image information includes information of multiple products; a first determination unit, configured to determine target scene information matching the multiple products based on the information of the multiple products in the image information; and a second determination unit, configured to determine a target product to be recommended matching the target scene information based on the target scene information, and push the information of the target product to the target object.
[0014] Furthermore, the second determining unit includes: a first calling module, used to call a search term recommendation skill tool to determine at least one candidate search query term based on the target scenario information; a second calling module, used to call a relevance analysis skill tool to score the relevance of the at least one candidate search query term to obtain a target score value; and a search module, used to perform a search based on the at least one candidate search query term to obtain the target product to be recommended if the target score value is greater than a preset threshold.
[0015] Furthermore, the device further includes: a first invocation unit, configured to, after invoking a relevance analysis skill tool to score the relevance of the at least one candidate search query term and obtaining a target score value, if there is a candidate search query term whose target score value is less than or equal to the preset threshold, invoking a search term rewriting skill tool to rewrite the candidate search query term to obtain a rewritten candidate search query term; and a second invocation unit, configured to invoke the search term recommendation skill tool to score the rewritten candidate search query term again, until the rewritten candidate search query term is greater than the preset threshold.
[0016] Furthermore, the first determining unit includes: a third calling module, used to call a search tool to search for multiple products contained in the image information to obtain product description information corresponding to the multiple products; and a fourth calling module, used to call a scene extraction skill tool to perform relationship reasoning on the product description information and the image information to obtain the target scene information.
[0017] Furthermore, the second calling module includes: a first scoring submodule, used to score the relevance of the at least one candidate search query term and the target scene information to obtain a first score; a second scoring submodule, used to score the relevance of the at least one candidate search query term and the image information to obtain a second score; and a determining submodule, used to obtain the target score based on the first score and the second score.
[0018] Furthermore, the device further includes: a third calling unit, configured to, after acquiring the image information input by the target object, call a recognition skill tool to identify the products contained in the image information and obtain the proportion area corresponding to the products in the image information; and a third determining unit, configured to determine the information of multiple products in the image information based on the proportion area.
[0019] Furthermore, the fourth invocation module includes: a first identification module, used to identify the product description information to obtain corresponding product feature information, wherein the product feature information includes at least: product category information, product attribute information, and association information between products; a second identification module, used to identify the image information to obtain global semantic information; and a reasoning module, used to reason based on the product feature information and the global semantic information to obtain the target scene information.
[0020] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein, when the program is executed, the device where the storage medium is located executes the above-described method for recommending product information.
[0021] According to another aspect of the present application, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the above-described method for recommending product information during runtime.
[0022] According to another aspect of the embodiments of this application, a computer program product is provided, including a computer program or instructions, wherein the computer program or instructions, when executed by a processor, implement the above-described method for recommending product information.
[0023] In this embodiment, the following steps are adopted: obtaining image information input by the target object, wherein the image information includes information on multiple products; determining target scene information matching the multiple products based on the information on the multiple products in the image information; determining the target product to be recommended based on the target scene information, and pushing the information of the target product to the target object, thereby solving the technical problem of comparing the accuracy of product recommendations in related technologies.
[0024] In this application, by treating the information of multiple products in an image as a whole, rather than processing them in isolation, joint modeling of the semantic relationships between products is achieved, thereby generating target scene information that aligns with the user's potential intent. Then, based on the target scene information, target products matching the target scene information are identified and pushed to the target audience. This avoids the fragmentation and intent bias caused by relying solely on single features. The pushed products are functionally, stylistically, and contextually compatible with the image content, thus improving the accuracy of recommendations. Attached Figure Description
[0025] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0026] Figure 1 This is a hardware structure block diagram of a computer terminal provided according to Embodiment 1 of this application;
[0027] Figure 2 This is a flowchart of a method for recommending product information according to Embodiment 1 of this application;
[0028] Figure 3 This is an illustration of a method for recommending product information according to Embodiment 1 of this application. Figure 1 ;
[0029] Figure 4 This is an illustration of a method for recommending product information according to Embodiment 1 of this application. Figure 2 ;
[0030] Figure 5 This is a flowchart of a method for recommending product information according to Embodiment 2 of this application;
[0031] Figure 6 This is a schematic diagram of a device for recommending product information according to Embodiment 3 of this application;
[0032] Figure 7 This is a structural block diagram of an electronic device provided according to Embodiment 4 of this application. Detailed Implementation
[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws and standards of the relevant regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0036] Example 1
[0037] According to an embodiment of this application, a method for recommending product information is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0038] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a method of recommending product information is shown. Figure 1 As shown, the computer terminal (or mobile device) 10 may include a processor set 102 (the processor set 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA, and the processor set 102 may include a processor set, Figure 1The data is illustrated using 102a, 102b, ..., 102n. A memory 104 is used for storing data, and a transmission module 106 is used for communication functions. In addition, it may include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0039] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the product information recommendation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned product information recommendation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0041] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0042] The display may be, for example, a touchscreen LCD display that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0043] Under the aforementioned operating environment, this application provides the following: Figure 2 The recommended method for the product information shown. Figure 2 This is a flowchart of a product information recommendation method according to Embodiment 1 of this application. The product information recommendation method includes:
[0044] Step S201: Obtain the image information input by the target object, wherein the image information includes information about multiple products.
[0045] Optionally, image information input by the user (i.e., the target object mentioned above) can be received through an image search portal. This image information includes details of multiple products; for example, an image may simultaneously contain multiple product entities such as trekking poles, shoes, sleeping bags, and tents. The image information input by the user can originate from snapshots taken by the user in real-world usage scenarios, such as camping or holiday decorations.
[0046] It should be noted that the product can be a physical product or a virtual product, such as digital illustrations or in-game items.
[0047] In an optional embodiment, the aforementioned multiple products can be products to be analyzed. That is, after receiving an image input by the user, the information contained in the image is identified to determine the main product contained therein, and the main product is used as the product to be analyzed. It should be noted that the main product can be a product that occupies the majority of the image, for example, a product whose proportion in the image exceeds a threshold.
[0048] Step S202: Based on the information of multiple products in the image information, determine the target scene information that matches the multiple products.
[0049] Optionally, the corresponding target scene information can be obtained by reasoning based on the information of multiple products in the image information.
[0050] For example, a detection module can be invoked to locate and identify products in image information. For example, based on the visual understanding capabilities of a multimodal large model, each subject in the image can be detected, and the bounding box, product category (such as hiking poles, Santa hats, sleeping bags), attribute features (such as color, material, style: red and white color scheme, plush, outdoor functionality, etc.) and confidence score of each subject can be output.
[0051] For virtual products (such as deer dolls in digital illustrations), visual semantic alignment technology can also be used to match them to the corresponding categories in the product knowledge graph, thereby achieving a unified representation of virtual and physical products.
[0052] Based on information from multiple products, output a structured, natural language expression of the target scene information. For example, analyze the spatial layout, visual synergy, and functional complementarity among the various entities to obtain the corresponding target scene information.
[0053] In an optional embodiment, the target scene information can be in the format of: [core theme] + [style / atmosphere] + [usage context].
[0054] For example: Input items: Santa hat, red sweater, reindeer doll, fairy lights, plush boots;
[0055] Output target scene information: Christmas-themed family party outfits and decorations.
[0056] Input items: trekking poles, tent, sleeping bag, sleeping mat, headlamp;
[0057] Output target scene information: "Outdoor hiking and camping equipment kit".
[0058] In an optional embodiment, if the image uploaded by the user has a blurry, occluded, or low-resolution subject, semantic completion can be performed based on the identified subject. For example, based on the identified subject, scene-related products that may be missing can be retrieved, and virtual subjects (labeled as inferred subjects) can be generated with high probability completion items to participate in subsequent relationship modeling.
[0059] In an optional embodiment, when generating target scene information, the user's recent search actions can also be taken into account. For example, if the user has recently searched or browsed content such as "Christmas gifts" or "parent-child outfits", then sub-scenes such as "family gathering" and "parent-child matching" will be given higher weight. Furthermore, the rationality of the generated scene can be verified by the "scene-product" relationship in the associated product knowledge graph (such as "Christmas party" often being paired with "sequin dress" and "furry socks") to avoid semantic conflicts (such as "desert camping" and "Christmas blanket" coexisting).
[0060] Step S203: Based on the target scenario information, determine the target products to be recommended that match the target scenario information, and push the information of the target products to the target objects.
[0061] Optionally, based on the target scenario information obtained in step S202, a target product to be recommended that matches the target scenario information is determined, and the information of the target product is pushed to the user.
[0062] For example, target scene information is input into a multimodal semantic encoder, which transforms it into a dense semantic vector. Then, based on the product's existing product library on the platform and multidimensional information such as product images, titles, detail page text, and user reviews, a semantic embedding vector for each product is pre-constructed. By calculating the semantic similarity between the target scene vector and the product vector, candidate products that highly match the target scene information are quickly selected.
[0063] To avoid overly simplistic or misleading recommendations, a scenario-adaptive filtering mechanism can be further introduced: retaining products that form a logical loop with the target scenario in terms of category, style, and functionality. For example, if the scenario is "Christmas-themed family gathering outfits and decorations," then clothing, accessories, and home decorations with holiday elements should be prioritized, while sports equipment or office supplies unrelated to the scenario should be excluded.
[0064] In addition, the recommendations can be personalized by taking into account the user's browsing history, favorites, and purchase behavior, making the recommendations more in line with individual preferences.
[0065] In an optional embodiment, the execution entity of the product information recommendation method provided in Embodiment 1 of this application can be an Agent plus skills architecture. This architecture is based on a multimodal large model and completes the entire closed loop from image understanding to product recommendation through modular and programmable skill units. It should be noted that an Agent is an intelligent entity capable of perceiving the environment, making autonomous decisions, and taking actions to achieve specific goals. The agent created by this scheme is a multimodal agent, based on a multimodal large model, possessing not only text understanding capabilities but also image perception capabilities. Skills are specific capability modules that the Agent can invoke, which are functional units for the Agent to perform specific tasks. For example, the Agent, as the intelligent decision-making center, receives product images containing multiple subjects uploaded by the user as initial input and autonomously schedules multiple predefined skills to complete the task step by step according to a predetermined logical process.
[0066] In summary, by treating information from multiple products in an image as a whole, rather than processing them in isolation, joint modeling of semantic relationships between products is achieved, thereby generating target scene information that aligns with the user's potential intent. Then, based on the target scene information, target products matching the target scene information are identified and pushed to the target audience. This avoids the fragmentation and intent bias caused by relying solely on single features. The pushed products are functionally, stylistically, and contextually compatible with the image content, thus improving the accuracy of recommendations.
[0067] To improve the accuracy of target products, the product information recommendation method provided in Embodiment 1 of this application determines the target products to be recommended based on target scenario information, including: calling a search term recommendation skill tool to determine at least one candidate search query term based on the target scenario information; calling a relevance analysis skill tool to score the relevance of the at least one candidate search query term to obtain a target score value; if the target score value is greater than a preset threshold, then a search is performed based on the at least one candidate search query term to obtain the target products to be recommended.
[0068] Optionally, the Agent invokes the search term recommendation skill (i.e., the search term recommendation tool mentioned above), taking structured target scene information as input, such as: Christmas-themed family party outfits and decorations, and outputs at least one candidate search query term, such as: Christmas red sweater, reindeer pattern plush socks, glowing Christmas tree ornaments, matching Christmas hat sets for parents and children, and sequined party dresses, etc.
[0069] Then, the Agent invokes the relevance analysis skill (i.e., the relevance analysis tool mentioned above) to score the candidate query terms, that is, to determine the relevance of the candidate query terms to the target scene information, and obtain the target score value mentioned above.
[0070] The agent determines whether the target score of a candidate query term is higher than a preset threshold (e.g., 0.70). If it is higher than the threshold, it is a valid query term. If at least one candidate query term is determined to have a target score greater than the preset threshold, a search is performed based on at least one candidate query term to obtain the target product to be recommended.
[0071] For example, multiple candidate search queries can be treated as a set of semantically coordinated search intents, prioritizing the recall of products that simultaneously satisfy the semantic features of multiple keywords. For instance, a product like a matching Christmas hat set for parents and children not only matches the terms "matching" and "Christmas hat," but also implies scenarios such as family gatherings and outfit coordination. Therefore, it can receive higher weight in multi-word joint ranking and is thus prioritized for inclusion in the recommendation list.
[0072] Since different query terms may point to the same product (e.g., "Christmas red sweater 1" and "red Christmas sweater" refer to the same product), duplicate recommendations are avoided by using product IDs. Additionally, to prevent recommendations from being overly concentrated in a single category (e.g., only recommending sweaters), the Agent can also invoke the scenario coverage enhancement skill to limit the recommended results to cover at least two functional subcategories based on the multi-dimensional attributes of the target scenario (e.g., "outfit," "decoration," "family").
[0073] In an optional embodiment, the search term recommendation skill, based on target scenario information, determines at least one candidate search query, including: taking structured target scenario information as input, such as "Christmas-themed family party outfits and decorations," and breaking it down into semantic atomic units: core theme (Christmas), style atmosphere (family party, warm and cute), and usage context (outfits, decorations). Candidate query terms can then be generated in parallel through the following three paths:
[0074] Direct semantic conversion path: Directly convert keywords in the scenario into commonly used search terms by users.
[0075] For example, "Christmas theme" is converted to "Christmas"; "outfit" is converted to "sweater" or "coat"; and "decoration" is converted to "ornament" or "hanging ornament". Combining these elements generates items such as "Christmas sweater for women" and "Christmas decorative lights".
[0076] Co-occurrence reasoning path: Based on the triple relationship of "product-scenario-user behavior" in the knowledge graph, we can mine product words that co-occur frequently in the same scenario.
[0077] For example, historical data shows that when users search for Christmas sweaters, 87% of users also click on "reindeer pattern socks", thus inferring that "reindeer pattern plush socks" is a strongly related candidate keyword.
[0078] Path for scene expansion and context completion: Combine the context to complete the implicit needs that users have not explicitly stated but are in line with their habits.
[0079] For example, if the target scenario includes "family gathering", it can expand to sub-intents such as matching parent-child outfits, family outfits, and photo props, and generate query terms such as parent-child Christmas hat sets and Christmas-themed group photo backdrops.
[0080] Through the multi-stage mechanism consisting of the Agent-coordinated scheduling of search term recommendation skills and relevance analysis skills, the system no longer relies on a single model's coarse-grained understanding of images to directly match products. Instead, it improves the accuracy of recommendations by transforming scene semantics into real user search language, filtering highly relevant search terms, jointly performing multi-word semantic recall, deduplication and redundancy prevention, and enhancing the fit between recommendation results and users' true intentions.
[0081] To improve the accuracy of candidate search query terms, in the product information recommendation method provided in Embodiment 1 of this application, after calling a relevance analysis skill tool to score the relevance of at least one candidate search query term and obtaining a target score value, the method further includes: if the target score value corresponding to a candidate search query term is less than or equal to a preset threshold, then calling a search term rewriting skill tool to rewrite the candidate search query term to obtain a rewritten candidate search query term; calling a search term recommendation skill tool to score the relevance of the rewritten candidate search query term again, until the rewritten candidate search query term is greater than the preset threshold.
[0082] Optionally, when the Agent calls the relevance analysis skill to score candidate search query terms (such as "party sequined dress"), if the target score is lower than a preset threshold (such as 0.70), it is determined that the term is semantically relevant but inaccurate or ambiguous. For example, sequined dresses are often associated with non-family holiday scenarios such as parties and New Year's Eve parties in searches, which is semantically different from the intention of Christmas-themed family gatherings.
[0083] The agent triggers the search term rewriting skill (i.e., the relevance analysis skill tool mentioned above) to rewrite candidate search query terms, for example, semantically reconstructing low-scoring query terms from three dimensions:
[0084] Scene-aligned rewriting: Replace vague or generalized words with more specific terms that better fit the target scene. Example: Rewrite the original phrase "party sequined dress" as "Christmas party dress";
[0085] User habit alignment rewriting: Transform machine-generated or unnatural expressions into frequently searched terms within the platform. Example: Rewrite the original phrase "glowing Christmas tree ornament" as "glowing Christmas tree decoration";
[0086] Category standardization rewriting: Correcting colloquial and non-standard category terms to standard category / attribute terms. Example: The original term "deer pattern plush socks" is rewritten as "deer pattern warm socks".
[0087] After obtaining the rewritten candidate search terms, the search term recommendation tool is called again to score the relevance of the rewritten candidate search terms until the rewritten candidate search terms exceed the preset threshold.
[0088] By introducing search term rewriting skills and a multi-round re-evaluation mechanism, the accuracy of candidate search terms was improved, thereby increasing their matching success rate in product searches. Without increasing the number of recalled terms, the quality distribution of candidate terms was optimized, reducing missed recommendations due to inaccurate expression, making the recommended results closer to the user's true search intent, and improving the system's robustness and recommendation coverage in complex contexts.
[0089] To improve the accuracy of target scene information, in the product information recommendation method provided in Embodiment 1 of this application, target scene information matching multiple products is determined based on the information of multiple products in the image information. This includes: calling a search tool to search for multiple products contained in the image information to obtain product description information corresponding to multiple products; and calling a scene extraction skill tool to perform relationship reasoning on the product description information and image information to obtain target scene information.
[0090] Optionally, after receiving an image containing multiple products uploaded by a user, the Agent can call the image detection and recognition module to locate and segment prominent product subjects in the image (such as: red sweater, reindeer doll, glowing lights, Santa hat, etc.), and generate independent image region cropping results for each subject.
[0091] Then, for each main area, the Agent can call the platform's internal image search tool, take the local image as input, search for matching or similar products in the product library, and return the structured product description information corresponding to the product, including: product title, category (such as women's clothing, sweaters, home furnishings, holiday decorations), brand, material, high-frequency words in user reviews (such as warm, cute, strong atmosphere), and platform-related scene tags (such as Christmas, family gathering, parent-child outfits).
[0092] Finally, the Agent invokes the scene extraction skill (i.e., the scene extraction tool mentioned above) to jointly model the text descriptions and original images of the multiple products, and then outputs the target scene information.
[0093] It should be noted that the scene extraction skill can be a joint modeling of the text descriptions and original images of the above-mentioned multiple products using a multimodal large model as the inference base.
[0094] By using image search tools to obtain structured descriptions of multiple products and combining them with multimodal large models to jointly model textual and visual information, we can extract scene expressions that closely resemble real user intentions from blurry images, thereby improving the accuracy of subsequent recommendations.
[0095] To further improve the accuracy of target scene information, in the product information recommendation method provided in Embodiment 1 of this application, a scene extraction skill tool is invoked to perform relational reasoning on product description information and image information to obtain target scene information. This includes: identifying product description information to obtain corresponding product feature information, wherein the product feature information includes at least: product category information, product attribute information, and association information between products; identifying image information to obtain global semantic information; and performing reasoning based on product feature information and global semantic information to obtain target scene information.
[0096] Optionally, the scene extraction skill performs semantic parsing on the structured product description information of each product to extract product feature information, including: product category information (such as sweaters, decorations, plush toys, toys), clarifying the product category to which each subject belongs; product attribute information (such as red, plush material, light-up, with a hat), capturing specific visual and functional characteristics; and relationship information between products, inferring the functional collaboration relationship between products (such as sweaters for wearing, plush toys for display, and colored lights for creating atmosphere) through co-occurrence frequency, user evaluation semantics (such as buying them together so the whole family can wear them), and category combination rules (such as clothing and decoration categories often appearing together in holiday scenes). This forms a preliminary logical chain of items, behaviors, and scenes.
[0097] Then, the scene extraction skill performs global visual analysis on the original image to extract global semantic information, including: overall color tone (such as warm red and golden yellow, which are in line with the festive atmosphere); lighting and spatial layout (such as indoor environment, people's standing posture, and the placement of objects); recognition of non-product elements (such as the outline of a Christmas tree in the background, a gift box on the windowsill, and text watermarks); and cues of people's behavior (such as multiple people in the same frame, smiling expressions, and holding objects), reflecting social and usage contexts.
[0098] By integrating the aforementioned product feature information with global semantic information, and through semantic weight aggregation and scene probability modeling, target scene information is output, such as Christmas-themed family party attire and decorations.
[0099] In an optional embodiment, joint modeling of the text descriptions and original images of the above-mentioned products includes: semantic alignment: cross-modal alignment of keywords in the product title (such as deer pattern, glowing, same style) with the visual features of the corresponding objects in the image (such as plush texture, light reflection, and how people wear it) to confirm the consistency between the text description and the visual content and eliminate misidentification or noise interference.
[0100] Functional synergy analysis: By combining product categories and user reviews, the usage relationships between multiple entities can be inferred. For example, a red sweater and a reindeer plush toy belong to the clothing and decoration categories, respectively. However, user reviews mention that a family of three wears them together and that they create a warm and cozy atmosphere at home. This suggests that the two items do not exist in isolation but rather together form a synergistic combination for creating a festive family atmosphere.
[0101] Scene-based inductive reasoning: With the support of a knowledge graph, the combination is mapped to a matching scene. For example, when four product categories—sweaters, dolls, lamps, and hats—are detected to co-occur, and their semantics are concentrated on keywords such as red, plush, holiday, and family, based on the statistical pattern that this combination was used by 150,000 users for family Christmas parties in December, the target scene information is output: Christmas-themed family party attire and decorations.
[0102] By aligning and integrating the categories, attributes, and relationships in product descriptions with the global semantic features of images in multiple dimensions, the one-sidedness and risk of misjudgment caused by relying solely on text tags or visual recognition are effectively mitigated. Without introducing additional manual rules, the accuracy of scene understanding is improved, making the recommendation starting point closer to the user's true intent.
[0103] To improve the accuracy of the scoring, the product information recommendation method provided in Embodiment 1 of this application calls a relevance analysis tool to score the relevance of at least one candidate search query term and obtain a target score value, including: scoring the relevance of at least one candidate search query term and target scene information to obtain a first score value; scoring the relevance of at least one candidate search query term and image information to obtain a second score value; and obtaining a target score value based on the first score value and the second score value.
[0104] Optionally, a semantic relevance match is performed between candidate search queries and target scenario information (such as Christmas-themed family party outfits and decorations) to calculate a first score. For example, the evaluation assesses whether candidate search queries semantically cover or extend the core elements of the scenario (such as matching Christmas party dresses with Christmas family party outfits, with a score of 0.85), ensuring that recommended terms are consistent with the user's intended scenario and avoiding generalized terms that deviate from the theme.
[0105] Then, visual and text alignment analysis is performed on the candidate search query terms and the original uploaded images (i.e., the image information mentioned above) to verify whether the query terms have image-level evidence and to prevent words that are generated out of thin air or semantically drifted.
[0106] Finally, the first score and the second score are weighted and fused to generate the final target score. The first score focuses on the semantic intent relevance and can be set with a higher weight (e.g., 70%) to ensure that the recommendations do not deviate from the user's potential needs. The second score focuses on the supporting image evidence and has a relatively lower weight (e.g., 30%) to filter out pseudo-related words that lack visual evidence.
[0107] By employing a dual-dimensional scoring mechanism, we ensure the semantic consistency between recommended words and users' abstract intentions, while also constraining them to possess interpretability and traceability of image content. This effectively reduces misjudgments and noise caused by relying solely on scene or image-based scoring, making the final scoring results more robust and interpretable, and improving the accuracy of candidate word selection and the credibility of recommendations.
[0108] To improve the accuracy of product identification, in the product information recommendation method provided in Embodiment 1 of this application, after obtaining the image information input by the target object, the method further includes: calling a recognition skill tool to identify the products contained in the image information, and obtaining the proportion area corresponding to the products in the image information; and determining the information of multiple products in the image information based on the proportion area.
[0109] Optionally, an object detection tool (such as a deep learning-based multi-object detection model) can be invoked to perform end-to-end object detection on the input image, identify all possible product entities in the image, and output the bounding box of each detected product and its confidence score.
[0110] Based on the pixel range of each bounding box, its area ratio in the entire image (i.e., the percentage of the product area to the total pixel area of the image) is calculated as a quantitative indicator of its significance. The identified products are then sorted and filtered based on this area ratio.
[0111] If multiple products occupy similar areas and all exceed a certain threshold (e.g., all greater than 10%), they are identified as multiple significant subjects and enter the multi-subject analysis process. Detection results with excessively small areas (e.g., less than 3%) or low confidence levels are considered background interference, local details, or non-product elements (e.g., hands, furniture, text) and are filtered to avoid misjudging them as valid subjects.
[0112] For example, in an image containing a Christmas sweater, a reindeer doll, fairy lights, and a coffee table, if the sweater occupies 28%, the doll 12%, the fairy lights 9%, while the coffee table only occupies 5% and the hands 4%, then the sweater, doll, and fairy lights will be identified as the three main products, while the coffee table and hands will be ignored. This mechanism helps avoid misidentification caused by relying solely on detection confidence or category probability (such as misidentifying background tables and chairs as products), thus improving the accuracy of subsequent recommendations.
[0113] In an alternative embodiment, the following can be employed: Figure 3 The diagram shown illustrates product recommendation, which is based on an Agent and Skills architecture that extracts the main image, understands the scene, and recommends relevant queries.
[0114] Step 1: Invoke the intent recognition skill and use it for intent recognition: After the user uploads a product image, perform subject bounding box detection on the product's main body. If there is only one salient e-commerce subject (a salient e-commerce subject is generally defined as an e-commerce subject occupying a large area of the entire image), then recommend similar products. If there are multiple subject bounding boxes for the product, and multiple salient subjects exist, then proceed to the multi-subject recommendation scenario.
[0115] Step 2: Based on the image uploaded by the user, use image search tools and / or similar tools to obtain the main information in the image.
[0116] Step 3: Introduce scene extraction skills to analyze the relationships between different subjects, and perform correlation analysis on relevant elements and text information in the image to find the user's core intent and output a multi-subject scene information that matches the user's intent.
[0117] Step 4: Introduce query recommendation skill (i.e., search term recommendation skill). Based on the scenario information, recommend four queries that match the scenario: {query1, query2, query3, query4}.
[0118] Step 5: Introduce the relevance skill (i.e., the relevance analysis skill) to score the relevance of the query.
[0119] Step 6: Introduce query rewriting skills. Rewrite queries below the relevant score threshold, and sort queries that meet the relevant score threshold from high to low according to the relevant score.
[0120] Step 7: Recall the product search results for the above query in the user interface.
[0121] In an optional embodiment, the user inputs an image such as... Figure 4 As shown, Step 1: Intent recognition skill identifies multiple e-commerce entities; Step 2: Call Taobao's image search tool to return a list of products such as Christmas hats, red Christmas-themed coats, Christmas snowman dolls, large Christmas reindeer dolls, and festive lights; Step 3: Based on the scene, extract the product information returned by the skill and product image search tool, parse the entity information, analyze the relationship between different e-commerce entities, and combine the semantic information of the whole image to perform multi-entity relationship reasoning. With the person's clothing as the core entity, it is inferred that the user has the intent to purchase Christmas-themed clothing, and the theme scene information: Christmas-themed clothing is extracted; Step 4: Based on the query The recommended skill (i.e., the search term recommendation skill) recommends related queries (search terms) based on Christmas-themed outfits, such as ['Christmas red sweater', 'reindeer pattern sweatshirt', 'Christmas plaid skirt', 'party sequined dress']. Step 5: Introduce the relevance skill to score the queries. The first three pass the relevance threshold, while the fourth, "party sequined dress", fails. Step 6: Introduce the query rewriting skill to rewrite the failed query, such as 'party sequined dress' ---> 'Christmas party dress'. Step 7: Recall the product search results for the above queries on the product page.
[0122] Compared to related technologies that can only understand a single subject, this application utilizes the full-image perception capability of a multimodal large model to perform multi-subject recognition and understanding of multiple subjects in an image. At the same time, based on the multimodal large model, it uses the current Agent plus skills architecture to understand and extract relationships between multiple subject products in the entire image, accurately abstract relevant scene information, mine the user's potential intent, and improve the accuracy of product recommendations.
[0123] In the product information recommendation method provided in Embodiment 1 of this application, image information input by the target object is obtained, wherein the image information includes information on multiple products; based on the information on multiple products in the image information, target scene information matching multiple products is determined; based on the target scene information, target products to be recommended that match the target scene information are determined, and the information on the target products is pushed to the target object, thus solving the technical problem of comparing the accuracy of product recommendations in related technologies.
[0124] In this application, by treating the information of multiple products in an image as a whole, rather than processing them in isolation, joint modeling of the semantic relationships between products is achieved, thereby generating target scene information that aligns with the user's potential intent. Then, based on the target scene information, target products matching the target scene information are identified and pushed to the target audience. This avoids the fragmentation and intent bias caused by relying solely on single features. The pushed products are functionally, stylistically, and contextually compatible with the image content, thus improving the accuracy of recommendations.
[0125] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0127] Example 2
[0128] According to embodiments of this application, a method for recommending product information is also provided, such as... Figure 5 As shown, the methods for recommending this product information include:
[0129] Step S501: Obtain the image information corresponding to the target object uploaded by the client, wherein the image information includes information about multiple products;
[0130] Step S502: Based on the information of multiple products in the image information, determine the target scene information that matches the multiple products in the cloud server; based on the target scene information, determine the target products to be recommended that match the target scene information.
[0131] Step S503: Push the target product information to the client.
[0132] It should be noted that the specific steps for processing image information on the cloud server are the same as in Example 1, and will not be repeated here.
[0133] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0135] Example 3
[0136] According to embodiments of this application, a product information recommendation apparatus for implementing the above-described product information recommendation method is also provided, such as... Figure 6 As shown, the device includes: an acquisition unit 601, a first determination unit 602, and a second determination unit 603.
[0137] The acquisition unit 601 is used to acquire image information input by the target object, wherein the image information includes information about multiple products;
[0138] The first determining unit 602 is used to determine target scene information that matches multiple products based on information about multiple products in the image information.
[0139] The second determining unit 603 is used to determine the target product to be recommended that matches the target scene information based on the target scene information, and push the information of the target product to the target object.
[0140] In the product information recommendation device provided in Embodiment 3 of this application, the acquisition unit 501 acquires the image information input by the target object, wherein the image information includes information on multiple products; the first determination unit 502 determines the target scene information matching the multiple products based on the information on the multiple products in the image information; the second determination unit 503 determines the target product to be recommended based on the target scene information and pushes the information of the target product to the target object, thereby solving the technical problem of comparing the accuracy of product recommendations in related technologies.
[0141] In this application, by treating the information of multiple products in an image as a whole, rather than processing them in isolation, joint modeling of the semantic relationships between products is achieved, thereby generating target scene information that aligns with the user's potential intent. Then, based on the target scene information, target products matching the target scene information are identified and pushed to the target audience. This avoids the fragmentation and intent bias caused by relying solely on single features. The pushed products are functionally, stylistically, and contextually compatible with the image content, thus improving the accuracy of recommendations.
[0142] Optionally, in the product information recommendation device provided in Embodiment 3 of this application, the second determining unit 603 includes: a first calling module, used to call a search term recommendation skill tool to determine at least one candidate search query term based on target scene information; a second calling module, used to call a relevance analysis skill tool to score the relevance of the at least one candidate search query term and obtain a target score value; and a search module, used to perform a search based on the at least one candidate search query term if the target score value is greater than a preset threshold, to obtain the target product to be recommended.
[0143] Optionally, in the product information recommendation device provided in Embodiment 3 of this application, the device further includes: a first calling unit, used to call a relevance analysis skill tool to score the relevance of at least one candidate search query term and obtain a target score value; if the target score value corresponding to a candidate search query term is less than or equal to a preset threshold, then call a search term rewriting skill tool to rewrite the candidate search query term to obtain a rewritten candidate search query term; and a second calling unit, used to call a search term recommendation skill tool to score the rewritten candidate search query term again until the rewritten candidate search query term is greater than the preset threshold.
[0144] Optionally, in the product information recommendation device provided in Embodiment 3 of this application, the first determining unit 602 includes: a third calling module, used to call a search tool to search for multiple products contained in the image information to obtain product description information corresponding to multiple products; and a fourth calling module, used to call a scene extraction skill tool to perform relationship reasoning on the product description information and image information to obtain target scene information.
[0145] Optionally, in the product information recommendation device provided in Embodiment 3 of this application, the second calling module includes: a first scoring submodule, used to score the relevance of at least one candidate search query term and target scene information to obtain a first score value; a second scoring submodule, used to score the relevance of at least one candidate search query term and image information to obtain a second score value; and a determining submodule, used to obtain a target score value based on the first score value and the second score value.
[0146] Optionally, in the product information recommendation device provided in Embodiment 3 of this application, the device further includes: a third calling unit, used to call a recognition skill tool to identify the products contained in the image information after obtaining the image information input by the target object, and obtain the proportion area corresponding to the products in the image information; and a third determining unit, used to determine the information of multiple products in the image information based on the proportion area.
[0147] Optionally, in the product information recommendation device provided in Embodiment 3 of this application, the fourth calling module includes: a first identification module, used to identify product description information to obtain corresponding product feature information, wherein the product feature information includes at least: product category information, product attribute information and association information between products; a second identification module, used to identify image information to obtain global semantic information; and a reasoning module, used to reason based on product feature information and global semantic information to obtain target scene information.
[0148] It should be noted that the aforementioned acquisition unit 601, first determination unit 602, and second determination unit 603 correspond to steps S201 to S203 in Embodiment 1. The three units and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the aforementioned units, as part of the device, can operate in the computer terminal 10 provided in Embodiment 1.
[0149] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0150] Example 4
[0151] Embodiments of this application may provide an electronic device, which may be any one of a group of electronic device terminals. Optionally, in this embodiment, the aforementioned electronic device may also be replaced by a terminal device such as a mobile terminal.
[0152] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0153] In this embodiment, the above-mentioned electronic device can execute the program code of the following steps in the product information recommendation method: obtaining image information input by the target object, wherein the image information includes information of multiple products; determining target scene information matching the multiple products based on the information of the multiple products in the image information; determining the target product to be recommended based on the target scene information, and pushing the information of the target product to the target object.
[0154] The aforementioned electronic device can execute the following steps in the product information recommendation method: Based on target scenario information, determine the target product to be recommended that matches the target scenario information, including: calling a search term recommendation skill tool to determine at least one candidate search query term based on the target scenario information; calling a relevance analysis skill tool to score the relevance of the at least one candidate search query term to obtain a target score value; if the target score value is greater than a preset threshold, then perform a search based on the at least one candidate search query term to obtain the target product to be recommended.
[0155] The aforementioned electronic device can execute the following steps in the product information recommendation method: After calling a relevance analysis skill tool to score the relevance of at least one candidate search query term and obtaining a target score, the method further includes: if the target score corresponding to a candidate search query term is less than or equal to a preset threshold, then calling a search term rewriting skill tool to rewrite the candidate search query term to obtain a rewritten candidate search query term; calling a search term recommendation skill tool to score the relevance of the rewritten candidate search query term again until the rewritten candidate search query term is greater than the preset threshold.
[0156] The aforementioned electronic device can execute the following steps in the product information recommendation method: Based on the information of multiple products in the image information, determine the target scene information that matches the multiple products, including: calling a search tool to search for the multiple products contained in the image information to obtain the product description information corresponding to the multiple products; calling a scene extraction skill tool to perform relationship reasoning on the product description information and the image information to obtain the target scene information.
[0157] The aforementioned electronic device can execute the following steps in the product information recommendation method: calling a relevance analysis skill tool to score the relevance of at least one candidate search query term and obtain a target score value, including: scoring the relevance of at least one candidate search query term and target scene information to obtain a first score value; scoring the relevance of at least one candidate search query term and image information to obtain a second score value; and obtaining the target score value based on the first score value and the second score value.
[0158] The aforementioned electronic device can execute the following steps in the product information recommendation method: After obtaining the image information input by the target object, the method further includes: calling a recognition skill tool to identify the products contained in the image information, and obtaining the proportion area corresponding to the products in the image information; based on the proportion area, determining the information of multiple products in the image information.
[0159] The aforementioned electronic device can execute the following steps in the product information recommendation method: calling a scene extraction skill tool to perform relational reasoning on product description information and image information to obtain target scene information, including: identifying product description information to obtain corresponding product feature information, wherein the product feature information includes at least: product category information, product attribute information, and association information between products; identifying image information to obtain global semantic information; and performing reasoning based on product feature information and global semantic information to obtain target scene information.
[0160] Optionally, Figure 7 This is a structural block diagram of an electronic device according to an embodiment of this application. Figure 7As shown, the electronic device 70 may include: one or more ( Figure 7 (Only one is shown in the image) Processor 702 and memory 704. The electronic device 70 may also include a memory controller to control and manage the memory 704; the electronic device 70 may also include a peripheral interface to connect to a radio frequency module, an audio module, and a display screen, etc.
[0161] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the product information recommendation method and apparatus in this embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned product information recommendation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the electronic device 70 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0162] The processor can invoke information and applications stored in the memory via a transmission device to perform the following steps: acquiring image information input by the target object, wherein the image information includes information about multiple products; determining target scene information matching the multiple products based on the information about the multiple products in the image information; determining target products to be recommended based on the target scene information, and pushing the information of the target products to the target object.
[0163] Optionally, the processor may also execute program code for the following steps: determining target products to be recommended based on target scene information, including: calling a search term recommendation tool to determine at least one candidate search query term based on the target scene information; calling a relevance analysis tool to score the relevance of the at least one candidate search query term to obtain a target score value; if the target score value is greater than a preset threshold, then searching based on the at least one candidate search query term to obtain the target products to be recommended.
[0164] Optionally, the processor may also execute program code for the following steps: after calling a relevance analysis skill tool to score the relevance of at least one candidate search query term and obtaining a target score, the method further includes: if the target score corresponding to a candidate search query term is less than or equal to a preset threshold, then calling a search term rewriting skill tool to rewrite the candidate search query term to obtain a rewritten candidate search query term; calling a search term recommendation skill tool to score the relevance of the rewritten candidate search query term again until the rewritten candidate search query term is greater than the preset threshold.
[0165] Optionally, the processor may also execute program code that performs the following steps: based on information about multiple products in the image information, determine target scene information that matches the multiple products, including: calling a search tool to search for the multiple products contained in the image information to obtain product description information corresponding to the multiple products; and calling a scene extraction skill tool to perform relationship reasoning on the product description information and the image information to obtain target scene information.
[0166] Optionally, the processor may also execute program code that performs the following steps: calling a relevance analysis tool to score the relevance of at least one candidate search query term and obtain a target score, including: scoring the relevance of at least one candidate search query term and target scene information to obtain a first score; scoring the relevance of at least one candidate search query term and image information to obtain a second score; and obtaining the target score based on the first and second scores.
[0167] Optionally, the processor may also execute program code for the following steps: after obtaining the image information input by the target object, the method further includes: calling a recognition skill tool to recognize the products contained in the image information, and obtaining the proportion area corresponding to the products in the image information; based on the proportion area, determining the information of multiple products in the image information.
[0168] Optionally, the processor may also execute program code that performs the following steps: calling a scene extraction skill tool to perform relational reasoning on product description information and image information to obtain target scene information, including: identifying product description information to obtain corresponding product feature information, wherein the product feature information includes at least: product category information, product attribute information and association information between products; identifying image information to obtain global semantic information; and performing reasoning based on product feature information and global semantic information to obtain target scene information.
[0169] Those skilled in the art will understand that Figure 7The structure shown is for illustrative purposes only. Electronic device 70 can also be a smartphone, tablet computer, handheld computer, mobile internet device (MID), PAD and other terminal devices. Figure 7 This does not limit the structure of the aforementioned electronic device. For example, electronic device 70 may also include components that are more... Figure 7 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 7 The different configurations shown.
[0170] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0171] Example 5
[0172] Embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product can be used to store the program code executed by the recommended method for storing product information provided in Embodiment 1.
[0173] Optionally, in this embodiment, the computer program product may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0174] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0175] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0176] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0177] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0178] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0179] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0180] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for recommending product information, characterized in that, include: Obtain image information input by the target object, wherein the image information includes information about multiple products; Based on the information of multiple products in the image information, determine the target scene information that matches the multiple products; Based on the target scenario information, a target product to be recommended that matches the target scenario information is determined, and the information of the target product is pushed to the target object.
2. The method according to claim 1, characterized in that, Based on the target scenario information, determine the target products to be recommended that match the target scenario information, including: The search term recommendation tool is invoked to determine at least one candidate search query term based on the target scene information; Use a relevance analysis tool to score the relevance of the at least one candidate search query term and obtain a target score. If the target score is greater than a preset threshold, a search is performed based on the at least one candidate search query term to obtain the target product to be recommended.
3. The method according to claim 2, characterized in that, After invoking a relevance analysis tool to score the relevance of the at least one candidate search query term and obtaining a target score, the method further includes: If the target score value corresponding to a candidate search query term is less than or equal to the preset threshold, the search term rewriting tool is invoked to rewrite the candidate search query term to obtain the rewritten candidate search query term. The search term recommendation tool is invoked to re-evaluate the relevance of the rewritten candidate search terms until the rewritten candidate search terms exceed the preset threshold.
4. The method according to claim 1, characterized in that, Based on information about multiple products in the image, target scene information matching the multiple products is determined, including: The search tool is invoked to search for multiple products contained in the image information, and product description information corresponding to the multiple products is obtained; The scene extraction tool is invoked to perform relationship reasoning on the product description information and the image information to obtain the target scene information.
5. The method according to claim 2, characterized in that, The relevance analysis tool is invoked to score the relevance of the at least one candidate search query term, resulting in a target score, including: A relevance score is calculated between the at least one candidate search query term and the target scene information to obtain a first score. A relevance score is calculated for the at least one candidate search query term and the image information to obtain a second score. The target score is obtained based on the first score and the second score.
6. The method according to claim 1, characterized in that, After obtaining the image information input by the target object, the method further includes: The identification tool is invoked to identify the products contained in the image information, and the area ratio of the products in the image information is obtained. Based on the stated area proportion, information about multiple products in the image is determined.
7. The method according to claim 4, characterized in that, The scene extraction tool is invoked to perform relationship reasoning on the product description information and the image information to obtain the target scene information, including: The product description information is identified to obtain corresponding product feature information, wherein the product feature information includes at least: product category information, product attribute information and association information between products; The image information is identified to obtain global semantic information; The target scene information is obtained by reasoning based on the product feature information and the global semantic information.
8. A method for recommending product information, characterized in that, include: Obtain image information corresponding to the target object uploaded by the client, wherein the image information includes information about multiple products; Based on the information of multiple products in the image information, the cloud server determines the target scene information that matches the multiple products; based on the target scene information, it determines the target products to be recommended that match the target scene information. The information of the target product is pushed to the client.
9. A product information recommendation device, characterized in that, include: The acquisition unit is used to acquire image information input by the target object, wherein the image information includes information about multiple products; The first determining unit is used to determine target scene information that matches the multiple products based on the information of multiple products in the image information; The second determining unit is used to determine, based on the target scene information, a target product to be recommended that matches the target scene information, and to push the information of the target product to the target object.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the storage medium is located to perform the method for recommending product information as described in any one of claims 1 to 8.
11. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method for recommending product information as described in any one of claims 1 to 8.
12. A computer program product, characterized in that, Includes a computer program or instructions that, when executed by a processor, implement the method for recommending product information as described in any one of claims 1 to 8.