Method and device for recommending commodities, storage medium and electronic equipment
By acquiring user profiles and environmental information through smart glasses and generating structured product recommendation information using target models, the problem of insufficient efficiency and accuracy in e-commerce recommendations using smart glasses is solved. This enables efficient online product database-generated e-commerce recommendations and improves user experience.
Patent Information
- Application Number
- CN202510918853.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-11
AI Technical Summary
Existing smart glasses lack effective offline interaction methods in e-commerce recommendations, resulting in insufficient recommendation efficiency and accuracy, and a poor user experience.
By acquiring user profile information and environmental information of target users through smart glasses, multimodal retrieval is performed using the target model to generate structured product recommendation information, and the recommended products and attribute information are presented through screen or voice, combined with optical waveguide technology to realize augmented reality display.
It improves the efficiency and accuracy of e-commerce recommendations, enhances the user's interactive experience, and realizes an online product database generation e-commerce recommendation paradigm based on offline interaction.
Smart Images

Figure CN120931358A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to computer technology, and more particularly to a method, apparatus, storage medium, and electronic device for recommending products. Background Technology
[0002] Smart glasses are head-mounted devices that integrate microcomputer systems. Based on the form of eyeglasses, they merge digital information with the physical world through optical displays, voice interaction, and sensing technologies, expanding users' perception of their environment and ways to interact. In the early stages of smart glasses development, all smart glasses integrated microphone hardware, and voice interaction was a standard feature. Smart glasses were now fully equipped with multi-microphone arrays (e.g., five beamforming microphones), making voice interaction a standard function.
[0003] This year, most smart glasses manufacturers will integrate cameras, giving the glasses environmental awareness capabilities. Configuration parameters are evolving towards dual-camera setups (e.g., 1280P main camera + ToF depth sensor). With breakthroughs in waveguide display technology and interactive innovations, some glasses already possess waveguide capabilities, meaning they already have a certain level of display functionality. Waveguide solutions increase brightness to 3000 nits (40% more energy efficient than traditional solutions), supporting all-day use. Against the backdrop of the aforementioned developments in the payment industry and hardware for smart glasses, conditions have been created for seamless e-commerce recommendations using smart glasses. Summary of the Invention
[0004] The purpose of the embodiments in this specification is to provide a method, apparatus, storage medium, and electronic device for recommending products.
[0005] This specification provides a method for recommending products. Based on user profile information of a target user wearing smart glasses and environment-related information of the environment in which the smart glasses are located, product recommendation prompts are determined for the target user. These prompts are then input into a target model, which retrieves at least one recommended product from a product database. The product database includes structured understanding information for multiple products, including multiple attributes associated with the product and the values of each attribute. The method obtains the at least one recommended product and its corresponding attribute recommendation information output by the target model, enabling the smart glasses to provide the target user with the at least one recommended product and its corresponding attribute recommendation information based on their screen attributes. The method proposes a generative e-commerce recommendation paradigm based on offline interaction of smart glasses, in a scenario of light interaction with smart glasses. The online product library is structured using product attributes and attribute values for easy storage and retrieval. By integrating structured information understanding into a large model, the prompt information input to the large model not only links to static data such as user profile information (e.g., habits and occupation) but also integrates environmental information from the current real-world environment. This improves the efficiency and accuracy of e-commerce recommendations from the large model to users wearing smart glasses, significantly enhancing the user experience. The method includes:
[0006] Based on the user profile information of the target user who has already worn the smart glasses and the environmental information related to the environment in which the smart glasses are located, determine the product recommendation prompt information corresponding to the target user;
[0007] The product recommendation information is input into the target model, so that the target model retrieves at least one recommended product from the product database based on the product recommendation information. The product database includes structured understanding information corresponding to multiple products. The structured understanding information corresponding to the products includes multiple attributes associated with the products and the value of each attribute.
[0008] The smart glasses obtain at least one recommended product and the corresponding attribute recommendation information output by the target model, so that the smart glasses provide the at least one recommended product and the attribute recommendation information to the target user based on its screen attributes. The attribute recommendation information includes at least one attribute that matches the product recommendation prompt information from a plurality of attributes associated with the recommended product and the value of the at least one attribute.
[0009] Further, determining the product recommendation prompt information corresponding to the target user based on the user profile information of the target user who has already worn the smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located includes:
[0010] Based on the product recommendation input information, the user profile information corresponding to the target user, and the environment-related information corresponding to the environment in which the smart glasses are located, the product recommendation prompt information corresponding to the target user is determined, wherein the product recommendation input information includes the voice information input by the target user through the smart glasses.
[0011] Furthermore, the product recommendation input information also includes information about the target physical object;
[0012] The method further includes:
[0013] Extract the target physical object information from the current environment image corresponding to the smart glasses.
[0014] Further, the step of extracting the target physical object information from the current environment image corresponding to the smart glasses includes:
[0015] Based on the voice information, the target physical object information is extracted from the current environment screen corresponding to the smart glasses.
[0016] Further, the step of extracting the target physical object information from the current environment image corresponding to the smart glasses includes:
[0017] Based on the gesture interaction performed by the user, the target physical object information is extracted from the current environment screen corresponding to the smart glasses.
[0018] Furthermore, the structured understanding information includes text structured understanding information and image structured understanding information. The text structured understanding information includes multiple attributes associated with the product description text of the product and the value of each attribute. The image structured understanding information includes multiple attributes associated with the product description image of the product and the value of each attribute.
[0019] The method further includes:
[0020] Obtain product description text and product description images corresponding to the multiple products from at least one online store;
[0021] The structured understanding information of the text is obtained from the product description text, and the structured understanding information of the image is obtained from the product description image.
[0022] Further, obtaining the text structured understanding information based on the product description text includes:
[0023] The product description text is segmented into multiple word units;
[0024] Based on the multiple word units and the category attribute table corresponding to the product category to which the product belongs, multiple attributes associated with the product description text and the value of each attribute are obtained.
[0025] Furthermore, obtaining the image structured understanding information based on the product description image includes:
[0026] Based on the product description image, obtain multiple attributes associated with the product description image and the value of each attribute.
[0027] Furthermore, obtaining multiple attributes associated with the product description image and the value of each attribute based on the product description image includes:
[0028] Determine whether the product description image is a physical image of the product;
[0029] If so, extract the corresponding physical image of the product from the product description image;
[0030] Based on the extracted physical image and the category attribute table corresponding to the product category to which the product belongs, multiple attributes associated with the product description image and the value of each attribute are obtained.
[0031] Furthermore, the product database also includes at least one physical image corresponding to a product, enabling the target model to perform multimodal retrieval based on the structured understanding information corresponding to the multiple products and the physical image corresponding to the at least one product.
[0032] Furthermore, if the product description image is a physical image of the product, the image structured understanding information also includes the scene category corresponding to the product;
[0033] The method further includes:
[0034] Extract the background information corresponding to the product from the product description image;
[0035] The scenario category corresponding to the product is determined based on the background information.
[0036] Furthermore, the image structured understanding information also includes key image description information corresponding to the product;
[0037] The method further includes:
[0038] Based on multiple attributes associated with the product description image and the value of each attribute, the key description information of the image corresponding to the product is determined.
[0039] Furthermore, the image structured understanding information also includes multiple attributes associated with at least one text in the product description image and the value of each attribute;
[0040] The method further includes:
[0041] By performing optical character recognition on the product description image, at least one text located in the product description image can be obtained;
[0042] Obtain multiple attributes associated with the at least one text and the value of each attribute.
[0043] Furthermore, the method also includes:
[0044] Obtain the voice feedback information of the target user regarding the at least one recommended product provided by the smart glasses;
[0045] The target model re-outputs at least one latest recommended product and its corresponding attribute recommendation information based on the voice feedback information, so that the smart glasses re-provide the target user with the at least one latest recommended product and its corresponding attribute recommendation information based on their screen attributes.
[0046] Furthermore, the voice feedback information includes at least one attribute to be adjusted corresponding to the recommended product and / or attribute feedback corresponding to the at least one attribute to be adjusted.
[0047] Furthermore, the attribute recommendation information corresponding to the latest recommended product includes one or more of the at least one attribute to be adjusted and the value of the latest recommended product with respect to the one or more attributes to be adjusted.
[0048] This specification also provides an embodiment of a device for recommending products, comprising:
[0049] The prompt information determination module is used to determine the product recommendation prompt information corresponding to the target user based on the user profile information corresponding to the target user who has worn the smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located;
[0050] The retrieval module is used to input the product recommendation prompt information into the target model, so that the target model retrieves at least one recommended product from the product database based on the product recommendation prompt information. The product database includes structured understanding information corresponding to multiple products, and the structured understanding information corresponding to the products includes multiple attributes associated with the products and the value of each attribute.
[0051] An output module is used to obtain the at least one recommended product and the attribute recommendation information corresponding to the recommended product output by the target model, so that the smart glasses provide the at least one recommended product and the attribute recommendation information to the target user based on its screen attributes. The attribute recommendation information includes at least one attribute that matches the product recommendation prompt information among a plurality of attributes associated with the recommended product and the value of the at least one attribute.
[0052] This specification also provides a storage medium storing a computer program adapted to be loaded by a processor and to execute the steps of the method described above.
[0053] This specification also provides an electronic device, including a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method described above.
[0054] This specification also provides a computer program product that stores at least one instruction, characterized in that the at least one instruction, when executed by a processor, implements the steps of the above-described method.
[0055] According to the scheme of the embodiments of this specification, based on the user profile information corresponding to the target user who is wearing smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located, product recommendation prompts are determined for the target user; the product recommendation prompts are input into a target model, so that the target model retrieves at least one recommended product from a product database based on the product recommendation prompts, wherein the product database includes structured understanding information corresponding to multiple products, and the structured understanding information corresponding to the products includes multiple attributes associated with the products and the value of each attribute; the at least one recommended product and the attribute recommendation information corresponding to the recommended product output by the target model are obtained, so that the smart glasses provide the at least one recommended product and the attribute to the target user based on its screen attributes. The recommendation information, wherein the attribute recommendation information includes at least one attribute among multiple attributes associated with the recommended product that matches the product recommendation prompt information and the value of the at least one attribute, proposes an online product library generative e-commerce recommendation paradigm based on offline interaction of smart glasses in the scenario of light interaction of smart glasses. The online product library is structured by product attributes and attribute values to facilitate storage and retrieval, and is integrated into a large model by a structured understanding of information method. The prompt information input to the large model not only links to static data of user profile information (such as habits and occupations), but also integrates environmental information of the current real environment, which can improve the efficiency and accuracy of the large model in making e-commerce recommendations to users wearing smart glasses, and can significantly enhance the e-commerce recommendation user experience of users wearing smart glasses. Attached Figure Description
[0056] Figure 1 A flowchart illustrating a method for recommending products provided in an embodiment of this specification;
[0057] Figure 2 A flowchart illustrating a method for generating product recommendation prompts provided in an embodiment of this specification;
[0058] Figure 3 A flowchart illustrating a method for generating structured text understanding information provided in an embodiment of this specification;
[0059] Figure 4 A flowchart illustrating a method for generating structured understanding information from images, provided in an embodiment of this specification.
[0060] Figure 5 A flowchart illustrating a method for recommending products provided in an embodiment of this specification;
[0061] Figure 6 A schematic diagram of a device for recommending products provided in an embodiment of this specification;
[0062] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0064] Please see Figure 1 This is a flowchart illustrating a method for recommending products provided in an embodiment of this specification. In this embodiment, the method for recommending products is applied to a device for recommending products (hereinafter referred to as a "product recommendation device") or an electronic device equipped with a product recommendation device. The following will focus on... Figure 1 The process shown will be described in detail. The method for recommending products may specifically include the following steps:
[0065] S102, based on the user profile information of the target user who has worn the smart glasses and the environmental information related to the environment in which the smart glasses are located, determine the product recommendation prompt information corresponding to the target user.
[0066] In some embodiments, the product recommendation device in the embodiments of this specification may be smart glasses, or it may be a server that communicates with the smart glasses.
[0067] In some embodiments, user profile information includes, but is not limited to, any information related to the target user, such as spending power, occupational tags, spending amounts, cycles, category preferences, recent purchase records, and habits in the user's historical bills, etc. This example embodiment does not impose any special limitations on this. In some embodiments, user profile information may simultaneously include multiple pieces of information in a multimodal manner, such as simultaneously including multiple pieces of information such as spending power, occupational tags, and recent purchase records and habits.
[0068] In some embodiments, the environment-related information corresponding to the smart glasses' location (i.e., the real environment in which the smart glasses are currently located) includes, but is not limited to, any information related to that environment, such as weather information, location information, and POI (Point of Interest) information. This example embodiment does not impose any special limitations on this. In some embodiments, if the product recommendation device is a server, the smart glasses need to send the environment-related information to the server. In some embodiments, the environment-related information may simultaneously include multiple modal pieces of information related to the real environment, such as weather information, location information, and POI information.
[0069] In some embodiments, the product recommendation prompt is a prompt input to the target model (i.e., the model used to recommend products, whose output is products recommended to the target user). The prompt is an instruction or text input for interacting with the model, guiding the model to output specific content through natural language; its essence is to provide the model with task instructions or contextual hints. In some embodiments, the product recommendation prompt is generated based on user profile information and environment-related information, through multi-source data fusion. For example, the product recommendation prompt can be generated based on preset generation rules, or it can be obtained by inputting user profile information and environment-related information into a trained prompt generation model to obtain the product recommendation prompt output by that model. This example embodiment does not specifically limit the specific generation method of the product recommendation prompt. For example, the product recommendation prompt could be "The user is a tour guide, their usual consumption habit is to prefer casual clothing, today's weather is sunny, the user is in a shopping mall…". In some embodiments, the product recommendation prompt is input into the target model, and the product recommendation prompt is generated based on multimodal user profile information and environment-related information; that is, the input to the target model is essentially multimodal information.
[0070] S104, the product recommendation prompt information is input into the target model, so that the target model retrieves at least one recommended product from the product database based on the product recommendation prompt information. The product database includes structured understanding information corresponding to multiple products, and the structured understanding information corresponding to the products includes multiple attributes associated with the products and the value of each attribute.
[0071] In some embodiments, the target model is a generative large model for recommending products. The output of the target model is the products recommended to the target user. This example embodiment does not specifically limit the model structure and model parameters of the target model. In some embodiments, in addition to linking the static data of the target user through user profile information, this specification embodiment also integrates multimodal information of the real environment through environment-related information, enabling the target model to perform multimodal alignment and generative training based on a prompt generated from training data containing user profile information and environment-related information. In some embodiments, the product database obtains product information of multiple products located in at least one online shopping mall by accessing at least one online shopping mall. Then, it generates structured understanding information corresponding to each product based on the product information and stores it in the product database. The product information includes, but is not limited to, any information related to the product, such as product name, product title, product description, product price, product image, etc. This example embodiment does not specifically limit this. In some embodiments, the product database is connected to at least one online marketplace, meaning it is linked to at least one online marketplace. For example, the product database is connected to multiple online marketplaces. The product database has a corresponding cross-platform data federation access layer and a multimodal feature vectorization engine. Vectorization techniques such as BERT (Bidirectional Encoder Representations from Transformers) are used to semantically encode the product text information. Combined with the CLIP (Contrastive Language-Image Pre-training) model, cross-modal feature alignment between text and images is achieved. A billion-level vector database is built through indexing, supporting millisecond-level approximate nearest neighbor retrieval. In some embodiments, user profile information can be constructed and cross-platform jointly modeled with the product database of at least one online marketplace to ensure that the original data does not leave the domain.In some embodiments, structured understanding information is used to enable the target model to understand products in the product database in a structured manner. The structured understanding information corresponding to a product includes multiple attributes related to the product (e.g., material, safety level) and the value of the product for each attribute (e.g., the material value is 316 stainless steel, and the safety level value is food grade). That is, the structured understanding information can be directly the multiple attributes related to the product and the value of the product for each attribute. Alternatively, the structured understanding information can also be information generated based on the multiple attributes related to the product and the value of the product for each attribute. For example, the information can be a summary of the multiple attributes. For example, the multiple attributes and the value of each attribute can be concatenated using a preset concatenation rule to obtain the structured understanding information. Alternatively, the multiple attributes and the value of each attribute can be input into a trained model to obtain the structured understanding information output by the model. This example embodiment does not specifically limit the specific method of obtaining the structured understanding information. In some embodiments, not only is the structured understanding information corresponding to each product stored in the product database, but also the multiple attributes associated with each product and the value of each attribute are stored in the product database. In some embodiments, after the product recommendation prompt information is input into the target model, the target model will search for the structured understanding information corresponding to multiple products stored in the product database based on the product recommendation prompt information, retrieve the structured understanding information that matches the product recommendation prompt information, and use the product corresponding to the structured understanding information as the recommended product suitable for recommendation to the target user.
[0072] S106, obtain the at least one recommended product and the corresponding attribute recommendation information output by the target model, so that the smart glasses provide the at least one recommended product and the attribute recommendation information to the target user based on its screen attributes. The attribute recommendation information includes at least one attribute among multiple attributes associated with the recommended product that matches the product recommendation prompt information, and the value of the at least one attribute. In some embodiments, the target model outputs at least one recommended product and corresponding attribute recommendation information for each recommended product to the target user. The attribute recommendation information is used to characterize the reason for recommending the product to the target user. The attribute recommendation information includes at least one attribute among multiple attributes associated with the recommended product that matches the product recommendation prompt information (e.g., color, price, brand, material, etc.) and the value of the recommended product regarding the at least one attribute. In some embodiments, if the product recommendation device is a server, the server needs to send the at least one recommended product and the corresponding attribute recommendation information to the smart glasses. In some embodiments, smart glasses provide the target user with at least one recommended product and the attribute recommendation information based on their screen attributes. Screen attributes include both having a screen and not having a screen. That is, in smart glasses with a screen, the at least one recommended product and the corresponding attribute recommendation information are directly presented on the screen of the smart glasses. For example, augmented reality overlay technology is used to project the recommended product and the corresponding attribute recommendation information onto the screen of the smart glasses. Another example is to display the recommended product and the corresponding attribute recommendation information on the screen of the smart glasses through an optical waveguide. This example embodiment does not specifically limit the specific presentation method. In smart glasses without a screen, the at least one recommended product and the corresponding attribute recommendation information are broadcast through voice.
[0073] According to the scheme of the embodiments of this specification, based on the user profile information corresponding to the target user who is wearing smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located, product recommendation prompts are determined for the target user; the product recommendation prompts are input into a target model, so that the target model retrieves at least one recommended product from a product database based on the product recommendation prompts, wherein the product database includes structured understanding information corresponding to multiple products, and the structured understanding information corresponding to the products includes multiple attributes associated with the products and the value of each attribute; the at least one recommended product and the attribute recommendation information corresponding to the recommended product output by the target model are obtained, so that the smart glasses provide the at least one recommended product and the attribute to the target user based on its screen attributes. The recommendation information, wherein the attribute recommendation information includes at least one attribute among multiple attributes associated with the recommended product that matches the product recommendation prompt information and the value of the at least one attribute, proposes an online product library generative e-commerce recommendation paradigm based on offline interaction of smart glasses in the scenario of light interaction of smart glasses. The online product library is structured by product attributes and attribute values to facilitate storage and retrieval, and is integrated into a large model by a structured understanding of information method. The prompt information input to the large model not only links to static data of user profile information (such as habits and occupations), but also integrates environmental information of the current real environment, which can improve the efficiency and accuracy of the large model in making e-commerce recommendations to users wearing smart glasses, and can significantly enhance the e-commerce recommendation user experience of users wearing smart glasses.
[0074] In some embodiments, determining the product recommendation prompt information corresponding to the target user based on the user profile information corresponding to the target user wearing the smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located includes: determining the product recommendation prompt information corresponding to the target user based on product recommendation input information, the user profile information corresponding to the target user, and the environment-related information corresponding to the environment in which the smart glasses are located, wherein the product recommendation input information includes voice information input by the target user through the smart glasses. In some embodiments, product recommendation input information refers to information input by the target user to the target model through the smart glasses. The product recommendation input information is used by the target user to indicate to the target model that they need it to recommend corresponding products to them. The product recommendation input information includes voice information input by the target user through the smart glasses, for example, the voice information is "It's a bit cold, buy me a trendy coat this year." In some embodiments, the corresponding product recommendation prompt information is generated based on the product recommendation input information, user profile information, and environment-related information. In some embodiments, if the product recommendation device is a server, the smart glasses also need to send the product recommendation input information to the server. In some embodiments, the product recommendation input information also includes the current real-world environment image corresponding to the target user obtained by the smart glasses through their camera (the camera of the smart glasses will collect the surrounding real-world environment image in real time).
[0075] In some embodiments, the product recommendation input information further includes target physical object information; wherein, the method further includes: extracting the target physical object information from the current environment image corresponding to the smart glasses. In some embodiments, the product recommendation input information further includes target physical object information extracted from the current real-world environment image corresponding to the target user obtained by the smart glasses through its camera. The target physical object information may be an image containing a target physical object (e.g., a piece of clothing) captured from the current real-world environment image, or it may be relevant descriptive information (e.g., color, style, brand, etc.) corresponding to the target physical object obtained by performing image content recognition on the image or the current real-world environment image. In some embodiments, the smart glasses will use computer vision technology to analyze the real-world environment image captured by the camera in real time, and combine it with object detection algorithms (such as YOLO or Faster RCNN) to identify physical objects located in the real-world environment image. In some embodiments, the product recommendation input information may simultaneously include multiple real-time inputs of multiple modalities, such as simultaneously including at least two of the following: voice information, current real-world environment image, and target physical object information.
[0076] In some embodiments, extracting the target object information from the current environment image corresponding to the smart glasses includes: extracting the target object information from the current environment image corresponding to the smart glasses based on the voice information. In some embodiments, the target object information can be extracted from the current real environment image corresponding to the target user obtained by the smart glasses through its camera based on the target user's voice information. For example, if the voice information is "It's a bit cold, buy me a trendy coat this year, you can refer to the style in front of me," then the target object information corresponding to the coat can be extracted from the current real environment image based on this voice information.
[0077] In some embodiments, extracting the target object information from the current environment corresponding to the smart glasses includes: extracting the target object information from the current environment corresponding to the smart glasses based on the user's gesture interaction. In some embodiments, the user's current gesture interaction can be captured by the smart glasses' camera, and the target object information can be extracted from the target user's current real-world environment based on gesture recognition. For example, if the user's gesture interaction is a pointing gesture with their index finger, then the target object information corresponding to the object pointed to by the user's index finger can be extracted from the target user's current real-world environment based on this gesture interaction.
[0078] In some embodiments, the structured understanding information includes text structured understanding information and image structured understanding information. The text structured understanding information includes multiple attributes associated with the product description text of the product and the value of each attribute. The image structured understanding information includes multiple attributes associated with the product description image of the product and the value of each attribute. The method further includes: obtaining product description text and product description images corresponding to the multiple products from at least one online marketplace; obtaining the text structured understanding information based on the product description text; and obtaining the image structured understanding information based on the product description image. In some embodiments, the structured understanding information includes text structured understanding information and image structured understanding information. The text structured understanding information is used to enable the target model to understand the product description text of each product in the product database in a structured manner, including multiple attributes associated with the product description text and the value of each attribute. The image structured understanding information is used to enable the target model to understand the product description images of each product in the product database in a structured manner, including multiple attributes associated with the product description image and the value of each attribute. The product description text includes, but is not limited to, any text-type information used to describe the product, such as product introductions or product titles. The product description images include, but are not limited to, any image-type product images used to describe the product. In some embodiments, the text structured understanding information is obtained by concatenating multiple attributes associated with the product description text and the value of each attribute using preset concatenation rules. Alternatively, it can be obtained by inputting multiple attributes associated with the product description text and the value of each attribute into a trained model to obtain the text structured understanding information output by the model. In some embodiments, image structured understanding information is obtained by concatenating multiple attributes associated with a product description image and the values of each attribute using preset concatenation rules. Alternatively, it can be obtained by inputting multiple attributes associated with a product description image and the values of each attribute into a trained model to obtain the image structured understanding information output by the model. In some embodiments, it is necessary to first obtain product description text and product description images corresponding to multiple products from at least one online marketplace connected to the product database. Then, the text structured understanding information is obtained based on the product description text, and the image structured understanding information is obtained based on the product description images. For example, semantic understanding of the product description text can be used to obtain multiple attributes associated with the product description text and the values of each attribute. Alternatively, image content analysis of the product description images can be used to obtain multiple attributes associated with the product description images and the values of each attribute.
[0079] In some embodiments, obtaining the structured understanding information of the text based on the product description text includes: segmenting the product description text into multiple word units; and obtaining multiple attributes associated with the product description text and the value of each attribute based on the multiple word units and the category attribute table corresponding to the product category to which the product belongs. In some embodiments, the product description text can be segmented into multiple word units first. For example, for the product description text "316 stainless steel mixing basin, extra thick, household food-grade stainless steel basin, new kitchen vegetable washing basin", the word units obtained after segmentation include "316 stainless steel", "mixing basin", "extra thick", "household", "food grade", "stainless steel basin", "new", "kitchen", and "vegetable washing basin". This example embodiment does not specifically limit the specific segmentation method. In some embodiments, multiple category attribute tables corresponding to product categories are pre-defined. Different product categories typically correspond to different category attribute tables. Each category attribute table contains multiple attributes belonging to that product category. Based on the multiple word units and the multiple attributes belonging to that product category (e.g., kitchenware) contained in the category attribute table, an association relationship is established between the multiple word units and the multiple attributes. For each attribute, its value is determined based on at least one word unit associated with it from the multiple word units. This yields multiple attributes associated with the product description text and the value of each attribute. For example, the at least one word unit can be directly used as the attribute value, or the text content generated based on the at least one word unit can be used as the attribute value. In some embodiments, if the value of an attribute cannot be determined based on the multiple word units, the attribute value is defaulted. For example, the category attribute table for kitchenware includes "material", "use", "appearance characteristics", "suitable scenarios", "safety level" and "style". Based on the multiple word units obtained after word segmentation, the value of the "material" attribute is determined to be "316 stainless steel", the value of the "use" attribute is "dough basin, vegetable washing basin", the value of the "appearance characteristics" attribute is "extra thick", the value of the "suitable scenarios" attribute is "household, kitchen", the value of the "safety level" attribute is "food grade", and the value of the "style" attribute is "new model".
[0080] In some embodiments, obtaining the image structured understanding information based on the product description image includes: obtaining multiple attributes associated with the product description image and the value of each attribute based on the product description image. In some embodiments, multiple attributes associated with the product description image and the value of each attribute are determined from the corresponding image content by performing image content analysis on the product description image.
[0081] In some embodiments, obtaining multiple attributes associated with the product description image and the value of each attribute based on the product description image includes: determining whether the product description image belongs to a physical object type image of the product; if so, extracting a physical object extraction image corresponding to the product from the product description image; and obtaining multiple attributes associated with the product description image and the value of each attribute based on the physical object extraction image and the category attribute table corresponding to the product category to which the product belongs. In some embodiments, the product description image is first determined to be a physical object type image or a non-physical object type image of the product based on whether it contains the physical object of the product. If so, it is necessary to first extract a physical object extraction image containing the product (e.g., a block diagram containing the product) from the product description image. This physical object extraction image does not contain any other objects besides the product, or the physical object extraction image does not contain any other complete objects besides the product. In some embodiments, based on the extracted physical image and the category attribute table corresponding to the product category to which the product belongs, multiple attributes associated with the product description image and the value of each attribute are obtained. Specifically, image content analysis is performed on the extracted physical image to obtain the corresponding image content, and based on the multiple attributes belonging to the product category contained in the category attribute table corresponding to the product category, the value of each of the multiple attributes is determined according to the image content, thereby obtaining multiple attributes associated with the product description image and the value of each attribute. In some embodiments, if the product description image is a non-physical image of the product, multiple attributes associated with the product description image and the value of each attribute are directly obtained based on the product description image and the category attribute table corresponding to the product category to which the product belongs. Specifically, image content analysis is performed on the product description image to obtain the corresponding image content, and based on the multiple attributes belonging to the product category contained in the category attribute table corresponding to the product category, the value of each of the multiple attributes is determined according to the image content, thereby obtaining multiple attributes associated with the product description image and the value of each attribute.
[0082] In some embodiments, the product database further includes at least one physical image corresponding to a product, enabling the target model to perform multimodal retrieval based on the structured understanding information corresponding to the plurality of products and the physical image corresponding to the at least one product. In some embodiments, the product database further includes at least one physical image corresponding to a product (i.e., not all products have corresponding physical images), thereby enabling the target model to perform multimodal text-image retrieval based on the structured understanding information in text form corresponding to the plurality of products in the product database and the physical image in image form corresponding to at least one product, retrieving from the product database at least one recommended product that matches the product recommendation prompt information and is suitable for recommendation to the target user.
[0083] In some embodiments, if the product description image is a physical image of the product, the image structured understanding information also includes the scene classification corresponding to the product; wherein, the method further includes: extracting background information corresponding to the product from the product description image; and determining the scene classification corresponding to the product based on the background information. In some embodiments, if the product description image is a physical image of the product, the image structured understanding information also includes the scene classification corresponding to the product. It is necessary to first extract the background information corresponding to the product from the product description image (e.g., the background image remaining after removing the physical product from the product description image), and then determine the scene classification corresponding to the product based on the background information. The scene classification is used to characterize the category of the scene in which the physical product may be located. For example, the background information can be input into a trained classification model to obtain the scene classification output by the model. Alternatively, the scene classification corresponding to the product can be determined by performing color analysis and / or edge detection and / or texture analysis and / or object recognition on the background information. This example embodiment does not specifically limit the specific method of determining the scene classification based on the background information.
[0084] In some embodiments, the image structured understanding information further includes key image description information corresponding to the product; wherein, the method further includes: determining key image description information corresponding to the product based on multiple attributes associated with the product description image and the value of each attribute. In some embodiments, the image structured understanding information also includes key image description information corresponding to the product. The key image description information is used to summarize the key content of multiple attributes associated with the product description image and the values of each attribute to enhance the target model's understanding of the product. For example, multiple key attributes are selected from multiple attributes associated with the product description image, and the key image description information is generated based on the multiple key attributes and the values of the multiple key attributes. That is, the key image description information is a selective understanding of the product regarding multiple attributes associated with the product description image and the values of each attribute (i.e., key knowledge of the product description image, integrating the product description image into the target model through key knowledge). For example, several attributes with the highest correlation to the product are selected from the multiple attributes as key attributes. Or, several attributes that best highlight the selling points of the product are selected from the multiple attributes as key attributes. Or, several attributes with the highest importance to the product are selected from the multiple attributes as key attributes. This example embodiment does not impose any special limitations on this. In some embodiments, the key image description information output by the model can be obtained by inputting multiple attributes associated with the product description image and the values of each attribute into a trained model. Alternatively, the key image description information output by the model can be obtained by inputting multiple attributes associated with the product description image, the values of each attribute, and the scene classification corresponding to the product determined based on the background information as described above into the trained model. Alternatively, the key image description information output by the model can be obtained by inputting multiple attributes associated with the product description image, the values of each attribute, and some other related product information (such as product name, product title, product category, etc.) into the model. Alternatively, the model can be obtained by inputting multiple attributes associated with the product description image and the values of each attribute, the scene classification corresponding to the product determined based on the background information as described above, and some other related product information (such as product name, product title, product category, etc.) into the model. This example embodiment does not impose any special limitations on this, so that the target model can perform a search based on the image key description information corresponding to multiple products in the product database, retrieve the image key description information that matches the product recommendation prompt information from the product database, and use the product corresponding to the image key description information as a recommended product suitable for recommendation to the target user.
[0085] In some embodiments, the image structured understanding information further includes multiple attributes associated with at least one text in the product description image and the value of each attribute; wherein, the method further includes: obtaining at least one text located in the product description image by performing optical character recognition (OCR) on the product description image; obtaining multiple attributes associated with the at least one text and the value of each attribute. In some embodiments, the image structured understanding information further includes multiple attributes associated with at least one text in the product description image and the value of each attribute. At least one text located in the product description image is obtained by performing optical character recognition (OCR) on the product description image, and then multiple attributes associated with the at least one text and the value of each attribute are obtained by performing semantic understanding on the at least one text. In some embodiments, the at least one text can first be segmented into words, and then based on several word units obtained after segmentation and the category attribute table corresponding to the product category to which the product belongs, multiple attributes associated with the at least one text and the value of each attribute are obtained. The specific method of obtaining these attributes is the same as or similar to the method for obtaining multiple attributes associated with the product description text and the value of each attribute described above, and will not be repeated here.
[0086] In some embodiments, the method further includes: obtaining voice feedback information from the target user regarding the at least one recommended product provided by the smart glasses; causing the target model to re-output at least one latest recommended product and attribute recommendation information corresponding to the latest recommended product based on the voice feedback information; and causing the smart glasses to re-provide the at least one latest recommended product and attribute recommendation information corresponding to the latest recommended product to the target user based on its screen attributes. In some embodiments, obtaining voice feedback information input by the target user through the smart glasses regarding at least one recommended product currently provided to them by the smart glasses, i.e., voice feedback information, which is used to characterize the target user's feedback on the recommended product (i.e., negative feedback of the recommendation information output by the target model). In some embodiments, by inputting the voice feedback information into the target model, or by first converting the voice feedback information into corresponding text information and then inputting the text information into the target model, at least one latest recommended product and its corresponding attribute recommendation information re-output by the target model are obtained. This latest recommended product is different from the recommended product previously recommended to the target user by the target model, and enables the smart glasses to re-provide the at least one latest recommended product and its corresponding attribute recommendation information to the target user based on its screen attributes. For example, it can be directly presented on the screen of the smart glasses, or it can be broadcast through voice. The attribute recommendation information includes at least one attribute among multiple attributes associated with the latest recommended product that matches the product recommendation prompt information, and the value of the latest recommended product with respect to the at least one attribute.
[0087] In some embodiments, the voice feedback information includes at least one attribute to be adjusted corresponding to the recommended product and / or attribute feedback corresponding to the at least one attribute to be adjusted. In some embodiments, the target user can provide feedback in the voice feedback information regarding at least one attribute to be adjusted that they are dissatisfied with in the recommended products currently provided to them by the smart glasses. That is, the target user expects the latest recommended product re-recommended by the target model to have a different value for the at least one attribute to be adjusted compared to the current value of the recommended product currently provided to the target user by the smart glasses. In some embodiments, the voice feedback information may further include attribute feedback from the target user regarding the at least one attribute to be adjusted. This attribute feedback may include the expected value of the attribute to be adjusted, i.e., the target user expects the latest recommended product re-recommended by the target model to have the expected value for the attribute to be adjusted. This expected value differs from the current value of the recommended product currently provided to the target user by the smart glasses regarding the attribute to be adjusted. For example, if the attribute to be adjusted is color, the attribute feedback would be "expected color is red." Alternatively, the attribute feedback may also include the adjustment direction of the attribute to be adjusted, which characterizes how the attribute to be adjusted should be adjusted based on its current value. That is, the target user expects the target model to re-recommend the latest recommended product based on the adjustment direction of the attribute to be adjusted. For example, if the attribute to be adjusted is color, the attribute feedback would be "the color should be darker." Alternatively, the attribute feedback may also include negative feedback from the target user regarding the attribute to be adjusted, i.e., the target user expects the target model to re-recommend the latest recommended product based on the target user's negative feedback regarding the attribute to be adjusted. For example, if the attribute to be adjusted is price, the attribute feedback would be "the price is a bit expensive." In some embodiments, the target model will adjust based on at least one attribute to be adjusted from the target attribute feedback and / or the attribute feedback corresponding to the at least one attribute to be adjusted, so that the latest value of the recommended product re-recommended by the target model to the target user with respect to the at least one attribute to be adjusted is as different as possible from the current value of the recommended product currently provided to the target user by the smart glasses with respect to the at least one attribute to be adjusted, and the latest value of the at least one attribute to be adjusted is adjusted by the target model according to the attribute feedback.
[0088] In some embodiments, the attribute recommendation information corresponding to the latest recommended product includes one or more of the at least one attribute to be adjusted and the value of the latest recommended product with respect to the one or more attributes to be adjusted. In some embodiments, the attribute recommendation information corresponding to the latest recommended product re-output by the target model includes one or more of the at least one attribute to be adjusted and the value of the latest recommended product with respect to the one or more attributes to be adjusted. In some embodiments, the one or more attributes to be adjusted still need to match the product recommendation prompt information.
[0089] Figure 2 This is a flowchart illustrating a method for generating product recommendation prompts provided in an embodiment of this specification.
[0090] like Figure 2 As shown, the system obtains environmental information related to the smart glasses' location through a real-time interface. This environmental information includes weather information, LBS (Location Based Services) information, and POI (Point of Interest) information. It also obtains user profile information for the target user wearing the smart glasses from an internal database. This user profile information includes the target user's recent purchase history and habits, occupation, and other information. Then, based on the environmental information and user profile information, it generates corresponding product recommendation prompts, such as "The user is a tour guide, their usual consumption habit is to prefer casual clothing, today the weather is sunny, the user is in a shopping mall…".
[0091] Figure 3 This is a flowchart illustrating a method for generating structured textual understanding information, provided in an embodiment of this specification.
[0092] like Figure 3 As shown, by segmenting the product description text into words, multiple word units are obtained. Based on the multiple word units and the category attribute table corresponding to the product category to which the product belongs, the large model obtains the attribute summary of the product description text, that is, multiple attributes associated with the product description text and the value of each attribute.
[0093] Figure 4 This is a flowchart illustrating a method for generating structured understanding information from images, as provided in an embodiment of this specification.
[0094] like Figure 4As shown, by performing Optical Character Recognition (OCR) on the product description image, at least one text located in the product description image is obtained. Then, by performing semantic understanding on the at least one text, the corresponding text structure understanding is obtained, that is, multiple attributes associated with the at least one text and the value of each attribute. Then, it is determined whether the product description image belongs to the physical image of the product. If it belongs to the physical image, the physical image corresponding to the product is extracted from the product description image. Based on the physical image and the category attribute table corresponding to the product category to which the product belongs, the corresponding attribute structure understanding is obtained, that is, multiple attributes associated with the product description image and the value of each attribute. By extracting the background information corresponding to the product from the product description image, the scene classification corresponding to the product is determined according to the background information. And based on the attribute structure understanding and the scene classification, the product is classified accordingly. Scene classification determines the key descriptive information (i.e., caption) of the image corresponding to the product. If it is not a physical image, the corresponding attribute structured understanding is obtained based on the product description image and the category attribute table corresponding to the product category to which the product belongs. That is, multiple attributes associated with the product description image and the value of each attribute. The key descriptive information (i.e., caption) of the image corresponding to the product is determined based on the attribute structured understanding. Then, the image structured understanding corresponding to the product is obtained based on text structured understanding, attribute structured understanding, caption, and scene classification. Image structured understanding is used to enable the target model to understand the product description images of each product in the product database in a structured way. In addition, if the product description image is a physical image, the corresponding physical image will also be obtained.
[0095] Figure 5 This is a flowchart illustrating a method for recommending products as provided in an embodiment of this specification.
[0096] like Figure 5 As shown, the system extracts environmental information, including the user's voice, the surrounding environment, and gesture inputs. It also obtains environmental and user information, such as LBS location information and weather, and user habits. By integrating multimodal information including the user's purchase history, environmental information, user information, and the extracted voice, surrounding environment, and gesture inputs, the system generates a corresponding prompt. This prompt is then input into a generative model, which performs a multimodal search on a structured online product database based on the prompt and outputs recommended products. The user can then provide real-time feedback on the recommended products. The generative model adjusts the recommended products based on the user's feedback on at least one of the adjustable attributes (core attributes) and outputs new recommended products.
[0097] Figure 6 This is a schematic diagram of a device for recommending products, provided as an embodiment of this specification. This device (hereinafter referred to as "product recommendation device 1") can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, the product recommendation device 1 includes a prompt information determination module 11, a retrieval module 12, and an output module 13.
[0098] The prompt information determination module 11 is used to determine the product recommendation prompt information corresponding to the target user based on the user profile information corresponding to the target user who has worn the smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located;
[0099] The retrieval module 12 is used to input the product recommendation prompt information into the target model, so that the target model retrieves at least one recommended product from the product database based on the product recommendation prompt information. The product database includes structured understanding information corresponding to multiple products, and the structured understanding information corresponding to the products includes multiple attributes associated with the products and the value of each attribute.
[0100] Output module 13 is used to obtain the at least one recommended product and the attribute recommendation information corresponding to the recommended product output by the target model, so that the smart glasses provide the at least one recommended product and the attribute recommendation information to the target user based on its screen attributes. The attribute recommendation information includes at least one attribute that matches the product recommendation prompt information among a plurality of attributes associated with the recommended product and the value of the at least one attribute.
[0101] In some embodiments, determining the product recommendation prompt information corresponding to the target user based on the user profile information corresponding to the target user who is wearing smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located includes: determining the product recommendation prompt information corresponding to the target user based on product recommendation input information, the user profile information corresponding to the target user and the environment-related information corresponding to the environment in which the smart glasses are located, wherein the product recommendation input information includes the voice information input by the target user through the smart glasses.
[0102] In some embodiments, the product recommendation input information further includes target physical object information; wherein, the product recommendation device 1 is further configured to: extract the target physical object information from the current environment screen corresponding to the smart glasses.
[0103] In some embodiments, extracting the target physical object information from the current environment screen corresponding to the smart glasses includes: extracting the target physical object information from the current environment screen corresponding to the smart glasses based on the voice information.
[0104] In some embodiments, extracting the target physical object information from the current environment screen corresponding to the smart glasses includes: extracting the target physical object information from the current environment screen corresponding to the smart glasses based on the gesture interaction operation performed by the user.
[0105] In some embodiments, the structured understanding information includes text structured understanding information and image structured understanding information. The text structured understanding information includes multiple attributes associated with the product description text of the product and the value of each attribute. The image structured understanding information includes multiple attributes associated with the product description image of the product and the value of each attribute. The product recommendation device 1 is further configured to: obtain product description text and product description images corresponding to the multiple products from at least one online marketplace; obtain the text structured understanding information based on the product description text; and obtain the image structured understanding information based on the product description image.
[0106] In some embodiments, obtaining the text structured understanding information based on the product description text includes: segmenting the product description text into multiple word units; and obtaining multiple attributes associated with the product description text and the value of each attribute based on the multiple word units and the category attribute table corresponding to the product category to which the product belongs.
[0107] In some embodiments, obtaining the image structured understanding information based on the product description image includes: obtaining multiple attributes associated with the product description image and the value of each attribute based on the product description image.
[0108] In some embodiments, obtaining multiple attributes associated with the product description image and the value of each attribute based on the product description image includes: determining whether the product description image belongs to the physical type image of the product; if so, extracting the physical image corresponding to the product from the product description image; and obtaining multiple attributes associated with the product description image and the value of each attribute based on the physical image and the category attribute table corresponding to the product category to which the product belongs.
[0109] In some embodiments, the product database further includes at least one physical image corresponding to a product, enabling the target model to perform multimodal retrieval based on the structured understanding information corresponding to the plurality of products and the physical image corresponding to the at least one product.
[0110] In some embodiments, if the product description image is a physical image of the product, the image structured understanding information further includes the scene category corresponding to the product; wherein, the product recommendation device 1 is further configured to: extract the background information corresponding to the product from the product description image; and determine the scene category corresponding to the product based on the background information.
[0111] In some embodiments, the image structured understanding information further includes key image description information corresponding to the product; wherein, the product recommendation device 1 is further configured to: determine the key image description information corresponding to the product based on multiple attributes associated with the product description image and the value of each attribute.
[0112] In some embodiments, the image structured understanding information further includes multiple attributes associated with at least one text in the product description image and the value of each attribute; wherein, the product recommendation device 1 is further configured to: obtain at least one text located in the product description image by performing optical character recognition on the product description image; and obtain multiple attributes associated with the at least one text and the value of each attribute.
[0113] In some embodiments, the product recommendation device 1 is further configured to: obtain voice feedback information from the target user regarding the at least one recommended product provided by the smart glasses; cause the target model to re-output at least one latest recommended product and attribute recommendation information corresponding to the latest recommended product based on the voice feedback information; and cause the smart glasses to re-provide the at least one latest recommended product and attribute recommendation information corresponding to the latest recommended product to the target user based on its screen attributes.
[0114] In some embodiments, the voice feedback information includes at least one attribute to be adjusted corresponding to the recommended product and / or attribute feedback corresponding to the at least one attribute to be adjusted.
[0115] In some embodiments, the attribute recommendation information corresponding to the latest recommended product includes one or more of the at least one attribute to be adjusted and the value of the latest recommended product with respect to the one or more attributes to be adjusted.
[0116] The above-described apparatus embodiments correspond to the aforementioned method embodiments. For detailed descriptions, please refer to the description in the method embodiments section; further details will not be repeated here. The apparatus embodiments are derived from the corresponding method embodiments and have the same technical effects. For detailed descriptions, please refer to the corresponding method embodiments.
[0117] This specification also provides a computer storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in this specification.
[0118] This specification also provides a computer program product that stores at least one instruction, which is loaded by the processor and executes the method described in this specification embodiment.
[0119] This specification also provides an electronic device, including a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and execute the method described in the embodiments of this specification.
[0120] The embodiments in this specification also provide Figure 7 The diagram shows the structure of the electronic device. Figure 7 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the above method.
[0121] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0122] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0123] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0124] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0125] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0126] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0127] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0128] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0129] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for recommending products, comprising: Based on the user profile information of the target user who has already worn the smart glasses and the environmental information related to the environment in which the smart glasses are located, determine the product recommendation prompt information corresponding to the target user; The product recommendation information is input into the target model, so that the target model retrieves at least one recommended product from the product database based on the product recommendation information. The product database includes structured understanding information corresponding to multiple products. The structured understanding information corresponding to the products includes multiple attributes associated with the products and the value of each attribute. The smart glasses obtain at least one recommended product and the corresponding attribute recommendation information output by the target model, so that the smart glasses provide the at least one recommended product and the attribute recommendation information to the target user based on its screen attributes. The attribute recommendation information includes at least one attribute that matches the product recommendation prompt information from a plurality of attributes associated with the recommended product and the value of the at least one attribute.
2. The method according to claim 1, wherein determining the product recommendation prompt information corresponding to the target user based on the user profile information corresponding to the target user who has worn the smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located includes: Based on the product recommendation input information, the user profile information corresponding to the target user, and the environment-related information corresponding to the environment in which the smart glasses are located, the product recommendation prompt information corresponding to the target user is determined, wherein the product recommendation input information includes the voice information input by the target user through the smart glasses.
3. The method according to claim 2, wherein the product recommendation input information further includes target physical object information; in, The method further includes: Extract the target physical object information from the current environment image corresponding to the smart glasses.
4. The method according to claim 3, wherein extracting the target physical object information from the current environment image corresponding to the smart glasses includes: Based on the voice information, the target physical object information is extracted from the current environment screen corresponding to the smart glasses.
5. The method according to claim 3, wherein extracting the target physical object information from the current environmental image corresponding to the smart glasses includes: Based on the gesture interaction performed by the user, the target physical object information is extracted from the current environment screen corresponding to the smart glasses.
6. The method according to claim 1, wherein the structured understanding information includes text structured understanding information and image structured understanding information, wherein the text structured understanding information includes multiple attributes associated with the product description text of the product and the value of each attribute, and the image structured understanding information includes multiple attributes associated with the product description image of the product and the value of each attribute; in, The method further includes: Obtain product description text and product description images corresponding to the multiple products from at least one online store; The structured understanding information of the text is obtained from the product description text, and the structured understanding information of the image is obtained from the product description image.
7. The method according to claim 6, wherein obtaining the text structured understanding information based on the product description text includes: The product description text is segmented into multiple word units; Based on the multiple word units and the category attribute table corresponding to the product category to which the product belongs, multiple attributes associated with the product description text and the value of each attribute are obtained.
8. The method according to claim 7, wherein obtaining the image structured understanding information based on the product description image includes: Based on the product description image, obtain multiple attributes associated with the product description image and the value of each attribute.
9. The method according to claim 8, wherein obtaining multiple attributes associated with the product description image and the value of each attribute based on the product description image includes: Determine whether the product description image is a physical image of the product; If so, extract the corresponding physical image of the product from the product description image; Based on the extracted physical image and the category attribute table corresponding to the product category to which the product belongs, multiple attributes associated with the product description image and the value of each attribute are obtained.
10. The method according to claim 9, wherein the product database further includes at least one physical image corresponding to a product, such that the target model performs multimodal retrieval based on the structured understanding information corresponding to the plurality of products and the physical image corresponding to the at least one product.
11. The method according to claim 9, wherein if the product description image belongs to the physical type image of the product, the image structured understanding information further includes the scene classification corresponding to the product; in, The method further includes: Extract the background information corresponding to the product from the product description image; The scenario category corresponding to the product is determined based on the background information.
12. The method according to claim 9, wherein the image structured understanding information further includes key image description information corresponding to the product; in, The method further includes: Based on multiple attributes associated with the product description image and the value of each attribute, the key description information of the image corresponding to the product is determined.
13. The method according to claim 9, wherein the image structured understanding information further includes multiple attributes associated with at least one text in the product description image and the value of each attribute; in, The method further includes: By performing optical character recognition on the product description image, at least one text located in the product description image can be obtained; Obtain multiple attributes associated with the at least one text and the value of each attribute.
14. The method according to claim 1, further comprising: Obtain the voice feedback information of the target user regarding the at least one recommended product provided by the smart glasses; The target model re-outputs at least one latest recommended product and its corresponding attribute recommendation information based on the voice feedback information, so that the smart glasses re-provide the target user with the at least one latest recommended product and its corresponding attribute recommendation information based on their screen attributes.
15. The method according to claim 14, wherein the voice feedback information includes at least one attribute to be adjusted corresponding to the recommended product and / or attribute feedback corresponding to the at least one attribute to be adjusted.
16. The method according to claim 15, wherein the attribute recommendation information corresponding to the latest recommended product includes one or more of the at least one attribute to be adjusted and the value of the latest recommended product with respect to the one or more attributes to be adjusted.
17. An apparatus for recommending goods, comprising: The prompt information determination module is used to determine the product recommendation prompt information corresponding to the target user based on the user profile information corresponding to the target user who has worn the smart glasses and the environment-related information corresponding to the environment in which the smart glasses are located; The retrieval module is used to input the product recommendation prompt information into the target model, so that the target model retrieves at least one recommended product from the product database based on the product recommendation prompt information. The product database includes structured understanding information corresponding to multiple products, and the structured understanding information corresponding to the products includes multiple attributes associated with the products and the value of each attribute. An output module is used to obtain the at least one recommended product and the attribute recommendation information corresponding to the recommended product output by the target model, so that the smart glasses provide the at least one recommended product and the attribute recommendation information to the target user based on its screen attributes. The attribute recommendation information includes at least one attribute that matches the product recommendation prompt information among a plurality of attributes associated with the recommended product and the value of the at least one attribute.
18. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 16.
19. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method as claimed in any one of claims 1 to 16.
20. A computer program product having at least one instruction stored thereon, characterized in that, When the at least one instruction is executed by the processor, it implements the steps of the method according to any one of claims 1 to 16.