AI virtual human shopping guide system and method based on facial expression recognition
By obtaining and identifying user's facial expression information, and generating shopping guide information that is more in line with user needs, it solves the problem of inaccurate expression recognition in the existing virtual person shopping guide system, and improves user experience and sales conversion rate.
Patent Information
- Application Number
- CN202411907505.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The existing virtual shopping guide system cannot accurately identify the degree of expression of users' facial expressions, resulting in poor performance of the shopping guide information and interactive analysis, and low user satisfaction and sales conversion rate.
The face data acquisition module is used to obtain the face image sequence, and the expression information recognition module is combined with the user's real-time facial expression type and performance index, and the shopping guide information that matches the user's emotions and needs is generated through the shopping guide information generation module.
It improves the coherent analysis performance of shopping guide information, enhances users' satisfaction and acceptance of shopping guide services, and improves sales conversion rate and service quality.
Smart Images

Figure CN119379399B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of expression recognition, and in particular to an AI virtual human shopping guide system and method based on facial expression recognition. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, virtual human shopping guide systems are becoming a key application area in the commercial sector. To provide consumers with more attentive and personalized service, AI virtual human shopping guide systems based on facial expression recognition have emerged in e-commerce platforms and offline stores. Currently, these systems typically use cameras and other devices to collect facial image data from users. These systems analyze this data using algorithms and models, attempting to understand users' emotions and needs, thereby providing shopping recommendations that better meet their expectations.
[0003] In terms of technical implementation, existing facial expression recognition technologies are mostly based on traditional image processing and machine learning methods. These methods can, to a certain extent, identify some basic expression categories, such as happiness, sadness, and anger. However, existing virtual human shopping guide systems can only simply identify the type of user's expression, not the degree of expression. Furthermore, their performance in providing interactive and coherent analysis of shopping guide information is poor, resulting in low user satisfaction and acceptance of shopping guide services, as well as low sales conversion rates and service quality.
[0004] Therefore, the present invention proposes an AI virtual human shopping guide system and method based on facial expression recognition. Summary of the Invention
[0005] The present invention provides an AI virtual human shopping guide system and method based on facial expression recognition, which uses a facial data acquisition module to acquire facial image sequences, providing a data basis for subsequent analysis of user expressions, and ensuring the integrity and accuracy of the information. The facial expression information recognition module uses facial image sequences to accurately identify the user's real-time facial expressions and expression indexes, which can not only identify the user's expression type, but also the degree of expression of the user's expression type, which helps to gain a deep understanding of the user's emotions and reactions. Based on the accurate recognition of the user's real-time facial expressions and combined with the latest shopping guide information given by the AI virtual person, the shopping guide information generation module can generate the latest shopping guide information that is more in line with the user's current emotions and needs, improve the coherent analysis performance in the interactive process of shopping guide information, and improve the pertinence and effectiveness of shopping guides. It helps to improve the user's interactive experience with the AI virtual person and enhance the user's satisfaction and acceptance of shopping guide services. It can adjust the shopping guide strategy in a timely manner according to the user's real-time feedback, and improve sales conversion rate and service quality.
[0006] The present invention provides an AI virtual human shopping guide system based on facial expression recognition, comprising:
[0007] The facial data acquisition module is used to obtain facial image sequences during the interaction between the shopping guide target user and the AI virtual human;
[0008] An expression information recognition module is used to recognize the facial expressions of the target shopping guide user during the interaction with the AI virtual human based on the facial image sequence, and obtain the real-time facial expression information of the target shopping guide user, wherein the real-time facial expression information includes the real-time facial expression type of the target shopping guide user and the corresponding real-time expression index;
[0009] The shopping guide information generation module is used to generate the latest shopping guide information based on the real-time facial expression information of the shopping guide target user and the latest shopping guide information given by the AI virtual person.
[0010] Preferably, the facial data acquisition module includes:
[0011] The monitoring video acquisition submodule is used to obtain real-time monitoring video during the interaction between the user and the AI virtual human;
[0012] The facial image capture submodule is used to capture all complete facial images of the shopping guide target user in the real-time monitoring video;
[0013] The image sequence generation submodule is used to generate a facial image sequence of the shopping guide target user based on all the complete facial images of the shopping guide target user that have been intercepted.
[0014] Preferably, the facial image capture submodule includes:
[0015] The shopping guide target user extraction unit is used to extract all head images from the real-time monitoring video, determine whether there is a video frame in the real-time monitoring video containing more than one head image, and if so, determine the complete limb image sequence for each head image in the corresponding video frame; otherwise, extract the head image contained in the latest video frame in the real-time monitoring video as the head image of the shopping guide target user;
[0016] A shopping guide target user head image determination unit is configured to determine a shopping guide target user head image based on a complete limb image sequence of each head image in a video frame containing more than one head image in the real-time monitoring video;
[0017] A head image extraction unit is used to extract all head monitoring images of the shopping guide target user's head image from the real-time monitoring video based on the head morphological features of the shopping guide target user's head image;
[0018] The facial image extraction unit is used to extract all complete facial images of the shopping guide target user from all head monitoring images of the shopping guide target user's head images.
[0019] Preferably, the shopping guide target user head image determination unit includes:
[0020] Other information acquisition subunits, used to acquire touch screen interaction record sequences and online interaction record sequences in application scenarios;
[0021] a transaction dominance index determination subunit, configured to evaluate a transaction dominance index for each head image in a video frame containing more than one head image in the real-time surveillance video based on a touch screen interaction record sequence and an online interaction record sequence in the application scenario and a complete limb image sequence of each head image in a video frame containing more than one head image in the real-time surveillance video;
[0022] The shopping guide target user head image determination subunit is used to use the head image with the largest transaction dominance index as the shopping guide target user head image.
[0023] Preferably, the transaction leading index determination subunit includes:
[0024] a touch screen behavior matching degree determination terminal, configured to input a touch screen interaction record sequence in an application scenario and a complete limb image sequence of each head image in a video frame containing more than one head image in a real-time monitoring video into a touch screen behavior matching degree analysis model, and obtain a touch screen behavior matching degree of each head image in the video frame containing more than one head image in the real-time monitoring video;
[0025] An online interaction behavior matching degree determination terminal is used to input the online interaction record sequence in the application scenario and the complete limb image sequence of each head image in the video frame containing more than one head image in the real-time monitoring video into the online interaction behavior matching degree analysis model to obtain the online interaction behavior matching degree of each head image in the video frame containing more than one head image in the real-time monitoring video;
[0026] The transaction dominance index determination terminal is used to use the sum of the touch screen behavior matching degree and the online interactive behavior matching degree of each head image in the video frame containing more than one head image in the real-time monitoring video as the transaction dominance index of each head image in the video frame containing more than one head image in the real-time monitoring video.
[0027] Preferably, the expression information recognition module includes:
[0028] The expression type recognition submodule is used to analyze the real-time facial expression type of the shopping guide target user based on the facial image sequence;
[0029] The expression index recognition submodule is used to determine the real-time expression index of the real-time facial expression type of the shopping guide target user based on the current facial image sequence and the historical facial image sequence of the shopping guide target user.
[0030] Preferably, the performance index identification submodule includes:
[0031] a correction factor determination unit, configured to input a historical facial image sequence of a shopping guide target user into a facial expression macro-index relative correction factor determination model to obtain a facial expression macro-index relative correction factor of the shopping guide target user;
[0032] A macro-index determination unit is used to input the current facial image sequence into the expression expression macro-index determination model to obtain the expression expression macro-index of the shopping guide target user;
[0033] The expression index determination unit is used to use the sum of the expression expression macro index and the expression expression macro index relative correction factor of the shopping guide target user as the real-time expression index of the real-time facial expression type of the shopping guide target user.
[0034] Preferably, the shopping guide information generating module includes:
[0035] The feedback interpretation submodule is used to evaluate the user's feedback type and corresponding feedback performance index on the latest shopping guide information given by the AI virtual person based on the user's real-time facial expression information;
[0036] The shopping guide information generation submodule is used to generate the latest shopping guide information based on the user's feedback type and corresponding feedback performance index on the shopping guide information latest given by the AI virtual person and the shopping guide information latest given by the AI virtual person.
[0037] Preferably, the shopping guide information generation submodule includes:
[0038] A first shopping guide information generating unit is configured to generate the latest shopping guide information based on the latest touch screen feedback information and / or voice feedback information provided by the user, the user's feedback type and corresponding feedback performance index on the latest shopping guide information provided by the AI avatar, and the latest shopping guide information provided by the AI avatar when the user provides the latest touch screen feedback information and / or voice feedback information;
[0039] The second shopping guide information generation unit is used to generate the latest shopping guide information based on the user's feedback type and corresponding feedback performance index on the latest shopping guide information given by the AI virtual person and the latest shopping guide information given by the AI virtual person when the user has not given the latest touch screen feedback information and / or voice feedback information.
[0040] The present invention provides an AI virtual human shopping guide method based on facial expression recognition, which is applied to any of the above AI virtual human shopping guide systems based on facial expression recognition, comprising:
[0041] S1: Acquire facial image sequences during the interaction between the shopping guide target user and the AI virtual human;
[0042] S2: Recognize facial expressions of the target shopping guide user during the interaction with the AI virtual human based on the facial image sequence to obtain real-time facial expression information of the target shopping guide user, where the real-time facial expression information includes the real-time facial expression type of the target shopping guide user and the corresponding real-time expression index;
[0043] S3: Generate the latest shopping guide information based on the real-time facial expression information of the shopping guide target user and the latest shopping guide information given by the AI virtual person.
[0044] The beneficial effects of the present invention compared to the prior art are as follows: the facial data acquisition module can acquire facial image sequences, providing a data basis for subsequent analysis of user expressions, ensuring the integrity and accuracy of the information. The facial expression information recognition module processes the facial image sequence to accurately identify the user's real-time facial expression and expression index, which can not only identify the user's expression type, but also the degree of expression of the user's expression type, which helps to gain a deeper understanding of the user's emotions and reactions. Based on the accurate recognition of the user's real-time facial expression and combined with the latest shopping guide information given by the AI virtual person, the shopping guide information generation module can generate the latest shopping guide information that is more in line with the user's current emotions and needs, improve the coherent analysis performance in the interactive process of shopping guide information, and improve the pertinence and effectiveness of shopping guides. It helps to improve the user's interactive experience with the AI virtual person and enhance the user's satisfaction and acceptance of the shopping guide service. It can adjust the shopping guide strategy in a timely manner according to the user's real-time feedback, and improve the sales conversion rate and service quality.
[0045] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in this application document.
[0046] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0048] Figure 1 Schematic diagram of an AI virtual human shopping guide system based on facial expression recognition in an embodiment of the present invention;
[0049] Figure 2 This is a flow chart of the AI virtual human shopping guide method based on facial expression recognition in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0051] Example 1
[0052] The present invention provides an AI virtual human shopping guide system based on facial expression recognition, referring to Figure 1 ,include:
[0053] The facial data acquisition module is used to obtain facial image sequences during the interaction between the shopping guide target user and the AI virtual human;
[0054] An expression information recognition module is used to recognize the facial expressions of the target shopping guide user during the interaction with the AI virtual human based on the facial image sequence, and obtain the real-time facial expression information of the target shopping guide user, wherein the real-time facial expression information includes the real-time facial expression type of the target shopping guide user and the corresponding real-time expression index;
[0055] The shopping guide information generation module is used to generate the latest shopping guide information based on the real-time facial expression information of the shopping guide target user and the latest shopping guide information given by the AI virtual person.
[0056] In this embodiment, the target users of the shopping guide refer to customers who are interacting with the AI virtual person and have the intention to purchase products. They are also the main target users of this shopping guide promotion.
[0057] In this embodiment, the facial image sequence is a series of facial images of the shopping guide target user taken continuously.
[0058] In this embodiment, the real-time facial expression type of the shopping guide target user refers to the specific expression category presented by the shopping guide target user when interacting with the AI virtual person, such as happiness, surprise, confusion, etc.
[0059] In this embodiment, the real-time expression index of a real-time facial expression type is used to quantify the degree of obviousness or intensity of a certain real-time facial expression type. For example, if the customer's expression of surprise is very strong, the corresponding real-time expression index of surprise may be higher.
[0060] In this embodiment, the latest shopping guide information provided by the AI virtual person refers to the information about product introductions, recommendations, promotional activities, etc. that the AI virtual person has just provided to the shopping guide target user.
[0061] For example: The AI virtual person said, "This drink is on sale, it's a good deal."
[0062] In this embodiment, the latest shopping guide information is regenerated based on the real-time facial expression information of the shopping guide target user and the shopping guide information previously given by the AI virtual person, and is shopping guide content that is more in line with the user's current status and needs.
[0063] For example: When a user shows a puzzled expression, the latest shopping guide information may be a more detailed explanation of the previously recommended product.
[0064] The beneficial effects of the above technology are as follows: the facial data acquisition module can acquire facial image sequences, providing a data basis for subsequent analysis of user expressions, ensuring the integrity and accuracy of the information. The facial expression information recognition module accurately identifies the user's real-time facial expressions and expression index by processing the facial image sequence. It can not only identify the user's expression type, but also the degree of expression of the user's expression type, which helps to gain a deeper understanding of the user's emotions and reactions. Based on the accurate recognition of the user's real-time facial expressions and combined with the latest shopping guide information given by the AI virtual person, the shopping guide information generation module can generate the latest shopping guide information that is more in line with the user's current emotions and needs, improve the coherent analysis performance in the interactive process of shopping guide information, and improve the pertinence and effectiveness of shopping guides. It helps to improve the user's interactive experience with the AI virtual person and enhance the user's satisfaction and acceptance of the shopping guide service. It can adjust the shopping guide strategy in a timely manner according to the user's real-time feedback, and improve sales conversion rate and service quality.
[0065] Example 2
[0066] Based on Example 1, the facial data acquisition module includes:
[0067] The monitoring video acquisition submodule is used to obtain real-time monitoring video during the interaction between the user and the AI virtual human;
[0068] The facial image capture submodule is used to capture all complete facial images of the shopping guide target user in the real-time monitoring video;
[0069] The image sequence generation submodule is used to generate a facial image sequence of the shopping guide target user based on all the complete facial images of the shopping guide target user that have been intercepted.
[0070] In this embodiment, in the context of a vending machine, a complete facial image refers to a facial image that can clearly and comprehensively display the entire facial features of the target shopping guide user without being blocked or missing any important parts.
[0071] For example, the customer's face captured by the camera in front of the vending machine includes the eyes, nose, and mouth, and no key parts are blocked by items such as hats and masks.
[0072] The beneficial effects of the above technologies are as follows: The surveillance video acquisition submodule captures real-time surveillance video, comprehensively recording the interaction process and ensuring the richness and integrity of the data source. The facial image capture submodule captures complete facial images from real-time surveillance video, improving the accuracy and pertinence of the data and eliminating irrelevant interference. The image sequence generation submodule generates sequences based on the captured facial images, providing organized and coherent data support for subsequent expression recognition. This helps improve the efficiency and quality of facial data acquisition and lays a solid foundation for accurate analysis of user expressions. It can provide more reliable and valuable user information for subsequent shopping guide services, improving shopping guidance effectiveness and user experience.
[0073] Example 3
[0074] Based on Example 2, the facial image capture submodule includes:
[0075] The shopping guide target user extraction unit is used to extract all head images from the real-time monitoring video, determine whether there is a video frame in the real-time monitoring video containing more than one head image, and if so, determine the complete limb image sequence for each head image in the corresponding video frame; otherwise, extract the head image contained in the latest video frame in the real-time monitoring video as the head image of the shopping guide target user;
[0076] A shopping guide target user head image determination unit is configured to determine a shopping guide target user head image based on a complete limb image sequence of each head image in a video frame containing more than one head image in the real-time monitoring video;
[0077] A head image extraction unit is used to extract all head monitoring images of the shopping guide target user's head image from the real-time monitoring video based on the head morphological features of the shopping guide target user's head image;
[0078] The facial image extraction unit is used to extract all complete facial images of the shopping guide target user from all head monitoring images of the shopping guide target user's head images.
[0079] In this embodiment, the head image in the vending machine scenario refers to an image that only contains the user's head in the real-time surveillance video.
[0080] In this embodiment, the complete limb image sequence of the head image is a series of continuous image sets including the head and other limb parts (such as shoulders, limbs, torso, etc.).
[0081] In this embodiment, the head image of the target shopping guide user is a head image of the target shopping guide user. For example, in a group of people gathered in front of a vending machine, the head image of a user who is communicating with the AI virtual person and preparing to make a purchase.
[0082] In this embodiment, based on the head morphological features of the shopping guide target user's head image, all head monitoring images of the shopping guide target user's head image are extracted from the real-time monitoring video: according to the shape, size, outline and other features of the shopping guide target user's head that have been determined, all monitoring images containing the user's head are found from the real-time monitoring video.
[0083] For example: By identifying the unique head features of the shopping guide target user, such as hairstyle, head size, etc., all head images with this feature are filtered out from the video.
[0084] In this embodiment, all complete facial images of the target shopping guide user are extracted from all head surveillance images of the target shopping guide user's head. From the already screened surveillance images containing the target shopping guide user's head, further images are found that can clearly show the user's entire face (including eyes, nose, mouth, etc.) and are not blocked.
[0085] For example: Among a series of images containing the user's head, select those images in which the user's facial features can be fully seen and are not blocked by hands or other objects.
[0086] The beneficial effects of the above technologies are as follows: the shopping guide target user extraction unit is able to handle complex situations involving multiple head images, accurately judge and extract possible user head images, and improve the accuracy of target user identification. The shopping guide target user head image determination unit determines the target user head image based on a complete limb image sequence, which increases the reliability and accuracy of the judgment. The head image extraction unit extracts the head monitoring image based on the head morphological features, which improves the accuracy and pertinence of data extraction. The facial image extraction unit further accurately extracts the complete facial image from the head monitoring image, providing high-quality data for subsequent expression analysis. The entire facial image capture submodule process helps to improve the accuracy and completeness of facial image capture, and provides strong support for accurately analyzing user expressions and generating effective shopping guide information.
[0087] Example 4
[0088] Based on the third embodiment, the shopping guide target user head image determination unit includes:
[0089] Other information acquisition subunits, used to acquire touch screen interaction record sequences and online interaction record sequences in application scenarios;
[0090] a transaction dominance index determination subunit, configured to evaluate a transaction dominance index for each head image in a video frame containing more than one head image in the real-time surveillance video based on a touch screen interaction record sequence and an online interaction record sequence in the application scenario and a complete limb image sequence of each head image in a video frame containing more than one head image in the real-time surveillance video;
[0091] The shopping guide target user head image determination subunit is used to use the head image with the largest transaction dominance index as the shopping guide target user head image.
[0092] In this embodiment, the application scenario: in the context of a vending machine, refers to the specific environment and conditions in which a user interacts with the vending machine and the AI virtual person therein and generates a transaction.
[0093] In this embodiment, the touch screen interaction record sequence is a record of a series of operations performed by the user on the touch screen of the vending machine.
[0094] For example: including records such as the order in which users click on products and the number of times they view product details.
[0095] In this embodiment, the online interaction record sequence refers to a series of records of the user communicating and operating with the vending machine-related system through the network.
[0096] For example, consider a user browsing a vending machine's products on a mobile app.
[0097] In this embodiment, the transaction dominance index of the head image is used to measure the degree of dominance or participation of the user corresponding to each head image in the transaction process when there are multiple head images.
[0098] For example, if the user corresponding to a head image has more purchase intention operations on the touch screen and online, the transaction dominance index will be higher.
[0099] The beneficial effects of the above technology are as follows: the other information acquisition subunit acquires touchscreen interaction record sequences and online interaction record sequences, enriching the information sources used for judgment and improving the comprehensiveness of the analysis. The transaction dominance index determination subunit evaluates the transaction dominance index by integrating multiple records and body image sequences, making the judgment more scientific and objective. It can accurately determine the head image of the shopping guide target user with the highest transaction dominance index among multiple head images, improving the accuracy of target user identification. This helps to more effectively focus on key users and provide accurate targets for subsequent expression analysis and shopping guide information generation. It also enhances the overall system's ability to accurately identify target users in complex interaction scenarios, thereby improving the relevance and effectiveness of shopping guide services.
[0100] Example 5
[0101] Based on Example 4, the transaction leading index determination sub-unit includes:
[0102] a touch screen behavior matching degree determination terminal, configured to input a touch screen interaction record sequence in an application scenario and a complete limb image sequence of each head image in a video frame containing more than one head image in a real-time monitoring video into a touch screen behavior matching degree analysis model, and obtain a touch screen behavior matching degree of each head image in the video frame containing more than one head image in the real-time monitoring video;
[0103] An online interaction behavior matching degree determination terminal is used to input the online interaction record sequence in the application scenario and the complete limb image sequence of each head image in the video frame containing more than one head image in the real-time monitoring video into the online interaction behavior matching degree analysis model to obtain the online interaction behavior matching degree of each head image in the video frame containing more than one head image in the real-time monitoring video;
[0104] The transaction dominance index determination terminal is used to use the sum of the touch screen behavior matching degree and the online interactive behavior matching degree of each head image in the video frame containing more than one head image in the real-time monitoring video as the transaction dominance index of each head image in the video frame containing more than one head image in the real-time monitoring video.
[0105] In this embodiment, the touch screen behavior matching analysis model is a model used to analyze the matching degree between the body movements of a single head image in a real-time monitoring video and a touch screen interaction record sequence.
[0106] For example, determining the degree of compatibility between the hand movements of a user corresponding to a head image and operations such as clicking and sliding on a touch screen.
[0107] The training samples of the touch screen behavior matching analysis model may include:
[0108] A large number of video clips of users operating vending machine touch screens, including user head images and corresponding touch screen operation records;
[0109] These samples are labeled with different degrees of touch screen behavior matching, such as complete match, partial match, no match, etc.
[0110] For example, for a touch screen behavior matching analysis model, a training sample may be a user continuously clicking on several items on the touch screen of a vending machine, while the user's head image is clearly visible in the surveillance video and is labeled as a "highly matching" touch screen behavior.
[0111] In this embodiment, the touch screen behavior matching degree of the head image refers to a quantitative value of the degree of consistency between the body movements of the user represented by the specific head image and the touch screen interaction behavior.
[0112] For example, if a user looks at the screen and quickly clicks on multiple products with their hands, the touch screen behavior match of their head image is higher.
[0113] In this embodiment, the online interaction behavior matching analysis model is a model for evaluating the matching degree between the body movements of the head image in the real-time monitoring video and the online interaction record sequence.
[0114] Example: Analyze the correlation between user operations on a vending machine's online platform and user hand movements.
[0115] The training samples for the online interactive behavior matching analysis model may include:
[0116] Records of users' online interactions with vending machine-related systems (such as browsing products and checking promotions through mobile apps), combined with images of the user's head from real-time surveillance video;
[0117] These samples are also labeled with different levels of online interaction behavior matching.
[0118] For example, for an online interactive behavior matching analysis model, the training sample may be a user viewing information about a certain type of product on a mobile app for a long time, and the corresponding head image in the surveillance video is marked as a "moderate match" online interactive behavior.
[0119] In this embodiment, the online interactive behavior matching degree of the head image represents a numerical value of the matching degree between the action of the user corresponding to the specific head image and the online interactive behavior.
[0120] For example, if the user corresponding to a head image has a high number of online searches for specific products while looking at the mobile phone screen, the matching degree of their online interactive behavior is higher.
[0121] The beneficial effects of the above technology are as follows: the touch screen behavior matching determination end can accurately measure the degree of match between each head image and touch screen interaction behavior by inputting data into the touch screen behavior matching analysis model. The online interaction behavior matching determination end uses a similar method to determine the online interaction behavior matching degree, comprehensively covering different types of interaction behaviors. The transaction dominance index determination end comprehensively considers the matching degree of touch screen and online interaction behaviors to determine the transaction dominance index, making the index more comprehensive and accurate. This helps to more accurately assess the dominance of each head image in the transaction, thereby more accurately identifying the target users of shopping guides. It provides a critical and reliable basis for the subsequent generation of shopping guide information that better meets the needs of target users, thereby improving the quality and effectiveness of shopping guide services.
[0122] Example 6
[0123] Based on Example 1, the expression information recognition module includes:
[0124] The expression type recognition submodule is used to analyze the real-time facial expression type of the shopping guide target user based on the facial image sequence;
[0125] The expression index recognition submodule is used to determine the real-time expression index of the real-time facial expression type of the shopping guide target user based on the current facial image sequence and the historical facial image sequence of the shopping guide target user.
[0126] In this embodiment, the real-time facial expression type of the shopping guide target user is analyzed based on the facial image sequence: this means that by analyzing and judging a series of continuously taken facial images of the shopping guide target user, the specific expression category currently presented by the user, such as happiness, sadness, surprise, etc., is determined.
[0127] For example: From continuous facial images, it is found that the user's mouth corners are raised and the eyes are narrowed, so the real-time facial expression type is judged to be happy.
[0128] In this embodiment, the historical facial image sequence of the shopping guide target user refers to a series of facial images of the shopping guide target user during the interaction with the vending machine and the AI virtual human, which have been acquired and stored before the current interaction.
[0129] For example, a collection of facial images taken when a user shopped at this vending machine multiple times before.
[0130] The beneficial effects of the above technology are as follows: The expression type recognition submodule can accurately identify the real-time facial expression type of the target shopping guide user by analyzing facial image sequences, providing a basis for understanding the user's emotional state. The expression index recognition submodule combines current and historical facial image sequences to determine the real-time expression index, providing a more comprehensive and accurate assessment of the user's expression intensity. This helps to more deeply understand the degree of user emotional response during interaction, providing a key basis for providing personalized shopping guidance services. It can more promptly capture the changing trends of user emotions, providing real-time guidance for adjusting shopping guidance strategies. This improves the accuracy and reliability of expression information recognition, thereby enhancing the targeted and adaptable shopping guidance services.
[0131] Example 7
[0132] Based on Example 6, the performance index identification submodule includes:
[0133] a correction factor determination unit, configured to input a historical facial image sequence of a shopping guide target user into a facial expression macro-index relative correction factor determination model to obtain a facial expression macro-index relative correction factor of the shopping guide target user;
[0134] A macro-index determination unit is used to input the current facial image sequence into the expression expression macro-index determination model to obtain the expression expression macro-index of the shopping guide target user;
[0135] The expression index determination unit is used to use the sum of the expression expression macro index and the expression expression macro index relative correction factor of the shopping guide target user as the real-time expression index of the real-time facial expression type of the shopping guide target user.
[0136] In this embodiment, the facial expression macro-index relative correction factor determination model is an algorithm or model for calculating the relevant factor for correcting the facial expression macro-index based on a historical facial image sequence of a shopping guide target user. The training samples thereof include:
[0137] Multiple sets of historical facial image sequences of the same user at different times and scenes;
[0138] Each set of historical facial image sequences needs to be labeled with a corresponding correction factor, and the labeled value may be determined based on factors such as the user's facial expression habits, general reactions in specific scenarios, etc.
[0139] For example, a training sample might be a set of consecutive facial images of a user viewing different product recommendations at a vending machine, annotated with the corresponding macro-index or correction factor for their facial expression. This set of samples is then trained on a model, enabling it to accurately calculate the corresponding index or correction factor based on the input facial image sequence.
[0140] For example, by analyzing the changing patterns of a user's facial expressions during their past shopping trips, a computational model for adjusting the current macro-index of facial expression can be derived.
[0141] In this embodiment, the relative correction factor for the macro-index of facial expression is a value calculated based on the user's historical facial expression data and used to correct and adjust the current macro-index of facial expression. For example, if the user has typically expressed excitement when purchasing certain products in the past (i.e., compared to the average person), this correction factor may lower the current macro-index of a similar expression.
[0142] In this embodiment, the macro-index determination model for facial expression is an algorithm or model used to calculate the macro-level of a user's facial expression (i.e., the macro-level of a user's facial expression compared to the expressions of most ordinary people) based on a current facial image sequence. Specifically, during training, the model learns the different types and levels of facial expressions of a large number of people in front of vending machines to achieve a relative assessment of the level of facial expression of a single user in an input facial image sequence. The training samples include:
[0143] A large number of facial image sequences containing different users in different situations (such as different vending machine scenes, different time of day, etc.);
[0144] These facial image sequences should be annotated with corresponding macro-indices of facial expression, for example, through evaluation by professionals or based on certain known standards.
[0145] In this embodiment, the macro-index of facial expression is a comprehensive quantitative representation of the user's current facial expression, reflecting the intensity or obviousness of the expression. A larger value may indicate a stronger expression, such as extreme joy or extreme surprise; a smaller value may indicate a milder expression.
[0146] The beneficial effects of the above technology are as follows: the correction factor determination unit obtains a relative correction factor by inputting a historical facial image sequence, which can take into account the user's long-term facial expression habits and make the evaluation more personalized. The macro-index determination unit uses the current facial image sequence to determine the macro-index of facial expression, focusing on the current facial expression state. The performance index determination unit combines the macro-index and the correction factor to determine the real-time performance index, making the evaluation result more comprehensive, accurate and in line with the user's actual situation. It helps to more accurately quantify the intensity of the user's real-time facial expression, and provide strong support for the subsequent generation of accurate shopping guide information. It can improve the scientificity and reliability of expression index recognition, thereby optimizing the quality and effectiveness of shopping guide services and better meeting user needs.
[0147] Example 8
[0148] Based on Example 1, the shopping guide information generation module includes:
[0149] The feedback interpretation submodule is used to evaluate the user's feedback type and corresponding feedback performance index on the latest shopping guide information given by the AI virtual person based on the user's real-time facial expression information;
[0150] The shopping guide information generation submodule is used to generate the latest shopping guide information based on the user's feedback type and corresponding feedback performance index on the shopping guide information latest given by the AI virtual person and the shopping guide information latest given by the AI virtual person.
[0151] In this embodiment, the type of user feedback on the latest shopping guide information provided by the AI virtual person refers to the type of user response to the shopping guide information provided by the AI virtual person.
[0152] For example: The user's facial expression changes, verbal responses, operational actions, etc. reflect whether the user is interested in, confused or dissatisfied with the shopping guide information.
[0153] In this embodiment, the feedback performance index is used to quantify the degree of phenotypic expression of the user's feedback on the latest shopping guide information provided by the AI avatar. If the user's expression is very pleasantly surprised, the feedback performance index may be high; if the user only frowns slightly, the index may be low.
[0154] The beneficial effects of the above technology are as follows: the feedback interpretation submodule can evaluate the feedback type and feedback performance index from the user's real-time facial expression information, providing a quantitative basis for accurately understanding the user's response to shopping guide information. This quantitative assessment helps to more accurately grasp the user's attitude and emotional tendencies, and improve the accuracy of interpreting user feedback. The shopping guide information generation submodule generates new shopping guide information based on the accurate feedback evaluation and the original shopping guide information, making the new shopping guide more targeted and adaptable. It can better meet user needs and expectations, and improve user satisfaction and acceptance of the shopping guide service. It helps to optimize the shopping guide process, improve sales efficiency and service quality, and enhance the competitiveness and practicality of AI virtual human shopping guides.
[0155] Example 9
[0156] Based on Example 8, the shopping guide information generation submodule includes:
[0157] A first shopping guide information generating unit is configured to generate the latest shopping guide information based on the latest touch screen feedback information and / or voice feedback information provided by the user, the user's feedback type and corresponding feedback performance index on the latest shopping guide information provided by the AI avatar, and the latest shopping guide information provided by the AI avatar when the user provides the latest touch screen feedback information and / or voice feedback information;
[0158] The second shopping guide information generation unit is used to generate the latest shopping guide information based on the user's feedback type and corresponding feedback performance index on the latest shopping guide information given by the AI virtual person and the latest shopping guide information given by the AI virtual person when the user has not given the latest touch screen feedback information and / or voice feedback information.
[0159] In this embodiment, the touch screen feedback information refers to feedback data generated by the user performing operations by touching the screen, such as clicking on a product, sliding the screen, selecting an option, etc.
[0160] For example, a user clicks on a beverage icon on the touch screen of a vending machine, indicating interest in the beverage.
[0161] In this embodiment, the voice feedback information refers to the feedback content expressed by the user to the vending machine or the AI virtual person through language.
[0162] Example: User says "I want chocolate flavor."
[0163] In this embodiment, the latest shopping guide information is generated based on the latest touch screen feedback information and / or voice feedback information given by the user, the user's feedback type and corresponding feedback performance index on the latest shopping guide information given by the AI virtual person, and the latest shopping guide information given by the AI virtual person: this means that the latest feedback expressed by the user through the touch screen or voice, the feedback on the previous shopping guide information (such as type and performance intensity) and the shopping guide content given by the AI virtual person before are comprehensively considered to regenerate new shopping guide information that is more in line with user needs and reactions.
[0164] For example, if a user selects a drink through the touch screen and says "this is too expensive", and the feedback performance index shows a high level of dissatisfaction, new shopping guide information will be generated, such as recommending a similar drink with a more suitable price.
[0165] In this embodiment, the latest shopping guide information is generated based on the user's feedback type and corresponding feedback performance index on the latest shopping guide information given by the AI virtual person, as well as the latest shopping guide information given by the AI virtual person: when the user has no new touch screen or voice feedback, the new shopping guide information is regenerated only based on the feedback on multiple aspects of the previous shopping guide information (such as dimensions such as expression changes) and its intensity, as well as the previous shopping guide content.
[0166] For example: If the user does not speak or touch the screen, but his expression shows confusion and the feedback performance index is average, then new shopping guide information may be generated to further explain the characteristics of the previously recommended product.
[0167] The beneficial effects of the above technology include the ability to generate the latest shopping guide information in different ways depending on whether the user has the latest touchscreen and voice feedback information, thereby improving the flexibility and adaptability of the generation strategy. When the user has touchscreen and voice feedback, the shopping guide information is generated by integrating various feedback information, making the shopping guide more closely aligned with the user's explicit needs and expressions. When the user does not have such explicit feedback, the shopping guide information is generated based solely on facial expression feedback, fully utilizing all available feedback channels and ensuring the continuity and effectiveness of the shopping guide service. This helps to more comprehensively and accurately meet user needs, improving the user experience and the success rate of shopping guides. This refined processing method can optimize resource utilization and improve the efficiency and quality of shopping guide information generation.
[0168] Example 10:
[0169] The present invention provides an AI virtual human shopping guide method based on facial expression recognition, which is applied to the AI virtual human shopping guide system based on facial expression recognition described in any one of embodiments 1 to 9, with reference to Figure 2 ,include:
[0170] S1: Acquire facial image sequences during the interaction between the shopping guide target user and the AI virtual human;
[0171] S2: Recognize facial expressions of the target shopping guide user during the interaction with the AI virtual human based on the facial image sequence to obtain real-time facial expression information of the target shopping guide user, where the real-time facial expression information includes the real-time facial expression type of the target shopping guide user and the corresponding real-time expression index;
[0172] S3: Generate the latest shopping guide information based on the real-time facial expression information of the shopping guide target user and the latest shopping guide information given by the AI virtual person.
[0173] The beneficial effects of the above technology are as follows: the facial data acquisition module can acquire facial image sequences, providing a data basis for subsequent analysis of user expressions, ensuring the integrity and accuracy of the information. The facial expression information recognition module accurately identifies the user's real-time facial expressions and expression index by processing the facial image sequence. It can not only identify the user's expression type, but also the degree of expression of the user's expression type, which helps to gain a deeper understanding of the user's emotions and reactions. Based on the accurate recognition of the user's real-time facial expressions and combined with the latest shopping guide information given by the AI virtual person, the shopping guide information generation module can generate the latest shopping guide information that is more in line with the user's current emotions and needs, improve the coherent analysis performance in the interactive process of shopping guide information, and improve the pertinence and effectiveness of shopping guides. It helps to improve the user's interactive experience with the AI virtual person and enhance the user's satisfaction and acceptance of the shopping guide service. It can adjust the shopping guide strategy in a timely manner according to the user's real-time feedback, and improve sales conversion rate and service quality.
[0174] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. AI virtual human shopping guide system based on facial expression recognition, characterized by: include: The facial data acquisition module is used to obtain facial image sequences during the interaction between the shopping guide target user and the AI virtual human, including: The monitoring video acquisition submodule is used to obtain real-time monitoring video during the interaction between the user and the AI virtual human; The facial image capture submodule is used to capture all complete facial images of the shopping guide target user in the real-time monitoring video, including: The shopping guide target user extraction unit is used to extract all head images from the real-time monitoring video, determine whether there is a video frame in the real-time monitoring video containing more than one head image, and if so, determine the complete limb image sequence for each head image in the corresponding video frame; otherwise, extract the head image contained in the latest video frame in the real-time monitoring video as the head image of the shopping guide target user; A shopping guide target user head image determination unit is configured to determine a shopping guide target user head image based on a complete limb image sequence of each head image in a video frame containing more than one head image in a real-time monitoring video, comprising: Other information acquisition subunits, used to acquire touch screen interaction record sequences and online interaction record sequences in application scenarios; The transaction dominance index determination subunit is configured to evaluate the transaction dominance index of each head image in a video frame containing more than one head image in the real-time surveillance video based on a touch screen interaction record sequence and an online interaction record sequence in the application scenario and a complete limb image sequence of each head image in a video frame containing more than one head image in the real-time surveillance video, including: a touch screen behavior matching degree determination terminal, configured to input a touch screen interaction record sequence in an application scenario and a complete limb image sequence of each head image in a video frame containing more than one head image in a real-time monitoring video into a touch screen behavior matching degree analysis model, and obtain a touch screen behavior matching degree of each head image in the video frame containing more than one head image in the real-time monitoring video; An online interaction behavior matching degree determination terminal is used to input the online interaction record sequence in the application scenario and the complete limb image sequence of each head image in the video frame containing more than one head image in the real-time monitoring video into the online interaction behavior matching degree analysis model to obtain the online interaction behavior matching degree of each head image in the video frame containing more than one head image in the real-time monitoring video; a transaction dominance index determining terminal, configured to use the sum of the touch screen behavior matching degree and the online interaction behavior matching degree of each head image in a video frame containing more than one head image in the real-time monitoring video as the transaction dominance index of each head image in the video frame containing more than one head image in the real-time monitoring video; A shopping guide target user head image determination subunit, configured to use the head image with the largest transaction dominance index as the shopping guide target user head image; A head image extraction unit is used to extract all head monitoring images of the shopping guide target user's head image from the real-time monitoring video based on the head morphological features of the shopping guide target user's head image; A facial image extraction unit is used to extract all complete facial images of the shopping guide target user from all head monitoring images of the shopping guide target user's head image; An image sequence generation submodule, configured to generate a facial image sequence of a shopping guide target user based on all captured complete facial images of the shopping guide target user; An expression information recognition module is used to recognize the facial expressions of the target shopping guide user during the interaction with the AI virtual human based on the facial image sequence, and obtain the real-time facial expression information of the target shopping guide user, wherein the real-time facial expression information includes the real-time facial expression type of the target shopping guide user and the corresponding real-time expression index; The shopping guide information generation module is used to generate the latest shopping guide information based on the real-time facial expression information of the shopping guide target user and the latest shopping guide information given by the AI virtual person.
2. The AI virtual human shopping guide system based on facial expression recognition according to claim 1 is characterized in that: Expression information recognition module, including: The expression type recognition submodule is used to analyze the real-time facial expression type of the shopping guide target user based on the facial image sequence; The expression index recognition submodule is used to determine the real-time expression index of the real-time facial expression type of the shopping guide target user based on the current facial image sequence and the historical facial image sequence of the shopping guide target user.
3. The AI virtual human shopping guide system based on facial expression recognition according to claim 2 is characterized in that: Performance index identification submodule, including: a correction factor determination unit, configured to input a historical facial image sequence of a shopping guide target user into a facial expression macro-index relative correction factor determination model to obtain a facial expression macro-index relative correction factor of the shopping guide target user; A macro-index determination unit is used to input the current facial image sequence into the expression expression macro-index determination model to obtain the expression expression macro-index of the shopping guide target user; The expression index determination unit is used to use the sum of the expression expression macro index and the expression expression macro index relative correction factor of the shopping guide target user as the real-time expression index of the real-time facial expression type of the shopping guide target user.
4. The AI virtual human shopping guide system based on facial expression recognition according to claim 1, characterized in that: Shopping guide information generation module, including: The feedback interpretation submodule is used to evaluate the user's feedback type and corresponding feedback performance index on the latest shopping guide information given by the AI virtual person based on the user's real-time facial expression information; The shopping guide information generation submodule is used to generate the latest shopping guide information based on the user's feedback type and corresponding feedback performance index on the shopping guide information latest given by the AI virtual person and the shopping guide information latest given by the AI virtual person.
5. The AI virtual human shopping guide system based on facial expression recognition according to claim 4 is characterized in that: The shopping guide information generation submodule includes: A first shopping guide information generating unit is configured to generate the latest shopping guide information based on the latest touch screen feedback information and / or voice feedback information provided by the user, the user's feedback type and corresponding feedback performance index on the latest shopping guide information provided by the AI avatar, and the latest shopping guide information provided by the AI avatar when the user provides the latest touch screen feedback information and / or voice feedback information; The second shopping guide information generation unit is used to generate the latest shopping guide information based on the user's feedback type and corresponding feedback performance index on the latest shopping guide information given by the AI virtual person and the latest shopping guide information given by the AI virtual person when the user has not given the latest touch screen feedback information and / or voice feedback information.
6. The AI virtual human shopping guide method based on facial expression recognition is characterized by: The AI virtual human shopping guide system based on facial expression recognition applied to any one of claims 1 to 5 comprises: S1: Acquire facial image sequences during the interaction between the shopping guide target user and the AI virtual human; S2: Recognize facial expressions of the target shopping guide user during the interaction with the AI virtual human based on the facial image sequence to obtain real-time facial expression information of the target shopping guide user, where the real-time facial expression information includes the real-time facial expression type of the target shopping guide user and the corresponding real-time expression index; S3: Generate the latest shopping guide information based on the real-time facial expression information of the shopping guide target user and the latest shopping guide information given by the AI virtual person.
Citation Information
Patent Citations
Multidimensional visual shopping guide system and method
CN108573403A
Commodity promotion method, device and system and computer storage medium
CN108876427A
Method and device for determining target object
CN113076346A
Commodity recommendation method and device
CN116843428A