Active guidance intelligent retail method, apparatus, device, and medium

By analyzing user data from multiple dimensions, the system generates purchase intention characteristics and provides proactive guidance, solving the problem that existing smart retail systems cannot identify user intent and improving user experience and sales conversion rates.

CN122492214APending Publication Date: 2026-07-31MITA VISION (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MITA VISION (BEIJING) TECHNOLOGY CO LTD
Filing Date
2026-04-03
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing smart retail systems are unable to accurately identify user intent, resulting in lost sales opportunities due to user hesitation and a poor user experience.

Method used

By collecting users' facial expression data, gaze direction data, body movement data, and voice interaction data through pre-deployed sensors, multi-dimensional analysis is performed to generate users' purchase intention characteristics, and proactive guidance messages are provided when user decision paralysis is detected.

Benefits of technology

It enabled precise guidance for users, improved user experience, effectively seized sales opportunities that might have been lost due to hesitation, and increased conversion rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492214A_ABST
    Figure CN122492214A_ABST
Patent Text Reader

Abstract

This disclosure provides a proactively guided intelligent retail method, device, equipment, and medium, comprising: analyzing facial expression data of a target user to obtain a first visual feature; analyzing the target user's gaze direction data to obtain a second visual feature based on gaze placement and wandering frequency; analyzing the target user's body movement data to obtain a third visual feature; conducting proactive intervention assessment on the target user based on the first, second, and third visual features and purchase intention characteristics to obtain a proactive intervention score; if the proactive intervention score exceeds a preset threshold, determining candidate products based on the target user's gaze placement information; generating proactive guidance scripts based on the target user's profile information and current context information; and displaying the proactive guidance scripts through a retail terminal. This provides users with precise guidance to facilitate transactions and effectively enhances the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of new retail technology, and more specifically, to a proactive guided smart retail method, apparatus, device, and medium. Background Technology

[0002] Currently, most intelligent interactive systems applied in retail scenarios, whether online chatbots or offline physical intelligent terminals (such as smart vending machines and digital signage), have a core interaction mode that is passively triggered.

[0003] In related technologies, the system mainly responds to the user's explicit instructions through a set of preset rules or knowledge base, using keyword matching or simple natural language understanding; the user must actively initiate a query before the system can provide an answer based on the knowledge base.

[0004] However, existing technologies cannot accurately identify user intent, leading to lost sales opportunities due to user hesitation and a poor user experience. Summary of the Invention

[0005] The embodiments described herein provide a proactively guided smart retail method, apparatus, device, and medium that overcomes the aforementioned problems.

[0006] Firstly, based on the content of this disclosure, a proactive guided smart retail method is provided, including: Data on facial expressions, gaze direction, body movements, and voice interaction of target users relative to retail terminals are collected through pre-deployed sensors. Facial expression data of the target user relative to the retail terminal is analyzed to obtain a first visual feature; gaze direction data of the target user relative to the retail terminal is analyzed to determine the target user's focus and degree of hesitation, thus obtaining a second visual feature; body movement data of the target user relative to the retail terminal is analyzed to obtain a third visual feature. The system performs speech recognition and semantic analysis on the voice interaction data of the target user relative to the retail terminal, extracting the target user's voice content, intonation changes, and speech rate features; identifies the target user's product inquiry intent based on the voice content to generate a semantic understanding vector; generates the target user's emotional tendency vector based on the intonation changes; evaluates the target user's decision-making representation vector based on the target user's emotional tendency vector and speech rate features; and performs multi-dimensional fusion analysis on the semantic understanding vector, the emotional tendency vector, and the decision-making representation vector to obtain the target user's purchase intention features. Based on the first visual characteristics, second visual characteristics, third visual characteristics, and purchase intention characteristics of the target user, an active intervention assessment is conducted on the target user to obtain an active intervention score corresponding to the target user; If the active intervention score corresponding to the target user is detected to exceed a preset threshold, then based on the target user's gaze location information on the retail terminal, candidate products that the target user has the intention to purchase are determined. Based on the target user's profile information and current context information, a proactive guidance script related to the candidate product is generated; and the proactive guidance script related to the candidate product is displayed to the target user through the retail terminal to proactively guide the target user during the shopping process.

[0007] Secondly, according to the present disclosure, a proactively guided smart retail device is provided, comprising: The data acquisition module is used to collect facial expression data, gaze direction data, body movement data, and voice interaction data of the target user relative to the retail terminal through pre-deployed sensors; The analysis module is used to perform facial state analysis on the facial expression data of the target user relative to the retail terminal to obtain the first visual feature; to perform gaze direction analysis on the gaze direction data of the target user relative to the retail terminal to determine the focus and degree of hesitation of the target user to obtain the second visual feature; and to perform body state analysis on the body movement data of the target user relative to the retail terminal to obtain the third visual feature. The extraction module is used to perform speech recognition and semantic analysis on the voice interaction data of the target user relative to the retail terminal, extracting the target user's voice content, intonation changes, and speech rate features; identify the target user's product consultation intent based on the voice content to generate a semantic understanding vector; generate the target user's emotional tendency vector based on the intonation changes; evaluate the target user's decision-making representation vector based on the target user's emotional tendency vector and speech rate features; and perform multi-dimensional fusion analysis on the semantic understanding vector, the emotional tendency vector, and the decision-making representation vector to obtain the target user's purchase intention features. The evaluation module is used to conduct an active intervention evaluation of the target user based on the first visual feature, the second visual feature, the third visual feature, and the purchase intention feature, and obtain an active intervention score corresponding to the target user. The determination module is used to determine, based on the target user's gaze location information on the retail terminal, candidate products that the target user has the intention to purchase if the active intervention score corresponding to the target user exceeds a preset threshold. The generation module is used to generate proactive guidance messages related to the candidate products based on the target user's profile information and current context information; and to display the proactive guidance messages related to the candidate products to the target user through the retail terminal, so as to proactively guide the target user during the shopping process.

[0008] Thirdly, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the proactively guided smart retail method as described in any of the above embodiments.

[0009] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the proactively guided smart retail method as described in any of the above embodiments.

[0010] The proactive guided smart retail method provided in this application collects facial expression data, gaze direction data, body movement data, and voice interaction data of target users relative to the retail terminal through pre-deployed sensors; analyzes the facial expression data of target users relative to the retail terminal to obtain first visual features; analyzes the gaze direction data of target users relative to the retail terminal to determine the focus and hesitation level of target users, obtaining second visual features; analyzes the body movement data of target users relative to the retail terminal to obtain third visual features; performs speech recognition and semantic parsing on the voice interaction data of target users relative to the retail terminal to extract the voice content, tone changes, and speech rate features of target users; identifies the product consultation intent of target users based on the voice content to generate a semantic understanding vector; and generates the target user's speech based on tone changes. The system analyzes the target user's emotional tendency vector; evaluates the target user's decision-making representation vector based on the target user's emotional tendency vector and speech rate characteristics; performs multi-dimensional fusion analysis of the semantic understanding vector, emotional tendency vector, and decision-making representation vector to obtain the target user's purchase intention characteristics; conducts proactive intervention evaluation on the target user based on the target user's first visual characteristics, second visual characteristics, third visual characteristics, and purchase intention characteristics to obtain the corresponding proactive intervention score; if the proactive intervention score of the target user exceeds a preset threshold, it identifies candidate products that the target user has the intention to purchase based on the target user's gaze point information at the retail terminal; generates proactive guidance scripts related to the candidate products based on the target user's profile information and current context information; and displays the proactive guidance scripts related to the candidate products to the target user through the retail terminal to proactively guide the target user during the shopping process. In this way, by analyzing real-time collected user visual and voice interaction information to generate proactive guidance scripts adapted to user needs, the traditional passive interaction is transformed into proactive guided interaction, providing users with precise guidance to facilitate transactions and effectively improving the user experience.

[0011] The above description is merely an overview of the technical solutions of the embodiments of this application. In order to better understand the technical means of the embodiments of this application and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of this application more obvious and understandable, specific implementation methods of this application are described below. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. It should be understood that the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure, wherein: Figure 1 This is a flowchart illustrating a proactive guided smart retail method disclosed herein.

[0013] Figure 2 This is a schematic diagram of the structure of an active-guided smart retail device disclosed herein.

[0014] Figure 3 This is a schematic diagram of the structure of a computer device provided in this disclosure.

[0015] It should be noted that the elements in the attached diagram are schematic and not drawn to scale. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are also within the scope of protection of this disclosure.

[0017] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this subject matter pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the specification and in the relevant art, and shall not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, the statement of “connecting” or “coupling” two or more parts together shall mean that these parts are directly joined together or joined through one or more intermediate components.

[0018] The term "embodiment" as used herein means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of the phrase "embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0019] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists, A and B exist simultaneously, or B exists. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or part of a component) from another component (or another part of a component).

[0020] In the description of this application, unless otherwise stated, "multiple" means two or more (including two), and similarly, "multiple groups" means two or more (including two groups).

[0021] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0022] Figure 1 This is a flowchart illustrating a proactively guided intelligent retail method provided in an embodiment of this disclosure, as shown below. Figure 1 As shown, the specific process of the proactive guided smart retail method includes: S110: Collect facial expression data, gaze direction data, body movement data, and voice interaction data of the target user relative to the retail terminal through pre-deployed sensors.

[0023] Among these, users' facial expression data, gaze direction data, and body movement data can be collected in real time through cameras deployed on retail terminals (such as smart vending machines); users' voice interaction data can be collected in real time through microphones deployed on retail terminals.

[0024] Facial expression data consists of captured facial image sequences, including dynamic changes in the eye, eyebrow, and mouth regions, as well as the overall facial contours. Gaze direction data represents the pointing angle of the user's eye axis in three-dimensional space. Body movement data may include the user's head posture, upper limb joint angles, hand positions, and movement trajectories. Voice interaction data comprises acoustic feature sequences of the raw audio signals after noise reduction, endpoint detection, and feature extraction, and may also include text content obtained through speech recognition and conversion.

[0025] S120. Perform facial expression analysis on the target user's facial expression data relative to the retail terminal to obtain the first visual feature; perform gaze direction analysis on the target user's gaze direction data relative to the retail terminal to determine the target user's focus and degree of hesitation, and obtain the second visual feature; perform body movement analysis on the target user's body movement data relative to the retail terminal to obtain the third visual feature.

[0026] The first visual feature can be used to describe the user's facial expression, such as confusion, curiosity, or disappointment; the second visual feature can be used to describe the user's attention to the product and the degree of hesitation; and the third visual feature can be used to describe the user's body language, such as crossed arms indicating defensiveness or hesitation, and leaning forward indicating interest.

[0027] In some embodiments, facial state analysis is performed on the facial expression data of the target user relative to the retail terminal to obtain a first visual feature, including: extracting the coordinates of facial geometric feature points from the facial expression data; identifying the facial expression category of the target user based on the facial geometric feature point coordinates; and constructing the first visual feature corresponding to the target user based on the facial expression category of the target user.

[0028] Among them, facial geometric feature point coordinates can effectively represent the contraction and relaxation state of a user's facial muscles; facial geometric feature point coordinates can include key point information such as eyebrow position coordinates, eye corner position coordinates, mouth corner curvature, and nose wing width. Facial expression category recognition can be performed through a deep learning-based facial expression recognition model, outputting the probability distribution of the target user's current expression category, thereby determining the dominant expression category; for example, confusion, curiosity, disappointment, happiness, surprise, etc. The constructed first visual features are represented in vector form, which can include dimensional information such as expression type encoding, expression confidence, expression duration, and expression change speed to quantify the user's expression state.

[0029] In some embodiments, the gaze direction data includes: eye position coordinates, pupil center coordinates, and eye movement trajectory information. The gaze direction data of the target user relative to the retail terminal is analyzed for gaze placement and wandering frequency to determine the target user's focus and degree of hesitation, resulting in a second visual feature. This includes: calculating a gaze direction vector based on the eye position coordinates and pupil center coordinates; determining the gaze placement coordinates of the target user on the retail terminal's display interface based on the gaze direction vector and the retail terminal's display interface coordinate system; calculating the gaze wandering frequency based on the number of changes in the gaze placement coordinates per unit time; identifying the target user's interested product category based on the distribution area of ​​the gaze placement coordinates on the retail terminal's display interface, and determining the target user's degree of hesitation based on the gaze wandering frequency; and constructing the target user's corresponding second visual feature based on the target user's interested product category and degree of hesitation.

[0030] The second visual feature is represented as a vector, which may include dimensions such as the code of the product category of interest, the heat distribution of gaze points, the duration of gaze dwell, the frequency of gaze wandering, and the level of hesitation. The code of the product category of interest can be used to identify the type of product the target user is currently focusing on, such as daily necessities or electronic products. The heat distribution of gaze points is formed by statistically analyzing the density of gaze points clustered in different areas of the display interface per unit time, reflecting the spatial distribution characteristics of user interest. The duration of gaze dwell records the continuous time a user's gaze is fixed on a specific product area, which can be used to assess the depth of the user's interest in that product. The frequency of gaze wandering directly quantifies the degree of uncertainty in the user's decision-making process; a higher frequency indicates a more pronounced comparison and weighing of different products. The level of hesitation can be calculated based on a combination of gaze wandering frequency and dwell time, and is divided into three levels: low, medium, and high. Therefore, through refined analysis of gaze direction data, the visual attention distribution and decision-making psychological state of the target user can be accurately captured.

[0031] In some embodiments, the body state analysis of the target user's body movement data relative to the retail terminal is performed to obtain third visual features, including: identifying the target user's dynamic and static behavioral features based on the body movement data; performing fusion analysis on the dynamic and static behavioral features to obtain the target user's interaction intention intensity and body tilt representation; and constructing the target user's corresponding third visual features based on the target user's interaction intention intensity and body tilt representation.

[0032] The dynamic behavioral characteristics may include: the target user's moving speed, the curvature of the moving trajectory, and the amplitude of limb swing in front of the retail terminal; the static behavioral characteristics may include: the target user's standing posture stability, the degree of head orientation fixation, and the state of arms crossing.

[0033] The intensity of interaction intention can be calculated using a weighted fusion method. The weight of dynamic behavioral features increases as the distance between the user and the retail terminal decreases, while the weight of static behavioral features increases significantly after the user's position stabilizes. Body tilt is determined by comprehensively considering the angle between the torso's central axis and the vertical direction, as well as the tilt direction; a forward-leaning posture is usually associated with a positive intention to explore products, while a sideways tilt suggests a user's tendency to compare adjacent display areas. Thus, through multi-dimensional analysis of body movement data, the body language information of target users in retail scenarios can be comprehensively characterized.

[0034] S130. Perform speech recognition and semantic analysis on the voice interaction data of the target user relative to the retail terminal, and extract the voice content, tone changes and speech rate features of the target user; identify the target user's product consultation intent based on the voice content to generate a semantic understanding vector; generate the target user's emotional tendency vector based on tone changes; evaluate the target user's decision representation vector based on the target user's emotional tendency vector and speech rate features; perform multi-dimensional fusion analysis on the semantic understanding vector, emotional tendency vector and decision representation vector to obtain the target user's purchase intention features.

[0035] Specifically, an intent classification model based on the Transformer architecture can be used to identify the target user's product inquiry intent, extracting key entity information such as product category, price range, and functional requirements to construct a high-dimensional semantic understanding vector. By extracting fundamental frequency contours, energy envelopes, and formant trajectories, emotional states such as excitement, hesitation, and indifference can be identified, generating an emotional tendency vector; for example, an upward tone can indicate positive emotions and purchase interest, while a downward tone reflects doubt or comparison. Speech rate feature calculation can include average speech rate, speech rate fluctuation coefficient, and pause frequency; understandably, a fast and coherent speech rate often corresponds to a clear purchase decision, while frequent pauses and a slower speech rate suggest information needs or decision-making difficulties.

[0036] Thus, semantic understanding vectors accurately identify user needs, emotion tendency vectors identify user feelings, and decision representation vectors determine whether users are ready to buy, thereby achieving real-time dynamic perception of the target user's shopping status.

[0037] S140. Based on the target user's first visual characteristics, second visual characteristics, third visual characteristics, and purchase intention characteristics, conduct an active intervention assessment of the target user to obtain the corresponding active intervention score.

[0038] Among them, the proactive intervention score corresponding to the target user can quantitatively represent the appropriateness of the target user's current acceptance of proactive shopping guidance.

[0039] In some embodiments, an active intervention assessment is performed on the target user based on the target user's first visual features, second visual features, third visual features, and purchase intention features to obtain the active intervention score corresponding to the target user. This includes: pre-training a user status assessment model for evaluating the user's active intervention; inputting the target user's first visual features, second visual features, third visual features, and purchase intention features into the user status assessment model to obtain the active intervention score corresponding to the target user.

[0040] The user status assessment model outputs a quantitative proactive intervention score based on the input features, ranging from 0 to 10. This score measures the urgency and likelihood of the user needing proactive guidance. For example, prolonged eye movement and a furrowed brow significantly increase this score.

[0041] S150. If the proactive intervention score corresponding to the target user is detected to exceed the preset threshold, then based on the target user's gaze point information at the retail terminal, determine the candidate products that the target user has the intention to purchase; based on the target user's profile information and current context information, generate proactive guidance scripts related to the candidate products; and display the proactive guidance scripts related to the candidate products to the target user through the retail terminal.

[0042] When the proactive intervention score exceeds a preset threshold, it can be determined that the target user has entered a state of "decision paralysis". At this time, the proactive intervention mechanism is triggered to proactively guide the user's shopping.

[0043] By displaying proactive guidance scripts related to candidate products to target users at retail terminals, this approach proactively guides them during their shopping process. Specifically, this can be achieved through speech synthesis or screen display, initiating an active dialogue. For example, based on user profiles and the current context, a large-scale language model generates personalized, empathetic guidance scripts, such as, "Hello, we noticed you lingered in front of several XX wines for a while. Are you hesitating about which one to choose as a business gift or for your personal collection?..." In some embodiments, gaze placement information is used to describe the area where the target user's gaze lingers on the retail terminal. Based on the gaze placement information of the target user on the retail terminal, candidate products that the target user intends to purchase are determined, including: spatially mapping the gaze placement area to the product display area of ​​the retail terminal to obtain the product location corresponding to the target user's gaze placement; statistically analyzing the cumulative gaze duration and cumulative gaze frequency of the target user at each product location to obtain the high-attention products corresponding to the target user; and determining candidate products that the target user intends to purchase based on the attribute information of the high-attention products.

[0044] For example, a retail terminal screen displays multiple products, and the target user's gaze moves back and forth across the screen. The system collects real-time data on the user's gaze placement, matches it with the coordinates of the preset product display area, and identifies the specific product location where the user's gaze lingers. For each product location, the system records the start and end times of the user's gaze, as well as the number of gazes, calculating the cumulative gaze duration and cumulative gaze count for each product. Assuming the target user's cumulative gaze duration for product A reaches 12 seconds and the number of gazes is 5; the cumulative gaze duration for product B is 8 seconds and the number of gazes is 3; and the gaze duration for the remaining products is less than 3 seconds; then products A and B can be identified as the target user's high-attention products.

[0045] If a target user's historical purchase history shows a preference for mid-to-high-end skincare products, and among the currently highly viewed products, product B is a mid-to-high-end moisturizing serum, is on a limited-time promotion, and has ample stock, while product A, although viewed for a longer period, belongs to the affordable facial cleanser category, then product B will be given a higher purchase intention weight and identified as the primary candidate product, while product A will be included as a secondary candidate product. Simultaneously, for product combinations with complementary attributes, such as skincare products and matching makeup tools, they can be packaged into a combined candidate solution to increase average order value and user satisfaction.

[0046] In some embodiments, the method further includes: acquiring the target user's eye contact response and voice response information to the proactive guidance dialogue; generating product introduction information / comparison information of multiple products based on the target user's eye contact response and voice response information to the proactive guidance dialogue; and displaying the product introduction information / comparison information of multiple products to the target user through a retail terminal to continuously guide the target user's dialogue response.

[0047] This can be achieved by calling specific tools, such as "get_product_comparison(A, B)", to retrieve information on several products the user is considering. This allows for continuous monitoring of user needs during proactive interaction, enabling timely responses to user requests.

[0048] In this embodiment, pre-deployed sensors collect facial expression data, gaze direction data, body movement data, and voice interaction data of the target user relative to the retail terminal. Facial expression data of the target user relative to the retail terminal is analyzed to obtain first visual features. Gazing direction data of the target user relative to the retail terminal is analyzed for gaze placement and wandering frequency to determine the target user's focus and degree of hesitation, obtaining second visual features. Body movement data of the target user relative to the retail terminal is analyzed to obtain third visual features. Voice interaction data of the target user relative to the retail terminal is subjected to speech recognition and semantic parsing to extract the target user's speech content, intonation changes, and speech rate features. The target user's product inquiry intent is identified based on the speech content to generate a semantic understanding vector. An emotional tendency vector of the target user is generated based on intonation changes. The system assesses the target user's decision-making representation vector based on their emotional tendency vector and speech rate characteristics. It then performs multi-dimensional fusion analysis of the semantic understanding vector, emotional tendency vector, and decision-making representation vector to obtain the target user's purchase intention characteristics. Based on the target user's first visual characteristics, second visual characteristics, third visual characteristics, and purchase intention characteristics, it conducts proactive intervention assessments to obtain a corresponding proactive intervention score. If the proactive intervention score exceeds a preset threshold, it identifies candidate products that the target user intends to purchase based on the target user's gaze location information at the retail terminal. Based on the target user's profile information and current context information, it generates proactive guidance scripts related to the candidate products and displays these scripts to the target user at the retail terminal, proactively guiding them during their shopping process. By analyzing real-time collected user visual and voice interaction information to generate proactive guidance scripts tailored to user needs, it transforms traditional passive interaction into proactive guided interaction, providing precise guidance to facilitate transactions and effectively improving the user experience.

[0049] In summary, this embodiment endows the intelligent agent with the "reading between the lines" ability similar to human sales experts. By continuously analyzing users' multimodal behavioral data (visual, voice, etc.), it assesses users' psychological state and purchasing intentions in real time. When it identifies key moments such as "decision paralysis," it proactively initiates interaction, provides precise guidance, and ultimately facilitates a transaction. This fundamentally changes the interaction paradigm, no longer waiting for users to ask questions, but proactively identifying and intervening in users' decision-making dilemmas. It can effectively seize sales opportunities that may be lost due to hesitation, improving conversion rates. Decision-making based on real-time, dynamic user visual behavior is more accurate and immediate. Interaction behavior is highly correlated with the user's current psychological state, making guidance more targeted and the user experience more natural. By introducing calculable and quantifiable indicators to determine the timing of proactive intervention, the vague human skill of "reading between the lines" becomes modeled and engineered, making the decision-making process evidence-based, avoiding ineffective or erroneous interference, and improving the accuracy of proactive services.

[0050] Figure 2 This is a schematic diagram of the structure of an actively guided smart retail device provided in this embodiment. The actively guided smart retail device may include: The data acquisition module 210 is used to collect facial expression data, gaze direction data, body movement data, and voice interaction data of the target user relative to the retail terminal through pre-deployed sensors.

[0051] The analysis module 220 is used to perform facial state analysis on the facial expression data of the target user relative to the retail terminal to obtain the first visual feature; to perform gaze direction analysis on the gaze direction data of the target user relative to the retail terminal to determine the focus and degree of hesitation of the target user to obtain the second visual feature; and to perform body state analysis on the body movement data of the target user relative to the retail terminal to obtain the third visual feature.

[0052] The extraction module 230 is used to perform speech recognition and semantic analysis on the voice interaction data of the target user relative to the retail terminal, extracting the target user's voice content, intonation changes, and speech rate features; identifying the target user's product consultation intent based on the voice content to generate a semantic understanding vector; generating the target user's emotional tendency vector based on intonation changes; evaluating the target user's decision representation vector based on the target user's emotional tendency vector and speech rate features; and performing multi-dimensional fusion analysis on the semantic understanding vector, emotional tendency vector, and decision representation vector to obtain the target user's purchase intention features.

[0053] The evaluation module 240 is used to conduct proactive intervention evaluation of target users based on their first visual characteristics, second visual characteristics, third visual characteristics, and purchase intention characteristics, and to obtain the proactive intervention score corresponding to the target user.

[0054] The determination module 250 is used to determine candidate products that the target user has the intention to purchase based on the target user's gaze location information at the retail terminal if the active intervention score corresponding to the target user is detected to exceed a preset threshold.

[0055] The generation module 260 is used to generate proactive guidance messages related to candidate products based on the target user's profile information and current context information; and to display the proactive guidance messages related to candidate products to the target user through the retail terminal, so as to proactively guide the target user during the shopping process.

[0056] In some embodiments, the analysis module 220 is specifically used for: Extract facial geometric feature point coordinates from facial expression data; identify the facial expression category of the target user based on the facial geometric feature point coordinates; construct the first visual feature corresponding to the target user based on the facial expression category.

[0057] In some embodiments, the gaze direction data includes: eye position coordinates, pupil center coordinates, and eye movement trajectory information.

[0058] Analysis module 220 is specifically used for: Calculate the gaze direction vector based on the eye position coordinates and pupil center coordinates; determine the gaze point coordinates of the target user on the retail terminal's display interface based on the gaze direction vector and the display interface coordinate system of the retail terminal; calculate the gaze wandering frequency based on the number of times the gaze point coordinates change per unit time; identify the target user's interested product category based on the distribution area of ​​the gaze point coordinates on the retail terminal's display interface, and determine the target user's degree of hesitation based on the gaze wandering frequency; construct the target user's corresponding second visual features based on the target user's interested product category and degree of hesitation.

[0059] In some embodiments, the analysis module 220 is specifically used for: Based on body movement data, identify the dynamic and static behavioral characteristics of the target user; perform fusion analysis on the dynamic and static behavioral characteristics to obtain the target user's interaction intention intensity and body tilt representation; construct the target user's corresponding third visual features based on the target user's interaction intention intensity and body tilt representation.

[0060] In some embodiments, the evaluation module 240 is specifically used for: A user status assessment model is pre-trained to evaluate user initiative and intervention. The target user's first visual feature, second visual feature, third visual feature, and purchase intention feature are input into the user status assessment model to obtain the target user's initiative and intervention score.

[0061] In some embodiments, gaze placement information is used to describe the area where a target user's gaze lingers on the retail terminal.

[0062] Module 250 is specifically used for: By spatially mapping the area where the target user's gaze rests to the product display area of ​​the retail terminal, the product location corresponding to the target user's gaze point is obtained; statistical analysis is performed on the cumulative gaze duration and cumulative gaze frequency of the target user at each product location to obtain the high-attention products corresponding to the target user; based on the attribute information of the high-attention products, candidate products that the target user has the intention to purchase are determined.

[0063] In some embodiments, a boot module is also included.

[0064] The guidance module is used to acquire the target user's eye and voice response information to the proactive guidance script; based on the target user's eye and voice response information to the proactive guidance script, it generates the corresponding product introduction information / comparison information of multiple products; and displays the corresponding product introduction information / comparison information of multiple products to the target user through the retail terminal to continuously guide the target user's response to the script.

[0065] The proactively guided smart retail device provided in this disclosure can execute the above-described method embodiments. Its specific implementation principle and technical effects can be found in the above-described method embodiments, and will not be repeated here.

[0066] This application also provides a computer device. Please refer to the following for details. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.

[0067] The computer device includes a memory 310 and a processor 320 that are interconnected via a system bus. It should be noted that only a computer device with memory 310 and processor 320 is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0068] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0069] The memory 310 includes at least one type of readable storage medium, including non-volatile memory or volatile memory, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. RAM may include static RAM or dynamic RAM. In some embodiments, the memory 310 may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 310 may also be an external storage device of the computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, or flash card equipped on the computer device. Of course, the memory 310 may include both internal storage units and external storage devices of the computer device. In this embodiment, the memory 310 is typically used to store the operating system and various application software installed on the computer device, such as the program code of the method described above. In addition, the memory 310 can also be used to temporarily store various types of data that have been output or will be output.

[0070] Processor 320 is typically used to perform overall operations of a computer device. In this embodiment, memory 310 is used to store program code or instructions, including computer operation instructions, and processor 320 is used to execute the program code or instructions stored in memory 310 or process data, such as program code that runs the methods described above.

[0071] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus system can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0072] Another embodiment of this application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads computer-readable program code stored in the computer-readable medium, enabling the processor to execute the functional actions specified in each step or combination of steps in the above method; and to generate means for implementing the functional actions specified in each block or combination of blocks in the block diagram.

[0073] Computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any suitable combination thereof, wherein the memory is used to store program code or instructions, the program code including computer operation instructions, and the processor is used to execute the program code or instructions of the above-described methods stored in the memory.

[0074] The definitions of memory and processor can be found in the description of the foregoing computer device embodiments, and will not be repeated here.

[0075] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0076] In the various embodiments of this application, the functional units or modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0077] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0078] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" as described in this application does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several units of these means may be embodied by the same item of hardware. The use of "first," "second," and "third," etc., does not indicate any order and these words should be interpreted as names. Unless otherwise specified, the steps in the above embodiments should not be construed as limiting the order of execution.

[0079] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A proactively guided intelligent retail method, characterized in that, include: Data on facial expressions, gaze direction, body movements, and voice interaction of target users relative to retail terminals are collected through pre-deployed sensors. Facial state analysis is performed on the facial expression data of the target user relative to the retail terminal to obtain the first visual feature; The gaze direction data of the target user relative to the retail terminal is analyzed to determine the focus and degree of hesitation of the target user, and a second visual feature is obtained. The body movement data of the target user relative to the retail terminal is analyzed to obtain third visual features; The voice interaction data of the target user relative to the retail terminal is subjected to speech recognition and semantic analysis to extract the voice content, intonation changes and speech rate features of the target user; The system identifies the target user's product inquiry intent based on the voice content to generate a semantic understanding vector; it generates the target user's emotional tendency vector based on the tone changes; it evaluates the target user's decision representation vector based on the emotional tendency vector and speech rate characteristics; and it performs multi-dimensional fusion analysis on the semantic understanding vector, the emotional tendency vector, and the decision representation vector to obtain the target user's purchase intention characteristics. Based on the first visual characteristics, second visual characteristics, third visual characteristics, and purchase intention characteristics of the target user, an active intervention assessment is conducted on the target user to obtain an active intervention score corresponding to the target user; If the active intervention score corresponding to the target user is detected to exceed a preset threshold, then based on the target user's gaze location information on the retail terminal, candidate products that the target user has the intention to purchase are determined. Based on the target user's profile information and the current context information, generate proactive guidance scripts related to the candidate products; The retail terminal displays proactive guidance messages related to the candidate products to the target user, thereby proactively guiding the target user's shopping process.

2. The method according to claim 1, characterized in that, Facial state analysis is performed on the facial expression data of the target user relative to the retail terminal to obtain the first visual features, including: Extract facial geometric feature point coordinates from the facial expression data; identify the facial expression category of the target user based on the facial geometric feature point coordinates; Based on the facial expression category of the target user, a first visual feature corresponding to the target user is constructed.

3. The method according to claim 1, characterized in that, The gaze direction data includes: eyeball position coordinates, pupil center coordinates, and eye movement trajectory information; the gaze direction data of the target user relative to the retail terminal is analyzed for gaze placement and wandering frequency to determine the target user's focus and degree of hesitation, obtaining second visual features, including: Calculate the gaze direction vector based on the eyeball position coordinates and the pupil center coordinates; determine the gaze point coordinates of the target user on the retail terminal's display interface based on the gaze direction vector and the retail terminal's display interface coordinate system; Calculate the line-of-sight shift frequency based on the number of times the coordinates of the line of sight change per unit time. Based on the distribution area of ​​the coordinates of the gaze points on the display interface of the retail terminal, the target user's product categories of interest are identified, and the degree of hesitation of the target user is judged based on the frequency of gaze movement. Based on the target user's interest in product categories and level of hesitation, a second visual feature corresponding to the target user is constructed.

4. The method according to claim 1, characterized in that, The body movement data of the target user relative to the retail terminal is analyzed to obtain third visual features, including: Based on the described body movement data, the dynamic and static behavioral characteristics of the target user are identified; By fusing and analyzing the dynamic and static behavioral characteristics, the interaction intention intensity and body tilt characteristics of the target user are obtained. Based on the target user's interaction willingness intensity and body tilt characteristics, construct the third visual features corresponding to the target user.

5. The method according to claim 1, characterized in that, Based on the target user's first visual characteristics, second visual characteristics, third visual characteristics, and purchase intention characteristics, an active intervention assessment is conducted on the target user to obtain an active intervention score corresponding to the target user, including: Pre-train a user status assessment model to evaluate user initiative and intervention. The first visual feature, the second visual feature, the third visual feature, and the purchase intention feature of the target user are input into the user status assessment model to obtain the active intervention score corresponding to the target user.

6. The method according to claim 1, characterized in that, The gaze placement information is used to describe the area where the target user's gaze rests on the retail terminal; Based on the target user's line of sight at the retail terminal, candidate products that the target user intends to purchase are determined, including: By spatially mapping the area where the gaze rests to the merchandise display area of ​​the retail terminal, the location of the merchandise corresponding to the point where the target user's gaze falls is obtained; Statistical analysis was performed on the cumulative gaze duration and cumulative gaze frequency of the target user at each product location to obtain the high-attention products corresponding to the target user; Based on the attribute information of the high-attention products, candidate products that the target user has the intention to purchase are identified.

7. The method according to claim 1, characterized in that, Also includes: Obtain the target user's eye contact response and voice response to the proactive guidance dialogue; Based on the target user's eye contact and voice response to the proactive guidance script, generate product introduction information / comparison information for multiple products; The retail terminal displays product information and comparisons of multiple products to the target user, thereby continuously guiding the target user's response to the conversation.

8. A proactively guided intelligent retail device, characterized in that, include: The data acquisition module is used to collect facial expression data, gaze direction data, body movement data, and voice interaction data of the target user relative to the retail terminal through pre-deployed sensors; The analysis module is used to perform facial state analysis on the facial expression data of the target user relative to the retail terminal to obtain the first visual feature; The gaze direction data of the target user relative to the retail terminal is analyzed to determine the focus and degree of hesitation of the target user, and a second visual feature is obtained. The body movement data of the target user relative to the retail terminal is analyzed to obtain third visual features; The extraction module is used to perform speech recognition and semantic analysis on the voice interaction data of the target user relative to the retail terminal, and to extract the voice content, intonation changes and speech rate features of the target user; The system identifies the target user's product inquiry intent based on the voice content to generate a semantic understanding vector; it generates the target user's emotional tendency vector based on the tone changes; it evaluates the target user's decision representation vector based on the emotional tendency vector and speech rate characteristics; and it performs multi-dimensional fusion analysis on the semantic understanding vector, the emotional tendency vector, and the decision representation vector to obtain the target user's purchase intention characteristics. The evaluation module is used to conduct an active intervention evaluation of the target user based on the first visual feature, the second visual feature, the third visual feature, and the purchase intention feature, and obtain an active intervention score corresponding to the target user. The determination module is used to determine, based on the target user's gaze location information on the retail terminal, candidate products that the target user has the intention to purchase if the active intervention score corresponding to the target user exceeds a preset threshold. The generation module is used to generate proactive guidance scripts related to the candidate products based on the target user's profile information and current context information. The retail terminal displays proactive guidance messages related to the candidate products to the target user, thereby proactively guiding the target user's shopping process.

9. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the proactive guided smart retail method as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the proactive guided smart retail method as described in any one of claims 1 to 7.