Intelligent dialogue auxiliary method and device based on AR glasses and electronic equipment
AR glasses can be used to obtain multi-dimensional features of the conversation partner in real time, generate user portraits and emotional states, and dynamically adjust conversation strategies, thus solving the problem of poor communication in initial interactions and improving the accuracy and fluency of communication.
Patent Information
- Application Number
- CN202510894980.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-03
AI Technical Summary
In interpersonal communication, especially in the first interaction or when there is a lack of background information, it is difficult to accurately grasp the other party's interests and emotional state, which leads to incorrect topic selection and awkward communication. Traditional subjective judgment makes it difficult to achieve accurate matching.
AR glasses can be used to obtain multi-dimensional feature groups of conversation partners in real time, including visual and audio information, generate user portraits and emotional states, dynamically adjust conversation guidance strategies, and provide visual, tactile, and audio feedback.
It achieves comprehensive perception of the dialogue partner, reduces communication blindness, improves the fluency and accuracy of communication, enhances the sense of participation and recognition, and avoids communication deadlocks.
Smart Images

Figure CN120748009A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular to an intelligent conversation assistance method, device, and electronic device based on AR glasses. Background Art
[0002] In traditional interpersonal communication, information acquisition between people mainly relies on visual observation, language communication and experience judgment.
[0003] Currently, when users are interacting with their conversation partners for the first time or lack background information, it's difficult to quickly establish a comprehensive understanding of them. For example, when facing a stranger, users can only make subjective guesses based on superficial words and limited non-verbal cues, unable to accurately grasp the other person's true interests or emotional state, which can easily lead to poor topic selection and awkward communication. Furthermore, different people have individual differences in their expression habits and emotional responses, making it difficult to achieve accurate matching based solely on subjective judgment, which is detrimental to the conversation between users and their conversation partners.
[0004] Therefore, there is an urgent need for an intelligent conversation assistance method, device and electronic equipment based on AR glasses. Summary of the Invention
[0005] The present application provides an intelligent conversation assistance method, device, and electronic device based on AR glasses to facilitate conversations between users and conversation partners.
[0006] In a first aspect of the present application, an intelligent conversation assistance method based on AR glasses is provided, the method comprising: obtaining a conversation feature group for a second user during a conversation between a first user and a second user, wherein the first user wears AR glasses; generating a user portrait of the second user based on the conversation feature group; determining an emotional state of the second user based on the conversation feature group; determining a conversation guidance strategy for the second user based on the user portrait and the emotional state; and providing the conversation guidance strategy to the first user through the AR glasses.
[0007] By employing the above technical solution, AR glasses can capture the conversation partner's conversational features in real time, including language, facial expressions, and body movements. This provides the first user with a comprehensive understanding of the conversation partner's status. This eliminates information asymmetry between people, especially during initial interactions or unfamiliar situations. The first user can use the AR device to quickly understand the second user's characteristics and emotions, avoiding blind spots in communication. It automatically captures the other party's potential interests, focus, and emotional fluctuations, surpassing the limitations of human observation. Based on the conversational features, a user profile is generated for the conversation partner, including interests, preferences, behavioral habits, and personality traits. This user profile provides an accurate reference for conversation content, enabling personalized conversation suggestions based on the other party's characteristics and preferences. If the user profile is stored, it can be continuously refined over multiple exchanges, making subsequent interactions smoother and more familiar. This overcomes the drawback of traditional communication, where emotion recognition relies on subjective judgment, allowing the first user to adjust their tone, speed, and delivery through emotion analysis. If the second user expresses negativity or becomes distracted, the first user can adjust their strategy promptly, avoiding communication deadlocks or misunderstandings. Based on the guidance strategy, the first user can choose an expression method that best suits the other party's interests and emotional state, enhancing their conversation partner's sense of engagement and acceptance. During the conversation, as the feature set and emotional state change, the guidance strategy can also be dynamically adjusted to ensure that the conversation remains smooth and well-directed. This facilitates conversations between the user and the conversation partner.
[0008] Optionally, obtaining a conversation feature group for the second user during a conversation between the first user and the second user specifically includes: receiving original visual features for the second user sent by the AR glasses, the original visual features including the second user's facial expressions, clothing, accessories, movements, and the environment in which the conversation with the first user takes place; receiving original audio features for the second user sent by the AR glasses, the original audio features including the second user's speaking speed, intonation, and pause rhythm; performing data processing on the original visual features and the original audio features to obtain the conversation feature group, the data processing including denoising, filtering, feature extraction, and normalization.
[0009] By employing the above technical solution, a multi-dimensional perception of the second user is established by acquiring visual and audio features. Rather than relying on a single signal source, this system combines visual and audio information to generate a more complete and accurate set of conversational features. Key features can be flexibly extracted in various scenarios, whether face-to-face or remote. By integrating multimodal information, the user's state and behavioral characteristics are better restored, providing high-quality input for subsequent user profiling and sentiment analysis. The user's facial expressions, clothing, accessories, movements, and environmental characteristics are captured, not only facilitating individual behavior analysis but also integrating them with context. Micro-expression changes are captured to capture emotional fluctuations, enabling real-time emotional state recognition. By analyzing the user's appearance, interests, hobbies, or personality style are inferred, enriching the user profile. Body language is interpreted to reveal the other party's communication intent or current mood. By identifying background objects or scenes and integrating them with context, the content direction and emotional tone of conversation suggestions can be dynamically adjusted. By analyzing speech rate, intonation, and pause rhythm, the user's underlying intentions and emotional tendencies in their speech are captured. Identifying language organization by pause intervals and frequency reveals the depth of thought or hesitation in the other person's conversation, providing clues for subsequent guidance strategies. Removing environmental noise that may have entered the data collection process ensures the purity of feature data and reduces analysis bias. Improving the accuracy of key information by optimizing signal frequency bands. Extracting key features from high-dimensional raw data, such as muscle movement parameters for facial expressions and pitch curves for speech, simplifies computational complexity for subsequent modeling. Standardizing feature data from different sources ensures consistency during multimodal feature fusion, avoiding model bias due to differences in feature dimensions.
[0010] Optionally, the conversation feature group includes the second user's clothing features, accessory features, and features of the environment in which the conversation with the first user takes place. Generating a user portrait of the second user based on the conversation feature group specifically includes: determining the membership scores corresponding to the clothing features, accessory features, and features of the environment in which the conversation with the first user takes place, with one feature corresponding to one membership score; performing weighted fusion of the membership scores with the weights corresponding to each feature to calculate a fuzzy membership category of the second user; matching interest preferences and topic tags in a preset database based on the fuzzy membership category; and determining a user portrait of the second user based on the fuzzy membership category, the interest preferences, and the topic tags.
[0011] By employing this technical solution, clothing and accessories reflect the second user's style, taste, and financial status, while environmental features provide contextual clues. This multi-dimensional information integration makes the user profile more specific and realistic. Multimodal data supports real-time context perception, allowing the user's profile to be dynamically updated even as the conversation partner changes. Fuzzy logic is used to calculate a membership score for each feature, and fuzzy membership categories are obtained through weighted fusion. By weightedly merging the membership scores with the corresponding feature weights, the contribution of each feature to the profile is aligned with its actual importance. Fuzzy membership categories classify user features into multiple possible categories rather than a single label. Fuzzy classification more accurately reflects the complexity and multidimensionality of user characteristics in reality, avoiding recognition bias caused by overly strict classification boundaries. Based on fuzzy membership categories, the user's interests and preferences and topic tags are matched against a pre-set database. Based on the fuzzy membership categories, interests, preferences, and topic tags, a comprehensive user profile is generated, integrating the membership categories, interests, preferences, and topic tags to facilitate conversations between the user and the conversation partner.
[0012] Optionally, the dialogue feature group also includes facial expression features, movement features, speech speed features, intonation features and pause rhythm features of the second user, and determining the emotional state of the second user based on the dialogue feature group specifically includes: mapping the facial expression features, the movement features, the speech speed features, the intonation features and the pause rhythm features into an emotional state space, where the emotional state space is a two-dimensional space, the X-axis of the emotional state space is used to represent the temporal change from positive emotion to negative emotion, and the Y-axis of the emotional state space is used to represent the temporal change from high-energy emotion to low-energy emotion; calculating the emotional coordinate point of the second user based on the emotional state space; and determining the emotional state of the second user based on the position of the emotional coordinate point in the emotional state space.
[0013] By employing the above technical solution, which combines facial expression features, movement features, speech rate, intonation, and pause rhythm, comprehensive emotional cues are captured through various sensory channels. Facial expressions and movements reflect non-verbal emotional characteristics, while speech rate, intonation, and pause rhythm reveal emotional information in language. Multimodal fusion enhances the integrity of emotion analysis. Single features can be affected by external factors. Integrating multimodal signals reduces the probability of misjudgment and improves recognition reliability. Compared to simple single labels, the continuity of coordinate points can more subtly describe the intensity and changing trends of emotional states. Based on temporal changes, the emotion state space can monitor the transition of emotions from high-energy positive to low-energy negative in real time, making it suitable for dynamic conversational scenarios. Coordinate point calculation simplifies the analysis and storage of emotional states, enabling emotion recognition models to more efficiently process complex multimodal inputs. Quantified emotion coordinate points facilitate classification, regression, or clustering analysis in machine learning models, further improving the accuracy of emotion prediction. Using coordinate points to define emotion category boundaries reduces ambiguity caused by subjective definitions and improves the consistency of system judgments.
[0014] Optionally, determining a conversation guidance strategy for the second user based on the user portrait and the emotional state specifically includes: obtaining emotional features according to the emotional state using Mel spectrum analysis and RNN network extraction; determining interest correlation features between the second user and the target item according to the user portrait in combination with a time series; determining a target feature vector based on the emotional features and the interest correlation features; and determining the conversation guidance strategy according to the directionality of the target feature vector.
[0015] By employing this technical solution, a user's emotional state and interest preferences are jointly modeled, comprehensively considering the other person's current mood swings and long-term interest preferences. This avoids relying solely on interests or emotions. Even if the second user has clear interests, if they are currently in a negative mood, the system can prioritize adjusting the tone or topic atmosphere rather than directly addressing their interests. Based on both emotional and interest characteristics, the system dynamically adapts to the second user's state, achieving a truly personalized conversation guidance strategy. Mel spectrum analysis and a recurrent neural network are used to process emotion-related audio data, fully exploiting the emotional information contained in speech rate, intonation, and rhythm. Mel spectrum analysis captures subtle changes in the audio signal, and combined with the time series modeling capabilities of the recurrent neural network, it accurately extracts emotional dynamics. Compared to static analysis methods, the recurrent neural network can process the changing trends of user emotions throughout the conversation, providing a basis for dynamically adjusting the conversation strategy. By integrating emotional and interest-related features into a target feature vector, a high-dimensional data model is constructed for conversation strategy generation. By analyzing the directionality of the target feature vector, the system can generate more flexible conversation strategies. The directionality of the target feature vector enables the system to anticipate user needs and provide guidance, preemptively introducing relevant information or suggestions.
[0016] Optionally, providing the dialogue guidance strategy to the first user through the AR glasses specifically includes: providing visual guidance to the first user according to the dialogue guidance strategy, the visual guidance including focus highlighting and color tone fine-tuning; and providing tactile feedback and ambient sound feedback to the first user according to the dialogue guidance strategy.
[0017] By employing the above technical solution, highlighting helps the first user quickly locate the second user's points of interest or emotional expression. It eliminates environmental distractions, helping the first user focus on the most relevant information and avoid overlooking the second user's nonverbal communication signals. By changing the color tone of the visual environment, the first user's emotional perception or understanding of the surrounding environment is adjusted. Haptic feedback is a low-intrusion interaction method that conveys information to the first user through subtle vibrations or pulses. The first user can quickly perceive system prompts during a conversation and make corresponding adjustments. Non-disruptive audio prompts remind the first user of key moments. Different feedback forms are suitable for various situations and environments. Multiple prompts can complement each other, reducing the risk of misjudgment or omissions caused by a single prompt method. The integration of vision, touch, and sound creates a natural and smooth interactive experience that does not cause users to feel distracted or burdened. It provides intuitive and subtle guidance, helping the first user easily control the rhythm and direction of the conversation and enhance communication confidence.
[0018] Optionally, the method further includes: acquiring historical conversation features of the first user; and fusing the historical conversation features with the conversation guidance strategy to generate a personalized guidance strategy.
[0019] By adopting the above technical solution, historical features are combined with dialogue guidance strategies, and the generated guidance plan can be more in line with the habits and expressions of the first user. Personalized guidance avoids the discomfort that may be caused by a unified and standardized guidance strategy, allowing the first user to use the system more naturally and smoothly. Adjusting the strategy based on the habits and advantages of the first user can make the guidance content more easily adopted. Historical dialogue features can reflect the interaction patterns between the first user and different objects in different situations. By integrating these features with the real-time features of the current dialogue, the generated personalized guidance strategy can dynamically adapt to the changing dialogue environment. Personalized strategies can make users feel the thoughtful design of the system and avoid the alienation or dissatisfaction that may be caused by stereotyped guidance content. The guidance content is in line with the language style and expression habits of the first user, making them more relaxed and confident in the conversation.
[0020] In a second aspect of the present application, an intelligent dialogue assistance device based on AR glasses is provided, wherein the intelligent dialogue assistance device includes an acquisition module and a processing module, wherein the acquisition module is used to obtain a dialogue feature group for a second user during a dialogue between a first user and a second user, and the first user wears AR glasses; the processing module is used to generate a user portrait of the second user based on the dialogue feature group; the processing module is also used to determine the emotional state of the second user based on the dialogue feature group; the processing module is also used to determine a dialogue guidance strategy for the second user based on the user portrait and the emotional state; the processing module is also used to provide the dialogue guidance strategy to the first user through the AR glasses.
[0021] In a third aspect of the present application, an electronic device is provided, which includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs the method described above.
[0022] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions. When the instructions are executed, the method described above is executed.
[0023] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: AR glasses capture the conversational feature set of the conversation partner in real time, including information such as language, facial expressions, and body movements. This allows the first user to fully perceive the conversation partner's status. This eliminates information asymmetry between people, especially during initial interactions or unfamiliar situations. The first user can use the AR device to quickly understand the second user's characteristics and emotions, avoiding blind spots in communication. It automatically captures the other party's potential interests, attention spans, and emotional fluctuations, surpassing the limitations of human observation. Based on the conversational feature set, a user profile of the conversation partner is generated, including interests, preferences, behavioral habits, and personality traits. This user profile provides an accurate reference for conversation content, enabling personalized conversation suggestions based on the other party's characteristics and preferences. If the user profile is stored, it can be continuously refined over multiple exchanges, making subsequent interactions smoother and more familiar. This overcomes the drawback of traditional communication, where emotion recognition relies on subjective judgment, allowing the first user to adjust their tone, speed, and delivery through emotion analysis. If the second user expresses negativity or is distracted, the first user can adjust their strategy promptly to avoid communication deadlocks or misunderstandings. Based on the guidance strategy, the first user can choose a delivery method that best suits the other party's interests and emotional state, enhancing the conversation partner's sense of engagement and acceptance. During the conversation, as the feature set and emotional state change, the guidance strategy can also be dynamically adjusted to ensure that the conversation always maintains good fluency and direction, thus facilitating the conversation between the user and the conversation partner. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 A flowchart of an intelligent conversation assistance method based on AR glasses provided in an embodiment of the present application; Figure 2 Another flowchart of an intelligent conversation assistance method based on AR glasses provided in an embodiment of the present application; Figure 3 A schematic diagram of a module of an intelligent conversation assistance device based on AR glasses provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0025] Explanation of the reference numerals: 31, acquisition module; 32, processing module; 41, processor; 42, communication bus; 43, user interface; 44, network interface; 45, memory. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0027] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0028] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0029] In traditional interpersonal communication, information acquisition mainly relies on visual observation, language communication and experience judgment.
[0030] However, when users and conversation partners are meeting for the first time or lack sufficient background information, it's difficult to fully understand them quickly. For example, when facing strangers, users are often limited to subjective inferences based on superficial language and limited non-verbal cues. This makes it difficult to accurately grasp the other person's interests, emotions, or personality, which can easily lead to poor topic selection and awkward communication. Furthermore, individual differences in expression habits and emotional responses make it difficult to achieve accurate interaction matching based solely on subjective judgment, thus affecting the quality of communication between users and conversation partners.
[0031] In order to solve the above technical problems, this application provides an intelligent conversation assistance method based on AR glasses, referring to Figure 1 , Figure 1 This is a flow chart of an intelligent conversation assistance method based on AR glasses provided in an embodiment of the present application. The intelligent conversation assistance method is applied to a server and includes steps S110 to S150, which are as follows: S110: Obtain a conversation feature group for the second user during a conversation between the first user and the second user, where the first user wears AR glasses.
[0032] Specifically, the server is responsible for obtaining various data features about the second user from the AR glasses. These features include visual and audio information. The server processes and analyzes this data to form a conversation feature group, which helps the system better understand the second user's status, interests, etc., and thus optimize the quality of the conversation. The AR glasses worn by the first user act as a tool for information collection and processing. AR glasses use cameras, microphones, and sensors to capture the surrounding environment and the second user's behavior, expressions, voice, and other information, and transmit it to the server for processing. The conversation feature group refers to the collected information that describes the second user, such as the second user's facial expressions, body language, clothing, posture, etc.; the second user's voice tone, speaking speed, pauses, etc.; the interaction scene between the second user and the first user, such as whether it is in a quiet room, whether there is background noise, etc.
[0033] In one possible implementation, obtaining a conversation feature group for the second user during a conversation between the first user and the second user specifically includes: receiving original visual features for the second user sent by the AR glasses, the original visual features including the second user's facial expressions, clothing, accessories, movements, and the environment in which the conversation with the first user takes place; receiving original audio features for the second user sent by the AR glasses, the original audio features including the second user's speaking speed, intonation, and pause rhythm; performing data processing on the original visual features and the original audio features to obtain a conversation feature group, the data processing including denoising, filtering, feature extraction, and normalization.
[0034] Specifically, AR glasses use a camera to capture the second user's facial expressions, which can reveal their emotions and attitudes. For example, a smile may indicate friendliness, while a frown may indicate dissatisfaction or confusion. The camera also captures the second user's clothing and analyzes information such as the formality and color of the clothing. This helps the server infer the second user's identity, occasion, or psychological state. For example, a formal suit may indicate a business setting, while casual attire may indicate a casual occasion. AR glasses also capture the second user's accessories, such as glasses, watches, and jewelry, which can provide additional social information. For example, an expensive watch may indicate a preference for luxury goods. Movements include the user's body language and gestures. Body language and gestures can help infer the second user's emotions and engagement. For example, crossed arms may indicate defensiveness or uneasiness, while extensive gestures may indicate enthusiasm and engagement in the conversation. The server also analyzes the environment in which the second user and the first user are conversing, such as whether indoors or outdoors, and whether it is quiet or noisy. This helps understand the atmosphere and context of the conversation. AR glasses use a microphone to capture the second user's speaking rate. A faster speaking rate may indicate nervousness or excitement, while a slower speaking rate may indicate calmness or reflection. Intonation reflects the emotional fluctuations of the second user while speaking. A high-pitched tone may indicate anger or excitement, while a low tone may indicate calmness or seriousness. Pause rhythm refers to the duration of pauses the second user makes during the conversation. Longer pauses may indicate deliberation or hesitation, while shorter pauses may indicate smooth speech.
[0035] Denoising is used to remove noise that may be generated during the sensor data collection process, such as background noise or device errors. Denoising improves data accuracy. Filters remove unnecessary or irrelevant frequency components, retaining meaningful signals for more accurate analysis of user behavior and emotions. Feature extraction extracts features useful for analysis. For example, it can extract features such as "smile" or "frown" from facial expressions, or extract "speech rate" and "pause" data from audio signals. Normalization converts data with different features into a unified standard range to avoid bias caused by varying data scales. For example, it adjusts speech rate and facial expression values to the same scale to ensure accurate subsequent analysis. After the above processing, a conversation feature set is generated. This feature set contains all the visual and audio features displayed by the second user in the conversation, facilitating subsequent analysis and modeling.
[0036] S120: Generate a user profile of the second user based on the conversation feature group.
[0037] Specifically, the user profile provides an overview of the second user across multiple dimensions, including: Personal characteristics: Basic information such as age, gender, and occupation; Emotional state: Inferring the second user's emotions (e.g., positive, negative, anxious, etc.) based on facial expressions, tone of voice, and speaking speed during the conversation; Interest preferences: Inferring the second user's interests and hobbies by analyzing the second user's interests during the conversation, such as the topic content and tone of voice; Communication style: Including conversation rhythm, language style, and body language, reflecting the second user's communication style (e.g., formal or casual, open or reserved). The server comprehensively analyzes the conversation features collected by the AR glasses. For example, based on the second user's facial expressions and speaking speed, the system can infer their emotions. For example, a smile and moderate speaking speed may indicate friendliness and relaxation. Further enriching the profile based on tone of voice, clothing, and gestures: Formal attire may indicate a business setting. The server then combines and analyzes these features using a specific algorithm to generate a more comprehensive profile of the second user.
[0038] In one possible implementation, the conversation feature group includes the second user's clothing features, accessory features, and features of the conversation environment with the first user. Based on the conversation feature group, a user portrait of the second user is generated, specifically including: determining the membership scores corresponding to the clothing features, accessory features, and features of the conversation environment with the first user, with one feature corresponding to one membership score; performing weighted fusion of the membership scores with the weights corresponding to each feature to calculate the fuzzy membership category of the second user; based on the fuzzy membership categories, matching interest preferences and topic tags in a preset database; and determining the user portrait of the second user based on the fuzzy membership categories, interest preferences, and topic tags.
[0039] Specifically, the conversation feature groups include: the second user's clothing style, color, and design. For example, wearing a formal suit creates a different impression than wearing casual clothing. The second user's accessories, such as watches, glasses, and jewelry, can reveal a user's taste, financial status, or social status. Conversation context features include the location or setting of the conversation, which may reflect the formality and social context of the conversation. For example, whether the conversation takes place in a cafe, a business meeting room, or a social setting can all influence the conversation's atmosphere. For each feature, the server calculates a membership score using a standardized method. The membership score reflects the degree to which each feature belongs to a certain category, using fuzzy logic for reasoning. For example, if the second user is wearing a suit, the system can determine the formality of the clothing based on the clothing features, such as "formal" or "casual," assigning a higher membership score to "formal" and a lower score to "casual." Similarly, a user wearing an expensive watch might be considered to be in the "high-income" category, thus assigning a higher membership score to this feature. Membership scores are not simply added together, but rather weighted based on the importance or weight of each feature. This process takes into account the membership scores of different features to more accurately reflect the characteristics of the second user. For example, if the user's clothing is considered to be more representative of the user's personality than the conversation environment, then the clothing feature may be given a greater weight.
[0040] The server calculates the second user's fuzzy membership category by combining the weighted membership scores. This is a comprehensive category that represents certain attributes of the second user, such as interests, personality, and social style. For example, the fuzzy membership category might indicate that the second user is a "high-end business professional" or a "young, casual, and social person." Within a pre-set database, the server uses the fuzzy membership category to match the second user's interests and preferences with topic tags. These tags and preferences are predicted based on past data and experience, helping the server understand the user's likely topics of interest or discussion areas. For example, if the second user is identified as a "high-end business professional," the server might associate topic tags related to "finance" and "business trends." If the second user exhibits a more casual personality, the server might match interest tags such as "entertainment" and "travel." Ultimately, based on this analysis, the server generates a user profile for the second user, encompassing their interests, emotions, communication style, and other characteristics. This profile provides guidance for interactions with the second user, enabling the first user to adjust their communication style and topic selection.
[0041] S130: Determine the emotional state of the second user based on the conversation feature group.
[0042] Specifically, by analyzing facial expressions and voice features, the server can calculate the emotional coordinates of the second user and infer their emotional state. For example, if the second user's facial expression shows a smile, and their speech rate is moderate and their tone is pleasant, the server might judge their emotional state as "high energy, positive." If the second user frowns, speaks faster, and pauses frequently, the server might judge their emotional state as "low energy, negative." This inference of emotional state helps adjust the conversation, ensuring that the first user can appropriately respond to the second user's emotional changes.
[0043] In a possible implementation, the dialogue feature group also includes facial expression features, movement features, speech speed features, intonation features, and pause rhythm features of the second user. Based on the dialogue feature group, the emotional state of the second user is determined, specifically including: mapping the facial expression features, movement features, speech speed features, intonation features, and pause rhythm features into an emotional state space, where the emotional state space is a two-dimensional space, where the X-axis of the emotional state space is used to represent the temporal change from positive emotion to negative emotion, and the Y-axis of the emotional state space is used to represent the temporal change from high-energy emotion to low-energy emotion; based on the emotional state space, the emotional coordinate point of the second user is calculated; and based on the position of the emotional coordinate point in the emotional state space, the emotional state of the second user is determined.
[0044] Specifically, the server maps these conversation features into an emotional state space. This space is typically two-dimensional, with two coordinate axes: the X-axis represents the range of emotions from positive to negative, typically from "happy" to "sad" or from "joyful" to "angry." The Y-axis represents the range of emotions from high energy to low energy, typically from "active and excited" to "calm and tired." For example, the positive direction of the X-axis can represent positive emotions such as happiness and contentment, while the negative direction represents negative emotions such as anger and frustration. Areas above the Y-axis represent high energy, such as excitement and tension, while areas below the Y-axis represent low energy, such as calm and tired. The server calculates an emotional coordinate point based on features such as the second user's facial expression, speech rate, intonation, movements, and pauses. The specific location of this coordinate point in the emotional state space is determined by the combined performance of various features. For example, a smile might result in a positive X-axis value, indicating a happy emotion; a frown might result in a negative X-axis value, indicating a negative emotion. Fast speech, often indicative of tension or excitement, might increase the Y-axis value, indicating a higher energy emotion. A high-pitched voice tone may indicate excitement or anger, further pushing up the energy value on the Y-axis. Openness in body language may mean positive emotions, while defensive postures, such as crossed arms, may mean negative emotions. The server uses this data to convert it into coordinate points in the emotional space, and the specific location represents the emotional state of the user. The position of the emotional coordinate point in the emotional state space determines the user's specific emotional state. For example: If the coordinate point is in the positive direction of the X-axis and above the Y-axis, this may indicate that the user is in a positive and energetic mood, such as excitement or happiness. If the coordinate point is in the negative direction of the X-axis and below the Y-axis, this may indicate that the user is low, tired, or depressed. Through this coordinate point, the server can accurately determine the second user's current emotional state, such as "high energy, positive" or "low energy, negative."
[0045] S140: Determine a conversation guidance strategy for the second user based on the user portrait and emotional state.
[0046] Specifically, conversation guidance strategies refer to strategies provided by the server to the first user to help guide and optimize the conversation with the second user. These strategies are customized based on the emotional state and profile of the second user. By analyzing the second user's emotions and personality traits, the server can recommend how to adjust the topic, tone, interaction method, etc. to enhance the effectiveness of the conversation. For example, if the user is in high spirits, the server may suggest continuing to discuss points of interest in depth to maintain the lively atmosphere of the conversation; if the user is in a low mood, the server may suggest providing comfort or changing the topic through a more caring and understanding tone.
[0047] In one possible implementation, a conversation guidance strategy for the second user is determined based on the user profile and emotional state, specifically including: obtaining emotional features based on the emotional state using Mel spectrum analysis and RNN network extraction; determining the interest correlation features between the second user and the target item based on the user profile in combination with the time series; determining the target feature vector based on the emotional features and interest correlation features; and determining the conversation guidance strategy based on the directionality of the target feature vector.
[0048] Specifically, suppose the second user has shown a particular interest in sports, particularly basketball, in past conversations and wears branded sportswear. Based on this information, the server identifies that the second user may have a strong interest in sports-related topics and may prefer sports-related conversations. In this conversation, the second user speaks quickly, with a smile on their face and an upbeat tone, indicating a positive mood. Based on these characteristics, the server analyzes the second user's emotional state as "high energy, positive emotions." Given the second user's high spirits, the server may recommend that the first user maintain a positive and engaging conversation and choose light-hearted topics related to sports, such as basketball. Based on the second user's interest in basketball, the server may suggest that the first user mention recent basketball games, team performances, or inquire about the second user's basketball experience and preferences. If the second user expresses an interest in other topics—for example, if the second user wears sneakers—the server detects that they may be interested in brands and fashion and may steer the conversation toward topics related to fashion and sports brands.
[0049] For example, if the second user displays high energy and a cheerful mood during a conversation, and the server determines that they have a strong interest in basketball, the server can provide the following conversation guidance strategy for the first user: The server might prompt the first user with, "I heard there was an exciting NBA game recently. What did you think?" This maintains a positive conversational atmosphere. The server might also remind the first user to pay attention to the brand of sneakers the second user is wearing: "What new sneakers did you buy recently? I heard XXX brand recently released new models, and they look great." Therefore, by combining user profiles and emotional states, the server can provide the first user with a specific conversation guidance strategy, helping the first user better understand the second user's interests and emotions, thereby optimizing the conversation content and making it smoother, more interactive, and more engaging. This emotion- and interest-based guidance strategy can effectively improve conversation quality and the communication experience for both parties.
[0050] S150: Providing a conversation guidance strategy to the first user through the AR glasses.
[0051] Specifically, AR glasses, as an augmented reality device, can provide real-time feedback to the wearer through vision, touch, and even sound. In the embodiment of the present application, AR glasses become an information display tool that can feed back the guidance strategy generated by the server to the first user in an appropriate manner.
[0052] In one possible implementation, a conversation guidance strategy is provided to the first user through AR glasses, specifically including: visually guiding the first user according to the conversation guidance strategy, the visual guidance including focus highlighting and color tone fine-tuning; and providing tactile feedback and ambient sound feedback to the first user according to the conversation guidance strategy.
[0053] Specifically, AR glasses attract the wearer's attention by highlighting certain elements or information in the wearer's field of view. For example, if the second user mentions a certain topic or displays a change in mood, the AR glasses can highlight relevant information or icons, helping the wearer better understand the conversation and respond accordingly. By adjusting the color tone of displayed content, AR glasses can influence the wearer's mood and attention. For example, when the second user expresses positive emotions, the AR glasses might use warm colors to enhance the conversational atmosphere; conversely, when the second user is depressed, cool colors might be used to indicate that the conversation needs to be adjusted. AR glasses may also feature haptic feedback. When the atmosphere or mood of the conversation changes, the AR glasses may vibrate or other tactile means to remind the first user to adjust their tone or topic. For example, if the second user begins to express displeasure, the AR glasses may vibrate to remind the first user to change the topic or tone. In some cases, AR glasses can also provide strategic advice through audio cues. For example, if the first user is unsure whether to continue a certain topic, the AR glasses may play audio suggestions through headphones or speakers: "The second user may be interested in the topic of sneakers" or "It's time to move on to a more lighthearted topic."
[0054] In one possible implementation, refer to Figure 2 , Figure 2 Another flow chart of an intelligent conversation assistance method based on AR glasses provided in an embodiment of the present application includes steps S210 to S220, and the above steps are as follows: S210, obtaining the historical conversation features of the first user; S220, integrating the historical conversation features with the conversation guidance strategy to generate a personalized guidance strategy.
[0055] Specifically, historical conversation features refer to information such as the behavior, tone of voice, interests, and frequently used topics displayed by the first user in past interactions with various conversation partners, such as friends, colleagues, and family. This information can be analyzed through conversation content, voice features, and interaction patterns to form a comprehensive understanding of the user's communication style, interests, and emotional responses. For example, if the first user tends to steer conversations towards personal interests such as travel and movies in past conversations, or if they typically display high emotional resonance in conversations, these historical features can help the server identify the user's conversational preferences. In real-time conversations, the server needs to combine the first user's historical conversation features with the current conversation guidance strategy for the second user to generate personalized guidance recommendations. Specifically, historical conversation features provide a user's past behavioral patterns, while the current conversation guidance strategy proposes conversational directions and recommendations based on factors such as the second user's interests and emotional state. For example, if the first user has demonstrated a high ability to remain calm in emotionally charged conversations in the past, the server might incorporate this feature to recommend strategies for maintaining composure and appropriately changing the topic in the current conversation. By integrating historical features with the current guidance strategy, the generated personalized guidance strategy will better suit the first user's communication habits, helping them interact more smoothly with the second user in the current conversation. For example, if the server recognizes that the first user has a strong interest in the topic of "technology" and has historically tended to lead more in-depth conversations, it may provide a guidance statement such as "What are your views on the development of artificial intelligence?" in the current conversation to trigger in-depth conversation.
[0056] This application also provides an intelligent conversation assistance device based on AR glasses, referring to Figure 3 , Figure 3 This is a block diagram of an intelligent conversation assistance device based on AR glasses, provided in an embodiment of the present application. The intelligent conversation assistance device is a server, comprising an acquisition module 31 and a processing module 32. Acquisition module 31 acquires a conversation feature group for a first user during a conversation between the first user and the second user, where the first user is wearing AR glasses. Processing module 32 generates a user profile for the second user based on the conversation feature group. Processing module 32 determines the emotional state of the second user based on the conversation feature group. Processing module 32 determines a conversation guidance strategy for the second user based on the user profile and emotional state. Processing module 32 provides the conversation guidance strategy to the first user via the AR glasses.
[0057] In a possible implementation, the acquisition module 31 acquires a conversation feature group for the second user during a conversation between the first user and the second user, specifically including: the acquisition module 31 receives original visual features for the second user sent by the AR glasses, the original visual features include the second user's facial expressions, clothing, accessories, movements, and the conversation environment with the first user; the acquisition module 31 receives original audio features for the second user sent by the AR glasses, the original audio features include the second user's speaking speed, intonation, and pause rhythm; the processing module 32 performs data processing on the original visual features and the original audio features to obtain a conversation feature group, and the data processing includes denoising, filtering, feature extraction, and normalization.
[0058] In one possible implementation, the conversation feature group includes the second user's clothing features, accessory features, and features of the conversation environment with the first user. The processing module 32 generates a user portrait of the second user based on the conversation feature group, specifically including: the processing module 32 determines the membership scores corresponding to the clothing features, accessory features, and features of the conversation environment with the first user, with one feature corresponding to one membership score; the processing module 32 performs a weighted fusion of the membership scores and the weights corresponding to each feature to calculate the fuzzy membership category of the second user; the processing module 32 matches the interest preferences and topic tags in a preset database based on the fuzzy membership categories; the processing module 32 determines the user portrait of the second user based on the fuzzy membership categories, interest preferences, and topic tags.
[0059] In a possible embodiment, the dialogue feature group also includes facial expression features, movement features, speech speed features, intonation features and pause rhythm features of the second user. The processing module 32 determines the emotional state of the second user based on the dialogue feature group, specifically including: the processing module 32 maps the facial expression features, movement features, speech speed features, intonation features and pause rhythm features into the emotional state space, where the emotional state space is a two-dimensional space, the X-axis of the emotional state space is used to represent the temporal change from positive emotion to negative emotion, and the Y-axis of the emotional state space is used to represent the temporal change from high-energy emotion to low-energy emotion; the processing module 32 calculates the emotional coordinate point of the second user based on the emotional state space; the processing module 32 determines the emotional state of the second user based on the position of the emotional coordinate point in the emotional state space.
[0060] In one possible implementation, the processing module 32 determines a conversation guidance strategy for the second user based on the user portrait and emotional state, specifically including: the processing module 32 uses Mel spectrum analysis and RNN network extraction to obtain emotional features according to the emotional state; the processing module 32 determines the interest correlation features between the second user and the target item according to the user portrait in combination with the time series; the processing module 32 determines the target feature vector based on the emotional features and the interest correlation features; the processing module 32 determines the conversation guidance strategy according to the directionality of the target feature vector.
[0061] In one possible implementation, the processing module 32 provides a dialogue guidance strategy to the first user through the AR glasses, specifically including: the processing module 32 provides visual guidance to the first user according to the dialogue guidance strategy, the visual guidance including focus highlighting and color tone fine-tuning; the processing module 32 provides tactile feedback and ambient sound feedback to the first user according to the dialogue guidance strategy.
[0062] In a possible implementation, the acquisition module 31 acquires historical conversation features of the first user; the processing module 32 integrates the historical conversation features with the conversation guidance strategy to generate a personalized guidance strategy.
[0063] It should be noted that the above embodiments provide devices that implement their functions using only the division of the above functional modules as examples. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0064] This application also provides an electronic device, referring to Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include: at least one processor 41, at least one network interface 44, a user interface 43, a memory 45, and at least one communication bus 42.
[0065] The communication bus 42 is used to realize the connection and communication between these components.
[0066] The user interface 43 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 43 may also include a standard wired interface and a wireless interface.
[0067] The network interface 44 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0068] The processor 41 may include one or more processing cores. Using various interfaces and circuits, the processor 41 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 45, as well as accesses data stored in the memory 45, to perform various server functions and process data. Optionally, the processor 41 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 41 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may also be implemented as a separate chip, rather than integrated into the processor 41.
[0069] Among them, the memory 45 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 45 includes a non-transitory computer-readable storage medium. The memory 45 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 45 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 45 may also be optionally at least one storage device located away from the aforementioned processor 41. As Figure 4 As shown, the memory 45 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program for an intelligent conversation assistance method based on AR glasses.
[0070] exist Figure 4In the electronic device shown, the user interface 43 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 41 can be used to call an application stored in the memory 45 for an intelligent dialogue assistance method based on AR glasses. When executed by one or more processors, the electronic device executes one or more methods in the above embodiments.
[0071] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.
[0072] The present application also provides a computer-readable storage medium storing instructions, which, when executed by one or more processors, enable an electronic device to execute one or more of the methods described in the above embodiments.
[0073] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0074] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely schematic, such as the division of units, which is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0075] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0076] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0077] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.
[0078] The above is only an exemplary embodiment of the present disclosure and cannot be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made according to the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification and the truth of practice, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any variation, use or adaptive change of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary technical means in the art that are not recorded in the present disclosure. The description and examples are to be regarded as exemplary only, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. An intelligent conversation assistance method based on AR glasses, characterized in that: The method comprises: Obtaining a conversation feature group for a second user during a conversation between a first user and the second user, where the first user wears AR glasses; generating a user profile of the second user based on the conversation feature group; determining an emotional state of the second user based on the conversation feature group; determining a conversation guidance strategy for the second user based on the user profile and the emotional state; The conversation guidance strategy is provided to the first user through the AR glasses.
2. The intelligent conversation assistance method based on AR glasses according to claim 1, characterized in that: The acquiring of a conversation feature group for the second user during the conversation between the first user and the second user specifically includes: receiving original visual features for the second user sent by the AR glasses, the original visual features including the second user's facial expressions, clothing, accessories, actions, and a conversation environment with the first user; receiving original audio features for the second user sent by the AR glasses, the original audio features including a speaking speed, intonation, and pause rhythm of the second user; The original visual features and the original audio features are processed to obtain the dialogue feature group, wherein the data processing includes denoising, filtering, feature extraction, and normalization.
3. The intelligent conversation assistance method based on AR glasses according to claim 1, characterized in that: The conversation feature group includes clothing features and accessory features of the second user and features of the conversation environment with the first user. Generating a user profile of the second user based on the conversation feature group specifically includes: Determining membership scores corresponding to the clothing feature, the accessory feature, and the feature of the environment in which the conversation with the first user takes place, where each feature corresponds to one membership score; Performing weighted fusion on the membership score and the weight corresponding to each feature to calculate a fuzzy membership category of the second user; According to the fuzzy membership categories, interest preferences and topic tags are matched in a preset database; Based on the fuzzy membership category, the interest preference, and the topic tag, a user profile of the second user is determined.
4. The intelligent conversation assistance method based on AR glasses according to claim 3 is characterized in that: The conversation feature group further includes facial expression features, movement features, speech speed features, intonation features, and pause rhythm features of the second user. Determining the emotional state of the second user based on the conversation feature group specifically includes: Mapping the facial expression features, the action features, the speech rate features, the intonation features, and the pause rhythm features into an emotional state space, wherein the emotional state space is a two-dimensional space, wherein the X-axis of the emotional state space is used to represent the temporal change from positive emotion to negative emotion, and the Y-axis of the emotional state space is used to represent the temporal change from high-energy emotion to low-energy emotion; Calculating the emotion coordinate point of the second user according to the emotion state space; The emotional state of the second user is determined according to the position of the emotional coordinate point in the emotional state space.
5. The intelligent conversation assistance method based on AR glasses according to claim 4 is characterized in that: The determining, based on the user profile and the emotional state, a conversation guidance strategy for the second user specifically includes: According to the emotional state, emotional features are obtained by using Mel spectrum analysis and RNN network extraction; Determining, based on the user portrait and in combination with the time series, the interest correlation characteristics between the second user and the target item; Determining a target feature vector based on the emotion feature and the interest-related feature; The dialogue guidance strategy is determined according to the directionality of the target feature vector.
6. The intelligent conversation assistance method based on AR glasses according to claim 1, characterized in that: Providing the conversation guidance strategy to the first user through the AR glasses specifically includes: Performing visual guidance on the first user according to the dialogue guidance strategy, wherein the visual guidance includes highlighting focus and fine-tuning color tone; According to the dialogue guidance strategy, tactile feedback and environmental sound feedback are provided to the first user.
7. The intelligent conversation assistance method based on AR glasses according to claim 1, characterized in that: The method further comprises: Obtaining historical conversation features of the first user; The historical conversation features are integrated with the conversation guidance strategy to generate a personalized guidance strategy.
8. An intelligent conversation assistance device based on AR glasses, characterized in that: The intelligent dialogue assistance device includes an acquisition module (31) and a processing module (32), wherein: The acquisition module (31) is used to acquire a conversation feature group for a second user during a conversation between a first user and the second user, wherein the first user wears AR glasses; The processing module (32) is used to generate a user profile of the second user based on the conversation feature group; The processing module (32) is further configured to determine the emotional state of the second user based on the conversation feature group; The processing module (32) is further configured to determine a conversation guidance strategy for the second user based on the user portrait and the emotional state; The processing module (32) is further configured to provide the conversation guidance strategy to the first user via the AR glasses.
9. An electronic device, characterized in that: The electronic device comprises a processor (41), a memory (45), a user interface (43) and a network interface (44), wherein the memory (45) is used to store instructions, the user interface (43) and the network interface (44) are both used to communicate with other devices, and the processor (41) is used to execute the instructions stored in the memory (45) so that the electronic device executes the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is performed.
Citation Information
Cited By
Augmented reality interaction method and device, electronic equipment and storage medium
CN121326155A
Extended reality interaction method and apparatus, electronic device, and storage medium
CN121326155B