Immersive dialogue method, intelligent glasses and storage medium
Acquisition and analysis of user emotional data through smart glasses, generate emotional response types and display relevant scene special effects in the augmented reality space, solving the problems of insufficient immersion and lack of emotions in the prior art, and significantly improving the user's chat experience.
Patent Information
- Application Number
- CN202510423043.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
AI Technical Summary
The existing online chat and virtual interaction technologies have problems of insufficient immersion and lack of emotions, and cannot effectively reproduce changes in the real world and deeply understand users' emotions.
Emotional data is obtained through smart glasses, emotional characteristics are extracted and emotional response types are generated, target scene effects associated with emotional response types are determined, and these effects are dynamically displayed in the augmented reality space to improve the user's immersion and interactive experience.
It accurately and in real time to capture the user's emotional characteristics during the chat process, and dynamically display scene effects related to emotions through augmented reality technology, which significantly enhances the user's immersion and interactive experience.
Smart Images

Figure CN119937799A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technology, and in particular to an immersive dialogue method, smart glasses and a storage medium. Background Art
[0002] In traditional online chat and virtual interaction technologies, lack of immersion and lack of emotion are two major factors that affect user experience.
[0003] Lack of immersion refers to the lack of realism and engagement of users in a virtual environment. Traditional artificial intelligence (AI) chats, such as Candy.ai, usually rely on two-dimensional interfaces and lack the realism of three-dimensional space, which makes users feel isolated and unreal when communicating. Although some systems try to enhance immersion through virtual reality (VR) and augmented reality (AR) technologies, there are still many shortcomings. For example, existing systems often lack simulation of dynamic environments and cannot effectively reproduce changes in the real world, such as weather, lighting and object movement, changes in chat emotions, etc., which causes users to feel monotonous and boring in virtual environments.
[0004] Lack of emotion means that existing systems often lack the ability to deeply understand and identify user emotions. Although some systems try to identify user emotions through text analysis, these methods can only capture the surface emotional state and cannot deeply understand the user's true feelings and emotional changes. This limitation makes the system unable to provide a truly personalized and emotional interactive experience.
[0005] In summary, current online chat and virtual interaction technologies suffer from insufficient immersion and lack of emotion. Summary of the invention
[0006] An immersive conversation method, smart glasses, and storage medium are provided in the embodiments of the present application, which can generate an interactive effect that is more in line with the user's emotional state and increase the immersion of the user's voice interaction.
[0007] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions: In a first aspect, an embodiment of the present application provides an immersive conversation method, the method comprising: Acquiring emotional data involved in a conversation interaction function of the smart glasses; Extracting emotional features from the emotional data, and generating emotional response types based on the emotional features; determining a target scene effect associated with the type of emotional response; The target scene special effects are dynamically displayed in the augmented reality space constructed by the smart glasses.
[0008] The wearer's chat partner can be a digital person, or a person or animal in the real world. When the chat partner is a person in the real world, the wearer can have a face-to-face conversation with him or her, or a remote conversation.
[0009] The emotional data involved in the conversation interaction function of the smart glasses may be the emotional data of the wearer of the smart glasses during the conversation, or the emotional data of the wearer's chat partner during the conversation, or the emotional data of the wearer and the chat partner during the conversation.
[0010] In combination with the first aspect, in a possible design manner, the dynamically displaying the target scene special effects in the augmented reality space constructed by the smart glasses includes: Controlling the target scene special effects to move in the augmented reality space constructed by the smart glasses; During the movement of the target scene special effect, the real scene is captured, and when the target scene special effect interacts with the real scene, the display effect of the target scene special effect is dynamically adjusted.
[0011] In combination with the first aspect, in a possible design manner, dynamically adjusting the display effect of the target scene special effect when the target scene special effect interacts with the real scene includes: When the target scene special effect interacts with the real scene, identifying an interactive object in the real scene that interacts with the target scene special effect; Based on the movement speed of the target scene special effect and / or the physical properties of the interactive object, performing motion simulation on the target scene special effect to obtain a motion simulation result; Based on the motion simulation result, the display effect of the target scene special effect is dynamically adjusted.
[0012] In combination with the first aspect, in a possible design method, dynamically adjusting the display effect of the target scene special effect based on the motion simulation result includes: Based on the motion simulation result, rendering the target scene special effects to obtain a special effects rendering image; The special effect rendering image is superimposed on the augmented reality space to obtain a display effect of the target scene special effect.
[0013] In combination with the first aspect, in a possible design manner, before dynamically displaying the target scene special effect in the augmented reality space constructed by the smart glasses, the method further includes: Acquire posture information of a wearer of the smart glasses in a real scene; The effective range of the target special effect scene in the augmented reality space is determined based on the posture information.
[0014] In combination with the first aspect, in a possible design manner, obtaining the posture information of the wearer in the real scene includes: Acquire image information of a real scene captured by the smart glasses; The image information is processed by a SLAM algorithm to obtain posture information of the wearer in the real scene.
[0015] In combination with the first aspect, in a possible design method, the target scene special effects include at least one of snowflakes, bubbles, and raindrops, and the display effects include at least one of dissolving effects, rupture effects, and elastic deformation effects.
[0016] In combination with the first aspect, in a possible design manner, the emotion data includes voice data of a wearer of the smart glasses, and extracting emotion features from the emotion data and generating an emotion response type based on the emotion features includes: Recognizing the voice data to obtain a conversation text; Extracting features of the voice data and the dialogue text corresponding to the voice data through a pre-trained dialogue emotion generation network to obtain the emotional features of the wearer; The dialogue emotion generation network is used to generate an emotion response type based on the emotion feature.
[0017] In a second aspect, an embodiment of the present application provides an immersive conversation device, the device comprising: The receiving module is used to obtain the emotion data involved in the dialogue interaction function of the smart glasses.
[0018] The emotion generation module is used to extract emotion features from the emotion data and generate emotion response types based on the emotion features.
[0019] The special effects selection module is used to determine the target scene special effects associated with the emotional response type.
[0020] A display module is used to dynamically display the target scene special effects in the augmented reality space constructed by the smart glasses.
[0021] In a third aspect, an embodiment of the present application provides an MR device, comprising at least one processor, a memory, a camera module, a microphone module, and a display module, wherein the at least one processor is coupled to the memory for reading and executing instructions in the memory to execute the method of the first aspect and its possible design methods.
[0022] In a fourth aspect, an embodiment of the present application provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the method of the first aspect and its possible design manner when running.
[0023] Compared with the prior art, an immersive conversation method, smart glasses and storage medium provided in an embodiment of the present application, wherein the method is applied to smart glasses, in which the smart glasses extract emotional features from the emotional data by acquiring the emotional data involved in the conversation interaction function. Then, an emotional response type is generated based on the emotional features, the target scene special effects associated with the emotional response type are determined, and then the target scene special effects are dynamically displayed in the constructed augmented reality space. In this way, it is possible to accurately and in real time capture the emotional characteristics of the wearer, and dynamically display the scene special effects related to the emotions through augmented reality technology, thereby significantly improving the user's sense of immersion and the beneficial effects of the interactive experience during the chat process. This method can be applied to education, entertainment, medical care and other fields, bringing emotional value to users in chats and improving the user's device usage experience.
[0024] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A diagram showing an application environment of an immersive conversation method provided by an embodiment of the present application is shown; Figure 2 A flowchart of an immersive conversation method provided by an embodiment of the present application is shown; Figure 3 A flowchart of a method for generating emotions based on dialogue provided in an embodiment of the present application is shown; Figure 4 A schematic diagram of a dialogue-based emotion generation method provided in an embodiment of the present application is shown; Figure 5 A flow chart of a method for dynamically displaying target scene special effects provided by an embodiment of the present application is shown; Figure 6 A flow chart of a motion simulation method provided by an embodiment of the present application is shown; Figure 7 A flow chart of a scene special effects rendering method provided by an embodiment of the present application is shown; Figure 8A flow chart of a method for generating a scene special effect effective range provided by an embodiment of the present application is shown; Fig. 9 A schematic diagram of the hardware structure of an immersive conversation device provided in an embodiment of the present application is shown; Fig.10 A hardware structure block diagram of smart glasses provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0026] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0027] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by people with general skills in the technical field to which this application belongs. The words "one", "a", "the", "these" and the like in this application do not indicate a quantitative limitation, and they may be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusions; for example, a process, method and system, product or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "multiple" involved in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there may be three relationships, for example, "A and / or B" may mean: A exists alone, A and B exist at the same time, and B exists alone. Generally, the character " / " indicates that the objects associated with each other are in an "or" relationship. The terms "first", "second", "third", etc. in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.
[0028] The development of technology has brought great changes to online chat, and users have increased their demands for immersive chat. How to improve users' sense of reality and participation in a virtual chat environment, as well as deeply understand users' emotions and bring users satisfaction and immersion, has become a current research hotspot. The solution of this application can be applied in human-computer interaction scenarios, such as human-computer interaction in terminals and human-computer interaction scenarios in vehicle-mounted systems. Figure 1 An application environment diagram of an immersive conversation method provided by an embodiment of the present application is shown, such as Figure 1As shown, the terminal 102 communicates with the server 104 through the network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can obtain the emotional data involved in the conversation interaction function of the smart glasses; extract emotional features from the emotional data, and generate emotional response types based on the emotional features; determine the target scene special effects associated with the emotional response type; and dynamically display the target scene special effects in the augmented reality space constructed by the smart glasses. Or the above-mentioned process of extracting emotional features is put into the server 104 for execution, such as the server 104 extracts emotional features from the emotional data, and generates emotional response types based on the emotional features. The server 104 sends the emotional response type to the terminal 102, and the terminal 102 determines the target scene special effects associated with the emotional response type from the server 104, and dynamically displays the target scene special effects in the augmented reality space constructed by the terminal 102. Alternatively, the model for extracting emotional features is placed in the server 104 for training. After training, the server 104 sends the model to the terminal 102. The terminal extracts features from the emotional data based on the deployed model and generates an emotional response type.
[0029] The terminal may specifically include a virtual reality headset (VR headset), smart glasses, electronic display screens, and mixed reality (MR) devices, etc. MR devices may include MR glasses, MR helmets, MR cameras, etc. Smart glasses include augmented reality glasses (AR glasses), MR glasses, etc. The vehicle-mounted system may be a vehicle-mounted chip, a vehicle-mounted device (such as a vehicle machine with the ability to build an augmented reality space, a vehicle-mounted computer, etc.), etc.
[0030] Figure 2 A flowchart of an immersive conversation method provided by an embodiment of the present application is shown. Figure 2 The method shown is applied to the fields of education, entertainment, medical treatment, etc. The wearer wears smart glasses, which may be AR glasses. The method includes steps S201 to S204.
[0031] Step S201: The smart glasses obtain emotional data involved in the dialogue interaction function.
[0032] The wearer's chat partner can be a digital person or a user.
[0033] The emotional data involved in the conversation interaction function of the smart glasses may be the emotional data of the wearer of the smart glasses (and / or the chat partner) during the conversation. The following description will be made by taking the wearer as an example.
[0034] In one aspect, the emotion data may be the wearer's emotion words, such as sad, happy, excited.
[0035] On the other hand, emotional data can be emotional information expressed by the wearer during the conversation through voice, facial expressions, gestures, etc. If the wearer's facial expression is crying, the emotional data received by the smart glasses is sadness. Specifically, the smart glasses capture these emotional information in real time through the built-in microphone module, camera module and other sensors (such as heart rate sensors). For example, voice data is obtained by recording the wearer's voice by the microphone module of the smart glasses, including tone, speaking speed, intonation, semantics, etc. Facial expressions are obtained by capturing changes in the wearer's facial expressions through the camera module of the smart glasses, such as smiling, frowning, etc. Physiological data is obtained through the heart rate sensor, eye tracking module, etc. of the smart glasses, including heart rate and eye tracking data.
[0036] In some embodiments, the emotion data can be obtained by combining multiple data such as voice, expression, gesture, heart rate, etc. This embodiment uses multimodal data to comprehensively analyze the wearer's current emotion, making the analysis result more accurate.
[0037] In some embodiments, the weight of the emotion data can be adjusted dynamically by analyzing the context information of the conversation. For example, in a long conversation, the smart glasses can judge the wearer's emotional trend based on the context. When it is determined that the wearer's emotions fluctuate violently, the weight of the wearer's expression, gesture, and heart rate in the emotion data will be increased.
[0038] In some embodiments, the smart glasses capture real-world environmental information to assist in emotion analysis. For example, dim environment data is used as emotion data to infer that the wearer is in an easily emotional state, and combined with other data, the user's sadness can be identified.
[0039] Through step S201, the smart glasses capture emotional data in real time, thereby being able to quickly respond to the wearer's needs.
[0040] Step S202: The smart glasses extract emotional features from the emotional data and generate emotional response types based on the emotional features.
[0041] Emotional features refer to the key information extracted from emotional data that can reflect the wearer's emotional state. Emotional features include voice features, expression features, heart rate features, etc.
[0042] In some of the embodiments, speech is converted into text using speech recognition technology, and spectral features (such as pitch, intensity) and emotional keywords of the speech are analyzed to obtain speech features.
[0043] In some of the embodiments, computer vision technology is used to analyze key feature points of facial expressions (such as changes in the positions of mouth corners and eyebrows) to obtain facial features.
[0044] After obtaining the emotional features by one or more of the above methods, the smart glasses generate an emotional response type corresponding to the emotional features.
[0045] Please refer to Figure 3 , which shows a flowchart of a dialogue-based emotion generation method provided in an embodiment of the present application, the method comprising steps S301 to S303.
[0046] Step S301: Recognize the voice data to obtain the dialogue text.
[0047] Step S302: extract features of the voice data and the dialogue text corresponding to the voice data through a pre-trained dialogue emotion generation network to obtain the wearer's emotional features.
[0048] Step S303: using the dialogue emotion generation network to generate an emotion response type based on the emotion features.
[0049] For more information about this method, see Figure 4 , which shows a schematic diagram of a dialogue-based emotion generation method provided in an embodiment of the present application, Figure 4 In the process, the speech data is converted into dialogue text through the speech recognition Transformer network. Then the speech data and dialogue text are input into the dialogue emotion generation Transformer network to generate the emotion response type, which is the emotion type that the dialogue system wants to present. For example, when the wearer cries, the emotion type that the dialogue system wants to present is comfort emotion.
[0050] The following is an explanation of the speech recognition Transformer network.
[0051] ASR (automated speech recognition) algorithm is a technology that converts human speech into computer-readable text. Using transformer to achieve end-to-end speech recognition has strong performance and potential. Transformer not only improves the accuracy of recognition, but also provides an effective solution for real-time ASR systems by reducing inference time. The implementation of the algorithm includes two stages: training and inference. In the training stage, speech data with text labels is used as the training set, and the cross entropy loss function is used to calculate the loss. The parameters of the transformer network are updated through gradient back propagation. The calculation method is as follows: ; Among them, x represents input, i represents data number, and x irepresents the i-th input, i is greater than or equal to 0 and less than or equal to n-1, n represents the amount of data, y represents the label, and LOSS represents the loss.
[0052] The following is an explanation of the dialogue sentiment generation Transformer network.
[0053] The essence of dialogue emotion generation is a classification task, and the purpose is to output targeted response emotions based on the emotions contained in the dialogue. Its input is a continuous dialogue voice and text, and the output is the emotion of the speech that the chatbot needs to reply (equivalent to the emotional response type in the embodiment of the present application). The dialogue emotion generation algorithm is implemented by a deep learning network composed of transformers, and the algorithm includes two parts: training and reasoning. In the training stage, a large amount of text data with emotion labels is prepared first. In order to make the chat have a positive guidance, the generated emotion labels include neutral, pleasant, encouraging, comforting, fondness, and sympathy. Based on these labeled data, the cross entropy loss function is used to calculate the Loss, and the parameters of the transformer network are updated by gradient back propagation. For the structure of the network, please refer to the above formula and its related introduction.
[0054] In the inference stage, the conversation speech and text data to be predicted are input, and the emotional response type is generated using the trained network parameters.
[0055] Through the method recorded in step S301 to step S303, the smart glasses use speech recognition technology to capture the user's voice data, perform sentiment analysis through a deep learning model, and identify the user's emotional state. This process involves extracting acoustic features from speech signals, such as rhythm, sound quality, and spectral features, which are crucial for emotion recognition. Through these features, smart glasses can more accurately capture the user's emotional changes, thereby generating an emotional response type that is more in line with the user's emotional state. Due to the high accuracy of sentiment analysis, the generated emotional response type is also more in line with the current chat atmosphere, providing the wearer with an emotional chat experience.
[0056] The above describes how to generate an emotional response type. After determining the emotional response type, the smart glasses determine the display effects of the virtual environment according to the emotional response type. For details, please refer to the description of step S203.
[0057] Step S203: The smart glasses determine a target scene effect associated with the emotional response type.
[0058] In some of these embodiments, the smart glasses determine a target scene effect associated with an emotional response type from preset scene effects.
[0059] Specifically, the smart glasses can preset multiple scene effects, each of which is associated with a different type of emotional response, and the target scene feature refers to the visual effects associated with the emotional response type in the scene effect. In other words, the smart glasses will select the appropriate scene effect from the preset scene effect library according to the emotional response type currently generated.
[0060] Specifically, step S203 includes two parts: special effect library construction and special effect selection.
[0061] Special effects library construction: The preset scene special effects library includes a variety of special effects (such as snowflakes, bubbles, raindrops, ripple lines, etc.), each of which is associated with a specific type of emotional response.
[0062] Effect selection: Select the target scene effect from the effect library based on the emotional response type. For example, if the emotional response type is "comfort", select the "bubble" effect.
[0063] In some embodiments, the parameters of the target scene special effects can be dynamically generated according to the type of emotional response and the intensity of the emotion, and the parameters include color, size, speed, etc. For example, the higher the intensity of the emotion, the smart glasses can generate more vivid and active special effects.
[0064] In some embodiments, the special effects library can be customized by the user (such as the wearer), for example, the user draws the scene special effects; or the user selects the scene special effects from the existing materials; or the user uploads a picture, and the smart glasses generate the scene special effects based on the picture. As an example, the user can select the "starry sky" special effect as the target scene special effect of the "active response" type. This embodiment provides a richer and more personalized visual experience through user-defined scene special effects.
[0065] In some embodiments, the display effect of the special effect is adjusted according to the real time and / or environmental information (such as light intensity, background color). For example, in a dim real environment, the smart glasses can generate scene special effects with lower brightness to reduce visual impact. In this embodiment, the special effect is associated with the real scene, and the display effect of the special effect is dynamically adjusted when the scene changes, so that the scene special effect and the environment are more natural and harmonious.
[0066] In other embodiments, the smart glasses generate target scene effects associated with the emotional response type through generative AI. Exemplarily, a text prompt for describing the emotional response type is generated, and the target scene effects are output by inputting the text prompt into an existing generative AI tool.
[0067] Step S204: The smart glasses dynamically display the target scene special effects in the constructed augmented reality space.
[0068] This step involves motion control of special effects, capture of real scenes, and interaction between special effects and reality. It is understandable that scene special effects are not static in the augmented display space, but will be displayed dynamically, such as snowflake special effects will fall in the air, and bubble special effects will break after touching the ground. These dynamic effects can increase the immersion of the wearer brought by the scene special effects and improve the user's chat experience.
[0069] Through step S201 to step S204, the embodiment of the present application provides an immersive conversation method, which is applied to smart glasses. In this method, the smart glasses obtain the emotional data involved in the conversation interaction function and extract emotional features from the emotional data. Then, the emotional response type is generated based on the emotional features, the target scene special effects associated with the emotional response type are determined, and then the target scene special effects are dynamically displayed in the constructed augmented reality space. In this way, it is possible to accurately and real-time capture the emotional characteristics of the wearer, and dynamically display the scene special effects related to the emotions through augmented reality technology, thereby significantly enhancing the user's sense of immersion and the beneficial effects of the interactive experience during the chat process. This method can be applied to education, entertainment, medical care and other fields, bringing emotional value to users in chats and improving users' device usage experience.
[0070] In some embodiments, the above step S204 includes the following steps: Figure 5 Steps S501 to S502 are shown.
[0071] Step S501: Control the target scene special effects to move in the augmented reality space constructed by the smart glasses.
[0072] Step S502: During the movement of the target scene special effect, the real scene is captured, and when the target scene special effect interacts with the real scene, the display effect of the target scene special effect is dynamically adjusted.
[0073] The display effect includes at least one of a dissolving effect, a rupture effect, and an elastic deformation effect.
[0074] The interaction between the target scene special effects and the real scene can be that the target scene special effects contact or collide with the real scene, and the target scene special effects move dynamically. When the two interact, the interactive effect of the target scene special effects can be simulated to enhance the immersiveness of the special effects interaction.
[0075] In this embodiment, visual, auditory and tactile feedback may be combined to further enhance the sense of immersion. For example, when the bubble effect bursts, it may be accompanied by a bursting sound and vibration feedback.
[0076] In this embodiment, the interaction rules may be the default settings of the smart glasses system, or may support user customization. For example, the user may choose to have the snowflake special effect display a "melting" effect when it hits an object, and to have the bubble special effect display a "breaking" effect when it hits the ground.
[0077] This embodiment dynamically adjusts the display effect of special effects, so that users can feel the deep interaction between virtual content and real environment, enhancing their sense of participation and satisfaction. It can also combine multiple senses such as vision, hearing, and touch to further enhance the sense of immersion. In addition, the selection of scene special effects is associated with the user's emotions. By sensing the wearer's emotions and generating special effects associated with the response emotions, the problem of insufficient chat interactivity and emotions is solved.
[0078] In some embodiments, step S502 further includes: Figure 6 Steps S601 to S603 are shown.
[0079] Step S601: Identify an interactive object in a real scene that interacts with a target scene special effect.
[0080] Step S602: Based on the motion speed of the target scene special effect and / or the physical properties of the interactive object, motion simulation is performed on the target scene special effect to obtain a motion simulation result.
[0081] Step S603: dynamically adjust the display effect of the target scene special effects based on the motion simulation result.
[0082] The interactive object refers to an object in the real scene that collides or contacts with the target scene special effect, such as a wall, a table, a human body, etc. In step S601, the position and shape of the interactive object are determined by the SLAM (or simultaneous localization and mapping) algorithm. In step S602, motion simulation refers to simulating the collision or contact process between the special effect and the interactive object according to the movement speed and direction of the target scene special effect and the physical properties of the interactive object (such as hardness, elasticity and friction).
[0083] In step S602, based on the movement speed of the target scene special effects, the target scene special effects are simulated in motion, so when the bubble speed is fast and falls to the ground, a bursting effect will appear, and when the bubble speed is slow and falls to the ground, a rebound effect will appear. This embodiment simulates the movement of the target scene special effects to make it more in line with the laws of physics. It is worth mentioning that the same special effects unit corresponds to multiple target scene special effects in the scene, and the display effect of each target scene special effect can be the same or different. For example, multiple bubbles can all have a bursting effect when falling; or some bubbles burst, some bubbles disappear, and some bubbles rebound.
[0084] In step S602, based on the physical properties of the interactive objects, motion simulation is performed on the target scene special effects. When the bubble contacts the soft interactive object, a rebound effect will appear. When the bubble contacts the hard interactive object, a burst effect will appear. By modeling multiple physical properties of the interactive objects, complex interactive effects can be simulated. Interactive effects can refer to different degrees of the same effect, such as when raindrop special effects fall on surfaces of different materials, different degrees of splashing effects can be generated.
[0085] In some embodiments, the motion speed of the target scene special effect and the physical properties of the interactive object may be combined to perform motion simulation on the target scene special effect.
[0086] Through steps S601 to S603, the target scene special effects displayed in this embodiment are associated with the real environment of the wearer in the augmented reality space. For example, colorful bubbles are generated in the three-dimensional space, and the correct occlusion information is calculated based on the three-dimensional scene information scanned in advance and the posture information of the smart glasses. There will be feedback when the bubbles touch the ground or the wall, which can enhance the immersion of the chat.
[0087] In some embodiments, step S603 includes: Figure 7 Steps S701 to S702 are shown.
[0088] Step S701: Render the target scene special effects based on the motion simulation results to obtain a special effects rendering image.
[0089] Step S702: superimpose the special effect rendering image with the augmented reality space to obtain a display effect of the target scene special effect.
[0090] Specifically, through the motion simulation results, the collision and contact process between the target scene special effects and the interactive objects in different scenes can be reflected, making the special effects display more realistic. When the picture in the wearer's eyes is different, the dynamic display of the target scene special effects will also be adaptively adjusted to achieve adaptation to the environment.
[0091] In addition, when rendering according to the simulation results, the three-dimensional scene information can be used for depth judgment to achieve the correct occlusion relationship, obtain the special effects rendering image, and superimpose it with the image in the current wearer's eyes to achieve the overall special effects presentation.
[0092] In some embodiments, before step S204, the method further includes: Figure 8 Steps S801 to S802 are shown.
[0093] Step S801: Obtain posture information of the wearer in a real scene.
[0094] Step S802: Determine the effective range of the target special effect scene in the augmented reality space based on the posture information.
[0095] Specifically, by acquiring the image information of the real scene captured by the smart glasses, the image information is processed by the SLAM algorithm to obtain the posture information of the wearer in the real scene. Then, the area range where special effects need to be presented is determined based on the posture information, for example, the area range is within a sphere with a radius of 3m centered on the wearer.
[0096] An immersive conversation method provided in an embodiment of the present application is described in detail above. An immersive conversation device is introduced below.
[0097] Please refer to Fig. 9 , Fig. 9 A schematic diagram of the hardware structure of an immersive conversation device provided in an embodiment of the present application is shown. Fig. 9 As shown, the device comprises: The receiving module 901 is used to obtain the emotional data of the smart glasses in the dialogue interaction function.
[0098] The emotion generation module 902 is used to extract emotion features from the emotion data and generate emotion response types based on the emotion features.
[0099] The special effect selection module 903 is used to determine the target scene special effect associated with the emotional response type.
[0100] The display module 904 is used to dynamically display the target scene special effects in the augmented reality space constructed by the smart glasses.
[0101] In some of the embodiments, the emotion generation module 902 is also used to recognize voice information to obtain dialogue text; extract features of the voice information and the dialogue text corresponding to the voice information through a pre-trained dialogue emotion generation network to obtain the wearer's emotional characteristics; and use the dialogue emotion generation network to generate emotional response types based on the emotional characteristics.
[0102] In some of the embodiments, the display module 904 is also used to control the target scene special effects to move in the augmented reality space constructed by the smart glasses; during the movement of the target scene special effects, the real scene is captured, and when the target scene special effects interact with the real scene, the display effect of the target scene special effects is dynamically adjusted.
[0103] In some of the embodiments, the display module 904 is also used to identify the interactive objects in the real scene that interact with the target scene special effects when the target scene special effects interact with the real scene; perform motion simulation on the target scene special effects based on the movement speed of the target scene special effects and / or the physical properties of the interactive objects to obtain motion simulation results; and dynamically adjust the display effect of the target scene special effects based on the motion simulation results.
[0104] In some of the embodiments, the display module 904 is also used to render the target scene special effects based on the motion simulation results to obtain a special effects rendering image; and superimpose the special effects rendering image with the augmented reality space to obtain a display effect of the target scene special effects.
[0105] In some of the embodiments, the display module 904 is further used to obtain posture information of the wearer in the real scene; and determine the effective range of the target special effect scene in the augmented reality space based on the posture information.
[0106] In some of the embodiments, the display module 904 is also used to obtain image information of the real scene captured by the smart glasses; the image information is processed through the SLAM algorithm to obtain the posture information of the wearer in the real scene.
[0107] It should be noted that the above modules can be functional modules or program modules, and can be implemented by software or hardware. For modules implemented by hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.
[0108] The present application also provides a smart glasses, such as Fig.10 As shown, the smart glasses include a processor 1001 , a memory 1002 , a shooting module 1003 , a display module 1004 , an interaction module 1005 and a wireless communication module 1006 .
[0109] The memory 1002 can be used to store computer programs, such as software programs and modules of application software. The processor 1001 executes various functional applications and data processing by running the computer programs stored in the memory 1002, such as implementing the method provided in the embodiments of the present application.
[0110] The shooting module 1003 is mainly responsible for shooting the image information of the real environment and obtaining the user's face and gesture information, etc. The shooting module can be a camera.
[0111] The display module 1004 is used to display virtual images and augmented reality contents, and includes a micro-projection system and optical elements.
[0112] The interaction module 1005 includes a microphone, an eye tracking sensor, etc.
[0113] The wireless communication module 1006 includes: 4G / 5G full network access, WiFi, Bluetooth, etc.
[0114] It can be understood by those skilled in the art that Fig.10The structure shown is for illustration only and does not limit the structure of the smart glasses. Fig.10 More or fewer components as shown, or with Fig.10 Different configurations are shown.
[0115] The following is a brief introduction to the application scenarios of smart glasses.
[0116] In the voice interaction scenario of smart glasses, the wearer can chat with the digital person. During the chat, the smart glasses receive the user's voice data and output the chat text through the voice recognition Transformer network. The input of the dialogue emotion generation Transformer network is the chat text and voice data, and the output is the emotion label. The special effect superposition is roughly divided into four steps. Taking the bubble special effect as an example, first determine the area range where the special effect needs to be presented according to the current posture. Then generate at least one basic special effect unit (bubble, raindrop, etc.) that matches the emotion label, and then assign motion parameters to each special effect unit, and simulate the movement according to the three-dimensional scene (when hitting the wall, it bounces off slowly, bursts quickly, etc.), and finally render according to the state after the simulated movement to obtain the special effect rendering image, which is superimposed with the current user screen to achieve the overall special effect presentation.
[0117] In addition, in combination with the method provided in the above embodiments, a storage medium may be provided in this embodiment to implement the method. The storage medium stores a computer program; when the computer program is executed by a processor, any immersive conversation method in the above embodiments is implemented.
[0118] It should be understood that the specific embodiments described herein are only used to explain the application, rather than to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of this application.
[0119] Obviously, the drawings are only some examples or embodiments of the present application. For ordinary technicians in the field, the present application can also be applied to other similar situations based on these drawings without creative work. In addition, it is understandable that although the work done in this development process may be complicated and lengthy, for ordinary technicians in the field, certain changes in design, manufacturing or production based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient content disclosed in this application.
[0120] The term "embodiment" in this application refers to a specific feature, structure or characteristic described in conjunction with the embodiment that can be included in at least one embodiment of the present application. The appearance of this phrase in various locations in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is clearly or implicitly understood by those of ordinary skill in the art that the embodiments described in this application can be combined with other embodiments without conflict.
[0121] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of patent protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the attached claims.
Claims
1. An immersive dialogue method, characterized in that: Applied to smart glasses, the method comprises: Acquiring emotional data involved in a conversation interaction function of the smart glasses; Extracting emotional features from the emotional data, and generating emotional response types based on the emotional features; determining a target scene effect associated with the type of emotional response; The target scene special effects are dynamically displayed in the augmented reality space constructed by the smart glasses.
2. The immersive conversation method according to claim 1, characterized in that: The dynamically displaying the target scene special effects in the augmented reality space constructed by the smart glasses includes: Controlling the target scene special effects to move in the augmented reality space constructed by the smart glasses; During the movement of the target scene special effect, the real scene is captured, and when the target scene special effect interacts with the real scene, the display effect of the target scene special effect is dynamically adjusted.
3. The immersive conversation method according to claim 2, characterized in that: When the target scene special effect interacts with the real scene, dynamically adjusting the display effect of the target scene special effect includes: When the target scene special effect interacts with the real scene, identifying an interactive object in the real scene that interacts with the target scene special effect; Based on the movement speed of the target scene special effect and / or the physical properties of the interactive object, performing motion simulation on the target scene special effect to obtain a motion simulation result; Based on the motion simulation result, the display effect of the target scene special effect is dynamically adjusted.
4. The immersive conversation method according to claim 3, characterized in that: Based on the motion simulation result, dynamically adjusting the display effect of the target scene special effect includes: Based on the motion simulation result, rendering the target scene special effects to obtain a special effects rendering image; The special effect rendering image is superimposed on the augmented reality space to obtain a display effect of the target scene special effect.
5. The immersive conversation method according to any one of claims 1 to 4, characterized in that: Before dynamically displaying the target scene special effect in the augmented reality space constructed by the smart glasses, the method further includes: Acquire posture information of a wearer of the smart glasses in a real scene; The effective range of the target special effect scene in the augmented reality space is determined based on the posture information.
6. The immersive conversation method according to claim 5, characterized in that: The obtaining of posture information of the wearer in the real scene includes: Acquire image information of a real scene captured by the smart glasses; The image information is processed by a SLAM algorithm to obtain posture information of the wearer in the real scene.
7. The immersive conversation method according to any one of claims 1 to 4, characterized in that: The target scene special effects include at least one of snowflakes, bubbles, and raindrops, and the display effects include at least one of dissolving effects, rupture effects, and elastic deformation effects.
8. The immersive conversation method according to claim 1, characterized in that: The emotion data includes voice data of the wearer of the smart glasses, and extracting emotion features from the emotion data and generating an emotion response type based on the emotion features includes: Recognizing the voice data to obtain a conversation text; Extracting features of the voice data and the dialogue text corresponding to the voice data through a pre-trained dialogue emotion generation network to obtain the emotional features of the wearer; The dialogue emotion generation network is used to generate an emotion response type based on the emotion feature.
9. A smart glasses, characterized in that: It includes at least one processor, a memory, a camera module, a microphone module and a display module. The at least one processor is coupled to the memory and is used to read and execute instructions in the memory to perform the immersive conversation method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the immersive dialogue method according to any one of claims 1 to 8 when running.
Citation Information
Patent Citations
Interaction control method and system
CN108509043A
Wearable augmented reality remote video system and a video call method
CN109803109A
Method, program product and system for visualizing an emotion of a vehicle user
CN113283379A
Scene control method and device, equipment and storage medium
CN115543078A
AR-based emotion data processing method and apparatus, and electronic device
CN118587757A
Cited By
Emotion analysis method and device, wearable equipment and storage medium
CN121167153A