Machine model interaction method and device, electronic equipment and storage medium
By collecting and processing multimodal information, extracting image and time series features, determining interaction modalities, and generating emotional responses, the problem of weak adaptability and insufficient emotion recognition of robots when the environment changes is solved, and more efficient user interaction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FIBOCOM WIRELESS
- Filing Date
- 2025-12-10
- Publication Date
- 2026-05-05
AI Technical Summary
Existing robots have weak adaptability when the environment changes, their speech recognition and visual perception are easily interfered with, and they lack dynamic switching of interaction modes, resulting in low interaction reliability. They also have difficulty capturing complex emotions or inferring reasons from context, and their companionship effect is poor.
Multimodal information is collected, including environmental objects, light intensity, user facial expressions and body movements. Image and time series features are extracted through convolutional neural networks and long short-term memory networks to determine the interaction modality. Responses are generated by combining multimodal emotional features and historical interactive voice.
It improves the interaction experience between machine models and users, enabling them to accurately identify and respond to user needs in different environments and emotional states, thus enhancing the naturalness and intelligence of the interaction.
Smart Images

Figure CN121979969A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more particularly to a method and apparatus for interacting with a machine model, an electronic device, and a storage medium. Background Technology
[0002] Currently, companion robot technology focuses on multimodal perception and interaction, aiming to improve the naturalness and intelligence of communication with users. However, current robots have weak adaptability to environmental changes (such as noise, lighting, and cluttered scenes), their speech recognition and visual perception are easily interfered with, and they lack the ability to dynamically switch interaction modalities, resulting in low interaction reliability. Moreover, current robot emotion recognition is limited to basic emotions, making it difficult to capture complex emotions or infer reasons from context, resulting in superficial responses lacking empathy. Therefore, existing technologies for robot companionship are inadequate and fail to meet user needs.
[0003] There is currently no effective solution to the aforementioned technical problems in the existing technology. Summary of the Invention
[0004] This application provides a method and apparatus for interacting with a machine model, an electronic device, and a storage medium to solve the problem of poor companionship effect of existing companion robots.
[0005] In a first aspect, this application provides a method for interacting with a machine model, comprising: collecting multimodal information associated with a user, wherein the multimodal information represents information associated with the user's environment and attributes; normalizing all information in the multimodal information to obtain corresponding feature vectors; extracting image spatial features and time series features from the feature vectors, and determining the interaction modality between the user and the machine model based on the image spatial features and the time series features, wherein the image spatial features represent illumination intensity and scene complexity, and the time series features represent noise level; extracting multimodal emotional features from the feature vectors, and combining the multimodal emotional features with the user's historical interactive speech with the machine model to determine the machine model's interactive response to the user.
[0006] Secondly, this application provides an interaction device with a machine model, comprising: a data acquisition module for acquiring multimodal information associated with a user, wherein the multimodal information represents information associated with the user's environment and attributes; a first processing module for normalizing all information in the multimodal information to obtain corresponding feature vectors; a second processing module for extracting image spatial features and time series features from the feature vectors, and determining the interaction modality between the user and the machine model based on the image spatial features and the time series features, wherein the image spatial features represent illumination intensity and scene complexity, and the time series features represent noise level; and a third processing module for extracting multimodal emotional features from the feature vectors, and determining the machine model's interactive response to the user by combining the multimodal emotional features and the user's historical interactive voice with the machine model.
[0007] Thirdly, this application provides an electronic device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the processor is configured to execute the interaction method with a machine model described in the first aspect of this application.
[0008] Fourthly, this application also provides a computer storage medium storing computer-executable instructions for performing the interaction method with the machine model described in the first aspect of this application.
[0009] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application first collects multimodal information associated with the user. This multimodal information represents information related to the user's environment and attributes. Then, all information in the multimodal information is normalized to obtain corresponding feature vectors. Next, image spatial features and time-series features are extracted from the feature vectors. Based on these features, the interaction modality between the user and the machine model is determined. Based on this, multimodal emotional features are extracted from the feature vectors, and combined with the multimodal emotional features and the user's historical voice interactions with the machine model, the machine model's interactive response to the user is determined. It can be seen that in this application, before the machine model responds to the user, the current state of the user's environment is comprehensively considered to identify the optimal interaction modality, ensuring that the machine model's interactive response is accurately received by the user. Furthermore, this application comprehensively considers multiple current features, such as user facial expressions, body language, and voice tone, accurately identifying the user's emotions. Therefore, it can accurately respond to the user based on the identified emotions, improving the interactive experience between the machine model and the user. Attached Figure Description The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0012] Figure 1 A flowchart illustrating a method for interacting with a machine model, as provided in an embodiment of this application; Figure 2 A schematic diagram of the structure of an interaction device with a machine model provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0014] The following disclosure provides numerous different embodiments or examples for implementing various structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of the invention. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0015] To address the issue of poor companionship effectiveness in existing companion robots, this application provides a method for interacting with a machine model, such as... Figure 1 As shown, the steps of this method include: Step 101: Collect multimodal information associated with the user, wherein the multimodal information represents information related to the user's environment and personal attributes; In specific examples, information associated with the user's environment and personal attributes may include information about environmental objects, light intensity, facial expressions, body movements, and touch information. It is evident that this application collects a considerable amount of multimodal information related to the user, which allows for more accurate determination of the robot's interactive responses—that is, combining all available information to accurately address the user's needs.
[0016] It should be noted that each modality in the multimodal information has a corresponding form. For example, the visual modality information corresponding to a user's facial expressions can be stored in the form of image data, while the auditory modality information corresponding to ambient sound and speech can be stored in the form of audio data. Furthermore, the machine model in this application embodiment can be a model built into a physical robot or a model built into other intelligent devices.
[0017] Step 102: Normalize all information in the multimodal information to obtain the corresponding feature vector; In this embodiment of the application, since there are multiple modal information in the multimodal information, each modal information needs to be normalized separately to ensure that the multimodal information is processed under a unified scale. Subsequently, when determining the robot's response, the multimodal information processed under a unified scale can obtain a more accurate response.
[0018] Step 103: Extract image spatial features and time series features from the feature vector, and determine the interaction modality between the user and the machine model based on the image spatial features and time series features. The image spatial features represent the illumination intensity and scene complexity, and the time series features represent the noise level. In specific implementations of this application, spatial features of images can be extracted using a Convolutional Neural Network (CNN), while time-series features can be extracted using a Long Short-Term Memory (LSTM) network. Based on this, when extracting features from feature vectors, data corresponding to various modalities can be extracted. For example, spatial features are extracted from image data, such as user expressions, body movements, environmental object boundaries, texture distribution, etc.
[0019] Furthermore, in this embodiment, determining the interaction mode between the user and the machine model based on image spatial features and time series features means determining the noise level and light intensity of the current environment in which the user and the machine model are located, based on image spatial features and time series features, in order to determine the interaction mode between the machine model and the user. This interaction mode includes voice, text, vibration prompts, etc. If the current light is insufficient, the interaction mode may be a voice prompt; if the current environment is noisy, the interaction mode may be a text display or vibration prompt. In other words, this application can select the most suitable interaction mode in real time according to the environment to ensure the user's interaction experience with the machine model.
[0020] Step 104: Extract multimodal sentiment features from the feature vector, and combine the multimodal sentiment features with the historical voice interaction between the user and the machine model to determine the machine model's interactive response to the user.
[0021] In this embodiment of the application, the multimodal emotional features extracted from the feature vector include facial expressions, body movements, and voice tone. Combining these multimodal emotional features can accurately identify the current user's emotions. The machine model can then respond to the user's emotions based on the identified emotions. For example, if the current user is sad, the machine model can play the user's favorite songs or books based on the user's actual situation. If the current user is happy, the machine model can control the device to play music and dance along.
[0022] Through steps 101 to 104 above, multimodal information associated with the user is first collected. This multimodal information represents information related to the user's environment and personal attributes. Then, all information in the multimodal information is normalized to obtain corresponding feature vectors. Next, image spatial features and time series features are extracted from the feature vectors, and the interaction modality between the user and the machine model is determined based on these features. Based on this, multimodal emotional features are extracted from the feature vectors, and the machine model's interactive response to the user is determined by combining the multimodal emotional features with the historical voice interaction between the user and the machine model. It is evident that in this application, before the machine model responds to the user, the current state of the user's environment is comprehensively considered to identify the optimal interaction modality, ensuring that the machine model's interactive response is accurately received by the user. Furthermore, this application also comprehensively considers multiple current features, such as user facial expressions, body language, and voice tone, to accurately identify the user's emotions. Therefore, it can accurately respond to the user based on the identified emotions, improving the interactive experience between the machine model and the user.
[0023] In an optional embodiment of this application, the method for collecting multimodal information associated with the user involved in step 101 above may further include: Step 11: Collect the user's facial expressions, body movements and two-dimensional visual features of objects in the surrounding environment through the camera to obtain corresponding image data, and generate corresponding visual modal information from the image data; To this end, the camera is used to capture the user's facial expressions (such as micro-expressions of the eyebrows, eyes, and mouth) and body movements (such as gestures and postures), generating two-dimensional image data. Simultaneously, the camera can also assist in capturing visual information of environmental objects (such as furniture and obstacles). The user's facial expressions can include 68 facial key points (such as eye opening and closing, and the curvature of the corners of the mouth) for emotion analysis. Body movements can include the 2D contours of the user's upper limbs, lower limbs, or full-body postures (such as gesture trajectories and body tilt). Environmental objects can include the 2D visual features of static or dynamic objects (such as color, shape, and edge detection).
[0024] In this embodiment, the image resolution is moderate (e.g., 720p or higher) and contains time-series frames, which facilitates CNN extraction of spatial features.
[0025] Step 12: Generate corresponding depth visual modal information by using the skeletal coordinates of the user's limb movements, facial depth contours, and point cloud representations of environmental objects collected by the depth sensor. To address this, depth sensors are used to detect the 3D position and distance of objects in the environment (such as object depth and spatial layout) and assist in capturing depth information of user limb movements (avoiding occlusion issues in 2D images). Furthermore, depth sensors support 3D modeling of user facial expressions (such as facial contour depth). The point cloud representation of user limb movements can be 3D skeletal points (such as joint position coordinates, x / y / z axes). The point cloud representation of environmental objects can include the point cloud representation of objects in space (such as depth values and density distribution of each point), used for scene complexity assessment (e_3 metric). The point cloud representation of the user's depth contour is used to assist in 3D facial meshing (such as depth contours), enhancing expression robustness.
[0026] In this embodiment of the application, the point cloud contains tens of thousands of points (e.g., 10^4~10^5 points / frame), has timestamps, and supports LSTM time series analysis.
[0027] It should be noted that image data can provide texture / color (2D), while point clouds can provide geometry / depth (3D). Both can be used to solve occlusion or lighting problems (such as facial expression recognition in low light).
[0028] Step 13: Generate corresponding auditory modal information from the user's voice and ambient sound collected by the microphone array; In response, the user's voice (e.g., dialogue content, tone of voice) and ambient sounds (e.g., background noise, echo) collected by the microphone array are beamformed and denoised to form a clean audio signal.
[0029] Step 14: Generate corresponding tactile modal information from the user touch information detected by the tactile sensor; In response, haptic sensors detect user touches (e.g., force, location, duration) and use them for interactive feedback (such as handshakes or touch responses). Step 15: Generate corresponding environmental physical mode information from the light intensity monitored by the ambient light sensor.
[0030] In this regard, ambient light sensors monitor light intensity (such as lux values) to assess the impact of lighting conditions on visual perception.
[0031] In the embodiments of this application, the aforementioned multimodal information can be collected in time series form (e.g., 30fps frame rate). After preprocessing, the total data stream can be extracted into feature vectors (e.g., facial key points, MFCC speech features), with a total volume of several GB / hour, emphasizing real-time performance and low latency.
[0032] In an optional embodiment of this application, the method of normalizing all information in the multimodal information to obtain the corresponding feature vector in step 102 above may further include: Step 21: Determine the mean value of each modality in the multimodal information; Step 22: Determine the difference between each modal information and its corresponding mean, and determine the ratio of the difference to the standard deviation as the feature vector corresponding to that modal information.
[0033] In a specific embodiment, steps 21 and 22 can be implemented using the following formula:
[0034] in, The original feature vector, The mean, Standard deviation The sum of normalized eigenvectors.
[0035] In an optional embodiment of this application, the method for determining the user-machine model interaction modality based on image spatial features and time series features involved in step 103 above may further include: Step 31: Perform pooling processing on the image spatial features and time series features respectively; Step 32: The spatial features and temporal series features of the pooled image are fused using an attention mechanism to obtain an intermediate environment description vector; Step 33: Map the intermediate environment description vector to three components using a regression function, where the three components are illumination intensity, scene complexity, and noise level, respectively. Step 34: Compare the sum of the products of the three components and their corresponding preset weight values with the interaction modality threshold. Step 35: Determine the interaction modality based on the comparison results, where the interaction modality represents the interaction form between the machine model and the user.
[0036] For steps 31 to 35 above, in a specific example, the environment analyzer uses CNN to extract image spatial features, LSTM to analyze the time series features of speech and ambient sound, and generates an environment state vector. ,in, Indicates the noise level (dB). Indicates light intensity (lux). This represents scene complexity (object density). Environmental state assessment formula:
[0037] in, The weights are (optimized through reinforcement learning, with initial values of 0.4, 0.3, 0.3). An environmental adaptation score is used to determine the threshold for switching interaction modalities. The interaction strategy generator is based on... Select the interaction modality (such as voice, text, light), for example when When there is noise, switch to text display or voice prompts.
[0038] In this regard, the extracted image spatial features are directly derived from the image data generated by the multimodal perception module. Specifically, the multimodal perception module generates image data (RGB or grayscale images) through a camera, including user facial expressions, body movements, and 2D visual information of environmental objects (such as color, shape, and edges). This raw image data is then input into a CNN (convolutional neural network, such as a ResNet variant) to extract spatial features, such as object boundaries, texture distribution, and illumination gradients. These features capture the static spatial structure of the image and are used to assess environmental complexity. ) and lighting conditions ( Instead of user-specific expressions (which are mainly used in the emotion module), this ensures a continuous data flow from perception to decision-making, avoiding additional data collection overhead.
[0039] Specifically, spatial features (CNN output): high-dimensional vectors extracted from the image (e.g., dimension 512, containing object detection probabilities, edge density, and illumination histogram). These features primarily contribute to... (Light intensity, lux) and (Scene complexity, object density). Process: CNN convolutional layers capture local patterns, fully connected layers aggregate them into a global description; then, through fully connected layers or MLP, they are mapped to low-dimensional metrics (such as object count / density).
[0040] Time-series features (LSTM output): Extracted from time series of speech and ambient sound (e.g., sequence length 100, including MFCC or spectral variations). LSTM handles temporal dependencies and primarily contributes to... (Noise level, dB). Process: LSTM gated units capture long-term dependencies (such as noise persistence) and output hidden state vectors; noise peaks are quantized via softmax or regression heads.
[0041] Based on this, steps 31 to 35 above can specifically involve first reducing the dimensionality of the spatial features and time-series features through pooling (e.g., average / max pooling). Then, based on an attention mechanism, an intermediate environment description vector is formed, and the fused features are further mapped to the three components of E through a regression function (e.g., a linear layer). Calculated from the energy spectrum of the time series; Average of illumination channels based on spatial characteristics; The process involves segmenting and counting objects based on spatial features, and finally assembling the vectors to determine the above-mentioned vector input formula. .
[0042] In a specific example, this could be: Assume the robot is placed in a living room scene: the image shows a sofa, television, and window (spatial features: CNN detects 3 objects, density = 0.4 objects / pixel, illumination gradient height = 800 lux); the audio captures the user's speech + television noise (time series features: LSTM analysis shows an average noise peak of 65 dB, with noise accounting for 30% of the sequence). Generation process: Spatial features → =800 (lux, high light intensity) =0.4 (medium density); Time series characteristics → =65 (dB, medium noise). Therefore, E=[65, 800, 0.4]. Based on this, =0.4·65 + 0.3·800 + 0.3·0.4 ≈ 26 +240 + 0.12 = 266.12 (after normalization, it is greater than the 0.7 threshold), triggering modal switching (such as voice + text display, to avoid noise interference).
[0043] In an optional embodiment of this application, the method for extracting multimodal sentiment features from the feature vector involved in step 104 above may further include: Step 41: Extract target vectors associated with the user's facial expressions, voice tone, and body movements from the feature vectors, and identify all target vectors as multimodal emotion features.
[0044] In an optional embodiment of this application, the method of determining the machine model's interactive response to the user by combining multimodal emotional features and the historical voice interaction between the user and the machine model in step 104 may further include: Step 51: Determine the value of each target vector in the multimodal sentiment features; Step 52: Compare the sum of the values multiplied by their corresponding weights with a preset mapping table to determine the user's emotional category. The preset mapping table represents the mapping relationship between emotional categories and values. Step 53: Combine the emotion category with historical voice interactions to determine the interactive response.
[0045] In this specific example, facial expressions (68 key points, extracted by ResNet-50), speech intonation (MFCC features, processed by LSTM), and body movements (skeleton points, extracted by OpenPose) are fused together to calculate the emotional state using the following formula:
[0046] in, These are the emotional features of each modality. The weights are (optimized through transfer learning, with initial values of 0.5, 0.3, and 0.2). The system synthesizes an emotion vector and maps it to an emotion category (such as sadness or joy). It then combines this vector with context (dialogue history, stored in a local database) to infer the emotional cause and generate a personalized response.
[0047] Furthermore, the sentiment feature extraction process involves extracting the following data after inputting multimodal data: Facial expressions ResNet-50 extracts features from camera images and detects low-set eyes (drooping eyes 0.6, downturned corners of the mouth 0.4), with a feature vector ≈ [-0.7, -0.5] (indicating a tendency towards depression).
[0048] intonation LSTM analyzes the user's utterance "Eating alone again today" from MFCC, with a slow speaking speed (0.8 times the normal) and a low tone (a decrease of 15%), and the feature vector is approximately [-0.6, -0.8] (dejected tone).
[0049] Body movements OpenPose extracts features from point cloud, showing the user's arm hanging down (opening angle 0.3) and slightly shaking their head (frequency 2Hz), with a feature vector ≈ [-0.4, -0.6] (passive pose).
[0050] Combining the above formulas: .
[0051] Furthermore, a classifier (such as an SVM or softmax layer) is used to... Projected onto the emotional space. Threshold setting: negative values > 0.5 are mapped to "sadness" (probability 0.85), "neutral" (0.1), or "fatigue" (0.05). Due to Mapped to the sadness category (85% confidence), because the vector falls into the negative cluster.
[0052] Further, the local database is queried, such as the conversation history: ID=001, user: "The children are busy and won't be back." Timestamp: yesterday; ID=002, user: "The food is cold." Timestamp: this morning. Then, association analysis is performed: an NLP model (e.g., BERT embedding) is used to match the current "sadness" with historical keywords ("alone," "children"), inferring the cause: loneliness stems from family separation (similarity > 0.7). This is not a generalized inference (e.g., "bad weather"), but rather a personalized one (the theme "going home" was repeated 3 times in the history). The emotional reason is "the user may feel lonely because their children are neglecting them" (confidence based on historical frequency).
[0053] Based on this, we obtain: Emotion category "sadness" + Reason "loneliness" + Context (elderly preferences: slow speech, large font, obtained from the personalization module). Generate an interactive response: The robot outputs in slow speech + on-screen text: "I understand you're eating alone again today, and you must feel empty inside. I remember you said your children are busy, but they love you the most. Come on, let's listen to an old song together and reminisce, okay? (Plays user history preferences: Teresa Teng songs)." In this embodiment of the application, the machine model can be further optimized and adjusted. Specifically, the machine model can be continuously optimized and adjusted based on the user's satisfaction with the interaction response and the interaction duration of the interaction response.
[0054] To address this, the learning engine can continuously record user interaction data, and reinforcement learning can be used to optimize the user model. The reward function is:
[0055] in, For user satisfaction (collected through feedback ratings). For interaction duration, The weights are set to 0.6 and 0.4 respectively. The interaction optimizer adjusts the interaction strategy based on the user model, such as providing gamified content for children and a loud, slow-speech interface for the elderly.
[0056] In a specific example: an 8-year-old child (user) plays with the robot in their bedroom in the afternoon. The machine model is initially based on the first interaction (age inferred through voice / facial recognition, preferences from historical data: likes animal stories, but has short attention span). Current interaction: the child says, "Tell me a story!" The robot captures the emotion (joy, (positive value) and environment (suitable) <0.3).
[0057] Reward calculation: The interaction ends, and the child participates with laughter and play. =6min, =4.5 / 5), R = 0.6·4.5+ 0.4·6 ≈ 2.7 + 2.4 = 5.1 (positive reward, above the threshold of 4).
[0058] Before interaction: No optimization, potentially outputting a boring story, causing children to walk away midway. <2min, R<2).
[0059] After interaction: Children become immersed (satisfaction score +1), model iterates (preference updated to "animals").
[0060] Corresponding to the above Figure 1 This application provides an interaction device for a machine model, such as... Figure 2 As shown, the device includes: The acquisition module 202 is used to acquire multimodal information associated with the user, wherein the multimodal information represents information related to the user's environment and personal attributes; The first processing module 204 is used to normalize all information in the multimodal information to obtain the corresponding feature vector; The second processing module 206 is used to extract image spatial features and time series features from the feature vector, and determine the interaction mode between the user and the machine model based on the image spatial features and time series features, wherein the image spatial features represent the illumination intensity and scene complexity, and the time series features represent the noise level. The third processing module 208 is used to extract multimodal emotion features from the feature vector and combine the multimodal emotion features with the historical voice interaction between the user and the machine model to determine the machine model's interactive response to the user.
[0061] The apparatus of this application first collects multimodal information associated with the user. This multimodal information represents information related to the user's environment and attributes. Then, all information in the multimodal information is normalized to obtain corresponding feature vectors. Next, image spatial features and time-series features are extracted from the feature vectors. Based on these features, the interaction modality between the user and the machine model is determined. Based on this, multimodal emotional features are extracted from the feature vectors. Combined with the multimodal emotional features and the user's historical voice interactions with the machine model, the machine model's interactive response to the user is determined. It is evident that in this application, before the machine model responds to the user, the current state of the user's environment is comprehensively considered to identify the optimal interaction modality, ensuring that the machine model's subsequent interactive response is accurately received by the user. Furthermore, this application comprehensively considers multiple current features, such as user facial expressions, body language, and voice tone, to accurately identify the user's emotions. Therefore, it can accurately respond to the user based on the identified emotions, improving the interactive experience between the machine model and the user.
[0062] In an optional embodiment of this application, the acquisition module may further include: a first processing unit, configured to acquire corresponding image data by acquiring the user's facial expressions, body movements, and two-dimensional visual features of objects in the surrounding environment through a camera, and generate corresponding visual modal information from the image data; a second processing unit, configured to generate corresponding depth visual modal information by acquiring the skeletal point coordinates of the user's body movements, facial depth contours, and point cloud representations of objects in the environment through a depth sensor; a third processing unit, configured to generate corresponding auditory modal information from the user's voice and ambient sounds acquired by a microphone array; a fourth processing unit, configured to generate corresponding tactile modal information from the user's touch information detected by a tactile sensor; and a fifth processing unit, configured to generate corresponding environmental physical modal information from the light intensity monitored by an ambient light sensor.
[0063] In an optional embodiment of this application, the first processing module in this application embodiment may further include: a first determining unit, used to determine the mean of each modality information in the multimodal information; and a sixth processing unit, used to determine the difference between each modality information and the corresponding mean, and to determine the ratio of the difference to the standard deviation as the feature vector corresponding to the modality information.
[0064] In an optional embodiment of this application, the second processing module may further include: a seventh processing unit, configured to perform pooling processing on the image spatial features and time series features respectively; an eighth processing unit, configured to fuse the pooled image spatial features and time series features through an attention mechanism to obtain an intermediate environment description vector; a ninth processing unit, configured to map the intermediate environment description vector to three components through a regression function, wherein the three components are illumination intensity, scene complexity, and noise level respectively; a tenth processing unit, configured to compare the sum of the three components multiplied by their corresponding preset weight values with an interaction modality threshold; and a second determining unit, configured to determine the interaction modality based on the comparison result, wherein the interaction modality characterizes the interaction form between the machine model and the user.
[0065] In an optional embodiment of this application, the third processing module may further include: an eleventh processing unit, used to extract target vectors associated with the user's facial expressions, voice tone and body movements from the feature vectors, and to determine all target vectors as multimodal emotion features.
[0066] In an optional embodiment of this application, the third processing module in this application includes: a third determining unit, used to determine the value of each target vector in the multimodal emotional features; a twelfth processing unit, used to compare the sum of the values multiplied by the corresponding weights with a preset mapping table to determine the user's emotional category, wherein the preset mapping table represents the mapping relationship between the emotional category and the values; and a fourth determining unit, used to combine the emotional category with historical interactive voice to determine the interactive response.
[0067] In an optional embodiment of this application, the apparatus further includes an adjustment module, configured to continuously optimize and adjust the machine model based on user satisfaction with the interaction response and the interaction duration of the interaction response.
[0068] like Figure 3 As shown in the figure, this application provides an electronic device, including a processor 311, a communication interface 312, a memory 313, and a communication bus 314, wherein the processor 311, the communication interface 312, and the memory 313 communicate with each other through the communication bus 314. Memory 313 is used to store computer programs; In one embodiment of this application, when the processor 311 executes the program stored in the memory 313, it implements the interaction method with the machine model provided in any of the foregoing method embodiments, and its function is similar, so it will not be described again here.
[0069] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the interaction method with the machine model as provided in any of the foregoing method embodiments.
[0070] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0071] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0072] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.
[0073] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for interacting with a machine model, characterized in that, include: Collect multimodal information associated with the user, wherein the multimodal information represents information related to the user's environment and personal attributes; All information in the multimodal information is normalized to obtain the corresponding feature vector; Image spatial features and time series features are extracted from the feature vector, and the interaction mode between the user and the machine model is determined based on the image spatial features and the time series features, wherein the image spatial features characterize illumination intensity and scene complexity, and the time series features characterize noise level; Multimodal sentiment features are extracted from the feature vector, and the interaction response of the machine model to the user is determined by combining the multimodal sentiment features with the historical voice interaction between the user and the machine model.
2. The method according to claim 1, characterized in that, Collect multimodal information associated with users, including: The camera captures the user's facial expressions, body movements, and the two-dimensional visual features of objects in the surrounding environment to obtain corresponding image data, and then generates corresponding visual modal information from the image data. The corresponding depth visual modal information is generated by representing the skeletal coordinates of the user's limb movements, facial depth contours, and point cloud representations of environmental objects collected by the depth sensor. The microphone array collects user speech and ambient sound to generate corresponding auditory modal information; The tactile sensor detects user touch information and generates corresponding tactile modal information; The ambient light sensor monitors the light intensity to generate corresponding environmental physical mode information.
3. The method according to claim 2, characterized in that, Normalization is performed on all information in the multimodal information to obtain the corresponding feature vector, including: Determine the mean value of each modality in the multimodal information; The difference between each modal information and its corresponding mean is determined, and the ratio of the difference to the standard deviation is determined as the feature vector corresponding to that modal information.
4. The method according to claim 1, characterized in that, Determining the interaction modality between the user and the machine model based on the image spatial features and the time series features includes: The image spatial features and the time series features are respectively subjected to pooling processing; The spatial and temporal features of the pooled image are fused using an attention mechanism to obtain an intermediate environment description vector; The intermediate environment description vector is mapped to three components by a regression function, wherein the three components are light intensity, scene complexity and noise level, respectively. The sum of the products of the three components and their corresponding preset weight values is compared with the interaction modality threshold. The interaction modality is determined based on the comparison results, wherein the interaction modality characterizes the interaction form between the machine model and the user.
5. The method according to claim 1, characterized in that, Extracting multimodal sentiment features from the feature vector includes: Extract target vectors associated with the user's facial expressions, voice tone, and body movements from the feature vectors, and determine all the target vectors as the multimodal emotion features.
6. The method according to claim 5, characterized in that, Determining the machine model's interactive response to the user by combining the multimodal emotional features and the user's historical voice interactions with the machine model includes: Determine the value of each of the target vectors in the multimodal emotion features; The sum of the product of the value and the corresponding weight is compared with a preset mapping table to determine the user's emotion category, wherein the preset mapping table represents the mapping relationship between emotion category and value; The interaction response is determined by combining the emotion category with the historical voice interactions.
7. The method according to claim 1, characterized in that, The method further includes: The machine model is continuously optimized and adjusted based on the user's satisfaction with the interaction response and the interaction duration of the interaction response.
8. An interaction device for a machine model, characterized in that, include: The acquisition module is used to acquire multimodal information associated with the user, wherein the multimodal information represents information related to the user's environment and personal attributes; The first processing module is used to normalize all information in the multimodal information to obtain the corresponding feature vector; The second processing module is used to extract image spatial features and time series features from the feature vector, and determine the interaction mode between the user and the machine model based on the image spatial features and the time series features, wherein the image spatial features characterize illumination intensity and scene complexity, and the time series features characterize noise level; The third processing module is used to extract multimodal emotion features from the feature vector, and combine the multimodal emotion features with the historical voice interaction between the user and the machine model to determine the interaction response of the machine model to the user.
9. An electronic device, characterized in that, include: The processor, communication interface, memory, and communication bus are connected, with the processor, communication interface, and memory communicating with each other via the communication bus. The memory is used to store computer programs; the processor is used to execute the computer programs to implement the interaction method with the machine model as described in any one of claims 1-7.
10. A storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method of interacting with the machine model as described in any one of claims 1-7.