Multi-modal man-machine interaction system based on hybrid perception

By using multimodal data fusion and dynamic prioritization mechanisms, a hybrid sensing system based on cameras, microphones, and infrared sensors solves the problems of noise interference and multi-user identification in human-computer interaction systems under complex environments, and achieves an efficient and continuous interaction process.

CN120877773AInactive Publication Date: 2025-10-31SHENZHEN LANZHENG TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510860207.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, human-computer interaction systems are susceptible to noise interference in complex environments, leading to a decrease in recognition accuracy. Furthermore, in multi-user scenarios, there is a lack of effective target object filtering mechanisms, resulting in chaotic interaction processes and poor feedback targeting.

Method used

A multimodal human-computer interaction system based on hybrid perception is adopted. Video and audio streams are acquired simultaneously through cameras and microphones. The distance to the user is calculated by combining infrared sensors, an interaction priority scoring system is constructed, the interaction intentions of near-field users are identified and processed first, and the system automatically switches to other voice sources when the target object has no voice signal.

Benefits of technology

It significantly improves the environmental adaptability and response efficiency of interactive terminals, avoids mislabeling and interaction interruption, and enhances the continuity and intelligence of interaction, making it particularly suitable for multi-person collaboration or consultation scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877773A_ABST
    Figure CN120877773A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode man-machine interaction system based on hybrid perception, and relates to the technical field of man-machine interaction. According to the invention, through multi-modal data fusion and a dynamic priority mechanism, the environmental adaptability and response efficiency of the interactive terminal are significantly improved; the data acquisition module synchronously acquires a video stream and an audio stream by using a camera and a microphone, and calculates a user distance in combination with an infrared sensor, so that a basic object can be accurately identified in a complex scene, and false marking is avoided; the interaction analysis module constructs a priority scoring system based on distance characteristics and voice energy, so that the interaction priority of near-field users is ensured, active interaction intentions can be captured through the voice energy, and the problem of interaction conflicts in a multi-user scene is solved; and when the target object has no voice signal within the second set time, the system is automatically switched to other sounding basic objects, so that interaction interruption is avoided, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human-computer interaction technology, specifically a multimodal human-computer interaction system based on hybrid perception. Background Technology

[0002] In existing technologies, human-computer interaction systems typically rely on a single modality (such as voice signals) to recognize user intent. This involves collecting audio streams through microphones and analyzing user needs based on speech recognition technology. However, this approach has significant drawbacks in complex environments: firstly, environmental noise can easily distort voice signals, drastically reducing recognition accuracy, especially in noisy scenarios such as shopping malls and exhibitions, where interactive terminals often misinterpret user commands due to noise interference; secondly, when multiple users are simultaneously within the interaction range, existing technologies lack effective target object filtering mechanisms, failing to clearly define the strength of the user's interaction intent and their spatial relationship with the interactive terminal, leading to chaotic interaction processes, a surge in data processing volume, and poor feedback targeting.

[0003] This invention provides a multimodal human-computer interaction system based on hybrid perception to solve the above-mentioned technical problems. Summary of the Invention

[0004] This invention aims to solve at least one of the technical problems existing in the prior art. To this end, this invention proposes a multimodal human-computer interaction system based on hybrid perception. Through multimodal data fusion and a dynamic priority mechanism, this invention significantly improves the environmental adaptability and response efficiency of the interactive terminal. The data acquisition module uses a camera and microphone to simultaneously acquire video and audio streams, and combines infrared sensors to calculate user distance, enabling accurate identification of basic objects in complex scenes and avoiding mislabeling. The interaction analysis module constructs a priority scoring system based on distance features and voice energy, which not only ensures the interaction priority of near-field users, but also captures active interaction intentions through voice energy, solving the interaction conflict problem in multi-user scenarios. When the target object has no voice signal within a set time period, the system automatically switches to other voice-emitting basic objects to avoid interaction interruption and improve user experience.

[0005] To achieve the above objectives, a first aspect of the present invention provides a multimodal human-computer interaction system based on hybrid sensing, applied to an interactive terminal, comprising: Data acquisition module: used to acquire video streams through the camera set in the interactive terminal and audio streams through the microphone set after the interactive terminal is woken up; Interaction Analysis Module: Used to identify several basic objects based on the video stream; set interaction priorities for these basic objects based on distance features and speech energy; select the object with the highest interaction priority as the target object; where the distance feature is the distance between the basic object and the interactive terminal; and, It is used to extract lip-shape features from the video stream corresponding to the target object and speech features from the audio stream; and to determine the interactive content of the target object based on the lip-shape features and speech features.

[0006] Preferably, several basic objects are identified based on the video stream, including: Identify user characteristics in the video stream; determine the number of users based on user characteristics; and simultaneously calculate the distance between each user and the display screen using infrared sensors and cameras. Determine if the user is within the identifiable range of the interactive terminal; if yes, mark the user as a base object; otherwise, do not mark the user.

[0007] Preferably, the identifiable range of the interactive terminal refers to the limit range of the voice signal in the audio stream that the microphone can collect.

[0008] Preferably, interaction priorities are set for several basic objects based on distance features and speech energy, including: The relative distances between several basic objects and the interactive terminal are calculated using a camera and microphone, serving as distance features; the voice energy of several basic objects is identified using the microphone. Priority scores for each basic object are calculated based on distance features and speech energy, and interaction priorities are set for each basic object based on these priority scores.

[0009] Preferably, before calculating the priority score of each basic object based on distance features and speech energy, it is determined whether the speech energy of several basic objects is empty; if yes, the distance features are used as the priority score of the basic objects; otherwise, the priority score of each basic object is calculated based on the distance features and speech energy.

[0010] Preferably, priority scores for each basic object are calculated based on distance features and speech energy, including: The relative distances between several basic objects and the interactive terminal are calculated using a camera and microphone, serving as distance features; the voice energy of several basic objects is identified using the microphone. Preset weights are extracted for distance features and speech energy; the distance features and speech energy are weighted based on the preset weights to obtain a priority score; wherein the preset weight of distance features is greater than the preset weight of speech energy.

[0011] Preferably, when there are two base objects with equal and highest interaction priorities, the base object with the smaller distance feature is selected as the target object.

[0012] Preferably, after identifying the target object, it is determined whether the target object's voice signal is collected within a set time period; if yes, the interaction content of the target object is determined based on lip-shape features and voice features; if no, the target object is replaced.

[0013] Preferably, changing the target object includes: If a voice signal of a basic object is detected within a set time period of two, the basic object corresponding to the voice signal is changed to the target object; otherwise, an interactive reminder is given through the display screen.

[0014] Preferably, the interactive content of the target object is determined based on lip shape features and speech features, including: Extract lip features of the target object from the video stream, extract speech features of the target object from the speech signal; and perform spatiotemporal alignment of the lip features and speech features. Calculate the confidence scores of lip shape features and speech features, and construct a weighting function using the sigmoid function; calculate the weighting coefficients of lip shape features and speech features using the weighting function. After weighted fusion of lip shape features and speech features based on weight coefficients, the results are input into a content generation model to obtain interactive content; the content generation model includes a Transformer model or a large language model.

[0015] Preferably, after the interactive terminal provides feedback information based on the interactive content of the target object, it detects whether the voice signal of the target object is received within a set time period of three. Yes, then feedback information is generated based on the interactive content generated from the received voice signal; If not, then the auxiliary object is determined based on the voice signal of the basic object within the set time period, and the interactive content feedback information is generated based on the voice signal of the auxiliary object.

[0016] Preferably, the auxiliary object is determined based on the speech signal of the basic object within a set time period, including: Extract the voice signal of the basic object within a set time period; Analyze the correlation between the speech signals corresponding to each basic object and the interactive content corresponding to the target object; select the basic object with the strongest correlation as the auxiliary object.

[0017] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention significantly improves the environmental adaptability and response efficiency of interactive terminals through multimodal data fusion and dynamic prioritization mechanisms. The data acquisition module uses a camera and microphone to simultaneously acquire video and audio streams, and combines infrared sensors to calculate user distance, enabling accurate identification of basic objects in complex scenes and avoiding mislabeling. The interaction analysis module constructs a priority scoring system based on distance features and voice energy, ensuring the interaction priority of near-field users while capturing active interaction intentions through voice energy, thus solving the interaction conflict problem in multi-user scenarios. When the target object has no voice signal within a set time period, the system automatically switches to other voice-emitting basic objects to avoid interaction interruption and improve user experience.

[0018] 2. This invention proposes a dynamic identification mechanism for auxiliary objects to address the problem of interaction interruption, significantly enhancing the continuity and intelligence of interaction. When the interactive terminal fails to receive the target object's voice signal within a set time after providing feedback, it automatically analyzes the correlation between the voice signals of other basic objects and the original interactive content. It then uses text similarity calculation or semantic matching models to filter out the most relevant auxiliary objects, avoiding irrelevant voice interference. This mechanism maintains the continuity of the interaction process, reduces user waiting time, and is particularly suitable for multi-user collaboration or consultation scenarios. Furthermore, by combining preset time thresholds with correlation analysis, the system can dynamically adapt to the rhythm of multi-user interaction, improving the practicality of the interactive terminal in complex scenarios. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the method steps of the multimodal human-computer interaction method in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram illustrating the steps of the priority scoring method for basic objects in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram illustrating the method steps for determining whether an auxiliary object needs to be identified in Embodiment 2 of the present invention; Detailed Implementation

[0021] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Example 1: Please see Figure 1 The first aspect of this invention provides a multimodal human-computer interaction system based on hybrid sensing, applied to an interactive terminal, comprising: Data acquisition module: used to acquire video streams through the camera set in the interactive terminal and audio streams through the microphone set after the interactive terminal is woken up; Interaction Analysis Module: Used to identify several basic objects based on the video stream; set interaction priorities for these basic objects based on distance features and speech energy; select the object with the highest interaction priority as the target object; where the distance feature is the distance between the basic object and the interactive terminal; and, It is used to extract lip-shape features from the video stream corresponding to the target object and speech features from the audio stream; and to determine the interactive content of the target object based on the lip-shape features and speech features.

[0023] The interactive terminal (AI large screen) allows users to interact with the digital human. The terminal mainly includes a display screen, camera, microphone, storage device, and interfaces. The camera supports face detection and tracking, and low-light wake-up, such as a 16MP RGB + infrared dual-mode sensor. The microphone supports sound source localization and pickup, such as a 6-microphone circular array with a pickup distance of 5 meters. The main control chip handles data processing. The display screen shows the interaction between the digital human and the user, and provides feedback based on the user's interaction.

[0024] The data acquisition module connects to various data acquisition devices in the interactive terminal, such as cameras and microphones, and sends the acquired data to the interactive analysis module. The interactive analysis module analyzes the data sent by the data acquisition module according to preset rules, generates feedback information based on the analysis results, and displays the feedback information to the user through the screen.

[0025] Data acquisition module: Used to acquire video streams through the camera set in the interactive terminal and audio streams through the microphone set after the interactive terminal is woken up.

[0026] The interactive terminal can be woken up in several ways. It can be voice-activated, such as when the user speaks a wake word directly in front of the display screen, and the terminal is activated after successful verification. It can also be visually activated, such as through facial recognition, where the terminal is activated when the user is directly in front of the display screen, or by making a preset gesture. Finally, it can be activated via touchscreen, where the terminal is activated upon detecting a touch input from the user.

[0027] In existing interactive terminal solutions, most methods determine user needs by recognizing the user's voice signals and then generate interactive content for the terminal to analyze. However, analyzing user needs based on voice signals is easily affected by the surrounding environment. If there is significant ambient noise, it greatly affects the accuracy of voice signal recognition and analysis, making it difficult to guarantee the accuracy of the content fed back by the interactive terminal.

[0028] The interactive analysis module is used to identify several basic objects based on the video stream.

[0029] The basic user is the user within the corresponding interaction range of the interactive terminal, such as being within the camera's field of view and within the microphone's pickup range. If the user is within the camera's and microphone's capture range, and it is determined that the user has the intention to interact, then the user is marked as a basic user. Conversely, if it is determined that the user does not have the intention to interact, then the corresponding user does not need to be marked as a basic user.

[0030] The basic object can be determined by following these steps: The system uses object detection algorithms to identify user features in the video stream; determines the number of users based on these features; and simultaneously calculates the distance between each user and the display screen using infrared sensors and cameras. Determine if the user is within the identifiable range of the interactive terminal; if yes, mark the user as a base object; otherwise, do not use the user as a base object.

[0031] After the interactive terminal is woken up, a video stream is captured through the camera, and a target detection algorithm is used to identify user features (including facial features and human body features) in the video stream to determine several basic objects within the interactive range of the interactive terminal.

[0032] Once the interactive terminal is woken up, the camera is activated to capture the video stream in front of the interactive terminal's display screen. The target detection algorithm is used to detect whether user features, i.e., whether there are faces and bodies, are present in the video stream. If user features are present, the number of users within the interactive range of the interactive terminal is determined based on the user features, thus obtaining several basic objects.

[0033] Interactive terminals can be set up in various scenarios, such as shopping malls, office buildings, and hospitals. In many scenarios, many people may appear in front of the interactive terminal at the same time, such as a family shopping in a mall, all within the recognition range of the interactive terminal; or multiple colleagues interacting in front of the interactive terminal at the same time. In this case, it is necessary to accurately determine how many users are present to lay the foundation for subsequently determining the interaction content.

[0034] In a preferred embodiment, the identifiable range of the interactive terminal refers to the limit of the range of voice signals in the audio stream that the microphone can capture.

[0035] The identifiable range of an interactive terminal refers to the range within which video data and audio signals can be clearly captured. Since the acquisition and analysis of video data are less affected by the surrounding environment, the interactive range is determined by the distance at which the microphone can clearly detect the audio signal (in the audio stream). This range is the limit of the microphone's audio stream acquisition capability. Furthermore, the distance at which the microphone can clearly detect the audio signal decreases as the ambient noise of the interactive terminal increases; therefore, the determination of the interactive range is dynamic.

[0036] The audio stream captured by a microphone may contain multiple sound signals. Therefore, the limit of the captured audio stream cannot be used as the recognizable range of the interactive terminal. Instead, the limit of the microphone's captured speech signal should be used. The limit of the captured speech signal is greatly affected by the environment. A correlation between the interactive terminal's settings and the limit range can be established in advance. The corresponding limit range can then be matched to the actual settings of the interactive terminal to determine its recognizable range. Alternatively, the recognizable range of the interactive terminal in its environment can be determined through testing and simulation.

[0037] Interaction Analysis Module: Used to identify several basic objects based on the video stream; set interaction priorities for several basic objects based on distance features and speech energy; and select the object with the highest interaction priority as the target object.

[0038] The purpose of setting interaction priority is to facilitate the interactive terminal in identifying the target object, thereby determining the interaction content based on the interaction information of the target object, and providing targeted feedback information based on the interaction content.

[0039] If the interactive terminal identifies only one basic object, that basic object is the target object, and the interaction content only needs to be determined by analyzing that target object. However, if the interactive terminal treats several identified basic objects as interactive objects, it needs to process the video and audio streams of all basic objects simultaneously to determine the interaction content for each. If the interaction content of the various basic objects is unrelated, the interactive terminal will struggle to provide effective feedback, impacting user experience and increasing its data processing load.

[0040] Therefore, if an interactive terminal simultaneously identifies multiple basic objects, it needs to determine a target object from among them. The interactive content for that target object is then identified from the video and audio streams. Based on this interactive content, the terminal can generate feedback information and use a digital human to deliver this feedback to the target object. This improves the relevance of the feedback generated by the interactive terminal and avoids unnecessary data processing.

[0041] Based on distance features and speech energy, interaction priorities are set for several basic objects, including: The relative distances between several basic objects and the interactive terminal are calculated using a camera and microphone, serving as distance features; the voice energy of several basic objects is identified using the microphone. Priority scores for each basic object are calculated based on distance features and speech energy, and interaction priorities are set for each basic object based on these priority scores.

[0042] Distance characteristics refer to the distance between each basic object identified by the camera and the interactive terminal. The smaller the distance, the stronger the interaction intention between the basic object and the interactive terminal. Voice energy mainly refers to the voice energy in the voice signal emitted by each basic object. The greater the voice energy, the stronger the interaction intention between the basic object and the interactive terminal. Voice energy corresponds to the energy of the voice signal passing through a unit area per unit time during propagation. It is proportional to the square of the sound wave amplitude. In digital voice signals (time-domain waveforms), it is represented as the sum of the squares of the amplitudes of the sampling points.

[0043] After calculating the priority score of each basic object, the basic objects are sorted from lowest to highest priority score, and each basic object is assigned a number according to the sorting. This number serves as the interaction priority. For example, if the basic objects have already been sorted from lowest to highest priority score, the basic object with the lowest priority score is assigned an interaction priority of 1, and subsequent basic objects are assigned an interaction priority of 2, and so on. The larger the number, the higher the interaction priority.

[0044] In a preferred embodiment, before calculating the priority score of each basic object based on distance features and speech energy, it is determined whether the speech energy of several basic objects is empty; if yes, the distance features are used as the priority score of the basic object; otherwise, the priority score of each basic object is calculated based on the distance features and speech energy.

[0045] A blank speech energy value indicates that no speech signal from the underlying object has been detected. In this case, there is no need to perform various standardization processes on the distance features; the distance features (with consistent units for all distance features) are directly used as the priority score for the underlying object. Compared to the preset weighting of speech energy to 0 when no speech signal is detected, this technical solution sets up a judgment step. If the speech energy of all underlying objects is blank, no further processing of the distance features is required to determine the target object, avoiding unnecessary data processing and reducing the resource consumption of the interactive terminal.

[0046] The speech energy of each basic object can be calculated from its speech signal. If the interactive terminal detects the distance characteristics of each basic object within a set time period and also detects multiple speech signals, then the sound source localization technology is used to match the basic object corresponding to the speech signal, thereby calculating the speech energy of each basic object.

[0047] It should be noted that setting a time limit is to improve interaction efficiency and quickly determine the duration set for the target object. This could be three seconds or five seconds, for example, starting the timer when the interactive terminal is woken up. The voice signal collected within five seconds can be used to calculate the priority score of the basic object.

[0048] During the process of microphone detection of voice signals, voice signals from more than one basic object may be received simultaneously within a set time period. Since several basic objects are within the interaction range of the interactive terminal, and the interaction range is relatively small, it is difficult to determine the attribution of multiple voice signals. Therefore, based on the voice signal acquisition by a 6-microphone wake-up array, the MVDR beamforming algorithm is used to achieve sound source localization. Based on the sound source localization results, the basic object corresponding to each voice signal is determined, thereby accurately calculating the priority score of each basic object.

[0049] Please see Figure 2 In a preferred embodiment, priority scores for each basic object are calculated based on distance features and speech energy, including: The relative distances between several basic objects and the interactive terminal are calculated using a camera and microphone, serving as distance features; the voice energy of several basic objects is identified using the microphone. Preset weights are extracted for distance features and speech energy; the distance features and speech energy are weighted based on the preset weights to obtain a priority score; wherein the preset weight of distance features is greater than the preset weight of speech energy.

[0050] The formula for weighting distance features and speech energy based on preset weights is as follows: Calculate priority score ;in, The distance features are standardized. The speech energy has been standardized. and The preset weights are used. Priority scoring is used to determine the target object from the base objects, so the preset weights for distance features are set to be greater than the preset weights for speech energy.

[0051] Standardization processing refers to the extraction and normalization of outliers in distance features and speech energy. There are many publicly available methods for standardization processing, so they will not be elaborated on here.

[0052] It should be noted that distance features, specifically the distance between the base object and the interactive terminal, refer to the distance between the base object's face (or head) and the interactive terminal (or display screen) when the base object is in front of the display screen and facing the screen. Distance features can be used to identify user facial features or human contours in the RGB video stream using object detection algorithms (such as YOLOv8 or Dlib) to determine the user's pixel coordinates in the image. Combined with an infrared depth map, the 2D pixel coordinates are mapped to 3D space to obtain the three-dimensional coordinates of the user's key points, thus obtaining the corresponding vertical distance between the user and the display screen.

[0053] If, after the interactive terminal is woken up, it can only detect the distance features of each basic object within a set time period, while the voice energy of all basic objects cannot be detected, it indicates that none of the basic objects are emitting voice signals. In this case, when calculating the priority score, the preset weight of voice energy can be set to 0, and only the distance features are considered when calculating the priority score of each basic object. This can be understood as disregarding the voice energy of the basic object and only considering the distance between the basic object and the interactive terminal; the smaller the distance, the higher its priority score.

[0054] After calculating the priority scores of each basic object, the interaction priority of each basic object is set according to the priority scores; the higher the priority score, the higher the interaction priority. The basic object with the highest interaction priority is selected as the target object, and the interactive terminal identifies the interaction information of the target object to determine the interaction content.

[0055] When two base objects have the same and highest interaction priority, the base object with the smaller distance feature is selected as the target object.

[0056] The interaction analysis module is used to extract lip-shape features from the video stream corresponding to the target object and speech features from the audio stream; based on the lip-shape features and speech features, the interaction content of the target object is determined.

[0057] After identifying the target object, determine whether the target object's voice signal is collected within a set time period. If yes, determine the target object's interaction content based on lip-shape features and voice features. If no, change the target object.

[0058] In a preferred embodiment, replacing the target object includes: If a voice signal of a basic object is detected within a set time period of two, the basic object corresponding to the voice signal is changed to the target object; otherwise, an interactive reminder is given through the display screen.

[0059] If the voice signal of the target object is not detected in the audio stream collected in real time via the microphone within the set time period two, the target object should be replaced. Specifically, it can identify whether the voice signal of the basic object exists in the audio stream collected in real time within the set time period two. If it exists, the basic object corresponding to the voice signal is marked as the target object; otherwise, an interactive reminder is given through the display screen.

[0060] The second setting is a preset time length, such as three seconds or five seconds. If no voice signal from the target object is detected after the second setting time, starting from when the target object is determined, the target object will be considered for replacement.

[0061] It should be noted that the microphone of the interactive terminal does not only collect the voice signal of the target object, but also the voice signals of all basic objects. However, if the voice signal of the target object can be used to determine the interaction content, the voice signals of the basic objects can be disregarded. If there are voice signals of multiple basic objects within a given time period, the analysis will determine whether the voice signals conform to the interaction domain of the interactive terminal. If they conform to the interaction domain, the basic object that emitted the earliest voice signal will be changed to the target object.

[0062] When the interactive terminal is first activated, the target object is identified using distance and voice features, primarily voice energy. During the activation phase, the target object is unaware of the terminal's internal data processing flow, and anyone could potentially emit a voice signal. If a target object without interactive intent is unknowingly identified as the target object, it will not emit a voice signal after the terminal is activated, and the terminal may not receive any voice signals for an extended period. In this case, the target object needs to be adjusted promptly to ensure continued interaction.

[0063] If no voice signal from the target object is detected within the set time period two, but other basic objects with interactive intent may emit voice signals, the basic object emitting the voice signal will be changed to the target object, and the original target object will become the basic object. The interactive terminal can generate interactive content based on the updated voice signal of the target object within the set time period two and provide feedback based on the interactive content.

[0064] After identifying the target object, a video stream is captured in real time via a camera, and an audio stream is captured in real time via a microphone. Video data of the target object is extracted from the video stream, and lip-sync features are identified using an object detection algorithm. Simultaneously, speech signals are extracted from the audio stream. After preprocessing, the lip-sync features and speech signals are fused and identified to determine the interactive content, which represents the feedback the target object expects from the interactive terminal.

[0065] The interaction content of the target object is determined based on lip shape features and speech features, including: Extract lip features of the target object from the video stream, extract speech features of the target object from the speech signal; and perform spatiotemporal alignment of the lip features and speech features. Calculate the confidence scores of lip shape features and speech features, and construct a weighting function using the sigmoid function; calculate the weighting coefficients of lip shape features and speech features using the weighting function. After weighted fusion of lip shape features and speech features based on weight coefficients, the results are input into a content generation model to obtain interactive content; the content generation model includes a Transformer model or a large language model.

[0066] The process of determining the interactive content of a target object based on lip-reading and speech features essentially involves recognizing the target object's speech content, i.e., the interactive content, based on lip-reading and speech features while comprehensively considering environmental influences. For example, if the target object asks for the specific location of a store, the interactive content would be "Where is store X?", and the feedback information would be the store's location and navigation directions.

[0067] Large language models, including GPT, BERT, and Wenxin Yiyan, learn the structure, rules, and semantics of language by training on massive amounts of text data, generating natural language-style text or answering natural language questions.

[0068] Extracting lip features of the target object from the video stream, including: Based on Dlib, 68 facial key points were extracted, 19 lip key points were located, and 2D / 3D lip contour vectors were constructed. The lip ROI was grayscaled and normalized (64×64 pixels) to eliminate the influence of lighting and scale. C3D convolutional neural network is used to process continuous frames and extract (52-dimensional) lip features, which can capture dynamic parameters such as opening degree and lip corner displacement.

[0069] Extracting speech features of the target object from the speech signal includes: calculating the 13th order MFCC (Mel frequency cepstral coefficients) + 13th order first-order difference + 13th order second-order difference of the speech signal to generate a 39-dimensional feature vector as the speech feature.

[0070] The confidence score of lip shape features refers to the mean maximum cosine similarity between the lip shape features and the pre-trained Viseme library (i.e., a lip shape unit library, corresponding to lip shape templates for different pronunciations). It contains 52 standard lip shape feature templates, each a feature vector of equal dimension. The confidence score of speech features is the average log-posterior probability of the Conformer output token. Conformer is a deep learning model that combines convolutional neural networks (CNN) and Transformer architectures, primarily used for sequence modeling tasks such as speech recognition and speech processing. It combines the advantages of both, capturing both the temporal relationships of local features and handling long-range dependencies, demonstrating excellent performance in the speech domain.

[0071] The Sigmoid function constructs the weight function as follows: , The confidence level of the speech features. The confidence level for lip shape diagnosis, This is the control coefficient, defaulted to 5. A larger control coefficient results in more sensitive weight changes, while a smaller coefficient leads to smoother weight changes. The weight coefficient for speech features is... The weighting coefficient for lip shape features is 1- .

[0072] It is worth noting that, in order to reduce data processing volume and improve interaction efficiency, the voice signal in the audio stream corresponding to the target object can be verified. If the verification fails, it is considered that the voice signal has not been recognized; if the verification passes, it is considered that the voice signal has been recognized. The verification of the voice signal is mainly to avoid invalid interactions. The content of the voice signal can be matched with the preset interaction domain in the interactive terminal. If the two match successfully, the verification passes; if the two fail to match, no further processing is required.

[0073] Similarly, if the target object does not emit a valid (verifiable) voice signal, the target object needs to be re-determined from the base objects. If the voice signal emitted by the base object also fails verification, then the base object cannot be used as the target object. It should be noted that if interactive content can be generated based on the voice signal, then the voice signal should be verifiable.

[0074] Example 2: Compared to Example 1, this example provides a solution for the interactive terminal failing to maintain the interactive state for a period of time after receiving feedback information based on the interactive content of the target object.

[0075] Please see Figure 3 After the interactive terminal provides feedback based on the interactive content of the target object, it detects whether the voice signal of the target object is received within a set time of three. Yes, then feedback information is generated based on the interactive content generated from the received voice signal; If not, then the auxiliary object is determined based on the voice signal of the basic object within the set time period, and the interactive content feedback information is generated based on the voice signal of the auxiliary object.

[0076] Based on the speech signals of the base object within a set time period, auxiliary objects are determined, including: Extract the voice signal of the basic object within a set time period; Analyze the correlation between the speech signals corresponding to each basic object and the interactive content corresponding to the target object; select the basic object with the strongest correlation as the auxiliary object.

[0077] After the interactive terminal is woken up, it generates interactive content by fusing the video and audio streams of the target object and provides feedback information. If no next voice signal from the target object is received within a set time period of three, the terminal analyzes the correlation between the voice signals of other basic objects within the set time period and the interactive content of the target object. If the correlation is high, the corresponding basic object is used as an auxiliary object. The interactive terminal generates interactive content based on the voice signals and video data of the auxiliary object within the set time period and provides feedback information on the interactive content.

[0078] Relevance analysis can be performed by converting speech signals into text and calculating the text similarity between this text and the content of the interaction with the target object. Alternatively, semantic matching models can be used to analyze the semantic relevance between the speech signals and the corresponding text of the interaction with the target object.

[0079] In a preferred embodiment, timing begins after the interactive terminal provides feedback information based on the interactive content of the target object, while simultaneously collecting and analyzing the voice signal of the base object, and generating interactive content based on the video data and voice signal of the base object.

[0080] To improve the continuity of interaction, the interactive terminal starts timing after providing feedback on the interaction content to the target object. Simultaneously, it analyzes whether the underlying object has emitted a voice signal. If the content corresponding to the voice signal is related to the target object's previous interaction content, the voice signal is combined with the corresponding video data of the underlying object to generate interactive content. This allows for rapid feedback if no voice signal from the target object is detected within a set time period of three seconds. The set time period of three seconds is a pre-defined duration, such as three or five seconds.

[0081] If the set time period reaches three without receiving a new voice signal from the target object, the feedback information corresponding to the voice signal of the matching base object (auxiliary object) within the set time period will be displayed on the screen. If the set time period does not reach three and the interactive terminal receives a voice signal from the target object, the feedback information of the already matched base object voice signal will no longer be used; it can be stored or deleted.

[0082] After receiving feedback on the interactive content based on the auxiliary object's voice signal, a timer begins. If the timer has not reached the set time three, the interactive content is generated and fed back based on whichever of the target object or the auxiliary object emitted a voice signal during that time period. If both emitted voice signals, the interactive content is generated and fed back sequentially according to the order of detection. Furthermore, if none of the three emitted voice signals by the set time three, the voice signals of other basic objects within the set time three are verified, and the auxiliary object is re-determined.

[0083] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A multimodal human-computer interaction system based on hybrid sensing, applied to an interactive terminal, characterized in that, include: Data acquisition module: used to acquire video streams through the camera set in the interactive terminal and audio streams through the microphone set after the interactive terminal is woken up; Interaction analysis module: used to identify several basic objects based on video stream; set interaction priorities for several basic objects based on distance features and speech energy; select the object with the highest interaction priority as the target object; wherein, the distance feature is the distance between the basic object and the interactive terminal; and used to extract lip-shape features from the video stream corresponding to the target object and extract speech features from the audio stream; and determine the interaction content of the target object based on the lip-shape features and the speech features.

2. The multimodal human-computer interaction system based on hybrid perception according to claim 1, characterized in that, Identify several basic objects based on the video stream, including: Identify user characteristics in the video stream; determine the number of users based on user characteristics; and simultaneously calculate the distance between each user and the display screen using infrared sensors and cameras. Determine if the user is within the identifiable range of the interactive terminal; if yes, mark the user as a base object; otherwise, do not mark the user.

3. The multimodal human-computer interaction system based on hybrid perception according to claim 1, characterized in that, Based on distance features and speech energy, interaction priorities are set for several basic objects, including: The relative distances between several basic objects and the interactive terminal are calculated using a camera and microphone, serving as distance features; the voice energy of several basic objects is identified using the microphone. Priority scores for each basic object are calculated based on distance features and speech energy, and interaction priorities are set for each basic object based on these priority scores.

4. A multimodal human-computer interaction system based on hybrid perception according to claim 3, characterized in that, Before calculating the priority score of each basic object based on distance features and speech energy, it is determined whether the speech energy of several basic objects is empty; if yes, the distance features are used as the priority score of the basic object; otherwise, the priority score of each basic object is calculated based on the distance features and speech energy.

5. A multimodal human-computer interaction system based on hybrid perception according to claim 3 or 4, characterized in that, Priority scores for each basic object are calculated based on distance features and speech energy, including: The relative distances between several basic objects and the interactive terminal are calculated using a camera and microphone, serving as distance features; the voice energy of several basic objects is identified using the microphone. Preset weights are extracted for distance features and speech energy; the distance features and speech energy are weighted based on the preset weights to obtain a priority score; wherein the preset weight of distance features is greater than the preset weight of speech energy.

6. A multimodal human-computer interaction system based on hybrid perception according to claim 1, characterized in that, After identifying the target object, determine whether the target object's voice signal is collected within a set time period. If yes, determine the target object's interaction content based on the lip shape features and the voice features. If no, change the target object.

7. A multimodal human-computer interaction system based on hybrid perception according to claim 6, characterized in that, Replacing the target object includes: If a voice signal of a basic object is detected within a set time period of two, the basic object corresponding to the voice signal is changed to the target object; otherwise, an interactive reminder is given through the display screen.

8. A multimodal human-computer interaction system based on hybrid perception according to claim 6, characterized in that, Determining the interaction content of the target object based on the lip shape features and the speech features includes: Extract lip features of the target object from the video stream, extract speech features of the target object from the speech signal; and perform spatiotemporal alignment of the lip features and speech features. Calculate the confidence scores of lip shape features and speech features, and construct a weighting function using the sigmoid function; calculate the weighting coefficients of lip shape features and speech features using the weighting function. After weighted fusion of lip shape features and speech features based on weight coefficients, the results are input into the content generation model to obtain interactive content; the content generation model is used to generate interactive content based on the fused features.

9. A multimodal human-computer interaction system based on hybrid perception according to claim 1, characterized in that, After the interactive terminal provides feedback based on the interactive content of the target object, it detects whether the voice signal of the target object is received within a set time of three. Yes, then feedback information is generated based on the interactive content generated from the received voice signal; If not, then the auxiliary object is determined based on the voice signal of the basic object within the set time period, and the interactive content feedback information is generated based on the voice signal of the auxiliary object.

10. A multimodal human-computer interaction system based on hybrid perception according to claim 9, characterized in that, Based on the speech signals of the base object within a set time period, auxiliary objects are determined, including: Extract the voice signal of the basic object within a set time period; Analyze the correlation between the speech signals corresponding to each basic object and the interactive content corresponding to the target object; select the basic object with the strongest correlation as the auxiliary object.

Citation Information

Cited By

  • Multi-user interaction response strategy automatic switching method fusing behaviors and emotions

    CN121934719A