Exhibition hall voice interaction method and system

By combining a rotatable microphone array and a depth camera to dynamically track sound sources in the exhibition hall, generating scene state feature vectors, and selecting the optimal speech recognition model, the problems of noise interference and multi-device collaboration in the exhibition hall voice interaction system were solved, achieving high accuracy and personalized interaction.

CN120823832BActive Publication Date: 2026-01-16STATEGRID RUIJIA (TIANJIN) INTELLIGENT ROBOT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511323760.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2026-01-16
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

The existing exhibition hall voice interaction system is severely affected by noise in noisy environments, making it difficult to dynamically track sound sources. It also suffers from low voice recognition accuracy, lacks scene adaptability and personalized interaction, and has insufficient multi-device collaboration capabilities.

Method used

It employs a combination of a rotatable microphone array and a depth camera to dynamically track sound sources, fuses environmental and user data to generate scene state feature vectors, selects the optimal speech recognition model, combines exhibit knowledge base and dynamic knowledge graph to provide personalized answers, and achieves device collaboration through a distributed collaborative control platform.

Benefits of technology

It improved the accuracy of speech recognition, enhanced the user experience, enabled personalized interaction, and strengthened the system's real-time response and multi-device collaboration capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823832B_ABST
    Figure CN120823832B_ABST
Patent Text Reader

Abstract

A kind of exhibition hall voice interaction method and system, the method includes: collecting the environmental data and user sound source data of exhibition hall;According to environmental data and user sound source data, judge whether it needs to move tracking sound source;Based on user sound source data and user visual data, the user is positioned tracking;When not needing to move tracking sound source, environmental data, user related data and the product information data read by robot upper RFID reader are fused, and scene state feature vector is generated;Scene state feature vector is parsed, and after selecting optimal speech recognition model according to current exhibition hall environmental noise level, user visit behavior, user position, user interest weight obtained by parsing, the question of user is identified, and the matching answer to the question of user is generated, and the matching answer is output to user.This application can improve the dynamic tracking capability of sound source position, improve the performance and user experience of exhibition hall voice interaction system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of intelligent speech processing, and particularly relates to a hall voice interaction method and system combining robot mobile collection, multi-modal data processing and FunASR voice model. BACKGROUND

[0002] The existing hall voice interaction system has the following problems in actual application: the traditional fixed voice collection device is greatly disturbed by environmental noise, and it is difficult to dynamically track the sound source position, resulting in limited voice collection effect; the voice synthesis lacks scene adaptability, the output content is disconnected from the visit line, and personalized interactive experience cannot be provided; the data processing delay is high, and it is difficult to meet the real-time interaction demand; the collaboration ability among multiple devices is poor, and there is no unified control platform, resulting in low overall efficiency of the system. For example, in a noisy exhibition hall environment, a fixed microphone array is easily affected by background noise, and the signal-to-noise ratio decreases significantly, directly affecting the voice recognition accuracy; at the same time, the traditional voice synthesis technology is usually based on a static template, and cannot dynamically adjust the output content according to the specific position of the user, the exhibit information or the visit path, resulting in poor user experience. In addition, due to the large number of devices and wide distribution in the exhibition hall, the existing system often cannot realize efficient multi-device collaboration, further limiting the practicality and expandability of the system. Therefore, it is of great significance to develop an exhibition hall intelligent voice interaction system that can overcome the above defects.

[0003] Prior art document 1 (CN119309580A) discloses a hall robot visual language navigation method based on a large model. Prior art document 2 (CN116931728A) discloses a multi-modal human-computer interaction control system and method of a humanoid robot.

[0004] However, prior art document 1 mainly relies on static image recognition and a pre-set language model to generate navigation content, and does not involve the active collection and optimized processing of user voice signals by the robot during movement, lacks dynamic tracking ability of the sound source position, and relies on a fixed microphone array for voice interaction, which is seriously disturbed by background noise, resulting in a decrease in voice recognition accuracy; in addition, the system does not consider the adaptive adjustment of voice interaction strategies according to user behavior state, environmental changes and other factors, and the interactive experience is single, which cannot meet the personalized needs.

[0005] Although prior art document 2 realizes the integration of multi-modal data, its voice module is still based on a general voice recognition model, and has not been optimized and adjusted in structure for a specific exhibition hall scene, lacking a linkage mechanism for voice recognition and environmental perception; in addition, the system adopts a centralized control architecture, and has not established an efficient distributed communication mechanism, making it difficult to support multi-device collaboration in a large-scale exhibition hall. SUMMARY

[0006] The application provides an exhibition hall voice interaction method and system, aiming to solve the deficiencies of traditional voice interaction systems in noise suppression, scene adaptation, real-time response and multi-device collaboration by fusing dynamic acquisition technology, multi-modal data processing and adaptive voice models.

[0007] The application adopts the following technical solutions.

[0008] The first aspect of the application provides an exhibition hall voice interaction method, comprising: collecting environmental data and user sound source data of the exhibition hall;

[0009] According to the environmental data and the user sound source data, it is determined whether the sound source needs to be tracked;

[0010] When the sound source needs to be tracked, user visual data is collected, and the user is tracked based on the user sound source data and the user visual data;

[0011] When the sound source does not need to be tracked, the environmental data, user-related data and exhibit information data read by the RFID reader on the robot are fused to generate a scene state feature vector;

[0012] The scene state feature vector is analyzed, and the optimal voice recognition model is selected according to the current exhibition hall environmental noise level, user visit behavior, user position and user interest weight obtained by analysis, the user's question is recognized, the matched answer is generated for the user's question according to the voice information obtained by recognition combined with the exhibit knowledge base and the dynamic knowledge graph, and the matched answer is output to the user.

[0013] Optionally, the user is positioned based on the user sound source data and the user visual data, comprising:

[0014] The rotatable microphone array is rotated so that the normal direction of the rotatable microphone array is aligned with the maximum signal-to-noise ratio azimuth angle;

[0015] The depth camera is turned to the maximum signal-to-noise ratio azimuth angle, and point cloud data is output;

[0016] The maximum signal-to-noise ratio azimuth angle is converted into an initial three-dimensional coordinate in a visual SLAM coordinate system based on a preset coordinate mapping function;

[0017] The point cloud data and the initial three-dimensional coordinate are fused and estimated to obtain a three-dimensional coordinate of the user after fusion.

[0018] Optionally, the point cloud data and the initial three-dimensional coordinate are fused and estimated according to the following formula to obtain a three-dimensional coordinate of the user after fusion:

[0019]

[0020]

[0021]

[0022] Optionally, the corresponding weight is dynamically adjusted according to the signal-to-noise ratio of the sound source and the point cloud data density obtained by the rotatable microphone array, and the weight is specifically:

[0023] The signal-to-noise ratio of the sound source and the point cloud data density are both normalized and mapped into the interval of 0 to 1 to obtain the normalized signal-to-noise ratio of the sound source and the point cloud data density;

[0024] The difference between the normalized signal-to-noise ratio of the sound source and the point cloud data density is calculated; the ratio of the normalized signal-to-noise ratio of the sound source to the difference is determined as the weight of the initial three-dimensional position, and the ratio of the normalized point cloud data density to the difference is determined as the weight of the user position coordinates directly observed by the point cloud data.

[0025] Optionally, the method further comprises:

[0026] The face detection is used to calculate the face center direction angle of the user, and the maximum signal-to-noise ratio azimuth and the face center direction angle are used as observation vectors to estimate the face orientation state of the user to obtain an optimal face center direction angle.

[0027] The difference between the optimal face center direction angle and the face center direction angle in the current observation vector is calculated as the optimal compensation angle of the microphone array.

[0028] The gimbal rotation of the microphone array is adjusted according to the optimal compensation angle so as to be aligned with the face center direction of the user, thereby realizing the collaborative alignment of acoustics and vision.

[0029] Optionally, the face detection is used to calculate the face center direction angle of the user, and the maximum signal-to-noise ratio azimuth and the face center direction angle are used as observation vectors to estimate the face orientation state of the user to obtain an optimal face center direction angle, and the method comprises:

[0030] The state vector is set as and the observation vector is set as , wherein represents the state of the face orientation of the user at time k, represents the face center direction angle of the user, represents the change rate of the face orientation angle, i.e., the angular velocity, is the maximum signal-to-noise ratio azimuth, is the face center direction angle calculated by the face detection based on the face key points by the depth camera; and

[0031] The state vector of the last moment is fused with the state transition matrix and the process noise to obtain the state of the current moment, to obtain the prior state estimation, and the state estimation error covariance matrix of the last moment is added to the process noise covariance matrix to obtain the prior estimation error covariance matrix;

[0032] The new observation vector is multiplied by the observation matrix, and then the difference between the prior state estimation is obtained, the difference is multiplied by the gain matrix, and the sum of the prior state estimation is obtained to obtain the posterior state estimation, the gain matrix is multiplied by the observation matrix, and then the difference between the unit matrix is obtained, the difference is multiplied by the prior estimation error covariance matrix to obtain the posterior estimation error covariance matrix;

[0033] When the error covariance matrix tends to a constant positive definite matrix, the iteration is stopped, the optimal gain matrix is obtained, and then the optimal face center direction angle is obtained.

[0034] Optionally, the method further comprises:

[0035] Real-time detection of the noise level in the exhibition hall, adjusting the acquisition parameters of the microphone array according to the noise of different frequencies, specifically:

[0036] When the noise energy in the first frequency range exceeds the first preset energy threshold, the main lobe width is adjusted according to the difference between the noise energy and the first preset energy threshold, and / or, a zero point is added in the interference direction, and / or, the suppression amount is dynamically adjusted according to the noise energy, and / or, the weight of each antenna of the microphone array is adaptively adjusted by least square or recursive least square;

[0037] When the noise energy in the second frequency range exceeds the second preset energy threshold, the main lobe width is adjusted according to the difference between the noise energy and the second preset energy threshold, and / or, the compensation amount is dynamically adjusted according to the noise energy.

[0038] Optionally, the main lobe width is adjusted according to the difference between the noise energy and the first preset energy threshold as follows:

[0039]

[0040]

[0041] In the formula, The adjusted main lobe width is represented by The maximum width is represented by The minimum width is represented by The noise energy in the first frequency range is represented by The first preset energy threshold is represented by The normalized adjustment factor is represented by.

[0042] Optionally, the suppression amount is dynamically adjusted according to the noise energy as follows:

[0043]

[0044] In the formula, Indicates the amount of inhibition. This represents the noise energy within the first frequency band. This indicates the first preset energy threshold;

[0045] The compensation amount is dynamically adjusted according to the noise energy using the following formula:

[0046]

[0047] In the formula, Indicates the amount of inhibition. This represents the noise energy within the second frequency band. This indicates the second preset energy threshold.

[0048] Combination Figure 2 As shown, the second aspect of this application provides a voice interaction system for an exhibition hall, including: a mobile sound source tracking module, used to collect environmental data and user sound source data of the exhibition hall, and determine whether it is necessary to track the mobile sound source based on the environmental data and user sound source data; when it is necessary to track the mobile sound source, it collects user visual data, and performs location tracking of the user based on the user sound source data and user visual data;

[0049] The multimodal data processing module is used to fuse environmental data, user-related data, and exhibit information data read by the RFID reader on the robot to generate a scene state feature vector when it is not necessary to track the sound source.

[0050] The adaptive voice interaction engine is used to parse the scene state feature vector, and select the optimal voice recognition model based on the current exhibition hall environmental noise level, user visit behavior, user location, and user interest weights obtained from the parsing. Then, it recognizes the user's question, and generates a matching answer for the user's question based on the recognized voice information combined with the exhibit knowledge base and dynamic knowledge graph, and outputs the matching answer to the user.

[0051] Optionally, the system further includes:

[0052] The distributed collaborative control platform is used to build a device communication network using the MQTT protocol, connect robots and other devices in the exhibition hall to the MQTT communication network, and allocate tasks to each device and synchronize device status based on the MQTT communication network.

[0053] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when loaded onto the processor, implements the aforementioned exhibition hall voice interaction method.

[0054] The fourth aspect of the present application provides a computer readable storage medium, the computer readable storage medium stores a computer program, the computer program is executed by a processor to realize the above-mentioned exhibition hall voice interaction method.

[0055] Compared with the prior art, the beneficial effects of the present application at least include:

[0056] The present application determines whether to move and track the sound source according to the environmental data and the sound source data, and when it is necessary to move and track the sound source, the user visual data is collected, and the user is positioned and tracked based on the user sound source data and the user visual data. Compared with the prior art which lacks dynamic tracking ability of the sound source position and the voice interaction which relies on fixed microphone array and is seriously interfered by background noise, resulting in the decrease of voice recognition accuracy, the present application can fuse the sound source data and the user visual data, improve the dynamic tracking ability of the sound source position, and reduce the interference of background noise and improve the voice recognition accuracy before voice interaction according to the environmental data and the sound source data.

[0057] The present application fuses the environmental data, the user related data and the exhibit information data read by the RFID reader on the robot to generate a scene state feature vector, provides context support for voice recognition and feedback by constructing a more comprehensive scene state feature vector, generates a voice answer more in line with the user's individual needs, and thus improves the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor. Among them:

[0059] Figure 1 is a flowchart of an exhibition hall intelligent voice interaction method provided by an embodiment of the present application;

[0060] Figure 2 is a system architecture diagram of an exhibition hall intelligent voice interaction method provided by an embodiment of the present application;

[0061] Figure 3 is a working flowchart of a mobile sound source tracking module provided by an embodiment of the present application;

[0062] Figure 4 is a structure diagram of a multi-modal data fusion mechanism provided by an embodiment of the present application;

[0063] Figure 5Fig. 1 is a schematic diagram of an adaptive voice interaction engine architecture provided by an embodiment of the present application. DETAILED DESCRIPTION

[0064] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. The embodiments described in the present application are only a part of the embodiments of the present application, but not all the embodiments. Based on the spirit of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0065] In combination with Figure 1 As shown in Fig. 1, an embodiment 1 of the present application provides an intelligent voice interaction method for an exhibition hall, which comprises the following steps.

[0066] Step 1: A mobile sound source tracking module collects user sound source data and environmental data.

[0067] The user sound source data is collected by devices such as a microphone array, which is used for sound source positioning, and the environmental data is collected by devices such as a noise reduction sensor and a depth camera, which is used for understanding the environmental context.

[0068] Step 2: Whether mobile tracking of the sound source is needed is determined according to the user sound source and environmental data.

[0069] In some embodiments, a voice endpoint detection algorithm is used to analyze the stability of the sound source; whether the user presents a clear interaction intention is determined in combination with visual sensors such as a depth camera, and whether the system needs to be mobile tracked is determined according to whether the user is in a mobile state and the moving speed exceeds a speed threshold.

[0070] Step 3: When mobile tracking of the sound source is needed, the mobile sound source tracking module collects user visual data, locates and tracks the user based on the user sound source data and the user visual data, locates the sound source of the user, adjusts the posture of the robot according to the position change of the user, and when it is detected that the robot is blocked, re-plans a path to bypass the obstacle according to the azimuth angle of the sound source obtained by positioning to continue tracking the sound source of the user.

[0071] Optionally, locating the user based on the user sound source data and the user visual data comprises:

[0072] Rotating the rotatable microphone array so that the normal direction of the rotatable microphone array is aligned with the maximum signal-to-noise ratio azimuth angle;

[0073] Turning the depth camera to the maximum signal-to-noise ratio azimuth angle and outputting point cloud data;

[0074] Converting the maximum signal-to-noise ratio azimuth angle into an initial three-dimensional coordinate in a visual SLAM coordinate system based on a preset coordinate mapping function;

[0075] Fusion estimation is performed on the point cloud data and the initial three-dimensional coordinates to obtain three-dimensional coordinates of the user after fusion.

[0076] Specifically, the preset mapping function is represented by the following formula:

[0077]

[0078]

[0079]

[0080] wherein, is the maximum signal-to-noise ratio azimuth of the sound source measured by the microphone array, is the distance of the user estimated from the point cloud data, and the initial three-dimensional position of the user in the SLAM coordinate system is preliminarily calculated through conversion from the spherical coordinate system to the rectangular coordinate system . The maximum signal-to-noise ratio azimuth of the sound source measured by the microphone array is converted into the robot vision SLAM coordinate system through a preset coordinate mapping function, wherein is the horizontal angle, is the pitch angle. The preset coordinate mapping function is determined based on the relative installation position and the calibration parameters of the microphone array and the depth camera.

[0081] Optionally, fusion estimation is performed on the point cloud data and the initial three-dimensional coordinates according to the following formula to obtain the three-dimensional coordinates of the user after fusion:

[0082]

[0083]

[0084]

[0085] wherein, is the initial three-dimensional coordinate, is the user position coordinate directly observed through the point cloud data, is the three-dimensional coordinate of the user, is the weight of the initial three-dimensional coordinate , and is the weight of the user position coordinate directly observed through the point cloud data , and and sum to 1.

[0086] It can be understood that the visual data has a higher weight in position accuracy, and the acoustic data contributes more to the initial capture and tracking continuity, so in the embodiment, > The specific values of the above parameters can be set by the person skilled in the art according to actual conditions. , The specific values of the above parameters can be set by the person skilled in the art according to actual conditions.

[0087] As one of the prominent substantive features of the present application, by using the sound source azimuth angle and the point cloud data for fusion positioning, the acoustic positioning result is converted to the visual SLAM coordinate system through coordinate mapping, and the spatial geometry is calculated in combination with the depth information and the direction angle to obtain the three-dimensional coordinates and the orientation information of the target user. The fusion process makes up for the defects that the acoustic positioning is easily disturbed by environmental noise and the distance estimation deviation is large, and at the same time, the instantaneous error of visual positioning in a fast motion or occlusion scene is corrected, so that the coordinate and orientation data are more accurate, and the fast response capability of acoustic positioning and the high spatial precision of visual positioning are effectively combined, which not only improves the response speed of the positioning system, but also significantly improves the accuracy and robustness of position estimation in a complex environment. Moreover, it avoids the positioning jump or loss caused by the change of a single sensor environment including but not limited to sudden noise and light occlusion, and ensures continuous and stable tracking of the user during movement.

[0088] Optionally, the corresponding weights are dynamically adjusted according to the signal-to-noise ratio of the sound source and the density of the point cloud data obtained by the rotatable microphone array, and the specific method is as follows:

[0089] The signal-to-noise ratio of the sound source and the density of the point cloud data are both normalized and mapped to the interval of 0 to 1 to obtain the normalized signal-to-noise ratio of the sound source and the density of the point cloud data;

[0090] The difference between the normalized signal-to-noise ratio of the sound source and the density of the point cloud data is calculated; the ratio of the normalized signal-to-noise ratio of the sound source to the difference is determined as the weight of the initial three-dimensional position, and the ratio of the normalized density of the point cloud data to the difference is determined as the weight of the user position coordinates directly observed by the point cloud data.

[0091] In this embodiment, when the visual signal quality is high, that is, the point cloud density is large, the weight will increase, and the proportion of visual data in the fusion will increase; when the acoustic signal quality is high, that is, the signal-to-noise ratio of the sound source is large, the weight will increase, and the proportion of acoustic data in the fusion will increase, thereby achieving the purpose of dynamically adjusting the weight according to the signal quality.

[0092] Optionally, the method further comprises:

[0093] The direction angle of the center of the user's face is calculated through face detection, and the maximum signal-to-noise ratio azimuth angle and the direction angle of the center of the face are used as observation vectors to estimate the orientation state of the user's face to obtain the optimal direction angle of the center of the face.

[0094] The difference between the optimal face center direction angle and the face center direction angle in the current observation vector is calculated as the optimal compensation angle of the microphone array.

[0095] The gimbal rotation of the microphone array is adjusted according to the optimal compensation angle, so as to align with the user's face center direction, and the acoustic and visual cooperative alignment is realized.

[0096] Optionally, the user's face center direction angle is calculated through face detection, the maximum signal-to-noise ratio azimuth angle and the face center direction angle are taken as the observation vector, the user's face orientation state is estimated, and the optimal face center direction angle is obtained, including:

[0097] The state vector is set and the observation vector is , wherein represents the state of the user's face orientation at time k, represents the user's face center direction angle, represents the change rate of the face orientation angle, i.e. the angular velocity, is the maximum signal-to-noise ratio azimuth angle, is the face center direction angle calculated by the depth camera based on the face key points through face detection;

[0098] The state vector at the previous time is fused with the state transition matrix and the process noise to obtain the state at the current time, and the prior state estimation is obtained. The state estimation error covariance matrix at the previous time is added to the process noise covariance matrix to obtain the prior estimation error covariance matrix:

[0099]

[0100]

[0101] , wherein is the prior state estimation based on the k-1 time to the k time, is the true state of the system at the k-1 time, is the prior estimation error covariance matrix at the k time, representing the uncertainty measure of the predicted state , is the optimal state estimation at the previous time through the state transition matrix F, is the state transition matrix, is the sampling time interval, is the state estimation error covariance matrix, is the process noise covariance matrix, representing the uncertainty of state prediction, ~ represents the process noise, which is a normal distribution with 0 as the mean value and the covariance matrix .

[0102] The new observation vector is multiplied by the observation matrix, and the difference is subtracted from the prior state estimate. The difference is multiplied by the gain matrix, and the sum is obtained with the prior state estimate to obtain the posterior state estimate. The gain matrix is multiplied by the observation matrix, and the difference is subtracted from the unit matrix. The difference is multiplied by the prior estimation error covariance matrix to obtain the posterior estimation error covariance matrix:

[0103]

[0104]

[0105]

[0106]

[0107]

[0108] wherein, represents the observation matrix, the matrix that maps the state vector to the observation vector, represents the observation noise covariance matrix, representing the error of acoustic and visual positioning, represents the noise variance of acoustic positioning, represents the noise variance of visual positioning, and it can be understood that, is calibrated in advance, is the gain matrix, used for weighted fusion of predicted value and observed value. It can be understood that by the above formula gradually converges, thereby obtaining the optimal gain matrix, and further obtaining the optimal estimated face direction angle, thereby obtaining the optimal compensation angle.

[0109] When the error covariance matrix tends to a constant positive definite matrix, the iteration is stopped, and the optimal gain matrix is obtained, and further the optimal face center direction angle is obtained.

[0110] Optimal estimated face direction angle is given by the first component of the state vector , and the sound source positioning direction angle is the current acoustic observation value. The difference between the optimal estimated face direction angle and the current acoustic observation value is calculated as the optimal compensation angle of the microphone array. According to the optimal compensation angle, the gimbal rotation of the microphone array is adjusted to align with the user's face center direction, realizing the collaborative alignment of acoustic and visual.

[0111] As one of the prominent substantive features of the present application, the multi-modal data is fused and estimated by comparing the angle difference between the direction of the visual detected user face center and the direction obtained by sound source positioning, to obtain the optimal compensation angle, so as to realize the accurate alignment of the microphone array. The efficient cooperation of the microphone array and the depth camera improves the system cooperation capability, and further improves the overall efficiency of the system.

[0112] Step 4: When the sound source does not need to be tracked, the noise reduction sensor collects environmental data at a fixed position in the exhibition hall, and extracts features from the collected environmental data to obtain environmental sound features. The multi-modal data processing module fuses the environmental sound features, user-related data, and exhibit information data read by the RFID reader to generate a scene state feature vector. The user-related data includes user position and posture data obtained by the mobile sound source tracking module, user head orientation, user dwell time, and user speed.

[0113] It should be noted that the environmental sound features are obtained by the noise reduction sensor; the user position and posture data are obtained from the mobile sound source tracking module and the depth camera; the exhibit information data is obtained from the RFID tag and the exhibit knowledge base; and the head orientation is obtained from the visual SLAM and MediaPipe algorithms.

[0114] The modal data is normalized, denoised, and feature extracted; the modal feature vectors are aligned by timestamp and spliced into a high-dimensional feature vector; the weighting coefficients are set according to the confidence of each modal data to improve the fusion accuracy; when some modal data is abnormal or missing, a default model or cached data is used for replacement; and finally a scene state feature vector containing environmental state, user state, and exhibit state is generated for the adaptive voice interaction engine.

[0115] The multi-modal data processing module is deployed in the edge computing node in the robot body, and the edge computing node is a computing unit in the robot body. The data integration and analysis are performed in the robot body, which does not depend on cloud computing, avoiding network delay problems, thereby ensuring the real-time and low delay of voice interaction. In combination Figure 4 As shown, the generated scene vector is sent to the speech recognition module and the navigation module according to the data category.

[0116] Optionally, the method further comprises:

[0117] Real-time detection of noise level in the exhibition hall, adjustment of microphone array acquisition parameters according to different frequency noises, specifically:

[0118] adjusting the main lobe width according to the difference between the noise energy and the first preset energy threshold, and / or adding a zero point in the interference direction, and / or dynamically adjusting the suppression amount according to the noise energy, and / or adaptively adjusting the weight of each antenna of the microphone array by least mean square or recursive least square;

[0119] adjusting the main lobe width according to the difference between the noise energy and the second preset energy threshold, and / or dynamically adjusting the compensation amount according to the noise energy.

[0120] It should be noted that the main interference sound source direction in the environment is estimated by sound source positioning of the microphone array, and if the energy of a certain direction is significantly higher than the average background noise (such as more than 3dB), it is determined as the interference direction; when there is more low-frequency noise, the energy exceeds the second preset energy threshold without adding an additional zero point, and the default beam pattern is maintained.

[0121] Optionally, the main lobe width is adjusted according to the difference between the noise energy and the first preset energy threshold as follows:

[0122]

[0123]

[0124] In the formula, denotes the adjusted main lobe width, denotes the maximum width, which is used in a low-frequency dominant environment, and can be 60°, denotes the minimum width, which is used in a high-frequency severe interference environment, and can be 15°, denotes the noise energy in the first frequency range, denotes the first preset energy threshold, denotes a normalized adjustment factor.

[0125] It should be noted that the first frequency range is 3kHz to 8kHz, the second frequency range is 20Hz to 300Hz, the noise energy in the first frequency range is high-frequency noise, the noise energy in the second frequency range is low-frequency noise, the first preset energy threshold can be dynamically set according to the background energy of the environmental noise, for example, it can be set to 1.5 times the historical average energy, and the specific values of the first preset energy threshold and the second preset energy threshold can be calibrated according to the actual scene.

[0126] In this embodiment, as long as the high-frequency noise exceeds the first preset energy threshold, even if it exceeds a little, a certain adjustment effect can be produced, avoiding too weak adjustment.

[0127] Optionally, the suppression amount is dynamically adjusted according to the noise energy as follows:

[0128]

[0129] in the formula, represents the suppression amount, represents the noise energy in the first frequency band range, represents the first preset energy threshold value;

[0130] The compensation amount is dynamically adjusted according to the noise energy according to the following formula:

[0131]

[0132] in the formula, represents the suppression amount, represents the noise energy in the second frequency band range, represents the second preset energy threshold value.

[0133] In this embodiment, the spectrum analysis result is uploaded to the edge computing node through the noise reduction sensor, the node analyzes the current frequency band energy distribution, and selects the beam strategy according to the above-mentioned preset rule, and the beam forming parameter is issued to the microphone array controller, the controller adjusts the beam pattern in real time, realizes dynamic optimization, reduces the influence of high frequency interference on voice recognition, improves the clarity and directivity of the collected voice, and improves the capture ability of low frequency voice signal.

[0134] Further, when the RFID tag is lost, the appearance characteristics of the exhibits can be identified by the depth camera to infer information, the depth camera is used to collect the images of the exhibits, Mask R-CNN segmentation, ResNet feature extraction, and feature library matching are used to identify the exhibits; multi-view fusion and semantic auxiliary identification are combined to improve the accuracy.

[0135] The Azure Kinect DK collects RGB images and depth maps; extracts the image of the exhibit area in the user's line of sight direction; uses an instance segmentation model to segment the exhibits in the image pixel by pixel; extracts the boundary box and mask of the exhibit to realize individual identification of the exhibit; uses a CNN feature extraction network to encode the features of the segmented exhibit area, constructs an exhibit feature database, and stores the appearance feature vector of each exhibit in the exhibit feature database, including the exhibit ID, name, and description information; the matching degree of the current exhibit and the exhibit features in the feature database is calculated by cosine similarity or Euclidean distance; if the matching degree is higher than the set matching threshold 0.85, it is determined that the exhibit is correct. In this way, in the case of abnormal individual data, the system also has a perfect error correction mechanism, and the preliminary processing and subsequent depth fusion are organically combined to ensure the efficiency and reliability of multi-modal data processing.

[0136] Step 5: The adaptive voice interaction engine module parses the scene state feature vector output by the multi-modal data processing module, selects the optimal voice recognition model according to the current exhibition hall environment noise level, user visit behavior, user location, and user interest weight obtained by parsing, recognizes the user's question, generates a matching answer to the user's question based on the voice information obtained by recognition, the exhibition knowledge base, and the dynamic knowledge graph, and outputs the matching answer to the user.

[0137] In combination Figure 5 As shown in the figure, the adaptive voice interaction engine module is the core of the entire system, developed based on the open-source framework FunASR, and can dynamically adjust the voice recognition and feedback strategy according to the actual scene requirements. A rich exhibition-related information base is built in, which can instantly generate a matching answer according to the user's question content. For example, when the user asks about the development history of a robot, the system will provide detailed and personalized voice commentary in combination with the exhibition attributes and user visit progress.

[0138] Specifically, the exhibition hall environment noise level is calculated based on the equivalent sound pressure in a 200ms window of the noise reduction sensor, the user visit behavior includes in-depth visit or quick browsing, which is determined based on the residence time and moving speed, the user location is a three-dimensional coordinate in the exhibition coordinate system; the user interest weight is a dynamic interest weight matrix calculated based on the residence time, head orientation, and frequency of occurrence of keywords in interactive questions.

[0139] According to the voice information obtained by recognition, the exhibition knowledge base, and the dynamic knowledge graph, a matching answer to the user's question is generated and output to the user, including:

[0140] According to the user's question content, keywords are matched; BERT semantic similarity model is used to search for related entities in the dynamic knowledge graph; the most relevant content segment is selected in combination with the dynamic interest weight matrix.

[0141] The exhibition knowledge base includes exhibition ID, exhibition name, exhibition type attribute, development history, technical parameters, and application scenarios. The exhibition type attribute is the robot type, including live-line work robots and lifting robots. The entity nodes in the dynamic knowledge graph structure include exhibitions, people, events, and technical terms; the relationship edges include "belongs to", "is associated with", "is applied to", and "develops in".

[0142] The exhibition attributes, user interest weight, and current interactive context are input into the NLG model BART based on Transformer to output natural language commentary text, where the natural language commentary text includes an extended explanation mode and an abstract broadcast mode. The extended explanation mode outputs detailed content, including history, technology, and application, and the abstract broadcast outputs concise content, including function and purpose.

[0143] According to the current exhibition hall environment noise level, the environmental characteristics of the current exhibition hall environment are judged. Specifically, when Leq < 30 dB, it is determined that the current exhibition hall environment belongs to a quiet area, and the sensitivity parameter range of voice recognition is set to -30 dB~0 dB; when 30 dB ≤ Leq ≤ 60 dB, it is determined that the current exhibition hall environment belongs to a normal area, and the sensitivity parameter range of voice recognition is set to -10 dB~+10 dB; when Leq > 60 dB, it is determined that the current exhibition hall environment belongs to a noisy area, and the sensitivity parameter range of voice recognition is set to +10 dB~+30 dB; wherein, Leq represents the current exhibition hall environment noise level.

[0144] Further, the generated relevant voice interpretation content is provided to the user through the dual-mode output device equipped by the system, including:

[0145] When the user distance ≤ 1 m, the holder automatically pops out the bone conduction earphone, the positioning accuracy is ± 5 cm, and a private audio channel is established;

[0146] When the user distance > 1 m or multiple people interact, the directional speaker is enabled, the beam width is controlled within ± 15°, and the volume and direction are dynamically adjusted to ensure clear and audible voice; voice in multiple user scenarios is supported.

[0147] In combination with Figure 2 As shown in FIG. 2, the embodiment 2 of the present application provides an exhibition hall voice interaction system for executing the exhibition hall voice interaction method provided by the embodiment 1, and the system comprises:

[0148] The mobile sound source tracking module is used to collect environmental data and user sound source data of the exhibition hall, and to determine whether the sound source needs to be tracked according to the environmental data and the user sound source data; when the sound source needs to be tracked, the user visual data is collected, the user is tracked based on the user sound source data and the user visual data, and when it is detected that the robot is blocked, the path is re-planned according to the sound source azimuth angle obtained by positioning to bypass the obstacle to continue tracking the sound source of the user.

[0149] It can be understood that the robot posture not only includes the moving coordinates (x, y, z) of the whole robot, wherein x represents the lateral coordinate, y represents the longitudinal coordinate, and z represents the height, but also includes the rotation angle and the pitch angle of the holder, and the robot posture is adjusted to ensure that the microphone array and the camera can be aligned with the target.

[0150] The mobile sound source tracking module is integrated into the autonomous navigation robot platform, has flexible moving ability, and can actively approach the user and track the sound source of the user in real time. In a noisy exhibition hall environment, the module can effectively reduce the background noise interference and improve the quality of the voice signal by dynamically tracking the sound source of the user.

[0151] In combination withFigure 3 As shown, preferably but not limitedly, the moving sound source tracking module includes a rotatable microphone array and a depth camera, the user's sound source is located, and the robot posture is adjusted according to the change of the user's position, which includes:

[0152] When the rotatable microphone array detects that the sound pressure exceeds the sound pressure threshold, the rotatable microphone array rotates so that the normal direction of the rotatable microphone array is aligned with the maximum signal-to-noise ratio azimuth angle; the depth camera is turned to the same azimuth as the rotatable microphone array at the same time, and outputs point cloud data;

[0153] Based on the point cloud data, visual confirmation is performed; if the visual confirmation successfully locks the target, the visual dominant tracking mode is entered; otherwise, the user's sound source is tracked by using the rotatable microphone array, the three-dimensional coordinates and the orientation information of the target user are generated through acoustic positioning based on the maximum signal-to-noise ratio azimuth angle and the point cloud data;

[0154] When the rotatable microphone array detects that the sound pressure is less than or equal to the sound pressure threshold, the visual tracking mode is entered;

[0155] Wherein, in the visual tracking mode, the depth camera is used to adjust the robot posture according to the change of the user's position.

[0156] In some embodiments, the sound pressure threshold can be 60dB, and those skilled in the art can set the specific value of the sound pressure threshold according to the actual application scene.

[0157] In some embodiments, the autonomous navigation robot platform is based on a TurtleBot3 chassis carrier, uses a 48V lithium battery power supply system, and realizes omnidirectional movement function through double-wheel differential drive. The overall structure of the autonomous navigation robot platform is divided into two layers: the lower layer is the driving system and the computing unit, which is used to support motion control and data processing; the upper layer is a liftable gimbal structure, which integrates a ring-shaped 4-microphone array and an Azure Kinect DK depth camera at the top. Among them, the diameter of the ring-shaped 4-microphone array is 15cm, and the microphone array has a horizontal 270° rotation capability, while the depth camera can realize horizontal 360°, pitch ±45° full-range rotation, which can obtain the distance between each pixel point in the scene and the camera and output a depth map. The gimbal structure of the microphone array and the depth camera is controlled by two independent servo control systems, one controls the rotation of the microphone array, and the other controls the rotation of the camera, which work together to ensure accurate positioning and tracking in multiple degrees of freedom.

[0158] Specifically, under the ROS framework, the microphone array and the camera are two independent nodes, each controlling its own gimbal; the cooperation mechanism is realized by the upper layer control, for example: in the sound source initial positioning stage: the microphone dominates the rotation search; after visual locking: the camera takes over the tracking, and the microphone gimbal compensates the angle according to the visual feedback.

[0159] It should be noted that the liftable gimbal structure in the embodiments of the present disclosure is an electric mechanical type liftable gimbal, which is simple in structure, high in control precision, suitable for the ROS control framework, and improves the accuracy of tracking the user's voice and vision in the exhibition hall environment.

[0160] Specifically, the orientation information is estimated by analyzing the geometric features of the user's face in the point cloud data and the feature point direction under the visual SLAM framework.

[0161] Combined with the coarse-grained direction information provided by the sound source azimuth, the orientation vector calculated by vision is verified and fine-tuned, and finally the orientation angle of the user relative to the robot is output. Specifically:

[0162] According to the sound source azimuth, a coarse-grained direction range is determined, the acoustic azimuth is converted into an initial direction vector in the visual SLAM coordinate system, representing the user's general orientation based on acoustics, the included angle between the initial direction vector and the visual initial orientation vector is calculated, and if the included angle is not in the preset included angle range, the corresponding weight is combined for fine-tuning.

[0163] Preferably but not limitedly, the maximum signal-to-noise ratio azimuth is obtained by the following steps: the microphone array collects multi-channel voice: the TDOA (Time Difference of Arrival) algorithm is used to calculate the time difference of the user sound source arriving at the microphone of each channel; according to the geometric structure of the microphone array, a mapping relationship between direction and time difference is established; the signal is delayed and added in different directions to simulate beam scanning; the signal-to-noise ratio of the voice signal is calculated in different directions by using the beam forming algorithm, and the signal-to-noise ratio of the output signal of each direction is evaluated; the maximum signal-to-noise ratio direction angle is selected, that is, the general direction of the sound source.

[0164] It can be understood that the ratio of the user output signal power to the noise power is calculated, and the horizontal direction angle of the microphone array corresponding to the maximum ratio is calculated to obtain the maximum signal-to-noise ratio direction angle.

[0165] The present application combines sound source positioning technology and visual SLAM technology, so that the robot can automatically adjust its posture according to the position change of the user, and ensure that it is always in the best acquisition distance.

[0166] Preferably but not limitedly, the microphone array cooperates with the depth camera, specifically: after the microphone array rotates to search for the sound source direction of the user, the depth camera tracks the facial features of the user based on the MediaPipe framework face detection algorithm.

[0167] In some embodiments, the microphone array and the depth camera are configured as ROS nodes, and a motion control framework is constructed based on the ROS nodes, and the microphone array and the depth camera cooperate under the motion control framework. Specifically, in the initial positioning stage of the sound source direction, the microphone array dominates the rotation search; once the visual system successfully locks the target, the depth camera will take over the tracking task, continuously tracks the facial features of the user based on the MediaPipe framework face detection algorithm, and sends a compensation angle instruction to the microphone gimbal when it detects that the user's face orientation deviates or the relative position between the robot and the user changes.

[0168] Preferably but not limitedly, when the robot is detected to be blocked, the path is re-planned, including: using a double path planning mechanism, the main path generates the shortest approach route according to the sound source azimuth angle to obtain a first positioning result, and the auxiliary path generates a second positioning result based on the 3D semantic map constructed by the depth camera.

[0169] When the deviation between the first positioning result and the second positioning result exceeds the deviation threshold 0.5m, multi-modal sensor fusion calibration is performed, and the error of the first positioning result is corrected based on the second positioning result.

[0170] It should be noted that the visual positioning data is ultimately expressed based on the SLAM coordinate system, which is a global coordinate system set by the SLAM system at initialization. All user positions, obstacle positions, and robot poses detected by vision are expressed in three-dimensional coordinates (x, y, z) and directions (yaw, pitch, roll) in this coordinate system. The visual SLAM coordinate system is the spatial reference basis for all visual positioning data in the system, and is used to realize multi-target tracking, map construction, and path planning.

[0171] As one of the prominent substantive features of the present application, the present application performs multi-modal sensor fusion calibration when the deviation between the two positioning results exceeds the threshold, corrects the acoustic positioning error based on the visual positioning data, realizes adaptive path re-planning and multi-modal fusion calibration, and compared with the NAV2 navigation stack in the prior art, the present application introduces path evaluation and switching logic at the semantic level, and improves the autonomous decision-making ability of the robot in a dynamic environment.

[0172] When the deviation between the first positioning result and the second positioning result exceeds the deviation threshold 0.5m, multi-modal sensor fusion calibration is performed, and the error of the first positioning result is corrected based on the second positioning result, specifically including:

[0173] The acoustic positioning result including azimuth angle, distance and three-dimensional position of the visual SLAM output are fused using Kalman filter. The acoustic positioning is regarded as an observation, and the visual positioning is regarded as a high-precision reference quantity. The confidence weight of the acoustic positioning is dynamically adjusted through the filter. The fused optimal position estimation is output as the corrected acoustic positioning result.

[0174] In the visual SLAM system, if the robot has positioning drift, the system quickly recovers the global pose using image retrieval-based repositioning. The current image is matched with the key frame in the map to redetermine the accurate position of the robot in the SLAM coordinate system, which is used to correct the cumulative error of the acoustic positioning.

[0175] It should be noted that the image retrieval-based repositioning is a commonly used technology in SLAM or robot navigation, which is used to redetermine the current position in a known environment. The core idea is to find the most similar image by querying the current image and the reference image in the database, so as to estimate the current position and attitude. This application does not repeat it here.

[0176] In some embodiments, in a multi-user scenario, a deep clustering voiceprint separation scheme is adopted. By extracting GFCC features including but not limited to the duration and decibel size of the sound source, the system can effectively distinguish different speakers. In combination with the visual detection module to identify the mouth movement state, the effective sound source is further confirmed. By default, the system preferentially tracks the stable sound source with a continuous sound lasting more than 2 seconds; if there are competing sound sources, the tracking object is selected according to the principle of spatial proximity; when the target user stops speaking for more than 5 seconds, the system automatically switches to a pure visual tracking mode, and uses the gait recognition algorithm based on the OpenPose framework and the face orientation prediction to maintain tracking. If the visual loss time exceeds 10 seconds, the acoustic wake-up function is triggered, and a directional prompt sound is played to guide the user to re-establish the interaction.

[0177] As one of the prominent substantive features of the present application, in a single-user scenario, the acoustic optimal orientation is preferentially maintained; in a multi-user scenario, different sound sources are distinguished by voiceprint feature separation technology, and the continuous speaker is preferentially tracked. Compared with the prior art, the speech synthesis lacks scene adaptability, the output content is disconnected with the visiting route, and personalized interactive experience cannot be provided. The present application can adapt to different scenes to provide personalized interactive experience and improve user experience.

[0178] The multi-modal data processing module is used to fuse the environment data, user-related data and exhibit information data read by the RFID reader on the robot to generate a scene state feature vector when the sound source does not need to be moved.

[0179] An adaptive voice interaction engine is used to analyze the scene state feature vector, and select the optimal speech recognition model according to the current exhibition hall environment noise level, user visit behavior, user location, and user interest weight obtained by analysis, and then recognize the user's question. According to the voice information obtained by recognition, combined with the exhibit knowledge base and dynamic knowledge graph, a matched answer is generated for the user's question, and the matched answer is output to the user.

[0180] A distributed collaborative control platform is used to construct a device communication network using the MQTT protocol, connect robots and devices in the exhibition hall to the MQTT communication network, and distribute tasks and synchronize device states based on the MQTT communication network.

[0181] The distributed collaborative control module includes a central scheduler, which is the core of the entire network and is responsible for coordinating the tasks of mobile robot clusters, voice broadcast terminals, and environmental sensors. For example, when multiple users ask questions at the same time, the scheduler will reasonably allocate resources to ensure that each user can receive timely responses.

[0182] Specifically, a communication network based on the MQTT protocol is built. The distributed collaborative control module is the core management unit of the entire network, responsible for coordinating task allocation and state synchronization of various devices.

[0183] In combination Figure 5 As shown in the figure, the distributed collaborative control module is used to coordinate the task allocation and state synchronization of various devices in the exhibition hall, including:

[0184] (1) First, connect mobile robots, microphone arrays, noise reduction sensors, RFID readers, depth cameras, directional speakers, bone conduction headphones, central scheduling servers, edge computing nodes, network relay devices, fixed voice broadcast terminals, and environmental sensors to the MQTT communication network. Then, periodically send their state information to the distributed collaborative control module through the MQTT protocol.

[0185] The state information includes power, location, working state, etc.

[0186] The communication network is built through the MQTT protocol, which has lightweight characteristics and is particularly suitable for resource-constrained embedded devices. Using the publish / subscribe mode, such as 1:N message broadcasting, can meet the group communication of exhibition hall devices, and the MQTT protocol has QoS level guarantee and high bandwidth efficiency, saving about 70% of traffic compared to HTTP, improving data interaction security and transmission rate.

[0187] (2) When the user triggers an interaction event, based on the global coordinate synchronization of the multi-robot SLAM system, the service area is dynamically divided by Voronoi diagram, the robot closest to the user is selected as the main service robot, and the next two robots are selected as the auxiliary robots. The control instructions of the distributed collaborative control module are quickly received through the MQTT protocol network to interact with the user.

[0188] Specifically, the global coordinate synchronization based on the robot SLAM system is to realize the accurate positioning of the robot in the global coordinate system through the SLAM technology, with an accuracy of ±5 cm, to ensure the consistency of the robot's cognition of the environment and the positions of each other. The service area is divided based on the Voronoi diagram, which includes constructing a Voronoi diagram based on the real-time coordinates of the robot, dividing the exhibition hall space into multiple areas, each area corresponding to a robot, to ensure that each robot serves the nearest user. The dynamic division method specifically refers to updating the coordinates of each robot in real time, recalculating the Voronoi diagram, and dynamically adjusting the service area boundary when the user's position or the robot's state changes, to ensure the efficiency and balance of the service.

[0189] In a single-user scenario, the response time from voice triggering to the first response does not exceed 500 ms, including voice endpoint detection VAD, semantic analysis, and scheduling delay.

[0190] Preferably but not limitedly, in a multi-user concurrent scenario, a cluster collaboration mechanism is adopted, which allows one robot to serve up to three users, sorts and allocates tasks based on spatial distance, time priority, and permission level, and switches service objects. Specifically: spatial priority: preferentially serving the nearest user; time priority: preferentially serving the user who initiates the request earliest;

[0191] In the embodiments of the present disclosure, through the cluster collaboration mechanism and based on the spatial distance, time priority, and permission level, the tasks are sorted and allocated to avoid overloading a single robot and to avoid resource contention, so that the response interval does not exceed 800 ms, and each user can obtain timely and effective service.

[0192] Preferably but not limitedly, a robot is assigned to exclusively serve a high-attention user, which is a user who stays for more than a threshold of stay time or asks questions more than a threshold of question-asking times.

[0193] Preferably but not limitedly, the main service robot is selected according to the Euclidean distance and the angle between the user's face and the robot, and the auxiliary robot maintains a preset following distance of 3 m from the main service robot.

[0194] The main service robot selection criteria: prefer to select the nearest Euclidean distance; the smallest angle between the user's front direction and the robot, that is, the robot is within the user's field of view; if the distance is the same, prefer to select the smallest angle; if the angle is also the same, prefer to select the robot with higher power and lighter task load.

[0195] If the Euclidean distance of the two robots is the same and the user's front is not within the 120° range, select the smaller angle; if the angle is also the same, select the robot with higher power or lighter task load. If the Euclidean distances are not the same, the user's front is not within the 120° range, prefer to select the closer Euclidean distance; if the distances are similar, such as the difference is less than 0.5m, then further compare the angle and the load. Further, by using CBOR binary serialization for message compression, assigning independent VLANs for control instructions and guaranteeing bandwidth, deploying FPGA at the edge node to realize hardware offloading of the MQTT protocol stack, using PTPv2 precision clock protocol for clock synchronization, and deploying network probes to monitor and dynamically adjust the QoS level in real time, the system realizes a lightweight and efficient message delivery mechanism, which can achieve a response time of 50-300ms, improving real-time performance and millisecond-level response.

[0196] Further, when detecting that the communication delay exceeds the communication time threshold, a local cache mode is started to call a preset high-frequency question and answer knowledge base, trigger bandwidth preemption, suspend non-critical data transmission, and enable mesh self-organization when the central node of the distributed collaborative control module fails, establish an ad-hoc network through Wi-Fi Direct to realize mesh self-organization,

[0197] It should be noted that all devices that need to participate in networking, including mobile robots, fixed voice terminals, etc., need to support Wi-Fi Direct mode in order to quickly network when the distributed collaborative control module fails, which can be applied to temporary and fast recovery network requirements to ensure the continuity and stability of the service.

[0198] As one of the prominent substantive features of the present application, the present application builds a lightweight message middleware based on the MQTT protocol; designs a task scheduling and resource allocation mechanism, supports multi-device collaboration of voice recognition devices, robot bodies, display terminals, etc.; supports dynamic load balancing and failover mechanisms. Compared with the existing centralized control architecture, the multi-device collaborative response delay in the present application is reduced; the system supports multiple terminal device access and has strong scalability; the overall system running stability and fault tolerance capability are significantly improved.

[0199] In one embodiment, as Figure 2As shown, the whole system can be divided into an interaction layer, a processing layer and a perception layer, the mobile sound source tracking module is located in the perception layer, the multi-modal data processing module is located in the processing layer, the adaptive voice interaction module is located in the processing layer, and the distributed collaborative control platform collaboratively controls the modules and devices in the above three layers.

[0200] Embodiment 3 of the present application provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the computer program implements the exhibition hall intelligent voice interaction method of embodiment 1 when loaded into the processor.

[0201] Embodiment 4 of the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the exhibition hall intelligent voice interaction method according to embodiment 1.

[0202] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0203] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0204] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application rather than limit them, and although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered within the protection scope of the claims of the present application.

Claims

1. A method for voice interaction in an exhibition hall, characterized in that, The method comprises the following steps: Collecting environmental data and user sound source data of the exhibition hall; Judging whether it is necessary to move and track the sound source according to the environmental data and the user sound source data; When it is necessary to move and track the sound source, collecting visual data of the user, and positioning and tracking the user based on the user sound source data and the visual data, specifically: rotating the rotatable microphone array so that the normal direction of the rotatable microphone array is aligned with the maximum signal-to-noise ratio azimuth angle; the depth camera is turned to the maximum signal-to-noise ratio azimuth angle, and point cloud data is output; the maximum signal-to-noise ratio azimuth angle is converted into an initial three-dimensional coordinate in a visual SLAM coordinate system based on a preset coordinate mapping function; the point cloud data and the initial three-dimensional coordinate are fused and estimated to obtain a three-dimensional coordinate of the user after fusion; When it is not necessary to move and track the sound source, fusing the environmental data, user-related data and exhibit information data read by the RFID reader on the robot to generate a scene state feature vector; wherein the user-related data includes user position and attitude data, user head orientation, user stay time and user speed; Analyzing the scene state feature vector, and selecting an optimal speech recognition model according to the current exhibition hall environmental noise level, user visiting behavior, user position and user interest weight obtained by analysis, identifying the user's question, generating a matched answer for the user's question according to the speech information obtained by identification combined with the exhibit knowledge base and the dynamic knowledge graph, and outputting the matched answer to the user.

2. The exhibition hall voice interaction method of claim 1, wherein the point cloud data and the initial three-dimensional coordinate are fused and estimated according to the following formula to obtain a three-dimensional coordinate of the user after fusion:

3. The exhibition hall voice interaction method of claim 2, wherein the corresponding weights are dynamically adjusted according to the sound source signal-to-noise ratio obtained by the rotatable microphone array and the point cloud data density, specifically: wherein, is the initial three-dimensional coordinate, is the user position coordinate directly observed through the point cloud data, is the three-dimensional coordinate of the user, is the initial three-dimensional coordinate is the weight of the initial three-dimensional coordinate, is the weight of the user position coordinate directly observed through the point cloud data is the weight of the user position coordinate directly observed through the point cloud data and the sum of which is 1. The sound source signal-to-noise ratio and the point cloud data density are both normalized and mapped to the interval of 0 to 1 to obtain normalized sound source signal-to-noise ratio and point cloud data density; The difference value of the normalized sound source signal-to-noise ratio and the point cloud data density is calculated; the ratio of the normalized sound source signal-to-noise ratio to the difference value is determined as the weight of the initial three-dimensional position, and the ratio of the normalized point cloud data density to the difference value is determined as the weight of the user position coordinate directly observed by the point cloud data.

4. The exhibition hall voice interaction method of claim 1, wherein the method further comprises: calculating a face center direction angle of the user through face detection, and taking the maximum signal-to-noise ratio azimuth angle and the face center direction angle as an observation vector to estimate the face orientation state of the user to obtain an optimal face center direction angle; calculating the difference between the optimal face center direction angle and the face center direction angle in the current observation vector as the optimal compensation angle of the microphone array; adjusting the gimbal rotation of the microphone array according to the optimal compensation angle to align it with the face center direction of the user, realizing the collaborative alignment of acoustics and vision.

5. The exhibition hall voice interaction method of claim 4, wherein ​ ​ ​ The face center direction angle of the user is calculated through face detection, and the maximum signal-to-noise ratio azimuth angle and the face center direction angle are taken as observation vectors to estimate the face orientation state of the user, and the optimal face center direction angle comprises: Set state vector and observation vector wherein, represents the state of the user's face orientation at time k, represents the face center direction angle of the user, represents the change rate of the face orientation angle, i.e. the angular velocity, is the maximum signal-to-noise ratio azimuth angle, is the face center direction angle calculated by the depth camera based on the face key points through face detection. The state vector at the last moment is fused with the process noise through the state transition matrix to obtain the state at the current moment, and the prior state estimation error covariance matrix is obtained by adding the process noise covariance matrix to the state estimation error covariance matrix at the last moment; The new observation vector is multiplied by the observation matrix, and then the difference between the result and the prior state estimation is obtained, and the difference is multiplied by the gain matrix, and then the sum of the result and the prior state estimation is taken to obtain the posterior state estimation, and the gain matrix is multiplied by the observation matrix, and then the difference between the result and the unit matrix is obtained, and the difference is multiplied by the prior estimation error covariance matrix to obtain the posterior estimation error covariance matrix; When the error covariance matrix tends to be a constant positive definite matrix, the iteration is stopped, the optimal gain matrix is obtained, and then the optimal face center direction angle is obtained.

6. The exhibition hall voice interaction method of claim 1, wherein: The method further comprises: Real-time detection of noise level in the exhibition hall, adjustment of microphone array collection parameters according to noise of different frequencies, specifically: When the noise energy in the first frequency band exceeds the first preset energy threshold, the main lobe width is adjusted according to the difference between the noise energy and the first preset energy threshold, and / or, a zero point is added in the interference direction, and / or, the suppression amount is dynamically adjusted according to the noise energy, and / or, the weight of each antenna of the microphone array is adaptively adjusted by the least mean square or recursive least square; When the noise energy in the second frequency band exceeds the second preset energy threshold, the main lobe width is adjusted according to the difference between the noise energy and the second preset energy threshold, and / or, the compensation amount is dynamically adjusted according to the noise energy.

7. The exhibition hall voice interaction method of claim 6, wherein: The main lobe width is adjusted according to the difference between the noise energy and the first preset energy threshold as follows: In the formula, represents the adjusted main lobe width, represents the maximum width, represents the minimum width, represents the noise energy in the first frequency band range, represents the first preset energy threshold, represents the normalized adjustment factor.

8. The exhibition hall voice interaction method of claim 6, wherein: The suppression amount is dynamically adjusted according to the noise energy as follows: In the formula, represents the inhibition amount, represents the noise energy in the first frequency range, represents the first preset energy threshold; The compensation amount is dynamically adjusted according to the noise energy as follows: In the formula, represents the inhibition amount, represents the noise energy in the second frequency band range, represents the second preset energy threshold.

9. A showroom voice interaction system using the showroom voice interaction method according to any one of claims 1 to 8, characterized by, The system comprises: A mobile sound source tracking module for collecting environmental data and user sound source data of the exhibition hall, determining whether mobile tracking of the sound source is needed according to the environmental data and the user sound source data, collecting user visual data when mobile tracking of the sound source is needed, and positioning and tracking the user based on the user sound source data and the user visual data; A multi-modal data processing module for fusing the environmental data, user-related data and exhibit information data read by the RFID reader on the robot to generate a scene state feature vector when mobile tracking of the sound source is not needed. The adaptive voice interaction engine is used for analyzing the scene state feature vector, and selecting an optimal voice recognition model according to a current exhibition hall environment noise level, user visiting behavior, user position and user interest weight obtained through the analysis, recognizing a user question, generating a matched answer for the user question according to voice information obtained through the recognition and combining an exhibit knowledge base and a dynamic knowledge graph, and outputting the matched answer to the user. 10.The exhibition hall voice interaction system of claim 9, characterized in that: The system further comprises: A distributed cooperative control platform is used for constructing a device communication network by using an MQTT protocol, connecting the robot and devices in the exhibition hall to the MQTT communication network, distributing tasks of the devices and synchronizing states of the devices based on the MQTT communication network. 11.An electronic device comprising a processor and a storage medium; characterized in that: The storage medium is used for storing instructions; The processor is used for operating according to the instructions to perform steps of the exhibition hall voice interaction method according to any one of claims 1-8.

12. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement steps of the exhibition hall voice interaction method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Multi-mode man-machine interaction control system and method of humanoid robot

    CN116931728A

  • Exhibition hall robot visual language navigation method based on large model

    CN119309580A

  • Exhibition hall display method and device based on digital human interaction, equipment and medium

    CN120234404A