An XR interaction method, device and equipment based on voice control

By integrating multimodal sensor technology that combines voice signals and non-voice behavioral information, the problem of unstable interaction of XR devices in mobile environments has been solved, achieving high-precision and reliable voice control and improving the user experience.

CN120913566BActive Publication Date: 2026-03-27HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing XR devices lack precision, convenience, and robustness in human-computer interaction when users are moving or in complex environments. Traditional voice control struggles to accurately identify specific interface elements that users intend to operate, and dynamic changes in the coordinate system lead to unstable control.

Method used

By acquiring multimodal interaction data streams through multimodal sensors, fusing voice signals and non-voice behavior information, determining the target interaction space region, and dynamically adapting to changes in user and device coordinate systems, accurate identification and stable control can be achieved.

Benefits of technology

It improves the accuracy and stability of XR device interaction in mobile scenarios, reduces misoperation, expands the scope of application, and enhances the smoothness of user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913566B_ABST
    Figure CN120913566B_ABST
Patent Text Reader

Abstract

The application provides an XR interaction method, device and equipment based on voice control, and belongs to the technical field of human-computer interaction. The method comprises the following steps: receiving a multi-modal interaction data stream acquired by a multi-modal sensor; identifying and analyzing a voice signal in the multi-modal interaction data stream to determine a corresponding voice control instruction; wherein the voice control instruction comprises an operation intention and a space description; the space description is space constraint information of a user on a target operation object; based on the voice control instruction and one or more non-voice behavior information in the multi-modal interaction data stream synchronized with the voice signal, a corresponding target interaction space region is determined; according to the information of each operable object in the target interaction space region, a target operation object corresponding to the voice control instruction is matched, and an XR interaction operation corresponding to the operation intention is executed according to the target operation object. Thus, a voice control interaction with high precision, strong dynamic adaptability and high reliability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of human-computer interaction, and in particular to an XR interaction method and device based on voice control and equipment. BACKGROUND

[0002] Currently, there are still significant limitations in the human-computer interaction of extended reality (XR) devices: the mainstream interaction relies on gestures, eye movement tracking or physical controllers, and in the case of user movement or complex environment, the accuracy, convenience and robustness of these methods are greatly reduced.

[0003] Traditional voice control technology has more prominent problems: on the one hand, relying only on voice signal analysis to determine the operation intent lacks deep perception of user behavior and environmental context, making it difficult to accurately identify specific interface elements of user intent operation, resulting in frequent misoperations. On the other hand, when the user moves in the XR environment, the user's own coordinate system (a space reference system centered on the user) and the XR device interface coordinate system (a global space reference system defined by the device) change dynamically, so that spatial descriptions such as "front" and "right side" in the voice instruction cannot be accurately mapped to the device display space, causing coordinate drift, which seriously affects the stability and continuity of control. SUMMARY

[0004] The embodiments of the present application provide an XR interaction method, device and equipment based on voice control, which are used to solve the technical problem of how to realize high-precision, dynamically adaptive and high-reliability voice control interaction in an XR environment.

[0005] In one aspect, the embodiments of the present application provide an XR interaction method based on voice control, which comprises:

[0006] receiving a multi-modal interaction data stream obtained by a multi-modal sensor;

[0007] identifying and analyzing a voice signal in the multi-modal interaction data stream to determine a corresponding voice control instruction; wherein the voice control instruction includes an operation intent and a spatial description; the spatial description is spatial constraint information of a target operation object by a user;

[0008] determining a corresponding target interaction space region based on the voice control instruction and one or more non-voice behavior information in the multi-modal interaction data stream synchronized with the voice signal;

[0009] According to the information of each operable object in the target interaction space region, the target operation object corresponding to the voice control instruction is matched, and an XR interaction operation corresponding to the operation intent is performed according to the target operation object.

[0010] In an implementation form of the application, the target interaction space region is determined based on the voice control instruction and one or more non-voice behavior information in the multi-modal interaction data stream synchronized with the voice signal, and specifically comprises:

[0011] The non-voice behavior information is converted into a behavior feature vector in a three-dimensional space; wherein the non-voice behavior information at least includes one or more of the following: head orientation information, line-of-sight direction information, hand pointing information;

[0012] The space description in the voice control instruction is converted into a space constraint condition; the space constraint condition includes a constraint rule parameter group representing the space constraint information;

[0013] Based on the fusion calculation result of the behavior feature vector and the space constraint condition, the target interaction space region boundary in a three-dimensional space coordinate system is determined to obtain the target interaction space region according to the target interaction space region boundary.

[0014] In an implementation form of the application, the target interaction space region boundary in a three-dimensional space coordinate system is determined based on the fusion calculation result of the behavior feature vector and the space constraint condition, and specifically comprises:

[0015] The preset fusion weight group, the behavior feature vector and the reference vector corresponding to the space constraint condition are weighted and calculated to generate a target direction vector after fusion calculation;

[0016] The target direction vector corresponding to the direction is taken as a reference to generate the target interaction space region boundary according to the space constraint condition; wherein the space constraint condition at least includes a distance constraint parameter and a range constraint parameter.

[0017] In an implementation form of the application, before the preset fusion weight group, the behavior feature vector and the reference vector corresponding to the space constraint condition are weighted and calculated to generate a target direction vector after fusion calculation, the method further comprises:

[0018] Based on the motion state detection information from the multi-modal sensor, the motion state of the user is determined; the motion state includes a static state and a moving state;

[0019] When the motion state is a static state, a first preset fusion weight group is determined according to the information type of the non-voice behavior information, and weighted calculation is performed according to the first preset fusion weight group;

[0020] When the motion state is a moving state, a user motion intensity value is determined according to the motion state detection information, and the user motion intensity value is matched with a plurality of preset intensity threshold intervals.

[0021] According to the preset intensity threshold interval where the user motion intensity value is located, a corresponding dynamic weight adjustment factor is determined, and a baseline preset fusion weight set is loaded;

[0022] According to the dynamic weight adjustment factor, the baseline preset fusion weight set is corrected to obtain a corresponding second preset fusion weight set, and a corresponding weighted calculation is performed.

[0023] In an implementation manner of the present application, after determining the motion state of the user, the method further comprises:

[0024] In the case where the motion state is a moving state, a predicted user pose is determined based on a preset pose prediction model and the current motion state detection information;

[0025] According to the pose information and the position information corresponding to the predicted user pose, a dynamic mapping matrix from the user's own coordinate system to the XR device interface coordinate system is constructed; the pose information includes rotation parameters, and the position information includes translation parameters;

[0026] According to the dynamic mapping matrix, the space description is converted into coordinates in the XR device interface coordinate system corresponding to the predicted user pose, and the converted coordinates are updated to the space constraint condition.

[0027] In an implementation manner of the present application, the target operation object corresponding to the voice control instruction is matched according to the information of each operable object in the target interaction space region, specifically comprising:

[0028] According to each of the operable object information, the spatial coordinates and the functional attributes of each operable object are determined; the operable object is an XR device interface element;

[0029] According to the spatial coordinates, the functional attributes and the voice control instruction, the matching degrees of each of the operable objects and the voice control instruction are calculated; the matching degrees at least include spatial coordinate matching degrees and functional attribute matching degrees;

[0030] Each of the matching degrees is compared with a preset matching degree threshold value, so as to determine the target operation object according to the comparison result.

[0031] In an implementation manner of the present application, the target operation object is determined according to the comparison result, specifically comprising:

[0032] In the case where the comparison result determines that there are multiple pending operable objects, each of the pending operable objects is added to an XR interface display list and displayed to the user.

[0033] playing a voice prompt information so as to select a display label of each of the pending operable objects in the XR interface display list by the user;

[0034] determining the target operation object based on the secondary voice instruction of the user.

[0035] In an implementation manner of the present application, before performing the XR interaction operation corresponding to the operation intention according to the target operation object, the method further comprises:

[0036] generating a visual marking signal of the target operation object; wherein the visual marking signal comprises at least one or more of the following: a highlight frame, dynamic flashing, color gradient;

[0037] performing the XR interaction operation after receiving a confirmation instruction of the user; wherein the confirmation instruction comprises at least one of the following: a voice confirmation instruction, a preset gesture instruction or a preset action instruction.

[0038] In a second aspect, the embodiments of the present application further provide an XR interaction device based on voice control, which can perform the XR interaction method based on voice control described above; the device comprises:

[0039] a receiving module configured to receive a multi-modal interaction data stream acquired by a multi-modal sensor;

[0040] a first determining module configured to identify and analyze a voice signal in the multi-modal interaction data stream to determine a corresponding voice control instruction; wherein the voice control instruction comprises an operation intention and a spatial description; the spatial description is spatial constraint information of a target operation object by a user;

[0041] a second determining module configured to determine a corresponding target interaction space region based on the voice control instruction and one or more non-voice behavior information in the multi-modal interaction data stream synchronized with the voice signal;

[0042] a matching and performing module configured to match the target operation object corresponding to the voice control instruction according to the information of each operable object in the target interaction space region, and perform an XR interaction operation corresponding to the operation intention according to the target operation object.

[0043] In a third aspect, the embodiments of the present application further provide an XR interaction device based on voice control, which comprises:

[0044] at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the voice control-based XR interaction method described above.

[0045] Compared with the prior art, the present application has the following remarkable effects:

[0046] 1) Through the technical solution described above, the present application provides a voice control-based XR interaction control scheme, which fuses multi-modal interaction data streams to obtain a target interaction space region determined in combination with voice signals and non-voice behavior information, breaks through the limitation of traditional voice control relying only on semantic analysis, greatly reduces the misoperation caused by ambiguous space description, and realizes accurate identification of a target operation object. At the same time, the target interaction space region obtained by fusing user behavior information and space description in the voice control instruction is dynamically adapted to the relative change of the user's own coordinate system and the XR device interface coordinate system, effectively avoiding the coordinate drift problem when the user moves, ensuring that even in dynamic scenes such as walking and turning the head, the voice instruction can still be stably mapped to the target position of the device display space, significantly improving the continuity and reliability of the control.

[0047] 2) Through the cooperative processing of multi-modal information, the user can interact with the XR device through the intuitive way of "voice + natural behavior" without relying on additional controllers, which expands the application range of the XR device in mobile scenarios and improves the smoothness of the user experience. BRIEF DESCRIPTION OF DRAWINGS

[0048] The accompanying drawings described herein are used to provide further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:

[0049] Figure 1 An application scenario diagram of the voice control-based XR interaction method in the embodiments of the present application;

[0050] Figure 2 A flowchart of the voice control-based XR interaction method in the embodiments of the present application;

[0051] Figure 3 A flowchart of the method for determining a target interaction space region in the embodiments of the present application;

[0052] Figure 4 A flowchart of the method for determining a dynamic weight adjustment factor and performing weight correction in the embodiments of the present application;

[0053] Figure 5 A flowchart of an embodiment of the application for updating a spatial constraint condition;

[0054] Figure 6 A flowchart of an embodiment of the application for performing a workflow of XR device operation;

[0055] Figure 7 A structural diagram of an embodiment of the application for an XR interaction device based on voice control;

[0056] Figure 8 A structural diagram of an embodiment of the application for an XR interaction device based on voice control. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described below in detail with reference to the embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0058] The human-computer interaction mode of XR devices mainly relies on gestures, eye movement tracking or physical controllers. These interaction means perform well in specific scenarios, but their accuracy, convenience and robustness still face challenges in user movement or complex environments. In particular, traditional voice control technology often fails to accurately identify specific interface elements of user intent operations due to the lack of deep perception of user behavior and environmental context, which easily leads to misoperation. In addition, when the user moves in the XR environment, the dynamic changes between the device coordinate system and the user's own coordinate system make it difficult to accurately map the spatial description contained in the voice instruction, thereby affecting the stability and continuity of control.

[0059] Based on this, the embodiments of the present application provide a voice control-based XR interaction method, device and equipment to solve the technical problem of how to achieve high-precision, dynamically adaptive and high-reliability voice control interaction in an XR environment.

[0060] The various embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0061] The embodiments of the present application provide a voice control-based XR interaction method, which can be applied to, for example Figure 1In the application environment shown, the XR device of the present application can be set as a virtual reality (VR) glasses, augmented reality (AR) glasses, mixed reality (MR) glasses and the like. For convenience of description, the XR device is taken as an XR smart glasses in the embodiments of the present application for explanation and description. As shown in the figure, Figure 1 As shown, the application environment can include: a smart glasses 101 worn by a user 107, a communication network 102, a server 103, a voice semantic processing server 104, a multi-modal behavior data processing server 105, and a database 106. The communication network 102 can serve as a channel for data transmission, providing a communication link for communication between the smart glasses 101 and the server 103, so as to facilitate the smart glasses 101 to transmit the multi-modal interactive data stream such as various voice signals, visual angle information, image information and the like collected by the smart glasses 101 to the server 103 based on the communication network 102, and also to feedback the information processed by the server 103 to the smart glasses 101 for rendering display. The server 103 is a server that can provide various services, and in order to improve the processing efficiency, the server 103 is connected with the voice semantic processing server 104 and the multi-modal behavior data processing server 105, so as to be able to send the voice signal recognized to the voice semantic processing server 104 for processing after receiving the information collected by the smart glasses 101 based on the communication network 102; and send the non-voice behavior information to the multi-modal behavior data processing server 105 for processing. After the processing is completed, the server 103 determines the target operation object in combination with the voice signal detected by the smart glasses 101, and performs rendering display related processing on the target operation object, and finally feeds back to the smart glasses 101 based on the communication network 102 to present to the user.

[0062] The database 106 is connected with the voice semantic processing server 104 and the multi-modal behavior data processing server 105 respectively, and can use, store and manage the voice signal data of the voice semantic processing server 104 and the behavior data of the multi-modal behavior data processing server 105. The database 106 can be integrated on the server, or placed on the cloud or other network server.

[0063] The XR interaction method based on voice control provided by the embodiments of the present application will be described in detail below. Figure 2 The flowchart of the XR interaction method based on voice control provided by the embodiments of the present application is shown in the figure, Figure 2 As shown, the method can include steps S201-S204:

[0064] S201, receiving multi-modal interactive data stream acquired by a multi-modal sensor.

[0065] A cloud server is arranged, which can process the corresponding multi-modal data, and store voice control instructions, target interaction space regions, operable object information, preset fusion weight sets, preset pose prediction models and other related data through a database, and can interact with the XR glasses to realize corresponding data transmission and processing functions.

[0066] The multi-modal sensor is a combination of hardware that can synchronously collect multiple types of interaction information, and the multi-modal interaction data stream is a time-synchronized data set containing voice, behavior, motion and other dimensions. The multi-modal sensor can be arranged inside the smart glasses or externally electrically connected with the smart glasses, and includes but is not limited to an accelerometer, a gyroscope, a visual odometry, a visual sensor, an eye tracking sensor, a sound sensor, etc. The specific sensor types and quantities can be set during actual use, and are not specifically limited here.

[0067] The multi-modal sensor synchronously and continuously collects voice, vision (head orientation, line of sight direction), motion (displacement, acceleration, angular velocity) and gesture information of the user in real time, obtains multiple interaction data streams, and performs time synchronization processing on all the data streams to ensure the accuracy of information fusion.

[0068] S202, recognizing and analyzing the voice signal in the multi-modal interaction data stream to determine the corresponding voice control instruction.

[0069] The voice control instruction includes an operation intention and a space description. The space description is the spatial constraint information of the user on the target operation object.

[0070] The voice signal in the multi-modal interaction data stream can be recognized and analyzed by a voice processing module composed of an automatic speech recognition (ASR) unit and a natural language understanding (NLU) unit. The ASR converts the voice into text, and the NLU extracts the core operation intention (such as “open”, “select”, “close”, “adjust size”) and any explicit space description (such as “front”, “right side”, “this”, “that”, “third”) from the text. For example, when the user says “open the second application on the right”, the system will identify that the operation intention is “open”, the space description is “the second on the right”, and the target type is “application”.

[0071] The above space description is the space constraint information of the target operation object directly expressed by the user through the voice instruction, which is the explicit limitation of the user on the spatial attributes such as the position and range of the target object based on the self-cognition. For example, "1 meter on the right side" is the space limitation directly expressed by the user through the voice, which includes the direction constraint "right side" (limiting the target to the right side of the user) and the distance constraint "1 meter (m)" (limiting the distance range of the target from the user). Meanwhile, the application can also quantify the space description and convert the space constraint information directly expressed by the user into machine-identifiable parameters, for example, "right side" corresponds to the positive direction of the Y-axis of the user coordinate system (perpendicular to the right of the head orientation), and "1 meter" corresponds to the distance constraint parameter d = 1.0 m (allowing an error of ± 0.2 m, which reflects the user's tolerance range for distance accuracy).

[0072] In addition, the application also sets an instruction classifier, which can preliminarily classify the above obtained voice control instruction to quickly execute the corresponding control operation.

[0073] S203, determining a corresponding target interaction space region based on the voice control instruction and one or more non-voice behavior information in the multi-modal interaction data stream synchronized with the voice signal.

[0074] The application also analyzes the non-voice behavior information synchronized with the voice signal while determining the voice control instruction, and determines the target interaction space region in combination with the voice control instruction. Thus, in combination with the voice control instruction pointing to the XR device interface coordinate system and the non-voice behavior information directly associated with the user's own coordinate system state, the voice control instruction and the non-voice control instruction are synchronously fused to construct the target interaction space region, thereby realizing the dynamic association of the user's own coordinate system and the XR device interface coordinate system to realize the stability and continuity of the interaction control.

[0075] In the embodiment of the application, the above determination of the corresponding target interaction space region based on the voice control instruction and one or more non-voice behavior information in the multi-modal interaction data stream synchronized with the voice signal includes the following steps: Figure 3 As shown in the figure, the steps include:

[0076] S301, converting the non-voice behavior information into a behavior feature vector in a three-dimensional space. The non-voice behavior information includes at least one or more of the following: head orientation information, line-of-sight direction information, and hand pointing information.

[0077] S302, converting the space description in the voice control instruction into a space constraint condition. The space constraint condition includes a constraint rule parameter group representing the space constraint information.

[0078] S303, based on the fusion calculation result of the behavior feature vector and the space constraint condition, determine the target interaction space region boundary in the three-dimensional space coordinate system, to obtain the target interaction space region according to the target interaction space region boundary.

[0079] Specifically, the application can obtain the head posture data of the user in real time through the inertial measurement unit (IMU) and the visual sensor built in the intelligent glasses, obtain the head orientation information, and convert it into a head orientation vector in the XR device interface coordinate system. Using the eye tracking sensor, the user's eye movement is accurately captured to obtain the gaze direction information, and the gaze direction vector of the user in the XR device interface coordinate system is calculated. This usually involves analysis and geometric mapping of eye images. Through the depth sensor or gesture recognition algorithm in the server, the key points of the user's hand are detected in real time, and the vector of the user's hand pointing is calculated according to the gesture (such as pointing with the index finger). For non-directional gestures (such as grabbing, pinching), they are used as auxiliary confirmation or operation instructions. After obtaining the vector corresponding to one or more non-speech behavior information, the application constructs a behavior feature vector. That is, the non-speech behavior information is obtained based on the user's body movement, eye movement and body posture features.

[0080] Further, the space description is converted into a space constraint condition to generate a constraint rule parameter group, which directly corresponds to the space constraint information in the user's voice and can include: direction constraint: based on the user's "right side" to determine the reference direction vector S=(0, 1, 0); Distance constraint: based on the user's "1 meter" to set d=1.0m (allowing an error of ±0.2m to match the user's habit of ambiguous expression of distance); Range constraint: combine common interaction scenarios to quantify the user's implicit "nearby area" as a cubic region (length x width x height=0.5m x 0.5m x 0.5m) centered on the target point.

[0081] Subsequently, the behavior feature vector and the space constraint condition are fused to obtain a fusion calculation result. The fusion calculation can be weighted average, Bayesian network, or attention mechanism based on deep learning, and the present application does not make specific limitations. After fusion calculation, one or more target space interaction regions can be defined in the XR virtual space. For example, if the voice instruction is "open this", the system will give priority to the area where the user's visual focus is located; if the instruction is "right", a fan-shaped or conical region is determined in the user's right field of view in combination with the head orientation and the visual direction. If the voice instruction contains a mention of a specific type of element (such as "open application" or "select file"), the range of the region is further narrowed down to only consider the type of interactive elements. This multi-modal fusion mechanism significantly improves the context understanding ability of voice instructions and the accuracy of target positioning, effectively avoiding misrecognition caused by semantic ambiguity or complex environment in traditional voice control.

[0082] Based on the above scheme, the present application fuses voice and non-voice behavior information to calculate, constructs a target interaction space region relying on data constructed in two different dimensions, and the two complement each other to avoid relying on voice or behavior information alone during interaction.

[0083] Further, in some embodiments of the present application, the fusion calculation result based on the behavior feature vector and the space constraint condition determines the target interaction space region boundary in the three-dimensional space coordinate system, specifically including:

[0084] The preset fusion weight set, the behavior feature vector, and the reference vector corresponding to the space constraint condition are weighted to generate a target direction vector after fusion calculation (i.e., the fusion calculation result). Taking the direction corresponding to the target direction vector as the reference, the target interaction space region boundary is generated according to the space constraint condition. The space constraint condition at least includes a distance constraint parameter and a range constraint parameter.

[0085] In other words, the present application presets the fusion weight set, that is, the initial weights can be assigned to different modalities, such as a speech semantic weight 0.4, a line-of-sight direction weight 0.3, a head orientation weight 0.2, and a hand pointing weight 0.1. The head orientation vector H (x1, y1, z1), the line-of-sight direction vector E (x2, y2, z2), and G (x3, y3, z3) in the behavior feature vector are respectively multiplied by the corresponding line-of-sight direction weight 0.3, the head orientation weight 0.2, and the hand pointing weight 0.1 to obtain a first parameter; a reference vector is extracted according to the spatial constraint condition, which can include its {direction angle range, distance threshold, region shape parameter}, and the product value of the reference vector and the speech semantic weight 0.4 is calculated to obtain a second parameter; the sum of the first parameter and the second parameter is taken as the target direction vector after fusion calculation. Subsequently, the target direction vector is taken as a reference to determine the target interaction space region boundary in combination with the direction constraint, distance constraint, and range constraint in the spatial constraint condition. For example, a conical region with the target direction vector corresponding direction as the center, an angle range of ±20°, and a distance of 20-50 cm from the user is taken as the target interaction space region boundary.

[0086] Further, the behavior feature vector and the corresponding spatial constraint condition of the speech are flexibly fused to demarcate the target interaction space region boundary, so as to realize accurate positioning in the three-dimensional space.

[0087] In addition, in some embodiments of the present application, before the preset fusion weight set, the behavior feature vector, and the reference vector corresponding to the spatial constraint condition are weighted and calculated to generate the target direction vector after fusion calculation, as shown in Figure 4 The present application also provides the following embodiments, and the specific steps are as follows:

[0088] S401, determining the motion state of the user based on the motion state detection information from the multi-modal sensor. The motion state includes a static state and a moving state;

[0089] S402, when the motion state is the static state, determining a first preset fusion weight set according to the information type of the non-speech behavior information, and performing weighted calculation according to the first preset fusion weight set;

[0090] S403, when the motion state is the moving state, determining a user motion intensity value according to the motion state detection information, and matching the user motion intensity value with a plurality of preset intensity threshold intervals;

[0091] S404, determining a corresponding dynamic weight adjustment factor according to the preset intensity threshold interval in which the user motion intensity value is located, and loading a reference preset fusion weight set;

[0092] S405, according to the dynamic weight adjustment factor, the reference preset fusion weight set is corrected to obtain a corresponding second preset fusion weight set, and corresponding weighted calculation is performed.

[0093] In other words, the application will further determine whether the user is in the process of motion, which can specifically continuously obtain the motion state detection information of the user through the IMU data, and determine whether the user is in a stationary, uniform motion, accelerated motion, rotation, etc. state by analyzing the amplitude and rate of change of acceleration and angular velocity. For example, when the acceleration and angular velocity are both lower than a preset threshold, it is determined to be a stationary state; otherwise, it is a moving state. The preset threshold can be set by the user or the developer according to expert experience, which is not specifically limited here. When it is detected that the user is in a stationary state, the application can use the above-mentioned default initial: voice semantic weight 0.4, line of sight direction weight 0.3, head orientation weight 0.2, and hand pointing weight 0.1 as the first preset fusion weight set, in order to execute the above-mentioned S303. When it is detected that the user is in a moving state, the application will start the adaptive attention adjustment mechanism, intelligently dynamically adjust the weights of each modality according to the user motion intensity, and effectively avoid the misoperation caused by unstable sensor data or user unconscious action. The modality weights dynamically adjusted according to the user motion intensity are located in the reference preset fusion weight set, and the weights in the reference preset fusion weight set can be the same as the weight values in the above-mentioned first preset fusion weight set, or can be set by the user or the developer in the actual scene based on expert experience, which is not specifically limited here.

[0094] Specifically, the application can use real-time acceleration amplitude as user motion intensity value, or other data as user motion intensity value, such as combining real-time acceleration amplitude, speed, heart rate, etc. to obtain user motion intensity value by weighted calculation, which can be set by the user or the developer according to expert experience, which is not specifically limited here. The application takes real-time acceleration amplitude as user motion intensity value as an example, defines the real-time acceleration amplitude as accel_magnitude, and sets multiple motion intensity thresholds, such as: Threshold_low_accel (light motion threshold, such as 1.5 m / s 2 (meters per square second)): below this value, the user is considered to be in a relatively stable state. Threshold_high_accel (intense motion threshold, such as 4.0 m / s 2): above this value, the user is considered to be in a state of intense exercise. The preset intensity threshold interval includes: (0, Threshold_low_accel), [Threshold_low_accel, Threshold_high_accel], (Threshold_high_accel, Threshold_high_accel2). Threshold_high_accel2 is an upper limit of the intensity threshold set by the user or the developer, which can be obtained by expert experience and is not specifically limited herein.

[0095] In the determination of the dynamic weight adjustment factor adjust ment_factor, the specific determination strategy can be:

[0096] 1. When accel_magnitude is in the first interval (0, Threshold_low_accel), adjust ment_factor is 1.0, indicating that no adjustment is performed, and the weight of the non-speech behavior mode remains the original value.

[0097] 2. When accel_magnitude is in the second interval [Threshold_low_accel, Threshold_high_accel], adjust ment_factor is linearly interpolated, and decreases linearly with the increase of the user's exercise intensity value. For example, adjust ment_factor = 1.0 - 0.8 * (accel_magnitude - Threshold_low_accel) / (Threshold_high_accel - Threshold_low_accel).

[0098] 3. When accel_magnitude is in the third interval (Threshold_high_accel, Threshold_high_accel2), adjust ment_factor is 0.2 (or lower), indicating that the non-speech behavior mode is greatly reduced, and at this time more reliance will be placed on speech semantic information.

[0099] For the adjustment strategy of the dynamic weight adjustment factor, the above is only an exemplary existence, and the specific adjustment strategy can be set by the user or the developer according to expert experience or scene, and is not specifically limited herein.

[0100] After obtaining the dynamic weight adjustment factor, the behavior-related weight in the preset fusion weight set can be multiplied by the dynamic weight adjustment factor to obtain an adjusted behavior weight, and at the same time, the weight of speech semantics can be correspondingly increased to compensate for the part of the behavior weight that is reduced, so that the total weight of all modalities remains unchanged (for example, ensuring that the sum of the weights in the preset fusion weight set is 1), to obtain a second preset fusion weight set. Thus, the part of the behavior weight that is reduced is allocated to the speech semantics weight, so that it occupies a dominant position when the user is in a violent motion. The adjusted weight is applied to the S303 step in real time.

[0101] In addition, when the accel_magnitude is greater than Threshold_high_accel2, at this time the user can be in an extreme violent motion state, the dependence on accurate spatial positioning can be temporarily reduced, or the user can be prompted through vision / speech that "your current motion amplitude is large, please try voice control again after stabilizing", or even temporarily suspend the voice control function, to avoid misoperation caused by sensor data jitter or user unconscious action. Through the above adaptive attention adjustment mechanism, the application ensures that the system can maintain high robustness and stability in various complex motion scenarios, and provides a safer and more reliable interactive experience for the user.

[0102] Through the above scheme, the weighted fusion algorithm can dynamically adjust the weight of the behavior feature vector (such as reducing the hand pointing weight when the motion is violent), solve the problem of interference of the user motion state on the sensor data (such as pointing deviation caused by hand shaking when running), and dynamically adapt the weight to ensure that the system can maintain stable performance in static, moving, and violent motion scenarios, avoid misoperation caused by environmental interference, and enhance the robustness of the system.

[0103] In some embodiments of the application, the user is in a relatively violent motion process, which can make the coordinate drift more serious, so after determining the motion state of the user, as shown in Figure 5 The application also provides the following embodiments, specifically including steps S501-S503:

[0104] S501, in the case where the motion state is a moving state, determining a predicted user pose based on a preset pose prediction model and current motion state detection information;

[0105] S502, constructing a dynamic mapping matrix from the user's own coordinate system to the XR device interface coordinate system according to the attitude information and the position information corresponding to the predicted user pose; the attitude information includes rotation parameters, and the position information includes translation parameters;

[0106] S503, convert the space description into coordinates in the XR device interface coordinate system corresponding to the predicted user pose according to the dynamic mapping matrix, and update the converted coordinates to the spatial constraint condition.

[0107] That is, the present application predicts the user's pose using a preset pose prediction model when the user moves, and constructs a dynamic mapping matrix that can perform coordinate mapping, so as to ensure that the user's expressed space description is accurately mapped in the XR device interface coordinate system, avoiding coordinate drift.

[0108] In a popular way, during the voice instruction processing process, the present application can continuously and real-time detect the user's motion state, including displacement speed, acceleration and other parameters. Based on these motion data, when it is detected that the user is in a moving state, the system starts the preset pose prediction model to predict the user's spatial pose (position and attitude) in a very short future time based on the user's accurate pose at the current time (provided by Simultaneous Localization and Mapping (SLAM) or Visual-Inertial Odometry (VIO)), real-time speed, acceleration and angular velocity and other kinematic parameters, to obtain the predicted user pose. When the voice instruction contains spatial description words such as "front", "right side", "upper side", etc., the system will map these relative spatial descriptions to the XR device interface coordinate system corresponding to the predicted user pose. This means that even if the user is walking, turning his head or making other body movements, the system can accurately understand and execute the instruction according to his predicted future position and orientation, ensuring the stability and continuity of voice control, greatly improving the interaction experience in a moving scene.

[0109] The above-mentioned preset pose prediction model can adopt, for example: Kalman Filter or Extended Kalman Filter (EKF): suitable for linear or weakly nonlinear systems, which can optimally estimate and predict the user's pose by fusing IMU data and visual odometry data. For example: Unscented Kalman Filter (UKF) or Particle Filter: suitable for strongly nonlinear systems, which can more accurately handle pose prediction under complex user motion patterns. For example: deep learning-based sequence prediction model: such as Recurrent Neural Network (RNN), Long Short-Term Memory Network or Transformer model, which can predict the future pose sequence by learning a large amount of user motion data, especially suitable for predicting user intentional motion. The specific optional model can be set by the developer based on expert experience or use scenarios, which is not limited here.

[0110] Subsequently, based on the predicted user pose corresponding attitude information and position information, a dynamic mapping matrix from the user's own coordinate system to the XR device interface coordinate system is constructed, the attitude information includes rotation parameters, and the position information includes translation parameters. The dynamic mapping matrix is used to realize the real-time conversion of three-dimensional coordinates from the user's own coordinate system to the XR device interface coordinate system.

[0111] When the spatial description in the voice control instruction contains direction words relative to the user's own, the dynamic mapping matrix is used to convert the spatial description into the coordinates in the XR device display space corresponding to the predicted user pose. The direction words include "front", "right", "up", and "left front". The coordinates include absolute coordinates or relative coordinates. This means that even if the user issues a "front" instruction while walking, turning his head, or making other body movements, the system can accurately point to the user's predicted future front, rather than the front at the time of instruction issuance, thereby completely eliminating coordinate drift caused by user movement and ensuring the stability and continuity of voice control.

[0112] It should be noted that the above-mentioned user's own coordinate system is a three-dimensional coordinate system established with a specific part of the user's body as the origin, and the specific part of the body includes the head or the torso, and the coordinate axis direction corresponds to the front, back, left, right, and up directions of the user's own. The XR device interface coordinate system is a three-dimensional global coordinate system established by the XR device (smart glasses) for positioning interface elements, and its origin and coordinate axis direction correspond to the physical space or virtual space perceived by the XR device, and is used to uniformly describe the spatial positions of all interface elements. In the user's own coordinate system, the direction words relative to the user's own correspond to specific coordinate axis directions. "Front" corresponds to the front coordinate axis of the user's own coordinate system, "right" corresponds to the right coordinate axis of the user's own coordinate system, and "up" corresponds to the up coordinate axis of the user's own coordinate system.

[0113] Through the above scheme, the dynamic mapping matrix constructed in the user movement state accurately converts the relative position in the user's own coordinate system into the absolute position in the XR device interface coordinate system, ensuring that even in dynamic scenarios such as walking and turning the head, the spatial description in the voice instruction can be accurately mapped to the corresponding virtual interface element, avoiding operation deviation caused by coordinate drift.

[0114] S204, according to the information of each operable object in the target interaction space region, matching the target operation object corresponding to the voice control instruction, and performing the XR interaction operation corresponding to the operation intention according to the target operation object.

[0115] The operable object in the present application can be understood as an XR interface element in a display interface corresponding to an XR device interface coordinate system, such as a button, a menu, an information card, a 3D model, etc., and the operable object information is structured data used to describe the relevant attributes of these XR interface elements in the XR device interface coordinate system, such as spatial attribute information, functional attribute information, state attribute information, and identification attribute information. The user can operate the XR interface element through voice or gesture operations. The XR interaction operation, for example, performs specific operations such as “opening an application” and “selecting a file”.

[0116] In the embodiments of the present application, the above-mentioned matching of the target operation object corresponding to the voice control instruction according to the information of each operable object in the target interaction space region specifically includes:

[0117] According to the information of each operable object, the spatial coordinates and functional attributes of each operable object are determined. The operable object is an XR device interface element. According to the spatial coordinates, the functional attributes, and the voice control instruction, the matching degrees of each operable object and the voice control instruction are calculated. The matching degrees at least include spatial coordinate matching degrees and functional attribute matching degrees. Each matching degree is compared with a preset matching degree threshold value to determine the target operation object according to the comparison result.

[0118] That is, the present application obtains the spatial coordinates and functional attributes of each operable object from the operable object information, and then calculates the matching degrees for each operable object with the operation intention and spatial description of the voice control instruction. The matching degrees can include spatial coordinate matching degrees, which can be calculated by calculating the cosine similarity, Euclidean distance, etc., to determine whether the spatial coordinates of the operable object are in the target interaction space region; and functional attribute matching degrees, which can be obtained by calculating the cosine similarity of the text semantics between the operation intention and the functional attributes, so as to quantify the degree of coincidence between the functional attributes of the element and the operation intention of the voice instruction. Among them, the spatial coordinate matching degree can further include geometric matching degree and spatial matching degree. The geometric matching degree: the alignment degree of the position and size of the element with the fused attention vector (head, line of sight, hand pointing direction), for example, the smaller the included angle between the element center and the line of sight direction vector, the higher the matching degree; the spatial matching degree: whether the element meets the spatial description in the voice instruction (such as “front”, “second from the right”).

[0119] The matching degrees are obtained by comprehensively considering the above-mentioned spatial coordinate matching degrees and functional attribute matching degrees, and then the matching degrees are compared with a preset matching degree threshold value. The operable object with a matching degree greater than the preset matching degree threshold value is taken as the target operation object. The preset matching degree threshold value can be set by the user, which is not limited in the present application. In addition, if there are multiple operable objects with matching degrees greater than the preset matching degree threshold value, the operable object corresponding to the maximum value in the matching degrees can be taken as the target operation object.

[0120] When each matching degree is less than the preset matching degree threshold, it can be determined that the recognition fails, at this time, a voice prompt such as "I'm sorry, I didn't understand your instruction, please say it again" can be given or a candidate list can be displayed on the XR interface to guide the user to reissue the instruction or make a selection. If there are multiple operable objects with a matching degree greater than the preset matching degree threshold, and the multiple matching degrees are similar, at this time, multi-target conflict processing is entered.

[0121] Through the multi-dimensional matching degree calculation of the spatial coordinates and the functional attributes, multi-dimensional verification of the target operation object can be realized, the matching degree threshold can be set, and the optimal object can be selected, so that the error selection caused by single-dimensional matching is avoided, and the accuracy of target matching is improved.

[0122] In some embodiments of the present application, in the case of multi-target conflict processing, the present application determines the target operation object according to the comparison result, specifically including:

[0123] In the case where it is determined according to the comparison result that there are multiple pending operable objects, each pending operable object is added to the XR interface display list and displayed to the user. Voice prompt information is played to allow the user to select the display label of each pending operable object in the XR interface display list. Based on the secondary voice instruction of the user, the target operation object is determined.

[0124] In other words, the present application provides a secondary selection function for the user to explicitly select the target operation object. For example, the XR interface display list displays the display labels of multiple pending operable objects, and the display label can be a number assigned to each pending operable object. The secondary voice instruction of the user can include the display label, such as "number 1", "number 2", etc. The present application can also present the display labels to the user in the form of thumbnails, and the voice prompt information can be "there are multiple matching items, please say the number you want to select".

[0125] The present application addresses the multi-target conflict scenario, guides the user to make a secondary selection through list display + voice prompt, solves the pain point that the traditional multi-target conflict cannot be decided, and through the explicit target of the secondary voice instruction, avoids random selection errors caused by close matching degrees, improves the processing capability of the system for complex interface scenarios, and ensures the determinacy of interaction.

[0126] In an embodiment of the present application, if the target operation object is matched, a direct interaction operation can be performed, which can cause a misoperation, especially in a motion state. Therefore, before performing the XR interaction operation corresponding to the operation intention according to the target operation object, the following embodiments are provided, including:

[0127] A visual feedback signal is generated to indicate the target operation object. The visual feedback signal includes at least one of the following: a highlighted border, a dynamic flashing, a color gradient. Upon receiving a confirmation instruction from the user, the XR interaction operation is executed. The confirmation instruction includes at least one of the following: a voice confirmation instruction, a preset gesture instruction, or a preset action instruction.

[0128] In other words, after the multimodal target positioning technology successfully determines the unique or most likely XR interface element controlled by the user's intention, the system does not immediately execute the operation. The intelligent feedback generation unit of the feedback control module immediately generates and displays a clear visual feedback signal to intuitively inform the user of the target recognized by the system and provide the user with a confirmation opportunity. These visual feedback methods include but are not limited to: highlight display: significantly improving the color, brightness or saturation of the target interface element to make it stand out in the XR environment; border flashing / outline: adding a dynamic flashing or continuous outline around the target interface element; color / transparency change: changing the color of the target element or making its transparency periodically change; zooming / floating effect: slightly enlarging or slightly floating the target element in space to attract the user's attention.

[0129] The duration of the visual feedback can be set to 1.5 to 2.5 seconds to ensure that the user has enough time to observe and understand the feedback information. During the display of the visual feedback, the system waits for the user to confirm. The user can confirm in various ways, increasing the flexibility of the interaction: such as issuing a short voice command again: for example, the user can say "confirm", "yes", "execute", etc. For example, a specific non-voice action: the user can perform a preset nodding action (identified by a head pose sensor) or a specific gesture (such as the "OK" gesture).

[0130] In addition, in some scenarios (such as when the user has explicitly expressed a preference), if the user does not perform any other operation or issue a cancellation instruction within a preset short period of time (for example, 3 seconds) after the visual feedback is displayed, the system can automatically consider the operation as a confirmation and execute it. Once the user's confirmation signal is received, the instruction execution unit sends the final control instruction to the XR device, triggering the corresponding interface operation (such as opening an application, selecting a menu item, adjusting a parameter, submitting a form, etc.).

[0131] If the user finds that the system has made a mistake and is not the target operation object that the user wants to operate after the visual feedback is displayed, the user can immediately issue a cancellation instruction (such as a voice control instruction of "cancel", "no", "return") or perform a specific cancellation gesture. Upon receiving the cancellation instruction, the system will cancel the current operation to be executed and may prompt the user to re-enter or provide alternative solutions.

[0132] The application allows users to visually confirm the target operation object through visual marking signals, provides an intuitive and effective voice control feedback mechanism, solves the problem that users cannot instantly confirm whether the system correctly understands their voice instructions, adds a confirmation link, allows users to correct errors in time, significantly reduces the risk of misoperation, enhances the user's trust and operation safety of the system, and improves the interaction reliability and user experience.

[0133] In addition, the application can also continuously monitor the entire response time from the issuance of the voice instruction to the execution of the operation. If the response time exceeds the preset time threshold, the system will judge it as a delay, which can be preset by the user or the developer, and is not specifically limited here. At this time, the system will provide the user with a prompt that "the system is processing, please wait", and try to optimize the processing flow or prompt the user to check the network connection, so as to execute the XR interaction operation after the processing system delay.

[0134] Based on the above-mentioned voice control-based XR interaction method, the application in the working process of executing XR device operation, as shown in Figure 6 includes three parts of initialization phase, real-time processing phase and abnormal processing mechanism, which realize the XR device operation process through information interaction.

[0135] The initialization phase specifically includes: sensor module startup and self-check, user personalized profile loading / establishment, calibration of head tracking and eye tracking system, initialization of voice recognition engine and natural language processing module, establishment of three-dimensional space coordinate system of current environment.

[0136] More specifically, when the smart glasses are started for the first time or the user switches, the system will perform a series of initialization operations to ensure the accuracy and personalization of subsequent voice control. Sensor module startup and self-check: activate all built-in sensors (microphone array, camera, eye tracker, IMU, etc.) and perform self-check to ensure normal operation of the hardware. User personalized profile loading / establishment: load the user's personalized voice model, interaction preference settings, and historical interaction data. For new users, guide them to complete voice model training and calibration. Calibration of head tracking and eye tracking system: through a specific calibration program (such as the user staring at a point on the screen), accurately calibrate the mapping relationship between head and eye movement and XR display space, and ensure the accuracy of the line-of-sight direction and head orientation data. Initialization of voice recognition engine and natural language processing module: load the pre-trained ASR and NLU models, and make adaptive adjustments according to the user's language and accent. Establishment of three-dimensional space coordinate system of current environment: use SLAM technology to construct a three-dimensional map of the environment where the user is located, and establish a stable global coordinate system for positioning XR interface elements and user pose.

[0137] The real-time processing stage specifically includes: input collection, speech processing, motion detection and pose prediction, target positioning, feedback generation, user confirmation and instruction execution, XR device operation.

[0138] More specifically, the real-time processing stage is the core stage of the system's continuous operation, responsible for responding to the user's voice instructions in real time and conducting corresponding XR interactions. Input collection: The multi-modal input collection module collects the user's speech, vision (head orientation, gaze direction), motion (displacement, acceleration, angular velocity), and gesture information in parallel and continuously. All data streams are time-synchronized to ensure the accuracy of information fusion. Speech processing: The collected speech signals are sent to the speech processing module in real time. The ASR engine converts them into text, and the NLU unit analyzes the operation intent and spatial description. At the same time, the instruction classifier preliminarily classifies the instructions. Motion detection and pose prediction: The motion perception unit continuously monitors the user's motion state. If it detects that the user is in a moving state, the pose prediction module will predict the user's pose in the future very short time according to the current motion data, and generate a dynamic coordinate mapping matrix. Target positioning: The target positioning module fuses the semantic information of the voice instruction, the user's behavior (head orientation, gaze direction, hand pointing) under the predicted pose, and the dynamic coordinate mapping matrix, accurately calculates the "target space region" controlled by the user's intent, and matches the XR interface element that best matches the intent in the region. The adaptive attention adjustment algorithm dynamically adjusts the weights of each behavior signal in this process. Feedback generation: Once the target interface element is determined, the feedback control module will immediately generate visual feedback signals (such as highlighting, flashing, color change) to intuitively prompt the user on the XR display interface that the system has identified the target. This step is crucial as it provides the user with an opportunity to confirm. User confirmation and instruction execution: The system waits for the user to confirm the visual feedback. The user can complete the confirmation by issuing a short "confirmation" voice instruction again, a specific nodding action, or by not having other operations within a preset short time (automatic confirmation). Once confirmed, the execution confirmation unit sends the final control instruction to the XR device, triggering the corresponding interface operation (such as opening an application, selecting a menu item, adjusting a parameter, etc.).

[0139] The exception handling mechanism specifically includes: recognition failure handling, motion interference handling, multi-target conflict handling, system delay handling.

[0140] To improve the robustness and user experience of the system, the application designs a perfect abnormality processing mechanism. Specifically, the recognition failure processing: when the highest confidence calculated by the target positioning confidence assessment unit is lower than the preset threshold, the system will judge as recognition failure. At this time, the system will guide the user to reissue the instruction or make a selection through voice prompt (such as "I'm sorry, I didn't understand your instruction, please say it again") or display the alternative list on the XR interface. Motion interference processing: when the motion sensing unit detects that the user's acceleration or angular velocity exceeds the threshold of violent motion, the system will temporarily reduce the attention weight of the user's behavior (head, line of sight, gesture), and even can temporarily suspend the voice control function, and prompt the user through vision or voice "your current motion amplitude is large, please try voice control after stabilizing", in order to avoid misoperation caused by violent motion. Multi-target conflict processing: in some cases, the target positioning module may identify multiple candidate target interface elements with similar confidence. At this time, the system will not immediately execute the operation, but will present these candidate elements in the form of number or thumbnail to the user on the XR interface, and prompt the user through voice "there are multiple matching items, please say the number you want to select", and wait for the user to further specify the selection. System delay processing: the system will continuously monitor the entire response time from the issuance of voice instruction to the execution of operation. If the response time exceeds the preset threshold, the system will judge as delay. At this time, the system will provide the user with the prompt "the system is processing, please wait", and try to optimize the processing flow or prompt the user to check the network connection.

[0141] Through the above technical solution, the application provides an XR interaction control scheme based on voice control, which fuses multi-modal interaction data streams to obtain a target interaction space region determined by combining voice signals and non-voice behavior information, breaks through the limitation of traditional voice control relying only on semantic analysis, greatly reduces the misoperation caused by ambiguous space description, and realizes accurate identification of the target operation object. At the same time, the target interaction space region obtained by fusing user behavior information and space description in voice control instruction dynamically adapts to the relative change of the user's own coordinate system and the XR device interface coordinate system, effectively avoiding the problem of coordinate drift when the user moves, ensuring that even in dynamic scenes such as walking and turning the head, the voice instruction can still be stably mapped to the target position of the device display space, significantly improving the continuity and reliability of the control.

[0142] In addition, through the cooperative processing of multi-modal information, the user can interact with the XR device through the intuitive way of "voice + natural behavior" without relying on additional controllers, which expands the application range of XR devices in mobile scenarios and improves the smoothness of user experience. A high-precision, dynamically adaptive and reliable voice control interaction is realized.

[0143] Based on the same inventive concept, the embodiment of the present application also provides an XR interaction device based on voice control, which can execute the XR interaction method based on voice control described above. As shown in Figure 7 The XR interaction device based on voice control 700 includes:

[0144] The receiving module 701 is configured to receive the multi-modal interaction data stream acquired by the multi-modal sensor. The first determining module 702 is configured to identify and analyze the voice signal in the multi-modal interaction data stream to determine the corresponding voice control instruction. The voice control instruction includes an operation intention and a spatial description. The spatial description is the spatial constraint information of the user on the target operation object. The second determining module 703 is configured to determine the corresponding target interaction space region based on the voice control instruction and one or more non-voice behavior information in the multi-modal interaction data stream synchronized with the voice signal. The matching execution module 704 is configured to match the target operation object corresponding to the voice control instruction according to the information of each operable object in the target interaction space region, and execute the XR interaction operation corresponding to the operation intention according to the target operation object.

[0145] The second determining module 703 is specifically configured to:

[0146] The non-voice behavior information is converted into a behavior feature vector in a three-dimensional space. The non-voice behavior information includes at least one or more of the following: head orientation information, line-of-sight direction information, and hand pointing information. The spatial description in the voice control instruction is converted into a spatial constraint condition. The spatial constraint condition includes a constraint rule parameter group representing the spatial constraint information. Based on the fusion calculation result of the behavior feature vector and the spatial constraint condition, the boundary of the target interaction space region in the three-dimensional space coordinate system is determined, so as to obtain the target interaction space region according to the boundary of the target interaction space region.

[0147] The second determining module 703 is specifically configured to:

[0148] The preset fusion weight set, the behavior feature vector, and the reference vector corresponding to the spatial constraint condition are weighted and calculated to generate a target direction vector after fusion calculation. The direction corresponding to the target direction vector is taken as a reference to generate the boundary of the target interaction space region according to the spatial constraint condition. The spatial constraint condition includes at least a distance constraint parameter and a range constraint parameter.

[0149] Before the preset fusion weight set, the behavior feature vector, and the reference vector corresponding to the spatial constraint condition are weighted and calculated to generate the target direction vector after fusion calculation, the system can also:

[0150] The motion state of the user is determined based on motion state detection information from the multi-modal sensor. The motion state includes a stationary state and a moving state. When the motion state is the stationary state, a first preset fusion weight set is determined according to the information type of the non-speech behavior information, and weighted calculation is performed according to the first preset fusion weight set. When the motion state is the moving state, a user motion intensity value is determined according to the motion state detection information, and the user motion intensity value is matched with a plurality of preset intensity threshold intervals. According to the preset intensity threshold interval in which the user motion intensity value is located, a corresponding dynamic weight adjustment factor is determined, and a baseline preset fusion weight set is loaded. The baseline preset fusion weight set is modified according to the dynamic weight adjustment factor to obtain a corresponding second preset fusion weight set, and corresponding weighted calculation is performed.

[0151] After determining the motion state of the user, the system can also:

[0152] When the motion state is the moving state, a predicted user pose is determined based on a preset pose prediction model and current motion state detection information. A dynamic mapping matrix from the user's own coordinate system to the XR device interface coordinate system is constructed according to the pose information and position information corresponding to the predicted user pose. The pose information includes rotation parameters, and the position information includes translation parameters. The spatial description is converted into coordinates in the XR device interface coordinate system corresponding to the predicted user pose according to the dynamic mapping matrix, and the converted coordinates are updated to the spatial constraint condition.

[0153] The matching execution module 704 is specifically configured to:

[0154] The spatial coordinates and functional attributes of each operable object are determined according to each operable object information. The operable object is an XR device interface element. The matching degrees of each operable object and the voice control instruction are calculated according to the spatial coordinates, the functional attributes, and the voice control instruction. The matching degrees include at least spatial coordinate matching degrees and functional attribute matching degrees. Each matching degree is compared with a preset matching degree threshold to determine a target operation object according to the comparison result.

[0155] The matching execution module 704 is specifically configured to:

[0156] In a case where there are a plurality of pending operable objects determined according to the comparison result, each pending operable object is added to an XR interface display list and displayed to the user. Voice prompt information is played to allow the user to select a display label of each pending operable object in the XR interface display list. The target operation object is determined based on a secondary voice instruction of the user.

[0157] Before performing the XR interaction operation corresponding to the operation intention according to the target operation object, the system can also:

[0158] A visual marking signal of the target operation object is generated. The visual marking signal includes at least one or more of the following: a highlighted border, dynamic flashing, and color gradient. After receiving a confirmation instruction of the user, an XR interaction operation is performed. The confirmation instruction includes at least one of the following: a voice confirmation instruction, a preset gesture instruction, or a preset action instruction.

[0159] Based on the same inventive concept, the present application also provides an XR interaction device based on voice control. Referring to Figure 8 , Figure 8 A structural schematic diagram of an XR interaction device based on voice control is provided for the embodiments of the present application. The device includes:

[0160] at least one processor; and a memory connected to the at least one processor in communication. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:

[0161] receive a multi-modal interaction data stream acquired by a multi-modal sensor. The voice signal in the multi-modal interaction data stream is identified and analyzed to determine a corresponding voice control instruction. The voice control instruction includes an operation intention and a spatial description. The spatial description is spatial constraint information of the target operation object by the user. Based on the voice control instruction and one or more non-voice behavior information in the multi-modal interaction data stream synchronized with the voice signal, a corresponding target interaction space region is determined. According to the information of each operable object in the target interaction space region, the target operation object corresponding to the voice control instruction is matched, and the XR interaction operation corresponding to the operation intention of the target operation object is performed.

[0162] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, the device and equipment embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0163] The device and equipment provided by the embodiments of the present application correspond to the method, so the device and equipment also have similar beneficial technical effects as the method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and equipment will not be described here.

[0164] It should also be noted that the terms "comprising," "including," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0165] The above description is merely illustrative of the application, and not restrictive. Various modifications and changes can become apparent to those skilled in the art. Incorporating any modification, equivalent substitution, improvement, etc. within the spirit and principle of the application, shall be included in the scope of the claims of the application.

Claims

1. A voice-controlled XR interaction method, characterized in that, The method includes: Receive multimodal interactive data streams acquired by multimodal sensors; The speech signal in the multimodal interactive data stream is recognized and parsed to determine the corresponding voice control command; wherein, the voice control command includes the operation intention and spatial description; the spatial description is the spatial constraint information of the user on the target operation object; Based on the voice control command and one or more non-voice behavior information synchronized with the voice signal in the multimodal interaction data stream, a corresponding target interaction space region is determined; this includes: converting the spatial description in the voice control command into spatial constraints; the spatial constraints include a set of constraint rule parameters characterizing the spatial constraint information; based on the fusion calculation result of the behavior feature vector and the spatial constraints, the boundary of the target interaction space region in a three-dimensional spatial coordinate system is determined, so as to obtain the target interaction space region according to the boundary of the target interaction space region, including: determining the user's motion state based on motion state detection information from the multimodal sensor. The motion state includes a stationary state and a moving state. After determining the user's motion state, if the motion state is moving, the predicted user pose is determined based on a preset pose prediction model and the current motion state detection information. A dynamic mapping matrix is ​​constructed from the user's own coordinate system to the XR device interface coordinate system based on the posture information and position information corresponding to the predicted user pose. The posture information includes rotation parameters, and the position information includes translation parameters. Based on the dynamic mapping matrix, the spatial description is converted into coordinates in the XR device interface coordinate system corresponding to the predicted user pose, and the converted coordinates are updated to the spatial constraints. Based on the information of each operable object in the target interactive space area, the target operation object corresponding to the voice control command is matched, and the XR interactive operation corresponding to the operation intention is executed according to the target operation object.

2. The method according to claim 1, characterized in that, The method further includes: The non-voice behavior information is converted into a behavior feature vector in three-dimensional space; wherein the non-voice behavior information includes at least one or more of the following: head orientation information, gaze direction information, and hand pointing information.

3. The method according to claim 2, characterized in that, Based on the fusion calculation result of the behavioral feature vector and the spatial constraints, the boundary of the target interaction space region in the three-dimensional spatial coordinate system is determined, specifically including: The preset fusion weight recombination, the behavioral feature vector, and the reference vector corresponding to the spatial constraint are weighted and calculated to generate the target direction vector after fusion calculation. Based on the direction corresponding to the target direction vector, the boundary of the target interaction space region is generated according to the spatial constraints; wherein, the spatial constraints include at least: distance constraint parameters and range constraint parameters.

4. The method according to claim 3, characterized in that, Before generating the fused target direction vector by weighting the preset fusion weight recombination, the behavioral feature vector, and the reference vector corresponding to the spatial constraint conditions, the method further includes: When the motion state is a stationary state, a first preset fusion weight reorganization is determined according to the information type of the non-voice behavior information, and a weighted calculation is performed according to the first preset fusion weight reorganization. When the motion state is a moving state, the user's motion intensity value is determined based on the motion state detection information, and the user's motion intensity value is matched with multiple preset intensity threshold intervals; Based on the preset intensity threshold range in which the user's exercise intensity value is located, a corresponding dynamic weight adjustment factor is determined, and a benchmark preset fusion weight recombination is loaded. Based on the dynamic weight adjustment factor, the benchmark preset fusion weight reorganization is modified to obtain the corresponding second preset fusion weight reorganization, and the corresponding weighted calculation is performed.

5. The method according to claim 1, characterized in that, Based on the information of each operable object in the target interaction space region, the target operable object corresponding to the voice control command is matched, specifically including: Based on the information of each operable object, determine the spatial coordinates and functional attributes of each operable object; the operable object is an XR device interface element. Based on the spatial coordinates, the functional attributes, and the voice control commands, calculate the matching degree between each operable object and the voice control commands; the matching degree includes at least the spatial coordinate matching degree and the functional attribute matching degree. Each matching degree is compared with a preset matching degree threshold to determine the target operation object based on the comparison result.

6. The method according to claim 5, characterized in that, The target operation object is determined based on the comparison results, specifically including: If, based on the comparison results, it is determined that there are multiple undetermined operable objects, each of the undetermined operable objects is added to the XR interface display list and displayed to the user. Play voice prompts so that the user can select the display labels of each of the pending operable objects in the XR interface display list; The target operation object is determined based on the user's secondary voice command.

7. The method according to claim 1, characterized in that, Before executing the XR interactive operation corresponding to the operation intent based on the target operation object, the method further includes: Generate a visual marker signal for the target object; wherein the visual marker signal includes at least one or more of the following: a highlighted border, dynamic flashing, and color gradient; Upon receiving a confirmation command from the user, the XR interaction operation is executed; wherein the confirmation command includes at least: a voice confirmation command, a preset gesture command, or a preset action command.

8. A voice-controlled XR interactive device, characterized in that, The device is capable of executing the voice-controlled XR interaction method according to any one of claims 1-7; the device comprises: The receiving module is used to receive the multimodal interactive data stream acquired by the multimodal sensor; The first determining module is used to identify and parse the voice signal in the multimodal interactive data stream to determine the corresponding voice control command; wherein, the voice control command includes the operation intention and spatial description; the spatial description is the user's spatial constraint information on the target operation object; The second determining module is used to determine a corresponding target interaction space region based on the voice control command and one or more non-voice behavior information synchronized with the voice signal in the multimodal interaction data stream; including: converting the spatial description in the voice control command into spatial constraints; the spatial constraints include a set of constraint rule parameters characterizing the spatial constraint information; determining the boundary of the target interaction space region in a three-dimensional spatial coordinate system based on the fusion calculation result of the behavior feature vector and the spatial constraints, so as to obtain the target interaction space region according to the boundary of the target interaction space region, including: determining the user based on motion state detection information from the multimodal sensor. The motion state includes a stationary state and a moving state. After determining the user's motion state, if the motion state is moving, the predicted user pose is determined based on a preset pose prediction model and the current motion state detection information. A dynamic mapping matrix is ​​constructed from the user's own coordinate system to the XR device interface coordinate system according to the posture information and position information corresponding to the predicted user pose. The posture information includes rotation parameters, and the position information includes translation parameters. According to the dynamic mapping matrix, the spatial description is converted into coordinates in the XR device interface coordinate system corresponding to the predicted user pose, and the converted coordinates are updated to the spatial constraints. The matching execution module is used to match the target operation object corresponding to the voice control command based on the information of each operable object in the target interaction space area, and to execute the XR interactive operation corresponding to the operation intention based on the target operation object.

9. A voice-controlled XR interactive device, characterized in that, The device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform a voice-controlled XR interaction method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Voice interaction method containing fuzzy anaphora, related device and communication system

    CN118824241A