Kinect voice tracking positioning method and system fusing depth information
By constructing a human spatial positioning coordinate system and fusing Kinect microphone sound source information with human skeletal data, the problem of Kinect's voice positioning in indoor multi-sound source scenarios was solved, achieving fast and accurate voice recognition and positioning, and improving the recognition accuracy of human-computer interaction.
Patent Information
- Application Number
- CN202310477397.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-04-27
AI Technical Summary
In indoor multi-sound-source scenarios, Kinect devices struggle to accurately locate the source of speech, leading to reduced accuracy in speech recognition. Existing methods also suffer from issues such as inconsistent comparison standards and inconsistent acquisition of speech information by microphones.
By constructing a human spatial positioning coordinate system, utilizing Kinect microphone sound source information and human skeletal data, and combining kinematic knowledge and time delay estimation, the orientation and distance values of the human body and sound source signals are fused to achieve voice tracking and positioning.
It achieves rapid and accurate recognition and localization of Kinect voice in indoor multi-sound source scenarios, improving the targeting and efficiency of voice recognition, and is suitable for human-computer interaction systems.
Smart Images

Figure CN116559781B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech recognition technology, and in particular to a Kinect speech tracking and localization method and system that integrates depth information. Background Technology
[0002] With the rapid development of human-computer interaction technology, interaction methods are becoming increasingly diverse, and capturing body movements has become an indispensable part. Human posture recognition is an important research branch in the field of computer vision. Meanwhile, the most primitive form of communication—language—is becoming increasingly popular. Kinect, a motion-sensing device for small-scale human-computer interaction, has emerged to meet this need, primarily used in motion-sensing games and exhibition halls (museums, science museums), etc. Kinect can not only perform human posture recognition using depth and color cameras, but also has four linearly evenly distributed microphone arrays for small-scale speech recognition. However, in enclosed indoor environments, sound reflection causes Kinect to struggle to identify the true source of sound when faced with multiple sound sources, resulting in a lack of specificity in speech recognition. This significantly reduces the effectiveness of Kinect's speech localization and recognition in indoor multi-sound-source scenarios.
[0003] For Kinect speech recognition in complex and specific spatial environments, methods based on Kinect primarily determine the location of the sound source by comparing the speech information acquired by the four microphones of the Kinect. This is followed by noise reduction and enhancement processing of the determined sound source information, ultimately achieving speech recognition through template matching algorithms. However, this method, which determines the sound source location through comparison of four microphones, suffers from problems such as difficulty in unifying comparison standards and inconsistent speech information acquired by the microphones. Another example is the speech recognition method that integrates a deep-information Chinese multimodal corpus. This method utilizes the Kinect color camera and depth camera to acquire color and depth images of the speaker. After preprocessing the acquired multimodal data, multimodal features are extracted to construct a deep-information-integrated Chinese multimodal corpus for multimodal speech recognition. While this method fully utilizes multimodal data from depth cameras, the comparison process is complex, recognition speed is reduced, and the recognition corpus is small, exhibiting limitations in speech recognition. For example, a robot noisy speech recognition device and method is used for complex and specific speech recognition scenarios. By reconstructing lip data and performing multi-stream data fusion, speech recognition is achieved through HMM model modeling, which improves speech recognition efficiency to a certain extent. However, reconstructing lip data not only has high time complexity, but is also limited to the recognition of short words. Summary of the Invention
[0004] To address this issue, the present invention provides a Kinect voice tracking and localization method and system that integrates depth information, solving the problem that Kinect cannot track and locate the true sound source when performing Kinect voice recognition in indoor multi-sound source scenarios.
[0005] According to the design scheme provided by the present invention, a Kinect voice tracking and localization method incorporating depth information is provided, comprising:
[0006] A human body spatial positioning coordinate system is constructed with the Kinect origin as the origin of the spatial coordinate system.
[0007] The target human skeleton data and Kinect microphone sound source information in the spatial scene are obtained by using the human spatial positioning coordinate system. The target human skeleton data includes the spatial coordinates of the human head skeleton points and the spatial coordinates of the neck skeleton points.
[0008] Based on human skeletal data and kinematic knowledge, the orientation value of the target human body and its distance from the Kinect origin are obtained. Based on the sound source information and time delay estimation, the orientation value of the sound source signal and its distance from the Kinect origin are obtained.
[0009] The location and distance values of both the human body and the sound source signal are fused to locate and identify the sound source.
[0010] As a Kinect voice tracking and localization method that integrates depth information according to the present invention, a human body spatial positioning coordinate system is further constructed in the spatial coordinate system with the Kinect origin as the origin of the spatial coordinate system, including:
[0011] First, using the sound source as a reference point, the distance ratio between the sound source's position and the bone points of the head and neck is set according to the human body's physical structure.
[0012] Then, based on the distance ratio between the set human head bone points and neck bone points, human spatial coordinate points are established in the spatial coordinate system.
[0013] As a Kinect voice tracking and localization method that integrates depth information according to the present invention, further, target human skeleton data in a spatial scene is obtained using a human spatial positioning coordinate system, including: obtaining skeleton data at the target time through a Kinect depth camera; and in the case of data loss at the target time, using the method of calculating the average value of the change amount to predict the human skeleton data at the same timestamp to obtain the average displacement change of human skeleton point data in consecutive frames at the time of data loss, and using the average value and the human skeleton point data of the previous time to compensate for the human skeleton data at the target time.
[0014] As a Kinect voice tracking and localization method that integrates depth information according to the present invention, further, it utilizes the human body spatial positioning coordinate system to obtain Kinect microphone sound source information in the spatial scene, including:
[0015] First, the time delay value of sound source detection between the microphones is obtained based on the Kinect microphone array;
[0016] Then, based on the time delay value, the angle between the sound source and the X-axis in the spatial positioning coordinate system and the distance between the sound source and each microphone are calculated.
[0017] As a Kinect voice tracking and positioning method that integrates depth information according to the present invention, further, the orientation value of the target human body and its distance from the Kinect origin are obtained based on human skeletal data and kinematic knowledge, including: obtaining human body spatial positioning coordinates based on the spatial coordinates of the head bone points and neck bone points in the human skeletal data at the target time, and using the human body spatial positioning coordinates to calculate the orientation value of the target human body and its distance from the Kinect origin.
[0018] As a Kinect voice tracking and localization method that integrates depth information according to the present invention, further, obtaining the azimuth value of the sound source signal and its distance from the Kinect origin based on the sound source information and time delay estimation includes: calculating the azimuth value of the sound source signal and its distance from the Kinect origin by using the angle between the sound source and the X-axis and the distance between the sound source and each microphone.
[0019] As a Kinect voice tracking and localization method that integrates depth information according to the present invention, the method further integrates the orientation and distance values of the human body and the sound source signal to locate and identify the sound source, including: using preset data rules to integrate the orientation and distance value data of the human body and the sound source signal, and determining the sound source that conforms to the preset data rules as the target sound source for voice tracking and localization, wherein the preset data fusion rules are expressed as follows: d 阈 and θ 阈 These represent the preset distance threshold and the preset azimuth threshold, respectively. d1 is the human body distance value, θ1 is the human body orientation value, d is the sound source distance value, and θ is the sound source orientation value.
[0020] Furthermore, the present invention also provides a Kinect voice tracking and positioning system that integrates depth information, comprising: a data acquisition module, a data processing module, and a data output module, wherein,
[0021] The data acquisition module is used to construct a human spatial positioning coordinate system with the Kinect origin as the origin in the spatial coordinate system; and to acquire target human skeleton data and Kinect microphone sound source information in the spatial scene using the human spatial positioning coordinate system. The target human skeleton data includes the spatial coordinates of the human head skeleton points and the spatial coordinates of the neck skeleton points.
[0022] The data processing module is used to obtain the orientation value of the target human body and its distance from the Kinect origin based on human skeletal data and kinematic knowledge, and to obtain the orientation value of the sound source signal and its distance from the Kinect origin based on sound source information and time delay estimation.
[0023] The data output module is used to fuse the orientation and distance values of both the human body and the sound source signal to locate and identify the sound source.
[0024] The beneficial effects of this invention are:
[0025] This invention collects Kinect voice data and skeletal data, extracts the location information of the voice data and the location information of the skeletal data, and obtains accurate sound source data based on the fusion positioning strategy. It can realize fast and accurate Kinect voice recognition and positioning in indoor multi-sound source scenarios, which provides convenience for improving the targeted recognition of Kinect voice recognition in human-computer interaction in the future and has good application prospects. Attached image description:
[0026] Figure 1 This is a schematic diagram of the Kinect voice tracking and localization process that incorporates depth information in the embodiment.
[0027] Figure 2 This is a schematic diagram of a Kinect scene in the embodiment;
[0028] Figure 3 This is a schematic diagram illustrating the principle of the Kinect voice tracking and localization algorithm in the embodiment;
[0029] Figure 4 This is a schematic diagram of the Kinect spatial scene coordinates in the embodiment;
[0030] Figure 5 This is a schematic diagram of Kinect positioning for tracking and positioning in the embodiment. Detailed implementation method:
[0031] To make the objectives, technical solutions, and advantages of this invention clearer and more understandable, the invention will be further described in detail below with reference to the accompanying drawings and technical solutions.
[0032] In this embodiment of the invention, see Figure 1 As shown, a Kinect voice tracking and localization method incorporating depth information is provided, comprising:
[0033] S101. Construct a human spatial positioning coordinate system with the Kinect origin as the origin of the spatial coordinate system.
[0034] S102. Use the human body spatial positioning coordinate system to obtain target human skeleton data and Kinect microphone sound source information in the spatial scene. The target human skeleton data includes the spatial coordinates of the human head bone points and the spatial coordinates of the neck bone points.
[0035] S103. Based on human skeletal data and kinematic knowledge, obtain the orientation value of the target human body and its distance from the Kinect origin, and based on the sound source information and time delay estimation, obtain the orientation value of the sound source signal and its distance from the Kinect origin.
[0036] S104. The location and distance values of the human body and the sound source signal are fused to locate and identify the sound source.
[0037] See Figure 2 As shown, a Kinect data acquisition system can be set up indoors. The Kinect device 2 is positioned upright, facing the shooting area 1. The computer 4 is connected to the Kinect via an adapter 3 for data acquisition. Once the system is turned on, the computer 4 acquires real-time skeletal data of the human body within the scene. (See also...) Figure 3 In the algorithm described, since the microphones in Kinect are arranged linearly, they have an advantage in indoor voice interaction in near-field scenarios. Therefore, a spatial coordinate system is established with the midpoint of the Kinect sensor as the origin. The orientation information of the voice data and the orientation information of the skeletal data are extracted through the acquisition of Kinect voice data and skeletal data. The human body localization process based on Kinect does not require the sensor to acquire information on all skeletal points of the human body, ignoring a significant portion of skeletal point information irrelevant to localization. The throat position is selected as the sound source point, so only the Head and Neck skeletal points need to be considered for human body localization, reducing the complexity of data processing and improving recognition efficiency. Each skeletal data point is G(x). i ,y i ,z i ), G(x) i ,y i ,z i ) represents the three-dimensional spatial data of the i-th skeletal point. To improve acquisition efficiency, and considering the characteristics of the pharyngeal vocalization area, the head skeletal point Head(x) is selected. h ,y h ,z h ), Neck bone point (x) n ,y n ,z n), where i represents the number of the bone point, and 0 < i ≤ the number of bone points. At the same timestamp, the computer 4 can collect the sound source information Q in real time from four microphones i , Q i is the sound source information collected by the i-th microphone.
[0038] The human body positioning process based on Kinect does not require the sensor to obtain all the bone point information of the human body. By ignoring a considerable part of the bone point information irrelevant to positioning, the voice positioning recognition efficiency can be improved. Therefore, in the embodiments of this case, in the spatial coordinate system, the origin of Kinect is used as the origin of the spatial coordinate system to construct the human body spatial positioning coordinate system, which can be designed to include the following content:
[0039] First, taking the sound source as the reference point and combining the physical structure characteristics of the human body, the position of the human body sound source can be determined only by the head bone and the neck bone. Therefore, the distance ratio of the sound source emission position can be set according to the physical structure of the human body between the head bone point and the neck bone point;
[0040] Then, based on the set distance ratio between the head bone point and the neck bone point of the human body, a human body spatial coordinate point is established in the spatial coordinate system.
[0041] Based on the sound source as the reference point and considering the physical structure of the human body and the distribution of the bone information collected by the sensor, the distance ratio between the head bone point Head and Neck where the sound source emission position should be located is about 1:2. Therefore, in the spatial coordinate system, assuming the coordinates of the head bone point Head are (x h , y h , z h ), and the coordinates of the neck bone point Neck are (x n , y n , z n ), and setting the human body positioning coordinate point as (x, y, z), according to the above theoretical rules, the human body spatial positioning coordinate point can be obtained from Equation (1-3).
[0042] x = x h = x n (1)
[0043]
[0044] z = z y = z n (3)
[0045] Furthermore, by using the human spatial positioning coordinate system to obtain target human skeleton data in a spatial scene, the following can be designed: acquire skeleton data at the target time using a Kinect depth camera; and, in the case of data loss at the target time, use the method of averaging the changes to predict the human skeleton data at the same timestamp to obtain the average displacement change of human skeleton point data in consecutive frames at the time of data loss, and use the average value and the human skeleton point data of the previous time to compensate for the human skeleton data at the target time.
[0046] At the same timestamp, if the head bone point Head(x) h ,y h ,z h If the data does not meet the calculation standards, the average value of the changes in the skeletal point data will be used for prediction.
[0047] For frame T, calculate the average displacement changes of the skeleton point data x, y, z over n consecutive frames.
[0048]
[0049]
[0050]
[0051] At frame T, the skeleton point G(x) T-1 ,y T-1 ,z T-1 )+(x p ,y p ,z p ), where n < 10
[0052] Based on the spatial coordinates of the head and neck bones in the human skeleton data at the target time, the spatial positioning coordinates of the human body are obtained, and the orientation value of the target human body and its distance from the Kinect origin are calculated using the spatial positioning coordinates of the human body.
[0053] For Kinect's skeletal spatial coordinates, the positions of each human bone can be directly represented using (x, y, z). The unit of measurement in skeletal spatial coordinates is meters. Therefore, the x, y, and z axes correspond to the x, y, and z spatial coordinate axes of the Kinect sensor's physical space. A spatial coordinate system is created with the Kinect origin as the origin, and the X, Y, and Z axes perpendicular to each other, as shown below. Figure 4 As shown.
[0054]
[0055]
[0056] In the formula, θ represents the angle between the human body position and the X-axis of the coordinate system. Since the sound source is a point coordinate in the coordinate system, in order to reduce the error, the human body recognition skeleton point is optimized to be represented as a point coordinate, and the distance d from the human body to the origin of the spatial coordinate and the angle θ value are obtained.
[0057] As a preferred embodiment, further, the acquisition of Kinect microphone sound source information in the spatial scene using the human body spatial positioning coordinate system can be designed to include the following:
[0058] First, the time delay value of sound source detection between the microphones is obtained based on the Kinect microphone array;
[0059] Then, based on the time delay value, the angle between the sound source and the X-axis in the spatial positioning coordinate system and the distance between the sound source and each microphone are calculated.
[0060] Kinect obtains the location of audio information in a spatial scene through a time-delay estimation method and compares it with the position of a person in the scene. Next, it preprocesses the appropriate audio data, extracts features, and uses template matching to achieve voice command recognition. The relationship between the sound source and the coordinate system, such as... Figure 5 As shown. Kinect can detect the location of a sound source through a microphone array. A and D represent an array of four microphones. Assume the distances from the sound source to the microphones are r, r1, r2, and r3, respectively. The distances from A to B and B to C are both l, and the distance from C to D is L. Here, A represents the first microphone with index Q = 1, and so on, with D being the fourth microphone, so Q = 4. The time delay t between the microphones is then calculated. AC t BC .
[0061] r2-r=t AC *C (6)
[0062] r1-r=t BC *C (7)
[0063]
[0064] r1 2 =r 2 +l 2 +2r*l*cosα (9)
[0065] r2 2 =r 2 +4l 2 +4r*l*cosα (10)
[0066] Where C is the speed of sound, which is 340 m / s, α is the angle between the sound source and the X-axis, and r, r1, r2 represent the distances from the sound source to microphones A, B, and C, respectively. tAC and tBC represent the time differences between the sound source and microphones A and C, and B and C, respectively. Based on equation (6-10) above, we can obtain:
[0067]
[0068]
[0069]
[0070]
[0071] in:
[0072] A = C * (t BC -t AC (15)
[0073] B = C * (t) BC +t AC (16)
[0074] The azimuth value of the sound source signal and its distance from the Kinect origin are calculated using the angle between the sound source and the X-axis and the distance between the sound source and each microphone.
[0075] t can be obtained through time delay estimation. BC ,t AC Therefore, A and B are known quantities. From equation (12-16), the values of cosα, r, r1, and r2 can be obtained. Figure 5 The distance from OC to the origin is L-2l, which is a known quantity. The coordinate system for sound source localization is the same as the coordinate system for human body localization. Therefore, from the known quantities cosα, L-2l, and r, we can obtain the distance d1 from the sound source to the origin of the coordinate system and the azimuth angle θ1 of the sound source.
[0076] d1 2 =(L-2l) 2 +r 2 -2(L-2l)*r*cosα (17)
[0077]
[0078] Furthermore, preset data rules can be used to fuse the orientation and distance values of both the human body and the sound source signals. Sound sources that conform to the preset data rules are identified as the target sound sources for voice tracking and localization. The preset data fusion rules are expressed as follows: d 阈 and θ 阈These represent the preset distance threshold and the preset azimuth threshold, respectively. d1 is the human body distance value, θ1 is the human body orientation value, d is the sound source distance value, and θ is the sound source orientation value.
[0079] Voice tracking first obtains the human body's location information by acquiring speech through a microphone array and determining the sound source's location using a time delay estimation method. Secondly, based on the human body's position information in the scene, the speech information is compared and judged. According to a set threshold, it is determined whether the audio information is emitted by a human body recognized by Kinect, thus achieving seamless tracking and recognition of human sound sources within the target area.
[0080] Furthermore, based on the above method, this embodiment of the invention also provides a Kinect voice tracking and positioning system that integrates depth information, comprising: a data acquisition module, a data processing module, and a data output module, wherein...
[0081] The data acquisition module is used to construct a human spatial positioning coordinate system with the Kinect origin as the origin in the spatial coordinate system; and to acquire target human skeleton data and Kinect microphone sound source information in the spatial scene using the human spatial positioning coordinate system. The target human skeleton data includes the spatial coordinates of the human head skeleton points and the spatial coordinates of the neck skeleton points.
[0082] The data processing module is used to obtain the orientation value of the target human body and its distance from the Kinect origin based on human skeletal data and kinematic knowledge, and to obtain the orientation value of the sound source signal and its distance from the Kinect origin based on sound source information and time delay estimation.
[0083] The data output module is used to fuse the orientation and distance values of both the human body and the sound source signal to locate and identify the sound source.
[0084] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of the invention.
[0085] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0086] The units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations are not considered to be beyond the scope of this invention.
[0087] Those skilled in the art will understand that all or part of the steps in the above methods can be implemented by a program instructing related hardware, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk. Optionally, all or part of the steps in the above embodiments can also be implemented using one or more integrated circuits. Accordingly, each module / unit in the above embodiments can be implemented in hardware or as a software functional module. This invention is not limited to any particular combination of hardware and software.
[0088] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A Kinect voice tracking and localization method incorporating depth information, characterized in that, Include: A human body spatial positioning coordinate system is constructed with the Kinect origin as the origin of the spatial coordinate system. The target human skeleton data and Kinect microphone sound source information in the spatial scene are obtained by using the human spatial positioning coordinate system. The target human skeleton data includes the spatial coordinates of the human head skeleton points and the spatial coordinates of the neck skeleton points. Based on human skeletal data and kinematic knowledge, the orientation value of the target human body and its distance from the Kinect origin are obtained. Based on the sound source information and time delay estimation, the orientation value of the sound source signal and its distance from the Kinect origin are obtained. The location and distance values of both the human body and the sound source signal are fused to locate and identify the sound source. Using the human body spatial positioning coordinate system, obtain Kinect microphone sound source information in the spatial scene, including: First, the time delay value of sound source detection between the microphones is obtained based on the Kinect microphone array; Then, based on the time delay value, the angle between the sound source and the X-axis in the spatial positioning coordinate system, as well as the distance from the sound source to each microphone, are calculated, specifically expressed by the following formula: Where C represents the speed of sound, and A, B, and C represent different microphones. Let r be the angle between the sound source and the X-axis. and tBC represents the distance from the sound source to microphones A, B, and C, respectively; d represents the time difference between the sound source and microphones B and C; and d represents the distance to the human body. The time delay value between microphones B and C. This represents the time delay between microphones A and C; The azimuth value of the sound source signal and its distance from the Kinect origin are obtained based on the sound source information and time delay estimation. This includes calculating the azimuth value of the sound source signal and its distance from the Kinect origin using the angle between the sound source and the X-axis and the distance between the sound source and each microphone, specifically expressed by the following formula: in, L represents the distance to the sound source, where L is the distance between microphones C and D. Let A be the distance from microphone A to B and microphone B to C. This represents the orientation value of the human body.
2. The Kinect voice tracking and localization method fused with depth information according to claim 1, characterized in that, A human spatial positioning coordinate system is constructed using the Kinect origin as the origin, which includes: First, using the sound source as a reference point, the distance ratio between the sound source's position and the bone points of the head and neck is set according to the human body's physical structure. Then, based on the distance ratio between the set human head bone points and neck bone points, human spatial coordinate points are established in the spatial coordinate system.
3. The Kinect voice tracking and localization method fusion depth information according to claim 1, characterized in that, The method utilizes a human spatial positioning coordinate system to acquire target human skeleton data in a spatial scene, including: acquiring skeleton data at the target time using a Kinect depth camera; and, in the case of data loss at the target time, using the method of averaging changes to predict the average displacement change of human skeleton data at the same timestamp to obtain the average displacement change of human skeleton point data in consecutive frames at the time of data loss, and using the average value and human skeleton point data from the previous time to compensate for the human skeleton data at the target time.
4. The Kinect voice tracking and localization method fused with depth information according to claim 1 or 3, characterized in that, Based on human skeletal data and kinematic knowledge, the orientation value of the target human body and its distance from the Kinect origin are obtained. This includes: obtaining the spatial positioning coordinates of the human body based on the spatial coordinates of the head and neck bones in the human skeletal data at the target time, and using the spatial positioning coordinates of the human body to calculate the orientation value of the target human body and its distance from the Kinect origin.
5. The Kinect voice tracking and localization method fused with depth information according to claim 1, characterized in that, The method of fusing the orientation and distance values of both the human body and the sound source signal to locate and identify the sound source includes: using preset data rules to fuse the orientation and distance values of both the human body and the sound source signal, and identifying sound sources that conform to the preset data rules as the target sound sources for speech tracking and localization. The preset data fusion rules are expressed as follows: , and These represent the preset distance threshold and the preset azimuth threshold, respectively. This represents the distance to the sound source. This represents the orientation value of the human body. This represents the location of the sound source.
6. A Kinect voice tracking and positioning system that integrates depth information, characterized in that, It includes: a data acquisition module, a data processing module, and a data output module, among which, The data acquisition module is used to construct a human spatial positioning coordinate system with the Kinect origin as the origin of the spatial coordinate system. The target human skeleton data and Kinect microphone sound source information in the spatial scene are obtained by using the human spatial positioning coordinate system. The target human skeleton data includes the spatial coordinates of the human head skeleton points and the spatial coordinates of the neck skeleton points. The data processing module is used to obtain the orientation value of the target human body and its distance from the Kinect origin based on human skeletal data and kinematic knowledge, and to obtain the orientation value of the sound source signal and its distance from the Kinect origin based on sound source information and time delay estimation. The data output module is used to fuse the orientation and distance values of both the human body and the sound source signal to locate and identify the sound source. Using the human body spatial positioning coordinate system, obtain Kinect microphone sound source information in the spatial scene, including: First, the time delay value of sound source detection between the microphones is obtained based on the Kinect microphone array; Then, based on the time delay value, the angle between the sound source and the X-axis in the spatial positioning coordinate system, as well as the distance from the sound source to each microphone, are calculated, specifically expressed by the following formula: Where C represents the speed of sound, and A, B, and C represent different microphones. Let r be the angle between the sound source and the X-axis. and tBC represents the distance from the sound source to microphones A, B, and C, respectively; d represents the time difference between the sound source and microphones B and C; and d represents the distance to the human body. The time delay value between microphones B and C. This represents the time delay between microphones A and C; The azimuth value of the sound source signal and its distance from the Kinect origin are obtained based on the sound source information and time delay estimation. This includes calculating the azimuth value of the sound source signal and its distance from the Kinect origin using the angle between the sound source and the X-axis and the distance between the sound source and each microphone, specifically expressed by the following formula: in, L represents the distance to the sound source, where L is the distance between microphones C and D. Let A be the distance from microphone A to B and microphone B to C. This represents the orientation value of the human body.
7. An electronic device, characterized in that, The system includes a memory and a processor, which communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor can execute the steps of the method as described in any one of claims 1 to 5 by calling the program instructions.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent following shooting method and intelligent following shooting device
CN106647423A
A multi-mode information acquisition system based on Kinect V2
CN109814718A
Microphone tracking system and method combining image recognition and voice positioning
CN111932619A
Multi-target motion capturing skeleton key point tracking method
CN115359098A