A sound source tracking method, device, equipment, system and storage medium

By converting the acoustic signal stream collected by the microphone array into visual data and using an image recognition model to track the sound source, the problems of insufficient robustness and generalization ability in existing technologies are solved, and higher sound source tracking accuracy and adaptability to complex environments are achieved.

CN114355286BActive Publication Date: 2025-09-12ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011086519.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-12
Publication Date
2025-09-12
Estimated Expiration
2040-10-12

AI Technical Summary

Technical Problem

Existing sound source tracking technology has poor robustness, insufficient generalization ability and insufficient accuracy in multiple sound sources or noisy environments.

Method used

The acoustic signal stream collected by the microphone array is converted into visual data of the sound source orientation information, and the sound source is tracked through an image recognition model, which subverts the traditional acoustic signal processing method and uses visual analysis to improve accuracy and adaptability.

Benefits of technology

It improves the accuracy of sound source tracking, enhances adaptability to complex environments, avoids noise interference, and ensures the accuracy and comprehensiveness of visualization data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114355286B_ABST
    Figure CN114355286B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a sound source tracking method, apparatus, device, system and storage medium. The method includes: obtaining an acoustic signal stream collected by a microphone array in at least one time frame; performing sound source orientation estimation based on the acoustic signal stream to obtain an information stream containing sound source orientation information in the at least one time frame; converting the information stream into visual data describing the azimuth distribution state of the sound source; and performing sound source tracking based on the visual data. In the embodiments of the present application, the information stream containing the sound source orientation information is converted into visual data describing the azimuth distribution state of the sound source, and sound source tracking is performed based on the visual data. This subverts the traditional way of tracking sound sources from the acoustic signal processing level, and instead performs sound source tracking from the visual analysis level. Accordingly, in the embodiments of the present application, the accuracy of sound source tracking can be effectively improved, and the adaptability to various complex environments can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a sound source tracking method, device, equipment, system and storage medium. Background Art

[0002] Sound source tracking based on microphone arrays has been a popular technology in the field of acoustic signal processing in recent years. Currently, sound source tracking technology typically involves performing signal-level processing on the microphone array, such as filtering, extracting extrema, calculating fundamental frequency, and calculating azimuth angles, to track the sound source.

[0003] However, this type of processing method has poor robustness and insufficient generalization ability, especially in environments with multiple sound sources or noisy environments, the accuracy of sound source tracking is insufficient. Summary of the Invention

[0004] Various aspects of the present application provide a sound source tracking method, apparatus, device, system, and storage medium to improve the accuracy of sound source tracking.

[0005] The present invention provides a method for tracking a sound source, including:

[0006] Acquiring an acoustic signal stream collected by a microphone array in at least one time frame;

[0007] Performing sound source orientation estimation based on the acoustic signal stream to obtain an information stream containing sound source orientation information in the at least one time frame;

[0008] Converting the information stream into visual data describing the azimuth distribution state of the sound source;

[0009] Sound source tracking is performed based on the visualization data.

[0010] The present application also provides a sound source tracking method, including:

[0011] Determining the sound source location information in at least one time frame within the target period;

[0012] Converting the sound source orientation information in the at least one time frame into at least one set of image data describing the orientation distribution state of the sound source to form an image stream;

[0013] An image recognition model is used to perform image recognition on the image stream to track the sound source within the target time period.

[0014] The present application also provides a sound source tracking device, including:

[0015] An acquisition module, configured to acquire an acoustic signal stream collected by a microphone array in at least one time frame;

[0016] a calculation module, configured to perform sound source orientation estimation based on the acoustic signal stream to obtain an information stream containing sound source orientation information in the at least one time frame;

[0017] A conversion module, configured to convert the information stream into visual data describing the azimuth distribution state of the sound source;

[0018] The tracking module is used to track the sound source according to the visualization data.

[0019] An embodiment of the present application further provides a computing device, including a memory and a processor;

[0020] The memory is used to store one or more computer instructions;

[0021] The processor is coupled to the memory and configured to execute the one or more computer instructions for:

[0022] Acquiring an acoustic signal stream collected by a microphone array in at least one time frame;

[0023] Performing sound source orientation estimation based on the acoustic signal stream to obtain an information stream containing sound source orientation information in the at least one time frame;

[0024] Converting the information stream into visual data describing the azimuth distribution state of the sound source;

[0025] Sound source tracking is performed based on the visualization data.

[0026] The present application also provides a sound source tracking device, including:

[0027] A determination module, configured to determine the sound source position information in at least one time frame within a target period;

[0028] a conversion module, configured to convert the sound source orientation information in the at least one time frame into at least one set of image data describing the orientation distribution state of the sound source, so as to form an image stream;

[0029] The tracking module is used to perform image recognition on the image stream using an image recognition model to track the sound source within the target time period.

[0030] An embodiment of the present application further provides a computing device, including a memory and a processor;

[0031] The memory is used to store one or more computer instructions;

[0032] The processor is coupled to the memory and configured to execute the one or more computer instructions for:

[0033] Determining the sound source location information in at least one time frame within the target period;

[0034] Converting the sound source orientation information in the at least one time frame into at least one set of image data describing the orientation distribution state of the sound source to form an image stream;

[0035] An image recognition model is used to perform image recognition on the image stream to track the sound source within the target time period.

[0036] An embodiment of the present application further provides a sound source tracking system, comprising: a microphone array and a computing device, wherein the microphone array is communicatively connected to the computing device;

[0037] The microphone array is used to collect acoustic signals;

[0038] The computing device is used to obtain an acoustic signal stream collected by a microphone array in at least one time frame; perform sound source orientation estimation based on the acoustic signal stream to obtain an information stream containing sound source orientation information in the at least one time frame; convert the information stream into visual data describing the orientation distribution state of the sound source; and perform sound source tracking based on the visual data.

[0039] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the aforementioned sound source tracking method.

[0040] In an embodiment of the present application, the acoustic signal stream collected by the microphone array in at least one time frame can be used to perform acoustic orientation estimation to determine the acoustic orientation information in at least one time frame respectively, and the information stream containing the sound source orientation information is converted into visual data describing the azimuth distribution state of the sound source, and the sound source is tracked based on the visual data. In this way, in an embodiment of the present application, the traditional method of tracking the sound source from the acoustic signal processing level is overturned, and the sound source is tracked from the visual analysis level. Since the visual data in this embodiment can accurately and comprehensively reflect the azimuth distribution state of the sound source, this ensures the accuracy and comprehensiveness of the basis of the visual analysis and avoids the robustness problem; moreover, in the process of visual analysis, the field of view of the analysis can cover more time frames, so the noise in the field of view can be found, thereby avoiding noise interference; accordingly, in an embodiment of the present application, the accuracy of sound source tracking can be effectively improved, and the adaptability to various complex environments can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0042] Figure 1 A flowchart of a sound source tracking method provided by an exemplary embodiment of the present application;

[0043] Figure 2 A logical diagram of a sound source tracking solution provided by an exemplary embodiment of the present application;

[0044] Figure 3 A schematic diagram of sound source orientation information provided by an exemplary embodiment of the present application;

[0045] Figure 4 A schematic diagram of a heat map of the azimuth distribution of a sound source provided by an exemplary embodiment of the present application;

[0046] Figure 5 A schematic structural diagram of a sound source tracking device provided as an exemplary embodiment of the present application;

[0047] Figure 6 A schematic structural diagram of a computing device provided as another exemplary embodiment of the present application;

[0048] Figure 7 A flowchart of another sound source tracking method provided by an exemplary embodiment of the present application;

[0049] Figure 8 A schematic structural diagram of another sound source tracking device provided as an exemplary embodiment of the present application;

[0050] Figure 9 A schematic structural diagram of another computing device provided as an exemplary embodiment of the present application;

[0051] Figure 10 A structural diagram of a sound source tracking system provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0052] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0053] In response to technical issues such as poor robustness and insufficient generalization capabilities in existing sound source tracking solutions, some embodiments of the present application can convert an information stream containing sound source orientation information into visual data describing the orientation distribution state of the sound source, and perform sound source tracking based on the visual data. This subverts the traditional method of tracking sound sources from the acoustic signal processing level, and instead performs sound source tracking from the visual analysis level. Accordingly, in the embodiments of the present application, the accuracy of sound source tracking can be effectively improved, and the adaptability to various complex environments can be improved.

[0054] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0055] Figure 1 A flowchart of a sound source tracking method provided by an exemplary embodiment of the present application. Figure 2 This is a logical diagram of a sound source tracking solution provided by an exemplary embodiment of the present application. The sound source tracking method provided by this embodiment can be performed by a sound source tracking device, which can be implemented as software or a combination of software and hardware. The sound source tracking device can be integrated into a computing device. Figure 1 As shown, the method includes:

[0056] Step 100: Acquire an acoustic signal stream collected by a microphone array in at least one time frame;

[0057] Step 101: performing sound source orientation estimation based on an acoustic signal stream to obtain an information stream containing sound source orientation information in at least one time frame;

[0058] Step 102: convert the information stream into visual data describing the azimuth distribution state of the sound source;

[0059] Step 103: Track the sound source based on the visualization data.

[0060] The sound source tracking method provided in this embodiment can be applied to various scenarios, such as voice control scenarios, audio and video conferencing scenarios, or other scenarios requiring sound source tracking. This embodiment does not limit the application scenarios. In different application scenarios, the sound source tracking method provided in this embodiment can be integrated into a variety of scene devices. For example, in a voice control scenario, the scene devices can be smart speakers, smart robots, etc., and in an audio and video conferencing scenario, the scene devices can be various conference terminals, etc.

[0061] In this embodiment, in step 100, a microphone array may be used to collect the acoustic signal stream. The microphone array may be a group of array elements, and this embodiment does not limit the number of array elements in the microphone array. This embodiment also does not limit the arrangement of the microphone array, and the microphone array may be a circular array, a linear array, a planar array, or a three-dimensional array, etc. In different application scenarios, the microphone array can be installed in various types of scene equipment as needed.

[0062] The signal collection process of the microphone array is usually a continuous process. Therefore, in this embodiment, subsequent processing can be performed in the form of an acoustic signal stream.

[0063] In this embodiment, based on the recognition accuracy, at least one time frame can be selected within a single recognition period, and a single time frame can be used as a processing unit. The length of a single recognition period is adapted to the recognition accuracy. For example, if the recognition accuracy is 1s, that is, the sound source tracking result is shown once every 1s, then the length of a single recognition period can be set to 1s. Then, in step 100, the acoustic signal stream formed by the acoustic signal in at least one time frame within 1s can be obtained as the processing object of the subsequent steps. In actual applications, in different application scenarios, at least one time frame can be selected as needed within the target period. For example, when the acoustic signal changes little, at least one time frame can be selected within the target period by using frame skipping or variable frame rate sampling. Of course, in most cases, all time frames within the target period can be selected, and this embodiment does not limit this.

[0064] In this embodiment, the frame length of the time frame can be configured according to the actual request. For example, the frame length of a single time frame can be configured to 20ms. In addition, the number of at least one time frame in the recognition period can also be set as needed. For example, if the recognition period is 1s, 3 time frames can be selected in the recognition target period. In this way, the sound source tracking in the target period is performed based on these 3 time frames. Of course, the frame length of the time frame and the number of time frames in the recognition period in this embodiment are not limited to this. In addition, the frame lengths of different time frames in the recognition period may not be exactly the same, and this embodiment does not limit this.

[0065] Based on this, the sound source tracking method provided in this embodiment can be applied to real-time sound source tracking scenarios as well as offline sound source tracking scenarios. Sound source tracking is performed continuously in each recognition period according to recognition accuracy.

[0066] In practical applications, each element in the microphone array can be used to collect time domain signals separately. Taking the microphone array containing M elements as an example, M time domain signal streams can be collected in at least one time frame as the acoustic signal stream in step 100.

[0067] refer to Figure 1 and Figure 2 In step 101, the sound source orientation may be estimated based on the acoustic signal stream to obtain an information stream containing sound source orientation information in at least one time frame.

[0068] In this embodiment, a sound source orientation estimation technology can be used to perform signal processing on the acoustic signal stream to determine the sound source orientation information in at least one time frame. The sound source orientation information is used to characterize the orientation data of the sound source in the time frame. In this embodiment, the orientation data can be the confidence level that the sound source is in each orientation. In this way, the sound source orientation information can at least include the confidence level that the sound source is in each orientation in the time frame. In this embodiment, the orientations involved in the sound source orientation information can be configured according to actual needs. For example, 360 orientations, 120 orientations, 60 orientations, etc. can be configured for the full circumference of the microphone array. Of course, the microphone array can also be non-full circumference, for example, 180 orientations can be configured in a 180° range on the front, etc. This embodiment does not limit this.

[0069] Figure 3 A schematic diagram of sound source orientation information provided by an exemplary embodiment of the present application. Figure 3 The sound source orientation information is visualized in the video, but it should be understood that Figure 3 This description is provided for convenience only and should not limit the data format of the sound source location information in this embodiment. In practical applications, the sound source location information can be in any other data format understandable by a computing device, such as [1, 3, 5, 60, 70, 80, 90, 80, 70, ..., 0]. In this example, each number in [] represents the confidence level of the sound source in any of the 360 ​​directions.

[0070] Furthermore, in step 101, the sound source direction information is determined in time frames. As previously mentioned, in step 100, the acoustic signal streams collected by each of the M array elements in at least one time frame are acquired. Here, in step 101, the sound source direction can be estimated for the acoustic signals collected by the M array elements in the time frame, thereby determining the sound source direction information in the time frame.

[0071] In an optional implementation, the time domain signal stream collected by each array element can be converted into time-frequency domain signals respectively; and the sound source direction estimation technology is used to determine the sound source direction information in at least one time frame based on the time-frequency domain signals of each array element.

[0072] In this implementation, a target time frame from at least one time frame is used as an example. The target time frame can be any one of the at least one time frame. The time-domain signals collected by each array element in the target time frame can be converted into time-frequency domain signals. For example, the time-domain signals can be decomposed into subbands to obtain time-frequency domain signals. The subband decomposition process can be implemented based on, for example, an end-time Fourier transform and / or a filter bank, which is not limited here. Based on this, the time-frequency domain signals corresponding to each array element in the target time frame can be obtained.

[0073] On this basis, the time-frequency domain signals corresponding to each array element in the target time frame can be used to estimate the sound source direction to output the sound source direction information in the target time frame. Among them, sound source direction estimation technologies include but are not limited to steerable beam response power-phase transform (SRP-PHAT), generalized cross correlation phase transform (GCC-PHAT), or multiple signal classification (MUSIC). The principle of sound source direction estimation can be: based on the acoustic signals collected by different microphones in the microphone array at the same time, the azimuth range of the sound source is calculated separately, and then the sound source direction is estimated based on the multiple azimuth ranges. Of course, this is merely exemplary, and this embodiment is not limited to this. This embodiment does not limit the sound source direction estimation technology used, nor does it elaborate on the processing procedures of various sound source direction estimation technologies.

[0074] In addition, in this embodiment, other implementations may be used to estimate the direction of the sound source based on the acoustic signal stream, and this embodiment is not limited to the above implementations.

[0075] Based on this, an information stream can be obtained, which contains the sound source orientation information in at least one time frame.

[0076] refer to Figure 1 and Figure 2 On this basis, in step 102, the information stream can be converted into visual data describing the azimuth distribution state of the sound source.

[0077] In this embodiment, the sound source location information may include descriptive information in dimensions such as time and location. Based on this, the sound source location information can be converted into descriptive information about the location distribution. For example, the sound source location information may include confidence levels for the sound source at different locations. The confidence levels can then be converted into adaptive display brightness, with the display brightness at different locations serving as descriptive information about the location distribution. It is worth noting that during the conversion of the visual data, no content in the sound source location information is lost; only the representation of the sound source location information is converted. This ensures that the visual data in this embodiment accurately and comprehensively describes the location distribution of the sound source.

[0078] The visualization data can be a heat map of the sound source's azimuth distribution. Continuing with the previous example, after converting the confidence level of the sound source at different azimuths into appropriate display brightness, the display brightness at each azimuth within the time frame can be obtained. This allows the heat map of the sound source's azimuth distribution to be determined from the three dimensions of time frame, azimuth, and display brightness.

[0079] Of course, in this embodiment, the visualization data is not limited thereto. For example, the visualization data may also be a three-dimensional image used to represent the sound source orientation information in at least one time frame. In practical applications, the visualization data may be obtained by mapping the corresponding image of at least one time frame to a three-dimensional image. Figure 3 The display curves of the sound source orientation information are arranged in time to obtain a three-dimensional stereogram.

[0080] In this embodiment, the information stream can be converted into various forms of visual data. During the visualization process, the position information of each sound source can be fully retained. Therefore, in step 102, the process of sound source tracking can be converted from the acoustic signal processing level to the visualization processing level.

[0081] In step 103, the sound source may be tracked based on the visualization data. In this embodiment, the number of sound sources is not limited, and the number of sound sources may be one or more.

[0082] Based on this, in this embodiment, visual analysis of the visual data can be performed to track the sound source within at least one time frame, thereby converting the acoustic signal processing problem into a visual analysis problem. Because the visual data in this embodiment accurately and comprehensively reflects the azimuthal distribution of the sound source, this ensures the accuracy and comprehensiveness of the foundation for the visual analysis and avoids robustness issues. Furthermore, during the visual analysis process, the analysis field of view can cover more time frames, rather than being limited to a single time frame. Therefore, noise within the field of view can be detected, thus avoiding noise interference. This effectively circumvents the shortcomings of traditional acoustic signal processing, such as poor robustness and insufficient generalization capabilities.

[0083] Based on this, in this embodiment, an information stream containing sound source position information can be converted into visual data describing the azimuthal distribution of the sound source, and sound source tracking can be performed based on this visual data. This subverts the traditional approach of tracking sound sources from the perspective of acoustic signal processing, and instead performs sound source tracking from the perspective of visual analysis. As a result, in this embodiment of the application, the accuracy of sound source tracking can be effectively improved, and the adaptability to various complex environments can be enhanced.

[0084] In the above or following embodiments, the information stream can be converted into a directional distribution heat map of the sound source in at least one time frame as a basis for tracking the sound source in the at least one time frame. The directional distribution heat map is used to describe the distribution heat of the sound source in different directions in the at least one time frame.

[0085] In this embodiment, the sound source orientation information includes the confidence level of the sound source at each orientation. Based on this confidence level and the corresponding relationship between display brightness, the display brightness corresponding to each orientation in at least one time frame can be determined according to the confidence level of the sound source at each orientation in at least one time frame. Based on the display brightness, a heat map of the sound source's orientation distribution in at least one time frame is generated. Different display brightness levels represent different distribution heat levels.

[0086] In practical applications, higher confidence levels can correspond to higher display brightness, representing a higher distribution heat. Of course, this embodiment is not limited to this; higher confidence levels can also correspond to lower display brightness. However, typically, confidence levels and display brightness are proportional, allowing the display brightness to accurately reflect confidence levels.

[0087] Figure 4 A schematic diagram of a heat map of the azimuth distribution of a sound source provided by an exemplary embodiment of the present application. Figure 4 ,The vertical axis of the heat map is the time frame, and the horizontal axis is the direction. Figure 4 In the embodiment, the number of at least one time frame is 800 and the number of orientations is configured to be 120, which are used to characterize the omnidirectional space of the microphone array.

[0088] In an optional implementation, the image content corresponding to at least one time frame can be determined based on the display brightness corresponding to each direction in at least one time frame; and the image content corresponding to at least one time frame can be arranged in sequence according to the time sequence between at least one time frame to generate an orientation distribution heat map. Figure 4 The sound source orientation information in each time frame can be converted into a horizontal line in the heat map. For example, the sound source orientation information in the 400th frame can be converted into a straight line y=400 in the heat map, and the display brightness of the pixels corresponding to each orientation on this line is determined according to the corresponding confidence level. The higher the confidence level, the brighter the corresponding pixel brightness. For example, combined with Figure 3Schematic diagram of the sound source orientation information shown in FIG. Figure 3 The peak position in the image has the highest confidence level. When converted to the heat map, the direction corresponding to the peak position has the brightest brightness.

[0089] Of course, in this embodiment, other implementation methods can also be used to generate an orientation distribution heat map. For example, based on the sound source orientation information in at least one time frame, the confidence level of the sound source existing in the target orientation in different time frames can be obtained, thereby determining the display brightness corresponding to each time frame in the target orientation to generate image content corresponding to the target orientation, wherein the target orientation can be any one of the orientations. This can obtain the image content corresponding to each orientation, and thus the image content corresponding to each orientation can be arranged in sequence according to the orientation order to generate an orientation distribution heat map. This embodiment does not limit the method for generating an orientation distribution heat map.

[0090] Based on this, in this embodiment, the sound source location information for at least one time frame can be converted into a heat map of the sound source's location distribution. Furthermore, the conversion process preserves all the information contained in the sound source location information, providing an accurate basis for visualization analysis and ensuring the accuracy of tracking results.

[0091] In the above or below embodiments, a machine learning model and visual data may be used to track the sound source.

[0092] In this embodiment, regardless of the form of visual data, a machine learning model can be used for visual analysis to track sound sources. In practical applications, different types of machine learning models can be used for different forms of visual data, and model training methods that are adapted to the data format can be used to improve the performance of the machine learning model.

[0093] The following uses the heat map as an example to illustrate the visualization analysis process.

[0094] In this embodiment, in the machine learning model, image features in the azimuth distribution heat map can be extracted; based on the mapping relationship between the image features and the sound source attribute parameters and the image features extracted from the azimuth distribution heat map, the target sound source attribute parameters in at least one time frame are determined to perform sound source tracking.

[0095] Among them, the mapping relationship between image features and sound source attribute parameters can be configured into the machine learning model through model training.

[0096] An exemplary model training process may also be:

[0097] Obtain sample heat maps corresponding to several sample time frame groups; annotate sound source attribute parameters for each sample heat map to obtain annotation information corresponding to each sample heat map; input each sample heat map and its corresponding annotation information into a machine learning model so that the machine learning model can learn the mapping relationship between image features and sound source attribute parameters.

[0098] The number of time frames in the sample time frame can be consistent with the number of at least one time frame used to collect acoustic signals in step 100. That is, the processing units during model training and model use can remain consistent. Thus, during model training, annotation is performed using groups of sample time frames as units, while during model use, tracking results can be output using at least one time frame of the same order.

[0099] In this embodiment, the sound source attribute parameters include one or more of location, number, duration of sound emission, and time frame covered. Accordingly, in this embodiment, after visual analysis, the machine learning model can output information such as the number of sound sources, their location, duration of sound emission, and time frame covered in at least one time frame as tracking results.

[0100] When a machine learning model is introduced, in this embodiment, the step of converting the information flow into visual data describing the azimuthal distribution state of the sound source can be performed in the machine learning model or outside the machine learning model.

[0101] In one possible implementation, during model use, the process of converting the information stream into visual data describing the azimuthal distribution of sound sources can be performed outside of the machine learning model. This visual data can then be used as an input parameter for the machine learning model. Accordingly, during model training, the information stream corresponding to the sample time frame groups can be pre-converted into sample heatmaps as a basis for model training.

[0102] In another possible implementation, during the use of the model, the information stream can be input into the machine learning model; in the machine learning model, the information stream is converted into visual data that describes the azimuthal distribution state of the sound source.

[0103] In this implementation, a functional module can be configured within the machine learning model to convert an information stream into visual data describing the azimuthal distribution of the sound source, allowing the information stream to serve as an input parameter for the machine learning model. Upon receiving the information stream, the machine learning model can convert it into visual data describing the azimuthal distribution of the sound source, which can then be used for visual analysis.

[0104] Accordingly, the model training process will be slightly different from the previous implementation. During the model training process of this implementation, sample information streams corresponding to each sample time frame group can be obtained; sound source attribute parameters are annotated for each sample information stream to obtain the corresponding annotation information for each sample information stream; each sample information stream and its corresponding annotation information are input into the machine learning model, which then converts each sample information stream into visual data describing the azimuthal distribution of the sound source and learns the mapping relationship between image features and sound source attribute parameters.

[0105] Accordingly, in this embodiment, after training the machine learning model with sufficient sample data, the machine learning model can learn the accurate mapping relationship between image features and sound source attribute parameters. Consequently, the trained machine learning model can be used to perform visual analysis on the visualized data to output sound source attribute information for at least one time frame, thereby tracking one or more sound sources based on this sound source attribute information. This sound source tracking method can eliminate various noise interferences during the tracking process, eliminate the need for a separate search for the origin of the sound, and avoid the shortcomings of other acoustic signal processing methods. Furthermore, it can effectively improve the accuracy of tracking results and enhance adaptability to various complex environments.

[0106] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 101 to 103 can be device A; for another example, the execution entity of steps 101 and 102 can be device A, and the execution entity of step 103 can be device B; and so on.

[0107] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The sequence numbers of the operations, such as 101, 102, etc., are merely used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel.

[0108] Figure 5 A schematic diagram of a sound source tracking device provided by an exemplary embodiment of the present application. Figure 5 , the sound source tracking device comprises:

[0109] An acquisition module 50 is configured to acquire an acoustic signal stream collected by the microphone array in at least one time frame;

[0110] A calculation module 51 is configured to estimate the direction of a sound source based on the acoustic signal stream to obtain an information stream containing the direction information of the sound source in at least one time frame;

[0111] a conversion module 52 for converting the information stream into visual data describing the azimuth distribution state of the sound source;

[0112] The tracking module 53 is used to track the sound source according to the visualization data.

[0113] In an optional embodiment, when converting the information stream into visual data describing the azimuth distribution state of the sound source, the conversion module 52 is configured to:

[0114] The information stream is converted into an azimuth distribution heat map of the sound source in at least one time frame, where the azimuth distribution heat map is used to describe the distribution heat of the sound source in different directions in the at least one time frame.

[0115] In an optional embodiment, the sound source position information includes the confidence level of the sound source in each position; when converting the information stream into a heat map of the sound source position distribution in at least one time frame, the conversion module 52 is configured to:

[0116] Based on the correspondence between confidence and display brightness, according to the confidence of the sound source in each direction in at least one time frame, the display brightness corresponding to each direction is determined in at least one time frame, and different display brightness represents different distribution heat;

[0117] Generate a heat map of the azimuth distribution of the sound source in at least one time frame according to the display brightness.

[0118] In an optional embodiment, when generating the azimuth distribution heat map of the sound source in at least one time frame according to the display brightness, the conversion module 52 is configured to:

[0119] Determining the image content corresponding to each of the at least one time frame according to the display brightness corresponding to each direction in the at least one time frame;

[0120] According to the time sequence between at least one time frame, the image content corresponding to each of the at least one time frame is arranged in sequence to generate an orientation distribution heat map.

[0121] In an optional embodiment, when tracking the sound source based on the visualization data, the tracking module 53 is configured to:

[0122] Use machine learning models and visualized data to track sound sources.

[0123] In an optional embodiment, if the visualization data is a heat map of the azimuth distribution of the sound source in at least one time frame, the tracking module 53, when using the machine learning model and the visualization data to track the sound source, is configured to:

[0124] In the machine learning model, image features are extracted from the orientation distribution heat map;

[0125] Based on the mapping relationship between image features and sound source attribute parameters and the image features extracted from the azimuth distribution heat map, the target sound source attribute parameters in at least one time frame are determined to perform sound source tracking.

[0126] In an optional embodiment, the sound source attribute parameters include one or more of direction, quantity, sound emission duration and covered time frame.

[0127] In an optional embodiment, the tracking module 53 is further configured to:

[0128] Obtaining sample heat maps corresponding to each of several sample time frame groups, where the sample heat maps are used to describe the distribution heat of the sound source at different directions in the sample time frames;

[0129] Mark the sound source attribute parameters for each sample heat map to obtain the annotation information corresponding to each sample heat map;

[0130] The heat map of each sample and its corresponding annotation information are input into the machine learning model so that the machine learning model can learn the mapping relationship between image features and sound source attribute parameters.

[0131] In an optional embodiment, when converting the information stream into visual data describing the azimuth distribution state of the sound source, the tracking module 53 is configured to:

[0132] Feeding information streams into machine learning models;

[0133] In the machine learning model, the information flow is converted into visual data describing the azimuthal distribution of the sound source.

[0134] In an optional embodiment, the tracking module 53 is further configured to:

[0135] Get the sample information stream corresponding to each sample time frame group;

[0136] Labeling sound source attribute parameters for each sample information stream to obtain labeling information corresponding to each sample information stream;

[0137] Each sample information stream and its corresponding annotation information are input into the machine learning model so that the machine learning model can convert each sample information stream into visual data describing the azimuthal distribution state of the sound source and learn the mapping relationship between image features and sound source attribute parameters.

[0138] In an optional embodiment, the acoustic signal stream includes a time domain signal stream collected by each array element in the microphone array. When the calculation module 51 performs sound source orientation estimation based on the acoustic signal stream to obtain an information stream including sound source orientation information in at least one time frame, it is configured to:

[0139] Convert the time domain signal stream collected by each array element into time-frequency domain signals respectively;

[0140] The sound source direction estimation technology is used to determine the sound source direction information in at least one time frame based on the time-frequency domain signals of each array element.

[0141] In an optional embodiment, the sound source direction estimation technology includes one or more of the steerable beam response phase transform technology SRP-PHAT, the generalized cross-correlation phase transform technology GCC-PHAT or the multiple signal classification technology MUSIC.

[0142] It is worth noting that the technical details in the above-mentioned embodiments of the sound source tracking device can be referred to the relevant descriptions in the above-mentioned embodiments of the sound source tracking method. In order to save space, they will not be repeated here, but this should not cause any loss of the protection scope of this application.

[0143] Figure 6 A schematic diagram of a computing device provided as another exemplary embodiment of the present application is shown in FIG. Figure 6 As shown, the computing device includes a memory 60 and a processor 61 .

[0144] The processor 61 is coupled to the memory 60 and is configured to execute the computer program in the memory 60 to:

[0145] Acquiring an acoustic signal stream collected by the microphone array 62 in at least one time frame;

[0146] Performing sound source orientation estimation based on the acoustic signal stream to obtain an information stream containing sound source orientation information in at least one time frame;

[0147] Convert the information stream into visual data describing the azimuth distribution of the sound source;

[0148] Track the sound source based on the visual data.

[0149] In an optional embodiment, when converting the information stream into visual data describing the azimuth distribution state of the sound source, the processor 61 is configured to:

[0150] The information stream is converted into an azimuth distribution heat map of the sound source in at least one time frame, where the azimuth distribution heat map is used to describe the distribution heat of the sound source in different directions in the at least one time frame.

[0151] In an optional embodiment, the sound source position information includes the confidence level of the sound source in each position; when converting the information stream into a heat map of the sound source position distribution in at least one time frame, the processor 61 is configured to:

[0152] Based on the correspondence between confidence and display brightness, according to the confidence of the sound source in each direction in at least one time frame, the display brightness corresponding to each direction is determined in at least one time frame, and different display brightness represents different distribution heat;

[0153] Generate a heat map of the azimuth distribution of the sound source in at least one time frame according to the display brightness.

[0154] In an optional embodiment, when generating the azimuth distribution heat map of the sound source in at least one time frame according to the display brightness, the processor 61 is configured to:

[0155] Determining the image content corresponding to each of the at least one time frame according to the display brightness corresponding to each direction in the at least one time frame;

[0156] According to the time sequence between at least one time frame, the image content corresponding to at least one time frame is arranged in sequence to generate an orientation distribution heat map.

[0157] In an optional embodiment, when performing sound source tracking based on the visualization data, the processor 61 is configured to:

[0158] Use machine learning models and visualized data to track sound sources.

[0159] In an optional embodiment, if the visualization data is a heat map of the azimuth distribution of the sound source in at least one time frame, the processor 61, when using the machine learning model and the visualization data to track the sound source, is configured to:

[0160] In the machine learning model, image features are extracted from the orientation distribution heat map;

[0161] Based on the mapping relationship between image features and sound source attribute parameters and the image features extracted from the azimuth distribution heat map, the target sound source attribute parameters in at least one time frame are determined to perform sound source tracking.

[0162] In an optional embodiment, the sound source attribute parameters include one or more of direction, quantity, sound emission duration and covered time frame.

[0163] In an optional embodiment, the processor 61 is further configured to:

[0164] Obtaining sample heat maps corresponding to each of several sample time frame groups, where the sample heat maps are used to describe the distribution heat of the sound source at different directions in the sample time frames;

[0165] Mark the sound source attribute parameters for each sample heat map to obtain the annotation information corresponding to each sample heat map;

[0166] The heat map of each sample and its corresponding annotation information are input into the machine learning model so that the machine learning model can learn the mapping relationship between image features and sound source attribute parameters.

[0167] In an optional embodiment, when converting the information stream into visual data describing the azimuth distribution state of the sound source, the processor 61 is configured to:

[0168] Feeding information streams into machine learning models;

[0169] In the machine learning model, the information flow is converted into visual data describing the azimuthal distribution of the sound source.

[0170] In an optional embodiment, the processor 61 is further configured to:

[0171] Get the sample information stream corresponding to each sample time frame group;

[0172] Labeling sound source attribute parameters for each sample information stream to obtain labeling information corresponding to each sample information stream;

[0173] Each sample information stream and its corresponding annotation information are input into the machine learning model so that the machine learning model can convert each sample information stream into visual data describing the azimuthal distribution state of the sound source and learn the mapping relationship between image features and sound source attribute parameters.

[0174] In an optional embodiment, the acoustic signal stream includes a time domain signal stream collected by each array element in the microphone array. When the processor 61 performs sound source orientation estimation based on the acoustic signal stream to obtain an information stream including sound source orientation information in at least one time frame, the processor 61 is configured to:

[0175] Convert the time domain signal stream collected by each array element into time-frequency domain signals respectively;

[0176] The sound source direction estimation technology is used to determine the sound source direction information in at least one time frame based on the time-frequency domain signals of each array element.

[0177] In an optional embodiment, the sound source direction estimation technology includes one or more of the steerable beam response phase transform technology SRP-PHAT, the generalized cross-correlation phase transform technology GCC-PHAT or the multiple signal classification technology MUSIC.

[0178] It is worth noting that the technical details in the above-mentioned embodiments of the computing device can be referred to the relevant descriptions in the above-mentioned embodiments of the sound source tracking method. In order to save space, they will not be repeated here, but this should not cause any loss of the scope of protection of this application.

[0179] Further, if Figure 6 As shown, the computing device also includes: a communication component 63, a power supply component 64 and other components. Figure 6 Only some components are shown schematically, and it does not mean that the computing device only includes Figure 6 Components shown.

[0180] Figure 7 This is a flowchart of another sound source tracking method provided by an exemplary embodiment of the present application. The sound source tracking method provided by this embodiment can be performed by a sound source tracking device, which can be implemented as software or a combination of software and hardware. The sound source tracking device can be integrated into a computing device. Figure 7 As shown, the method includes:

[0181] Step 700: Determine the sound source position information in at least one time frame within the target period.

[0182] Step 701: Convert the sound source orientation information in at least one time frame into at least one set of image data describing the orientation distribution state of the sound source to form an image stream;

[0183] Step 702: Use an image recognition model to perform image recognition on the image stream to track the sound source within the target time period.

[0184] Among them, step 700 can refer to Figure 1 Related descriptions in the associated embodiments. Step 701 may further include obtaining an acoustic signal stream collected by the microphone array in at least one time frame; and performing sound source orientation estimation based on the acoustic signal stream to obtain sound source orientation information in at least one time frame. To save space, the specific process is not repeated here.

[0185] In step 701, the sound source orientation information for at least one time frame can be converted into at least one set of image data. The image data can be an orientation distribution heat map. In step 701, at least one orientation distribution heat map can be obtained to form an image stream that is input into the image recognition model.

[0186] Based on step 701, the sound source tracking solution provided in this embodiment can be applied to scenarios such as real-time tracking or offline tracking. In the offline tracking scenario, the sound source position information for at least one time frame can be obtained at one time. The sound source position information for at least one time frame can then be grouped according to recognition accuracy, and the operation of converting the acoustic position information into image data is performed on a group-by-group basis.

[0187] In the online tracking scenario, the target time frame within the current recognition period can be determined from at least one time frame based on the preset recognition accuracy;

[0188] Convert the sound source orientation information under each target time frame into a set of image data describing the orientation distribution state of the sound source within the current recognition period;

[0189] Continue to determine the time frame and image data within the next recognition period in the target period until image data corresponding to all recognition periods in the target period are generated.

[0190] In an online tracking scenario, as acoustic signals are continuously generated, the operation of converting acoustic position information into image data can be performed continuously in each recognition period, and then continuously input into the image recognition model in step 702 .

[0191] For example, if the recognition accuracy is 1s, an orientation distribution heat map can be generated based on the sound source orientation information of N time frames within the current recognition period (1s). After that, the orientation distribution heat map of the next recognition period (1s) can be generated, and the orientation distribution heat map of each subsequent recognition period can be generated continuously, and each orientation distribution heat map can be provided to the image recognition model in a streaming manner.

[0192] In this embodiment, the image recognition model can be pre-trained. The image recognition model can adopt a machine learning model. The training process of the image recognition model can refer to Figure 1 Related description in the associated embodiments.

[0193] It is worth noting that the recognition accuracy remains consistent during the training and application stages of the image recognition model.

[0194] For technical details of each embodiment of the above-mentioned sound source tracking method, please refer to Figure 1 To save space, the relevant descriptions in each embodiment of the associated sound source tracking method will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0195] Figure 8 A schematic diagram of another sound source tracking device provided by an exemplary embodiment of the present application. Figure 8 , the sound source tracking device comprises:

[0196] A determination module 80 is configured to determine the sound source position information in at least one time frame within a target period;

[0197] The conversion module 81 is used to convert the sound source position information in at least one time frame into at least one set of image data describing the position distribution state of the sound source to form an image stream;

[0198] The tracking module 82 is configured to perform image recognition on the image stream using an image recognition model to track the sound source within a target time period. In an optional embodiment, when converting the sound source position information in at least one time frame into at least one set of image data describing the position distribution of the sound source to form the image stream, the conversion module 81 is configured to:

[0199] Based on a preset recognition accuracy, determining a target time frame within a current recognition period from at least one time frame;

[0200] Convert the sound source orientation information under each target time frame into a set of image data describing the orientation distribution state of the sound source within the current recognition period;

[0201] Continue to determine the time frame and image data within the next recognition period in the target period until image data corresponding to all recognition periods in the target period are generated to form an image stream.

[0202] In an optional embodiment, the determination module 80 includes an acquisition module 83 and a calculation module 84;

[0203] An acquisition module 83 is configured to acquire an acoustic signal stream collected by the microphone array in at least one time frame;

[0204] The calculation module 84 is configured to perform sound source direction estimation based on the acoustic signal stream to obtain sound source direction information in at least one time frame.

[0205] In an optional embodiment, when converting the sound source position information in each target time frame into a set of image data describing the position distribution state of the sound source in the current recognition period, the conversion module 81 is configured to:

[0206] The sound source orientation information in each target time frame is converted into an orientation distribution heat map of the sound source in at least one time frame. The orientation distribution heat map is used to describe the distribution heat of the sound source in different orientations in at least one time frame.

[0207] In an optional embodiment, the sound source orientation information includes the confidence level of the sound source in each orientation. When converting the sound source orientation information in each target time frame into a heat map of the orientation distribution of the sound source in at least one time frame, the conversion module 81 is configured to:

[0208] Based on the correspondence between confidence and display brightness, according to the confidence of the sound source in each direction in at least one time frame, the display brightness corresponding to each direction is determined in at least one time frame, and different display brightness represents different distribution heat;

[0209] Generate a heat map of the azimuth distribution of the sound source in at least one time frame according to the display brightness.

[0210] In an optional embodiment, when generating the azimuth distribution heat map of the sound source in at least one time frame according to the display brightness, the conversion module 81 is configured to:

[0211] Determining the image content corresponding to each of the at least one time frame according to the display brightness corresponding to each direction in the at least one time frame;

[0212] According to the time sequence between at least one time frame, the image content corresponding to each of the at least one time frame is arranged in sequence to generate an orientation distribution heat map.

[0213] In an optional embodiment, if the image data is a heat map of the azimuth distribution of the sound source in at least one time frame, the tracking module 82, when performing image recognition on the image stream using the image recognition model to track the sound source within the target time period, is configured to:

[0214] In the image recognition model, image features are extracted from the orientation distribution heat map;

[0215] Based on the mapping relationship between image features and sound source attribute parameters and the image features extracted from the azimuth distribution heat map, the target sound source attribute parameters in at least one time frame are determined to perform sound source tracking.

[0216] In an optional embodiment, the sound source attribute parameters include one or more of direction, quantity, sound emission duration and covered time frame.

[0217] In an optional embodiment, the tracking module 82 is further configured to:

[0218] Obtaining sample heat maps corresponding to each of several sample time frame groups, where the sample heat maps are used to describe the distribution heat of the sound source at different directions in the sample time frames;

[0219] Mark the sound source attribute parameters for each sample heat map to obtain the annotation information corresponding to each sample heat map;

[0220] The heat map of each sample and its corresponding annotation information are input into the image recognition model so that the image recognition model can learn the mapping relationship between image features and sound source attribute parameters.

[0221] In an optional embodiment, when converting the sound source position information in each target time frame into a heat map of the sound source position distribution in at least one time frame, the tracking module 82 is configured to:

[0222] Feed the information stream into the image recognition model;

[0223] In the image recognition model, the sound source orientation information in each target time frame is converted into a heat map of the orientation distribution of the sound source in at least one time frame.

[0224] In an optional embodiment, the tracking module 82 is further configured to:

[0225] Get the sample information stream corresponding to each sample time frame group;

[0226] Labeling sound source attribute parameters for each sample information stream to obtain labeling information corresponding to each sample information stream;

[0227] Each sample information stream and its corresponding annotation information are input into the image recognition model so that the image recognition model can convert each sample information stream into visual data describing the azimuth distribution state of the sound source and learn the mapping relationship between image features and sound source attribute parameters.

[0228] In an optional embodiment, the acoustic signal stream includes a time domain signal stream collected by each array element in the microphone array. When the calculation module 84 performs sound source orientation estimation based on the acoustic signal stream to obtain sound source orientation information in at least one time frame, it is configured to:

[0229] Convert the time domain signal stream collected by each array element into time-frequency domain signals respectively;

[0230] The sound source direction estimation technology is used to determine the sound source direction information in at least one time frame based on the time-frequency domain signals of each array element.

[0231] In an optional embodiment, the sound source direction estimation technology includes one or more of the steerable beam response phase transform technology SRP-PHAT, the generalized cross-correlation phase transform technology GCC-PHAT or the multiple signal classification technology MUSIC.

[0232] It is worth noting that the technical details of the above-mentioned embodiments of the sound source tracking device can be referred to the aforementioned Figure 1 and Figure 7 To save space, the relevant descriptions in each embodiment of the associated sound source tracking method will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0233] Figure 9 A schematic diagram of another computing device provided by an exemplary embodiment of the present application, referring to Figure 9 , the computing device includes: a memory 90 and a processor 91.

[0234] The processor 91 is coupled to the memory 90 and is configured to execute the computer program in the memory 90 to:

[0235] Determining the sound source location information in at least one time frame within the target period;

[0236] Converting the sound source orientation information in at least one time frame into at least one set of image data describing the orientation distribution state of the sound source to form an image stream;

[0237] An image recognition model is used to perform image recognition on the image stream to track the sound source within the target time period.

[0238] In an optional embodiment, when converting the sound source position information in at least one time frame into at least one set of image data describing the position distribution state of the sound source to form an image stream, the processor 91 is configured to:

[0239] Based on a preset recognition accuracy, determining a target time frame within a current recognition period from at least one time frame;

[0240] Convert the sound source orientation information under each target time frame into a set of image data describing the orientation distribution state of the sound source within the current recognition period;

[0241] Continue to determine the time frame and image data within the next recognition period in the target period until image data corresponding to all recognition periods in the target period are generated to form an image stream.

[0242] In an optional embodiment, when the processor 91 determines the sound source position information in at least one time frame within the target time period, it is configured to:

[0243] Acquiring an acoustic signal stream collected by a microphone array in at least one time frame;

[0244] The sound source direction is estimated based on the acoustic signal stream to obtain the sound source direction information in at least one time frame.

[0245] In an optional embodiment, when converting the sound source position information in each target time frame into a set of image data describing the position distribution state of the sound source in the current recognition period, the processor 91 is configured to:

[0246] The sound source orientation information in each target time frame is converted into an orientation distribution heat map of the sound source in at least one time frame. The orientation distribution heat map is used to describe the distribution heat of the sound source in different orientations in at least one time frame.

[0247] In an optional embodiment, the sound source orientation information includes the confidence level of the sound source in each orientation; when converting the sound source orientation information in each target time frame into a heat map of the orientation distribution of the sound source in at least one time frame, the processor 91 is configured to:

[0248] Based on the correspondence between confidence and display brightness, according to the confidence of the sound source in each direction in at least one time frame, the display brightness corresponding to each direction is determined in at least one time frame, and different display brightness represents different distribution heat;

[0249] Generate a heat map of the azimuth distribution of the sound source in at least one time frame according to the display brightness.

[0250] In an optional embodiment, when generating the azimuth distribution heat map of the sound source in at least one time frame according to the display brightness, the processor 91 is configured to:

[0251] Determining the image content corresponding to each of the at least one time frame according to the display brightness corresponding to each direction in the at least one time frame;

[0252] According to the time sequence between at least one time frame, the image content corresponding to each of the at least one time frame is arranged in sequence to generate an orientation distribution heat map.

[0253] In an optional embodiment, if the image data is a heat map of the azimuth distribution of the sound source in at least one time frame, the processor 91, when performing image recognition on the image stream using the image recognition model to track the sound source within the target time period, is configured to:

[0254] In the image recognition model, image features are extracted from the orientation distribution heat map;

[0255] Based on the mapping relationship between image features and sound source attribute parameters and the image features extracted from the azimuth distribution heat map, the target sound source attribute parameters in at least one time frame are determined to perform sound source tracking.

[0256] In an optional embodiment, the sound source attribute parameters include one or more of direction, quantity, sound emission duration and covered time frame.

[0257] In an optional embodiment, the processor 91 is further configured to:

[0258] Obtaining sample heat maps corresponding to each of several sample time frame groups, where the sample heat maps are used to describe the distribution heat of the sound source at different directions in the sample time frames;

[0259] Mark the sound source attribute parameters for each sample heat map to obtain the annotation information corresponding to each sample heat map;

[0260] The heat map of each sample and its corresponding annotation information are input into the image recognition model so that the image recognition model can learn the mapping relationship between image features and sound source attribute parameters.

[0261] In an optional embodiment, when converting the sound source orientation information in each target time frame into a heat map of the orientation distribution of the sound source in at least one time frame, the processor 91 is configured to:

[0262] Feed the information stream into the image recognition model;

[0263] In the image recognition model, the sound source orientation information in each target time frame is converted into a heat map of the orientation distribution of the sound source in at least one time frame.

[0264] In an optional embodiment, the processor 91 is further configured to:

[0265] Get the sample information stream corresponding to each sample time frame group;

[0266] Labeling sound source attribute parameters for each sample information stream to obtain labeling information corresponding to each sample information stream;

[0267] Each sample information stream and its corresponding annotation information are input into the image recognition model so that the image recognition model can convert each sample information stream into visual data describing the azimuth distribution state of the sound source and learn the mapping relationship between image features and sound source attribute parameters.

[0268] In an optional embodiment, the acoustic signal stream includes a time domain signal stream collected by each array element in the microphone array. When the processor 91 performs sound source orientation estimation based on the acoustic signal stream to obtain sound source orientation information in at least one time frame, it is configured to:

[0269] Convert the time domain signal stream collected by each array element into time-frequency domain signals respectively;

[0270] The sound source direction estimation technology is used to determine the sound source direction information in at least one time frame based on the time-frequency domain signals of each array element.

[0271] In an optional embodiment, the sound source direction estimation technology includes one or more of the steerable beam response phase transform technology SRP-PHAT, the generalized cross-correlation phase transform technology GCC-PHAT or the multiple signal classification technology MUSIC.

[0272] It is worth noting that for the technical details of the above embodiments of the computing device, please refer to the aforementioned Figure 1 and Figure 7 To save space, the relevant descriptions in each embodiment of the associated sound source tracking method will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0273] Further, if Figure 9 As shown, the computing device also includes: a communication component 93, a power supply component 94 and other components. Figure 9 Only some components are shown schematically, and it does not mean that the computing device only includes Figure 9 Components shown.

[0274] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be executed by a computing device in the above method embodiment.

[0275] Figure 10 A schematic diagram of a sound source tracking system provided by an exemplary embodiment of the present application. Figure 10 The sound source tracking system may include: a microphone array 10 and a computing device 20, and the microphone array 10 and the computing device 20 are communicatively connected.

[0276] The sound source tracking system provided in this embodiment can be applied in various scenarios, such as voice control scenarios, audio and video conferencing scenarios, or other scenarios requiring sound source tracking. This embodiment does not limit the application scenarios. In different application scenarios, the sound source tracking system provided in this embodiment can be integrated and deployed in a variety of scene devices. For example, in voice control scenarios, it can be deployed in smart speakers and smart robots, and in audio and video conferencing scenarios, it can be deployed in various conference terminals.

[0277] The microphone array 10 can be used to collect acoustic signals. In this embodiment, there is no limitation on the number and arrangement of elements in the microphone array 10.

[0278] For technical details about computing devices, please refer to Figure 6 and Figure 9 To save space, the relevant descriptions in the associated embodiments will not be repeated here, but this should not cause any loss to the scope of protection of this application.

[0279] above Figure 6 and Figure 9 The memory in the computing platform is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination of them, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0280] above Figure 6 and Figure 9The communication component in is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0281] above Figure 6 and Figure 9 The power supply component in a device provides power to various components of the device in which the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component is located.

[0282] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0283] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0284] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0285] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0286] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0287] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0288] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0289] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0290] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included in the protection scope of the present application.

Claims

1. A sound source tracking method, characterized in that: include: Acquiring an acoustic signal stream collected by a microphone array in at least one time frame; Performing sound source orientation estimation based on the acoustic signal stream to obtain an information stream including sound source orientation information in the at least one time frame, wherein the sound source orientation information includes confidence that the sound source is in each orientation; Based on the correspondence between confidence and display brightness, according to the confidence of the sound source in each orientation in the at least one time frame, the display brightness corresponding to each orientation in the at least one time frame is determined, and different display brightness represents different distribution heat; generating a heat map of the azimuth distribution of the sound source in the at least one time frame according to the display brightness; In a machine learning model, extracting image features from the orientation distribution heat map; Based on the mapping relationship between the image features and the sound source attribute parameters, the target sound source attribute parameters in the at least one time frame are determined to perform sound source tracking.

2. The method according to claim 1, characterized in that Generating a heat map of the azimuth distribution of the sound source in the at least one time frame according to the display brightness includes: Determining the image content corresponding to each of the at least one time frame according to the display brightness corresponding to each direction in the at least one time frame; The image contents corresponding to the at least one time frame are arranged in sequence according to the time sequence between the at least one time frame to generate the orientation distribution heat map.

3. The method according to claim 1, characterized in that The sound source attribute parameters include one or more of direction, quantity, sounding duration and covered time frame.

4. The method according to claim 1, wherein Also includes: Obtaining a sample heat map corresponding to each of the plurality of sample time frame groups, wherein the sample heat map is used to describe the distribution heat of the sound source at different directions in the sample time frame; Mark the sound source attribute parameters for each sample heat map to obtain the annotation information corresponding to each sample heat map; The sample heat maps and their corresponding annotation information are input into the machine learning model so that the machine learning model can learn the mapping relationship between the image features and the sound source attribute parameters.

5. The method according to claim 1, wherein Also includes: Get the sample information stream corresponding to each sample time frame group; Labeling sound source attribute parameters for each sample information stream to obtain labeling information corresponding to each sample information stream; The sample information streams and their corresponding annotation information are input into the machine learning model so that the machine learning model can convert the sample information streams into visual data describing the azimuth distribution state of the sound source and learn the mapping relationship between the image features and the sound source attribute parameters.

6. The method according to claim 5, characterized in that The converting of the information stream into visual data describing the azimuth distribution state of the sound source comprises: inputting the information stream into a machine learning model; In the machine learning model, the information flow is converted into visual data describing the azimuth distribution state of the sound source.

7. The method according to claim 1, characterized in that The acoustic signal stream includes a time domain signal stream collected by each array element in the microphone array, and the sound source orientation estimation is performed based on the acoustic signal stream to obtain an information stream including the sound source orientation information in the at least one time frame, including: Convert the time domain signal stream collected by each array element into time-frequency domain signals respectively; A sound source direction estimation technology is adopted to determine the sound source direction information in the at least one time frame according to the time-frequency domain signals of each array element.

8. The method according to claim 7, characterized in that The sound source direction estimation technology includes one or more of the following: steerable beam response phase transform technology SRP-PHAT, generalized cross-correlation phase transform technology GCC-PHAT or multiple signal classification technology MUSIC.

9. The method according to claim 1, characterized in that Also includes: An image stream is formed based on the azimuth distribution heat map of the sound source in the at least one time frame.

10. The method according to claim 9, characterized in that The method forms an image stream based on the azimuth distribution heat map of the sound source in the at least one time frame, including: Based on a preset recognition accuracy, determining a target time frame within a current recognition period from the at least one time frame; Converting the sound source orientation information in each target time frame into a set of orientation distribution heat maps describing the orientation distribution state of the sound source in the current recognition period; Continue to determine the time frame and the orientation distribution heat map within the next recognition period in the target period until the orientation distribution heat maps corresponding to all recognition periods in the target period are generated to form the image stream.

11. A sound source tracking device, characterized in that: include: An acquisition module, configured to acquire an acoustic signal stream collected by a microphone array in at least one time frame; a calculation module, configured to perform sound source orientation estimation based on the acoustic signal stream to obtain an information stream including sound source orientation information in the at least one time frame, wherein the sound source orientation information includes confidence that the sound source is in each orientation; a conversion module, configured to determine, based on a correspondence between confidence and display brightness, display brightness corresponding to each orientation in the at least one time frame according to the confidence of the sound source in each orientation in the at least one time frame, where different display brightnesses represent different distribution heats; generating a heat map of the azimuth distribution of the sound source in the at least one time frame according to the display brightness; The tracking module extracts image features from the azimuth distribution heat map in a machine learning model; and determines the target sound source attribute parameters in the at least one time frame based on the mapping relationship between the image features and the sound source attribute parameters to perform sound source tracking.

12. A computing device, characterized in that including memory and processor; The memory is used to store one or more computer instructions; The processor is coupled to the memory and configured to execute the one or more computer instructions for: Acquiring an acoustic signal stream collected by a microphone array in at least one time frame; Performing sound source orientation estimation based on the acoustic signal stream to obtain an information stream including sound source orientation information in the at least one time frame, wherein the sound source orientation information includes confidence that the sound source is in each orientation; Based on the correspondence between confidence and display brightness, according to the confidence of the sound source in each orientation in the at least one time frame, the display brightness corresponding to each orientation in the at least one time frame is determined, and different display brightness represents different distribution heat; generating a heat map of the azimuth distribution of the sound source in the at least one time frame according to the display brightness; In a machine learning model, extracting image features from the orientation distribution heat map; Based on the mapping relationship between the image features and the sound source attribute parameters, the target sound source attribute parameters in the at least one time frame are determined to perform sound source tracking.

13. A sound source tracking device, characterized in that: include: a determination module, configured to determine, in at least one time frame within a target period, sound source position information, wherein the sound source position information includes a confidence level that the sound source is at each position; a conversion module, configured to convert the sound source orientation information in the at least one time frame into at least one set of orientation distribution heat maps describing the orientation distribution state of the sound source; The tracking module uses an image recognition model to extract image features from the azimuth distribution heat map; based on the mapping relationship between the image features and the sound source attribute parameters, determines the target sound source attribute parameters in the at least one time frame to track the sound source within the target time period.

14. A sound source tracking system, characterized in that: include: a microphone array and a computing device, wherein the microphone array is communicatively connected to the computing device; The microphone array is used to collect acoustic signals; The computing device is configured to obtain an acoustic signal stream collected by the microphone array in at least one time frame; Performing sound source orientation estimation based on the acoustic signal stream to obtain an information stream including sound source orientation information in the at least one time frame, wherein the sound source orientation information includes confidence that the sound source is in each orientation; Based on the correspondence between confidence and display brightness, according to the confidence of the sound source in each orientation in the at least one time frame, the display brightness corresponding to each orientation in the at least one time frame is determined, and different display brightness represents different distribution heat; Based on the display brightness, an azimuth distribution heat map of the sound source in the at least one time frame is generated. Then, in the machine learning model, image features in the azimuth distribution heat map are extracted; based on the mapping relationship between the image features and the sound source attribute parameters, the target sound source attribute parameters in the at least one time frame are determined to perform sound source tracking.

15. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the sound source tracking method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Acoustic imaging method, device and equipment and readable storage medium

    CN111443330A

  • Acoustic scene reconstruction device, acoustic scene reconstruction method, and program

    US20200066023A1