Sound source localization method based on array sensing terminal data

By collecting sound source environmental data within a sliding time window, updating the sound source feature relationship database, and training the sound source localization model, the problem of inaccurate sound source localization in scenarios with frequent changes in multiple sound sources is solved, achieving more accurate sound source localization and fault detection.

CN120908753APending Publication Date: 2025-11-07HUANENG YANGPU THERMAL POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511026791.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing sound source localization technologies are inaccurate in complex scenarios with multiple sound sources and frequent changes in the types of sound sources.

Method used

By collecting sound source environmental data within a sliding time window, obtaining sound source change data, updating the sound source feature relationship database, and training a sound source localization model, sound source localization in the target scene can be achieved.

Benefits of technology

It improves the accuracy of sound source localization, enables timely detection of equipment malfunctions, and solves the problem of inaccurate sound source localization in scenarios with frequent changes in multiple sound sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120908753A_ABST
    Figure CN120908753A_ABST
Patent Text Reader

Abstract

The invention discloses a sound source positioning method based on array sensing terminal data, and relates to the technical field of sound source positioning. The method comprises the following steps: acquiring target scene sound source environment data in a sliding time window, and acquiring sound source change data based on the sound source environment data; the sound source change data comprises sound source increase and decrease data and sound source position change data; updating a sound source characteristic relation database of the target scene based on the sound source change data; training a sound source positioning model based on the updated sound source characteristic relation database of the target scene; and positioning each target sound source in the target scene through the trained sound source positioning model. The technical problem of inaccurate sound source positioning in a complex scene with multiple sound sources and frequent change of sound source types in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound source positioning, in particular to a sound source positioning method based on array sensor terminal data. BACKGROUND

[0002] Sound source positioning is to obtain sound source signals through a microphone array, analyze and process the signals, and finally estimate the position of the sound source. It has been widely used in professional fields such as industrial production, military, aerospace, etc. The existing sound source positioning technology is mainly aimed at single sound source or fixed sound source type scene.

[0003] However, in a complex scene with multiple sound sources and frequent changes in sound source types, the positioning performance of the existing sound source positioning technology will be affected.

[0004] Therefore, the inaccurate sound source positioning in a complex scene with multiple sound sources and frequent changes in sound source types is a technical problem that needs to be solved at present. SUMMARY

[0005] The present application aims to provide a sound source positioning method based on array sensor terminal data, which solves the technical problem of inaccurate sound source positioning in a complex scene with multiple sound sources and frequent changes in sound source types in the prior art.

[0006] The present application provides a sound source positioning method based on array sensor terminal data, which comprises:

[0007] Collecting sound source environment data of the target scene in a sliding time window, and obtaining sound source change data based on the sound source environment data; the sound source change data includes sound source increase / decrease data and sound source position change data;

[0008] Updating the sound source feature relationship database of the target scene based on the sound source change data;

[0009] Training a sound source positioning model based on the updated sound source feature relationship database of the target scene;

[0010] Positioning each target sound source in the target scene through the trained sound source positioning model.

[0011] Further, collecting sound source environment data of the target scene in a sliding time window, and obtaining sound source change data based on the sound source environment data, including:

[0012] Comparing the image data of the target scene at the beginning and end of the sliding time window to obtain device change data; the device change data includes device increase / decrease data and device position change data; each device corresponds to several sound sources;

[0013] Obtaining sound source increase / decrease data based on device increase / decrease data;

[0014] Based on the device position change data, obtain the sound source position change data.

[0015] Further, the device increase and decrease data includes a plurality of newly added devices, and the device model and device feature point scene coordinates of each newly added device in the target scene; based on the device increase and decrease data, obtain the sound source increase and decrease data, including:

[0016] Based on the device model of the newly added device, obtain a plurality of device feature sound sources, and the position data of each device feature sound source relative to the device feature point;

[0017] Based on the position data of the device feature sound source relative to the device feature point and the feature point scene coordinates, obtain the sound source coordinates of the device feature sound source in the target scene and mark the newly added sound source label; the newly added sound source label includes the device name and the sound source name.

[0018] Further, the device increase and decrease data further includes a plurality of newly reduced devices, and a plurality of sound source data pairs corresponding to each newly reduced device; the sound source data pair includes the sound source feature, the sound source coordinates, the sound source label, and the corresponding relationship among the sound source feature, the sound source coordinates, and the sound source label; based on the device increase and decrease data, obtain the sound source increase and decrease data, including:

[0019] Mark the newly reduced label in the sound source label of each sound source data pair of the newly reduced device to obtain the newly reduced sound source data pair.

[0020] Further, the device position change data includes a plurality of mobile devices, and the sound source data pair before the position change and the current sound source coordinates of each feature sound source after the position change of each mobile device; based on the device position change data, obtain the sound source position change data, including:

[0021] Exchange the original sound source coordinates in the sound source data pair before the position change of the device feature sound source to the current sound source coordinates to obtain the mobile sound source data pair; the mobile sound source data pair includes the original sound source feature, the current sound source coordinates, the original sound source label, and the corresponding relationship among the original sound source feature, the current sound source coordinates, and the original sound source label;

[0022] The sound source label includes the device name, the sound source name, and the sound source type.

[0023] Further, based on the sound source change data, update the sound source feature relationship database of the target scene, including:

[0024] Perform a deletion operation on the plurality of newly reduced sound source data pairs in the sound source feature relationship database, and train the sound source positioning neural network through the deleted sound source feature relationship database;

[0025] The trained sound source localization neural network is used to locate several sound sources in the target scene collected by the array sensor terminal, and obtain the sound source coordinates and sound source features of several known sound sources and several unknown sound sources.

[0026] Based on the coordinates and features of unknown sound sources, the data pairs of newly added or moving sound sources are matched to obtain the data pairs of sound sources to be updated.

[0027] Update the sound source data pairs to be updated to the sound source feature relation database.

[0028] Furthermore, based on the coordinates and features of the unknown sound sources, newly added sound sources are matched to obtain sound source data pairs to be updated, including:

[0029] Based on the coordinates of the unknown sound source and the coordinates of the newly added sound source, obtain the coordinate similarity.

[0030] When the coordinate similarity is greater than the preset similarity, the sound source features of the unknown sound source are associated with the sound source coordinates of the newly added sound source and the newly added sound source label.

[0031] The associated sound source features, the coordinates of newly added sound sources, and the labels of newly added sound sources are verified and the sound source types are marked to obtain the sound source data pairs to be updated.

[0032] Furthermore, based on the coordinates and features of the unknown sound sources, the moving sound source data pairs are matched to obtain the sound source data pairs to be updated, including:

[0033] Based on the coordinates of the unknown sound source and the current coordinates of the moving sound source in the data pair, obtain the coordinate similarity.

[0034] Based on the sound source features of the unknown sound source and the original sound source features in the moving sound source data pair, the feature similarity is obtained.

[0035] Based on coordinate similarity and feature similarity, determine whether a match is successful;

[0036] If so, update the sound source characteristics of the unknown sound source to the moving sound source data pair, and obtain the sound source data pair to be updated.

[0037] Furthermore, updating the sound source feature relationship database of the target scene based on sound source variation data also includes:

[0038] Whenever there is a pair of sound source data to be updated and updated to the sound source feature relationship database, the sound source localization neural network is trained through the sound source feature relationship database.

[0039] Furthermore, based on the updated target scene's sound source feature relationship database, a sound source localization model is trained, including:

[0040] When the newly added sound source data in the sound source feature relationship database meets the preset updating requirement, the sound source positioning model is trained.

[0041] Compared with the prior art, the present application has the following advantages:

[0042] In this embodiment, the sound source variation data is obtained by analyzing the sound source environment data of the target scene in the sliding time window, to obtain the sound source variation in the target scene; wherein the sound source variation includes sound source increase / decrease and sound source position variation. When sound source increase / decrease or sound source position variation occurs in the target scene, the sound source feature relationship database of the target scene needs to be updated in a timely manner according to the sound source variation, so as to further train the sound source positioning model in a timely manner based on the updated sound source feature relationship data of the target scene, so that the positioning of each target sound source in the next stage of the target scene is more accurate. Thus, the fault of the discovery device is timely and accurately obtained. The technical problem of inaccurate sound source positioning in the complex scene with multiple sound sources and frequent sound source type changes in the prior art is solved. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 The method steps of the sound source positioning method based on array sensor terminal data of the present application are shown in the figure. DETAILED DESCRIPTION

[0044] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme of the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0045] As shown in the figure, the sound source positioning method based on array sensor terminal data includes the following steps: Figure 1

[0046] S1: collecting sound source environment data of a target scene in a sliding time window, and obtaining sound source variation data based on the sound source environment data; the sound source variation data includes sound source increase / decrease data and sound source position variation data;

[0047] S2: updating the sound source feature relationship database of the target scene based on the sound source variation data;

[0048] S3: training a sound source positioning model based on the updated sound source feature relationship database of the target scene;

[0049] S4: positioning each target sound source in the target scene through the trained sound source positioning model.

[0050] ​The specific implementation process of the embodiment includes:

[0051] In the embodiment, the sound source variation data is obtained by analyzing the sound source environment data of the target scene in the sliding time window, to obtain the sound source variation in the target scene; the sound source variation includes sound source increase / decrease and sound source position variation. When the sound source increase / decrease or the sound source position variation occurs in the target scene, the sound source feature relationship database of the target scene needs to be updated in time according to the sound source variation, so as to further train the sound source positioning model through the updated sound source feature relationship data of the target scene, so that the positioning of each target sound source in the next stage of the target scene is more accurate. Thus, the fault of the discovery device is timely and accurately obtained. The technical problem of inaccurate sound source positioning in the complex scene with multiple sound sources and frequent sound source type variation in the prior art is solved.

[0052] In the embodiment, the sound source environment data of the target scene is collected in the sliding time window, and the sound source variation data is obtained based on the sound source environment data, including:

[0053] S11: comparing the image data of the target scene at the beginning and the end of the sliding time window to obtain device variation data; the device variation data includes device increase / decrease data and device position variation data; each device corresponds to a plurality of sound sources;

[0054] In the embodiment, the image data of the target scene is collected once for each movement of the sliding time window; the beginning and the end of the sliding time window in the embodiment refer to the time point corresponding to the beginning of the sliding time window and the time point corresponding to the end of the sliding time window. Comparing the image data at the beginning and the end of the sliding time window avoids the situation that the device variation is not obvious due to the similar image data collected at adjacent times. In some embodiments, when the device in the target scene is changed, the management personnel controls the image collection device to actively collect the image data before and after the device variation of the target scene. Before comparison, the device data in the image data is identified and obtained, including the device model and the feature point scene coordinates; the feature point scene coordinates include the coordinate data of the device feature points in the space rectangular coordinate system constructed based on the target scene; the device feature points in the embodiment include the vertices of the device surface. When comparing, if the device in the original target scene does not exist in the current target scene, it is judged that the device is newly decreased; if the device in the original target scene has inconsistent position in the current target scene, it is judged that the device has position variation; if a device exists in the current target scene but does not exist in the original target scene, it is judged that the device is newly added.

[0055] S12: obtaining sound source increase / decrease data based on the device increase / decrease data;

[0056] S13: obtaining sound source position variation data based on the device position variation data.

[0057] For step S12, the device increase / decrease data in this embodiment includes a plurality of newly added devices, and device model and feature point scene coordinates of feature points of each newly added device in the target scene; based on the device increase / decrease data, sound source increase / decrease data is obtained, including:

[0058] S121: based on the device model of the newly added device, a plurality of device feature sound sources are obtained, and position data of each device feature sound source relative to the device feature point;

[0059] In this embodiment, each device of each device model is pre-provided with a plurality of corresponding device feature sound sources; the plurality of device feature sound sources are distributed at a plurality of positions of the device; the position data of each device feature sound source relative to the device feature point is unchanged, and the position data includes distance data, pitch angle data, etc.

[0060] S122: based on the position data of the device feature sound source relative to the device feature point and the feature point scene coordinates, the sound source coordinates of the device feature sound source in the target scene are obtained and a newly added sound source label is marked; the newly added sound source label includes the device name and the sound source name.

[0061] In this embodiment, based on the position data of the device feature sound source relative to a plurality of device feature points and the coordinate data of each device feature point, the sound source coordinates of the device feature sound source in the target scene are obtained; the coordinates of the newly added sound source are marked with a newly added sound source label, and the newly added sound source label includes the device name and the sound source name.

[0062] In this embodiment, the device increase / decrease data further includes a plurality of newly reduced devices, and a plurality of sound source data pairs corresponding to each newly reduced device; the sound source data pair includes sound source features, sound source coordinates, sound source labels, and the corresponding relationship of the sound source features, sound source coordinates, and sound source labels; based on the device increase / decrease data, sound source increase / decrease data is obtained, including:

[0063] S123: a newly reduced label is marked in the sound source label of each sound source data pair of the newly reduced device, and a newly reduced sound source data pair is obtained.

[0064] In this embodiment, the sound source feature relationship database includes a plurality of sound source data pairs, and the sound source data pair includes sound source features, sound source coordinates, sound source labels, and the corresponding relationship of the sound source features, sound source coordinates, and sound source labels.

[0065] The sound source label includes the sound source name, the device name, and the sound source type. In this embodiment, the sound source type includes normal and abnormal reasons. Each device feature sound source corresponds to at least one sound source data pair.

[0066] For step S13, the device position change data in this embodiment includes a plurality of mobile devices, and the sound source data pair before the position change of each mobile device and the current sound source coordinates of each feature sound source after the position change; based on the device position change data, the sound source position change data is obtained, including:

[0067] S131: the original sound source coordinates in the sound source data pair before the position change of the device feature sound source are exchanged into the current sound source coordinates to obtain a mobile sound source data pair; the mobile sound source data pair includes the original sound source feature, the current sound source coordinates, the original sound source label, and the corresponding relationship among the original sound source feature, the current sound source coordinates, and the original sound source label;

[0068] S132: the sound source label includes the device name, the sound source name, and the sound source type.

[0069] In this embodiment, the sound source feature includes the voiceprint feature, and for the device feature sound source after the position change, the coordinates in the sound source data pair before the position change are temporarily exchanged into the current sound source coordinates; it is necessary to collect the sound source feature after the position change to update the mobile sound source data pair to construct the formal sound source data pair.

[0070] In this embodiment, based on the sound source change data, the sound source feature relationship database of the target scene is updated, including:

[0071] S21: a plurality of new sound source data pairs are deleted in the sound source feature relationship database, and the sound source positioning neural network is trained through the sound source feature relationship database after the deletion;

[0072] In this embodiment, for the new sound source, it is directly deleted in the sound source feature relationship database, and the sound source positioning neural network is retrained to avoid the output error of the sound source positioning neural network. In this embodiment, the output of the sound source positioning neural network includes the sound source coordinates and the sound source label, and the input of the sound source positioning neural network includes the sound source feature. At the same time, it can also avoid missing unknown sound sources.

[0073] S22: the sound source positioning neural network is trained to perform sound source positioning on a plurality of sound sources collected by the array sensing terminal in the target scene to obtain the sound source coordinates and the sound source feature of a plurality of known sound sources and a plurality of unknown sound sources;

[0074] In this embodiment, the array sensing terminal includes an array composed of a plurality of microphones, which is used to collect a plurality of sound sources in the target scene in real time, including obtaining the sound source feature of each sound source in the target scene.

[0075] In addition to obtaining the audio data of the target scene, the array sensing terminal also includes obtaining the sound source feature of each sound source based on the audio data. The data processing process is realized through related technologies, which will not be described here.

[0076] In this embodiment, the sound source features include a voiceprint feature vector and a positioning feature vector; the positioning feature vector includes a phase difference feature and a spatial covariance feature; and the voiceprint feature vector includes a spectral envelope feature, a time-domain transient feature, and a time-frequency detail feature.

[0077] S23: Based on the sound source coordinates of the unknown sound source and the sound source features, match the new sound source or the mobile sound source data pair to obtain a to-be-updated sound source data pair.

[0078] S24: Update the to-be-updated sound source data pair to the sound source feature relationship database.

[0079] In this embodiment, each time the to-be-updated sound source data pair is updated to the sound source feature relationship database, the sound source positioning neural network is trained through the sound source feature relationship database. The sound source positioning neural network after training is used for sound source positioning of subsequent sound sources collected by the array sensing terminal.

[0080] For step S23, based on the sound source coordinates of the unknown sound source and the sound source features, match the new sound source to obtain a to-be-updated sound source data pair, including:

[0081] S231: Based on the sound source coordinates of the unknown sound source and the sound source coordinates of the new sound source, obtain a coordinate similarity.

[0082] In this embodiment, obtaining the coordinate similarity includes calculating the Euclidean distance according to the sound source coordinates of the unknown sound source and the sound source coordinates of the new sound source, and taking the reciprocal of the Euclidean distance as the coordinate similarity.

[0083] S232: When the coordinate similarity is greater than a preset similarity, associate the sound source features of the unknown sound source with the sound source coordinates of the new sound source, and the new sound source label.

[0084] S233: Verify the associated sound source features, the sound source coordinates of the new sound source, and the new sound source label, and mark the sound source type to obtain a to-be-updated sound source data pair.

[0085] In this embodiment, the associated sound source features, the sound source coordinates of the new sound source, and the new sound source label are verified and the sound source type is marked through expert experience. The sound source label in the to-be-updated sound source data pair is more accurate.

[0086] For step S23, in this embodiment, based on the sound source coordinates of the unknown sound source and the sound source features, match the mobile sound source data pair to obtain a to-be-updated sound source data pair, including:

[0087] Based on the sound source coordinates of the unknown sound source and the current sound source coordinates in the mobile sound source data pair, obtain a coordinate similarity.

[0088] obtain a feature similarity based on the sound source feature of the unknown sound source and the original sound source feature in the mobile sound source data pair;

[0089] determine whether the matching is successful based on the coordinate similarity and the feature similarity;

[0090] when the coordinate similarity is greater than a preset similarity and the feature similarity is greater than a feature similarity threshold, determine that the sound source feature of the unknown sound source and the sound source coordinates and the sound source label in the mobile sound source data pair are matched successfully;

[0091] if yes, update the sound source feature of the unknown sound source to the mobile sound source data pair to obtain a to-be-updated sound source data pair.

[0092] In this embodiment, the calculation process of the feature similarity is as follows:

[0093] The voiceprint feature vector of the unknown sound source is represented as A=[As, At, Atf], wherein As is the spectral envelope feature, At is the time-domain transient feature, and Atf is the time-frequency detail feature;

[0094] The voiceprint feature vector of the mobile sound source data pair is represented as B=[Bs, Bt, Btf], wherein Bs is the spectral envelope feature, Bt is the time-domain transient feature, and Btf is the time-frequency detail feature;

[0095] The Mahalanobis distances of the spectral envelope feature, the time-domain transient feature, and the time-frequency detail feature are calculated based on the voiceprint feature vectors A and B respectively to obtain the spectral envelope distance, the time-domain transient distance, and the time-frequency detail distance, which are represented as d(As, Bs), d(At, Bt), and d(Atf, Btf) respectively.

[0096] Then, the weight coefficients α, β, and γ corresponding to the spectral envelope feature, the time-domain transient feature, and the time-frequency detail feature are obtained based on the device model of the mobile sound source, wherein α+β+γ=1.

[0097] The feature similarity is calculated, and the formula is as follows:

[0098]

[0099] wherein D is the feature similarity.

[0100] In this embodiment, based on the sound source feature relationship database of the updated target scene, the sound source positioning model is trained, including:

[0101] When the newly added sound source data pair in the sound source feature relationship database meets the preset update requirement, the sound source positioning model is trained.

[0102] The preset update requirement in the embodiment includes that the number of newly added sound source data pairs in the sound source feature relationship database is greater than a preset number of additions, and the mobile sound source data pair update reaches a first preset percentage.

[0103] It should be noted that in the embodiment, the sound source positioning model includes a neural network model; the training process of the sound source positioning neural network or the sound source positioning model can refer to the manner in the related art, including continuously optimizing the network model through the gradient direction propagation algorithm, and the like, which will not be described herein again.

[0104] It should be noted that in this document, the terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0105] The above embodiments are only used to illustrate the technical method of the present application but not limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical method of the present application.

Claims

1. A method for acoustic source localization based on array sensor terminal data, characterized in that: The method comprises: Collecting target scene sound source environment data in a sliding time window, and obtaining sound source variation data based on the sound source environment data; the sound source variation data comprises sound source increase / decrease data and sound source position variation data; Updating the sound source feature relationship database of the target scene based on the sound source variation data; Training a sound source positioning model based on the updated sound source feature relationship database of the target scene; Positioning each target sound source in the target scene through the trained sound source positioning model.

2. The method of claim 1, wherein: Collecting target scene sound source environment data in a sliding time window, and obtaining sound source variation data based on the sound source environment data, comprising: Comparing image data of the target scene at the beginning and end of the sliding time window to obtain device variation data; the device variation data comprises device increase / decrease data and device position variation data; each device corresponds to a plurality of sound sources; Obtaining sound source increase / decrease data based on the device increase / decrease data; Obtaining sound source position variation data based on the device position variation data.

3. The method of claim 2, wherein: The device increase / decrease data comprises a plurality of newly added devices, and the device model and device feature point scene coordinates of each newly added device in the target scene; based on the device increase / decrease data, the sound source increase / decrease data is obtained, comprising: Based on the device model of the newly added device, a plurality of device feature sound sources and the position data of each device feature sound source relative to the device feature point are obtained; Based on the position data of the device feature sound source relative to the device feature point and the feature point scene coordinates, the sound source coordinates of the device feature sound source in the target scene are obtained and a newly added sound source label is marked; the newly added sound source label comprises a device name and a sound source name.

4. The method of claim 3, wherein: The device increase / decrease data further comprises a plurality of newly removed devices, and a plurality of sound source data pairs corresponding to each newly removed device; the sound source data pair comprises sound source features, sound source coordinates, sound source labels, and the corresponding relationship among the sound source features, sound source coordinates, and sound source labels; Based on the device increase / decrease data, the sound source increase / decrease data is obtained, comprising: Marking a newly removed label in the sound source label of each sound source data pair of the newly removed device to obtain a newly removed sound source data pair; The sound source label comprises a device name, a sound source name, and a sound source type.

5. The method of claim 4, wherein: The device position variation data comprises a plurality of mobile devices, and the sound source data pair before the position variation and the current sound source coordinates of each feature sound source after the position variation of each mobile device; based on the device position variation data, the sound source position variation data is obtained, comprising: Swapping the original sound source coordinates in the sound source data pair before the position variation of the device feature sound source to the current sound source coordinates to obtain a mobile sound source data pair; the mobile sound source data pair comprises original sound source features, current sound source coordinates, original sound source labels, and the corresponding relationship among the original sound source features, current sound source coordinates, and original sound source labels.

6. The method of claim 5, wherein: Based on the sound source variation data, the sound source feature relationship database of the target scene is updated, comprising: Performing a deletion operation on a plurality of newly removed sound source data pairs in the sound source feature relationship database, and training a sound source positioning neural network through the deleted sound source feature relationship database; Performing sound source positioning on a plurality of sound sources of the target scene collected by the array sensing terminal through the trained sound source positioning neural network to obtain the sound source coordinates and sound source features of a plurality of known sound sources and a plurality of unknown sound sources; The unknown sound source coordinate and the sound source feature of the unknown sound source are matched with the newly added sound source or the moving sound source data pair, and a to-be-updated sound source data pair is obtained; The to-be-updated sound source data pair is updated to the sound source feature relationship database.

7. The method of claim 6, wherein: The unknown sound source coordinate and the sound source feature of the unknown sound source are matched with the newly added sound source, and a to-be-updated sound source data pair is obtained, including: The coordinate similarity is obtained based on the sound source coordinate of the unknown sound source and the sound source coordinate of the newly added sound source; When the coordinate similarity is greater than a preset similarity, the sound source feature of the unknown sound source is associated with the sound source coordinate of the newly added sound source and the newly added sound source label; The associated sound source feature, the sound source coordinate of the newly added sound source and the newly added sound source label are checked and marked with a sound source type, and a to-be-updated sound source data pair is obtained.

8. The method of claim 6, wherein the method further comprises: determining a direction of the sound source based on the array sensor terminal data. The unknown sound source coordinate and the sound source feature of the unknown sound source are matched with the moving sound source data pair, and a to-be-updated sound source data pair is obtained, including: The coordinate similarity is obtained based on the sound source coordinate of the unknown sound source and the current sound source coordinate in the moving sound source data pair; The feature similarity is obtained based on the sound source feature of the unknown sound source and the original sound source feature in the moving sound source data pair; Whether the matching is successful is judged based on the coordinate similarity and the feature similarity; If yes, the sound source feature of the unknown sound source is updated to the moving sound source data pair, and a to-be-updated sound source data pair is obtained.

9. The method of claim 6, wherein the method further comprises: determining a direction of the sound source based on the terminal data. Based on the sound source change data, the sound source feature relationship database of the target scene is updated, and further comprising: Whenever there is a to-be-updated sound source data pair updated to the sound source feature relationship database, the sound source positioning neural network is trained through the sound source feature relationship database.

10. The method of claim 6, wherein the method further comprises: determining a direction of the sound source based on the terminal data. Based on the updated sound source feature relationship database of the target scene, a sound source positioning model is trained, including: When the newly added sound source data pair in the sound source feature relationship database meets a preset updating requirement, the sound source positioning model is trained.