Position coding method for unquantized error of two-dimensional sound source localization

By employing an unbiased label distribution algorithm for encoding and decoding in sound source localization, the quantization error problem is solved, improving the accuracy and noise resistance of sound source localization, especially maintaining high-precision localization even in harsh environments.

CN117686976BActive Publication Date: 2026-08-04NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHWESTERN POLYTECHNICAL UNIV
Filing Date
2023-11-28
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing sound source localization technologies, the location encoding and decoding methods have large quantization errors, resulting in large errors in the sound source coordinate decoding process, especially under harsh conditions such as noise and reverberation, the localization accuracy is insufficient.

Method used

An unbiased label distribution algorithm is used to encode the sound source location. The space is gridded in the Cartesian coordinate system and encoded using an unbiased label distribution matrix. During decoding, a weighted approximation is performed considering the peak category and its neighboring categories to reduce quantization error.

Benefits of technology

It significantly improves the accuracy of sound source localization and reduces quantization errors, maintaining good localization performance even under noise and reverberation conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117686976B_ABST
    Figure CN117686976B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of sound source positioning, and disclose a position coding method for two-dimensional sound source positioning without quantization error, which comprises: dividing a space where a sound source is located into a plurality of grids; based on a preset resolution, respectively discretizing the space into a plurality of segments in the x-axis direction and the y-axis direction according to the lengths of the space in the x-axis direction and the y-axis direction, and determining the category of the sound source in the x-axis direction and the category of the sound source in the y-axis direction according to the coordinates of the sound source; based on the category of the sound source in the x-axis direction, using an unbiased label distribution vector to perform position coding of the sound source in the x-axis direction, and based on the category of the sound source in the y-axis direction, using an unbiased label distribution vector to perform position coding of the sound source in the y-axis direction; and based on and generating a two-dimensional unbiased label distribution matrix ρ, completing the position coding of the sound source, which can eliminate quantization error and greatly improve the accuracy of sound source positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sound source localization technology, and in particular to a position encoding and decoding method for two-dimensional sound source localization without quantization error. Background Technology

[0002] Sound source localization is a technique that uses multi-channel signals received by a microphone array to estimate the location of a sound source. It can serve as an auxiliary technology in many applications, such as human-computer interaction, drone applications, speech separation, and target speaker extraction. Sound source localization determines the spatial location of a sound source by analyzing signals received from multiple microphones.

[0003] A key step in sound source localization technology is location encoding and decoding. A commonly used method in the industry divides the room into multiple grids, treating each grid as a category, and uses one-hot encoding to label the grids. In this method, the grid containing the sound source is labeled 1, while other grids are labeled 0. During decoding, the center of the grid with the highest probability is taken as the sound source location. However, the inventors of this application have discovered that this location encoding method has a large quantization error, resulting in even greater errors in the sound source coordinates obtained during decoding. Summary of the Invention

[0004] The purpose of this application is to provide a position encoding and decoding method for two-dimensional sound source localization without quantization error, which can eliminate quantization error and greatly improve the accuracy of sound source localization, and has a good localization effect even under harsh conditions such as noise and reverberation.

[0005] To address the aforementioned technical problems, embodiments of this application provide a quantization-error-free position encoding method for two-dimensional sound source localization, comprising the following steps: establishing a Cartesian coordinate system in the space where the sound source is located, and dividing the space into several grids; based on a preset resolution, encoding the coordinates of the space according to the coordinates of the sound source in the grid. Length along the axis and in The length along the axial direction, which divides the space in axial direction and The sound source is discretized into several segments along the axial direction, and the location of the sound source is determined based on its coordinates. Category in the axial direction and in Category in the axial direction; based on the sound source in Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source Position encoding in the axial direction, and based on the sound source in Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source Position encoding in the axial direction; based on the and stated Generate a two-dimensional unbiased label distribution matrix The location encoding of the sound source is completed.

[0006] Embodiments of this application also provide a quantization-error-free location decoding method for two-dimensional sound source localization, comprising the following steps: acquiring the target sound source signals received by each microphone, inputting the target sound source signals received by each microphone into a pre-trained decoding network, and obtaining the predicted two-dimensional unbiased label distribution matrix output by the decoding network. and the target sound source in Peak category in the axial direction and in Peak category in the axial direction; wherein, the The encoding method used is the same as the quantization-error-free position encoding method for two-dimensional sound source localization described above; based on the above... Obtain the predicted unbiased label distribution vector. And predict the unbiased label distribution vector According to the above The target sound source is in Peak category and its adjacent categories in the axial direction, the and the target sound source in By analyzing the peak values ​​along the axis and their adjacent values, the coordinates of the target sound source are determined, thus completing the localization of the target sound source.

[0007] Embodiments of this application also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described error-free position encoding method for two-dimensional sound source localization, or to perform the above-described error-free position decoding method for two-dimensional sound source localization.

[0008] Embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described error-free position encoding method for two-dimensional sound source localization, or implements the above-described error-free position decoding method for two-dimensional sound source localization.

[0009] The embodiments of this application provide a quantization-error-free location encoding and decoding method for two-dimensional sound source localization. During the encoding process, an unbiased label distribution algorithm is used to encode the sound source location, transforming the problem of finding the sound source location into a classification problem. A Cartesian coordinate system is established in the space where the sound source is located to grid the space, and the space is then divided into categories. axial direction and The sound source is discretized into several segments along the axial direction, and then the sound source is analyzed based on its horizontal and vertical coordinates. axial direction and in Classification along the axis allows for the use of unbiased label distribution vectors after classification. and Encoding the location of the sound source involves a different approach. Compared to one-hot coding, the unbiased label distribution algorithm doesn't use a single number to encode the source location; instead, it uses a two-dimensional unbiased label distribution matrix containing more precise information, significantly reducing quantization errors. During decoding, a weighted adjacency decoding algorithm is used. While quantization errors are unavoidable if only peak categories are considered, this application considers multiple adjacent categories in addition to peak categories, providing a more accurate weighted approximation for sound source localization. This overcomes quantization errors during decoding, greatly reducing sound source localization errors and achieving excellent localization results even under harsh conditions such as noise and reverberation.

[0010] In some optional embodiments, the step of basing the resolution on a preset value, respectively, according to the space in... Length along the axis and in The length along the axial direction, which divides the space in axial direction and The axis is discretized into several segments, including those based on a preset resolution. According to the space respectively Length in the axial direction and in Length in the axial direction Determine the space in Segment length in the axial direction and in Segment length in the axial direction Based on the above And the space in the I mentioned above. Discretized into several segments along the axial direction, the space in Several segments along the axial direction are represented as Based on the above And the space in the I mentioned above. Discretized into several segments along the axial direction, the space in Several segments along the axial direction are represented as The step of determining the location of the sound source based on its coordinates is as follows: Category in the axial direction and in The category along the axis is determined by the following formula:

[0011] ;

[0012] ;

[0013] in, Let x be the x-coordinate of the sound source. Let be the ordinate of the sound source. This indicates that the sound source is in Categories in the axial direction, This indicates that the sound source is in Categories along the axes. Discretizing the space, that is, discretizing the space in... axial direction and The sound source location is divided into several segments along each axis. Each segment has the same length along the same axis. The sound source location can be classified by dividing its x and y coordinates by the corresponding segment length. and It is a real number, not necessarily an integer. This classification is more scientific than one-hot encoding.

[0014] In some alternative embodiments, the Represented as The Represented as The sound source based on Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source The position encoding along the axis is achieved using the following formula:

[0015] ;

[0016] ;

[0017] in, For the decimal function, For the floor function, For the The first in Item; the item based on the sound source in Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source The position encoding along the axis is achieved using the following formula:

[0018] ;

[0019] ;

[0020] in, For the The first in Item. Use and When encoding the location of a sound source, for each direction, two adjacent integers are used to approximate a real number located between them. One-hot encoding, on the other hand, rounds a real number to the nearest integer. Therefore, the unbiased label distribution algorithm proposed in this application has a much higher accuracy than one-hot encoding.

[0021] In some alternative embodiments, the statement based on the and stated Generate a two-dimensional unbiased label distribution matrix , including: the Transpose column vectors and the Transpose row vectors ; will the Multiplied by the above Generate a two-dimensional unbiased label distribution matrix The Represented as:

[0022] .

[0023] In some alternative embodiments, the for The matrix, based on the Obtain the predicted unbiased label distribution vector. And predict the unbiased label distribution vector This can be achieved through the following formula:

[0024] ;

[0025] in, Indicates the The OK, Indicates the The List.

[0026] In some alternative embodiments, the statement according to the The target sound source is in Peak category in the axial direction, the and the target sound source in The peak category along the axis is used to determine the coordinates of the target sound source, including: based on the... The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the abscissa of the target sound source; based on the... The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the ordinate of the target sound source.

[0027] In some alternative embodiments, the statement according to the The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the x-coordinate of the target sound source, which is achieved through the following formula:

[0028] ;

[0029] in, For the target sound source in Peak category in the axial direction, For the target sound source in The category to the left of the peak category along the axis. For the target sound source in The category adjacent to the right of the peak category along the axis. For the preset Segment length in the axial direction, The x-coordinate of the target sound source is obtained; the step according to the... The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the ordinate of the target sound source, which is achieved through the following formula:

[0030] ;

[0031] in, For the target sound source in Peak category in the axial direction, For the target sound source in The category to the left of the peak category along the axis. For the target sound source in The category adjacent to the right of the peak category along the axis. For the preset Segment length in the axial direction, The solution represents the ordinate of the target sound source. Attached Figure Description

[0032] One or more embodiments are illustrated by way of example with reference to the accompanying drawings, and these illustrative descriptions do not constitute a limitation on the embodiments.

[0033] Figure 1 This is a flowchart of a quantization-error-free position coding method for two-dimensional sound source localization provided in one embodiment of this application;

[0034] Figure 2 This is a schematic diagram of the space where a meshed sound source is located, provided in one embodiment of this application;

[0035] Figure 3 This is a flowchart of a quantization-error-free position decoding method for two-dimensional sound source localization provided in another embodiment of this application;

[0036] Figure 4 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0038] One embodiment of this application proposes a quantization-error-free position encoding method for two-dimensional sound source localization, applied to an electronic device, wherein the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described using a server as an example. The implementation details of the quantization-error-free position encoding method for two-dimensional sound source localization of this embodiment are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this solution.

[0039] The specific process of the error-free position encoding method for two-dimensional sound source localization provided in this embodiment can be described as follows: Figure 1 As shown, it includes:

[0040] Step 101: Establish a Cartesian coordinate system in the space where the sound source is located, and divide the space into several grids.

[0041] In the specific implementation, when the server performs location encoding, it first needs to establish a Cartesian coordinate system in the space where the sound source is located. That is, it selects the lower left corner of the space as the origin and establishes a two-dimensional rectangular coordinate system, thereby dividing the space into several grids and making the space where the sound source is located gridded.

[0042] In one example, the space containing the meshed sound source can be as follows: Figure 2 As shown.

[0043] Step 102, based on the preset resolution, according to the space in Length along the axis and in The length along the axial direction, which divides the space in axial direction and The sound source is discretized into several segments along the axial direction, and the position of the sound source is determined based on its coordinates. Category in the axial direction and in Category in the axial direction.

[0044] In a specific implementation, after dividing the space into several grids, the server, based on a preset resolution, calculates the values ​​of each grid according to the space's... Length along the axis and in The length along the axial direction, which divides the space in axial direction and The space is discretized into several segments along the axial direction. It is worth noting that the space... The space is divided into several segments of equal length along the axial direction. The segments divided along the axis are all of equal length. The server completes the space... axial direction and After discretization along the axial direction, the location of the sound source can be determined based on its coordinates. Category in the axial direction and in The categories along the axis transform the problem of finding the sound source location into a classification problem.

[0045] In one example, the server, based on a preset resolution, determines the spatial values ​​according to the specified values. Length along the axis and in The length along the axial direction, which divides the space in axial direction and When discretizing into several segments along the axis, first based on a preset resolution. According to the space respectively Length in the axial direction and in Length in the axial direction Determine the space in Segment length in the axial direction and in Segment length in the axial direction ,in, , .for The server is based on the axis direction. And the space in the I mentioned above. Discretized into several segments along the axial direction, the space in Several segments along the axial direction are represented as ;for The server is based on the axis direction. And the space in the I mentioned above. Discretized into several segments along the axial direction, the space in Several segments along the axial direction are represented as .

[0046] The server completes the space in axial direction and After discretization along the axial direction, the coordinates of the sound source are used to determine the position of the sound source in the following formula. Category in the axial direction and in Categories in the axial direction:

[0047] ;

[0048] ;

[0049] In the formula, Let x be the x-coordinate of the sound source. Let be the vertical coordinate of the sound source. Indicates the sound source is at Categories in the axial direction, Indicates the sound source is at Categories along the axis. It's worth noting that categories... and It is a real number, not necessarily an integer. This classification is more scientific than one-hot encoding.

[0050] Step 103, based on the sound source in Categories along the axes are represented by unbiased label distribution vectors. To the sound source Position encoding in the axial direction, and based on the sound source in Categories along the axes are represented by unbiased label distribution vectors. To the sound source Position encoding in the axial direction.

[0051] In the specific implementation, the server determines the sound source. Category in the axial direction and in After classifying the categories along the axis, an unbiased label distribution algorithm can be used, based on the sound source. Categories along the axes are represented by unbiased label distribution vectors. To the sound source Position encoding in the axial direction, and based on the sound source in Categories along the axes are represented by unbiased label distribution vectors. To the sound source Position encoding in the axial direction.

[0052] In one example, the Represented as The server is based on the sound source. Categories along the axes are represented by unbiased label distribution vectors. To the sound source Position encoding along the axis can be achieved using the following formula:

[0053] ;

[0054] ;

[0055] In the formula, For the decimal function, For the floor function, For the The first in This formula defines the unbiased label distribution vector. Each element in the array takes values ​​from 0 to 1, and the sum of all element values ​​is 1, so it can also be viewed as a probability distribution.

[0056] for example, ,So exist and The value will be obtained from the item, and the sound source is at The position code along the axis can be represented as: .

[0057] In one example, the Represented as The server is based on the sound source. Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source Position encoding along the axis can be achieved using the following formula:

[0058] ;

[0059] ;

[0060] In the formula, For the The first in This formula defines the unbiased label distribution vector. Each element in the array takes values ​​from 0 to 1, and the sum of all element values ​​is 1, so it can also be viewed as a probability distribution.

[0061] for example, ,So exist and The value will be obtained from the item, and the sound source is at The position code along the axis can be represented as: .

[0062] Understandably, using and When encoding the location of a sound source, for each direction, two adjacent integers are used to approximate a real number located between them. One-hot encoding, on the other hand, rounds a real number to the nearest integer. Obviously, approximating a real number with two integers is more accurate than approximating a real number with one integer. In other words, the accuracy of the unbiased label distribution algorithm is much higher than that of one-hot encoding.

[0063] Step 104, based on the above and stated Generate a two-dimensional unbiased label distribution matrix This completes the location encoding of the sound source.

[0064] In a specific implementation, the server uses the aforementioned and stated Complete the sound source at each Position encoding in the axial direction and in After the position is encoded along the axis, the... Transpose column vectors and the Transpose row vectors Then the above Multiplied by the above Generate a two-dimensional unbiased label distribution matrix This completes the location encoding of the sound source. It can be represented as:

[0065] .

[0066] In this embodiment, during the encoding process, an unbiased label distribution algorithm is used to encode the location of the sound source, transforming the problem of finding the sound source location into a classification problem. A Cartesian coordinate system is established in the space where the sound source is located to grid the space, and the space is then divided into categories. axial direction and The sound source is discretized into several segments along the axial direction, and then the sound source is analyzed based on its horizontal and vertical coordinates. axial direction and in Classification along the axis allows for the use of unbiased label distribution vectors after classification. and Encoding the location of the sound source: Compared with one-hot coding, the unbiased label distribution algorithm does not use a single number to encode the sound source location, but uses a two-dimensional unbiased label distribution matrix containing more and more accurate information, which can greatly reduce quantization error.

[0067] Another embodiment of this application proposes a quantization-error-free position decoding method for two-dimensional sound source localization. The implementation details of the quantization-error-free position decoding method for two-dimensional sound source localization in this embodiment are described in detail below. The following implementation details are provided for ease of understanding and are not necessary for implementing this solution.

[0068] The specific process of the quantization-error-free position decoding method for two-dimensional sound source localization provided in this embodiment can be described as follows: Figure 3 As shown, it includes:

[0069] Step 201: Acquire the target sound source signals received by each microphone, and input the target sound source signals received by each microphone into the pre-trained decoding network to obtain the predicted two-dimensional unbiased label distribution matrix output by the decoding network. and the target sound source in Peak category in the axial direction and in Peak category in the axial direction.

[0070] In the specific implementation, a microphone array consisting of multiple microphones is set up in the space where the target sound source is located. When locating the target sound source, the server acquires the signal of the target sound source received by each microphone and inputs the signal of the target sound source received by each microphone into a pre-trained decoding network to obtain the predicted two-dimensional unbiased label distribution matrix output by the decoding network. and the target sound source in Peak category in the axial direction and in Peak category in the axial direction Among them, the aforementioned It is a type of encoded information, and the encoding method used is the quantization-error-free position encoding method for two-dimensional sound source localization described above.

[0071] In one example, the output of the decoding network depends on its output layer. If the output layer of the decoding network is a fully connected layer, then its output... In reality, it's a stretched one-dimensional vector; at this point, the server needs to reshape the output of the decoding network into a... The matrix is ​​obtained from the above. .

[0072] Step 202, based on the above Obtain the predicted unbiased label distribution vector. And predict the unbiased label distribution vector According to the above The target sound source is in Peak category and its adjacent categories in the axial direction, the and the target sound source By analyzing the peak values ​​along the axis and their adjacent values, the coordinates of the target sound source are determined, thus completing the localization of the target sound source.

[0073] In the specific implementation, the server obtains Afterwards, it can be based on the above. , respectively along the acquisition Axial direction and Summing along the axes yields the predicted unbiased label distribution vector. And predict the unbiased label distribution vector The for The matrix, and It can be obtained using the following formula:

[0074] ;

[0075] In the formula, Indicates the The i-th row, Indicates the The i-th column.

[0076] In a specific implementation, the server can, according to the... The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the x-coordinate of the target sound source, which can be expressed by the formula:

[0077] ;

[0078] In the formula, For the target sound source in Peak category in the axial direction, For the target sound source in The category to the left of the peak category along the axis. For the target sound source in The category adjacent to the right of the peak category along the axis. Preset The segment length along the axial direction, which is the segment length in the above embodiment of the position coding method for two-dimensional sound source localization without quantization error. This represents the x-coordinate of the target sound source. This category is the same as the category in the above embodiment of the position encoding method for two-dimensional sound source localization without quantization error, and therefore has a fixed order.

[0079] In a specific implementation, the server can, according to the... The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the ordinate of the target sound source, which can be expressed by the formula:

[0080] ;

[0081] In the formula, For the target sound source in Peak category in the axial direction, For the target sound source in The category to the left of the peak category along the axis. For the target sound source in The category adjacent to the right of the peak category along the axis. Preset The segment length along the axial direction, which is the segment length in the above embodiment of the position coding method for two-dimensional sound source localization without quantization error. The vertical coordinate of the target sound source is obtained.

[0082] In this embodiment, a weighted adjacency decoding algorithm is used for decoding. Traditional decoding methods only consider peak categories, and quantization errors are still unavoidable. However, the weighted adjacency decoding algorithm considers multiple adjacent categories in addition to peak categories, and performs a more accurate weighted approximation of sound source localization. This overcomes quantization errors during the decoding process, thereby greatly reducing the error in sound source localization. It also has a good localization effect even under harsh conditions such as noise and reverberation.

[0083] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this patent.

[0084] In one embodiment, the server simulates and evaluates the quantization-error-free location encoding / decoding method for two-dimensional sound source localization proposed in this application. Three sets of simulated datasets are generated using the Pyroomacoustics library, with the localization difficulty of these datasets increasing progressively, named L-1, Ad-1, and Ad-2, respectively. All sound source speech is from the LibriSpeech corpus, and all sound source noise is from a large-scale noise set. The height of all rooms is fixed at 4.2 meters, the microphone height at 0.9 meters, and the sound source height at 0.95 meters.

[0085] The L-1 dataset is configured as a simple single-source dataset. All speech segments are in the same room free of echo or noise. Two 4-channel linear arrays are placed on the left and bottom walls of the room, respectively. The sound source is randomly placed at different locations within the room. Each speech segment has a duration of 160 milliseconds. The training, validation, and test sets of this dataset contain 100,000, 5,000, and 5,000 speech segments, respectively.

[0086] The Ad-1 and Ad-2 datasets have slightly different configurations, representing single-source and dual-source scenarios, respectively. For each speech in the dataset, a room with a random length and width within the range of [4, 10] m was generated. The reverberation time T60 was randomly selected between [0.2, 1.2] s. In the training set, the signal-to-noise ratio (SNR) was randomly selected between [0, 50] dB, while in the validation and test sets, the SNR was selected between [10, 20] dB. For each speech, there were 30 microphone nodes in its room, with each node containing only one microphone. The room was divided into 16×16 grids, some of which were randomly selected, and each grid was used to place either a microphone or a speaker. Each speech lasted 2 seconds. The training, validation, and test sets of this dataset contained 18,000, 1,800, and 1,800 speech samples, respectively. Implicit separation techniques were used to assist in multi-source localization, which required the use of the PIT (Permutation Invariant Training) algorithm.

[0087] The unbiased label distribution algorithm in the quantization-error-free location encoding and decoding method for two-dimensional sound source localization proposed in this application is denoted as ULD. It is compared with three different labeling methods: HM, RG, and One-hot. Notably, the baseline labels for HM, RG, and One-hot use I×I categories, while ULD uses (I + 1)×(I + 1) categories. In the experiments, the AdamW optimizer was used, with a maximum training epoch of 50. The batch size was set to 128 for training on the L-1 dataset, and 32 for other experiments. The initial learning rate was 0.001, which was reduced to 0.0001 if the validation loss did not decrease within 3 epochs. Training was prematurely terminated if the model's loss on the validation set did not decrease for 10 consecutive epochs.

[0088] Since all experiments are based on classification models, the most common and intuitive metric is ACC (Accuracy). For sound source localization, the most intuitive metrics are MAD (Mean Absolute Deviation), the straight-line distance between the predicted and true locations, and QE (Quantization Error), which represents the mean absolute error when using one-hot encoding. Experimental results on the L-1 dataset are shown in Table 1, on the Ad-1 dataset in Table 2, and on the Ad-2 dataset in Table 3.

[0089] Table 1: Experimental results on the L-1 dataset, with the network outputting 6×6 classes.

[0090]

[0091] Table 2: Experimental results on the Ad-1 dataset, where " / " indicates that model training cannot converge or is infeasible.

[0092]

[0093] Table 3: Experimental results on the Ad-2 dataset, where " / " indicates that model training cannot converge or is infeasible.

[0094]

[0095] As shown in Table 1, it is clear that in a clean environment, the classification model achieves a very high classification accuracy (ACC). In this case, one-hot encoding is mainly limited by quantization error, while RG and ULD significantly overcome this limitation. ULD performs exceptionally well, reducing MAE by 91.93% compared to one-hot encoding.

[0096] Table 2 shows that as the resolution I increases, the quantization error decreases, but the number of categories increases rapidly, leading to a rapid decline in classification accuracy. When I increases to 32, models trained with any label supervision fail to converge. This limits the performance of one-hot encoding, with a MAE of only 0.361 m. RG has There are several categories, while ULD has slightly more, with There are several categories. Therefore, in the lower... At this value, RG's ACC is slightly higher than ULD's. However, when RG is used as a label, a total of There are output neurons, when When the ACC increased from 8 to 16, it decreased significantly. For Both methods significantly overcome the limitations of quantization error. ULD achieves the best performance relative to One-hot by reducing MAE by 46.54%.

[0097] According to Table 3, when implicit separation is used, the performance of One-hot encoding is very close to that of the single-source case. However, due to the complexity of RG implementation and its further complexity when combined with PIT, its ACC is significantly lower than that of One-hot. On the other hand, ULD can be seen as a slightly modified, smoothed version of One-hot. It inherits the ease of training of One-hot, thus its performance in multi-source scenarios is similar to that in single-source scenarios. At different resolutions, it achieves a significantly smaller MAE compared to One-hot, maintaining a clear advantage in overcoming quantization errors; the MAE is reduced by 33.56% compared to One-hot.

[0098] This demonstrates that ULD possesses a strong ability to overcome quantization errors.

[0099] Another embodiment of this application relates to an electronic device, such as... Figure 4 As shown, it includes: at least one processor 301; and a memory 302 communicatively connected to the at least one processor 301; wherein the memory 302 stores instructions executable by the at least one processor 301, the instructions being executed by the at least one processor 301 to enable the at least one processor 301 to execute the quantization error-free position encoding and decoding method for two-dimensional sound source localization in the above embodiments.

[0100] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0101] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.

[0102] Another embodiment of this application relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.

[0103] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.

Claims

1. A position coding method for two-dimensional sound source localization without quantization error, characterized in that, include: Establish a Cartesian coordinate system in the space where the sound source is located, and divide the space into several grids; Based on the preset resolution, according to the space in Length along the axis and in The length along the axial direction, which divides the space in axial direction and The sound source is discretized into several segments along the axial direction, and the location of the sound source is determined based on its coordinates. Category in the axial direction and in Category along the axis; Based on the sound source Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source Position encoding in the axial direction, and based on the sound source in Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source Position encoding along the axis; Based on the above and stated Generate a two-dimensional unbiased label distribution matrix This completes the location encoding of the sound source; The preset resolution is used to determine the spatial values ​​in relation to the target area. Length along the axis and in The length along the axial direction, which divides the space in axial direction and Discretized along the axis into several segments, including: Based on the preset resolution I, and according to the space in Length in the axial direction and in Length in the axial direction Determine the space in Segment length in the axial direction and in Segment length in the axial direction ; Based on the above And the space in the I mentioned above. Discretized into several segments along the axial direction, the space in Several segments along the axial direction are represented as ; Based on the above And the space in the I mentioned above. Discretized into several segments along the axial direction, the space in Several segments along the axial direction are represented as ; The step of determining the location of the sound source based on its coordinates is as follows: Category in the axial direction and in The category along the axis is determined by the following formula: ; ; in, Let x be the x-coordinate of the sound source. Let be the ordinate of the sound source. This indicates that the sound source is in Categories in the axial direction, This indicates that the sound source is in Category along the axis; The Represented as The Represented as The sound source based on Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source The position encoding along the axis is achieved using the following formula: ; ; in, For the decimal function, For the floor function, For the The first in item; The sound source is based on Categories along the axes are represented by unbiased label distribution vectors. Perform on the sound source The position encoding along the axis is achieved using the following formula: ; ; in, For the The first in item; The basis of and stated Generate a two-dimensional unbiased label distribution matrix ,include: The Transpose column vectors and the Transpose row vectors ; The Multiplied by the above Generate a two-dimensional unbiased label distribution matrix The Represented as: 。 2. A quantization-error-free position decoding method for two-dimensional sound source localization, characterized in that, include: The signals from the target sound source received by each microphone are acquired, and these signals are input into a pre-trained decoding network to obtain the predicted two-dimensional unbiased label distribution matrix output by the decoding network. and the target sound source in Peak category in the axial direction and in Peak category in the axial direction; wherein, the The encoding method used is the same as the error-free position encoding method for two-dimensional sound source localization as described in claim 1; Based on the above Obtain the predicted unbiased label distribution vector. And predict the unbiased label distribution vector According to the above The target sound source is in Peak category and its adjacent categories in the axial direction, the and the target sound source in By analyzing the peak values ​​along the axis and their adjacent values, the coordinates of the target sound source are determined, thus completing the localization of the target sound source.

3. The quantization-error-free position decoding method for two-dimensional sound source localization according to claim 2, characterized in that, The for The matrix, based on the Obtain the predicted unbiased label distribution vector. And predict the unbiased label distribution vector This can be achieved through the following formula: ; in, Indicates the The i-th row, Indicates the The i-th column.

4. The quantization-error-free position decoding method for two-dimensional sound source localization according to claim 3, characterized in that, According to the The target sound source is in Peak category in the axial direction, the and the target sound source in The peak category along the axis is used to determine the coordinates of the target sound source, including: According to the above The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the abscissa of the target sound source; According to the above The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the ordinate of the target sound source.

5. The quantization-error-free position decoding method for two-dimensional sound source localization according to claim 4, characterized in that, According to the The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the x-coordinate of the target sound source, which is achieved through the following formula: ; in, For the target sound source in Peak category in the axial direction, For the target sound source in The category to the left of the peak category along the axis, For the target sound source in The category adjacent to the right of the peak category along the axis. For the preset Segment length in the axial direction, The x-coordinate of the target sound source is obtained from the solution. According to the The target sound source is in Peak categories along the axial direction and their adjacent categories, as well as preset categories. The segment length along the axis is used to solve for the ordinate of the target sound source, which is achieved through the following formula: ; in, For the target sound source in Peak category in the axial direction, For the target sound source in The category to the left of the peak category along the axis, For the target sound source in The category adjacent to the right of the peak category along the axis. For the preset Segment length in the axial direction, The vertical coordinate of the target sound source is obtained from the solution.

6. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform either the quantization-error-free position encoding method for two-dimensional sound source localization as described in claim 1, or the quantization-error-free position decoding method for two-dimensional sound source localization as described in any one of claims 2 to 5.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the position encoding method for two-dimensional sound source localization without quantization error as described in claim 1, or implements the position decoding method for two-dimensional sound source localization without quantization error as described in any one of claims 2 to 5.