In-vehicle speaker positioning method and device, storage medium and vehicle

By collecting and analyzing the acoustic signals in the car, combining the characteristics of ultrasonic signals and audio signals, and using positioning models to position the speakers in the car, solving the problem of low positioning accuracy in the existing technology and achieving higher positioning accuracy.

CN120214696APending Publication Date: 2025-06-27XIAOMI EV TECH CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510502464.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the accuracy of positioning of the speaker in the car is low, and it is easy to misidentify the non-human entity in the car as the speaker in the car. The ultrasonic wave can only identify the movable entity and cannot distinguish whether it is the speaker in the car.

Method used

By collecting the acoustic signals in the interior space of the vehicle, including the first ultrasonic signal and the first audio signal corresponding to the speaker, the characteristic of broadcasting by the in-car speaker is used, and positioning is performed in combination with the voice characteristics unique to the speaker in the first audio signal. The specific steps include obtaining the Mel spectrum of the audio signal and the channel impulse response spectrum of the ultrasonic signal, and inputting it to the trained positioning model to determine the speaker's positioning result.

Benefits of technology

It improves the accuracy of positioning of speakers in the car, can effectively distinguish whether the movable entity is a person in the car, and reduces the occurrence of misidentification of non-human entities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120214696A_ABST
    Figure CN120214696A_ABST
Patent Text Reader

Abstract

The invention provides a method for positioning a speaker in a vehicle, and the method comprises the steps: collecting an acoustic signal containing a first ultrasonic signal and a first audio signal corresponding to the speaker in a space in the vehicle, employing the characteristic that the first ultrasonic signal is broadcasted by a loudspeaker in the vehicle, combining the specific voice features of the speaker in the first audio signal, and carrying out the positioning of the speaker in the vehicle. In the positioning process, initial positioning is carried out by means of the spatial propagation characteristic of the ultrasonic signal, and whether the movable entity is the person in the vehicle or not is effectively distinguished through the first audio signal, so that the positioning accuracy of the speaker in the vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of vehicles, and in particular, to a method and apparatus for locating an in-vehicle speaker, a storage medium, and a vehicle. Background Art

[0002] With the continuous development of vehicle intelligence, people have higher and higher requirements for the intelligent cockpit of vehicles. By locating the in-vehicle speaker among the in-vehicle personnel, better intelligent services can be provided for the in-vehicle personnel.

[0003] In the related art of in-vehicle speaker location, usually only a single ultrasonic wave is used for in-vehicle speaker location. Since ultrasonic waves can only identify movable entities and cannot distinguish whether the movable entity is an in-vehicle speaker, it is easy to misidentify a non-human entity in the vehicle as an in-vehicle speaker, thereby resulting in a low accuracy of in-vehicle speaker location. Summary of the Invention

[0004] The present disclosure provides one to solve the problems in the related art.

[0005] A first aspect embodiment of the present disclosure proposes a method for locating an in-vehicle speaker, the method including:

[0006] Collecting acoustic signals in the in-vehicle space; the acoustic signals include a first ultrasonic signal and a first audio signal corresponding to the speaker; the ultrasonic signal is broadcast by a speaker in the vehicle when the vehicle performs the location of the in-vehicle speaker;

[0007] Determining a location result of the speaker according to the acoustic signals.

[0008] In some embodiments, the determining the location result of the speaker according to the acoustic signals includes:

[0009] Obtaining a Mel spectrum corresponding to the first audio signal and a channel impulse response spectrum corresponding to the first ultrasonic signal;

[0010] Inputting the Mel spectrum corresponding to the first audio signal and the channel impulse response spectrum corresponding to the first ultrasonic signal into a trained location model to obtain a location result output by the trained location model.

[0011] In some embodiments, the method further includes:

[0012] Obtaining an acoustic signal sample, the acoustic signal sample including a second ultrasonic signal and a second audio signal;

[0013] Training an initial location model based on the acoustic signal sample to obtain the trained location model;

[0014] Among them, the acoustic signal samples satisfy at least one of the following:

[0015] The speakers corresponding to the second audio signals in different acoustic signal samples are not exactly the same;

[0016] The positions of the speakers corresponding to the second audio signals in different acoustic signal samples in the vehicle are not exactly the same;

[0017] The vehicles to which the speakers corresponding to the second audio signals in different acoustic signal samples belong are not exactly the same;

[0018] The scenes in which the vehicles to which the speakers corresponding to the second audio signals in different acoustic signal samples are located are not exactly the same;

[0019] The scenes in which the speakers corresponding to the second audio signals in different acoustic signal samples are located are not exactly the same;

[0020] The second audio signals in at least some of the acoustic signal samples correspond to scenes without speakers;

[0021] The second audio signals in at least some of the acoustic signal samples correspond to at least one speaker.

[0022] In some embodiments, training the initial positioning model based on the acoustic signal samples to obtain the trained positioning model includes:

[0023] Based on the acoustic signal samples, obtaining enhanced acoustic signal samples, where the enhanced acoustic signal samples include a third ultrasonic signal and a third audio signal;

[0024] Based on the acoustic signal samples and / or the enhanced acoustic signal samples, training the initial positioning model to obtain the trained positioning model;

[0025] Among them, the enhanced acoustic signal samples satisfy at least one of the following:

[0026] The third audio signals in at least some of the enhanced acoustic signal samples are obtained by adding noise signals to the second audio signals, where different third audio signals correspond to not exactly the same second audio signals, and / or, different third audio signals correspond to not exactly the same noise signals;

[0027] The third ultrasonic signals in at least some of the enhanced acoustic signal samples are obtained by rotating the second ultrasonic signals in the complex plane, where different third ultrasonic signals correspond to not exactly the same second ultrasonic signals, and / or, different third ultrasonic signals correspond to not exactly the same rotation angles;

[0028] At least a partially enhanced acoustic signal sample is obtained by combining a third ultrasonic signal with a first weight and a third audio signal with a second weight, where different third ultrasonic signals correspond to not exactly the same first weights, and / or different third audio signals correspond to not exactly the same second weights.

[0029] In some embodiments, training an initial positioning model based on the acoustic signal sample includes:

[0030] Obtaining a Mel spectrum corresponding to the second audio signal and a channel impulse response spectrum corresponding to the second ultrasonic signal;

[0031] Inputting the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal into the initial positioning model for training;

[0032] Training an initial positioning model based on the enhanced acoustic signal sample includes:

[0033] Obtaining a Mel spectrum corresponding to the third audio signal and a channel impulse response spectrum corresponding to the third ultrasonic signal;

[0034] Inputting the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal into the initial positioning model for training.

[0035] In some embodiments, the inputting the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal into the initial positioning model for training includes:

[0036] Extracting a first Mel spectrum feature from the Mel spectrum corresponding to the second audio signal and extracting a first channel impulse response spectrum feature from the channel impulse response spectrum corresponding to the second ultrasonic signal;

[0037] Concatenating the first Mel spectrum feature and the first channel impulse response spectrum feature according to a preset feature dimension to obtain a first concatenated feature;

[0038] Inputting the first concatenated feature into a core neural network structure of the initial positioning model to obtain a multi-label output of a network output layer of the positioning model; adjusting parameters of the positioning model based on a loss corresponding to the multi-label output;

[0039] Iteratively executing at least one first training process, where the first training process includes: inputting the first concatenated feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain a multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on a loss corresponding to the multi-label output;

[0040] Inputting the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal into the initial positioning model for training includes:

[0041] Extracting the second Mel spectrum feature from the Mel spectrum corresponding to the third audio signal, and extracting the second channel impulse response spectrum feature from the channel impulse response spectrum corresponding to the third ultrasonic signal;

[0042] Concatenating the second Mel spectrum feature and the second channel impulse response spectrum feature according to the preset feature dimension to obtain a second concatenated feature;

[0043] Inputting the second concatenated feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output;

[0044] Iteratively execute the second training process at least once. The second training process includes: inputting the second concatenated feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0045] In some embodiments, the obtaining of the acoustic signal sample includes:

[0046] Obtaining an acoustic signal to be processed; the acoustic signal to be processed includes an ultrasonic signal to be processed and an audio signal to be processed;

[0047] Preprocessing the acoustic signal to be processed to obtain the acoustic signal sample;

[0048] Wherein, the preprocessing includes at least one of the following:

[0049] Noise reduction processing;

[0050] Filtering processing.

[0051] The second aspect embodiment of the present disclosure proposes a positioning device for an in-vehicle speaker. The device includes:

[0052] An acquisition unit for acquiring acoustic signals in the in-vehicle space; the acoustic signals include a first ultrasonic signal and a first audio signal corresponding to the speaker; the ultrasonic signal is broadcast by a speaker in the vehicle when the in-vehicle speaker positioning is performed;

[0053] A determination unit for determining the positioning result of the speaker according to the acoustic signals.

[0054] In some embodiments, the determination unit includes:

[0055] An acquisition module, configured to acquire the Mel spectrum corresponding to the first audio signal and the channel impulse response spectrum corresponding to the first ultrasonic signal;

[0056] An input module, configured to input the Mel spectrum corresponding to the first audio signal and the channel impulse response spectrum corresponding to the first ultrasonic signal into the trained positioning model, and obtain a positioning result output by the trained positioning model.

[0057] In some embodiments, the apparatus further includes:

[0058] An acquisition unit, configured to acquire an acoustic signal sample, where the acoustic signal sample includes a second ultrasonic signal and a second audio signal;

[0059] A training unit, configured to train an initial positioning model based on the acoustic signal sample to obtain the trained positioning model;

[0060] Wherein, the acoustic signal sample satisfies at least one of the following:

[0061] The speakers corresponding to the second audio signals in different acoustic signal samples are not completely the same;

[0062] The positions of the speakers corresponding to the second audio signals in different acoustic signal samples in the vehicle are not completely the same;

[0063] The vehicles to which the speakers corresponding to the second audio signals in different acoustic signal samples belong are not completely the same;

[0064] The scenes where the vehicles to which the speakers corresponding to the second audio signals in different acoustic signal samples belong are not completely the same;

[0065] The scenes where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same;

[0066] At least part of the second audio signals in the acoustic signal samples correspond to scenes without speakers;

[0067] At least part of the second audio signals in the acoustic signal samples correspond to at least one speaker.

[0068] In some embodiments, the training unit is further configured to:

[0069] Based on the acoustic signal sample, obtain an enhanced acoustic signal sample, where the enhanced acoustic signal sample includes a third ultrasonic signal and a third audio signal;

[0070] Based on the acoustic signal sample and / or the enhanced acoustic signal sample, train the initial positioning model to obtain the trained positioning model;

[0071] Among them, the enhanced acoustic signal samples satisfy at least one of the following:

[0072] The third audio signal in at least part of the enhanced acoustic signal samples is obtained by adding a noise signal to the second audio signal, where different third audio signals correspond to not completely identical second audio signals, and / or different third audio signals correspond to not completely identical noise signals;

[0073] The third ultrasonic signal in at least part of the enhanced acoustic signal samples is obtained by rotating the second ultrasonic signal in the complex plane, where different third ultrasonic signals correspond to not completely identical second ultrasonic signals, and / or different third ultrasonic signals correspond to not completely identical rotation angles;

[0074] At least part of the enhanced acoustic signal samples are obtained by combining the third ultrasonic signal with a first weight and the third audio signal with a second weight, where different third ultrasonic signals correspond to not completely identical first weights, and / or different third audio signals correspond to not completely identical second weights.

[0075] In some embodiments, the training unit is further configured to:

[0076] Obtain the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal;

[0077] Input the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal into the initial positioning model for training;

[0078] The training unit is further configured to:

[0079] Obtain the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal;

[0080] Input the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal into the initial positioning model for training.

[0081] In some embodiments, the training unit is further configured to:

[0082] Extract the first Mel spectrum feature from the Mel spectrum corresponding to the second audio signal, and extract the first channel impulse response spectrum feature from the channel impulse response spectrum corresponding to the second ultrasonic signal;

[0083] Stitch the first Mel spectrum feature and the first channel impulse response spectrum feature according to a preset feature dimension to obtain a first stitched feature;

[0084] Input the first splicing feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output;

[0085] Iteratively execute the first training process at least once. The first training process includes: inputting the first splicing feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output;

[0086] The training unit is further configured to:

[0087] Extract the second Mel spectrum feature in the Mel spectrum corresponding to the third audio signal, and extract the second channel impulse response spectrum feature in the channel impulse response spectrum corresponding to the third ultrasonic signal;

[0088] Splice the second Mel spectrum feature and the second channel impulse response spectrum feature according to the preset feature dimension to obtain a second splicing feature;

[0089] Input the second splicing feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output;

[0090] Iteratively execute the second training process at least once. The second training process includes: inputting the second splicing feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0091] In some embodiments, the obtaining unit is further configured to:

[0092] Obtain an acoustic signal to be processed; the acoustic signal to be processed includes an ultrasonic signal to be processed and an audio signal to be processed;

[0093] Preprocess the acoustic signal to be processed to obtain the acoustic signal sample;

[0094] Wherein, the preprocessing includes at least one of the following:

[0095] Noise reduction processing;

[0096] Filtering processing;

[0097] Filtering process.

[0098] A third aspect embodiment of the present disclosure provides a vehicle, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the method described in the first aspect embodiment of the present disclosure.

[0099] A fourth aspect embodiment of the present disclosure provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect embodiment of the present disclosure.

[0100] A fifth aspect embodiment of the present disclosure provides a non-transitory computer-readable storage medium, which, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to execute the method described in the first aspect embodiment of the present disclosure.

[0101] In summary, according to the in-vehicle speaker localization method proposed by the present disclosure, the method includes collecting acoustic signals in the in-vehicle space that include first ultrasonic signals and first audio signals corresponding to the speaker, using the characteristic that the first ultrasonic signals are broadcast by in-vehicle speakers, and combining the unique voice characteristics of the speaker in the first audio signals. During the localization process, not only the spatial propagation characteristics of the ultrasonic signals are used for initial localization, but also the first audio signals are used to effectively distinguish whether the movable entity is an in-vehicle person, improving the accuracy of in-vehicle speaker localization.

[0102] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used in conjunction with the specification to explain the principles of the present disclosure and do not constitute an undue limitation to the present disclosure.

[0104] Figure 1 It is a flowchart of a method for localizing an in-vehicle speaker provided by an embodiment of the present disclosure;

[0105] Figure 2 It is a flowchart of another method for localizing an in-vehicle speaker provided by an embodiment of the present disclosure;

[0106] Figure 3 It is a flowchart of a method for training a localization model provided by an embodiment of the present disclosure;

[0107] Figure 4 It is a flowchart of another method for training a localization model provided by an embodiment of the present disclosure;

[0108] Figure 5 Flowchart of another method for training a positioning model provided by an embodiment of the present disclosure;

[0109] Figure 6 Flowchart of another method for training a positioning model provided by an embodiment of the present disclosure;

[0110] Figure 7 Schematic flowchart of another method for training a positioning model provided by an embodiment of the present disclosure;

[0111] Figure 8 Flowchart of another method for training a positioning model provided by an embodiment of the present disclosure;

[0112] Figure 9 Flowchart of another method for training a positioning model provided by an embodiment of the present disclosure;

[0113] Figure 10 Flowchart of a method for obtaining an acoustic signal sample provided by an embodiment of the present disclosure;

[0114] Figure 11 Flowchart of the entire process of positioning an in-vehicle speaker provided by an embodiment of the present disclosure;

[0115] Figure 12 Schematic structural diagram of a device for positioning an in-vehicle speaker provided by an embodiment of the present disclosure;

[0116] Figure 13 Schematic structural diagram of another device for positioning an in-vehicle speaker provided by an embodiment of the present disclosure;

[0117] Figure 14 Block diagram of a vehicle provided by an embodiment of the present disclosure. Detailed implementation manners

[0118] Some embodiments of the present disclosure will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, modifications, and equivalents of the methods, devices, and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of operations described herein is merely an example and is not limited to the orders set forth herein. Rather, except for operations that must be performed in a specific order, changes can be made as will be apparent after understanding the present disclosure. Additionally, for the sake of clarity and conciseness, descriptions of features known in the art may be omitted.

[0119] The embodiments described in some embodiments of the present disclosure do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0120] With the continuous development of vehicle intelligence, people have higher and higher requirements for the intelligent cockpit of vehicles. By positioning the in-vehicle speaker of the people in the vehicle, better intelligent services can be provided for the people in the vehicle.

[0121] In the related art of in-vehicle speaker positioning, usually only a single ultrasonic wave is used for in-vehicle speaker positioning. Since ultrasonic waves can only identify movable entities and cannot distinguish whether the movable entity is an in-vehicle speaker, it is easy to misidentify non-human entities in the vehicle as in-vehicle speakers, resulting in a low accuracy rate of in-vehicle speaker positioning.

[0122] Therefore, in order to solve the problems existing in the related art, the present disclosure proposes a method for positioning an in-vehicle speaker. By collecting acoustic signals in the in-vehicle space that include a first ultrasonic signal and a first audio signal corresponding to the speaker, using the characteristic that the first ultrasonic signal is broadcast by the in-vehicle speaker, and combining the unique voice characteristics of the speaker in the first audio signal, not only the initial positioning is carried out by means of the spatial propagation characteristics of the ultrasonic signal during the positioning process, but also the first audio signal is used to effectively distinguish whether the movable entity is a person in the vehicle, improving the accuracy rate of in-vehicle speaker positioning.

[0123] The embodiments of the present disclosure are not exhaustive, but only schematic illustrations of some embodiments, and do not constitute a specific limitation on the protection scope of the present disclosure. Without contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily. For example, the solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be exchanged arbitrarily. In addition, the optional implementation manners in a certain embodiment can be combined arbitrarily; furthermore, the embodiments can be combined arbitrarily. For example, some or all of the steps of different embodiments can be combined arbitrarily, and a certain embodiment can be combined arbitrarily with the optional implementation manners of other embodiments.

[0124] In each embodiment of the present disclosure, if there is no special explanation and logical conflict, the terms and / or descriptions among the embodiments are consistent and can be cited from each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.

[0125] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended as a limitation on the present disclosure.

[0126] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular form, such as "a", "an", "the", "above-mentioned", "said", "aforementioned", "this", etc., may mean "one and only one", or may also mean "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English translation, the noun after the article can be understood as a singular expression or a plural expression.

[0127] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "when...", "while...", "if...", "if...", etc. may be used interchangeably.

[0128] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not lower than", "above", etc. may be used interchangeably, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", "below", etc. may be used interchangeably.

[0129] Prefix words such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different described objects, and do not limit the position, order, priority, quantity, content, etc. of the described objects. The description of the described objects refers to the description in the context of the claims or embodiments, and should not constitute an unnecessary limitation due to the use of prefix words.

[0130] In the embodiments of the present disclosure, "a plurality of" means two or more.

[0131] In the embodiments of the present disclosure, terms such as "import", "input", "read in", etc. may be used interchangeably.

[0132] In some embodiments, a device, etc. can be interpreted as physical or virtual, and its name is not limited to the name recorded in the embodiments. Terms such as "device", "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", "subject", etc. may be used interchangeably.

[0133] Figure 1The flowchart of a method for locating an in-vehicle speaker provided by an embodiment of the present disclosure. This method can be applied to terminals with in-vehicle location functions, such as electric vehicles, hybrid vehicles, electric two-wheelers, and other terminals with in-vehicle location functions, which are not limited in the present disclosure. As Figure 1 shown, the method for locating an in-vehicle speaker includes steps 101-102.

[0134] Step 101, collect acoustic signals in the in-vehicle space; the acoustic signals include a first ultrasonic signal and a first audio signal corresponding to the speaker; the ultrasonic signal is broadcast by a speaker in the vehicle when the vehicle performs the location of the in-vehicle speaker.

[0135] The in-vehicle space refers to the in-vehicle cockpit environment inside the vehicle, including but not limited to the driver's cab and the passenger cabin area. Acoustic signals refer to sound wave signals propagating in the in-vehicle space, such as human voices, vehicle running noises, and broadcast sounds. The first ultrasonic signal is an ultrasonic signal broadcast by a speaker in the vehicle, which is used to assist in locating the in-vehicle speaker. The first ultrasonic signal usually has specific frequencies and waveforms, and can provide spatial location information through acoustic characteristics (such as propagation time, reflection characteristics, etc.). The first audio signal is an audible sound signal emitted by the in-vehicle speaker, which contains the speech characteristics of the speaker (such as frequency, pitch, loudness, etc.), and is used to identify the presence and location of the speaker. Among them, the first audio signal can be an audio signal with a frequency less than a first preset frequency, and the first audio signal can be a low-frequency audio signal. The low-frequency audio signal is mainly generated by the actions of in-vehicle personnel (such as speaking, limb movement, breathing) or object collisions (such as closing the car door). The first ultrasonic signal can be an audio signal with a frequency greater than a second preset frequency broadcast in the in-vehicle space, and the first ultrasonic signal can be a high-frequency sound signal. For example, it is an ultrasonic signal of a self-correlation baseband signal modulated by a Zadoff-Chu (ZC) sequence. Among them, the first preset frequency can be 8 kHz, and the second preset frequency can be 18 kHz. When the first preset frequency is 8 kHz, the first audio signal is an audio signal within the audible range of the human ear. When the second preset frequency is 18 kHz, the first ultrasonic signal is an ultrasonic signal inaudible to the human ear. However, it should be clear that this statement is not intended to limit the first preset frequency to only 8 kHz and the second preset frequency to only 18 kHz. The first preset frequency and the second preset frequency can also take other values, and the specific values of the first preset frequency and the second preset frequency are not limited in the embodiments of the present disclosure.

[0136] The acoustic signals in the vehicle interior space can be collected through the built-in microphone of the vehicle, or through a device with a sound collection function connected to the vehicle's in-vehicle computer. The first ultrasonic signal can be played in the vehicle interior space through the built-in speaker of the vehicle, such as the overhead speaker of the vehicle, or through a device with a sound playback function connected to the vehicle's in-vehicle computer. By performing frequency-domain analysis on the collected acoustic signals in the vehicle interior space, the first audio signal with a frequency lower than the first preset frequency in the audio signal of the vehicle interior space is extracted, or after downsampling the audio signal of the vehicle interior space, the first audio signal is extracted; after the first audio signal is extracted, white noise may exist in the first audio signal, and the white noise will interfere with the first audio signal. It is necessary to perform noise reduction processing on the first audio signal to remove the white noise in the first audio signal. The first audio signal can be denoised through a filter to reduce the influence of the white noise on the first audio signal. By performing frequency-domain analysis on the collected acoustic signals in the vehicle interior space, the frequency band where the specific baseband signal is located is obtained, demodulation is performed, the channel state information is obtained according to the baseband signal, and the obtained channel state information is filtered using the path length and signal strength to exclude the interference of non-target paths. Further, filtering processing and background elimination processing are performed on the channel state information to obtain the first ultrasonic signal.

[0137] The first preset frequency and the second preset frequency are preset frequencies, which can be adjusted according to actual needs and can flexibly adapt to the positioning requirements in different vehicle models and different noise environments.

[0138] The first ultrasonic signal and the first audio signal corresponding to the speaker are collected. The first audio signal can clearly identify the voice characteristics of the speaker. Combining the spatial position information carried by the first ultrasonic signal, the position of the speaker can be determined more accurately, improving the accuracy of speaker localization in the vehicle interior.

[0139] Step 102: Determine the localization result of the speaker according to the acoustic signal.

[0140] The localization result refers to determining the specific position information of the speaker in the vehicle interior space by analyzing the acoustic signal. The localization result is usually represented in the form of coordinates or regional division. Among them, the regional division can be made according to the seats in the vehicle. The number of seats in the vehicle determines how many regions the cockpit in the vehicle is divided into.

[0141] Multiple microphone arrays on the vehicle are used to collect acoustic signals. The microphone arrays can be evenly distributed at different positions inside the vehicle (such as the roof, doors, instrument panel, etc.) to cover the entire interior space of the vehicle. The vehicle's built-in speaker broadcasts ultrasonic signals, and the microphone arrays collect acoustic signals in real time, including the first ultrasonic signal and the first audio signal emitted by the speaker. The collected acoustic signals are denoised to remove background noise (such as wind noise, engine noise, etc.) to improve the signal-to-noise ratio. The Mel spectrum of the first audio signal and the channel impulse response spectrum of the first ultrasonic signal are obtained, and the Mel spectrum of the first audio signal and the channel impulse response spectrum of the first ultrasonic signal are input into the trained localization model, and the localization result of the speaker is output through the trained localization model.

[0142] In addition to determining the localization result of the speaker through the localization model, the position of the speaker can also be determined by calculating the time delay difference between the acoustic signals received by different microphones and combining geometric calculations. The microphone arrays can also be used to perform weighted summation on the acoustic wave signals in different directions, and the direction with the maximum beam output power is taken as the sound source direction. Combining the characteristics of ultrasonic signals and audio signals can more accurately locate the speaker inside the vehicle and avoid misidentifying non-human entities. The localization of the speaker inside the vehicle can be achieved only by collecting audio signals through the vehicle's built-in microphones, without the need for additional sensors for the localization of the speaker inside the vehicle, reducing the hardware cost.

[0143] According to the method for localizing a speaker inside a vehicle proposed by the present disclosure, the method includes collecting acoustic signals in the interior space of the vehicle that include a first ultrasonic signal and a first audio signal corresponding to the speaker, and using the characteristic that the first ultrasonic signal is broadcast by the in-vehicle speaker, and combining the unique voice characteristics of the speaker in the first audio signal. During the localization process, not only the spatial propagation characteristics of the ultrasonic signal are used for initial localization, but also the first audio signal is used to effectively distinguish whether the moving entity is a vehicle occupant, improving the accuracy of localizing the speaker inside the vehicle.

[0144] For a better understanding of determining the localization result of the speaker according to the acoustic signals, Figure 2 A flowchart of a method for localizing a speaker inside a vehicle proposed by the present disclosure is further shown. Based on Figure 1 the embodiment shown, step 102 is further explained, Figure 2 It may include the following steps (that is, determining the localization result of the speaker according to the acoustic signals may include the following steps):

[0145] Step 201, obtain the Mel spectrum corresponding to the first audio signal and the channel impulse response spectrum corresponding to the first ultrasonic signal.

[0146] The Mel-spectrum is a spectral representation based on the Mel frequency scale. The Mel frequency scale is closer to the frequency scale of human auditory perception. The Mel-spectrum converts linear frequency to Mel frequency, then performs a Fourier transform on the signal, and finally obtains the energy distribution on the Mel frequency scale through a Mel filter bank. The Mel-spectrum can highlight the features related to human auditory perception in the audio signal, reduce the influence of high-frequency noise, and more effectively represent the spectral characteristics of the audio signal.

[0147] The Channel Impulse Response (CIR) spectrum is a spectral representation that describes the response characteristics of the channel to the ultrasonic signal when the ultrasonic signal propagates in the channel. In the in-vehicle environment, the first ultrasonic signal will experience multiple reflections and scatterings during propagation, forming multipath propagation. The channel impulse response spectrum records the multipath propagation effect, reflects the propagation path and attenuation of the first ultrasonic signal from the transmitter to the receiver, and contains the spatial information related to the speaker's position.

[0148] The first audio signal can be processed by a preset Mel-spectrum transformation algorithm to obtain the Mel-spectrum corresponding to the first audio signal. The preset Mel-spectrum transformation algorithm includes performing a short-time Fourier transform on the first audio signal to convert the first audio signal from the time domain to the frequency domain, obtaining a frequency-domain signal, filtering the frequency-domain signal through a Mel filter bank, and converting the filtered spectrogram to a Mel-spectrum to obtain the frequency characteristics of the signal on the Mel scale. The frequency characteristics contain the audio content of the signal, such as the intonation and timbre of speech.

[0149] The propagation path of the ultrasonic signal can be determined by analyzing the time-delay difference between the first ultrasonic signal from the speaker to the microphone. By comparing the transmitted first ultrasonic signal with the received first ultrasonic signal, the frequency response of the in-vehicle space to the first ultrasonic signal is calculated, including characteristics such as the time delay, attenuation, and frequency shift of the first ultrasonic signal. According to the characteristics of the frequency response, the channel impulse response of the first ultrasonic signal is obtained through methods such as inverse Fourier transform.

[0150] The Mel-spectrum can reflect the speech characteristics of the first audio signal, while the channel impulse response spectrum contains the spatial information of the first ultrasonic signal propagating in the vehicle. Combining the Mel-spectrum with the channel impulse response spectrum can provide richer and more accurate feature information for in-vehicle speaker localization, thereby improving the accuracy of in-vehicle speaker localization.

[0151] Step 202: Input the Mel-spectrum corresponding to the first audio signal and the channel impulse response spectrum corresponding to the first ultrasonic signal into the trained localization model to obtain the localization result output by the trained localization model.

[0152] The positioning model is a deep learning-based model used to predict the position of a speaker in a vehicle based on the input Mel spectrogram and channel impulse response spectrogram. The positioning model can be constructed by including but not limited to convolutional neural networks and recurrent neural networks. After the Mel spectrogram corresponding to the first audio signal and the channel impulse response spectrogram corresponding to the first ultrasonic signal are input into the trained positioning model, the positioning model will perform feature extraction on the Mel spectrogram corresponding to the first audio signal and the channel impulse response spectrogram corresponding to the first ultrasonic signal. The feature extraction includes independently normalizing each channel of the Mel spectrogram corresponding to the first audio signal, then dividing it into blocks, and performing feature extraction on each block using the feature extraction layer of the positioning model. After splicing the features of all channels, the features of the Mel spectrogram are formed. Among them, the feature extraction layer can be a neural network. For the channel impulse response spectrogram, by independently normalizing each channel of the channel impulse response spectrogram, then dividing it into blocks, performing feature extraction on each block using the feature extraction layer of the positioning model, and splicing the features of all channels, the features of the channel impulse response spectrogram are formed. After splicing the above two features, they are input into the core neural network structure of the positioning model. The network output layer of the core neural network structure is set to multi-label, and each label represents whether there is someone in a certain position in the vehicle. The core neural network outputs the positioning result of the speaker.

[0153] The features of the Mel spectrogram corresponding to the first audio signal and the features of the channel impulse response spectrogram corresponding to the first ultrasonic signal can be spliced according to a preset feature dimension to obtain the first spliced feature.

[0154] The trained positioning model is obtained by training with a large number of acoustic signal samples in different scenarios. In practical applications, when the Mel spectrogram and the channel impulse response spectrogram are input into the positioning model, the positioning model can accurately position different in-vehicle environments and speaker situations according to the learned mapping relationship, and has strong generalization ability.

[0155] In practical applications, in order to improve the accuracy of the positioning process of the positioning model, it is necessary to pre-train the positioning model. As Figure 3 shown, the embodiments of the present disclosure also provide a flowchart of a training method for a positioning model, including:

[0156] Step 301, obtain acoustic signal samples, where the acoustic signal samples include a second ultrasonic signal and a second audio signal.

[0157] The acoustic signal sample refers to the acoustic signal collected in the in-vehicle space of different test vehicles for training the position positioning model; among them, the test vehicle refers to the specific vehicle used to collect training data, and the test vehicle should be representative and cover different vehicle models, different in-vehicle spaces and other conditions.

[0158] The second ultrasonic signal refers to the ultrasonic signal that is emitted by the in-vehicle speaker of the test vehicle during the acquisition of the acoustic signal samples and is captured by the in-vehicle microphone of the test vehicle after being reflected and scattered through the in-vehicle environment. The second audio signal is the audio signal emitted by the in-vehicle speaker during the acquisition of the acoustic signal samples and is collected by the in-vehicle microphone of the test vehicle.

[0159] By obtaining acoustic signal samples covering various situations such as different vehicles, different scenarios, and different speaker positions, the trained positioning model can adapt to various complex in-vehicle environments. In practical applications, regardless of how the vehicle type and scenario change, the positioning model can perform accurate positioning and analysis based on the learned features, improving the generalization ability of the positioning model.

[0160] Step 302, the acoustic signal samples satisfy at least one of the following: the speakers corresponding to the second audio signals in different acoustic signal samples are not completely the same; the positions of the speakers corresponding to the second audio signals in different acoustic signal samples in the vehicle are not completely the same; the vehicles where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same; the scenarios where the vehicles where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same; the scenarios where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same; at least some of the acoustic signal samples have the second audio signal corresponding to a scenario without speakers; at least some of the acoustic signal samples have the second audio signal corresponding to at least one speaker.

[0161] The scenario refers to different states of the in-vehicle environment, such as a quiet driving environment, a multi-person conversation environment, a music-playing environment, etc. Different speakers include speakers of different genders, ages, and accents. Different positions refer to different positions in the vehicle, including but not limited to the driver's seat, the front passenger seat, and the rear seats. Different vehicles refer to vehicles of different models and brands.

[0162] The speaker corresponding to the second audio signal refers to the sound-producing entity that generates the second audio signal, including personnel such as the driver and passengers in the vehicle. The speakers not being completely the same means that the audio signals in different acoustic signal samples come from at least two different speakers. For example, assume there are 3 acoustic signal samples, namely acoustic signal sample 1, acoustic signal sample 2, and acoustic signal sample 3. Among them, the speaker of acoustic signal sample 1 is male driver a, the speaker of acoustic signal sample 2 is female passenger b, and the speaker of acoustic signal sample 3 is child passenger c.

[0163] The position of the speaker corresponding to the second audio signal in the vehicle refers to the physical coordinates or seat area in the vehicle of the person generating the second audio signal. For example, area labels such as the driver's seat, the front passenger seat, the left rear passenger seat, the right rear passenger seat, the middle seat, etc., or the precise position based on a coordinate system. The position reflects the specific orientation of the speaker in the vehicle space. That the positions are not exactly the same means that the speakers in different acoustic signal samples are from at least two different positions in the vehicle. For example, suppose there are 3 acoustic signal samples, namely acoustic signal sample 1, acoustic signal sample 2, and acoustic signal sample 3. Among them, the position of the speaker in acoustic signal sample 1 in the vehicle is the driver's seat, the position of the speaker in acoustic signal sample 2 in the vehicle is the front passenger seat, and the position of the speaker in acoustic signal sample 3 in the vehicle is the left rear passenger seat.

[0164] That the vehicles where the speakers are located are not exactly the same means that the vehicles where the speakers corresponding to the second audio signal in different acoustic signal samples are from at least two different vehicles. For example, suppose there are 3 acoustic signal samples, namely acoustic signal sample 1, acoustic signal sample 2, and acoustic signal sample 3. Among them, the vehicle where the speaker in acoustic signal sample 1 is located is vehicle a, the vehicle where the speaker in acoustic signal sample 2 is located is vehicle b, and the vehicle where the speaker in acoustic signal sample 3 is located is vehicle c.

[0165] The scene where the vehicle is located refers to the real-time state and external environment of the vehicle generating the second audio signal during data collection, including but not limited to the driving state, environmental noise, and environmental conditions. That the scenes where the vehicles where the speakers are located are not exactly the same means that the scenes where the vehicles where the speakers corresponding to the second audio signal in different acoustic signal samples are from at least two different scenes. For example, suppose there are 3 acoustic signal samples, namely acoustic signal sample 1, acoustic signal sample 2, and acoustic signal sample 3. Among them, the scene where the vehicle where the speaker in acoustic signal sample 1 is located is scene a, the scene where the vehicle where the speaker in acoustic signal sample 2 is located is scene b, and the scene where the vehicle where the speaker in acoustic signal sample 3 is located is scene c.

[0166] The scene where the speaker is located refers to the specific environmental state when the person generating the second audio signal is speaking, including but not limited to the spatial environment, environmental noise, and reverberation characteristics. That the scenes where the speakers are located are not exactly the same means that the scenes where the speakers corresponding to the second audio signal in different acoustic signal samples are from at least two different scenes. For example, suppose there are 3 acoustic signal samples, namely acoustic signal sample 1, acoustic signal sample 2, and acoustic signal sample 3. Among them, the scene where the speaker in acoustic signal sample 1 is located is scene d, the scene where the speaker in acoustic signal sample 2 is located is scene e, and the scene where the speaker in acoustic signal sample 3 is located is scene f.

[0167] A scene without a speaker refers to a state where no one in the vehicle actively makes a sound during the acquisition of acoustic signals. At this time, the second audio signal only contains environmental noise (such as the running sound of the air conditioner, tire noise during vehicle driving, road bump vibration noise), equipment self-noise (such as the background noise of ultrasonic emission, microphone thermal noise), or other non-human voice signals (such as in-vehicle system prompt sounds, friction sounds generated by passenger movements), but does not contain the effective vocalization of human speech (such as conversations, instructions, humming, etc.). An acoustic signal sample without a speaker refers to a sample composed of the second ultrasonic signal (the echo of ultrasonic waves reflected by in-vehicle entities) and the second audio signal (environmental signal without speech) collected in a scene without a speaker.

[0168] Different acoustic signal samples can correspond to different speakers. The voice characteristics of each speaker are different. Therefore, by analyzing the second audio signals generated by different speakers, more samples can be provided for the localization of speakers in the vehicle. Different positions of the in-vehicle microphone array will result in different captured second audio signals. Therefore, the second audio signals generated by the same speaker at different positions (such as the driver's seat, the back seat, etc.) will also be different. The multi-position acquisition method can help analyze the acoustic environment in the vehicle. Acoustic signal samples can come from different vehicles, and different vehicles can collect second audio signals in different scenarios. Different scenarios where the vehicle is located will generate different external noises. The second audio signal can come from in-vehicle audio data in different scenarios. The scenario where the speaker is located may also affect the characteristics of the audio signal. For example, the vehicle interior may sometimes be a quiet environment, while sometimes there may be more background noise. Analyzing the second audio signal samples in different scenarios can optimize processing methods such as noise suppression and echo cancellation. In some cases, there may be no speaker in the vehicle interior, but there are still other sounds (such as the noise of the vehicle itself, wind noise, etc.). Samples without a speaker are of great significance for analyzing the characteristics of in-vehicle environmental noise and optimizing the background noise suppression algorithm. In some acoustic signal samples, there is at least one speaker present. Analyzing the second audio signal with a speaker helps improve the accuracy of voice source localization.

[0169] Using acoustic signal samples under different conditions (such as differences in speakers, positions, and scenarios) can enhance the diversity of the training data of the localization model, enabling the localization model to better handle various situations that may occur in real life.

[0170] Step 303: Based on the acoustic signal samples, train the initial localization model to obtain the trained localization model.

[0171] The initial positioning model is a model that has not been fully trained. Usually, at the initial stage of model construction, the parameters of the positioning model have not been optimized. Using acoustic signal samples, the model parameters of the initial positioning model are adjusted through an optimization algorithm (such as the gradient descent algorithm), enabling the initial positioning model to gradually learn to accurately identify the positions of vehicle occupants from the acoustic signal samples. The trained positioning model obtained after training can accurately output the positions of vehicle occupants based on the acoustic signal samples. The network output layer of the initial positioning model is set to multi-label, where each label represents whether there is someone in a position; the loss of the multi-label is measured, and preferably, cross-entropy loss and circle loss can be selected. The initial positioning model is trained using the acoustic signal samples until the loss function of the initial positioning model converges and stops, obtaining the trained positioning model.

[0172] The trained positioning model obtained by training the initial positioning model can make full use of a large amount of training data, enabling the positioning model to learn the mapping relationship between the characteristics of acoustic signal samples at different positions and the corresponding positions. Compared with the untrained model, the positioning model can more accurately determine the position of the speaker in the vehicle, improving the accuracy and reliability of positioning.

[0173] As a refinement of step 303, when performing the step of training the initial positioning model based on the acoustic signal samples to obtain the trained positioning model, the following methods can be used but are not limited to, such as Figure 4 as shown Figure 4 The present disclosure also provides a flowchart of a method for training a positioning model, including (that is, training the initial positioning model based on the acoustic signal samples, including):

[0174] Step 401, obtaining enhanced acoustic signal samples based on the acoustic signal samples, where the enhanced acoustic signal samples include a third ultrasonic signal and a third audio signal.

[0175] The acoustic signal samples are processed for signal diversification to obtain enhanced acoustic signal samples; the third ultrasonic signal refers to the ultrasonic signal obtained by processing the second ultrasonic signal through signal diversification (such as complex plane rotation, filtering, etc.). On the basis of retaining the spatial characteristics of the original ultrasonic signal, more variations are introduced to simulate different actual situations. The third audio signal is the audio signal obtained by diversifying the second audio signal (such as adding noise, adjusting volume, etc.), enabling the model to adapt to changes in different noise environments and speech characteristics.

[0176] Step 402, the enhanced acoustic signal samples satisfy at least one of the following: the third audio signal in at least part of the enhanced acoustic signal samples is obtained by adding a noise signal to the second audio signal, where different third audio signals correspond to not entirely identical second audio signals, and / or different third audio signals correspond to not entirely identical noise signals; the third ultrasonic signal in at least part of the enhanced acoustic signal samples is obtained by rotating the second ultrasonic signal in the complex plane, where different third ultrasonic signals correspond to not entirely identical second ultrasonic signals, and / or different third ultrasonic signals correspond to not entirely identical rotation angles; at least part of the enhanced acoustic signal samples are obtained by combining the third ultrasonic signal with the first weight and the third audio signal with the second weight, where different third ultrasonic signals correspond to not entirely identical first weights, and / or different third audio signals correspond to not entirely identical second weights.

[0177] That different third audio signals correspond to not entirely identical second audio signals means that different third audio signals can be derived from different second audio signals or the same second audio signal. For example, assume there are 3 third audio signals, namely the third audio signal 1, the third audio signal 2, and the third audio signal 3. The third audio signal 1 is derived from the second audio signal 1, the third audio signal 2 is derived from the second audio signal 2, and the third audio signal 3 is derived from the second audio signal 1.

[0178] That different third audio signals correspond to not entirely identical noise signals means that different third audio signals can be derived from the same noise signal or different noise signals. For example, assume there are 3 third audio signals, namely the third audio signal 1, the third audio signal 2, and the third audio signal 3. The third audio signal 1 is derived from the noise signal 1, the third audio signal 2 is derived from the noise signal 2, and the third audio signal 3 is derived from the noise signal 3.

[0179] Different third audio signals can also be simultaneously derived from the same second audio signal and different noise signals, or from different second audio signals and the same noise signal, or from different second audio signals and different noise signals. For example, assume there are 3 third audio signals, namely the third audio signal 1, the third audio signal 2, and the third audio signal 3. The third audio signal 1 is derived from the second audio signal 1 and the noise signal 1, the third audio signal 2 is derived from the second audio signal 1 and the noise signal 2, and the third audio signal 3 is derived from the second audio signal 2 and the noise signal 1.

[0180] Different third ultrasonic signals can be derived from different second ultrasonic signals or the same second ultrasonic signal. For example, assume there are 3 third ultrasonic signals, namely the third ultrasonic signal 1, the third ultrasonic signal 2, and the third ultrasonic signal 3. Among them, the third ultrasonic signal 1 is derived from the second ultrasonic signal 1, the third ultrasonic signal 2 is derived from the second ultrasonic signal 2, and the third ultrasonic signal 3 is derived from the second ultrasonic signal 3.

[0181] The rotation angle refers to the angular parameter used when rotating the ultrasonic signal in the complex plane. The rotation angles of different third ultrasonic signals can be the same or different. For example, assume there are 3 third ultrasonic signals, namely the third ultrasonic signal 1, the third ultrasonic signal 2, and the third ultrasonic signal 3. Among them, the rotation angle of the third ultrasonic signal 1 is the rotation angle a, the rotation angle of the third ultrasonic signal 2 is the rotation angle b, and the rotation angle of the third ultrasonic signal 3 is the rotation angle a.

[0182] Different third ultrasonic signals can also be simultaneously derived from the same second ultrasonic signal and different rotation angles, or from different second ultrasonic signals and the same rotation angle, or from different second ultrasonic signals and different rotation angles. For example, assume there are 3 third ultrasonic signals, namely the third ultrasonic signal 1, the third ultrasonic signal 2, and the third ultrasonic signal 3. Among them, the third ultrasonic signal 1 is derived from the second ultrasonic signal 1 and the rotation angle a, the third ultrasonic signal 2 is derived from the second ultrasonic signal 1 and the rotation angle b, and the third ultrasonic signal 3 is derived from the second ultrasonic signal 3 and the rotation angle c.

[0183] The first weights of different third ultrasonic signals can be the same or different. For example, assume there are 3 third ultrasonic signals, namely the third ultrasonic signal 1, the third ultrasonic signal 2, and the third ultrasonic signal 3. Among them, the first weight of the third ultrasonic signal 1 is the first weight a, the first weight of the third ultrasonic signal 2 is the second weight b, and the first weight of the third ultrasonic signal 3 is the first weight a; the second weights of different third audio signals can be the same or different. For example, assume there are 3 third audio signals, namely the third audio signal 1, the third audio signal 2, and the third audio signal 3. Among them, the second weight of the third audio signal 1 is the second weight a, the second weight of the third audio signal 2 is the second weight b, and the second weight of the third audio signal 3 is the second weight a.

[0184] The noise signal is an analog noise signal, which is an analog of the noise signal in the interior space of a test vehicle under different noise scenarios; the complex plane rotation operation can adjust the phase and amplitude of the second ultrasonic signal, increasing the data diversity of the third ultrasonic signal. The third ultrasonic signal and the third audio signal are weighted and combined according to different weights. The first weight is used to adjust the contribution of the third ultrasonic signal, and the second weight is used to adjust the contribution of the third audio signal. The different weights can be adjusted according to the actual situation. By adjusting the first weight and the second weight, the third ultrasonic signal with the first weight and the third audio signal with the second weight can be combined into different acoustic signal samples, increasing the diversity of the acoustic signal samples; where "not completely the same" includes partially the same and partially different, or completely different.

[0185] Step 403: Based on the acoustic signal samples and / or the enhanced acoustic signal samples, train the initial positioning model to obtain the trained positioning model.

[0186] The initial positioning model can be trained only using acoustic signal samples, or only using enhanced acoustic signal samples, or using both acoustic signal samples and enhanced acoustic signal samples at the same time. The collected acoustic signal samples and enhanced acoustic signal samples can be divided according to a certain ratio (such as 70% training set, 15% validation set, 15% test set) to obtain a training set, a validation set, and a test set. The training set is used for updating the parameters of the positioning model, the validation set is used for adjusting hyperparameters and evaluating the performance of the positioning model on unseen data, and the test set is used for finally evaluating the generalization ability of the positioning model.

[0187] The enhanced acoustic signal samples have higher diversity and complexity, covering various possible actual situations. When the positioning model learns the acoustic signal samples, it can come into contact with richer acoustic features and scenarios, thereby improving the generalization ability of the positioning model.

[0188] As a refinement of step 403, when training the initial positioning model based on the acoustic signal samples, it can be implemented in, but not limited to, the following ways, such as Figure 5 as shown, Figure 5 This disclosure embodiment also provides a flowchart of a method for training a positioning model, including:

[0189] Step 501: Obtain the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal.

[0190] Step 502: Input the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal into the initial positioning model for training.

[0191] As a refinement of step 403, when training the initial positioning model based on the enhanced acoustic signal samples, the following methods can be adopted but are not limited to them. Please continue to refer to Figure 5 , including:

[0192] Step 503: Obtain the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal.

[0193] Step 504: Input the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal into the initial positioning model for training to obtain a trained positioning model or a positioning model to be further trained. Further training can be implemented using the training methods provided in the embodiments of the present disclosure.

[0194] When training the initial positioning model, the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal can be input into the initial positioning model for training once first, and then the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal can be input into the initial positioning model for training once.

[0195] When training the initial positioning model, the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal can also be input into the initial positioning model for training once first, and then the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal can be input into the initial positioning model for training once.

[0196] When training the initial positioning model, the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal can be input into the initial positioning model for training once first, and then the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal can be input into the initial positioning model for training once, and then the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal can be input into the initial positioning model for training once again, and then the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal can be input into the initial positioning model for training once again, and so on, until the output result of the verification acoustic signal sample or the enhanced acoustic signal sample is similar to the label.

[0197] The initial positioning model can also be trained by inputting the Mel spectrum corresponding to the second audio signal, the channel impulse response spectrum corresponding to the second ultrasonic signal, the Mel spectrum corresponding to the third audio signal, and the channel impulse response spectrum corresponding to the third ultrasonic signal into the initial positioning model until the output result of the verified acoustic signal sample or the enhanced acoustic signal sample is similar to the label. The training end conditions of the positioning model include but are not limited to the loss convergence of the positioning model, the number of iterations of the positioning model reaching a preset number, and the accuracy of the output result of the positioning model reaching a preset threshold.

[0198] The positioning model extracts features from the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal, or the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal. Feature extraction includes independently normalizing each channel of the Mel spectrum corresponding to the second audio signal and / or the Mel spectrum corresponding to the third audio signal, then dividing it into blocks, and using the feature extraction layer of the positioning model to extract features for each block. After splicing the features of all channels, the features of the Mel spectrum are formed. By independently normalizing each channel of the channel impulse response spectrum corresponding to the second ultrasonic signal or the channel impulse response spectrum corresponding to the third ultrasonic signal, then dividing it into blocks, using the feature extraction layer of the positioning model to extract features for each block, and splicing the features of all channels, the features of the channel impulse response spectrum are formed. After splicing the features of the Mel spectrum and the features of the channel impulse response spectrum of the same acoustic signal sample, it is input into the core neural network structure of the positioning model. The network output layer of the core neural network structure is set to multi-label, and each label represents whether there is someone in a vehicle interior. The core neural network outputs the positioning result of the speaker.

[0199] By using acoustic signal samples and enhanced acoustic signal samples for training, the positioning model can learn richer and more comprehensive acoustic features, thereby improving the accuracy of speaker positioning.

[0200] As a refinement of step 502, when performing the training by inputting the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal into the initial positioning model, it can be implemented in but not limited to the following ways, such as Figure 6 shown Figure 6 The present disclosure embodiment also provides a flowchart of a training method for a positioning model, including:

[0201] Step 601, extracting the first Mel spectrum feature in the Mel spectrum corresponding to the second audio signal, and extracting the first channel impulse response spectrum feature in the channel impulse response spectrum corresponding to the second ultrasonic signal.

[0202] The first Mel-spectrum feature refers to the key features extracted from the Mel-spectrum of the second audio signal, which contains information such as the frequency distribution and energy change of speech, reflecting the speech characteristics of the speaker and the spatial position correlation. The first channel impulse response spectrum feature refers to the features extracted from the channel impulse response spectrum of the second ultrasonic signal, which contains information such as the multipath effect and signal attenuation of the acoustic wave propagation path, and is used for the localization model to analyze the spatial structure.

[0203] The Mel-spectrum can effectively extract the time-frequency features of the audio signal, reduce the computational complexity, and at the same time retain the key information of the speech signal. The channel impulse response spectrum can describe the propagation characteristics of the acoustic signal in a complex environment and provide a more accurate basis for localization.

[0204] Step 602: Concatenate the first Mel-spectrum feature and the first channel impulse response spectrum feature according to a preset feature dimension to obtain a first concatenated feature.

[0205] The preset feature dimension refers to the dimension standard of the feature vector preset before feature concatenation. The preset feature dimension determines the vector length and structure of the concatenated feature. For example, if the dimension of the first Mel-spectrum feature is 100 and the dimension of the first channel impulse response spectrum feature is 50, the preset feature dimension can be set to 150.

[0206] The first concatenated feature refers to the feature formed after concatenating the first Mel-spectrum feature and the first channel impulse response spectrum feature. The concatenated feature contains both the low-frequency acoustic information reflected by the first Mel-spectrum feature and the high-frequency signal propagation information reflected by the first channel impulse response spectrum feature, which is the result of the fusion of the two features.

[0207] Step 603: Input the first concatenated feature into the core neural network structure of the initial localization model to obtain the multi-label output of the network output layer of the localization model; adjust the parameters of the localization model based on the loss corresponding to the multi-label output.

[0208] The loss refers to a function used to measure the difference between the predicted value and the true value of the localization model, guiding the update direction of the model parameters to optimize the model performance. The model parameters refer to the adjustable parameters in the localization model, which are used to optimize the performance of the localization model, such as weights and biases. The localization model uses the first concatenated feature to obtain the localization prediction value of the speaker, and generates a loss according to the prediction value and the true localization value of the speaker. The loss contains model parameters, and the parameters are continuously adjusted through an optimization algorithm to enable the model to better fit the training data.

[0209] Step 604, iteratively execute the first training process at least once to obtain a trained positioning model or a positioning model to be further trained. The further training can be implemented by using the training methods provided in the embodiments of the present disclosure. The first training process includes: inputting the first concatenated feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0210] The first concatenated feature recorded in each first training process can be all the Mel spectrograms corresponding to all the second audio signals and all the channel impulse response spectrograms corresponding to all the second ultrasonic signals, or a part of all the Mel spectrograms corresponding to all the second audio signals and all the channel impulse response spectrograms corresponding to all the second ultrasonic signals. The Mel spectrograms corresponding to the second audio signals and the channel impulse response spectrograms corresponding to the second ultrasonic signals input in each first training process can be the same part of the Mel spectrograms corresponding to the second audio signals and the channel impulse response spectrograms corresponding to the second ultrasonic signals, or different parts of the Mel spectrograms corresponding to the second audio signals and the channel impulse response spectrograms corresponding to the second ultrasonic signals.

[0211] By optimizing the loss function, the model parameters are adjusted to make the model output more accurate. The gradual convergence of the loss function ensures the continuous optimization of the model during the training process, thereby improving the accuracy and stability of the prediction.

[0212] As a refinement of step 504, when performing the training by inputting the Mel spectrogram corresponding to the third audio signal and the channel impulse response spectrogram corresponding to the third ultrasonic signal into the initial positioning model, it can be implemented by but not limited to the following methods, such as Figure 7 shown Figure 7 The present disclosure also provides a flowchart of a training method for a positioning model, including:

[0213] Step 701, extract the second Mel spectrogram feature from the Mel spectrogram corresponding to the third audio signal, and extract the second channel impulse response spectrogram feature from the channel impulse response spectrogram corresponding to the third ultrasonic signal.

[0214] Step 702, concatenate the second Mel spectrogram feature and the second channel impulse response spectrogram feature according to the preset feature dimension to obtain a second concatenated feature.

[0215] Step 703, input the second concatenated feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0216] Step 704: Iteratively execute the second training process at least once to obtain a trained positioning model or a positioning model to be further trained. The further training can be implemented by using the training methods provided in the embodiments of the present disclosure. The second training process includes: inputting the second concatenated feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0217] The training method of the embodiments of the present disclosure is the same as that of the foregoing embodiments and will not be repeated here.

[0218] In practical applications, the training method of the initial positioning model can also be implemented but is not limited to the following methods. For example Figure 8 as shown Figure 8 FIG. is a flowchart of a training method for a positioning model provided by an embodiment of the present disclosure, including:

[0219] Step 801: Extract the first Mel spectrum feature in the Mel spectrum corresponding to the second audio signal, extract the first channel impulse response spectrum feature in the channel impulse response spectrum corresponding to the second ultrasonic signal, extract the second Mel spectrum feature in the Mel spectrum corresponding to the third audio signal, and extract the second channel impulse response spectrum feature in the channel impulse response spectrum corresponding to the third ultrasonic signal.

[0220] Step 802: Concatenate the first Mel spectrum feature and the first channel impulse response spectrum feature according to a preset feature dimension to obtain a first concatenated feature, and concatenate the second Mel spectrum feature and the second channel impulse response spectrum feature according to the preset feature dimension to obtain a second concatenated feature.

[0221] Step 803: Input the first concatenated feature and the second concatenated feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0222] Step 804: Iteratively execute the third training process at least once to obtain a trained positioning model or a positioning model to be further trained. The further training can be implemented by using the training methods provided in the embodiments of the present disclosure. The third training process includes: inputting the first concatenated feature and the second concatenated feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0223] In practical applications, the training method of the initial positioning model can also be implemented by, but not limited to, the following methods. For example, Figure 9 as shown, Figure 9 This disclosure embodiment also provides a flowchart of a training method for a positioning model, including:

[0224] Step 901, extract the first Mel spectrum feature in the Mel spectrum corresponding to the second audio signal, and extract the first channel impulse response spectrum feature in the channel impulse response spectrum corresponding to the second ultrasonic signal.

[0225] Step 902, splice the first Mel spectrum feature and the first channel impulse response spectrum feature according to a preset feature dimension to obtain a first spliced feature.

[0226] Step 903, input the first spliced feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0227] Step 904, iteratively execute at least one first training process. The first training process includes: input the first spliced feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0228] Step 905, extract the second Mel spectrum feature in the Mel spectrum corresponding to the third audio signal, and extract the second channel impulse response spectrum feature in the channel impulse response spectrum corresponding to the third ultrasonic signal.

[0229] Step 906, splice the second Mel spectrum feature and the second channel impulse response spectrum feature according to the preset feature dimension to obtain a second spliced feature.

[0230] Step 907, input the second spliced feature into the core neural network structure of the initial positioning model trained based on the first Mel spectrum feature in the Mel spectrum corresponding to the second audio signal and the first channel impulse response spectrum feature in the channel impulse response spectrum corresponding to the second ultrasonic signal to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0231] Step 908: Iteratively execute the fourth training process at least once to obtain a trained positioning model or a positioning model to be further trained. The further training can be implemented by using the training methods provided in the embodiments of the present disclosure. The fourth training process includes: inputting the second spliced feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0232] In an implementable manner of the embodiments of the present disclosure, the Mel spectrum corresponding to the second audio signal, the channel impulse response spectrum corresponding to the second ultrasonic signal, the Mel spectrum corresponding to the third audio signal, and the channel impulse response spectrum corresponding to the third ultrasonic signal can be simultaneously input into the initial positioning model for training, including: inputting the first spliced feature and the second spliced feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output; iteratively executing the third training process at least once. The third training process includes: inputting the first spliced feature and the second spliced feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0233] The training method of the embodiments of the present disclosure is the same as the training method of the foregoing embodiments, and will not be repeated here.

[0234] As a refinement of step 301, when performing the acquisition of the acoustic signal sample, it can be implemented by, but not limited to, the following methods, such as Figure 10 as shown Figure 10 This disclosure also provides a flowchart of a method for acquiring an acoustic signal sample, including:

[0235] Step 1001: Acquire an acoustic signal to be processed; the acoustic signal to be processed includes an ultrasonic signal to be processed and an audio signal to be processed.

[0236] The acoustic signal to be processed refers to the raw acquired acoustic signal that has not been processed.

[0237] Step 1002: The preprocessing includes at least one of the following: noise reduction processing; filtering processing; filtering processing.

[0238] Noise reduction processing refers to removing white noise in the audio signal to be processed. The ultrasonic signal to be processed can be denoised through a filter to reduce the impact of white noise on the ultrasonic signal to be processed. Filtering processing refers to designing a band-pass filter for a specific frequency of the ultrasonic signal (such as 40 kHz) to suppress interference in other frequency bands. Filtering treatment refers to using a digital signal processing algorithm to remove abnormal spikes or periodic interference in the ultrasonic signal.

[0239] To better understand the filtering process and the filtering treatment, an example is provided to obtain the path length of the ultrasonic signal to be processed. The ultrasonic signal to be processed corresponding to the target path length within the preset path length threshold in the path length is determined as the target ultrasonic signal to be processed. The ultrasonic signal with path fluctuations in the target ultrasonic signal is determined as the second ultrasonic signal. Among them, the path length refers to the length of the path that the ultrasonic signal passes through when propagating in the vehicle interior space. The path length can be a physical distance or an equivalent distance of signal propagation. The path length of the ultrasonic signal is calculated using the position information of the sound source and the receiving point, combined with the propagation characteristics of the ultrasonic signal (such as speed, reflection, refraction, etc.). The preset path length threshold refers to a range set according to actual application requirements and experience, used to screen out ultrasonic signals that meet specific propagation distance requirements. For example, the upper and lower limits of the path length can be set according to the size of the vehicle interior space and the possible position range of the personnel. The target path length refers to the path length within the preset path length threshold, that is, the path length that meets the screening conditions. The target ultrasonic signal refers to the sample audio signal corresponding to the target path length. By setting an appropriate path length threshold, ultrasonic signals that have undergone long-term propagation or complex processing can be effectively excluded. By screening ultrasonic signals with appropriate path lengths, the impact of noise is reduced. By screening out the second ultrasonic signal with path fluctuations, signals with relatively stable paths and lacking dynamic information are excluded, avoiding interference from these invalid data on the training of the positioning model.

[0240] Step 1003, preprocess the acoustic signal to be processed to obtain the acoustic signal sample.

[0241] Through noise reduction and filtering processing, background noise and interference signals can be effectively removed, improving the signal-to-noise ratio of the acoustic signal.

[0242] To better understand the entire process of locating a speaker in a vehicle, as Figure 11 shown, Figure 11The figure is a flowchart of the entire process for locating an in-vehicle speaker provided by an embodiment of the present disclosure. The data acquisition module is used to collect audio signals, the data preprocessing module is used to preprocess the audio signals, the data feature enhancement module is used to increase the diversity of the audio signals, the offline deep learning training module for passenger position is used to train the positioning model, the deep learning model for online predicting the passenger position is used to perform positioning prediction, and the upper-layer application module is used to report the obtained position to the application layer for use. For example, it can adjust the zone control of the air conditioner and adjust the playback position of music.

[0243] All operations performed in all embodiments of the present disclosure are carried out under the authorization of the user and strictly comply with relevant laws and regulations such as privacy and security.

[0244] Corresponding to the above-mentioned method for locating an in-vehicle speaker, the present invention also proposes an in-vehicle speaker positioning device. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment, and will not be elaborated herein.

[0245] Figure 12 The figure is a schematic structural diagram of an in-vehicle speaker positioning device 1100 provided by an embodiment of the present disclosure. The in-vehicle speaker positioning device includes:

[0246] An acquisition unit 111, configured to acquire acoustic signals in the in-vehicle space; the acoustic signals include a first ultrasonic signal and a first audio signal corresponding to the speaker; the ultrasonic signal is broadcast by a speaker in the vehicle when performing in-vehicle speaker positioning.

[0247] A determination unit 112, configured to determine a positioning result of the speaker according to the acoustic signals.

[0248] According to the in-vehicle speaker positioning device proposed by the present disclosure, the device acquires acoustic signals in the in-vehicle space that include a first ultrasonic signal and a first audio signal corresponding to the speaker, and utilizes the characteristic that the first ultrasonic signal is broadcast by the in-vehicle speaker. Combining with the unique voice characteristics of the speaker in the first audio signal, during the positioning process, it not only performs initial positioning by means of the spatial propagation characteristics of the ultrasonic signal, but also effectively distinguishes whether the movable entity is an in-vehicle person through the first audio signal, improving the accuracy of in-vehicle speaker positioning.

[0249] Further, in a possible implementation manner of an embodiment of the present disclosure, as Figure 13 shown, the determination unit 112 includes:

[0250] An acquisition module 1121, configured to acquire a Mel spectrum corresponding to the first audio signal and a channel impulse response spectrum corresponding to the first ultrasonic signal.

[0251] An input module 1122, configured to input the Mel spectrum corresponding to the first audio signal and the channel impulse response spectrum corresponding to the first ultrasonic signal into the trained positioning model, and obtain a positioning result output by the trained positioning model.

[0252] Further, in a possible implementation manner of the embodiments of the present disclosure, as Figure 13 shown, the apparatus further includes:

[0253] An acquisition unit 113, configured to acquire an acoustic signal sample, where the acoustic signal sample includes a second ultrasonic signal and a second audio signal;

[0254] A training unit 114, configured to train an initial positioning model based on the acoustic signal sample to obtain the trained positioning model;

[0255] Wherein, the acoustic signal sample satisfies at least one of the following:

[0256] The speakers corresponding to the second audio signals in different acoustic signal samples are not completely the same;

[0257] The positions of the speakers corresponding to the second audio signals in different acoustic signal samples in the vehicle are not completely the same;

[0258] The vehicles where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same;

[0259] The scenes where the vehicles where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same;

[0260] The scenes where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same;

[0261] At least some of the second audio signals in the acoustic signal samples correspond to a scene without a speaker;

[0262] At least some of the second audio signals in the acoustic signal samples correspond to at least one speaker.

[0263] Further, in a possible implementation manner of the embodiments of the present disclosure, the training unit 114 is further configured to:

[0264] Based on the acoustic signal sample, obtain an enhanced acoustic signal sample, where the enhanced acoustic signal sample includes a third ultrasonic signal and a third audio signal;

[0265] Based on the acoustic signal sample and / or the enhanced acoustic signal sample, train the initial positioning model to obtain the trained positioning model;

[0266] Among them, the enhanced acoustic signal samples satisfy at least one of the following:

[0267] The third audio signal in at least part of the enhanced acoustic signal samples is obtained by adding a noise signal to the second audio signal, where different third audio signals correspond to not completely identical second audio signals, and / or different third audio signals correspond to not completely identical noise signals;

[0268] The third ultrasonic signal in at least part of the enhanced acoustic signal samples is obtained by rotating the second ultrasonic signal in the complex plane, where different third ultrasonic signals correspond to not completely identical second ultrasonic signals, and / or different third ultrasonic signals correspond to not completely identical rotation angles;

[0269] At least part of the enhanced acoustic signal samples are obtained by combining the third ultrasonic signal with a first weight and the third audio signal with a second weight, where different third ultrasonic signals correspond to not completely identical first weights, and / or different third audio signals correspond to not completely identical second weights.

[0270] Furthermore, in a possible implementation manner of the embodiments of the present disclosure, the training unit 114 is further configured to:

[0271] Obtain the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal;

[0272] Input the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal into the initial positioning model for training;

[0273] The training unit 114 is further configured to:

[0274] Obtain the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal;

[0275] Input the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal into the initial positioning model for training.

[0276] Furthermore, in a possible implementation manner of the embodiments of the present disclosure, the training unit 114 is further configured to:

[0277] Extract the first Mel spectrum feature in the Mel spectrum corresponding to the second audio signal, and extract the first channel impulse response spectrum feature in the channel impulse response spectrum corresponding to the second ultrasonic signal;

[0278] Stitch the first Mel spectrum feature and the first channel impulse response spectrum feature according to a preset feature dimension to obtain a first stitched feature;

[0279] Input the first stitching feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output;

[0280] Iteratively execute the first training process at least once. The first training process includes: inputting the first stitching feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output;

[0281] The training unit 114 is further configured to:

[0282] Extract the second Mel-spectrum feature from the Mel-spectrum corresponding to the third audio signal, and extract the second channel impulse response spectrum feature from the channel impulse response spectrum corresponding to the third ultrasonic signal;

[0283] Stitch the second Mel-spectrum feature and the second channel impulse response spectrum feature according to the preset feature dimension to obtain a second stitching feature;

[0284] Input the second stitching feature into the core neural network structure of the initial positioning model to obtain the multi-label output of the network output layer of the positioning model; adjust the parameters of the positioning model based on the loss corresponding to the multi-label output;

[0285] Iteratively execute the second training process at least once. The second training process includes: inputting the second stitching feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time to obtain the multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

[0286] Further, in a possible implementation manner of the embodiment of the present disclosure, the obtaining unit 113 is further configured to:

[0287] Obtain an acoustic signal to be processed; the acoustic signal to be processed includes an ultrasonic signal to be processed and an audio signal to be processed;

[0288] Preprocess the acoustic signal to be processed to obtain an acoustic signal sample;

[0289] Wherein, the preprocessing includes at least one of the following:

[0290] Noise reduction processing;

[0291] Filtering processing;

[0292] Filtering process.

[0293] Since the device provided in the embodiments of the present disclosure corresponds to the methods provided in the above several embodiments, the implementation manners of the methods are also applicable to the device provided in this embodiment and will not be described in detail in this embodiment.

[0294] In the above embodiments provided in this application, the methods and devices provided in the embodiments of this application are introduced. To implement each function in the methods provided in the embodiments of this application, an electronic device may include a hardware structure and software modules, and implement the above functions in the form of a hardware structure, software modules, or a combination of a hardware structure and software modules. A certain function among the above functions may be executed in the manner of a hardware structure, software modules, or a combination of a hardware structure and software modules.

[0295] Figure 14 is a block diagram of a vehicle 1200 shown according to an exemplary embodiment. For example, the vehicle 1200 may be a hybrid vehicle, or a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. The vehicle 1200 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.

[0296] Refer to Figure 14 , the vehicle 1200 may include various subsystems. For example, the infotainment system 1212, the perception system 1220, the decision control system 1230, the drive system 1240, and the computing platform 1250. Among them, the vehicle 1200 may further include more or fewer subsystems, and each subsystem may include multiple components. In addition, each subsystem and each component of the vehicle 1200 may be interconnected in a wired or wireless manner.

[0297] In some embodiments, the infotainment system 1212 may include a communication system, an entertainment system, a navigation system, and the like.

[0298] The perception system 1220 may include several sensors for sensing information about the environment around the vehicle 1200. For example, the perception system 1220 may include a global positioning system (the global positioning system may be a GPS system, or a Beidou system, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter wave radar, ultrasonic radar, and a camera device.

[0299] The decision control system 1230 may include a computing system, a vehicle controller, a steering system, an accelerator, and a braking system.

[0300] The drive system 1240 may include components that provide motive power for the vehicle 1200. In one embodiment, the drive system 1240 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of an internal combustion engine, an electric motor, and an air compression engine. The engine is capable of converting the energy provided by the energy source into mechanical energy.

[0301] Some or all of the functions of the vehicle 1200 are controlled by the computing platform 1250. The computing platform 1250 may include at least one processor 1251 and a memory 1252, and the processor 1251 may execute instructions 1253 stored in the memory 1252.

[0302] The processor 1251 may be any conventional processor, such as a commercially available CPU. The processor may also include, for example, a Graphic Process Unit (GPU), a Field Programmable Gate Array (FPGA), a System on Chip (SOC), an Application Specific Integrated Circuit (ASIC), or a combination thereof.

[0303] The memory 1252 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0304] In addition to the instructions 1253, the memory 1252 may also store data, such as road maps, route information, data on the position, direction, speed, etc. of the vehicle. The data stored in the memory 1252 can be used by the computing platform 1250.

[0305] In an embodiment of the present disclosure, the processor 1251 may execute the instructions 1253 to complete all or part of the steps of the method for locating an in-vehicle speaker described above.

[0306] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of the method for locating an in-vehicle speaker provided by the present disclosure are implemented.

[0307] In addition, as used herein, the word "exemplary" is used to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the word exemplary is intended to present concepts in a concrete fashion. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless specified otherwise, or clear from the context, "X applies A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then "X applies A or B" is satisfied under any of the foregoing instances. Additionally, unless specified otherwise or clear from the context that it refers to the singular form, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".

[0308] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. Specifically with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. Additionally, although a particular feature of the present disclosure may have been disclosed with respect to only one of several implementations, such a feature may, as may be desired and advantageous for any given or particular application, be combined with one or more other features of other implementations. Further, with respect to the use of "comprises", "comprising", "has", "having", "includes", or variants thereof in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "including".

[0309] Other embodiments of the present disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known in the art or conventional technical means not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0310] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

[0311] In the foregoing detailed description, reference has been made to the accompanying drawings, in which specific aspects in which the present disclosure may be practiced are shown by way of illustration. In this regard, directional or positional relationship-indicating terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. may be used with reference to the orientation of the described figures. Since the components of the described devices may be positioned in a number of different orientations, the directional terms may be used for illustrative purposes and not for purposes of limitation. It should be understood that other aspects may be utilized and structural or logical changes may be made without departing from the concepts of the present disclosure. Accordingly, the following detailed description should not be taken in a limiting sense.

[0312] It should be understood that, unless otherwise specifically stated, the features of some embodiments of the various present disclosures described herein may be combined with each other. As used herein, the term "and / or" includes any one of the related listed items and any combination of any two or more of them; similarly, "at least one of... " includes any one of the related listed items and any combination of any two or more of them.

[0313] Although terms such as "first", "second", and "third" may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. Rather, these terms are only used to distinguish one component, part, region, layer, or section from another. Thus, the first component, part, region, layer, or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer, or section without departing from the teachings of the examples. Additionally, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description herein, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically and explicitly defined.

[0314] It should be understood that, as used herein, spatial relative terms, such as "above", "upper", "below", and "lower", are used to describe the relationship of one element shown in the figures to another element. In addition to the orientation depicted in the figures, such spatial relative terms are also intended to encompass different orientations of the device during use or operation. For example, if the device in the figures is flipped over, an element described as "above" or "upper" relative to another element will then be "below" or "lower" relative to that other element. Thus, depending on the spatial orientation of the device, the term "above" encompasses both the above and below orientations. The device may have other orientations (e.g., rotated 90 degrees or in other orientations), and the spatial relative terms used herein should be interpreted accordingly.

Claims

1. A method for locating a speaker in a car, characterized in that: The method comprises: Acquiring an acoustic signal in the vehicle interior; the acoustic signal comprising a first ultrasonic signal and a first audio signal corresponding to the speaker; the ultrasonic signal is broadcast by a speaker in the vehicle when the vehicle performs positioning of the speaker in the vehicle; A positioning result of the speaker is determined according to the acoustic signal.

2. The method according to claim 1, characterized in that: Determining the positioning result of the speaker according to the acoustic signal includes: Acquire a Mel spectrum corresponding to the first audio signal and a channel impulse response spectrum corresponding to the first ultrasonic signal; The Mel spectrum corresponding to the first audio signal and the channel impulse response spectrum corresponding to the first ultrasonic signal are input into the trained positioning model to obtain a positioning result output by the trained positioning model.

3. The method according to claim 2, characterized in that The method further comprises: Acquiring an acoustic signal sample, wherein the acoustic signal sample includes a second ultrasonic signal and a second audio signal; Based on the acoustic signal samples, an initial positioning model is trained to obtain the trained positioning model; The acoustic signal sample satisfies at least one of the following: The speakers corresponding to the second audio signals in different acoustic signal samples are not completely the same; The positions of the speakers in the car corresponding to the second audio signals in different acoustic signal samples are not completely the same; The cars where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same; The scenes where the cars of the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same; The scenes where the speakers corresponding to the second audio signals in different acoustic signal samples are located are not completely the same; The second audio signal in at least some of the acoustic signal samples corresponds to a scene without a speaker; The second audio signal in at least a portion of the acoustic signal samples corresponds to at least one speaker.

4. The method according to claim 3, characterized in that: The step of training an initial positioning model based on the acoustic signal sample to obtain the trained positioning model includes: Based on the acoustic signal sample, an enhanced acoustic signal sample is obtained, wherein the enhanced acoustic signal sample includes a third ultrasonic signal and a third audio signal; Based on the acoustic signal samples and / or the enhanced acoustic signal samples, training the initial positioning model to obtain the trained positioning model; The enhanced acoustic signal sample satisfies at least one of the following: The third audio signal in the at least partially enhanced acoustic signal sample is obtained by adding a noise signal to the second audio signal, wherein different third audio signals correspond to different second audio signals and / or different third audio signals correspond to different noise signals; The third ultrasonic signal in at least part of the enhanced acoustic signal sample is obtained by rotating the second ultrasonic signal in a complex plane, wherein different third ultrasonic signals correspond to different second ultrasonic signals and / or different third ultrasonic signals correspond to different rotation angles; At least part of the enhanced acoustic signal samples are obtained by combining a third ultrasonic signal with a first weight and a third audio signal with a second weight, wherein different third ultrasonic signals correspond to different first weights and / or different third audio signals correspond to different second weights.

5. The method according to claim 4, characterized in that Based on the acoustic signal samples, an initial positioning model is trained, including: Acquire a Mel spectrum corresponding to the second audio signal and a channel impulse response spectrum corresponding to the second ultrasonic signal; Inputting the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal into the initial positioning model for training; Based on the enhanced acoustic signal samples, an initial positioning model is trained, including: Acquire a Mel spectrum corresponding to the third audio signal and a channel impulse response spectrum corresponding to the third ultrasonic signal; The Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal are input into the initial positioning model for training.

6. The method according to claim 5, characterized in that The step of inputting the Mel spectrum corresponding to the second audio signal and the channel impulse response spectrum corresponding to the second ultrasonic signal into the initial positioning model for training includes: Extracting a first mel spectrum feature from the mel spectrum corresponding to the second audio signal, and extracting a first channel impulse response spectrum feature from the channel impulse response spectrum corresponding to the second ultrasonic signal; According to a preset feature dimension, the first mel spectrum feature and the first channel impulse response spectrum feature are spliced ​​to obtain a first spliced ​​feature; Inputting the first splicing feature into the core neural network structure of the initial positioning model to obtain a multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output; Iteratively executing at least one first training process, the first training process comprising: inputting the first splicing feature into a core neural network structure of a positioning model obtained by adjusting the parameters of the positioning model in a previous time, to obtain a multi-label output of a network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output; The step of inputting the Mel spectrum corresponding to the third audio signal and the channel impulse response spectrum corresponding to the third ultrasonic signal into the initial positioning model for training includes: Extracting a second Mel spectrum feature from the Mel spectrum corresponding to the third audio signal, and extracting a second channel impulse response spectrum feature from the channel impulse response spectrum corresponding to the third ultrasonic signal; According to the preset feature dimension, concatenate the second mel spectrum feature and the second channel impulse response spectrum feature to obtain a second concatenated feature; Inputting the second splicing feature into the core neural network structure of the initial positioning model to obtain a multi-label output of the network output layer of the positioning model; adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output; Iteratively execute the second training process at least once, wherein the second training process includes: inputting the second splicing feature into the core neural network structure of the positioning model obtained by adjusting the parameters of the positioning model in the previous time, to obtain a multi-label output of the network output layer of the positioning model; and adjusting the parameters of the positioning model based on the loss corresponding to the multi-label output.

7. The method according to claim 3, characterized in that The acquiring of acoustic signal samples comprises: Acquiring an acoustic signal to be processed; the acoustic signal to be processed includes an ultrasonic signal to be processed and an audio signal to be processed; Preprocessing the acoustic signal to be processed to obtain the acoustic signal sample; Wherein, the preprocessing includes at least one of the following: Noise reduction processing; Filter processing; Filtration process.

8. A device for locating a speaker in a car, characterized in that: The device comprises: A collection unit, used for collecting acoustic signals in the vehicle interior space; the acoustic signals include a first ultrasonic signal and a first audio signal corresponding to the speaker; the ultrasonic signal is broadcast by a speaker in the vehicle when the vehicle performs positioning of the speaker in the vehicle; A determination unit is used to determine a positioning result of the speaker according to the acoustic signal.

9. The device according to claim 8, characterized in that The determining unit comprises: an acquisition module, configured to acquire a Mel spectrum corresponding to the first audio signal and a channel impulse response spectrum corresponding to the first ultrasonic signal; An input module is used to input the Mel spectrum corresponding to the first audio signal and the channel impulse response spectrum corresponding to the first ultrasonic signal into the trained positioning model to obtain a positioning result output by the trained positioning model.

10. A vehicle, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to: Implement the method according to any one of claims 1 to 7.

11. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, the mobile terminal is enabled to execute the method according to any one of claims 1 to 7.