Methods, apparatus, devices and computer-readable storage media for processing audio signals

By determining the user's location inside the vehicle and controlling the activation of the audio acquisition device, the voiceprint feature information of non-target users is processed and eliminated, thus solving the limitations of in-vehicle microphones and achieving greater flexibility and accuracy in vehicle interaction.

CN119255155BActive Publication Date: 2025-10-31CHERY AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411355886.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-10-31
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

The in-vehicle microphone can only collect voice data from the driver's position, resulting in poor flexibility and effectiveness of vehicle interaction.

Method used

The location of each user is determined by acquiring images inside the vehicle, and the audio acquisition device at the corresponding location is activated. After the audio signal is acquired, it is processed to remove the voiceprint feature information of non-target users, and the audio signal of the target user is obtained for interaction.

Benefits of technology

This allows users to interact with the vehicle from anywhere inside, improving the flexibility and accuracy of the interaction while saving resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119255155B_ABST
    Figure CN119255155B_ABST
Patent Text Reader

Abstract

This application discloses an audio signal processing method, apparatus, device, and computer-readable storage medium, belonging to the field of in-vehicle intelligent interaction technology. The method includes: acquiring an interior image of a target vehicle, the interior image being used to determine the locations of various users inside the target vehicle; controlling the activation of audio acquisition devices corresponding to the locations of each user; acquiring audio signals acquired by each audio acquisition device; for a first user, processing the audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain a reference audio signal; removing the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user, the audio signal corresponding to the first user being used for interaction between the first user and the target vehicle. This method improves the flexibility of vehicle control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of in-vehicle intelligent interaction technology, and in particular to an audio signal processing method, apparatus, device, and computer-readable storage medium. Background Technology

[0002] With the continuous development of in-vehicle intelligent interaction technology, vehicles can offer more and more functions, such as interacting with the vehicle through audio signals.

[0003] In related technologies, a vehicle microphone is installed at the vehicle's main control unit or the vehicle's dome light. When a user wants to interact with the vehicle, the user inputs voice data, the vehicle microphone collects the audio signal, processes the audio signal to obtain the interaction command, and then interacts with the vehicle according to the interaction command.

[0004] However, since users in all positions within the vehicle have interaction needs, the in-vehicle microphone can only collect voice data from the user in the driver's seat, which limits the vehicle interaction and reduces its flexibility, resulting in poor interaction performance. Summary of the Invention

[0005] This application provides a method, apparatus, device, and computer-readable storage medium for processing audio signals, which can be used to solve problems in related technologies. The technical solution is as follows:

[0006] On one hand, embodiments of this application provide a method for processing audio signals, the method comprising:

[0007] Acquire interior images of the target vehicle, which are used to determine the location of each user inside the target vehicle.

[0008] The audio acquisition device corresponding to the location of each user is activated.

[0009] Acquire audio signals collected by each audio acquisition device;

[0010] For the first user among the users, the audio signal collected by the audio acquisition device corresponding to the location of the first user is processed to obtain a reference audio signal;

[0011] The target audio signal is removed from the reference audio signal to obtain the audio signal corresponding to the first user. The target audio signal is an audio signal containing the voiceprint feature information of the second user. The second user is a user other than the first user among the users. The audio signal corresponding to the first user is used for the first user to interact with the target vehicle.

[0012] In one possible implementation, before processing the audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain the reference audio signal, the method further includes:

[0013] The audio signals acquired by each audio acquisition device are subjected to echo cancellation processing to obtain the echo-cancelled audio signals acquired by each audio acquisition device.

[0014] Based on the echo-cancelled audio signals acquired by each audio acquisition device, the voiceprint feature information corresponding to each user is determined, and the voiceprint feature information corresponding to any user is used to characterize the timbre of any user.

[0015] The step of processing the audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain a reference audio signal includes:

[0016] Beamforming is performed on the echo-cancelled audio signal acquired by the audio acquisition device at the location of the first user to obtain a reference audio signal.

[0017] In one possible implementation, determining the voiceprint feature information corresponding to each user based on the echo-cancelled audio signals acquired by each audio acquisition device includes:

[0018] For any one of the users, a first audio signal is determined from the echo-cancelled audio signals acquired by each audio acquisition device. The first audio signal is the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of any user.

[0019] A second audio signal is determined from the first audio signal. The second audio signal is the echo-cancelled audio signal collected by the audio acquisition device located at the target location in the audio acquisition device corresponding to the location of any user. The target location is determined based on the location of any user.

[0020] The second audio signal is processed to obtain the voiceprint feature information corresponding to any user.

[0021] In one possible implementation, processing the second audio signal to obtain the voiceprint feature information corresponding to any user includes:

[0022] The second audio signal is input into the voiceprint feature extraction model;

[0023] The output of the voiceprint feature extraction model is determined to be the voiceprint feature information corresponding to any user.

[0024] In one possible implementation, before removing the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user, the method further includes:

[0025] The reference audio signal is segmented to obtain multiple candidate audio signals;

[0026] Determine the voiceprint feature information corresponding to each candidate audio signal;

[0027] The candidate audio signal whose corresponding voiceprint feature information is not the voiceprint feature information corresponding to the first user is determined as the target audio signal.

[0028] In one possible implementation, after removing the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user, the method further includes:

[0029] Determine the control command that matches the audio signal corresponding to the first user;

[0030] The target vehicle is controlled according to the control command.

[0031] On the other hand, embodiments of this application provide an audio signal processing apparatus, the apparatus comprising:

[0032] The acquisition module is used to acquire images of the interior of the target vehicle, and the images are used to determine the location of each user inside the target vehicle.

[0033] The control module is used to control the activation of the audio acquisition device corresponding to the location of each user;

[0034] The acquisition module is also used to acquire audio signals acquired by each audio acquisition device;

[0035] The processing module is used to process the audio signal collected by the audio acquisition device at the location of the first user among the users to obtain a reference audio signal.

[0036] The acquisition module is further configured to remove the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user. The target audio signal is an audio signal containing the voiceprint feature information of the second user. The second user is a user other than the first user among the users. The audio signal corresponding to the first user is used for the first user to interact with the target vehicle.

[0037] In one possible implementation, the processing module is further configured to perform echo cancellation processing on the audio signals acquired by each audio acquisition device to obtain the echo-cancelled audio signals acquired by each audio acquisition device.

[0038] The device further includes:

[0039] The determination module is used to determine the voiceprint feature information corresponding to each user based on the echo-cancelled audio signals collected by each audio acquisition device. The voiceprint feature information corresponding to any user is used to characterize the timbre of any user.

[0040] The processing module is used to perform beamforming processing on the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain a reference audio signal.

[0041] In one possible implementation, the determining module is used to determine a first audio signal for any one of the users from the echo-cancelled audio signals acquired by the respective audio acquisition devices, wherein the first audio signal is the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of any one user.

[0042] A second audio signal is determined from the first audio signal. The second audio signal is the echo-cancelled audio signal collected by the audio acquisition device located at the target location in the audio acquisition device corresponding to the location of any user. The target location is determined based on the location of any user.

[0043] The second audio signal is processed to obtain the voiceprint feature information corresponding to any user.

[0044] In one possible implementation, the determining module is used to input the second audio signal into the voiceprint feature extraction model;

[0045] The output of the voiceprint feature extraction model is determined to be the voiceprint feature information corresponding to any user.

[0046] In one possible implementation, the device further includes:

[0047] The determination module is used to segment the reference audio signal to obtain multiple candidate audio signals;

[0048] Determine the voiceprint feature information corresponding to each candidate audio signal;

[0049] The candidate audio signal whose corresponding voiceprint feature information is not the voiceprint feature information corresponding to the first user is determined as the target audio signal.

[0050] In one possible implementation, the device further includes:

[0051] The determining module is used to determine the control command that matches the audio signal corresponding to the first user;

[0052] The control module is also used to control the target vehicle according to the control command.

[0053] On the other hand, embodiments of this application provide a computer device, the computer device including a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to enable the computer device to implement any of the above-described audio signal processing methods.

[0054] On the other hand, a computer-readable storage medium is also provided, wherein at least one piece of program code is stored in the computer-readable storage medium, the at least one piece of program code being loaded and executed by a processor to enable a computer to implement any of the above-described audio signal processing methods.

[0055] On the other hand, a computer program or computer program product is also provided, wherein the computer program or computer program product stores at least one computer instruction, which is loaded and executed by a processor to enable the computer to implement any of the above-described audio signal processing methods.

[0056] The technical solution provided in this application has at least the following beneficial effects:

[0057] The technical solution provided in this application activates the audio acquisition device corresponding to each user's location. This allows the user's audio signal to be acquired regardless of their position within the target vehicle, enabling each user to interact with the vehicle and thus increasing the flexibility of vehicle interaction. Furthermore, activating only the audio acquisition device corresponding to the user's location saves resources. Since the audio acquisition device at the user's location may also acquire audio signals from users in other locations, processing these signals to obtain the user's own audio signal eliminates interference from other users, allowing interaction with the target vehicle solely through the user's own audio signal, thereby improving the accuracy of vehicle interaction. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 This is a schematic diagram illustrating the implementation environment of an audio signal processing method provided in an embodiment of this application;

[0060] Figure 2 This is a flowchart of an audio signal processing method provided in an embodiment of this application;

[0061] Figure 3 This is an internal schematic diagram of a target vehicle provided in an embodiment of this application;

[0062] Figure 4 This is an internal schematic diagram of another target vehicle provided in an embodiment of this application;

[0063] Figure 5 This is a schematic diagram of the tonic and consonant registers at any position provided in an embodiment of this application;

[0064] Figure 6 This is a schematic diagram of the tonic and consonant registers at various positions provided in an embodiment of this application;

[0065] Figure 7 This is a schematic diagram of the tonic and consonant registers at various positions provided in an embodiment of this application;

[0066] Figure 8 This is a schematic diagram of the structure of an audio signal processing device provided in an embodiment of this application;

[0067] Figure 9 This is a schematic diagram of an audio signal processing system provided in an embodiment of this application;

[0068] Figure 10 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application;

[0069] Figure 11 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0070] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0071] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0072] Figure 1 This is a schematic diagram illustrating the implementation environment of an audio signal processing method provided in this application embodiment, such as... Figure 1 As shown, the implementation environment includes a computer device 101, which can be a terminal device or a server; this embodiment does not limit the specific type of device. The computer device 101 can be a device installed and running on the target vehicle, or a device capable of remotely controlling the target vehicle; this embodiment does not limit the specific type of device either. The computer device 101 is used to execute the audio signal processing method provided in this embodiment.

[0073] Optionally, computer device 101 is a terminal device. A terminal device can be any electronic device that allows human-computer interaction with a user through one or more methods such as a keyboard, touchpad, remote control, voice interaction, or handwriting device. Examples include PCs (Personal Computers), mobile phones, smartphones, PDAs (Personal Digital Assistants), wearable devices, PPCs (Pocket PCs), tablets, smart car systems, smart TVs, smart speakers, and smartwatches.

[0074] A terminal device can refer to one of multiple terminal devices; this embodiment uses only one terminal device as an example. Those skilled in the art will understand that the number of terminal devices can be more or less. For example, there may be only one terminal device, or there may be dozens or hundreds, or even more. This application embodiment does not limit the number or type of terminal devices.

[0075] When computer device 101 is a server, the server can be a single server, a server cluster consisting of multiple servers, or any of the following: a cloud computing platform or a virtualization center. This application embodiment does not limit this. The server and terminal devices communicate via a wired or wireless network. The server has data receiving, data processing, and data sending functions. Of course, the server may also have other functions, which this application embodiment does not limit.

[0076] Those skilled in the art should understand that the above-described terminal devices and servers are merely illustrative examples. Other existing or future terminal devices or servers that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.

[0077] This application provides an audio signal processing method, which can be applied to the above-mentioned... Figure 1 The implementation environment shown is as follows: Figure 2 The flowchart shown in this embodiment of the present application illustrates an audio signal processing method. This method can be implemented by... Figure 1 The computer device 101 in the middle performs the operation. For example... Figure 2 As shown, the method includes the following steps 201 to 205.

[0078] In step 201, an interior image of the target vehicle is acquired, which is used to determine the location of each user inside the target vehicle.

[0079] In an exemplary embodiment of this application, an in-vehicle camera module is installed at a reference position on the target vehicle. The reference position can be located on the rearview mirror or on the upper side of the left A-pillar; this embodiment does not limit the location. The left A-pillar is the pillar between the windshield and the driver's side window. The in-vehicle camera module can be a camera or other modules capable of image acquisition; this embodiment also does not limit the type of module.

[0080] In one possible implementation, when the target vehicle is started and all its doors are closed, the in-vehicle camera module is activated to capture images of the vehicle's interior. Once the in-vehicle camera module has captured the images, it sends them to a computer device, enabling the computer device to obtain the images of the vehicle's interior.

[0081] The in-vehicle camera module and computer equipment can communicate with each other via wired or wireless networks.

[0082] It should be noted that when the target vehicle stops and restarts, or when one of the target vehicle's doors is opened and then closed, it is necessary to re-capture images of the vehicle's interior.

[0083] In step 202, the audio acquisition device corresponding to the location of each user is turned on.

[0084] In one possible implementation, after acquiring an image of the interior of the target vehicle, the image is parsed to determine the location of each user inside the target vehicle. The location of any user can be the driver's seat, the front passenger seat, a rear seat position behind the driver's seat, a rear seat position behind the front passenger seat, or a middle rear seat position; this application does not limit the specific location.

[0085] Optionally, an audio acquisition device is installed on the left A-pillar, right A-pillar, left B-pillar, right B-pillar, left C-pillar, right C-pillar, the front dome light position, the rear dome light position, and important locations on the roof of the target vehicle. Specifically, the right A-pillar refers to the pillar between the windshield and the passenger-side side window; the left B-pillar refers to the pillar between the driver-side side window and the left rear side window; the right B-pillar refers to the pillar between the passenger-side side window and the right rear side window; the left C-pillar refers to the pillar between the left rear side window and the rear windshield; and the right C-pillar refers to the pillar between the right rear side window and the rear windshield. The audio acquisition device can be a microphone or other devices capable of acquiring audio; this application does not limit the specific device used.

[0086] In one possible implementation, after determining the location of each user, the corresponding audio acquisition device for each user's location is identified, and then the corresponding audio acquisition device for each user's location is activated. Optionally, the computer device sends an activation command to the audio acquisition device corresponding to each user's location. The audio acquisition device corresponding to each user's location powers on after receiving the activation command.

[0087] like Figure 3 This is a schematic diagram of the interior of a target vehicle provided in an embodiment of this application. Figure 3 It is known that the target vehicle is equipped with audio acquisition devices A, B, C, D, E, F, G, H, and I. Among them, audio acquisition devices A, B, D, and E correspond to location A. Audio acquisition devices B, C, E, and F correspond to location B. Audio acquisition devices D, E, G, and H correspond to location C. Audio acquisition devices E, F, H, and I correspond to location D.

[0088] like Figure 4This is an internal schematic diagram of another target vehicle provided in an embodiment of this application. Figure 4 It is known that the target vehicle is equipped with audio acquisition devices A, B, C, D, E, F, G, H, and I. Among them, audio acquisition devices A, B, D, and E correspond to location A. Audio acquisition devices B, C, E, and F correspond to location B. Audio acquisition devices D, E, and G correspond to location C. Audio acquisition devices E, G, H, and I correspond to location D. Audio acquisition devices E, F, and I correspond to location M.

[0089] In step 203, the audio signals collected by each audio acquisition device are acquired.

[0090] In one possible implementation, after the audio acquisition device corresponding to each user's location is turned on, the audio acquisition device corresponding to each user's location acquires audio signals in real time. After the audio acquisition device corresponding to each user's location acquires the audio signals, it sends the acquired audio signals to the computer device so that the computer device can obtain the audio signals acquired by each audio acquisition device.

[0091] In step 204, for the first user among all users, the audio signal collected by the audio acquisition device corresponding to the location of the first user is processed to obtain a reference audio signal.

[0092] In one possible implementation, after acquiring the audio signals from each audio acquisition device in step 203 above, before processing the audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain the reference audio signal, echo cancellation processing is performed on the audio signals acquired by each audio acquisition device to obtain echo-cancelled audio signals acquired by each audio acquisition device. Based on the echo-cancelled audio signals acquired by each audio acquisition device, the voiceprint feature information corresponding to each user is determined, and the voiceprint feature information corresponding to any user is used to characterize the timbre of any user. The process of processing the audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain the reference audio signal includes: performing beamforming processing on the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain the reference audio signal.

[0093] Echo refers to the sound signal hearing one's own voice again after a series of reflections. Optionally, echo cancellation processing can be performed on the audio signals acquired by each audio acquisition device using an adaptive filtering algorithm. Echo cancellation processing can also be performed on the audio signals acquired by each audio stimulation device using other methods. This application embodiment does not limit the specific methods used.

[0094] Optionally, the process of determining the voiceprint feature information corresponding to each user based on the echo-cancelled audio signals acquired by each audio acquisition device includes: for any user, determining a first audio signal from the echo-cancelled audio signals acquired by each audio acquisition device, wherein the first audio signal is the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of any user; determining a second audio signal from the first audio signal, wherein the second audio signal is the echo-cancelled audio signal acquired by the audio acquisition device located at the target location among the audio acquisition devices corresponding to the location of any user, wherein the target location is determined based on the location of any user; and processing the second audio signal to obtain the voiceprint feature information corresponding to any user.

[0095] Optionally, the process of processing the second audio signal to obtain the voiceprint feature information corresponding to any user includes: inputting the second audio signal into a voiceprint feature extraction model; and determining that the output of the voiceprint feature extraction model is the voiceprint feature information corresponding to any user. The voiceprint feature extraction model can be any model capable of acquiring voiceprint feature information.

[0096] When any user is located at position A, the target position is the upper left position. That is, the second audio signal is the echo-cancelled audio signal acquired by the audio acquisition device located at the upper left position of the audio acquisition device corresponding to position A. In other words, the second audio signal is the echo-cancelled audio signal acquired by audio acquisition device A.

[0097] When any user is located at position B, the target position is the upper right position. That is, the second audio signal is the echo-cancelled audio signal acquired by the audio acquisition device located at the upper right position of the audio acquisition device corresponding to position B. In other words, the second audio signal is the echo-cancelled audio signal acquired by audio acquisition device C.

[0098] When any user is located at position C, the target position is the lower left position. That is, the second audio signal is the echo-cancelled audio signal acquired by the audio acquisition device located at the lower left position of the audio acquisition device corresponding to position C. In other words, the second audio signal is the echo-cancelled audio signal acquired by audio acquisition device G.

[0099] When any user is located at position D, the target position is the lower right position. That is, the second audio signal is the echo-cancelled audio signal acquired by the audio acquisition device located at the lower right position of the audio acquisition device corresponding to position D. In other words, the second audio signal is the echo-cancelled audio signal acquired by audio acquisition device I.

[0100] In one possible implementation, the audio acquisition devices corresponding to each location can form an audio acquisition device array for each location. This array can divide each location into a dominant frequency range and a consonant frequency range. The dominant and consonant frequency ranges at any location can be interchanged as needed. This application embodiment does not limit the positions of the dominant and consonant frequency ranges at each location. The topology of the audio acquisition device array can be triangular, quadrilateral, or rectangular; this application embodiment does not limit this. Audio acquisition devices at the intersection of frequency ranges are shared audio acquisition devices, which belong to multiple frequency ranges.

[0101] like Figure 5 This is a schematic diagram of the main and consonant regions at any position provided in an embodiment of this application. Wherein, 501 is the audio acquisition device array corresponding to position A, 502 is the main region of position A, and 503 is the consonant region of position A. Alternatively, 501 is the audio acquisition device array corresponding to position A, 502 is the consonant region of position A, and 503 is the main region of position A.

[0102] The tonic and consonant registers are different for each location on the target vehicle. For example... Figure 6 This is a schematic diagram of the tonic and consonant registers at various positions provided in an embodiment of this application. Figure 6 It can be seen that the upper left region of position A is the main sound region of position A, and the lower right region is the consonant region of position A; the upper right region of position B is the main sound region of position B, and the lower left region is the consonant region of position B; the lower right region of position C is the main sound region of position C, and the upper left region is the consonant region of position C; the lower right region of position D is the main sound region of position D, and the upper left region is the consonant region of position D. This distinguishes the main sound regions of each position, effectively differentiates the speech content of each position, and reduces cross-noise.

[0103] like Figure 7 This is another schematic diagram of the tonic and consonant registers at various positions provided in the embodiments of this application. Figure 7It can be seen that the upper left area of ​​position A is the consonant region of position A, and the lower right area is the tonic region of position A; the upper right area of ​​position B is the consonant region of position B, and the lower left area is the tonic region of position B; the lower left area of ​​position C is the consonant region of position C, and the upper right area is the tonic region of position C; the lower right area of ​​position D is the consonant region of position D, and the upper left area is the tonic region of position D. This way, when the external environment of the target vehicle is noisy, the audio acquisition device located at the edge of the vehicle body may receive more ambient noise, resulting in a noisy audio signal. By distributing the tonic regions of each position in the middle of the vehicle body, noise can be effectively reduced.

[0104] In one possible implementation, when there are many users inside the target vehicle (e.g., 5 users), each position in the target vehicle does not distinguish between the main vocal range and the consonant range.

[0105] In step 205, the target audio signal is removed from the reference audio signal to obtain the audio signal corresponding to the first user. The target audio signal is an audio signal containing the voiceprint feature information of the second user.

[0106] The second user is any user other than the first user, and the audio signal corresponding to the first user is used for the first user to interact with the target vehicle.

[0107] In one possible implementation, before removing the target audio signal from the reference audio signal, it is necessary to first determine the target audio signal from the reference audio signal. This application does not limit the process of determining the target audio signal from the reference audio signal. For example, the process of determining the target audio signal from the reference audio signal includes: segmenting the reference audio signal to obtain multiple candidate audio signals; determining the voiceprint feature information corresponding to each candidate audio signal; and determining the candidate audio signal whose corresponding voiceprint feature information is not that of the first user as the target audio signal.

[0108] In this context, all candidate audio signals have the same length.

[0109] After the target audio signal is determined, the target audio signal is removed from the reference audio signal, and the audio signal after removing the target audio signal is taken as the audio signal corresponding to the first user.

[0110] It should be noted that the process of determining the audio signal for each user other than the first user is similar to the process of determining the audio signal for the first user, and will not be repeated here in the embodiments of this application.

[0111] In one possible implementation, after removing the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user, a control command matching the audio signal corresponding to the first user can be determined, and the target vehicle can be controlled according to the control command.

[0112] Optionally, the process of determining the control instruction that matches the audio signal corresponding to the first user includes: determining the text content corresponding to the audio signal corresponding to the first user; and determining that the instruction that matches the text content is a control instruction.

[0113] In one possible implementation, control commands matching the audio signal corresponding to the first user can also be determined in other ways, which will not be described in detail here.

[0114] This application does not limit the process of controlling a target vehicle according to control instructions. Optionally, the process of controlling a target vehicle according to control instructions includes: sending control instructions to the controller of the target vehicle, so that the controller of the target vehicle controls the target vehicle according to the control instructions.

[0115] For example, a control command might be to open the driver's side window. After sending the control command to the controller of the target vehicle, the controller will open the driver's side window. Another example is a control command to turn on the air conditioning. After sending the control command to the controller of the target vehicle, the controller will turn on the air conditioning.

[0116] It should be noted that when the target vehicle is in motion and the doors are not open or closed, but the users inside the vehicle have changed seats, the voiceprint characteristics of each user previously recorded for each location will appear in the other location. Therefore, the interior of the target vehicle needs to be photographed again to re-determine the voiceprint characteristics of each user in each location. This process is similar to the process of determining the voiceprint characteristics of each user mentioned above, and will not be repeated here.

[0117] The method described above activates the audio acquisition devices corresponding to each user's location. This ensures that the user's audio signal can be collected regardless of their position within the target vehicle, allowing each user to interact with the vehicle and thus increasing the flexibility of vehicle interaction. Furthermore, activating only the audio acquisition device corresponding to the user's location saves resources. Since the audio acquisition device at a user's location may also collect audio signals from users in other locations, the audio signals collected by the user's location are processed to obtain the user's own audio signal. This removes interference from other users, allowing interaction with the target vehicle solely through the user's own audio signal, thereby improving the accuracy of vehicle interaction.

[0118] Figure 8The diagram shown is a structural schematic of an audio signal processing device provided in an embodiment of this application. Figure 8 As shown, the device includes:

[0119] The acquisition module 801 is used to acquire images of the interior of the target vehicle, which are used to determine the location of each user inside the target vehicle.

[0120] Control module 802 is used to control the activation of the audio acquisition device corresponding to the location of each user;

[0121] The acquisition module 801 is also used to acquire audio signals acquired by each audio acquisition device;

[0122] The processing module 803 is used to process the audio signal collected by the audio acquisition device at the location of the first user among all users to obtain a reference audio signal.

[0123] The acquisition module 801 is also used to remove the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user. The target audio signal is an audio signal containing the voiceprint feature information of the second user. The second user is a user other than the first user among all users. The audio signal corresponding to the first user is used for the first user to interact with the target vehicle.

[0124] In one possible implementation, the processing module 803 is further configured to perform echo cancellation processing on the audio signals acquired by each audio acquisition device to obtain the echo-cancelled audio signals acquired by each audio acquisition device.

[0125] The device also includes:

[0126] The determination module is used to determine the voiceprint feature information corresponding to each user based on the echo-cancelled audio signals collected by each audio acquisition device. The voiceprint feature information corresponding to any user is used to characterize the timbre of any user.

[0127] The processing module 803 is used to perform beamforming processing on the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain a reference audio signal.

[0128] In one possible implementation, a determining module is used to determine a first audio signal for any user among the users from the echo-cancelled audio signals collected by each audio acquisition device. The first audio signal is the echo-cancelled audio signal collected by the audio acquisition device corresponding to the location of any user.

[0129] A second audio signal is determined from the first audio signal. The second audio signal is the echo-cancelled audio signal collected by the audio acquisition device located at the target location in the audio acquisition device corresponding to the location of any user. The target location is determined based on the location of any user.

[0130] The second audio signal is processed to obtain the voiceprint feature information corresponding to any user.

[0131] In one possible implementation, a module is defined for inputting the second audio signal into the voiceprint feature extraction model;

[0132] The output of the voiceprint feature extraction model is determined to be the voiceprint feature information corresponding to any user.

[0133] In one possible implementation, the device further includes:

[0134] The determination module is used to segment the reference audio signal to obtain multiple candidate audio signals;

[0135] Determine the voiceprint feature information corresponding to each candidate audio signal;

[0136] The candidate audio signal whose corresponding voiceprint feature information is not the voiceprint feature information of the first user is identified as the target audio signal among multiple candidate audio signals.

[0137] In one possible implementation, the device further includes:

[0138] The determination module is used to determine the control command that matches the audio signal corresponding to the first user;

[0139] The control module 802 is also used to control the target vehicle according to control commands.

[0140] The aforementioned device activates the audio acquisition device corresponding to each user's location. This ensures that the user's audio signal can be collected regardless of their position within the target vehicle, enabling each user to interact with the vehicle and thus increasing the flexibility of vehicle interaction. Furthermore, activating only the audio acquisition device corresponding to the user's location conserves resources. Since the audio acquisition device at a user's location may also collect audio signals from users in other locations, the audio signal collected at the user's location is processed to obtain the user's own audio signal. This removes interference from other users, allowing interaction with the target vehicle solely through the user's own audio signal, thereby improving the accuracy of vehicle interaction.

[0141] It should be understood that the above-described apparatus is only illustrated by the division of the functional modules described above when implementing its functions. In practical applications, the functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0142] Figure 9 This is a schematic diagram of an audio signal processing system provided in an embodiment of this application, as shown below. Figure 9 As shown, the system includes an in-vehicle camera module 901, an audio acquisition device 902, and a computer device 903. Among them,

[0143] The in-vehicle camera module 901 is used to capture images of the interior of the target vehicle when the target vehicle is in the start-up state and the doors of the target vehicle are in the closed state, and send the images of the interior of the vehicle to the computer device 903.

[0144] Computer device 903 is used to determine the location of each user inside the target vehicle based on the image inside the vehicle, and send an activation command to the audio acquisition device 902 corresponding to the location of each user.

[0145] The audio acquisition device 902 is used to turn on the audio acquisition device 902, acquire audio signals, and send the acquired audio signals to the computer device 903.

[0146] The computer device 903 is also used to process the audio signal collected by the audio acquisition device corresponding to the location of the first user among the users to obtain a reference audio signal; remove the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user, wherein the target audio signal is an audio signal containing the voiceprint feature information of the second user, the second user being the user other than the first user among the users, and the audio signal corresponding to the first user is used for the first user to interact with the target vehicle.

[0147] In one possible implementation, the computer device 903 is further configured to perform echo cancellation processing on the audio signals acquired by each audio acquisition device to obtain the echo-cancelled audio signals acquired by each audio acquisition device; and determine the voiceprint feature information corresponding to each user based on the echo-cancelled audio signals acquired by each audio acquisition device, wherein the voiceprint feature information corresponding to any user is used to characterize the timbre of any user.

[0148] Computer device 903 is used to perform beamforming processing on the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain a reference audio signal.

[0149] In one possible implementation, computer device 903 is configured to, for any user among various users, determine a first audio signal from the echo-cancelled audio signals acquired by various audio acquisition devices, wherein the first audio signal is the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of any user; determine a second audio signal from the first audio signal, wherein the second audio signal is the echo-cancelled audio signal acquired by the audio acquisition device located at a target location among the audio acquisition devices corresponding to the location of any user, wherein the target location is determined based on the location of any user; and process the second audio signal to obtain voiceprint feature information corresponding to any user.

[0150] In one possible implementation, computer device 903 is used to input a second audio signal into a voiceprint feature extraction model; and to determine that the output of the voiceprint feature extraction model is the voiceprint feature information corresponding to any user.

[0151] In one possible implementation, the computer device 903 is further configured to segment the reference audio signal to obtain multiple candidate audio signals; determine the voiceprint feature information corresponding to each candidate audio signal; and determine the candidate audio signal whose corresponding voiceprint feature information is not the voiceprint feature information corresponding to the first user as the target audio signal.

[0152] In one possible implementation, the computer device 903 is further configured to determine a control command that matches the audio signal corresponding to the first user; and control the target vehicle according to the control command.

[0153] In one possible implementation, the system also includes a controller 904;

[0154] Computer device 903 is used to send control commands to controller 904;

[0155] Controller 904 is used to control the target vehicle according to control commands.

[0156] Figure 10 This illustration shows a structural block diagram of a terminal device 1000 provided in an exemplary embodiment of this application. The terminal device 1000 can be any electronic device product capable of human-computer interaction with a user through one or more methods such as a keyboard, touchpad, remote control, voice interaction, or handwriting device. Examples include PCs (Personal Computers), mobile phones, smartphones, PDAs (Personal Digital Assistants), wearable devices, PPCs (Pocket PCs), tablet computers, smart car systems, smart TVs, smart speakers, and smartwatches.

[0157] Typically, terminal device 1000 includes a processor 1001 and a memory 1002.

[0158] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0159] The memory 1002 may include one or more computer-readable storage media, which may be non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one instruction, which is executed by the processor 1001 to implement the audio signal processing method provided in the method embodiments of this application.

[0160] In some embodiments, the terminal device 1000 may also optionally include: a peripheral device interface 1003 and at least one peripheral device. The processor 1001, memory 1002, and peripheral device interface 1003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1003 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1004, a display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1008.

[0161] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0162] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other terminal devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0163] Display screen 1005 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1001 for processing. In this case, display screen 1005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 1005 may be a single screen, disposed on the front panel of terminal device 1000; in other embodiments, display screen 1005 may be at least two, disposed on different surfaces of terminal device 1000 or in a folded design; in still other embodiments, display screen 1005 may be a flexible display screen, disposed on a curved or folded surface of terminal device 1000. Furthermore, display screen 1005 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0164] The camera assembly 1006 is used to acquire images or videos. Optionally, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal device 1000, and the rear-facing camera is located on the back of the terminal device 1000. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0165] The audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1001 for processing, or input to the radio frequency circuit 1004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal device 1000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1001 or the radio frequency circuit 1004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1007 may also include a headphone jack.

[0166] The power supply 1008 is used to power the various components in the terminal device 1000. The power supply 1008 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1008 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, while a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0167] In some embodiments, the terminal device 1000 further includes one or more sensors 1009. The one or more sensors 1009 include, but are not limited to: an acceleration sensor 1010, a gyroscope sensor 1011, a pressure sensor 1012, an optical sensor 1013, and a proximity sensor 1014.

[0168] Accelerometer 1010 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal device 1000. For example, accelerometer 1010 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1001 can control display screen 1005 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1010. Accelerometer 1010 can also be used for games or for acquiring user motion data.

[0169] The gyroscope sensor 1011 can detect the orientation and rotation angle of the terminal device 1000. The gyroscope sensor 1011 can work in conjunction with the accelerometer sensor 1010 to collect 3D motion data from the user on the terminal device 1000. Based on the data collected by the gyroscope sensor 1011, the processor 1001 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0170] The pressure sensor 1012 can be disposed on the side bezel of the terminal device 1000 and / or on the lower layer of the display screen 1005. When the pressure sensor 1012 is disposed on the side bezel of the terminal device 1000, it can detect the user's grip signal on the terminal device 1000, and the processor 1001 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1012. When the pressure sensor 1012 is disposed on the lower layer of the display screen 1005, the processor 1001 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1005. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0171] An optical sensor 1013 is used to collect ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the display screen 1005 based on the ambient light intensity collected by the optical sensor 1013. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1005 is increased; when the ambient light intensity is low, the display brightness of the display screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera assembly 1006 based on the ambient light intensity collected by the optical sensor 1013.

[0172] The proximity sensor 1014, also known as a distance sensor, is typically installed on the front panel of the terminal device 1000. The proximity sensor 1014 is used to detect the distance between the user and the front of the terminal device 1000. In one embodiment, when the proximity sensor 1014 detects that the distance between the user and the front of the terminal device 1000 is gradually decreasing, the processor 1001 controls the display screen 1005 to switch from a screen-on state to a screen-off state; when the proximity sensor 1014 detects that the distance between the user and the front of the terminal device 1000 is gradually increasing, the processor 1001 controls the display screen 1005 to switch from a screen-off state to a screen-on state.

[0173] Those skilled in the art will understand that Figure 10 The structure shown does not constitute a limitation on the terminal device 1000, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0174] Figure 11This is a schematic diagram of the server structure provided in the embodiments of this application. The server 1100 can vary considerably due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1101 and one or more memories 1102. The one or more memories 1102 store at least one line of program code, which is loaded and executed by the one or more processors 1101 to implement the audio signal processing methods provided in the various method embodiments described above. Of course, the server 1100 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server 1100 may also include other components for implementing device functions, which will not be elaborated here.

[0175] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one piece of program code that is loaded and executed by a processor to enable a computer to implement any of the above-described audio signal processing methods.

[0176] Optionally, the aforementioned computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.

[0177] In an exemplary embodiment, a computer program or computer program product is also provided, which stores at least one computer instruction that is loaded and executed by a processor to enable the computer to implement any of the above-described audio signal processing methods.

[0178] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0179] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0180] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0181] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.

Claims

1. A method for processing audio signals, characterized in that, The method includes: Acquire interior images of the target vehicle, which are used to determine the location of each user inside the target vehicle. The audio acquisition device corresponding to the location of each user is activated. Acquire audio signals collected by each audio acquisition device; The audio signals acquired by each audio acquisition device are subjected to echo cancellation processing to obtain the echo-cancelled audio signals acquired by each audio acquisition device. For any one of the users, a first audio signal is determined from the echo-cancelled audio signals collected by each audio acquisition device. The first audio signal is the echo-cancelled audio signal collected by the audio acquisition device corresponding to the location of any user. A second audio signal is determined from the first audio signal. The second audio signal is the echo-cancelled audio signal collected by the audio acquisition device located at the target location in the audio acquisition device corresponding to the location of any user. The target location is determined based on the location of any user. The second audio signal is processed to obtain the voiceprint feature information corresponding to any user, and the voiceprint feature information corresponding to any user is used to characterize the timbre of any user. For the first user among the users, the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of the first user is subjected to beamforming processing to obtain a reference audio signal; The target audio signal is removed from the reference audio signal to obtain the audio signal corresponding to the first user. The target audio signal is an audio signal containing the voiceprint feature information of the second user. The second user is a user other than the first user among the users. The audio signal corresponding to the first user is used for the first user to interact with the target vehicle.

2. The method according to claim 1, characterized in that, The process of processing the second audio signal to obtain the voiceprint feature information corresponding to any user includes: The second audio signal is input into the voiceprint feature extraction model; The output of the voiceprint feature extraction model is determined to be the voiceprint feature information corresponding to any user.

3. The method according to claim 1, characterized in that, Before removing the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user, the method further includes: The reference audio signal is segmented to obtain multiple candidate audio signals; Determine the voiceprint feature information corresponding to each candidate audio signal; The candidate audio signal whose corresponding voiceprint feature information is not the voiceprint feature information corresponding to the first user is determined as the target audio signal.

4. The method according to any one of claims 1 to 3, characterized in that, After removing the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user, the method further includes: Determine the control command that matches the audio signal corresponding to the first user; The target vehicle is controlled according to the control command.

5. An audio signal processing device, characterized in that, The device includes: An acquisition module is used to acquire images of the interior of a target vehicle, which are used to determine the location of each user inside the target vehicle. The control module is used to control the activation of the audio acquisition device corresponding to the location of each user; The acquisition module is also used to acquire audio signals acquired by each audio acquisition device; The processing module is used to perform echo cancellation processing on the audio signals acquired by each audio acquisition device to obtain echo-cancelled audio signals acquired by each audio acquisition device; for any user among the users, a first audio signal is determined from the echo-cancelled audio signals acquired by each audio acquisition device, the first audio signal being the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of the user; a second audio signal is determined from the first audio signal, the second audio signal being the echo-cancelled audio signal acquired by the audio acquisition device located at the target location among the audio acquisition devices corresponding to the location of the user, the target location being determined based on the location of the user; the second audio signal is processed to obtain the voiceprint feature information corresponding to the user, the voiceprint feature information corresponding to the user being used to characterize the timbre of the user; for the first user among the users, beamforming processing is performed on the echo-cancelled audio signal acquired by the audio acquisition device corresponding to the location of the first user to obtain a reference audio signal. The acquisition module is further configured to remove the target audio signal from the reference audio signal to obtain the audio signal corresponding to the first user. The target audio signal is an audio signal containing the voiceprint feature information of the second user. The second user is a user other than the first user among the users. The audio signal corresponding to the first user is used for the first user to interact with the target vehicle.

6. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one piece of program code, the at least one piece of program code being loaded and executed by the processor to enable the computer device to implement the audio signal processing method as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one piece of program code, which is loaded and executed by a processor to enable the computer to implement the audio signal processing method as described in any one of claims 1 to 4.

8. A computer program product, characterized in that, The computer program product stores at least one computer instruction, which is loaded and executed by a processor to enable the computer to implement the audio signal processing method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Voice interaction method and device for vehicle-mounted system, automobile and machine readable medium

    CN110070868A

  • In-vehicle wakeup-free voice interaction method, device and equipment and storage medium

    CN116013298A