Method, apparatus and computer readable storage medium for processing audio signal

By acquiring mid-to-high frequency signals of audio signals, determining feature information, and performing upmixing, combined with sparse coding and neural network models, the flexibility and sound field equalization problems of traditional audio upmixing algorithms are solved, improving the spatial distribution of in-vehicle audio signals and the auditory experience.

CN119743703BActive Publication Date: 2026-03-27NIO TECH ANHUI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, traditional spatial audio upmixing algorithms lack flexibility when processing different types of audio sources, resulting in unbalanced sound field performance and inaccurate positioning.

Method used

By acquiring the mid-to-high frequency signals of the audio signal, determining its characteristic information, and performing upmixing based on sparse coding technology and neural network model, channel signals corresponding to each speaker are generated, and enhancement processing is performed by combining physical and virtual speaker parameters and in-vehicle acoustic characteristics.

Benefits of technology

It improves the flexibility and sound field balance when processing different types of audio sources, and enhances the spatial distribution uniformity of in-vehicle audio signals and the listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119743703B_ABST
    Figure CN119743703B_ABST
Patent Text Reader

Abstract

The application provides a processing method and device of an audio signal, equipment and a computer readable storage medium. The method comprises the following steps: obtaining an audio signal, extracting a medium-high frequency signal of the audio signal, wherein the frequency of the medium-high frequency signal is equal to a preset frequency; determining feature information of the medium-high frequency signal based on the medium-high frequency signal; obtaining a channel signal corresponding to each loudspeaker in a smart device; and outputting the corresponding channel signal through each loudspeaker. The method can perform upmix processing based on the feature information, improve the flexibility when processing different types of sound sources, and improve the balance of the sound field in the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of audio signal processing, and particularly relates to an audio signal processing method and device, equipment and a computer readable storage medium. BACKGROUND

[0002] With the rapid development of intelligent electric vehicles and the demand of consumers for high-quality music experience in the vehicle cabin, the number of loudspeakers equipped in the vehicle cabin of electric vehicles has significantly increased. In the vehicle audio system, the traditional spatial audio upmixing technology cannot fully meet the needs of modern multi-channel sound systems. The upmixing algorithm provided in the related art usually relies on fixed rules or preset parameters to expand two-channel or few-channel sound sources into a multi-channel system. This upmixing method often lacks flexibility when processing different types of sound sources, and can result in uneven sound field performance and inaccurate positioning. SUMMARY

[0003] To solve the above problems, the embodiments of the present application provide an audio signal processing method, device, equipment and computer readable storage medium to solve the problem of lack of flexibility in upmixing different types of sound sources and uneven sound field performance in the related art.

[0004] In a first aspect, the embodiments of the present application provide an audio signal processing method, which comprises:

[0005] obtaining an audio signal;

[0006] extracting a mid-high frequency signal of the audio signal, wherein the frequency of the mid-high frequency signal is greater than or equal to a preset frequency;

[0007] determining feature information of the mid-high frequency signal based on the mid-high frequency signal;

[0008] performing upmixing processing on the mid-high frequency signal based on the feature information to obtain a channel signal corresponding to each loudspeaker in a smart device, wherein each loudspeaker is used to output a mid-high frequency signal;

[0009] outputting the corresponding channel signal through each loudspeaker.

[0010] In some embodiments, the determining of the feature information of the mid-high frequency signal based on the mid-high frequency signal comprises:

[0011] converting the mid-high frequency signal to a frequency domain to obtain a frequency signal;

[0012] extracting first feature information of the frequency signal using one or more audio processing models, wherein different audio processing models are used to extract different features in the frequency signal;

[0013] determine, according to the frequency signal, feature information of a historical audio signal as second feature information of the mid-high frequency signal, wherein the frequency signal of the historical audio signal matches the frequency signal of the mid-high frequency signal;

[0014] use the first feature information and / or the second feature information as the feature information of the mid-high frequency signal.

[0015] In some embodiments, the up-mixing the mid-high frequency signal based on the feature information to obtain the channel signal corresponding to each speaker in the smart device comprises:

[0016] determining a feature matrix based on the feature information using a sparse coding technique;

[0017] up-mixing the mid-high frequency signal into the channel signal corresponding to the number of speakers based on the feature matrix to obtain the channel signal corresponding to each speaker.

[0018] In some embodiments, the method further comprises:

[0019] updating a knowledge base and / or the one or more audio processing models based on the feature information and the frequency information of the audio signal, wherein the knowledge base comprises feature information of one or more historical audio signals.

[0020] In some embodiments, after outputting the channel signal corresponding to each speaker, the method further comprises:

[0021] obtaining parameters of physical speakers, parameters of virtual speakers, an absorption field and a reflection field of sound in the smart device;

[0022] enhancing the channel signal corresponding to each speaker based on the parameters of the physical speakers, the parameters of the virtual speakers, and the absorption field and the reflection field of sound in the smart device;

[0023] the outputting the channel signal corresponding to each speaker comprises:

[0024] outputting the enhanced channel signal corresponding to each speaker.

[0025] In some embodiments, the enhancing the channel signal corresponding to each speaker based on the parameters of the physical speakers, the parameters of the virtual speakers, and the absorption field and the reflection field of sound in the smart device comprises:

[0026] input the parameters of the physical loudspeaker, the parameters of the virtual loudspeaker, the absorption field and the reflection field of the sound in the smart device, and the channel signals corresponding to each loudspeaker into a neural network model to obtain enhanced channel signals corresponding to each loudspeaker, wherein the optimization target of the neural network model is that the uniformity of the distribution of the audio signals in the space of the smart device when each loudspeaker plays the enhanced channel signals is greater than a uniformity threshold.

[0027] In some embodiments, the enhancement processing of the channel signals corresponding to each loudspeaker based on the parameters of the physical loudspeaker, the parameters of the virtual loudspeaker, the absorption field and the reflection field of the sound in the smart device includes:

[0028] establishing a sound propagation model based on the parameters of the physical loudspeaker, the parameters of the virtual loudspeaker, the absorption field and the reflection field of the sound in the smart device;

[0029] simulating the emission of each channel signal from each loudspeaker based on the sound propagation model to obtain a simulation result;

[0030] performing enhancement processing on the channel signals corresponding to each loudspeaker based on the simulation result, and the uniformity of the distribution of the audio signals in the space of the smart device when each loudspeaker plays the enhanced channel signals is greater than a uniformity threshold.

[0031] In a second aspect, an embodiment of the present application provides an audio signal processing apparatus, including:

[0032] a first obtaining module configured to obtain an audio signal;

[0033] an extracting module configured to extract a mid-high frequency signal of the audio signal, wherein the frequency of the mid-high frequency signal is greater than or equal to a preset frequency;

[0034] a determining module configured to determine feature information of the mid-high frequency signal based on the mid-high frequency signal;

[0035] an up-mixing module configured to perform up-mixing processing on the mid-high frequency signal based on the feature information to obtain channel signals corresponding to each loudspeaker in a smart device, wherein each loudspeaker is configured to output a mid-high frequency signal;

[0036] an outputting module configured to output the corresponding channel signals through each loudspeaker.

[0037] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method provided in the first aspect when executing the computer program.

[0038] In a fourth aspect, the embodiments of the present application provide an intelligent device, comprising the electronic device of the fourth aspect.

[0039] In a fifth aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the method of the first aspect.

[0040] In a sixth aspect, the embodiments of the present application provide a computer program product, which comprises a computer program. The computer program is executed by a processor to implement at least the method of any one of the first aspect.

[0041] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0042] The method for processing an audio signal provided by the embodiments of the present application comprises the following steps: acquiring an audio signal, extracting a mid-high frequency signal of the audio signal, wherein the frequency of the mid-high frequency signal is greater than or equal to a preset frequency; determining feature information of the mid-high frequency signal based on the mid-high frequency signal; performing upmix processing on the mid-high frequency signal based on the feature information to obtain a channel signal corresponding to each loudspeaker in the intelligent device; and outputting the corresponding channel signal through each loudspeaker. The upmix processing can be performed based on the feature information, the flexibility in processing different types of sound sources can be improved, and the uniformity of the sound field in the vehicle can be improved.

[0043] It can be understood that the beneficial effects of the second aspect to the sixth aspect can be referred to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0045] Figure 1 A flowchart of the method for processing an audio signal provided by the embodiments of the present application is shown in the figure.

[0046] Figure 2 An implementation flowchart of the method for processing an audio signal provided by the embodiments of the present application is shown in the figure.

[0047] Figure 3 A processing flowchart of the adaptive upmix module provided by the embodiments of the present application is shown in the figure.

[0048] Figure 4A processing flow diagram of a swarm intelligence module is provided for an embodiment of the present application.

[0049] Figure 5 A structural diagram of an audio signal processing device is provided for an embodiment of the present application.

[0050] Figure 6 A structural diagram of an electronic device is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0051] In the following description, for purposes of explanation and not limitation, specific details are set forth, such as particular sequences of steps, techniques, etc., in order to provide a thorough understanding of the present embodiments. However, it will be apparent to one skilled in the art that the present embodiments can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, and circuits are omitted so as not to obscure the description of the present embodiments.

[0052] It is to be understood that the terminology "includes", "has", "holds", "contains" and / or "comprising", when used in this specification and in the following claims, indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0053] It is also to be understood that the terminology "and / or" when used in this specification and in the following claims, refers to at least one of the items, or any combination of the items, listed after the term in the various aspects.

[0054] As used in this specification and in the claims, the term "if" can be interpreted as meaning "when", or "once", or "in response to determining", or "in response to detecting". Similarly, the phrase "if determined", or "if detected" can be interpreted as meaning "once determined", or "in response to determining", or "once detected", or "in response to detecting".

[0055] In addition, in the description of the present application and in the following claims, the terms "first", "second", "third", etc. are only used to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0056] Reference within the specification of this application to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places within specified

[0057] Based on the technical problem of the related art, the embodiments of the present application provide an audio signal processing method which can be applied to a smart device. For example, the smart device can be a driving device, a vehicle, a robot, etc. The vehicle can include a smart electric vehicle, and a fuel vehicle or a hybrid vehicle. The vehicle includes but is not limited to a car, a truck, a bus, an electric vehicle, a motorcycle, a house car, a train, etc. The vehicle can be a vehicle driven by a person. In some other embodiments, the vehicle can also be a vehicle with certain automatic driving capability. Specifically, the audio signal processing method can be applied to a processing component in the smart device.

[0058] For example, taking the smart device as a smart electric vehicle, the audio signal processing method can be applied to a vehicle-mounted device in the smart electric vehicle. The vehicle-mounted device can include a controller in the electric vehicle.

[0059] Of course, the audio signal processing method provided by the embodiments of the present application can also be executed by an electronic device in communication with the smart device. For example, the electronic device can include a mobile phone, a tablet computer, a wearable device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), and the embodiments of the present application do not make any limitation on the specific type of the electronic device. Of course, the electronic device can also be a smart device.

[0060] The smart device in the embodiments of the present application has one or more loudspeakers. Each loudspeaker can be distributed in different positions of the smart device. Optionally, the smart device can also have an audio playing control device for playing audio signals. Taking a vehicle as an example, each loudspeaker in the vehicle can be distributed in positions such as vehicle doors, a central position of the vehicle, a position of the instrument panel, and both sides of the second-row seats. Different vehicle models and audio system configurations can also result in differences in the number and types of loudspeakers. For example, a common vehicle model can be equipped with only four loudspeakers, while a high-end vehicle model can have eight or even more loudspeakers.

[0061] The audio signals can be, for example, vehicle navigation audio, media audio, voice audio, call ring tones, prompt tones, and warning tones. The media audio can include, for example, audio from a music player, video soundtracks, or game music, etc. The voice audio can include, for example, call voice or voice-aided voice, such as voice issued by a user interface. The call ring tones can include, for example, dial-up call ring tones or call ring tones from network voice or video calls, etc. The prompt tones can include, for example, prompt tones for requesting to join or end a network voice or video call, prompt tones for instant messaging or non-instant messaging, prompt tones for events that are intended to attract the attention of a user, and prompt tones associated with a call, etc.

[0062] For example, the vehicle can play the audio signals through any one or more of a vehicle-mounted CD player, a radio, a Bluetooth-connected device, etc.

[0063] The embodiments of the present application provide a method for processing audio signals, Figure 1 A flowchart of a method for processing audio signals provided by the embodiments of the present application is shown in Figure 1 The method comprises the following steps:

[0064] In step S101, an audio signal is acquired, and a mid-high frequency signal of the audio signal is extracted, wherein the frequency of the mid-high frequency signal is greater than or equal to a preset frequency.

[0065] In the embodiments of the present application, the audio signals can be vehicle navigation audio, media audio, voice audio, call ring tones, prompt tones, and warning tones, etc. The media audio can be music audio signals, and the music audio signals can be audio signals in different channel formats, such as stereo (2.0 channels), 5.1 channels, 7.1 channels, 7.1.4 channels, etc. The audio signals contain frequency and amplitude information.

[0066] In the embodiments of the present application, the audio signal can be divided into low frequency, medium frequency and high frequency parts. The medium and high frequency signal generally refers to the part with higher frequency, which contains many details and clarity elements in the sound, such as the high pitch part of the human voice, the bright timbre of the musical instrument, etc. The preset frequency can be configured, for example, it can be configured as 200Hz.

[0067] In the embodiments of the present application, in the car audio system, the audio signal can come from the on-board CD player, radio, Bluetooth connected device, etc.

[0068] In the embodiments of the present application, the medium and high frequency signal of the audio signal can be extracted by digital signal processing (DSP) technology. The DSP algorithm can analyze the frequency components of the audio signal and extract the part with frequency greater than or equal to the preset frequency. In some embodiments, a high-pass filter with a cutoff frequency of the preset frequency can be used to separate the medium and high frequency signal of the audio signal.

[0069] In the embodiments of the present application, the medium and high frequency signal of the audio signal is extracted to adapt to the subsequent upmix processing, ensuring that the high frequency details and clarity will not be lost in the complex sound field.

[0070] Step S102, determining the feature information of the medium and high frequency signal based on the medium and high frequency signal.

[0071] In the embodiments of the present application, the feature information can include frequency characteristics, energy distribution, phase relationship, style type, dynamic range and detail level, etc.

[0072] In the embodiments of the present application, the medium and high frequency signal can be analyzed to extract the feature information. The feature information can include multiple features, and the feature information can include overall features, local features, feature distribution, etc. The feature information can be obtained by an audio processing model, or can be obtained by a knowledge base.

[0073] In some embodiments, the feature information includes first feature information and / or second feature information, wherein the first feature information is extracted from the frequency signal of the medium and high frequency signal based on one or more audio processing models, and the second feature information is the feature information of the historical audio signal matched with the frequency signal of the medium and high frequency signal.

[0074] Step S103, upmix processing the medium and high frequency signal based on the feature information to obtain the channel signal corresponding to each loudspeaker in the intelligent device, wherein each loudspeaker is used to output the medium and high frequency signal.

[0075] In the embodiments of the present application, the upmix processing is used to convert a mono or multi-channel signal into a signal of more channels (such as stereo, surround sound, etc.). In a car audio, the upmix processing can distribute the audio signal to different speakers to create a richer auditory experience.

[0076] In the embodiments of the present application, in a car, there are usually multiple speakers located at different positions (such as front door, rear door, roof, etc.). For example, there are N speakers in the car, and then N channel signals can be obtained after upmix processing, each of which can correspond to a speaker. The channel signal refers to the signal distributed to each speaker.

[0077] In the embodiments of the present application, the mid-high frequency signal can be converted into a signal of multiple channels based on the feature information by using the upmix algorithm, and distributed to each channel. After the signal is distributed, the channel signal can be delayed, phase adjusted, gain controlled, etc. to ensure that the sound output by each speaker is coordinated in time and space.

[0078] In step S104, the corresponding channel signal is output through each of the speakers.

[0079] In the embodiments of the present application, the channel signal processed by the upmix processing is sent to each speaker in the car. The speaker converts these electrical signals into sound, thereby creating the desired sound effect.

[0080] The processing method of the audio signal provided by the embodiments of the present application can obtain an audio signal, extract a mid-high frequency signal of the audio signal, wherein the frequency of the mid-high frequency signal is greater than or equal to a preset frequency; determine feature information based on the mid-high frequency signal; perform upmix processing on the mid-high frequency signal based on the feature information to obtain a channel signal corresponding to each speaker in the smart device; and output the corresponding channel signal through each of the speakers. The upmix processing can be performed based on the feature information, the flexibility in processing different types of sound sources can be improved, and the balance of the sound field in the smart device can be improved.

[0081] In some embodiments, step S102 can be implemented by the following steps:

[0082] In step S1021, the mid-high frequency signal is converted into a frequency signal in the frequency domain.

[0083] In the embodiments of the present application, the frequency signal describes the distribution of the signal in the frequency domain. The frequency signal can be obtained by performing Fourier transform on the time domain signal. The frequency signal can reveal the characteristics of the signal in the frequency domain, such as frequency spectrum distribution, frequency component, etc.

[0084] In step S1022, one or more audio processing models are used to extract first feature information of the frequency signal, and different audio processing models are used to extract different features in the frequency signal.

[0085] In an embodiment of the present application, one or more audio processing models can be stored in a model library. These audio processing models can be machine learning-based, deep learning-based, or other types of algorithmic models. In audio processing, each audio processing model can be used to extract different features. The first feature information can include overall features and local features of the frequency signal, for example, the overall features can include main frequency components, energy distribution, spectral shape, etc., and the local features can include frequency components in a local range and spectral changes.

[0086] In step S1023, the feature information of the historical audio signal is determined as second feature information of the mid-high frequency signal according to the frequency signal, wherein the frequency signal of the historical audio signal matches the frequency signal of the mid-high frequency signal.

[0087] In an embodiment of the present application, the feature information of the historical audio signal can be stored in a knowledge base. The feature information of the historical audio signal can be feature distribution and pattern.

[0088] In an embodiment of the present application, the frequency signal can be matched with the frequency signal of the historical audio signal, so as to extract the feature information related to the frequency signal. When matching, the similarity between the two frequency signals can be calculated, and the similarity is used to determine whether the two frequency signals match. By setting a similarity threshold, when the similarity is greater than the similarity threshold, it is determined that the two frequency signals match.

[0089] The method provided in the embodiment of the present application extracts the first feature information of the frequency signal through the audio processing model, and determines the feature information of the historical audio signal as the second feature information of the mid-high frequency signal according to the frequency signal, which can generate more rich multi-dimensional sampling features and ensure comprehensive understanding of the audio signal.

[0090] In some embodiments, step S103 can be implemented by the following steps: determining a feature matrix based on the feature information using a sparse coding technique; and upmixing the mid-high frequency signal into a corresponding number of channel signals of the loudspeakers based on the feature matrix to obtain the channel signals corresponding to each loudspeaker in the smart device.

[0091] In the embodiments of the present application, sparse coding is a signal processing technique, and sparse coding technology represents a signal as a linear combination of a few basis functions, which are usually predefined or obtained through learning. In sparse coding, the representation of a signal is sparse, that is, most of the coefficients are zero or close to zero, and only a few coefficients are non-zero. Such sparsity helps to extract important features in the signal while reducing redundant information. Each row or column of the feature matrix represents a feature of the signal. Through the sparse coding technology, the features can be compressed and expressed in a more flexible and accurate manner, retaining the uniqueness and discriminability of each feature, thereby improving the accuracy of subsequent signal allocation.

[0092] In the embodiments of the present application, a mapping strategy can be formulated according to the feature matrix using a deep learning algorithm to map the input audio signal to different output channels, thereby realizing upmixing of the mid-high frequency signal to the corresponding number of channel signals of the loudspeaker. The mapping strategy can be based on various factors, such as the layout of the loudspeaker, the characteristics of the in-vehicle sound field, etc. When formulating the mapping strategy, the deep learning algorithm can require that after being mapped to each channel, a consistent and balanced auditory experience is obtained at each position in the vehicle, and the spatial sense, positioning accuracy, and level of the sound field need to be guaranteed to meet the high-quality spatial audio presentation of the vehicle audio system in complex environments.

[0093] In some embodiments, after step S102, the method further comprises:

[0094] Step S105, updating the knowledge base and / or the one or more audio processing models based on the feature information and the frequency information of the audio signal, the knowledge base including feature information of one or more historical audio signals.

[0095] In the embodiments of the present application, the feature information and the frequency information of the audio signal can be taken as a new sample, and the similarity between the new sample and the existing samples in the knowledge base can be compared. The similarity measurement methods include Euclidean distance, cosine similarity, Manhattan distance, Jaccard similarity coefficient, etc. A similarity threshold can be set, and if the similarity between the new sample and the existing samples in the knowledge base is lower than the threshold, the new sample is added to the knowledge base. If the similarity between the new sample and the existing samples in the knowledge base is greater than the similarity threshold, the knowledge base can not be updated.

[0096] In the embodiments of the present application, when there is a new sample, the entire audio processing model does not need to be retrained, but an incremental learning algorithm can be used to update the audio processing model, and the parameters of the audio processing model with poor effect can be updated.

[0097] The method provided in the embodiments of the present application updates parameters through a poor effect model, thereby maintaining the dynamic adaptability of the audio processing model, and by updating the knowledge base and the audio processing model, more comprehensive feature information can be extracted.

[0098] In some embodiments, before step S104, the method further comprises:

[0099] Step S1041, parameters of a physical loudspeaker, parameters of a virtual loudspeaker, an absorption field and a reflection field of sound in the smart device are acquired.

[0100] In the embodiments of the present application, the physical loudspeaker is installed inside the smart device and is used to actually produce sound. The parameters of the physical loudspeaker can include the frequency response, sensitivity, impedance, power handling capability, etc. of the loudspeaker. The parameters of the physical loudspeaker determine the sound output capability and sound quality of the loudspeaker at different frequencies. The virtual loudspeaker is a loudspeaker position simulated by digital signal processing (DSP) technology and is usually used to create a surround sound effect or optimize the spatial distribution of sound. The parameters of the virtual loudspeaker can include the position, direction and relative relationship with other loudspeakers, etc. of the virtual loudspeaker. The absorption field of sound is used to represent the area information that has an absorbing effect on sound. The absorption field affects the reflection and propagation of sound, thereby affecting the balance and clarity of sound in the smart device. The reflection field is the information of the area or object that reflects sound. The reflection field determines how sound propagates and distributes in the space inside the smart device and can strengthen or weaken the sound of certain frequencies.

[0101] In the embodiments of the present application, the parameters of the physical loudspeaker can be directly acquired from each loudspeaker, the parameters of the virtual loudspeaker can be directly acquired from the electronic device, and the absorption field and the reflection field of sound are acquired by acoustic measurement to obtain the absorption and reflection characteristics of sound at different positions in the smart device, thereby obtaining the absorption field and the reflection field of sound.

[0102] Step S1042, each channel signal corresponding to each loudspeaker is enhanced based on the parameters of the physical loudspeaker, the parameters of the virtual loudspeaker, the absorption field and the reflection field of sound in the smart device.

[0103] In the embodiments of the present application, the enhancement processing of the channel signal can include phase adjustment, amplitude-frequency adjustment, etc. In some embodiments, each channel signal can also be filtered, gain-adjusted, delay-compensated, etc.

[0104] In the embodiments of the present application, the parameters of the physical speaker, the parameters of the virtual speaker, the absorption field, the reflection field, and the channel signals corresponding to each speaker can be input into a neural network model to obtain enhanced channel signals corresponding to each speaker, wherein the optimization target of the neural network model is that the uniformity of the distribution of the audio signals in the space of the smart device when each speaker plays the enhanced channel signals is greater than a uniformity threshold.

[0105] In the embodiments of the present application, the uniformity threshold can be determined according to the size of the space in the smart device, the speaker configuration, and the desired sound effect, and the neural network model can include an input layer, a hidden layer, and an output layer. The input layer is used to input the parameters of the physical speaker, the parameters of the virtual speaker, the absorption field, the reflection field, and the channel signals corresponding to each speaker, and the output layer is used to output the enhanced channel signals.

[0106] In the embodiments of the present application, the neural network model can be trained by sample data, and after training is completed, it can be deployed in an audio system to output enhanced channel signals based on the parameters of the physical speaker, the parameters of the virtual speaker, the absorption field, the reflection field, and the channel signals corresponding to each speaker.

[0107] In some embodiments, a sound propagation model can be established based on the parameters of the physical speaker, the parameters of the virtual speaker, the absorption field of the sound in the smart device, and the reflection field. Based on the sound propagation model, the emission of each channel signal from each speaker is simulated to obtain a simulation result. Based on the simulation result, the channel signals corresponding to each speaker are enhanced. When each speaker plays the enhanced channel signals, the uniformity of the distribution of the audio signals in the space of the smart device is greater than a uniformity threshold.

[0108] In the embodiments of the present application, the sound propagation process can be simulated by computer, considering the absorption field and the reflection field of the sound in the smart device, as well as the parameters and positions of the speakers. Through numerical simulation, a more accurate sound propagation model can be obtained. Then, based on the sound propagation model, the propagation process of the sound in the space of the smart device after being emitted from each speaker can be simulated. According to the simulation result, the sound intensity, phase, and other information at different positions in the smart device are calculated. According to the optimization target and the simulation result, the channel signals of each speaker are adjusted to improve the spatial distribution uniformity of the sound. Through multiple iterations of simulation and adjustment, the channel signals are gradually optimized until the set uniformity threshold is reached. In each iteration, the spatial distribution of the sound needs to be recalculated, and the channel signals are adjusted according to the result.

[0109] The method provided by the embodiment of the application can improve the audio performance of an audio signal in a space of a smart device, enhance the immersion and presence of music, and automatically adapt to different sound sources to provide richer and more delicate spatial sound performance in a vehicle, thereby greatly improving the immersion and level of auditory experience.

[0110] Based on the foregoing various embodiments, the embodiment of the application further provides an audio signal processing method, Figure 2 The implementation flowchart of the audio signal processing method provided by the embodiment of the application is shown in FIG. 1. Figure 2 As shown in FIG. 1, the audio source input in different channel formats is processed by a multi-band model to extract a mid-high frequency signal, is processed by an adaptive upmix module to obtain N mid-high sound channels, and is enhanced by a sound enhancement module based on dynamic parameters and static parameters to output N mid-high pressure channel signals.

[0111] In the embodiment of the application, a cutoff frequency can be set to separate the mid-high frequency component by a high-pass filter, for example, a cutoff frequency of 200 Hz is set to separate the mid-high frequency component by a 200-Hz high-pass filter.

[0112] In the embodiment of the application, the dynamic parameters are used to model the absorption field and the reflection field in the vehicle in real time, the environmental changes can be monitored and the sound wave propagation characteristics can be adjusted to solve the sound wave reflection and standing wave problems, and the static parameters can include physical loudspeaker parameters and virtual loudspeaker parameters. By combining the two types of parameters, the sound field enhancement module can achieve more rich sound performance in the limited cabin space, significantly enhance the auditory localization and immersion, and optimize the overall audio experience.

[0113] In the embodiment of the application, Figure 3 The processing flowchart of the adaptive upmix module provided by the embodiment of the application is shown in FIG. 2. Figure 3 As shown in FIG. 2, the I-path mid-high frequency signal can be changed in the time-frequency domain to convert the audio signal into a frequency representation, which is convenient for subsequent feature analysis. The feature information is extracted by a swarm intelligence module, the feature matrix is obtained by using a sparse coding technology for the feature information, the dynamic signal distribution is performed based on the feature matrix, and the signals of the N mid-high sound channels are output.

[0114] In the embodiment of the application, the sparse coding technology can compress and express the features in a more flexible and accurate manner, fully exert the advantages of multi-model collaborative learning and cross-sample learning, retain the uniqueness and discriminability of each feature, and thereby improve the accuracy of subsequent dynamic signal distribution.

[0115] In the embodiment of the present application, the dynamic signal distribution module intelligently remaps the signals to the N loudspeakers according to the key information in the feature matrix, ensures consistent and balanced auditory experience at each position in the vehicle, especially guarantees the spatial sense, positioning accuracy and level of the sound field, and meets the high-quality spatial audio presentation of the vehicle audio system in a complex environment.

[0116] In the embodiment of the present application, Figure 4 The processing flow diagram of the swarm intelligence module provided in the embodiment of the present application is shown in Figure 4 The swarm intelligence module consists of two key parts: one is collaborative learning relying on the model library, which deeply mines the overall and local features of the audio signal through multiple models with different characteristics and focuses; the other is cross-sample learning relying on the knowledge base, which realizes efficient analysis of new audio by referring to the feature distribution and pattern in historical audio data. Through the combination of the two, the system can generate more rich multi-dimensional abstract features to ensure comprehensive understanding of the audio signal. According to the final multi-dimensional abstract features, the knowledge base is updated using similarity matching, efficiently managing the update of the knowledge base and ensuring the efficiency of cross-sample learning. The model library updates the parameters of the models with poor performance through self-supervised learning mechanism to maintain the dynamic adaptability of the model library.

[0117] The method provided in the embodiment of the present application can dynamically and real-timely perform spatial audio upmixing on different songs through adaptive upmixing processing; combines virtual loudspeaker parameters with physical loudspeaker parameters, and models the reflection field and absorption field through the in-vehicle acoustic link to enhance the processing of the channels, which can enhance the spatial sound performance; can support spatial audio upmixing of different channel formats, and is more flexible in application.

[0118] It should be understood that the size of the serial number of each step in the above embodiment does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0119] According to the foregoing embodiments, the present application provides a processing device for an audio signal, each module included in the device and each unit included in the modules can be realized by a processor in a computer device; of course, it can also be realized by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA).

[0120] This application provides an audio signal processing apparatus. Figure 5 This is a schematic diagram of the structure of an audio signal processing device provided in an embodiment of this application, as shown below. Figure 5 As shown, the audio signal processing 500 includes:

[0121] The first acquisition module 501 is used to acquire audio signals;

[0122] Extraction module 502 is used to extract the mid-to-high frequency signal of the audio signal, wherein the frequency of the mid-to-high frequency signal is greater than or equal to a preset frequency;

[0123] The determining module 503 is used to determine the characteristic information of the medium- and high-frequency signal based on the medium- and high-frequency signal;

[0124] The upmixing module 504 is used to upmix the mid-to-high frequency signal based on the feature information to obtain the channel signal corresponding to each speaker in the smart device, wherein each speaker is used to output the mid-to-high frequency signal.

[0125] The output module 505 is used to output corresponding channel signals through each of the speakers.

[0126] In some embodiments, the determining module includes:

[0127] A conversion unit is used to convert the mid-to-high frequency signal to the frequency domain to obtain a frequency signal;

[0128] The first extraction unit is used to extract first feature information of the frequency signal using one or more audio processing models, wherein different audio processing models are used to extract different features in the frequency signal;

[0129] The second extraction unit is used to determine the feature information of the historical audio signal as the second feature information of the mid-to-high frequency signal based on the frequency signal, wherein the frequency signal of the historical audio signal matches the frequency signal of the mid-to-high frequency signal;

[0130] A determining unit is configured to use the first feature information and / or the second feature information as feature information of the mid-to-high frequency signal.

[0131] In some embodiments, the upmixing module includes:

[0132] A sparse coding unit is used to determine a feature matrix based on the feature information using sparse coding techniques.

[0133] The upmixing unit is used to upmix the mid-to-high frequency signal into the corresponding number of channel signals for each loudspeaker based on the feature matrix, so as to obtain the channel signal corresponding to each loudspeaker.

[0134] In some embodiments, the processing 500 of the audio signal further includes:

[0135] an updating module configured to update a knowledge base and / or the one or more audio processing models based on the feature information and frequency information of the audio signal, the knowledge base including feature information of one or more historical audio signals.

[0136] In some embodiments, the processing 500 of the audio signal further includes:

[0137] a second obtaining module configured to obtain parameters of a physical loudspeaker, parameters of a virtual loudspeaker, an absorption field and a reflection field of sound in the smart device;

[0138] an enhancing module configured to perform enhancement processing on the channel signal corresponding to each loudspeaker based on the parameters of the physical loudspeaker, the parameters of the virtual loudspeaker, the absorption field and the reflection field of sound in the smart device;

[0139] the outputting of the channel signal corresponding to each loudspeaker includes:

[0140] outputting the enhanced channel signal corresponding to each loudspeaker.

[0141] In some embodiments, the enhancing module includes:

[0142] a first enhancing unit configured to input the parameters of the physical loudspeaker, the parameters of the virtual loudspeaker, the absorption field, the reflection field and the channel signal corresponding to each loudspeaker into a neural network model to obtain the enhanced channel signal corresponding to each loudspeaker, wherein an optimization objective of the neural network model is that the uniformity of the distribution of the audio signal in the space of the smart device when each loudspeaker plays the enhanced channel signal is greater than a uniformity threshold.

[0143] In some embodiments, the enhancing module includes:

[0144] a establishing unit configured to establish a sound propagation model based on the parameters of the physical loudspeaker and the parameters of the virtual loudspeaker, the absorption field and the reflection field of sound in the smart device;

[0145] a simulating unit configured to simulate the sound propagation model based on the parameters of the physical loudspeaker, the parameters of the virtual loudspeaker, the absorption field and the reflection field of sound in the smart device;

[0146] a second enhancing unit configured to perform enhancement processing on the channel signal corresponding to each loudspeaker based on the simulation result, and the uniformity of the distribution of the audio signal in the space of the smart device when each loudspeaker plays the enhanced channel signal is greater than a uniformity threshold.

[0147] in addition, Figure 5 The processing of the audio signal shown can be a software unit, a hardware unit, or a combination of software and hardware built into existing electronic devices. It can also be integrated into electronic devices as a separate component or exist as a standalone terminal device.

[0148] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0149] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0150] This application provides an intelligent device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the audio signal processing method described in the above embodiments.

[0151] In this embodiment, the smart device can be a car.

[0152] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device 3 in this embodiment may include: at least one processor 30 ( Figure 6 Only one processor 30, memory 31, and computer program 32 stored in memory 31 and executable on at least one processor 30 are shown. When the processor 30 executes the computer program 32, it implements the steps in any of the above method embodiments, or the processor 30 executes the computer program 32 to implement the functions of each module / unit in the above device embodiments.

[0153] For example, the computer program 32 can be divided into one or more modules / units, one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program 32 instruction segments capable of completing a specific function, which is used to describe the execution process of the computer program 32 in the electronic device 3.

[0154] The embodiment of the present application further provides a computer readable storage medium, the computer readable storage medium stores the computer program 32, and the computer program 32 is executed by the processor 30 to realize the steps in each method embodiment.

[0155] The embodiment of the present application provides a computer program product, when the computer program product runs on the electronic device, so that the electronic device executes the steps in each method embodiment.

[0156] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. According to such understanding, the present application can realize all or part of the processes in the above-mentioned embodiment methods, which can be completed by instructing related hardware through the computer program 32. The computer program 32 can be stored in a computer readable storage medium, and the computer program 32 can realize the steps in each method embodiment when executed by the processor 30. The computer program 32 includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form. The computer readable medium at least includes any entity or device capable of carrying the computer program code to the terminal, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0157] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0158] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0159] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / network device and method can be implemented by other ways. For example, the apparatus / network device embodiments described above are only schematic, for example, the division of the modules or units is only a logical function division, and there can be another division in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the display or discussion of the mutual coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0160] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0161] The above-described embodiments are only used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalent; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method of processing an audio signal, characterized by, The method comprises: obtaining an audio signal; extracting a mid-high frequency signal of the audio signal, wherein the frequency of the mid-high frequency signal is greater than or equal to a preset frequency; determining feature information of the mid-high frequency signal based on the mid-high frequency signal, wherein the determining the feature information of the mid-high frequency signal based on the mid-high frequency signal comprises: converting the mid-high frequency signal to a frequency domain to obtain a frequency signal; extracting first feature information of the frequency signal by using one or more audio processing models, wherein different audio processing models are used to extract different features in the frequency signal; determining second feature information of the mid-high frequency signal according to the frequency signal, wherein the frequency signal of a historical audio signal matches the frequency signal of the mid-high frequency signal; and taking the first feature information and / or the second feature information as the feature information of the mid-high frequency signal; performing upmix processing on the mid-high frequency signal based on the feature information to obtain channel signals corresponding to each loudspeaker in a smart device; outputting the corresponding channel signals through each loudspeaker.

2. The method of claim 1, wherein, The method further comprises: updating a knowledge base and / or the one or more audio processing models based on the feature information and frequency information of the audio signal, wherein the knowledge base comprises feature information of one or more historical audio signals. After outputting the channel signals corresponding to each loudspeaker, the method further comprises:

3. The method of claim 1, wherein, obtaining parameters of physical loudspeakers, parameters of virtual loudspeakers, an absorption field and a reflection field of sound in the smart device; performing enhancement processing on the channel signals corresponding to each loudspeaker based on the parameters of the physical loudspeakers, the parameters of the virtual loudspeakers, the absorption field and the reflection field of sound in the smart device; 4. The method of claim 1, wherein, The method further comprises: outputting the enhanced channel signals corresponding to each loudspeaker. The method further comprises: inputting the parameters of the physical loudspeakers, the parameters of the virtual loudspeakers, the absorption field, the reflection field and the channel signals corresponding to each loudspeaker into a neural network model to obtain enhanced channel signals corresponding to each loudspeaker, wherein the optimization target of the neural network model is to make the uniformity of the distribution of the audio signal in the space of the smart device greater than a uniformity threshold when each loudspeaker plays the enhanced channel signals. The method further comprises:

5. The method of claim 4, wherein, ​ ​ 6. The method of claim 4, wherein, ​ establish a sound propagation model based on the parameters of the physical loudspeaker, the parameters of the virtual loudspeaker, the absorption field and the reflection field of the sound in the smart device; simulate emission of each channel signal from each loudspeaker based on the sound propagation model to obtain a simulation result; perform enhancement processing on the channel signal corresponding to each loudspeaker based on the simulation result, and when each loudspeaker plays the enhanced channel signal, the uniformity of the distribution of the audio signal in the space of the smart device is greater than a uniformity threshold.

7. An apparatus for processing an audio signal, characterized by comprising: a first acquisition module configured to acquire an audio signal; an extraction module configured to extract a mid-high frequency signal of the audio signal, wherein the frequency of the mid-high frequency signal is greater than or equal to a preset frequency; a determination module configured to determine feature information of the mid-high frequency signal based on the mid-high frequency signal, the determination of the feature information of the mid-high frequency signal based on the mid-high frequency signal comprising: converting the mid-high frequency signal to a frequency domain to obtain a frequency signal; extracting first feature information of the frequency signal using one or more audio processing models, different audio processing models being used to extract different features in the frequency signal; determining feature information of a historical audio signal as second feature information of the mid-high frequency signal according to the frequency signal, wherein the frequency signal of the historical audio signal matches the frequency signal of the mid-high frequency signal; and taking the first feature information and / or the second feature information as the feature information of the mid-high frequency signal; an upmixing module configured to perform upmixing processing on the mid-high frequency signal based on the feature information to obtain channel signals corresponding to each loudspeaker in the smart device; an output module configured to output the corresponding channel signal through each loudspeaker.

8. An electronic device, comprising: comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method of any one of claims 1 to 6 when executing the computer program.

9. A smart device, comprising: comprising: the electronic device of claim 8.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. the computer program is executed by the processor to implement the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Audio processing method, server, terminal equipment and storage medium

    CN118675496A

  • Display device and audio signal processing method thereof

    WO2024147370A1