Vehicle-mounted sound effect adaptation method and electronic equipment

By separating and identifying emotions in the car audio and calculating the output proportional coefficient of the audio clips, the problem that ordinary users find it difficult to adjust the in-car sound effects is solved, and the sound effects and riding experience are improved.

CN119148964BActive Publication Date: 2025-10-03ZHEJIANG LEAPMOTOR TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410955198.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-16
Publication Date
2025-10-03
Estimated Expiration
2044-07-16

AI Technical Summary

Technical Problem

It is difficult for ordinary users to adjust the in-car sound effects to suit their own preferences within the existing personalized settings, which affects the driving and riding experience.

Method used

By performing audio separation on the audio to be played, the emotions of the driver and passenger are identified, and the output scale coefficient of the audio clip is calculated based on the emotion recognition results to adjust the audio playback effect.

Benefits of technology

It can automatically adjust the audio playback effect according to the emotions of the driver and passengers, improve the adaptability of the sound effect and the driver and passengers, and enhance the driving or riding experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119148964B_ABST
    Figure CN119148964B_ABST
Patent Text Reader

Abstract

The present application discloses a method for adapting in-vehicle sound effects and an electronic device. The method comprises: performing audio separation on the audio to be played to obtain multiple audio segments, where different audio segments correspond to different types of audio data; performing emotion recognition on the driver or passenger corresponding to the target vehicle to obtain an emotion recognition result; wherein the target vehicle is equipped with an audio player; based on the emotion recognition result, calculating the output proportional coefficient corresponding to each audio segment; and outputting each audio segment to the audio player for playback according to the output proportional coefficient. The present application can automatically adjust the audio playback effect according to the driver or passenger's emotion, thereby improving the sound effect of the currently playing audio and the adaptability of the driver or passenger, thereby enhancing the driver or passenger's driving or riding experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of audio processing technology, and in particular to a vehicle-mounted sound effect adaptation method and electronic equipment. Background Art

[0002] Currently, in-car audio playback is a very common use case in cars, such as music playback and car radio. The sound quality of in-car audio playback will greatly affect the driving and riding experience of the vehicle.

[0003] However, tuning the car's sound effects is a very professional job that requires very high musical sense and acoustic knowledge. It is difficult for ordinary users to adjust the sound effects to their own preferences within the existing personalized settings. Summary of the Invention

[0004] In order to solve the above technical problems, the present application at least provides a vehicle-mounted sound effect adaptation method and electronic device.

[0005] In a first aspect, the present application provides a method for adapting in-vehicle sound effects, the method comprising: performing audio separation on the audio to be played to obtain multiple audio segments, where different audio segments correspond to different types of audio data; performing emotion recognition on the driver and passenger corresponding to the target vehicle to obtain an emotion recognition result; wherein the target vehicle is equipped with an audio player; based on the emotion recognition result, calculating the output proportional coefficient corresponding to each audio segment; and outputting each audio segment to the audio player for playback according to the output proportional coefficient.

[0006] In one embodiment, the types of audio data include direct sound, ambient sound and bass sound; audio separation is performed on the audio to be played to obtain multiple audio segments, including: performing short-time Fourier transform on the audio to be played, and separating direct sound segments and ambient sound segments from the audio to be played based on the short-time Fourier transform result; and filtering the audio to be played based on a preset low-pass filter to obtain a bass sound segment.

[0007] In one embodiment, emotion recognition is performed on the driver and passenger corresponding to the target vehicle to obtain an emotion recognition result, including: obtaining an object image obtained by image capture of the driver and passenger, an object voice obtained by voice capture of the driver and passenger, vehicle driving environment information of the target vehicle, and a driving destination of the target vehicle; based on one or more of the object image, the object voice, the vehicle driving environment information, and the driving destination, emotion recognition is performed on the driver and passenger to obtain an emotion recognition result.

[0008] In one embodiment, emotion recognition is performed on the driving object based on the object image to obtain an emotion recognition result, including: extracting the emotion-related image area in the object image; and performing emotion recognition on the emotion-related image area using a preset emotion recognition algorithm to obtain an emotion recognition result.

[0009] In one embodiment, there are multiple driving objects; emotion recognition is performed on the driving objects corresponding to the target vehicle to obtain emotion recognition results, including: emotion recognition is performed on each driving object separately to obtain an initial recognition result corresponding to each driving object; and the emotion recognition result is obtained by combining the initial recognition results corresponding to each driving object.

[0010] In one embodiment, an emotion recognition result is obtained in combination with the initial recognition results corresponding to each driving object, including: determining the role, riding position, and riding frequency of each driving object respectively; calculating the weight parameter of each driving object based on one or more of the role, riding position, and riding frequency of each driving object; and performing weighted fusion on the initial recognition results corresponding to each driving object according to the weight parameter of each driving object to obtain the emotion recognition result.

[0011] In one embodiment, there are multiple audio players, and different audio players are deployed at different positions in the target vehicle; based on the emotion recognition results, the output proportional coefficient corresponding to each audio clip is calculated, including: obtaining the deployment position of each audio player in the target vehicle; based on the deployment position of each audio player and the emotion recognition result, respectively calculating the output proportional coefficient of each audio clip relative to each audio player.

[0012] In one embodiment, multiple types of audio players are deployed in the target vehicle; based on the deployment location of each audio player and the emotion recognition result, the output proportional coefficient of each audio clip relative to each audio player is calculated respectively, including: obtaining the player type of each audio player; based on the deployment location of each audio player, the player type of each audio player and the emotion recognition result, the output proportional coefficient of each audio clip relative to each audio player is calculated respectively.

[0013] In one embodiment, based on the deployment location of each audio player, the player type of each audio player and the emotion recognition result, the output proportional coefficient of each audio clip relative to each audio player is calculated respectively, including: based on the deployment location of the current audio player, the player type of the current audio player and the audio type of the current audio clip, querying a matching preset proportional coefficient calculation formula; substituting the emotion recognition result into the preset proportional coefficient calculation formula to calculate the output proportional coefficient of the current audio clip relative to the current audio player.

[0014] A second aspect of the present application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-mentioned vehicle sound effect adaptation method.

[0015] A third aspect of the present application provides a computer-readable storage medium having program instructions stored thereon, which implement the above-mentioned vehicle-mounted sound effect adaptation method when the program instructions are executed by a processor.

[0016] The above scheme obtains multiple audio segments by performing audio separation on the audio to be played, and different audio segments correspond to different types of audio data; performs emotion recognition on the driver and passenger corresponding to the target vehicle to obtain emotion recognition results; wherein the target vehicle is equipped with an audio player; based on the emotion recognition results, calculates the output proportional coefficient corresponding to each audio segment; outputs each audio segment to the audio player for playback according to the output proportional coefficient, and can automatically adjust the audio playback effect according to the driver and passenger's emotion, so as to improve the sound effect of the currently played audio and the adaptability of the driver and passenger, thereby enhancing the driver and passenger's vehicle driving or riding experience.

[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0019] Figure 1 This is a schematic diagram of an implementation environment for a solution provided by an embodiment of the present application;

[0020] Figure 2 is a schematic diagram of in-vehicle audio playback according to an exemplary embodiment of the present application;

[0021] Figure 3 is a flow chart of a vehicle-mounted sound effect adaptation method shown in an exemplary embodiment of the present application;

[0022] Figure 4 is a schematic diagram of the deployment of an audio player shown in an exemplary embodiment of the present application;

[0023] Figure 5 is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present application;

[0024] Figure 6 It is a schematic diagram of the structure of a computer-readable storage medium shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0025] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0026] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0027] The term "and / or" in this article is merely information describing the association of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0028] The following describes the in-vehicle sound effect adaptation method provided in the embodiment of the present application.

[0029] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment of a solution provided by an embodiment of the present application. The implementation environment of the solution may include a vehicle 110 and a server 120, and the vehicle 110 and the server 120 are in communication connection with each other.

[0030] Server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0031] In one example, the vehicle 110 obtains the audio to be played from the server 120, and calculates the output proportional coefficient of the audio segment corresponding to the audio to be played based on the emotion recognition result of the driver, so as to output the audio to be played to the audio player for playback according to the output proportional coefficient.

[0032] For an example, see Figure 2 , Figure 2 This is a schematic diagram of an in-car audio player shown in an exemplary embodiment of the present application. Figure 2As shown, vehicle 110 uses a QAM8295 chip as a system-on-a-chip (SOC). The SOC runs the QNX (Quantum Nexus) system as a foundation to process hardware drivers. A virtual machine runs in the QNX system and boots the Android system in the virtual machine. The audio playback software runs in the Android system to obtain and play the audio to be played. It is understood that other types of system control chips and other audio playback software operation methods can also be selected, and this application is not limited to this.

[0033] The audio playback software runs on the Android system and decodes the audio into a two-channel pulse code modulation (PCM) data stream. This PCM data passes through the Android framework layer, the hardware abstraction (HAL) layer, and the audio driver (Tiny Advanced Linux Sound Architecture, TinyALSA) before being transferred to the QNX system. Within the QNX system, the Qualcomm framework is used to transmit the PCM data to the audio digital signal processor (ADSP). The ADSP module is connected to an AKM (Asahi Kasei Microsystems) audio-specific digital signal processing (DSP) chip. The in-vehicle sound adaptation algorithm runs in the AKM chip (such as AKM7709), and separates audio clips (such as direct sound clips and ambient sound clips) in the HIFI core, or transmits data to the AK core for active frequency division to obtain audio clips (such as treble sound clips and bass sound clips). The AK core can also perform sound effect pre-processing such as gain, delay, and equalizer on the audio. After the processing is completed, the AK core allocates the output proportional coefficient of each audio clip to each audio channel based on the emotion recognition results of the driver and passenger. Then, after inputting it into the audio amplifier (AMP) of the audio Kung Fu module, the amplified analog signal is input into each audio player. Different audio players correspond to different audio channels.

[0034] It should be noted that Figure 2 This is only an exemplary description of the application scenario of the vehicle-mounted sound effect adaptation method, which can be flexibly adjusted according to the actual application scenario and is not limited in this application.

[0035] In the vehicle-mounted sound effect adaptation method provided in the embodiment of the present application, the execution entity of each step can be the vehicle 110, or the server 120, or the vehicle 110 and the server 120 can cooperate with each other to execute the method, that is, part of the steps of the method are executed by the vehicle 110 and the other part of the steps are executed by the server 120.

[0036] See also Figure 3 , Figure 3 This is a flow chart of a vehicle sound effect adaptation method shown in an exemplary embodiment of the present application. The vehicle sound effect adaptation method can be applied to Figure 1 It should be understood that the method can also be applied to other exemplary implementation environments and be specifically executed by devices in other implementation environments, and this embodiment does not limit the implementation environment to which the method is applicable.

[0037] like Figure 3 As shown, the vehicle sound effect adaptation method includes at least steps S310 to S340, which are described in detail as follows:

[0038] Step S310: performing audio separation on the audio to be played to obtain multiple audio segments, where different audio segments correspond to different types of audio data.

[0039] Among them, the audio to be played can be stored in the vehicle's local storage, and the audio to be played can be directly obtained from the vehicle's local storage; the audio to be played can also be stored in a server, and a request for obtaining the audio to be played is sent to the server to obtain the audio to be played returned by the server.

[0040] The audio to be played is composed of different types of audio data. Audio separation is performed on each type of audio data in the audio to be played to obtain multiple audio segments.

[0041] For example, the types of audio data can be divided according to the different sound sources, and can be divided into direct sound, ambient sound, etc., that is, direct sound segments, ambient sound segments, etc. are separated from the audio to be played; the types of audio data can be divided according to the different frequencies, and can be divided into treble sound, mid-range sound, bass sound, etc., that is, treble sound segments, mid-range sound segments, bass sound segments, etc. are separated from the audio to be played.

[0042] The division method of audio data types can be flexibly set according to actual conditions, and this application does not limit the division method of audio data types.

[0043] For example, based on the source of the audio to be played, the audio data type to be classified can be determined to improve the accuracy of audio separation and subsequent sound effect adaptation. For example, if the audio to be played comes from music playback software, the audio to be played needs to be classified into direct sound, ambient sound, high-pitched sound, mid-pitched sound, and low-pitched sound; if the audio to be played comes from crosstalk broadcasting software, the audio to be played needs to be classified into direct sound and ambient sound.

[0044] Step S320: performing emotion recognition on the driver and passenger corresponding to the target vehicle to obtain an emotion recognition result; wherein the target vehicle is equipped with an audio player.

[0045] It should be noted that the driving object refers to the driving object or the riding object corresponding to the target vehicle.

[0046] The target vehicle is deployed with audio players, and the number and type of audio players can be flexibly set according to actual conditions, and this application does not limit this.

[0047] Emotion recognition is performed on the driver and passenger corresponding to the target vehicle to obtain an emotion recognition result, such as whether the driver and passenger's emotion is excitement, happiness, calmness, frustration or sadness.

[0048] Exemplarily, the emotion recognition of the driver and passenger can be performed based on one or more parameters such as the driver and passenger's image, the driver and passenger's voice, the driver and passenger's terminal usage, the target vehicle's driving environment information, the target vehicle's driving destination, etc., to obtain an emotion recognition result.

[0049] It should be noted that there is no order in which step S310 and step S320 are executed, that is, step S310 can be executed first and then step S320, or step S320 can be executed first and then step S310, or step S310 and step S320 can be executed simultaneously.

[0050] Step S330: Calculate the output proportional coefficient corresponding to each audio segment based on the emotion recognition result.

[0051] Adjust the output scaling factor of each audio clip based on the emotion recognition results.

[0052] The number of riding objects can be one or more. The following describes the cases where the number of riding objects is one and the number of riding objects is multiple:

[0053] For example, if there is only one passenger, the output scale factor of each audio segment is adjusted based on the emotion recognition result of that passenger. For example, if the emotion recognition result for the passenger is excitement, the output scale factor of the direct sound segment is adjusted to 1; if the emotion recognition result for the passenger is calm, the output scale factor of the direct sound segment is adjusted to 0.8.

[0054] For example, if there are multiple passengers, one of the passengers can be selected as the target passenger (e.g., the passenger with the lowest or highest emotions, the passenger with the highest frequency of riding, or the passenger in the driver's seat or the front passenger seat), and the output scale factor of each audio segment can be adjusted based on the emotion recognition result of the target passenger. The output scale factor of each audio segment can also be adjusted based on the emotion recognition result of each passenger. For example, if the emotion recognition results for passenger 1 and passenger 2 are excitement, the output scale factor of the direct sound segment can be adjusted to 1. If the emotion recognition result for passenger 1 is excitement and the emotion recognition result for passenger 2 is depression, the output scale factor of the direct sound segment can be adjusted to 0.7.

[0055] The number of audio players deployed in the target vehicle can be one or more. The following describes the cases where the number of audio players is one and the number of audio players is multiple:

[0056] For example, if there is only one audio player, the output ratio coefficients corresponding to different audio clips can be directly queried based on the emotion recognition results. For example, if the emotion recognition result is excitement, the output ratio of the direct sound clip of the audio player is adjusted to 1, and the output ratio of the ambient sound clip is adjusted to 0.5.

[0057] For example, if there are multiple audio players, the output scaling coefficients corresponding to different audio clips can be queried based on the emotion recognition results, and the output scaling coefficients can be unified as the output scaling coefficients of each audio player. Alternatively, the output scaling coefficients corresponding to different audio clips can be queried for each audio player based on the deployment location and / or player type of each audio player. This means that the output scaling coefficients corresponding to each audio player may be different. For example, for an audio player deployed in the left front door channel, the output scaling coefficient of the direct sound clip queried is 0.6, while for an audio player deployed in the center channel, the output scaling coefficient of the direct sound clip queried is 0.8.

[0058] Of course, in addition to the above embodiments, the spatial distance between each driving object and the audio player can also be comprehensively referred to to determine the degree of influence of the emotion recognition result of the driving object on the audio player based on the spatial distance between the driving object and the audio player, thereby improving the accuracy of the output proportional coefficient calculation.

[0059] Step S340: Output each audio clip to an audio player for playback according to the output scale factor.

[0060] After obtaining the output proportional coefficient corresponding to each audio clip, the strength of each audio clip in the audio player can be determined, so as to output each audio clip to the audio player for playback according to the output proportional coefficient.

[0061] In some embodiments, the output scaling factor can be calculated and modified when a change in the occupant's emotion is detected. Alternatively, the occupant's emotion can be periodically identified, and the output scaling factor can be calculated and modified based on the identified emotion. Alternatively, after calculating the current output scaling factor, the time interval between the last modification of the output scaling factor and the current time is detected. If the time interval is greater than a preset interval threshold, the current output scaling factor is modified to a new output scaling factor. Once the new output scaling factor is obtained, each audio clip is output to an audio player for playback according to the new output scaling factor.

[0062] This application can automatically adjust the audio playback effect according to the emotions of the driver and passenger, so as to improve the sound effect of the currently played audio and its adaptability to the driver and passenger, thereby enhancing the driver and passenger's vehicle driving or riding experience.

[0063] Next, some embodiments of the present application are described in detail.

[0064] In some embodiments, the types of audio data include direct sound, ambient sound, and bass sound; in step S310, audio separation is performed on the audio to be played to obtain multiple audio segments, including: performing short-time Fourier transform on the audio to be played, and separating direct sound segments and ambient sound segments from the audio to be played based on the short-time Fourier transform result; and filtering the audio to be played based on a preset low-pass filter to obtain a bass sound segment.

[0065] Taking the audio to be played as a stereo sound source as an example, the audio to be played consists of left channel audio data and right channel audio data. The left channel audio data is represented as X L , the right channel audio data is represented as X R , the left channel audio data X L The direct sound segment in is denoted as P L , the left channel audio data X L The ambient sound segment in is represented as U L , the right channel audio data is represented as X R The direct sound segment in is denoted as P R , the right channel audio data is represented as X R The ambient sound segment in is represented as U R, then we get the following formula 1 and formula 2:

[0066] X L =P L +U L (Formula 1)

[0067] X R =P R +U R (Formula 2)

[0068] Performing short-time Fourier transform on the left-channel audio data and the right-channel audio data in the audio to be played yields the following formulas 3 and 4:

[0069] X L (i,k)=P L (i,k)+U L (i,k) (Formula 3)

[0070] X R (i,k)=B(i,k)P R (i,k)+U R (i,k) (Formula 4)

[0071] Among them, i is the time frame index of the left channel audio data or the right channel audio data, k is the frequency index of the left channel audio data or the right channel audio data, and B is the correlation coefficient between the left channel audio data and the right channel audio data.

[0072] Due to X L and X R If we know, we can estimate B and P L 、U L 、U R The value of can complete the separation of the direct sound and the ambient sound of the left channel audio data and the right channel audio data, and estimate B, P L 、U L 、U R The steps for the value of can be:

[0073] X L The short-term energy estimation of is expressed as formula 5:

[0074]

[0075] Here, E{·} represents expectation.

[0076] Similarly, we get X R Short-term energy estimation of

[0077] In formula 3 and formula 4, the ambient sound of the left channel audio data and the right channel audio data has the same short-term energy, which is recorded as P A, the short-time energy of coherent sound is recorded as P S Combining the relevant assumptions, we can get Formula 6 and Formula 7:

[0078]

[0079] Combine the above formulas to define the normalized channel correlation coefficient of the audio to be played Expressed as formula 8:

[0080]

[0081] Substitute Formula 3 and Formula 4 into Formula 8, and combine Formula 6 and Formula 7 to obtain Formula

[0082] Formula 9:

[0083]

[0084] B.P S and P A It's about and function, so we can solve B, P S and P A :

[0085]

[0086] in:

[0087]

[0088] According to the above formulas 10 to 14, B and P are calculated. S and P A Afterwards, the same method is used to analyze the coherent sound S, the ambient sound U of the left channel audio data and the right channel audio data. L 、U R Estimate separately:

[0089] S=W L (P S +U L )+W R (P S +U R )(Formula 15)

[0090] Among them, W L 、W R To estimate the weight, the estimation error of S is expressed as formula 16:

[0091] σS=(1-W L -W R B)SW L U L -WR U R (Formula 16)

[0092] In the least squares algorithm, when the estimation error is completely unrelated to the audio to be played, the obtained weight is the optimal estimate:

[0093] E{σSX L}=0 (Formula 17)

[0094] E{σSX R}=0 (Formula 18)

[0095] At this time, the estimated weight of the optimal estimate is:

[0096]

[0097] Combining the above formula, we get P L 、U L 、P R 、U R value.

[0098] In addition, the audio to be played is filtered using a preset low-pass filter to obtain a bass sound segment. For example, the preset low-pass filter is used to filter out audio data below 120 Hz to obtain a bass sound segment.

[0099] In the above manner, the direct sound segment, the ambient sound segment and the bass sound segment are separated.

[0100] In some embodiments, step S320 performs emotion recognition on the driver and passenger corresponding to the target vehicle to obtain an emotion recognition result, including: obtaining an object image obtained by image capture of the driver and passenger, an object voice obtained by voice capture of the driver and passenger, vehicle driving environment information of the target vehicle, and a driving destination of the target vehicle; performing emotion recognition on the driver and passenger based on one or more of the object image, object voice, vehicle driving environment information, and driving destination to obtain an emotion recognition result.

[0101] Exemplarily, an image acquisition device is deployed in the target vehicle, and the image acquisition device is used to acquire an image of the driver and passenger to obtain an object image, and perform emotion recognition on the driver and passenger based on the object image to obtain an emotion recognition result.

[0102] For example, emotion recognition results are obtained by analyzing the facial expressions, body movements, age, gender, etc. of the driver in the object image.

[0103] For example: extract the emotion-related image area in the object image; use the preset emotion recognition algorithm to perform emotion recognition on the emotion-related image area to obtain the emotion recognition result.

[0104] The emotion-related image region may be a facial image region of the passenger; the emotion-related image region may also be a facial image region and a body image region of the passenger.

[0105] The preset emotion recognition algorithm can be an emotion recognition neural network model that has been pre-trained. The emotion-related image area is input into the emotion recognition neural network model for processing to obtain the emotion recognition result output by the emotion recognition neural network model.

[0106] Exemplarily, a voice collection device is deployed in the target vehicle, and the voice of the driver and passenger is collected by the voice collection device to obtain the target voice. Based on the target voice, emotion recognition is performed on the driver and passenger to obtain an emotion recognition result.

[0107] For example, emotion recognition results are obtained by analyzing the text content, intonation, volume, etc. of the driver's voice in the target voice.

[0108] Exemplarily, the target vehicle is deployed with environmental perception equipment, such as lidar, image acquisition equipment, audio acquisition equipment, speed sensor, etc., and the vehicle driving environment information is collected through the environmental perception equipment. The vehicle driving environment information includes but is not limited to the vehicle driving speed, vehicle driving scene (such as city, suburbs, desert, etc.), road congestion level, road flatness, etc. Based on the vehicle driving environment information, emotion recognition is performed on the driver and passenger to obtain emotion recognition results.

[0109] For example, if the vehicle driving environment information indicates that the road congestion level is lower, the driver's mood may be happier; if the vehicle driving environment information indicates that the road congestion level is higher, the driver's mood may be more depressed.

[0110] Exemplarily, the driving path of the target vehicle is obtained to obtain the driving destination of the target vehicle, and emotion recognition is performed on the driver and passenger based on the driving destination to obtain an emotion recognition result.

[0111] For example, if the driving destination is an entertainment venue such as a park or an amusement park, the driver and passengers may be in a happier mood.

[0112] Exemplarily, emotion recognition of the driver can be performed by combining one or more combinations of the object image, the object voice, the vehicle driving environment information, and the driving destination to obtain an emotion recognition result, so as to improve the accuracy of the emotion recognition result.

[0113] In some embodiments, there are multiple driving objects; in step S320, emotion recognition is performed on the driving objects corresponding to the target vehicle to obtain an emotion recognition result, including: emotion recognition is performed on each driving object separately to obtain an initial recognition result corresponding to each driving object; and the emotion recognition result is obtained by combining the initial recognition results corresponding to each driving object.

[0114] When there are multiple passengers in the target vehicle, the emotion corresponding to each passenger is determined based on the object image and / or object voice of each passenger, or in combination with the vehicle driving environment information, driving destination, etc., to obtain the initial recognition result.

[0115] The final emotion recognition result is obtained based on the initial recognition results corresponding to each driver and passenger.

[0116] Exemplarily, the emotion recognition result is obtained in combination with the initial recognition results corresponding to each driving object, including: determining the role, riding position, and riding frequency of each driving object respectively; calculating the weight parameter of each driving object based on one or more of the role, riding position, and riding frequency of each driving object; and performing weighted fusion on the initial recognition results corresponding to each driving object according to the weight parameter of each driving object to obtain the emotion recognition result.

[0117] For example, according to family relationships, the role of each driver and passenger can be divided into father, mother, child, etc.; the riding position can be divided into main driving position, co-pilot position, back seat position, etc.; the riding frequency refers to the number of times the driver and passenger rides or drives the target vehicle within a preset time period.

[0118] The weight parameter of each riding object is calculated according to one or more of the role, riding position, and riding frequency of each riding object.

[0119] For example, the roles belonging to different driving objects have different priorities. The higher the role priority, the higher the weight of the corresponding driving object; different riding positions have different priorities. The higher the position priority, the higher the weight of the corresponding driving object; and the higher the riding frequency, the higher the weight of the corresponding driving object.

[0120] According to the weight parameter of each driving object, the initial recognition results corresponding to each driving object are weighted fused to obtain the emotion recognition result.

[0121] For example, emotion recognition results are represented by an emotion value between 0 and 1. Values ​​closer to 1 indicate a more excited passenger, while values ​​closer to 0 indicate a sadder passenger. If the target vehicle's passengers include passenger 1 and passenger 2, and passenger 1 has an emotion value of 0.5 and a weight parameter of 0.7, and passenger 2 has an emotion value of 0.8 and a weight parameter of 0.8, then the final emotion value obtained by weighted fusion is 0.5*0.7+0.8*0.3=0.59.

[0122] In addition to obtaining the emotion recognition result by fusing the initial recognition results corresponding to each driver and passenger as described above, other methods may also be used to determine the final emotion recognition result.

[0123] For example, the emotion recognition result is represented by an emotion value between 0 and 1, and the emotion value with the largest value can be selected from multiple initial recognition results as the final emotion recognition result.

[0124] For another example, a weight parameter of each passenger is determined, and the initial recognition result corresponding to the passenger with the largest weight parameter is used as the final emotion recognition result.

[0125] For example, the emotion recognition result is an emotion classification (such as excitement, calmness, frustration, etc.). The initial recognition results of each driver and passenger are counted to obtain the number of drivers and passengers corresponding to each emotion category, and the emotion category with the largest number is selected as the final emotion recognition result.

[0126] The final emotion recognition result determined in the above embodiment is used to calculate the output proportional coefficient corresponding to each audio segment.

[0127] Of course, in another embodiment, after obtaining the initial recognition results corresponding to each driving object, the output proportional coefficient corresponding to each audio segment is directly determined according to the multiple initial recognition results.

[0128] For example, for audio player 1, the spatial distance between each passenger and audio player 1 is used to determine the impact of each passenger's emotions on audio player 1. The greater the spatial distance, the smaller the impact. Then, combining the initial recognition results and the impact of each passenger, the output proportional coefficient of each audio clip relative to audio player 1 is calculated.

[0129] For another example, the output proportional coefficient of each audio clip relative to audio player 1 can be calculated by combining the initial recognition result corresponding to each driver, the influence of each driver's emotion on the audio player, and the weight parameter of each driver.

[0130] In some embodiments, there are multiple audio players, and different audio players are deployed at different locations in the target vehicle. In step S330, based on the emotion recognition result, the output proportional coefficient corresponding to each audio clip is calculated, including: obtaining the deployment location of each audio player in the target vehicle; and based on the deployment location of each audio player and the emotion recognition result, respectively calculating the output proportional coefficient of each audio clip relative to each audio player.

[0131] Combined with the deployment location of each audio player and the emotion recognition results, the output proportional coefficient of each audio clip relative to each audio player is calculated respectively.

[0132] For example, if the emotion recognition result is calm, the output proportional coefficient of the direct sound segment of the audio player deployed at the left front door is calculated to be 0.6; the output proportional coefficient of the direct sound segment of the audio player deployed in the middle of the vehicle is calculated to be 0.8.

[0133] In some embodiments, multiple types of audio players are deployed in the target vehicle; in step S330, based on the deployment location of each audio player and the emotion recognition result, the output proportional coefficient of each audio clip relative to each audio player is calculated separately, including: obtaining the player type of each audio player; based on the deployment location of each audio player, the player type of each audio player and the emotion recognition result, the output proportional coefficient of each audio clip relative to each audio player is calculated separately.

[0134] For example, see Figure 4 , Figure 4 A schematic diagram of the deployment of an audio player is shown as an exemplary embodiment of the present application. Figure 4 As shown in the figure, the target vehicle is equipped with tweeters, mid-range speakers, woofers, full-range speakers, D-pillar surround speakers, etc. at different locations. Figure 4 This is just an example of deploying multiple audio players of various types. In actual application scenarios, the type and deployment location of the audio players can be flexibly selected according to the actual application situation. This application does not limit this.

[0135] Based on the deployment location of each audio player, the player type of each audio player, and the emotion recognition result, the output proportional coefficient of each audio clip relative to each audio player is calculated respectively.

[0136] For example, if the emotion recognition result is calm, the output proportional coefficient of the direct sound segment of the tweeter deployed in the left front door is calculated to be 0.6; the output proportional coefficient of the direct sound segment of the woofer deployed in the left front door is calculated to be 0.5.

[0137] In some embodiments, based on the deployment location of each audio player, the player type of each audio player, and the emotion recognition result, the output proportional coefficient of each audio clip relative to each audio player is calculated separately, including: based on the deployment location of the current audio player, the player type of the current audio player, and the audio type of the current audio clip, querying a matching preset proportional coefficient calculation formula; substituting the emotion recognition result into the preset proportional coefficient calculation formula to calculate the output proportional coefficient of the current audio clip relative to the current audio player.

[0138] That is, a preset proportional coefficient calculation formula is pre-set for audio players in different deployment locations, different types of audio players, and different types of audio clips. The preset proportional coefficient calculation formula that matches the deployment location of the current audio player, the player type of the current audio player, and the audio type of the current audio clip is queried, and the emotion recognition result is substituted into the preset proportional coefficient calculation formula to calculate the output proportional coefficient of the current audio clip relative to the current audio player.

[0139] Through the above method, the output proportional coefficient of each audio clip relative to each audio player is calculated.

[0140] For example, assuming that the emotion recognition result is Y, the left channel direct sound segment is represented by P L , the left channel ambient sound segment is represented as U L , the right channel direct sound segment is represented as P R , the right channel ambient sound segment is represented as U R , the bass sound segment is D, and the audio player includes a center full-range speaker, a left door full-range speaker, a right door full-range speaker, a left D-pillar surround speaker, a right D-pillar surround speaker, and a subwoofer. The calculation formula for determining the preset proportional coefficient of each audio segment relative to each audio player is:

[0141] 1. For the left channel direct sound clip and the right channel direct sound clip:

[0142] Center full-range speaker:

[0143] (P L +P R )(0.8+(Y-0.5)*2*0.2)(Formula 21)

[0144] Left door full-range speaker:

[0145]

[0146] Right door full-range speaker:

[0147]

[0148] 2. For the left channel ambient sound clip and the right channel ambient sound clip:

[0149] Center full-range speaker:

[0150]

[0151] Left door full-range speaker:

[0152]

[0153] Right door full-range speaker:

[0154]

[0155] Left D-pillar surround speaker:

[0156]

[0157] Right D-pillar surround speaker:

[0158]

[0159] 3. For bass sound clips:

[0160] Woofer:

[0161] D(0.6+(Y-0.5)*2*0.4)(Formula 29)

[0162] The output proportional coefficient of each audio clip relative to each audio player is calculated using the preset proportional coefficient calculation formula found above.

[0163] In another embodiment, the output proportional coefficient of each audio clip relative to each audio player can also be determined by looking up a table. For example, a proportional coefficient table is pre-stored, and the proportional coefficient table is used to store the output proportional coefficients corresponding to different audio clips, different audio players and different emotion recognition results.

[0164] For example, see Table 1 below, which is a proportional coefficient table shown in an exemplary embodiment:

[0165]

[0166] Table 1

[0167] By querying Table 1, the output proportional coefficient of each audio clip relative to each audio player can be obtained.

[0168] The in-vehicle sound effect adaptation method provided in the present application obtains multiple audio segments by performing audio separation on the audio to be played, and different audio segments correspond to different types of audio data; performs emotion recognition on the driver and passenger corresponding to the target vehicle to obtain an emotion recognition result; wherein the target vehicle is equipped with an audio player; based on the emotion recognition result, calculates the output proportional coefficient corresponding to each audio segment; outputs each audio segment to the audio player for playback according to the output proportional coefficient, and can automatically adjust the audio playback effect according to the emotion of the driver and passenger, so as to improve the sound effect of the currently played audio and the adaptability of the driver and passenger, and enhance the driving or riding experience of the driver and passenger.

[0169] This application also provides an electronic device, see Figure 5 , Figure 5 This is a schematic diagram of the structure of an embodiment of an electronic device of the present application. Electronic device 500 includes memory 501 and processor 502. Processor 502 is configured to execute program instructions stored in memory 501 to implement the steps of any of the aforementioned in-vehicle sound effect adaptation method embodiments. In a specific implementation scenario, electronic device 500 may include, but is not limited to, a microcomputer and a server. Furthermore, electronic device 500 may also include mobile devices such as laptops and tablet computers, which are not limited herein.

[0170] Specifically, the processor 502 is used to control itself and the memory 501 to implement the steps in any of the above-mentioned embodiments of the vehicle-mounted sound effect adaptation method. The processor 502 can also be called a central processing unit (CPU). The processor 502 may be an integrated circuit chip with signal processing capabilities. The processor 502 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 502 can be implemented by an integrated circuit chip.

[0171] This application also provides a computer-readable storage medium, see Figure 6 , Figure 6The computer-readable storage medium 600 stores program instructions 610 that can be executed by a processor, and the program instructions 610 are used to implement the steps of any of the above-mentioned vehicle-mounted sound effect adaptation method embodiments.

[0172] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0173] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0174] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0175] In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A vehicle sound effect adaptation method, characterized in that: include: Separate the audio to be played to obtain multiple audio segments, where different audio segments correspond to different types of audio data. Performing emotion recognition on a driver or passenger corresponding to a target vehicle to obtain an emotion recognition result; wherein the target vehicle is equipped with an audio player; Calculating an output proportional coefficient corresponding to each audio segment based on the emotion recognition result; According to the output proportional coefficient, each audio segment is output to the audio player for playing.

2. The method according to claim 1, characterized in that The types of audio data include direct sound, ambient sound, and bass sound; The audio to be played is separated to obtain multiple audio segments, including: Performing a short-time Fourier transform on the audio to be played, and separating a direct sound segment and an ambient sound segment from the audio to be played based on a result of the short-time Fourier transform; as well as, The audio to be played is filtered based on a preset low-pass filter to obtain a bass sound segment.

3. The method according to claim 1, characterized in that The emotion recognition of the driver and passenger corresponding to the target vehicle to obtain the emotion recognition result includes: Acquiring an object image obtained by collecting an image of the driving object, an object voice obtained by collecting a voice of the driving object, vehicle driving environment information of the target vehicle, and a driving destination of the target vehicle; Based on one or more of the object image, the object voice, the vehicle driving environment information, and the driving destination, emotion recognition is performed on the driver to obtain an emotion recognition result.

4. The method according to claim 3, characterized in that Performing emotion recognition on the driver based on the object image to obtain an emotion recognition result includes: extracting an emotion-related image region in the object image; Emotion recognition is performed on the emotion-related image region using a preset emotion recognition algorithm to obtain an emotion recognition result.

5. The method according to claim 1, wherein There are multiple driving objects; and performing emotion recognition on the driving objects corresponding to the target vehicle to obtain an emotion recognition result includes: Performing emotion recognition on each driver and passenger respectively to obtain an initial recognition result corresponding to each driver and passenger; The emotion recognition result is obtained by combining the initial recognition results corresponding to each driving object.

6. The method according to claim 5, characterized in that The emotion recognition result is obtained by combining the initial recognition results corresponding to each of the driving objects, including: Determining the role, riding position, and riding frequency of each riding object; Calculating a weight parameter for each riding object based on one or more of a role, a riding position, and a riding frequency of each riding object; According to the weight parameter of each driving object, the initial recognition results corresponding to each driving object are weightedly fused to obtain the emotion recognition result.

7. The method according to claim 1, characterized in that There are multiple audio players, and different audio players are deployed at different locations on the target vehicle; and calculating the output proportional coefficient corresponding to each audio clip based on the emotion recognition result includes: Obtaining a deployment position of each audio player in the target vehicle; Based on the deployment position of each audio player and the emotion recognition result, an output proportional coefficient of each audio clip relative to each audio player is calculated respectively.

8. The method according to claim 7, characterized in that The target vehicle is equipped with multiple types of audio players. Calculating the output proportional coefficient of each audio clip relative to each audio player based on the deployment location of each audio player and the emotion recognition result includes: Obtain the player type of each audio player; Based on the deployment location of each audio player, the player type of each audio player and the emotion recognition result, an output proportional coefficient of each audio clip relative to each audio player is calculated respectively.

9. The method according to claim 8, characterized in that The calculating, based on the deployment location of each audio player, the player type of each audio player, and the emotion recognition result, the output proportional coefficient of each audio clip relative to each audio player includes: Based on the deployment location of the current audio player, the player type of the current audio player, and the audio type of the current audio clip, a matching preset scaling coefficient calculation formula is searched; The emotion recognition result is substituted into the preset proportional coefficient calculation formula to calculate the output proportional coefficient of the current audio segment relative to the current audio player.

10. An electronic device, characterized in that: The electronic device includes a memory and a processor, and the processor is used to execute program instructions stored in the memory to implement the steps in the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Negative emotion relieving method and system for driver

    CN112137630A

  • Sound effect adjusting method and device and audio equipment

    CN114283854A