Audio adjustment method and device, electronic equipment and computer readable storage medium

By acquiring and adjusting multiple emotional information of audio, the problem that traditional audio systems cannot respond to emotional changes is solved, audio emotions are enhanced, and user experience is improved.

CN120631300APending Publication Date: 2025-09-12SHENZHEN TCL NEW-TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510645854.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional audio systems are unable to recognize and respond to emotional changes in program content, resulting in the audio output being unable to effectively enhance the user's emotional experience.

Method used

By obtaining the emotional values ​​and adjustment methods of multiple emotional information of the target audio, the audio is adjusted, including emotional information classification, difference threshold judgment, historical moment emotional value smoothing and audio parameter adjustment, such as dynamic range, reverberation effect and loudness adjustment.

Benefits of technology

It enhances the emotions conveyed by audio and improves the user's emotional experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120631300A_ABST
    Figure CN120631300A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an audio adjustment method and device, electronic equipment and a computer readable storage medium. The method comprises the steps of obtaining respective emotion values of multiple pieces of emotion information of target audio; adjusting modes corresponding to the multiple pieces of emotion information are obtained; and adjusting the target audio according to the respective emotion values of the plurality of pieces of emotion information and the respective adjustment modes corresponding to the plurality of pieces of emotion information. Therefore, according to the scheme, the emotional experience of the user can be enhanced by adjusting the audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing technology, and specifically to an audio adjustment method, device, electronic device, and computer-readable storage medium. Background Art

[0002] With the development of artificial intelligence technology, affective computing has gradually become an important means of improving user experience. Traditional audio systems are unable to recognize and respond to emotional changes in program content. As a result, when watching emotionally rich programs (such as movies and TV series) or listening to emotionally rich music, the audio output cannot effectively enhance the user's emotional experience. Summary of the Invention

[0003] The embodiments of the present application provide an audio adjustment method, device, electronic device, and computer-readable storage medium, which can enhance the user's emotional experience by adjusting the audio.

[0004] In a first aspect, an embodiment of the present application provides an audio adjustment method, comprising:

[0005] Obtaining the emotional values ​​of multiple emotional information of the target audio;

[0006] Obtaining adjustment methods corresponding to the plurality of emotion information;

[0007] The target audio is adjusted according to the emotional values ​​of the multiple emotional information and the adjustment methods corresponding to the multiple emotional information.

[0008] In one embodiment, adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes:

[0009] Classifying multiple emotional information of the target audio;

[0010] Get the difference between multiple sentiment values ​​in each category;

[0011] Get the difference threshold;

[0012] When the difference between multiple emotion information in the same category is not greater than the difference threshold, the target audio is adjusted according to the emotion values ​​of the multiple emotion information in the category and the adjustment methods corresponding to the multiple emotion information in the category.

[0013] In one embodiment, obtaining the emotion value of each of the plurality of emotion information of the target audio includes:

[0014] Obtaining a pre-processing emotion value of each of the plurality of emotion information of the target audio;

[0015] Obtaining respective emotion values ​​of a plurality of emotion information of audios at a plurality of historical moments corresponding to the target audio;

[0016] Get the weights of multiple historical moments;

[0017] According to the weights of each of the multiple historical moments and the emotional values ​​of each of the multiple emotional information of the audio at the multiple historical moments, the emotional values ​​of each of the multiple emotional information of the target audio before processing are smoothed to obtain the emotional values ​​of each of the multiple emotional information of the target audio.

[0018] In one embodiment, obtaining the pre-processing emotion value of each of the plurality of emotion information of the target audio includes:

[0019] According to multiple emotion value acquisition methods, respectively obtain the sub-emotional values ​​of the multiple emotion information of the target audio;

[0020] Obtaining weights corresponding to the plurality of sentiment value acquisition methods;

[0021] According to the weights corresponding to the multiple emotion value acquisition methods, the sub-emotion values ​​corresponding to the multiple emotion value acquisition methods are weighted and summed to obtain the emotion values ​​before processing of the multiple emotion information.

[0022] In one embodiment, the step of obtaining the sub-emotional values ​​of the plurality of emotion information of the target audio in accordance with the plurality of emotion value obtaining methods comprises at least one of the following steps:

[0023] Determining, according to the sound source label of the target audio, a first sub-emotion value of each of the plurality of emotion information of the target audio;

[0024] Performing audio analysis on the target audio to obtain a second sub-emotion value of each of the plurality of emotion information of the target audio;

[0025] Obtaining a video frame corresponding to the target audio;

[0026] Performing image analysis on the video frame to obtain a third sub-emotion value of each of the multiple emotion information of the target audio;

[0027] Get the emotion value set by the user;

[0028] According to the emotion value set by the user, the fourth sub-emotion value of each of the multiple emotion information of the target audio is determined.

[0029] In one embodiment, performing audio analysis on the target audio to obtain the second sub-emotional value of each of the plurality of emotion information of the target audio includes:

[0030] Converting the target audio into a frequency domain signal;

[0031] Sampling the frequency domain signal to obtain multiple frequency points;

[0032] Get multiple frequency ranges;

[0033] According to the ratio of the frequency points in the multiple frequency ranges, the second sub-emotional value of each of the multiple emotion information of the target audio is determined.

[0034] In one embodiment, adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes:

[0035] Adjusting the parameters of the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information;

[0036] The audio parameters include at least dynamic range, reverberation effect, loudness and / or impact.

[0037] In one embodiment, adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes:

[0038] According to the emotional value of the joy emotional information, the target audio is adjusted according to one or more of the following adjustment methods:

[0039] enhancing a dynamic range of an audio signal within a first target frequency range in the target audio;

[0040] reducing the reverberation effect of the target audio;

[0041] The loudness of the target audio is enhanced.

[0042] In one embodiment, adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes:

[0043] According to the emotional value of the sorrowful emotional information, the target audio is adjusted according to one or more of the following adjustment methods:

[0044] enhancing a dynamic range of an audio signal within a second target frequency range of the target audio;

[0045] enhancing the reverberation effect of the target audio;

[0046] The loudness of the target audio is reduced.

[0047] In one embodiment, adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes:

[0048] According to the emotional value of the dynamic emotional information, the target audio is adjusted in one or more of the following adjustment methods:

[0049] enhancing a dynamic range of an audio signal within a third target frequency range of the target audio;

[0050] enhancing the impact of an audio signal within a fourth target frequency range of the target audio;

[0051] The loudness of the target audio is enhanced.

[0052] In one embodiment, adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes:

[0053] According to the emotional value of the soothing emotional information, the target audio is adjusted in one or more of the following adjustment methods:

[0054] reducing the dynamic range of the audio signal within a fifth target frequency range of the target audio;

[0055] increasing the reverberation delay of the target audio;

[0056] The loudness of the target audio is reduced.

[0057] In a second aspect, an embodiment of the present application provides an audio adjustment device, comprising:

[0058] A first acquisition module is used to obtain the emotion value of each of the multiple emotion information of the target audio;

[0059] A second acquisition module is used to obtain the adjustment methods corresponding to the plurality of emotion information;

[0060] The adjustment module is used to adjust the target audio according to the emotional values ​​of the multiple emotional information and the adjustment methods corresponding to the multiple emotional information.

[0061] In one embodiment, the adjustment module includes:

[0062] A classification unit, configured to classify multiple emotional information of the target audio;

[0063] A difference acquisition unit, used to obtain the difference between multiple emotion values ​​in each category;

[0064] A threshold value obtaining unit, used for obtaining a difference threshold value;

[0065] The first adjustment unit is used to adjust the target audio according to the emotional values ​​of the multiple emotional information in the same category and the adjustment methods corresponding to the multiple emotional information in the category when the difference between the multiple emotional information in the same category is not greater than the difference threshold.

[0066] In one embodiment, the first acquisition module includes:

[0067] a pre-processing emotion value obtaining unit, configured to obtain a pre-processing emotion value of each of the plurality of emotion information of the target audio;

[0068] A historical emotion value acquisition unit, configured to acquire emotion values ​​of a plurality of emotion information of audios at a plurality of historical moments corresponding to the target audio;

[0069] A historical weight acquisition unit, used to obtain the weights of multiple historical moments;

[0070] A smoothing processing unit is used to smooth the pre-processed emotional values ​​of the multiple emotional information of the target audio according to the respective weights of the multiple historical moments and the respective emotional values ​​of the multiple emotional information of the audio at the multiple historical moments, so as to obtain the respective emotional values ​​of the multiple emotional information of the target audio.

[0071] In one embodiment, the pre-processing emotion value acquisition unit includes:

[0072] The sub-emotion value acquisition sub-unit is used to respectively acquire the sub-emotion values ​​of the plurality of emotion information of the target audio according to a plurality of emotion value acquisition methods;

[0073] A weight acquisition subunit, used to obtain weights corresponding to the plurality of sentiment value acquisition methods;

[0074] The pre-processing emotion value acquisition subunit is used to perform weighted summation on the sub-emotional values ​​corresponding to the multiple emotion value acquisition methods according to the weights corresponding to the multiple emotion value acquisition methods, so as to obtain the pre-processing emotion value of each of the multiple emotion information.

[0075] In one embodiment, the step of obtaining the sub-emotional values ​​of the plurality of emotion information of the target audio in accordance with the plurality of emotion value obtaining methods comprises at least one of the following steps:

[0076] Determining, according to the sound source label of the target audio, a first sub-emotion value of each of the plurality of emotion information of the target audio;

[0077] Performing audio analysis on the target audio to obtain a second sub-emotion value of each of the plurality of emotion information of the target audio;

[0078] Obtaining a video frame corresponding to the target audio;

[0079] Performing image analysis on the video frame to obtain a third sub-emotion value of each of the multiple emotion information of the target audio;

[0080] Get the emotion value set by the user;

[0081] According to the emotion value set by the user, the fourth sub-emotion value of each of the multiple emotion information of the target audio is determined.

[0082] In one embodiment, performing audio analysis on the target audio to obtain the second sub-emotional value of each of the plurality of emotion information of the target audio includes:

[0083] Converting the target audio into a frequency domain signal;

[0084] Sampling the frequency domain signal to obtain multiple frequency points;

[0085] Get multiple frequency ranges;

[0086] According to the ratio of the frequency points in the multiple frequency ranges, the second sub-emotional value of each of the multiple emotion information of the target audio is determined.

[0087] In one embodiment, the adjustment module includes:

[0088] A second adjustment unit is configured to adjust the parameters of the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information;

[0089] The audio parameters include at least dynamic range, reverberation effect, loudness and / or impact.

[0090] In one embodiment, the adjustment module includes:

[0091] The joy adjustment unit is configured to adjust the target audio according to the emotional value of the joy emotional information in one or more of the following adjustment methods:

[0092] enhancing a dynamic range of an audio signal within a first target frequency range in the target audio;

[0093] reducing the reverberation effect of the target audio;

[0094] The loudness of the target audio is enhanced.

[0095] In one embodiment, the adjustment module includes:

[0096] The complaint adjustment unit is configured to adjust the target audio according to one or more of the following adjustment methods based on the emotional value of the complaint emotional information:

[0097] enhancing a dynamic range of an audio signal within a second target frequency range of the target audio;

[0098] enhancing the reverberation effect of the target audio;

[0099] The loudness of the target audio is reduced.

[0100] In one embodiment, the adjustment module includes:

[0101] A dynamic adjustment unit is configured to adjust the target audio according to the emotional value of the dynamic emotional information in one or more of the following adjustment methods:

[0102] enhancing a dynamic range of an audio signal within a third target frequency range of the target audio;

[0103] enhancing the impact of an audio signal within a fourth target frequency range of the target audio;

[0104] The loudness of the target audio is enhanced.

[0105] In one embodiment, the adjustment module includes:

[0106] A soothing adjustment unit is configured to adjust the target audio according to one or more of the following adjustment methods based on the emotional value of the soothing emotional information:

[0107] reducing the dynamic range of the audio signal within a fifth target frequency range of the target audio;

[0108] increasing the reverberation delay of the target audio;

[0109] The loudness of the target audio is reduced.

[0110] In a third aspect, an embodiment of the present application further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps in the above-mentioned audio adjustment method are implemented.

[0111] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned audio adjustment method are implemented.

[0112] In a fifth aspect, embodiments of the present application further provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in the embodiments of the present application.

[0113] In summary, in the embodiments of the present application, the emotional values ​​of multiple emotional information of the target audio can be obtained; and the adjustment methods corresponding to the multiple emotional information can be obtained; because the emotional values ​​and adjustment methods are determined based on the multiple emotional information of the target audio, therefore, adjusting the target audio according to the emotional values ​​of the multiple emotional information and the adjustment methods corresponding to the multiple emotional information can effectively enhance the emotions conveyed by the target audio, thereby enhancing the user's emotional experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0114] In order to more clearly illustrate the technical solutions in this application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0115] Figure 1 This is a schematic diagram of the steps of the audio adjustment method provided in one embodiment of the present application;

[0116] Figure 2 This is a structural diagram of an audio adjustment device provided in one embodiment of the present application;

[0117] Figure 3 It is a structural diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0118] The following will be combined with the drawings in this application to clearly and completely describe the technical solutions in this application. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0119] In one embodiment, Figure 1As shown, an audio adjustment method is provided. Although a logical order is shown in the step diagram, in some cases, the steps shown or described may be performed in an order different from that shown in the accompanying drawings. Specifically, the audio adjustment method can be applied to a terminal or a server, wherein the terminal may include but is not limited to one or more of a smart phone, a tablet computer, a laptop computer, a desktop computer, and a car-mounted computer. The server may be a physical server or a cloud server that provides various cloud services. It is worth noting that the present application does not limit the number of terminals or servers. According to implementation needs, there can be any number of terminals or servers. For example, the server can be a single server or a server cluster consisting of multiple servers, etc.

[0120] It should be noted that the order of description of the following embodiments does not limit the priority order of the embodiments.

[0121] according to Figure 1 The audio adjustment method shown in FIG. 1 includes at least steps S110 to S130, which are described in detail as follows:

[0122] In step S110 , the emotion values ​​of the plurality of emotion information of the target audio are obtained.

[0123] The target audio can be any audio, pure audio, or audio in a video. The target audio can be a real-time audio stream or a complete audio file.

[0124] In one embodiment, the target audio may be an audio clip. Because the emotion conveyed by a longer audio clip typically changes over time, if the longer audio clip is adjusted using the same adjustment method, the emotion conveyed by the adjusted audio clip may be inaccurate. Therefore, in order to accurately enhance the emotion conveyed by the audio clip, the target audio may be an audio clip. The step size of the audio clip can be set according to actual needs, for example, it can be 100ms. When the audio is acquired, the audio can be segmented according to the preset step size to obtain multiple audio clips, and the audio clip is determined as the target audio.

[0125] The target audio can contain a variety of emotional signals, such as joy, sorrow, dynamism, and / or soothing. The emotional value represents the intensity of the emotional signal; the higher the emotional value, the stronger the emotional signal. A single audio signal can simultaneously contain one or more emotions.

[0126] The emotional value of the target audio's emotional information can be determined by analyzing its acoustic features, such as pitch, volume, speaking rate, and intonation. For example, a high-pitched voice with a fast speaking rate may convey joy or excitement, while a low-pitched voice with a slow speaking rate may convey sorrow or soothing emotions. By modeling and analyzing the target audio's acoustic features, the emotional value of each of the target audio's multiple emotional components can be determined.

[0127] In step S120, the adjustment methods corresponding to the multiple emotion information are obtained.

[0128] Adjustment methods corresponding to different emotional information can be preset. Based on multiple emotional information of the target audio, adjustment methods corresponding to each of the multiple emotional information can be obtained. For example, the adjustment method corresponding to the emotional information of joy may include, but is not limited to: enhancing the dynamic range of the audio signal in the mid-high frequency range of the target audio, reducing the reverberation effect of the target audio, and / or enhancing the loudness of the target audio.

[0129] In step S130, the target audio is adjusted according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information.

[0130] The intensity of the adjustment method corresponding to the emotional information can be determined based on the emotional value of the emotional information. The emotional value of the emotional information is proportional to the intensity of the adjustment method corresponding to the emotional information. The higher the emotional value, the greater the intensity of the corresponding adjustment method. For example, if the adjustment method is to enhance the loudness of the target audio, the higher the emotional value corresponding to the adjustment method, the greater the degree of enhancement of the target audio's loudness.

[0131] Therefore, the target audio can be adjusted according to the emotional values ​​of the multiple emotional information and the corresponding adjustment methods of the multiple emotional information.

[0132] In one embodiment, when there are multiple types of emotional information, the target audio may be adjusted according to the emotional values ​​of the multiple types of emotional information and the adjustment methods corresponding to the multiple types of emotional information. For example, when the emotional information of the target audio contains both joy and dynamism, the target audio may be adjusted according to the emotional value and adjustment method corresponding to the joyful emotional information, and the emotional value and adjustment method corresponding to the dynamic emotional information.

[0133] When the regulatory methods corresponding to multiple emotional information overlap, the regulatory effects can be superimposed; when the regulatory methods corresponding to multiple emotional information are contradictory, the regulatory effects can be offset.

[0134] In one embodiment, the emotional values ​​of multiple emotional information and the corresponding adjustment methods of the multiple emotional information can be combined to determine a final overall adjustment method, and the target audio can be adjusted according to the overall adjustment method. For example, if the loudness of the target audio needs to be increased by 20% based on the emotional information of joy, and the loudness of the target audio needs to be reduced by 10% based on the emotional information of sorrow, the overall adjustment method can be to increase the loudness of the target audio by 10%.

[0135] By adopting the technical solution of the embodiment of the present application, the emotional value of each of the multiple emotional information of the target audio can be obtained; and the adjustment method corresponding to each of the multiple emotional information can be obtained; because the emotional value and the adjustment method are determined based on the multiple emotional information of the target audio, therefore, adjusting the target audio according to the emotional value of each of the multiple emotional information and the adjustment method corresponding to each of the multiple emotional information can effectively enhance the emotions conveyed by the target audio, thereby enhancing the user's emotional experience.

[0136] On the basis of the above technical solution, as an embodiment, in order to obtain the emotional value of each of the multiple emotional information of the target audio, the sub-emotional value of each of the multiple emotional information of the target audio can be obtained respectively according to a plurality of emotional value acquisition methods. In the first emotional value acquisition method, the first sub-emotional value of each of the multiple emotional information of the target audio can be determined according to the sound source label of the target audio. In the second emotional value acquisition method, the target audio can be subjected to audio analysis to obtain the second sub-emotional value of each of the multiple emotional information of the target audio. In the third emotional value acquisition method, when the target audio is audio in a video, the video frame corresponding to the target audio can be obtained; the video frame can be subjected to image analysis to obtain the third sub-emotional value of each of the multiple emotional information of the target audio. In the fourth emotional value acquisition method, the emotional value set by the user can be obtained; and the fourth sub-emotional value of each of the multiple emotional information of the target audio can be determined according to the emotional value set by the user.

[0137] In the first method of obtaining the emotional value, the source tag of the target audio can be obtained. Optionally, the source tag of the target audio can be the source tag set by the uploader when uploading the source of the target audio. Optionally, the source tag of the target audio can be generated for the source of the target audio by the platform where the target audio is located. Optionally, the source tag of the target audio can be determined using natural language processing technology based on comments about the target audio.

[0138] In one embodiment, the sound source label is an emotion label, for example, the sound source label may be: joy, sorrow, dynamic and / or soothing.

[0139] In another embodiment, the sound source label is not an emotional label. For example, the sound source label may be "feelings of the protagonist," "fierce quarrel," or "cheerful background music." In this embodiment, the semantic information of the sound source label can be determined through natural language processing technology, and then the sound source label can be classified as an emotional label.

[0140] According to the emotional label corresponding to the sound source label, the amount of emotional information of the target audio can be determined; if the sum of the first sub-emotional values ​​of the multiple emotional information of the target audio is limited to 1, and the first sub-emotional values ​​of the multiple emotional information of the target audio are equal, then the first sub-emotional value of each of the multiple emotional information of the target audio can be determined according to the number of emotional information corresponding to the target audio.

[0141] In one embodiment, based on the emotion tag corresponding to the sound source tag, it can be determined that the target audio corresponds to four kinds of emotion information: joy, sorrow, dynamism, and soothing. Then, it can be determined that the first sub-emotion value of each is 0.25.

[0142] In another embodiment, based on the emotional label corresponding to the sound source label, it can be determined that the target audio corresponds to three emotional information: joy, sorrow, and soothing. Then, it can be determined that the first sub-emotional value of the joy, sorrow, and soothing emotional information is 0.33. Because there is no dynamic emotional information, the first sub-emotional value of the dynamic emotional information is 0.

[0143] In the second emotion value acquisition method, audio analysis can be performed on the target audio to obtain the second sub-emotion value of each of the multiple emotion information of the target audio.

[0144] In one embodiment, artificial intelligence (AI) or a neural network model may be used to perform audio analysis on the target audio to obtain the second sub-emotional value of each of the multiple emotion information of the target audio.

[0145] In another embodiment, the target audio can be converted into a frequency domain signal; the frequency domain signal is sampled to obtain multiple frequency points; multiple frequency ranges are obtained; and the second sub-emotional value of each of the multiple emotion information of the target audio is determined based on the ratio of the frequency points in the multiple frequency ranges.

[0146] The target audio signal can be converted from a time-domain signal to a frequency-domain signal using Fourier transform or other signal processing techniques. The frequency-domain signal is uniformly sampled to obtain multiple frequency points. The sampling step size can be set according to actual needs.

[0147] Acquire multiple pre-divided frequency ranges. Optionally, the multiple frequency ranges may include low frequency, medium frequency, and high frequency, wherein specific values ​​corresponding to each frequency range may refer to related technologies.

[0148] The ratio of the frequency points in the multiple frequency ranges is determined based on the frequencies corresponding to the multiple frequency points. The second sub-emotional value of each of the multiple emotion information of the target audio is determined based on the ratio of the frequency points in the multiple frequency ranges.

[0149] When the frequency points exceeding the first ratio are frequency points in the high frequency range, it can be determined that the target audio has joy emotional information. The second sub-emotional value of the joy emotional information can be determined based on the ratio of the frequency points in the high frequency range. The higher the ratio, the higher the second sub-emotional value of the joy emotional information. The first ratio can be set according to actual needs.

[0150] When the frequency points exceeding the second ratio are frequency points within the frequency range of the intermediate frequency, it can be determined that the target audio has soothing emotional information. The second sub-emotional value of the soothing emotional information can be determined based on the ratio of the frequency points within the frequency range of the intermediate frequency. The higher the ratio, the higher the second sub-emotional value of the soothing emotional information. The second ratio can be set according to actual needs.

[0151] When the frequency points exceeding the third ratio are in the low-frequency range, it can be determined that the target audio has sorrowful emotional information. The second sub-emotional value of the sorrowful emotional information can be determined based on the ratio of the frequency points in the low-frequency range. The higher the ratio, the higher the second sub-emotional value of the sorrowful emotional information. The third ratio can be set according to actual needs.

[0152] When the ratios of the frequency points in the low-frequency range and the frequency points in the high-frequency range both exceed the fourth ratio, it can be determined that the target audio has dynamic emotional information, and the second sub-emotional value of the dynamic emotional information can be determined based on the lower of the ratios of the frequency points in the low-frequency range and the frequency points in the high-frequency range. The higher the ratio corresponding to the lower one, the higher the second sub-emotional value of the dynamic emotional information. The fourth ratio can be set according to actual needs.

[0153] By converting the target audio into a frequency domain signal, the second sub-emotional value of each of the multiple emotional information of the target audio is determined according to the ratio of the frequency points in multiple frequency ranges. The frequency characteristics of the target audio are utilized. The frequency characteristics can be obtained through objective analysis technology, which is less affected by subjective factors. A relatively objective and accurate second sub-emotional value can be obtained, reducing the deviation of human judgment.

[0154] In the third method of obtaining the emotion value, when the target audio is the audio in the video, multiple video frames corresponding to the target audio can be obtained; image analysis is performed on the multiple video frames to obtain the third sub-emotion value of each of the multiple emotion information of the target audio.

[0155] In one embodiment, AI can be used to perform image analysis on video frames, and directly obtain the third sub-emotional value of each of the multiple emotional information of the target audio output by AI.

[0156] In another embodiment, the emotional information corresponding to each video frame can be determined by analyzing the face and action information of the video frame, as well as the color tone and other information of the video frame. Each video frame is determined to correspond to each piece of emotional information as one piece, and the sum of the third sub-emotional values ​​of the multiple pieces of emotional information corresponding to the multiple video frames of the target audio is 1, and the third sub-emotional values ​​of the multiple pieces of emotional information are equal. Then, the third sub-emotional values ​​of the multiple pieces of emotional information of the target audio can be determined based on the number of the multiple pieces of emotional information corresponding to the multiple video frames corresponding to the target audio.

[0157] For example, the target audio corresponds to two video frames, one of which corresponds to joyful emotional information and dynamic emotional information, and the other corresponds to joyful emotional information and soothing emotional information. There are a total of four pieces of emotional information, and the value corresponding to each piece of emotional information is 0.25. Joyful emotional information accounts for two pieces, and dynamic emotional information and soothing emotional information each account for one piece. The third emotional value of joyful emotional information is 0.5, the third emotional value of dynamic emotional information is 0.25, the third emotional value of soothing emotional information is 0.25, and the third emotional value of sad emotional information is 0.

[0158] In the fourth emotion value acquisition method, a user interface (UI) option may be provided for the user to set the fourth sub-emotion value of each of the plurality of emotion information.

[0159] By adopting the technical solution of the embodiment of the present application, multiple emotion value acquisition methods can be used to obtain multiple sub-emotional values ​​of the emotion information, thereby ensuring that the determined emotion value takes into account multiple factors and obtains a more accurate emotion value.

[0160] On the basis of the above technical solution, as an embodiment, after respectively obtaining the sub-emotional values ​​of multiple emotion information of the target audio according to multiple emotion value acquisition methods, the weights corresponding to the multiple emotion value acquisition methods can also be obtained; according to the weights corresponding to the multiple emotion value acquisition methods, the sub-emotional values ​​corresponding to the multiple emotion value acquisition methods are weightedly summed to obtain the emotion values ​​before processing of the multiple emotion information.

[0161] The weights corresponding to the various emotion value acquisition methods may be preset, and the weights corresponding to the various emotion value acquisition methods may be obtained. The weights corresponding to the emotion value acquisition methods may be determined based on the accuracy of the sub-emotion values ​​obtained by the emotion value acquisition methods.

[0162] In one embodiment, in order to enhance the user's emotional experience, the highest weight among the four emotion value acquisition methods can be set for the fourth emotion value acquisition method. Because the audio is actually adjusted, the lowest weight can be set for the third emotion value acquisition method. The sound source label can intuitively reflect the emotional information and is therefore more accurate. Therefore, the second highest weight among the four emotion value acquisition methods can be set for the first emotion value acquisition method. Therefore, the weight corresponding to the fourth emotion value acquisition method is higher than the weight corresponding to the first emotion value acquisition method, the weight corresponding to the first emotion value acquisition method is higher than the weight corresponding to the second emotion value acquisition method, and the weight corresponding to the second emotion value acquisition method is higher than the weight corresponding to the third emotion value acquisition method. Optionally, the weight corresponding to the first emotion value acquisition method is 0.3, the weight corresponding to the second emotion value acquisition method is 0.2, the weight corresponding to the third emotion value acquisition method is 0.1, and the weight corresponding to the fourth emotion value acquisition method is 0.4.

[0163] According to the weights corresponding to the multiple emotion value acquisition methods, the sub-emotion values ​​corresponding to the multiple emotion value acquisition methods are weighted and summed to obtain the emotion values ​​before processing of the multiple emotion information.

[0164] On the basis of the above embodiment, the emotional information of joy, sorrow, movement and relaxation are recorded as A, B, C and D respectively, A1 is recorded as the first sub-emotional value of the joy emotional information, and so on. The emotional value of the joy emotional information before processing = 0.3×A1+0.2×A2+0.1×A3+0.4×A4; the emotional value of the sorrow emotional information before processing = 0.3×B1+0.2×B2+0.1×B3+0.4×B4; and so on.

[0165] By adopting the technical solution of the embodiment of the present application, different weights are set for different methods of obtaining emotional values. The weight is determined based on the accuracy of the sub-emotional value obtained by the emotional value acquisition method. Therefore, by weighting the sub-emotional value according to the weight, a more accurate emotional value before processing can be obtained.

[0166] Based on the above technical solution, as an embodiment, the emotional values ​​of multiple emotional information of multiple historical moment audios corresponding to the target audio can be obtained; the weights of multiple historical moments can be obtained; and according to the weights of multiple historical moments and the emotional values ​​of multiple emotional information of multiple historical moment audios, the emotional values ​​of the multiple emotional information of the target audio before processing are smoothed to obtain the emotional values ​​of the multiple emotional information of the target audio.

[0167] When the target audio is an audio clip, the multiple historical moment audio clips corresponding to the target audio can be multiple audio clips preceding the target audio, where the historical moment audio clips and the target audio clips belong to the same sound source. For example, when the audio clip step size is 100ms, and the target audio clip is the audio clip from 300ms to 400ms of the first song, the historical moment audio clips can include: the audio clip from 0ms to 100ms of the song, the audio clip from 100ms to 200ms of the song, and the audio clip from 200ms to 300ms of the song.

[0168] Considering that the emotional information of audio clips at adjacent moments is similar, the emotional value of the emotional information of the target audio can be used to smooth the emotional value of the emotional information of the target audio. Multiple audio clips at the historical moments closest to the target audio can be selected and weights for each of these historical moments can be obtained. The greater the difference between the historical moment and the moment corresponding to the target audio, the lower the weight. Based on these weights, weighted emotional values ​​for each of the multiple emotional information clips at the multiple historical moments can be obtained.

[0169] For each piece of emotional information, the weighted emotional values ​​of the emotional information of audio at multiple historical moments are averaged with the emotional value of the emotional information of the target audio before processing, thereby achieving smoothing of the emotional value of the emotional information of the target audio before processing and obtaining the emotional value of the emotional information of the target audio.

[0170] For example, for joy emotional information, if the emotional values ​​of the joy emotional information of 9 historical moments are AT1, AT2,…, AT9 respectively, and the weights corresponding to the 9 historical moments are 0.1, 0.2,…, 0.9 respectively, and the emotional value of the joy emotional information of the target audio before processing is AT0, then the emotional value of the joy emotional information of the target audio = (AT1×0.1+AT2×0.2+…+AT9×0.9+AT0) / 10.

[0171] Optionally, the method for obtaining the emotional value of the audio at the historical moment may refer to the method for obtaining the emotional value of the target audio, wherein the emotional value of the audio at the historical moment may be an emotional value that has been or has not been smoothed.

[0172] By using the technical solutions of the embodiments of this application, the emotional values ​​of the target audio are smoothed using the emotional values ​​of the audio at historical moments. This allows the target audio's emotional values ​​to be corrected and supplemented with historical information, thereby improving the accuracy and reliability of the emotional values. This smoothing process makes the changes in the emotional values ​​more natural and fluid, reflecting the gradual transition of emotions.

[0173] Based on the above technical solution, as an embodiment, any sub-emotional value, or the emotion value before processing, can be directly determined as the emotion value of the emotion information of the target audio, thereby executing step S130.

[0174] Based on the above technical solution, as an embodiment, adjusting the target audio according to the emotional values ​​of multiple emotional information and the corresponding adjustment methods of multiple emotional information can include: classifying the multiple emotional information of the target audio; obtaining the difference between multiple emotional values ​​in each category; obtaining the difference threshold; when the difference between multiple emotional information in the same category is not greater than the difference threshold, adjusting the target audio according to the emotional values ​​of multiple emotional information in the category and the corresponding adjustment methods of multiple emotional information in the category.

[0175] Considering the actual situation, joyful and dynamic emotional information can be grouped into one category, while sorrowful and soothing emotional information can be grouped into another. It's understandable that the emotional values ​​of emotional information within the same category are generally close. Therefore, if the emotional values ​​of emotional information within the same category differ significantly, the emotional value of that category of emotional information may be incorrect, and the target audio will not be adjusted based on that category of emotional information.

[0176] Specifically, the difference between the emotion values ​​of multiple emotion information of the same category can be obtained. A difference threshold is obtained, and the difference threshold can be set according to actual needs. When the difference between the emotion values ​​of multiple emotion information of the same category is greater than the difference threshold, it is considered that the difference in the emotion value of the emotion information of this category is too large. When the difference between the emotion values ​​of multiple emotion information of the same category is not greater than the difference threshold, it is considered that there is no abnormality in the emotion value of the emotion information of this category, and therefore the target audio can be adjusted based on the emotion information of this category.

[0177] For example, if the difference threshold is not 0.4, the emotional value of the joyful emotional information of the target audio is 0.3, the emotional value of the dynamic emotional information is 0.4, the emotional value of the resentful emotional information is 0.7, and the emotional value of the soothing emotional information is 0.1; if the joyful emotional information and the dynamic emotional information are of the same category and the difference between their emotional values ​​is not greater than the difference threshold, the target audio can be adjusted based on the joyful emotional information and the dynamic emotional information; if the resentful emotional information and the soothing emotional information are of the same category and the difference between their emotional values ​​is greater than the difference threshold, the target audio will not be adjusted based on the resentful emotional information and the soothing emotional information.

[0178] By adopting the technical solution of the embodiment of the present application, by classifying the emotional information and judging whether the difference between the emotional values ​​of the emotional information in the same category is greater than the difference threshold, it is possible to judge whether the acquired emotional value is abnormal, and thus adjust the target audio according to the correct emotional value, so that the adjusted target audio can be accurately emotionally enhanced.

[0179] In another embodiment, when the difference between the emotion values ​​of multiple emotion information of any category is greater than the difference threshold, the target audio is not adjusted according to the emotion information of other categories even if the difference between the emotion values ​​of multiple emotion information of other categories is not greater than the difference threshold.

[0180] Based on the above technical solution, as an embodiment, adjusting the target audio according to the emotional values ​​of multiple emotional information and the adjustment methods corresponding to the multiple emotional information can include: adjusting the parameters of the target audio according to the emotional values ​​of multiple emotional information and the adjustment methods corresponding to the multiple emotional information; wherein the audio parameters include at least: dynamic range, reverberation effect, loudness and / or frequency.

[0181] Adjusting the target audio may include adjusting the dynamic range, reverberation effect, loudness, and / or impact of the target audio.

[0182] Enhancing the dynamic range can make the soft parts of the audio softer and the strong parts stronger, highlighting the detailed changes in the audio. For example, in a piece of classical music, subtle changes in volume can more delicately express the emotional ups and downs that the composer wants to convey, allowing the audience to feel the emotional connotation of the music more deeply. Reducing the dynamic range can make the overall volume of the audio more uniform, which is suitable for occasions that require a stable sound environment, such as the narration part in a movie. It allows the audience to receive information more clearly, while also creating a calm emotional atmosphere.

[0183] Appropriate reverb effects can simulate different spatial environments, creating a realistic sense of space and atmosphere for audio. For example, when recording a lyrical song, adding an appropriate amount of room reverb can make the vocals sound warmer and softer, as if performed in a cozy small room, enhancing the song's emotional appeal. When recording epic music, using reverb effects from larger spaces, such as church reverb or canyon reverb, can create a grand and solemn atmosphere, giving the audience a strong emotional impact and making them feel as if they are immersed in a specific scene.

[0184] Properly controlling audio loudness not only makes sound clearer and more audible, but also directly impacts the listener's emotional experience. Increasing the loudness of an audio track can enhance its expressiveness and appeal, making the audience more excited and agitated during climaxes. Reducing the loudness, on the other hand, can create a quieter, more intimate atmosphere, allowing listeners to focus more on the details and emotional expressions within the audio. For example, in quiet meditation music, lower loudness can help people relax and immerse themselves in a tranquil emotional atmosphere.

[0185] Adjusting the impact of audio can make the emotions expressed more vivid and intense. For example, in a sad piece of music, by appropriately increasing the volume and enhancing the impact of the bass, the listener can feel the sadness more deeply, as if they are drawn into a richer atmosphere of sorrow.

[0186] Based on the above technical solution, as an embodiment, according to the emotional value of the joy emotional information, the target audio may be adjusted in one or more of the following adjustment methods:

[0187] enhancing the dynamic range of an audio signal within a first target frequency range in the target audio;

[0188] Reduce the reverberation effect of the target audio;

[0189] Enhance the loudness of the target audio.

[0190] Based on the emotional value of the sorrowful emotional information, the target audio can be adjusted in one or more of the following ways:

[0191] enhancing a dynamic range of an audio signal within a second target frequency range of the target audio;

[0192] Enhance the reverberation effect of the target audio;

[0193] Reduce the loudness of the target audio.

[0194] Based on the emotional value of the dynamic emotional information, the target audio can be adjusted in one or more of the following ways:

[0195] enhancing a dynamic range of an audio signal within a third target frequency range of the target audio;

[0196] enhancing the impact of an audio signal within a fourth target frequency range of the target audio;

[0197] Enhance the loudness of the target audio.

[0198] Based on the emotional value of the soothing emotional information, the target audio may be adjusted in one or more of the following ways:

[0199] reducing the dynamic range of the audio signal within a fifth target frequency range of the target audio;

[0200] Increase the reverb delay of the target audio;

[0201] Reduce the loudness of the target audio.

[0202] In a possible implementation of an embodiment of the present application, the frequency range can be determined with reference to the frequency division in the music system, (40Hz-80Hz) can be determined as low frequency, (80Hz-160Hz) can be determined as medium-low frequency, (500Hz-2KHz) can be determined as medium frequency, (2KHz-4KHz) can be determined as medium-high frequency (2KHz-4KHz), and (5K-10KHz) can be determined as high frequency.

[0203] Among them, the first target frequency range can be medium and high frequency; the second target frequency range is medium and low frequency, the third target frequency range includes medium and high frequency and high frequency, the fourth target frequency range is low frequency, and the fifth target frequency range is low frequency and medium frequency.

[0204] In an embodiment of the present application, taking into account the impact of the dynamic range, reverberation effect, loudness and / or frequency of the audio on emotions mentioned above, as well as the frequency range on which music can actually convey emotions such as joy, sorrow, dynamism and soothing to the user, the above-mentioned multiple adjustment methods are provided to enhance the emotions of the target audio through multiple adjustment methods and improve the user's emotional experience.

[0205] To facilitate better implementation of the audio adjustment method of the present application, the present application also provides an audio adjustment device based on the above audio adjustment method. The meanings of the terms are the same as those in the above audio adjustment method, and the specific implementation details can be referred to the description in the method embodiment.

[0206] See also Figure 2 , Figure 2 : is a structural diagram of an audio adjustment device provided in an embodiment of the present application, wherein the audio adjustment device includes:

[0207] A first acquisition module 201 is used to obtain the emotion value of each of the plurality of emotion information of the target audio;

[0208] The second acquisition module 202 is used to obtain the adjustment methods corresponding to the plurality of emotion information;

[0209] The adjustment module 203 is used to adjust the target audio according to the emotional values ​​of the multiple emotional information and the adjustment methods corresponding to the multiple emotional information.

[0210] In one embodiment, the adjustment module 203 includes:

[0211] A classification unit, configured to classify multiple emotional information of the target audio;

[0212] A difference acquisition unit, used to obtain the difference between multiple emotion values ​​in each category;

[0213] A threshold value obtaining unit, used for obtaining a difference threshold value;

[0214] The first adjustment unit is used to adjust the target audio according to the emotional values ​​of the multiple emotional information in the same category and the adjustment methods corresponding to the multiple emotional information in the category when the difference between the multiple emotional information in the same category is not greater than the difference threshold.

[0215] In one embodiment, the first acquisition module 201 includes:

[0216] a pre-processing emotion value obtaining unit, configured to obtain a pre-processing emotion value of each of the plurality of emotion information of the target audio;

[0217] A historical emotion value acquisition unit, configured to acquire emotion values ​​of a plurality of emotion information of audios at a plurality of historical moments corresponding to the target audio;

[0218] A historical weight acquisition unit, used to obtain the weights of multiple historical moments;

[0219] A smoothing processing unit is used to smooth the pre-processed emotional values ​​of the multiple emotional information of the target audio according to the respective weights of the multiple historical moments and the respective emotional values ​​of the multiple emotional information of the audio at the multiple historical moments, so as to obtain the respective emotional values ​​of the multiple emotional information of the target audio.

[0220] In one embodiment, the pre-processing emotion value acquisition unit includes:

[0221] The sub-emotion value acquisition sub-unit is used to respectively acquire the sub-emotion values ​​of the plurality of emotion information of the target audio according to a plurality of emotion value acquisition methods;

[0222] A weight acquisition subunit, used to obtain weights corresponding to the plurality of sentiment value acquisition methods;

[0223] The pre-processing emotion value acquisition subunit is used to perform weighted summation on the sub-emotional values ​​corresponding to the multiple emotion value acquisition methods according to the weights corresponding to the multiple emotion value acquisition methods, so as to obtain the pre-processing emotion value of each of the multiple emotion information.

[0224] In one embodiment, the step of obtaining the sub-emotional values ​​of the plurality of emotion information of the target audio in accordance with the plurality of emotion value obtaining methods comprises at least one of the following steps:

[0225] Determining, according to the sound source label of the target audio, a first sub-emotion value of each of the plurality of emotion information of the target audio;

[0226] Performing audio analysis on the target audio to obtain a second sub-emotion value of each of the plurality of emotion information of the target audio;

[0227] Obtaining a video frame corresponding to the target audio;

[0228] Performing image analysis on the video frame to obtain a third sub-emotion value of each of the multiple emotion information of the target audio;

[0229] Get the emotion value set by the user;

[0230] According to the emotion value set by the user, the fourth sub-emotion value of each of the multiple emotion information of the target audio is determined.

[0231] In one embodiment, performing audio analysis on the target audio to obtain the second sub-emotional value of each of the plurality of emotion information of the target audio includes:

[0232] Converting the target audio into a frequency domain signal;

[0233] Sampling the frequency domain signal to obtain multiple frequency points;

[0234] Get multiple frequency ranges;

[0235] According to the ratio of the frequency points in the multiple frequency ranges, the second sub-emotional value of each of the multiple emotion information of the target audio is determined.

[0236] In one embodiment, the adjustment module 203 includes:

[0237] A second adjustment unit is configured to adjust the parameters of the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information;

[0238] The audio parameters include at least dynamic range, reverberation effect, loudness and / or impact.

[0239] In one embodiment, the adjustment module 203 includes:

[0240] The joy adjustment unit is configured to adjust the target audio according to the emotional value of the joy emotional information in one or more of the following adjustment methods:

[0241] enhancing a dynamic range of an audio signal within a first target frequency range in the target audio;

[0242] reducing the reverberation effect of the target audio;

[0243] The loudness of the target audio is enhanced.

[0244] In one embodiment, the adjustment module 203 includes:

[0245] The complaint adjustment unit is configured to adjust the target audio according to one or more of the following adjustment methods based on the emotional value of the complaint emotional information:

[0246] enhancing a dynamic range of an audio signal within a second target frequency range of the target audio;

[0247] enhancing the reverberation effect of the target audio;

[0248] The loudness of the target audio is reduced.

[0249] In one embodiment, the adjustment module 203 includes:

[0250] A dynamic adjustment unit is configured to adjust the target audio according to the emotional value of the dynamic emotional information in one or more of the following adjustment methods:

[0251] enhancing a dynamic range of an audio signal within a third target frequency range of the target audio;

[0252] enhancing the impact of an audio signal within a fourth target frequency range of the target audio;

[0253] The loudness of the target audio is enhanced.

[0254] In one embodiment, the adjustment module 203 includes:

[0255] A soothing adjustment unit is configured to adjust the target audio according to one or more of the following adjustment methods based on the emotional value of the soothing emotional information:

[0256] reducing the dynamic range of the audio signal within a fifth target frequency range of the target audio;

[0257] increasing the reverberation delay of the target audio;

[0258] The loudness of the target audio is reduced.

[0259] By adopting the technical solution of the embodiment of the present application, the emotional value of each of the multiple emotional information of the target audio can be obtained; and the adjustment method corresponding to each of the multiple emotional information can be obtained; because the emotional value and the adjustment method are determined based on the multiple emotional information of the target audio, therefore, adjusting the target audio according to the emotional value of each of the multiple emotional information and the adjustment method corresponding to each of the multiple emotional information can effectively enhance the emotions conveyed by the target audio, thereby enhancing the user's emotional experience.

[0260] For the specific definition of the audio adjustment device, please refer to the definition of the audio adjustment method above, which will not be repeated here. The various modules in the above-mentioned audio adjustment device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0261] In addition, the present application also provides an electronic device, such as Figure 3 As shown, it shows a schematic diagram of the structure of the electronic device involved in this application, specifically:

[0262] The electronic device may include one or more processing core processors 301 and one or more computer readable storage media memories 302 and other components. Those skilled in the art will understand that Figure 3 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0263] The processor 301 is the control center of the electronic device. It connects all parts of the electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 302 and accessing data stored in the memory 302, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 301.

[0264] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 302 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.

[0265] In one embodiment, the electronic device further includes a power supply 303 for supplying power to various components. Preferably, the power supply 303 can be logically connected to the processor 301 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 303 can also include one or more DC or AC power supplies, a recharging system, a power supply device debugging circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0266] In one embodiment, the electronic device may further include an input unit 304, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0267] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically, in this embodiment, the processor 301 in the electronic device loads the executable files corresponding to one or more application processes into the memory 302 according to the following instructions, and the processor 301 runs the application stored in the memory 302, thereby implementing the steps of any of the audio adjustment methods provided in the embodiments of the present application.

[0268] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the electronic device to which the solution of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0269] In one embodiment, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of the present application is implemented.

[0270] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.

[0271] In some embodiments, a computer program product is also proposed, including a computer program or instructions, which implements the method described in any embodiment of the present application when executed by a processor.

[0272] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0273] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0274] To this end, the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program can be loaded by a processor to execute the steps in any audio adjustment method provided in the present application.

[0275] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0276] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0277] Since the instructions stored in the computer-readable storage medium can execute the steps in any audio adjustment method provided in this application, the beneficial effects that can be achieved by any audio adjustment method provided in this application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0278] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements that are inherent to such process, method, article, or terminal device. In the absence of further restrictions, an element defined by the phrase "comprises a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0279] The above is a detailed introduction to the audio adjustment method, device, electronic device and computer-readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. An audio adjustment method, characterized in that: The method comprises: Obtaining the emotional values ​​of multiple emotional information of the target audio; Obtaining adjustment methods corresponding to the plurality of emotion information; The target audio is adjusted according to the emotional values ​​of the multiple emotional information and the adjustment methods corresponding to the multiple emotional information.

2. The method according to claim 1, characterized in that The adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes: Classifying multiple emotional information of the target audio; Get the difference between multiple sentiment values ​​in each category; Get the difference threshold; When the difference between multiple emotion information in the same category is not greater than the difference threshold, the target audio is adjusted according to the emotion values ​​of the multiple emotion information in the category and the adjustment methods corresponding to the multiple emotion information in the category.

3. The method according to claim 1, characterized in that The acquiring of the emotion values ​​of the plurality of emotion information of the target audio includes: Obtaining a pre-processing emotion value of each of the plurality of emotion information of the target audio; Obtaining respective emotion values ​​of a plurality of emotion information of audios at a plurality of historical moments corresponding to the target audio; Get the weights of multiple historical moments; According to the weights of each of the multiple historical moments and the emotional values ​​of each of the multiple emotional information of the audio at the multiple historical moments, the emotional values ​​of each of the multiple emotional information of the target audio before processing are smoothed to obtain the emotional values ​​of each of the multiple emotional information of the target audio.

4. The method according to claim 3, characterized in that The obtaining of the emotion value before processing of each of the plurality of emotion information of the target audio includes: According to multiple emotion value acquisition methods, respectively obtain the sub-emotional values ​​of the multiple emotion information of the target audio; Obtaining weights corresponding to the plurality of sentiment value acquisition methods; According to the weights corresponding to the multiple emotion value acquisition methods, the sub-emotion values ​​corresponding to the multiple emotion value acquisition methods are weighted and summed to obtain the emotion values ​​before processing of the multiple emotion information.

5. The method according to claim 4, characterized in that The step of respectively obtaining the sub-emotional values ​​of the plurality of emotion information of the target audio in accordance with the plurality of emotion value obtaining methods includes at least one of the following steps: Determining, according to the sound source label of the target audio, a first sub-emotion value of each of the plurality of emotion information of the target audio; Performing audio analysis on the target audio to obtain a second sub-emotion value of each of the plurality of emotion information of the target audio; Obtaining a video frame corresponding to the target audio; Performing image analysis on the video frame to obtain a third sub-emotion value of each of the multiple emotion information of the target audio; Get the emotion value set by the user; According to the emotion value set by the user, the fourth sub-emotion value of each of the multiple emotion information of the target audio is determined.

6. The method according to claim 5, characterized in that The performing audio analysis on the target audio to obtain respective second sub-emotional values ​​of the plurality of emotion information of the target audio includes: Converting the target audio into a frequency domain signal; Sampling the frequency domain signal to obtain multiple frequency points; Get multiple frequency ranges; According to the ratio of the frequency points in the multiple frequency ranges, the second sub-emotional value of each of the multiple emotion information of the target audio is determined.

7. The method according to claim 1, characterized in that The adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes: Adjusting the parameters of the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information; The audio parameters include at least dynamic range, reverberation effect, loudness and / or impact.

8. The method according to claim 1, characterized in that The adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes: According to the emotional value of the joy emotional information, the target audio is adjusted according to one or more of the following adjustment methods: enhancing a dynamic range of an audio signal within a first target frequency range in the target audio; reducing the reverberation effect of the target audio; The loudness of the target audio is enhanced.

9. The method according to claim 1, characterized in that The adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes: According to the emotional value of the sorrowful emotional information, the target audio is adjusted according to one or more of the following adjustment methods: enhancing a dynamic range of an audio signal within a second target frequency range of the target audio; enhancing the reverberation effect of the target audio; The loudness of the target audio is reduced.

10. The method according to claim 1, characterized in that The adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes: According to the emotional value of the dynamic emotional information, the target audio is adjusted in one or more of the following adjustment methods: enhancing a dynamic range of an audio signal within a third target frequency range of the target audio; enhancing the impact of an audio signal within a fourth target frequency range of the target audio; The loudness of the target audio is enhanced.

11. The method according to claim 1, wherein The adjusting the target audio according to the emotion values ​​of the plurality of emotion information and the adjustment methods corresponding to the plurality of emotion information includes: According to the emotional value of the soothing emotional information, the target audio is adjusted in one or more of the following adjustment methods: reducing the dynamic range of the audio signal within a fifth target frequency range of the target audio; Adding a reverberation delay to the target audio; The loudness of the target audio is reduced.

12. An audio adjustment device, characterized in that: The device comprises: A first acquisition module is used to obtain the emotion value of each of the multiple emotion information of the target audio; A second acquisition module is used to obtain the adjustment methods corresponding to the plurality of emotion information; The adjustment module is used to adjust the target audio according to the emotional values ​​of the multiple emotional information and the adjustment methods corresponding to the multiple emotional information.

13. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the audio adjustment method according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the audio adjustment method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Emotion classification method and device, server and storage medium

    CN109325124A

  • Speech emotion recognition method and system based on frequency domain F ratio analysis

    CN118918925A

  • Multi-modal classroom emotion recognition method and system based on modal adaptive learning

    CN119418725A

  • Method and apparatus for emotion change of music

    KR101823274B1