Method and device for dynamically adjusting animation sound based on character emotion
By using a dynamic sound adjustment method for animation based on character emotions, core emotional information is extracted and evaluated, a comprehensive emotional complexity model is constructed, and sound parameters are adaptively configured. This solves the automation and adaptation problems of traditional animation sound adjustment, and improves the artistic expression and production efficiency of animation works.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INSTITUTE OF GRAPHIC COMMUNICATION
- Filing Date
- 2025-11-04
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional methods of adjusting sound in animation rely on human experience, have a low degree of automation, and are difficult to accurately adapt to the complex dynamic emotional changes of characters, resulting in a stiff combination of sound and image and weakening the artistic appeal of the work.
The animation sound dynamic adjustment method based on character emotions extracts the core emotion type and intensity, evaluates the complexity of emotion mixing and change, constructs a comprehensive emotion complexity model, adaptively configures the sound parameter optimization process, and generates the optimal sound adjustment scheme.
It enables automated and precise adjustment of sound parameters for target characters in animated segments, enhancing the artistic expressiveness and production efficiency of audio-visual integration, and overcoming the drawbacks of traditional methods such as lag and rigid adaptation.
Smart Images

Figure CN121260188B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of voice sensing, in particular to an animation sound dynamic adjustment method and device based on character emotion. BACKGROUND
[0002] In animation production, sound is one of the core elements of artistic expression, and its organic combination with the picture plays an irreplaceable role in shaping the character personality, promoting the narrative development and setting the scene atmosphere.
[0003] The traditional animation sound adjustment method largely depends on the manual experience of sound engineers. Sound engineers need to manually set volume, equalization, reverb and other parameters according to the script and their subjective understanding of the picture. This method is not only low in efficiency, but also difficult to achieve precise, continuous and adaptive sound effect matching when facing complex animation segments with rapid conversion of character emotions or interweaving of multiple emotions. Fixed sound parameter templates are difficult to flexibly respond to dynamic emotional needs, which may lead to rigid combination of sound and picture and weaken the artistic appeal of the work. SUMMARY
[0004] The present application provides an animation sound dynamic adjustment method and device based on character emotion to solve the technical problems of animation sound adjustment relying on manual experience, low automation and difficulty in accurately adapting to complex dynamic emotional changes of characters in the prior art.
[0005] The technical solution of the present application to solve the above technical problems is as follows:
[0006] In a first aspect, the present application provides an animation sound dynamic adjustment method based on character emotion, comprising:
[0007] extracting core emotion types and core emotion intensities based on emotion attribute information sequences of target characters in a to-be-processed animation segment, obtaining core emotion type sequences and core emotion intensity sequences;
[0008] evaluating emotion mixing complexity and emotion change complexity according to the core emotion type sequences and the core emotion intensity sequences;
[0009] weighting and fusing the emotion mixing complexity and the emotion change complexity to obtain a comprehensive emotion complexity;
[0010] setting adaptation evaluation model selection parameters and adaptation optimization process parameters of sound parameter optimization according to the comprehensive emotion complexity, and generating a sound parameter optimization scheme;
[0011] performing sound parameter optimization according to the sound parameter optimization scheme, outputting an optimal sound adjustment scheme, and adjusting sound parameters of target characters in the to-be-processed animation segment according to the optimal sound adjustment scheme.
[0012] In a second aspect, the present application provides an animation sound dynamic adjustment device based on character emotion, comprising:
[0013] An emotion feature extraction module is configured to extract a core emotion type and a core emotion intensity based on an emotion attribute information sequence of a target character in a to-be-processed animation segment, and obtain a core emotion type sequence and a core emotion intensity sequence.
[0014] An emotion complexity evaluation module is configured to evaluate an emotion mixing complexity and an emotion change complexity according to the core emotion type sequence and the core emotion intensity sequence.
[0015] A comprehensive complexity calculation module is configured to obtain a comprehensive emotion complexity based on a weighted fusion of the emotion mixing complexity and the emotion change complexity.
[0016] An optimization scheme generation module is configured to set an adaptive evaluation model selection parameter and an adaptive optimization process parameter of sound parameter optimization according to the comprehensive emotion complexity, and generate a sound parameter optimization scheme.
[0017] A sound parameter adjustment module is configured to perform sound parameter optimization according to the sound parameter optimization scheme, output an optimal sound adjustment scheme, and adjust the sound parameter of the target character in the to-be-processed animation segment according to the optimal sound adjustment scheme.
[0018] The present application has the following advantages:
[0019] Compared with the prior art, the present application firstly quantitatively analyzes the sequence information of the emotion attribute of the animation character, accurately evaluates the emotion mixing complexity and the emotion change complexity, and thereby constructs a comprehensive emotion complexity model that can fully reflect the dynamic characteristics of the emotion. Secondly, according to the comprehensive emotion complexity, the key parameters in the sound parameter optimization process are adaptively configured, including the number of evaluation models and the iteration depth of the optimization process, ensuring the matching degree of the optimization strategy and the emotion scene. Thirdly, through the constructed sound parameter optimization scheme, efficient optimization is performed in the preset parameter space, and an optimal sound adjustment scheme highly consistent with the emotion state of the character can be automatically generated. Finally, the automatic and fine dynamic adjustment of the sound parameter of the target character in the animation segment is realized, effectively overcoming the disadvantages of the traditional method, such as dependence on artificial experience, response lag and rigid adaptation, and improving the artistic expression and production efficiency of the combination of sound and picture. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 A flowchart of an animation sound dynamic adjustment method based on character emotion provided by the present application is shown.
[0021] Figure 2 A structural diagram of an animation sound dynamic adjustment device based on character emotion provided by the present application is shown.
[0022] In the drawings, the components represented by the respective reference numerals are as follows:
[0023] The emotion feature extraction module 11, the emotion complexity evaluation module 12, the comprehensive complexity calculation module 13, the optimization scheme generation module 14, and the sound parameter adjustment module 15. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0025] In the description of the present application, the terms "first", "second" are used only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first", "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.
[0026] In the description of the present application, the term "for example" is used to indicate "as an example, illustration or explanation". Any embodiment described as "for example" in the present application is not necessarily interpreted as more preferred or more advantageous than other embodiments. The following description is given in order to enable any person skilled in the art to implement and use the present application. In the following description, details are listed for the purpose of explanation. It should be understood that those skilled in the art can realize the present application without using these specific details. In other examples, well-known structures and processes will not be described in detail to avoid unnecessary details making the description of the present application obscure. Therefore, the present application is not intended to be limited to the shown embodiments, but is consistent with the broadest scope in accordance with the principles and characteristics disclosed.
[0027] In one embodiment, as shown in Figure 1 The present application provides an animation sound dynamic adjustment method based on role emotion, which comprises the following steps:
[0028] S10: Based on the emotion attribute information sequence of the target role in the to-be-processed animation segment, the core emotion type and the core emotion intensity are extracted, and the core emotion type sequence and the core emotion intensity sequence are obtained;
[0029] Firstly, based on the emotion attribute information sequence of the target role in the to-be-processed animation segment, the core emotion type and the core emotion intensity are extracted, and the core emotion type sequence and the core emotion intensity sequence are obtained. Among them, the to-be-processed animation segment refers to an animation video paragraph containing complete or partial narrative content which needs to be adjusted in sound parameter; the target role refers to one or more specific animation roles whose emotional state needs to be analyzed and the sound parameter corresponding to the target role needs to be adjusted.
[0030] Based on the emotion attribute information sequence, the core emotion type and the core emotion intensity are extracted, wherein the emotion attribute information refers to a structured data set containing emotion categories and their corresponding confidence or quantitative values obtained by analyzing each moment or key time point of the target role on the time axis of the animation segment through emotion recognition technology; the core emotion type and the core emotion intensity refer to one or more emotion categories and their corresponding quantitative intensity values that best represent the main emotional state of the current role selected from the emotion attribute information of each moment or key time point according to a preset rule.
[0031] Since the emotional state of the animation role is complex and dynamic, by extracting the core emotion type, the most representative dominant emotion can be selected from the possible mixed emotions, ensuring that the subsequent sound adjustment has a clear emotional direction. At the same time, extracting the core emotion intensity can quantify the intensity of the dominant emotion, providing a quantitative basis for the adjustment range of the sound parameter.
[0032] Specifically, based on the emotion attribute information sequence of the target role in the to-be-processed animation segment, the core emotion type and the core emotion intensity are extracted, and the core emotion type sequence and the core emotion intensity sequence are obtained, including:
[0033] Configure an emotion type extraction strategy, wherein when the emotion attribute information is a single emotion type, the emotion type is determined as the core emotion type, and when the emotion attribute information is a mixture of multiple emotion types, the intensity values of each emotion type are sorted, and the top N emotion types with the highest intensity are determined as the core emotion types, N is greater than or equal to 2;
[0034] According to the emotion type extraction strategy, the emotion attribute information sequence of the target role in the to-be-processed animation segment is extracted, and the core emotion type sequence and the core emotion intensity sequence are obtained, wherein the emotion type at least includes joy, sadness, anger, fear, disgust, surprise, tension, calmness, love and contempt.
[0035] Firstly, the emotion type extraction strategy is configured. The emotion type extraction strategy defines the rules for filtering representative emotion types from the original emotion attribute information. Specifically, when the emotion attribute information of the target character in an animation segment is represented by a single emotion type, the single emotion type is directly determined as the core emotion type. When the emotion attribute information is represented by a complex state of mixed emotion types, the intensity values of each emotion type are sorted in descending order, and the top N emotion types with the highest intensity ranking are selected as the core emotion types, where N is an integer greater than or equal to 2. This emotion type extraction strategy can ensure that the extraction of core emotion types can reflect the dominant emotion and cover the key components in the mixed emotion state.
[0036] Secondly, the specific information extraction operation is performed according to the configured emotion type extraction strategy. The emotion attribute information sequence of the target character in the animation segment to be processed is analyzed frame by frame or time period by time period, and the core emotion type at each time is identified and recorded according to the above strategy, thereby forming a core emotion type sequence. At the same time, the emotion intensity values corresponding to the core emotion types are extracted and recorded, forming a core emotion intensity sequence. Among them, the emotion type is a predefined set, which at least includes basic emotion categories such as joy, sadness, anger, fear, disgust, surprise, tension, calm, love and contempt.
[0037] Among them, the quantification of emotion intensity is realized through a standardized data processing process. First, the facial expression, speech prosody and body movement of the target character are analyzed by the emotion recognition algorithm, and the initial confidence score of each basic emotion type in the range of 0-1 is generated. Second, the initial confidence of all existing emotion types in the same analysis unit is normalized, and the initial confidence of each emotion type is divided by the sum of the initial confidence of all emotion types to obtain the relative intensity proportion of each emotion type. Further, according to the preset emotion type extraction strategy, the core emotion types to be retained are selected from all emotion types. Finally, the relative intensity proportion corresponding to the selected core emotion type is determined as its final emotion intensity value.
[0038] For example, at a certain analysis time, the initial confidence output by the emotion recognition algorithm shows that anger is 0.72, sadness is 0.25, and fear is 0.03. The sum of the initial confidence of the three emotions is 1.00. After normalization calculation, the relative intensity proportion of anger is 0.72, that of sadness is 0.25, and that of fear is 0.03. When the top two core emotion types are set to be retained, the emotion intensity of anger is determined to be 0.72, and the emotion intensity of sadness is determined to be 0.25.
[0039] Through the above steps, subjective and continuous emotional fluctuations are converted into objective and discrete sequence data, and finally obtained core emotion type sequence and core emotion intensity sequence jointly constitute the basis for evaluating emotional complexity, which can provide standardized input for subsequent calculation.
[0040] S20: obtaining emotion mixing complexity and emotion change complexity according to the core emotion type sequence and the core emotion intensity sequence;
[0041] After obtaining the core emotion type sequence and the core emotion intensity sequence, further quantitative analysis and comprehensive evaluation are needed to obtain the emotion mixing complexity and the emotion change complexity.
[0042] The emotion mixing complexity is a static characteristic index, which is used to quantitatively represent the internal composition complexity of the target character's emotional state in the time range of the to-be-processed animation segment. The emotion mixing complexity comprehensively reflects the number diversity of the core emotion types experienced by the character, the inherent adjustment difficulty level of each emotion type, and the distribution relationship of the intensity of these emotion types. The higher the emotion mixing complexity, the more complex the character's emotional composition, and the more high-difficulty adjustment emotions are intertwined and coexist with significant intensity.
[0043] In addition, the emotion change complexity is a dynamic characteristic index, which is used to quantitatively represent the intensity and instability of the target character's emotional state fluctuation over time in the time range of the to-be-processed animation segment. The emotion change complexity comprehensively reflects the fluctuation of the number of core emotion types in the time sequence and the fluctuation of the comprehensive emotion complexity in the time sequence. The higher the emotion change complexity, the more unstable the character's emotional state, the more frequent the emotional transition, and the greater the fluctuation amplitude.
[0044] Specifically, obtaining emotion mixing complexity and emotion change complexity according to the core emotion type sequence and the core emotion intensity sequence includes:
[0045] Setting a plurality of initial emotion complexities based on a plurality of emotion types, wherein the initial emotion complexity is set based on the historical sound adjustment difficulty corresponding to the emotion type;
[0046] Extracting the number of types according to the core emotion type sequence to obtain a core emotion type number sequence, and calculating the mean value of the core emotion type number sequence to obtain a core emotion type number mean value;
[0047] According to the plurality of initial emotion complexities, performing weighted fusion according to the core emotion intensity sequence to obtain a core emotion complexity sequence, and calculating the mean value of the core emotion complexity sequence to obtain a core emotion complexity mean value;
[0048] The product of the core emotion type quantity average and the core emotion complexity average is set as emotion mixing complexity.
[0049] An emotion change complexity is evaluated according to the core emotion type quantity sequence and the core emotion complexity sequence.
[0050] First, an initial emotion complexity is set. A number of initial emotion complexity values are respectively assigned to a number of predefined emotion types. The setting of the initial emotion complexity values is based on the sound parameter adjustment difficulty corresponding to each emotion type in historical data. An emotion type with a higher adjustment difficulty is assigned a higher initial emotion complexity value, and vice versa.
[0051] For example, the initial emotion complexity values are set as: anger 0.9, sadness 0.8, fear 0.7, disgust 0.6, surprise 0.5, tension 0.7, calm 0.3, love 0.4, contempt 0.6, and joy 0.4.
[0052] Second, the core emotion type sequence is analyzed to extract the number of core emotion types contained in each time unit, forming a core emotion type quantity sequence. An arithmetic average calculation is performed on all values in the core emotion type quantity sequence to obtain a core emotion type quantity average. The core emotion type quantity average reflects the average level of the number of core emotion types experienced by the character in the entire animation segment.
[0053] For example, the core emotion type quantity sequence in an animation segment containing 3 time units is [3, 2, 3]. Through arithmetic average calculation, the core emotion type quantity average is (3+2+3) / 3=2.67.
[0054] Further, according to the pre-set initial emotion complexity of each emotion type, a weighted fusion calculation is performed in combination with the core emotion intensity sequence. For each time unit, the intensity values of the core emotion types in the unit are used as weights to perform a weighted summation on the initial emotion complexity corresponding to the intensity values, to obtain the emotion complexity value of the time unit, thereby forming a core emotion complexity sequence. An arithmetic average calculation is performed on all values in the core emotion complexity sequence to obtain a core emotion complexity average. The average comprehensively reflects the average contribution level of emotion types and their intensity to the overall complexity.
[0055] For example, assuming that in an animation segment containing 3 time units, the first time unit identifies an anger intensity of 0.6, a sadness intensity of 0.3, and a fear intensity of 0.1, the emotion complexity = 0.9*0.6+0.8*0.3+0.7*0.1=0.54+0.24+0.07=0.85; the second time unit identifies a joy intensity of 0.7 and a calm intensity of 0.3, the emotion complexity = 0.4*0.7+0.3*0.3=0.28+0.09=0.37; the third time unit identifies a tension intensity of 0.5, an aversion intensity of 0.3, and a contempt intensity of 0.2, the emotion complexity = 0.7*0.5+0.6*0.3+0.6*0.2=0.35+0.18+0.12=0.65. Then the core emotion complexity sequence is [0.85, 0.37, 0.65], and the average of the core emotion complexity = (0.85+0.37+0.65) / 3=1.87 / 3=0.623.
[0056] Further, the average of the calculated core emotion type quantity is multiplied by the average of the core emotion complexity, and the product is set as the emotion mixture complexity. The emotion mixture complexity takes into account the diversity of emotion types and the inherent adjustment difficulty and intensity of each emotion type, and is used to represent the internal composition complexity of the role emotion state.
[0057] For example, assuming that the average of the core emotion type quantity is 2.67, and the average of the core emotion complexity is 0.623, then the emotion mixture complexity = 2.67*0.623=1.663, indicating that the emotion composition of the animation segment has a high complexity.
[0058] Further, the emotion change complexity is evaluated based on the core emotion type quantity sequence and the core emotion complexity sequence, and the evaluation aims to quantify the volatility and instability of the role emotion state in the time dimension.
[0059] Specifically, the emotion change complexity is evaluated according to the core emotion type quantity sequence and the core emotion complexity sequence, including:
[0060] The type quantity fluctuation coefficient is calculated according to the core emotion type quantity sequence, wherein the type quantity fluctuation coefficient is the ratio of the quantity standard deviation to the quantity average in the core emotion type quantity sequence;
[0061] The emotion complexity fluctuation coefficient is calculated according to the core emotion complexity sequence;
[0062] The type quantity fluctuation coefficient and the emotion complexity fluctuation coefficient are weighted and fused to obtain the emotion change complexity.
[0063] First, the volatility of the core emotion type number sequence is analyzed, and the type number volatility coefficient is calculated. The type number volatility coefficient is obtained by calculating the ratio of the standard deviation to the mean of the core emotion type number sequence. The standard deviation reflects the dispersion of the core emotion type number in the time dimension, and the mean represents the average level of the core emotion type number. The ratio of the two, i.e. the type number volatility coefficient, is used to quantify the relative volatility amplitude of the core emotion type number over time. The larger the type number volatility coefficient value, the more unstable the number of emotions experienced by the character in the animation segment.
[0064] For example, based on the core emotion type number sequence [3, 2, 3] obtained in the preceding example, the sequence mean = (3 + 2 + 3) / 3 ≈ 2.667, and the standard deviation = √[((3 - 2.667) 2 +(2 - 2.667) 2 +(3 - 2.667) 2 ) / 3] ≈ 0.471. The type number volatility coefficient = 0.471 / 2.667 ≈ 0.177.
[0065] Second, the volatility of the core emotion complexity sequence is analyzed, and the emotion complexity volatility coefficient is calculated. The calculation of the emotion complexity volatility coefficient also uses the ratio of the standard deviation to the mean of the sequence. The standard deviation of the core emotion complexity sequence reflects the fluctuation of the overall complexity of the emotional state, and the mean represents the average level of the overall complexity. The ratio of the two, i.e. the emotion complexity volatility coefficient, quantifies the relative volatility intensity of the overall complexity of the emotion over time. The larger the emotion complexity volatility coefficient value, the more dramatic the fluctuation of the overall complexity of the emotional state.
[0066] For example, based on the core emotion complexity sequence [0.85, 0.37, 0.65] obtained in the preceding example, the sequence mean = (0.85 + 0.37 + 0.65) / 3 ≈ 0.623, and the standard deviation = √[((0.85 - 0.623) 2 +(0.37 - 0.623) 2 +(0.65 - 0.623) 2 ) / 3] ≈ 0.196. The emotion complexity volatility coefficient = 0.196 / 0.623 ≈ 0.315.
[0067] Finally, the calculated type quantity fluctuation coefficient and emotion complexity fluctuation coefficient are weighted and fused to obtain the final emotion change complexity. In the weighting process, the two coefficients are respectively assigned corresponding weights, and the weight configuration is adjusted according to the specific application scene, for example, in the scene where the emotion type stability is more important, the type quantity fluctuation coefficient can be assigned a weight of 0.4, and the emotion complexity fluctuation coefficient can be assigned a weight of 0.6. The emotion change complexity captures the fluctuation of the number of emotion types and the fluctuation of the change of the comprehensive emotion complexity, and can fully represent the dynamic change characteristics of the role emotion state in the time dimension.
[0068] For example, the type quantity fluctuation coefficient is 0.177, the emotion complexity fluctuation coefficient is 0.315, the type quantity fluctuation coefficient weight is 0.4, and the emotion complexity fluctuation coefficient weight is 0.6, then the emotion change complexity = 0.177*0.4 + 0.315*0.6 = 0.0708 + 0.189 ≈ 0.26. The emotion change complexity comprehensively reflects the dynamic change intensity of the role emotion state in the time dimension.
[0069] S30: obtaining a comprehensive emotion complexity by weighted fusion based on the emotion mixed complexity and the emotion change complexity;
[0070] Specifically, obtaining a comprehensive emotion complexity by weighted fusion based on the emotion mixed complexity and the emotion change complexity, comprising:
[0071] Based on the scene type of the to-be-processed animation segment, the corresponding adaptive weight proportion is called from the pre-defined weight mapping table, wherein the adaptive weight proportion includes an adaptive mixed complexity weight and an adaptive change complexity weight, and the scene type at least includes introspective drama type, intense action type, humorous comedy type, epic narrative type and ordinary scene.
[0072] After the emotion mixed complexity and the emotion change complexity are dimensionless processed, the adaptive mixed complexity weight and the adaptive change complexity weight are weighted and fused to obtain the comprehensive emotion complexity.
[0073] Firstly, based on the scene type of the to-be-processed animation segment, the corresponding adaptive weight proportion is called from the predefined weight mapping table. The weight mapping table establishes the corresponding relationship between different scene types and weight configurations, which specifically includes the following mapping rules: when the scene type is introspective drama type, the adaptive weight proportion is that the adaptive mixed complexity weight is greater than 0.6 and the adaptive change complexity weight is less than 0.4; when the scene type is intense action type, the adaptive weight proportion is that the adaptive mixed complexity weight is less than 0.4 and the adaptive change complexity weight is greater than 0.6; when the scene type is humorous comedy type, the adaptive weight proportion is that the adaptive mixed complexity weight is less than 0.5 and the adaptive change complexity weight is greater than 0.5; when the scene type is epic narrative type, the adaptive weight proportion is that the adaptive mixed complexity weight is between 0.4 and 0.6 and the adaptive change complexity weight is between 0.4 and 0.6; when the scene type is normal scene, the adaptive weight proportion is that the adaptive mixed complexity weight is equal to 0.5 and the adaptive change complexity weight is equal to 0.5. The distribution proportions of the two weights under different scene types are different, and in actual application, appropriate specific values should be assigned to each weight within the range specified in the table according to the preset rules or real-time strategies, so as to reflect the different emphasis of different scenes on emotional complexity and emotional dynamic change.
[0074] Secondly, weighted fusion calculation is performed. Before the weighted calculation, the emotional mixed complexity and the emotional change complexity are first processed dimensionless to ensure that the two types of complexity values are in the same order of magnitude and eliminate the calculation deviation that may be caused by the dimensional difference. Then, according to the adaptive mixed complexity weight and the adaptive change complexity weight obtained from the weight mapping table, the processed emotional mixed complexity and the processed emotional change complexity are weighted and summed to obtain the comprehensive emotional complexity. The comprehensive emotional complexity comprehensively reflects the static complexity degree of the role emotional state and the dynamic change characteristics.
[0075] For example, in the introspective drama type scene, the emotional mixed complexity is set to 1.663, the weight is 0.65, the emotional change complexity is 0.260, and the weight is 0.35, so the comprehensive emotional complexity = 1.663 x 0.65 + 0.260 x 0.35 = 1.081 + 0.091 = 1.172. This dynamic weight distribution mechanism based on scene type can ensure that the comprehensive emotional complexity can accurately adapt to the emotional performance needs of different scenes.
[0076] S40: According to the comprehensive emotional complexity, the adaptive evaluation model of sound parameter optimization selects parameters and adaptive optimization process parameters to generate a sound parameter optimization scheme;
[0077] The adaptive evaluation model is a scheme quality evaluation model based on a machine learning algorithm, which is used for quality scoring and effect prediction of the generated candidate sound adjustment scheme. Setting adaptive evaluation model selection parameters and adaptive optimization process parameters for the adaptive evaluation model is a process of dynamically configuring and optimizing the search strategy based on the comprehensive emotional complexity. Among them, the adaptive evaluation model selection parameter specifies the number of evaluation model instances that need to be used in parallel in the optimization process, which directly affects the comprehensiveness and reliability of scheme evaluation. The adaptive optimization process parameter specifies the upper limit of the number of iterations for random search in the parameter space, which determines the search depth and convergence quality of the optimization process.
[0078] For example, when the comprehensive emotional complexity is high, the adaptive evaluation model selection parameter needs to be automatically increased to adopt a more diversified evaluation perspective, and the adaptive optimization process parameter needs to be increased to perform more sufficient parameter space exploration, so as to ensure that the optimal sound adjustment scheme can still be found under complex emotional states.
[0079] Specifically, the adaptive evaluation model selection parameter and the adaptive optimization process parameter of the sound parameter optimization are set according to the comprehensive emotional complexity, including:
[0080] The ratio of the comprehensive emotional complexity to the historical maximum comprehensive emotional complexity in the historical time range is set as an evaluation model selection compensation coefficient;
[0081] The ratio of the evaluation model selection compensation coefficient to the preset evaluation model number K is rounded, which is used as the adaptive evaluation model selection parameter, wherein K is greater than or equal to 5 and less than or equal to 20, and the adaptive evaluation model selection parameter is not less than 2;
[0082] The ratio of the comprehensive emotional complexity to the preset standard comprehensive emotional complexity is set as an optimization number compensation coefficient;
[0083] The product of the optimization number compensation coefficient and the preset optimization convergence number is rounded, which is used as the adaptive optimization process parameter, wherein the preset optimization convergence number is 500, and the adaptive optimization process parameter is not less than 200.
[0084] First, the ratio of the comprehensive emotional complexity of the current animation segment to be processed to the historical maximum comprehensive emotional complexity recorded in the historical time range is set as a model selection compensation coefficient. The historical maximum comprehensive emotional complexity is the maximum value of the recorded comprehensive emotional complexity in the process of processing past animation segments, and the historical time range is a backtracking period set according to application requirements. According to the average production period of a typical animation project, for example, the data of animation production projects in the last 3 months is set.
[0085] For example, assuming that the historical maximum comprehensive emotion complexity is 2.5 and the comprehensive emotion complexity of the current segment is 1.172, the model selects a compensation coefficient = 1.172 / 2.5 = 0.4688.
[0086] Secondly, the calculated evaluation model selection compensation coefficient is multiplied by the preset evaluation model quantity K, and the product result is rounded up. The preset evaluation model quantity K is the total number of configured basic evaluation models, which is between 5 and 20, and is set according to the computing resources and evaluation accuracy requirements.
[0087] For example, the model selection compensation coefficient is 0.4688, and the preset evaluation model quantity K is 10. The calculation process is 0.4688*10 = 4.688, and the adaptive evaluation model selection parameter obtained after rounding up is 5. The finally obtained adaptive evaluation model selection parameter specifies the number of evaluation models that need to be called in the subsequent optimization process, and the adaptive evaluation model selection parameter should not be less than 2 to ensure that the optimization process has basic diversity.
[0088] Further, the optimization frequency compensation coefficient is calculated. The optimization frequency compensation coefficient is obtained by dividing the current comprehensive emotion complexity by the preset standard comprehensive emotion complexity. The preset standard comprehensive emotion complexity is a benchmark value obtained based on a large amount of typical animation scene data analysis, representing the regular level of emotion complexity. The optimization frequency compensation coefficient is used to quantify the deviation of the current segment emotion complexity from the standard benchmark.
[0089] Finally, the calculated optimization frequency compensation coefficient is multiplied by the preset optimization convergence frequency 500, and the product result is rounded. The preset optimization convergence frequency 500 is an iteration frequency that can ensure the basic convergence quality based on experience. The finally obtained adaptive optimization process parameter clearly specifies the number of iteration calculations that need to be performed in the subsequent sound parameter optimization process. The adaptive optimization process parameter value is set to have a lower limit constraint of not less than 200, so as to ensure that even in the face of a segment with low emotion complexity, the optimization process can also perform sufficient basic exploration.
[0090] For example, assuming that the preset standard comprehensive emotion complexity is 1 and the current segment comprehensive emotion complexity is 1.172, the optimization frequency compensation coefficient is 1.172, 1.172*500 = 586, and the adaptive optimization process parameter is determined to be 586 after rounding. The adaptive optimization process parameter value is greater than the lower limit 200, which meets the requirements, indicating that 586 iterations of optimization need to be performed in the sound parameter space.
[0091] S50: Perform sound parameter optimization according to the sound parameter optimization scheme, output the optimal sound adjustment scheme, and adjust the sound parameters of the target character in the to-be-processed animation segment according to the optimal sound adjustment scheme.
[0092] First, the sound parameter optimization scheme is optimized according to the sound parameter optimization scheme, and the optimal sound adjustment scheme is output, including:
[0093] Obtain the sound parameter adjustment space, wherein the sound parameter adjustment space includes a plurality of sound parameter adjustment thresholds, and the sound parameter at least includes a volume parameter, an equalization parameter, a sound image parameter, a reverberation sending amount, and a filter cutoff frequency;
[0094] Randomly select a first sound adjustment scheme in the sound parameter adjustment space, and generate a first emotion-sound adjustment scheme in combination with the emotion attribute information sequence;
[0095] Select a parameter calling scheme quality evaluation plug-in based on the adaptive evaluation model, and perform quality evaluation on the first emotion-sound adjustment scheme, and output the first scheme quality coefficient mean;
[0096] Continue to randomly select and evaluate the quality of the sound adjustment scheme in the sound parameter adjustment space until the optimization iteration number reaches the adaptive optimization process parameter, and output the sound adjustment scheme corresponding to the maximum scheme quality coefficient mean as the optimal sound adjustment scheme.
[0097] First, a predefined sound parameter adjustment space is obtained. The sound parameter adjustment space is obtained by analyzing the sound parameter configuration of historical excellent animation works and combining the experience of field experts, and clearly specifies the reasonable value range of various adjustable sound parameters. Its core parameters at least include the volume parameter for controlling the loudness, the equalization parameter for adjusting the intensity of each frequency band, the sound image parameter for positioning the direction of the sound source, the reverberation sending amount for shaping the sense of space, and the filter cutoff frequency for filtering specific frequencies. Each parameter has a clear upper and lower threshold, thus forming a multi-dimensional parameter search space.
[0098] For example, the volume parameter adjustment range is -20dB to +10dB; the equalization parameter center frequency adjustment range is 60Hz to 16000Hz; the sound image parameter direction adjustment range is -45 degrees to +45 degrees; the reverberation sending amount adjustment range is 0% to 100%; and the filter cutoff frequency adjustment range is 80Hz to 12000Hz.
[0099] Secondly, an initialization operation is performed in the sound parameter adjustment space, and a group of specific parameter values are randomly generated to form a first sound adjustment scheme. The first sound adjustment scheme is associated and integrated with the emotion attribute information sequence of the target role in the animation segment to be processed to generate a first emotion-sound adjustment scheme. This scheme completely describes the sound parameter configuration adopted under a specific emotion change sequence.
[0100] Further, according to the number specified by the adaptive evaluation model selection parameter, the corresponding number of parallel evaluation units in the scheme quality evaluation plug-in are called. The scheme quality evaluation plug-in is an integrated evaluation tool, which encapsulates multiple trained deep learning models for evaluating the matching quality of emotions and sound.
[0101] Specifically, the construction method of the scheme quality evaluation plug-in includes:
[0102] Based on historical animation sound adjustment records, a set of sample emotion-sound adjustment schemes is collected, and different sample emotion-sound adjustment schemes are marked for adjustment quality to obtain a set of sample scheme quality coefficients.
[0103] The set of sample emotion-sound adjustment schemes and the set of sample scheme quality coefficients are divided into K parts as training data, and K times of first sample training sets are constructed by randomly selecting from the K data sets, and K sample training sets are obtained by iterative selection.
[0104] The K sample training sets are used to train deep learning models to convergence, respectively, to generate K scheme quality evaluation models, and a scheme quality evaluation plug-in is obtained by combination.
[0105] First, based on historical animation sound adjustment records, a large number of sample emotion-sound adjustment schemes are collected to form a sample set. The historical animation sound adjustment records are a set of corresponding relationships between character emotion data and final adopted sound parameter configurations recorded in past successful animation projects, which are obtained through project archives and professional audio workstation logs. Each sample emotion-sound adjustment scheme obtained by collection contains a specific sequence of emotion attribute information and its corresponding sound parameter configuration. A professional evaluation is performed on different sample schemes by a senior sound engineer, and quality annotations are performed according to standards such as sound picture matching degree and artistic expression to obtain corresponding sample scheme quality coefficients, forming a set of sample scheme quality coefficients.
[0106] Secondly, the complete set of sample emotion-sound adjustment schemes and the corresponding set of sample scheme quality coefficients are divided into K data subsets as original training data. Through random sampling with replacement, K times of training sets are repeatedly extracted from the K data subsets, and each time a training set is constructed. This process is iterated for K rounds to finally generate K sample training sets that are different from each other, aiming to increase the diversity of the training set.
[0107] Finally, K deep learning models are independently trained using the K sample training sets obtained above. Each model is trained until the loss function converges, thereby obtaining K scheme quality evaluation models with emotion-audio scheme quality evaluation capabilities. The convergence condition is set according to the loss function change rate during training, for example, when the loss function value on the validation set decreases by less than one thousandth in twenty consecutive training periods, the model is determined to have converged. The K trained models are combined and packaged to form the final scheme quality evaluation plug-in. The scheme quality evaluation plug-in can effectively reduce the evaluation deviation of a single model by integrating the judgments of multiple models, thereby providing more stable and reliable quality evaluation results.
[0108] For example, because the quality evaluation of the emotion-audio adjustment scheme involves a complex nonlinear mapping between high-dimensional parameters and subjective listening, and deep learning models have a significant advantage in mining such deep nonlinear relationships, a deep learning model can be selected as the core architecture of the scheme quality evaluation model.
[0109] Specifically, the scheme quality evaluation model mainly consists of an input layer, a feature fusion layer, and an evaluation output layer. The input layer receives structured emotion-audio adjustment scheme data, which includes a standardized sequence of emotion attribute information and corresponding audio parameter vectors. The feature fusion layer uses a multi-layer perceptron structure, with the number of hidden layer neurons adaptively configured to 128 to 512 according to the input dimension. Each layer of neural network uses a ReLU activation function to enhance the nonlinear representation capability, and a Dropout layer is embedded between network layers with a dropout rate set to 0.2 to 0.3 to effectively suppress model overfitting and improve its generalization performance. The output layer uses a Sigmoid activation function to map the final abstract features to a continuous value between 0 and 1 as the scheme quality score.
[0110] During training, the key hyperparameters include a learning rate set to 0.001, a training round number set to 200, and a batch size set to 32. The learning rate is set based on the balance between training stability and convergence speed, the training round number is set to ensure that the model learns the quality evaluation pattern in the data, and the batch size is selected to balance training efficiency and gradient stability. Specifically, a supervised learning training method is used to collect sample emotion-audio adjustment schemes from historical animation audio adjustment records as input sample sets, and simultaneously obtain scheme quality coefficients labeled by experienced sound engineers to form sample label sets. The input sample sets and corresponding label sample sets are divided into training sets, validation sets, and test sets in a ratio of 7:2:1.
[0111] Further, the sample scheme data in the training set is taken as input, and the corresponding quality coefficient is taken as a supervision signal. Through a back propagation algorithm and in combination with an Adam optimizer, the network weight parameters are iteratively optimized. A mean square error loss function is used to measure the deviation between the predicted quality score and the labeled quality coefficient. The training process is monitored through the validation set. When the validation set loss function value does not decrease for 15 consecutive rounds and the correlation coefficient between the model prediction and the expert labeling reaches a predetermined threshold, such as 0.9, the training is terminated, and a converged scheme quality evaluation model is obtained. The scheme quality evaluation model can effectively capture the complex matching relationship between the emotional state and the sound parameter, and realize accurate scheme quality evaluation.
[0112] Specifically, in the analysis process, each invoked parallel evaluation unit independently analyzes the first emotion-sound adjustment scheme. The analysis is based on the sound parameter configuration and the sequence of emotional attribute information contained in the scheme, evaluates the fit degree between the two, and outputs a quantitative quality score. Further, the quality scores output by all parallel evaluation units are collected, and the arithmetic mean is calculated. The mean value is taken as the first scheme quality coefficient mean value. The first scheme quality coefficient mean value reflects the matching degree of the first emotion-sound adjustment scheme and the emotional state of the character.
[0113] Finally, the scheme generation and evaluation process is repeatedly executed, and new sound adjustment schemes are continuously randomly selected in the sound parameter adjustment space and evaluated. When the cumulative number of optimization iterations reaches a value specified by the adaptive optimization process parameter, the search process is terminated. From all the evaluated schemes, the scheme with the highest scheme quality coefficient mean value is selected as the final optimal sound adjustment scheme. The optimal sound adjustment scheme is the best parameter configuration obtained after sufficient optimization.
[0114] Further, according to the optimal sound adjustment scheme, the target character in the to-be-processed animation segment is adjusted in sound parameter, that is, the best parameter configuration obtained in the optimization process is applied to the actual animation production link. Specifically, the specific parameter values contained in the optimal sound adjustment scheme are analyzed, including the volume decibel value, the equalizer frequency band gain, the sound image positioning angle, the reverberation intensity ratio, and the filter cutoff frequency, and other core parameters. The parameter configuration is loaded into the audio track of the animation production system through an audio processing interface, and a corresponding binding relationship with the target character audio stream is established. In the parameter application process, the parameter values are dynamically adjusted according to the animation timeline to ensure that the sound effect is synchronized with the change of the character's emotional state, and a new animation segment with the best sound effect is generated. In summary, the implementation process realizes a complete closed loop from parameter optimization to actual application, so that the optimal solution in theory can be converted into specific artistic performance effect.
[0115] In summary, the embodiments of the present application have at least the following technical effects:
[0116] Compared with the prior art, the application firstly realizes accurate analysis and representation of the dynamic characteristics of the animation character emotion by constructing a quantitative evaluation system of mixed emotion complexity and emotion change complexity. Secondly, based on the adaptive configuration of the sound parameter optimization strategy based on the comprehensive emotion complexity, including dynamically adjusting the evaluation model quantity and the optimization iteration depth, the optimization process and the emotion characteristics are accurately matched. Thirdly, by performing directional optimization in the parameter space, the sound adjustment scheme highly consistent with the emotion fluctuation is automatically generated, realizing the adaptive adjustment of the sound parameters. Finally, a complete emotion-driven sound adjustment workflow is established, which improves the production efficiency and artistic expression of animation sound effects, and effectively overcomes the problem of rigid emotion expression caused by the dependence of traditional methods on fixed templates.
[0117] Embodiment two, as shown in Figure 2 based on the same inventive concept of the animation sound dynamic adjustment method based on the role emotion provided in embodiment one, the application embodiment also provides an animation sound dynamic adjustment device based on the role emotion, comprising:
[0118] An emotion feature extraction module 11 is configured to extract core emotion types and core emotion intensities based on the emotion attribute information sequence of the target character in the to-be-processed animation segment, and obtain a core emotion type sequence and a core emotion intensity sequence.
[0119] An emotion complexity evaluation module 12 is configured to evaluate the emotion mixed complexity and the emotion change complexity according to the core emotion type sequence and the core emotion intensity sequence.
[0120] A comprehensive complexity calculation module 13 is configured to obtain a comprehensive emotion complexity by weighted fusion based on the emotion mixed complexity and the emotion change complexity.
[0121] An optimization scheme generation module 14 is configured to set the adaptive evaluation model selection parameters and the adaptive optimization process parameters of the sound parameter optimization according to the comprehensive emotion complexity, and generate a sound parameter optimization scheme.
[0122] A sound parameter adjustment module 15 is configured to perform sound parameter optimization according to the sound parameter optimization scheme, output an optimal sound adjustment scheme, and adjust the sound parameters of the target character in the to-be-processed animation segment according to the optimal sound adjustment scheme.
[0123] The emotion feature extraction module 11 is specifically configured to:
[0124] extract core emotion types and core emotion intensities based on the emotion attribute information sequence of the target character in the to-be-processed animation segment, and obtain a core emotion type sequence and a core emotion intensity sequence, including:
[0125] The emotion type extraction strategy is configured, wherein when the emotion attribute information is a single emotion type, the emotion type is determined as a core emotion type; when the emotion attribute information is a mixture of multiple emotion types, the emotion types are sorted based on the intensity values of the emotion types, and the top N emotion types with the highest intensity are determined as the core emotion types, and N is greater than or equal to 2.
[0126] According to the emotion type extraction strategy, the emotion attribute information sequence of the target character in the to-be-processed animation segment is subjected to information extraction, and a core emotion type sequence and a core emotion intensity sequence are obtained, wherein the emotion types at least include joy, sadness, anger, fear, disgust, surprise, tension, calmness, love and contempt.
[0127] The emotion complexity evaluation module 12 is specifically configured to:
[0128] The emotion complexity evaluation module 12 is specifically configured to:
[0129] A plurality of initial emotion complexities are set based on a plurality of emotion types, wherein the initial emotion complexity is set based on a historical sound adjustment difficulty corresponding to the emotion type;
[0130] The type number extraction is performed according to the core emotion type sequence to obtain a core emotion type number sequence, and a core emotion type number mean value is obtained by performing mean value calculation on the core emotion type number sequence;
[0131] According to the plurality of initial emotion complexities, the core emotion intensity sequence is subjected to weighted fusion to obtain a core emotion complexity sequence, and a core emotion complexity mean value is obtained by performing mean value calculation on the core emotion complexity sequence;
[0132] The product of the core emotion type number mean value and the core emotion complexity mean value is set as the emotion mixing complexity;
[0133] The emotion change complexity is evaluated according to the core emotion type number sequence and the core emotion complexity sequence.
[0134] Specifically, the emotion change complexity is evaluated according to the core emotion type number sequence and the core emotion complexity sequence, including:
[0135] The number fluctuation calculation is performed according to the core emotion type number sequence to obtain a type number fluctuation coefficient, wherein the type number fluctuation coefficient is a ratio of a number standard deviation to a number mean value in the core emotion type number sequence;
[0136] The emotion complexity fluctuation calculation is performed according to the core emotion complexity sequence to obtain an emotion complexity fluctuation coefficient.
[0137] The type quantity fluctuation coefficient and the emotion complexity fluctuation coefficient are weighted and fused to obtain an emotion change complexity.
[0138] The comprehensive complexity calculation module 13 is specifically configured to:
[0139] The emotion mixed complexity and the emotion change complexity are weighted and fused to obtain a comprehensive emotion complexity, including:
[0140] Based on the scene type of the to-be-processed animation segment, an adaptive weight proportion is called from a predefined weight mapping table, wherein the adaptive weight proportion includes an adaptive mixed complexity weight and an adaptive change complexity weight, and the scene type at least includes introspective drama type, intense action type, humorous comedy type, epic narrative type, and ordinary scene.
[0141] After the emotion mixed complexity and the emotion change complexity are dimensionless processed, the adaptive mixed complexity weight and the adaptive change complexity weight are weighted and fused to obtain a comprehensive emotion complexity.
[0142] The optimization scheme generation module 14 is specifically configured to:
[0143] The adaptive evaluation model selection parameter and the adaptive optimization process parameter of the sound parameter optimization are set according to the comprehensive emotion complexity, including:
[0144] The ratio of the comprehensive emotion complexity to the historical maximum comprehensive emotion complexity in a historical time range is set as an evaluation model selection compensation coefficient;
[0145] The adaptive evaluation model selection parameter is obtained by rounding the ratio of the evaluation model selection compensation coefficient to a preset evaluation model quantity K, wherein K is greater than or equal to 5 and less than or equal to 20, and the adaptive evaluation model selection parameter is not less than 2;
[0146] The ratio of the comprehensive emotion complexity to a preset standard comprehensive emotion complexity is set as a number of optimization times compensation coefficient;
[0147] The adaptive optimization process parameter is obtained by rounding the product of the number of optimization times compensation coefficient and a preset number of optimization convergence times, wherein the preset number of optimization convergence times is 500, and the adaptive optimization process parameter is not less than 200.
[0148] The sound parameter adjustment module 15 is specifically configured to:
[0149] The sound parameter optimization is performed according to the sound parameter optimization scheme, and an optimal sound adjustment scheme is output, including:
[0150] acquire an acoustic parameter adjustment space, wherein the acoustic parameter adjustment space comprises a plurality of acoustic parameter adjustment thresholds, and the acoustic parameters comprise at least volume parameters, equalization parameters, sound image parameters, reverberation sending amounts, and filter cutoff frequencies;
[0151] randomly select a first acoustic adjustment scheme in the acoustic parameter adjustment space, and generate a first emotion-acoustic adjustment scheme in combination with the sequence of emotion attribute information;
[0152] select a parameter calling scheme quality evaluation plug-in based on the adaptive evaluation model, perform quality evaluation on the first emotion-acoustic adjustment scheme, and output a first scheme quality coefficient mean value;
[0153] continue to randomly select an acoustic adjustment scheme and perform quality evaluation on the acoustic adjustment scheme in the acoustic parameter adjustment space until the number of optimization iterations reaches the adaptive optimization process parameter, and output an acoustic adjustment scheme corresponding to the maximum scheme quality coefficient mean value as an optimal acoustic adjustment scheme.
[0154] Specifically, the construction method of the scheme quality evaluation plug-in comprises:
[0155] based on historical animation acoustic adjustment records, collect a sample emotion-acoustic adjustment scheme set, and perform adjustment quality labeling on different sample emotion-acoustic adjustment schemes to obtain a sample scheme quality coefficient set;
[0156] divide the sample emotion-acoustic adjustment scheme set and the sample scheme quality coefficient set into K parts as training data, randomly select K times from the K data sets, and iteratively select K times to obtain K sample training sets;
[0157] train the K sample training sets respectively until convergence to generate K scheme quality evaluation models, and combine the K scheme quality evaluation models to obtain a scheme quality evaluation plug-in.
[0158] It should be noted that the above sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. Moreover, the above describes specific embodiments of the present application. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0159] The above only describes the preferred embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0160] The specification and drawings are, of course, to be regarded in an illustrative rather than a restrictive sense. It is to be understood that any such modifications, variations, combinations or equivalents which fall within the scope of the application are intended to be embraced herein.
Claims
1. A method for dynamically adjusting animation sound based on character emotions, characterized in that, The methods include: Based on the emotional attribute information sequence of the target character in the animation clip to be processed, the core emotion type and core emotion intensity are extracted, and the core emotion type sequence and core emotion intensity sequence are obtained. The complexity of emotion mixing and the complexity of emotion change are evaluated based on the core emotion type sequence and the core emotion intensity sequence. The comprehensive emotional complexity is obtained by weighted fusion of the emotional mixing complexity and the emotional change complexity. Based on the comprehensive emotional complexity, the parameters of the sound parameter optimization adaptation evaluation model are selected and the adaptation optimization process parameters are set to generate a sound parameter optimization scheme; Based on the sound parameter optimization scheme, sound parameters are optimized, the optimal sound adjustment scheme is output, and the sound parameters of the target character in the animation segment to be processed are adjusted according to the optimal sound adjustment scheme. The selection of parameters and the optimization process parameters for the sound parameter optimization model, based on the comprehensive emotional complexity, include: The ratio of the overall emotional complexity to the historical maximum overall emotional complexity within the historical time range is set as the compensation coefficient for the evaluation model. The ratio of the compensation coefficient of the evaluation model to the number of preset evaluation models K is rounded down and used as the selection parameter for the adaptive evaluation model, wherein K is greater than or equal to 5 and less than or equal to 20, and the selection parameter for the adaptive evaluation model is not less than 2. The ratio of the overall emotional complexity to the preset standard overall emotional complexity is set as the optimization number compensation coefficient. The product of the optimization number compensation coefficient and the preset optimization convergence number is rounded down to obtain the adaptation optimization process parameter, wherein the preset optimization convergence number is 500 and the adaptation optimization process parameter is not less than 200.
2. The method for dynamic adjustment of animation sound based on character emotions according to claim 1, characterized in that, Based on the emotional attribute information sequence of the target character in the animation clip to be processed, the core emotion type and core emotion intensity are extracted, and the core emotion type sequence and core emotion intensity sequence are obtained, including: Configure an emotion type extraction strategy. When the emotion attribute information is a single emotion type, the emotion type is determined as the core emotion type. When the emotion attribute information is a mixture of multiple emotion types, the emotions are sorted based on their intensity values, and the top N emotion types with the highest intensity are determined as the core emotion types, where N is greater than or equal to 2. According to the emotion type extraction strategy, information is extracted from the sequence of emotional attributes of the target character in the animation clip to be processed, and the core emotion type sequence and core emotion intensity sequence are obtained. The emotion type includes at least joy, sadness, anger, fear, disgust, surprise, tension, calmness, love and contempt.
3. The method for dynamic adjustment of animation sound based on character emotions according to claim 2, characterized in that, The emotional mixing complexity and emotional change complexity are evaluated based on the core emotion type sequence and core emotion intensity sequence, including: Several initial emotional complexities are set based on several emotion types, wherein the initial emotional complexity is set based on the historical sound adjustment difficulty corresponding to the emotion type; The number of core emotion types is extracted based on the core emotion type sequence to obtain the core emotion type number sequence, and the mean of the core emotion type number sequence is calculated to obtain the mean of the number of core emotion types. Based on the aforementioned initial emotional complexity, a core emotional complexity sequence is obtained by weighted fusion of the core emotional intensity sequence, and the mean of the core emotional complexity sequence is calculated to obtain the mean of the core emotional complexity. The product of the average number of core emotion types and the average complexity of core emotions is set as the emotion mixture complexity. The complexity of emotion change is evaluated based on the sequence of the number of core emotion types and the sequence of core emotion complexity.
4. The method for dynamic adjustment of animation sound based on character emotions according to claim 3, characterized in that, The complexity of emotion change is evaluated based on the sequence of the number of core emotion types and the sequence of core emotion complexity, including: The fluctuation of the number of core emotion types is calculated based on the number sequence of the core emotion types to obtain the type number fluctuation coefficient, wherein the type number fluctuation coefficient is the ratio of the standard deviation of the number to the mean of the number in the number sequence of the core emotion types. The emotional complexity fluctuation is calculated based on the core emotional complexity sequence to obtain the emotional complexity fluctuation coefficient. The emotional change complexity is obtained by weighted fusion of the fluctuation coefficient of the type quantity and the fluctuation coefficient of the emotional complexity.
5. The method for dynamic adjustment of animation sound based on character emotions according to claim 1, characterized in that, The comprehensive emotional complexity is obtained by weighted fusion of the emotional mixing complexity and the emotional change complexity, including: Based on the scene type of the animation segment to be processed, the corresponding adaptation weight ratio is called from the predefined weight mapping table. The adaptation weight ratio includes the adaptation weight for mixed complexity and the adaptation weight for variable complexity. The scene type includes at least introspective drama, intense action, humorous comedy, epic narrative, and ordinary scene. After dimensionless processing of the emotional mixture complexity and emotional change complexity, a weighted fusion is performed based on the adaptive mixture complexity weight and the adaptive change complexity weight to obtain the comprehensive emotional complexity.
6. The method for dynamic adjustment of animation sound based on character emotions according to claim 1, characterized in that, Based on the aforementioned audio parameter optimization scheme, audio parameters are optimized, and the optimal audio adjustment scheme is output, including: Obtain the audio parameter adjustment space, wherein the audio parameter adjustment space includes several audio parameter adjustment thresholds, and the audio parameters include at least volume parameters, equalization parameters, sound panning parameters, reverberation delivery amount, and filter cutoff frequency; Within the sound parameter adjustment space, a first sound adjustment scheme is randomly selected, and a first emotion-sound adjustment scheme is generated by combining the emotion attribute information sequence. Based on the adaptation evaluation model, the parameter selection and calling scheme quality evaluation plugin are performed on the first emotion-sound adjustment scheme to evaluate its quality and output the average quality coefficient of the first scheme. Within the audio parameter adjustment space, the random selection and quality evaluation of audio adjustment schemes continue until the number of optimization iterations reaches the adaptation optimization process parameters. The audio adjustment scheme corresponding to the maximum average quality coefficient of the output scheme is set as the optimal audio adjustment scheme.
7. The method for dynamic adjustment of animation sound based on character emotions according to claim 6, characterized in that, The method for constructing the solution quality evaluation plugin includes: Based on historical animation sound adjustment records, a sample emotion-sound adjustment scheme set was collected, and the adjustment quality of different sample emotion-sound adjustment schemes was labeled to obtain a sample scheme quality coefficient set; The sample emotion-sound adjustment scheme set and the sample scheme quality coefficient set are used as training data, and are divided into K equal parts. The first sample training set is constructed by selecting K times with replacement from the K sets of data, and then the first sample training set is obtained by iteratively selecting K times. Using the K sample training sets, deep learning models are trained to convergence to generate K scheme quality evaluation models, which are then combined to obtain a scheme quality evaluation plugin.
8. An animation sound dynamic adjustment device based on character emotions, characterized in that, The method for dynamically adjusting animation audio based on character emotions as described in any one of claims 1-7 includes: The emotion feature extraction module is used to extract the core emotion type and core emotion intensity based on the emotion attribute information sequence of the target character in the animation clip to be processed, and to obtain the core emotion type sequence and core emotion intensity sequence. The emotion complexity assessment module is used to assess the emotion mixing complexity and emotion change complexity based on the core emotion type sequence and the core emotion intensity sequence. The comprehensive complexity calculation module is used to obtain the comprehensive emotion complexity based on the weighted fusion of the emotion mixing complexity and the emotion change complexity. The optimization scheme generation module is used to select parameters and optimization process parameters of the sound parameter optimization adaptation evaluation model based on the comprehensive emotional complexity, and generate a sound parameter optimization scheme. The audio parameter adjustment module is used to optimize audio parameters according to the audio parameter optimization scheme, output the optimal audio adjustment scheme, and adjust the audio parameters of the target character in the animation segment to be processed according to the optimal audio adjustment scheme. The selection of parameters and the optimization process parameters for the sound parameter optimization model, based on the comprehensive emotional complexity, include: The ratio of the overall emotional complexity to the historical maximum overall emotional complexity within the historical time range is set as the compensation coefficient for the evaluation model. The ratio of the compensation coefficient of the evaluation model to the number of preset evaluation models K is rounded down and used as the selection parameter for the adaptive evaluation model, wherein K is greater than or equal to 5 and less than or equal to 20, and the selection parameter for the adaptive evaluation model is not less than 2. The ratio of the overall emotional complexity to the preset standard overall emotional complexity is set as the optimization number compensation coefficient. The product of the optimization number compensation coefficient and the preset optimization convergence number is rounded down to obtain the adaptation optimization process parameter, wherein the preset optimization convergence number is 500 and the adaptation optimization process parameter is not less than 200.
Citation Information
Patent Citations
Sound effect adjusting method and device and audio equipment
CN114283854A
Animation data generation method and device, terminal equipment and storage medium
CN117830480A