Data processing and making system for dynamic recombination and automatic sound mixing of music elements

By building a data processing system for dynamic reorganization and automatic mixing of musical elements, the problems of poor tonality adaptability, broken emotional expression, and homogeneous mixing effects have been solved, cross-module collaborative optimization and efficient conversion of user intentions have been achieved, and the quality of music creation and emotional relevance have been improved.

CN120636376AInactive Publication Date: 2025-09-12ZIBO VOCATIONAL INST
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510953026.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology has problems in the process of reorganizing and mixing music elements, such as poor tonal adaptability, broken emotional expression, homogeneity of cross-style mixing effects, and insufficient collaborative optimization of creation modules, and the efficiency of user intention conversion is low.

Method used

Construct a data processing and production system for dynamic reorganization and automatic mixing of music elements, including a music information deep representation and semantic understanding module, a music element dynamic reorganization module, an intelligent adaptive mixing module, a multi-agent collaborative optimization and evaluation learning module, a user interaction and collaborative creation interface module, and a knowledge base and learning results storage module. Through multi-dimensional feature extraction, emotional semantic understanding, dynamic weight matrix adjustment and multi-objective evaluation optimization, cross-module collaborative optimization is achieved.

Benefits of technology

It improves the tonal adaptability, emotional fit and creative quality of music creation, solves the problems of tonal conflicts and homogeneity of mixing effects in the reorganization of music elements, enhances the efficiency of user intention conversion, and promotes the continuous improvement of creative quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636376A_ABST
    Figure CN120636376A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and digital audio processing, and discloses a music element dynamic recombination and automatic sound mixing data processing and making system. Comprising a music information deep representation and semantic understanding module, a music element dynamic recombination module, an intelligent adaptive sound mixing module, a multi-agent collaborative optimization and evaluation learning module, a user interaction and collaborative creation interface module and a knowledge base and learning result storage module. According to the method, deep understanding of music content is realized through multi-modal music feature extraction, structure analysis and emotion semantic dynamic modeling, sound mixing parameters are adaptively adjusted in real time according to context factors by utilizing a dynamic weight matrix mechanism, and a creation closed loop is formed by combining multi-agent collaborative optimization; the problems that in the prior art, music element recombination adaptability is poor, the sound mixing effect is homogenized, and cross-module collaboration is insufficient are solved, and the music creation efficiency and the emotion expression accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and digital audio processing, and in particular to a data processing and production system for dynamic reorganization and automatic mixing of musical elements. Background Art

[0002] Dynamic reorganization and automatic mixing of musical elements is a creative process that uses data processing technology to intelligently process musical materials. The core is to extract multi-dimensional features and understand emotional semantics of original music data, dynamically reorganize musical elements based on creative intent, and adaptively adjust mixing parameters based on contextual information, realizing full-process intelligence from material processing to finished product output.

[0003] In the existing technology, in order to meet the automation needs of music element reorganization and mixing, solutions based on rule templates or shallow machine learning are mainly adopted. For example, the element reorganization link mostly relies on preset harmonic patterns, rhythmic library or melody motive templates to complete material splicing, and realizes element adaptation through fixed mode conversion rules; the mixing link generally adopts standardized parameter configuration, such as a unified volume balance strategy or a preset equalizer curve. Some systems only achieve basic mixing processing through surface feature matching such as rhythm speed and timbre type.

[0004] Although existing technologies can meet the automation needs of simple scenarios, they generally remain at the level of parameterized operations and rule matching, and lack the ability to understand the deep semantics of music. However, due to the use of single-modal feature extraction, there is a lack of in-depth analysis of the emotional dynamics and structural roles of music, resulting in frequent tonal conflicts or emotional logic breaks in the reorganized elements; the mixing link relies on static parameter templates and cannot adjust parameters such as phase and reverberation in real time according to the music style or emotional arc, resulting in the "homogenization" of the mixing effects of different types of music; each functional module operates independently and lacks a cross-module collaborative optimization mechanism, which can neither convert user feedback into learning signals for system iteration nor coordinate the dynamic linkage requirements of element arrangement and mixing processing. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a data processing and production system for dynamic reorganization and automatic mixing of music elements, aiming to solve the problems of poor tonality adaptability and emotional expression disruption in the reorganization of music elements, homogeneity of cross-style mixing effects, insufficient collaborative optimization of creation modules, and low efficiency in converting user intentions.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a data processing and production system for dynamic reorganization and automatic mixing of musical elements, comprising: The music information deep representation and semantic understanding module is used to extract multi-dimensional features, analyze structure, and understand emotional semantics of externally input raw music data and user initial intentions, and output context information, context factor vectors, and semantic labels; A music element dynamic reorganization module is used to perform creative music element selection, arrangement and transformation based on received context information and high-level composition instructions, and output candidate music element reorganization schemes; An intelligent adaptive mixing module receives the music element reorganization scheme and the corresponding context factors, adjusts the mixing parameters in real time through a dynamic weight matrix mechanism, and outputs an audio stream or mixing parameter configuration that has undergone dynamic mixing processing; A multi-agent collaborative optimization and evaluation learning module is used to integrate information from each module to perform multi-objective evaluation and optimization on the complete music creation plan, generate high-level creation instructions, coordination instructions and learning signals, and drive knowledge updating; The user interaction and collaborative creation interface module is used to provide an interface for users to input creative intentions, browse system status, receive system feedback, and conduct iterative interactions; The knowledge base and learning results storage module is used to store and manage music theory rules, style models, successful creation patterns, user preferences, and various knowledge and parameters learned by the system, and provide decision support for other modules.

[0007] Preferably, the music information deep representation and semantic understanding module includes: a multimodal music feature extraction unit, for extracting acoustic features and symbolic features from the input raw music data; Music structure analysis and role identification unit, used to segment music data into segments, analyze the structural relationship between segments, and identify the structural role of each music segment; Music emotion semantic analysis and dynamic modeling unit, used to identify the emotional characteristics contained in music clips and model the dynamic arc of emotion changes over time; The context information synthesis and representation unit is used to integrate the various features, structural information and emotional semantics extracted by the above units, quantify and generate unified context information, context factor vectors and semantic labels.

[0008] Preferably, the music element dynamic reorganization module includes: The music element retrieval and screening unit is used to parse user instructions and contextual semantic information, and retrieve and screen candidate music elements that meet the target requirements from the music material library based on the calculated semantic similarity and structural matching degree; The music element transformation and adaptation unit is used to perform parameterized transformation on the selected candidate music elements, including unifying the mode and tonality, adjusting the rhythm and duration, and adapting the timbre to make them conform to the overall musical environment of the target music segment; The music fragment structure generation and arrangement unit is used to arrange the transformed and adapted music elements in time sequence, combine multiple parts and construct harmony according to music structure theory and arrangement rules, so as to generate a preliminary music fragment with a complete structure.

[0009] Preferably, the intelligent adaptive mixing module includes: A reorganization scheme and context receiving unit is used to receive an input music element reorganization scheme and corresponding context factors to provide a data basis for subsequent mixing processing; A dynamic weight parameter mapping unit, configured to calculate and generate dynamic adjustment weights of various mixing parameters based on context factors through a dynamic weight matrix mechanism; A multi-track audio intelligent mixing processing unit is used to adaptively adjust the volume, pan, balance, and effects of each track in the music element reorganization scheme in real time based on the generated dynamic adjustment weights; The mixing configuration and audio stream generation unit is used to integrate the adjusted mixing parameters of each audio track to form a complete mixing parameter configuration or output the final audio stream after dynamic mixing processing.

[0010] Preferably, the multi-agent collaborative optimization and evaluation learning module includes: The cross-module information fusion and user intention understanding unit is used to receive and integrate data from the music information deep representation and semantic understanding module, the music element dynamic reorganization module, the intelligent adaptive mixing module, and the user interaction and collaborative creation interface module, and analyze the creative intention proposed by the user; The music creation plan comprehensive evaluation unit is used to comprehensively evaluate the current music element reorganization plan and mixing configuration based on preset multi-dimensional indicators such as musicality, innovation, emotional expression accuracy, and user preferences; A reorganization and remixing collaborative optimization and instruction generation unit, which is used to execute the optimization algorithm based on the evaluation results, adjust the reorganization and remixing strategies, and generate high-level creation instructions for the dynamic reorganization module of the music elements and coordination instructions for each execution module; The system evolution learning and knowledge base feedback unit is used to extract learning signals from successful creation cases and optimization processes, drive the adaptive adjustment of system parameters, and feed back valuable experiences and patterns to the knowledge base and learning results storage module for updating.

[0011] Preferably, the user interaction and collaborative creation interface module includes: The creative intent semantic input and parsing unit is used to receive the creative intent input by the user through natural language, sample music or parameterized description, and convert it into instructions that the system can understand; A music creation process status visualization unit is used to display the internal status of the system, such as music information representation, element reorganization scheme, mixing parameter configuration, and the final generated music effect in real time; A multi-dimensional parameter interactive control unit allows users to intuitively fine-tune and intervene in the key parameters of musical element selection, reorganization logic, and mixing effects; The user feedback collection and iteration triggering unit is used to receive users' evaluations and modification suggestions on the current generation results, and convert these feedbacks into new iterative instructions to drive the system's re-creation.

[0012] Preferably, the knowledge base and learning achievement storage module includes: Music theory and style model storage unit, used for persistent storage and management of basic music theory rules, harmonic progression paradigms, and characteristic models of different music styles; A successful creation model and user preference recording unit, used to archive proven effective music element reorganization models, mixing strategy templates, and personalized creation preference data for different users; System learning parameter and evolutionary knowledge management unit, used to store and version control system model parameters, dynamic weight matrices, and knowledge generated during the evolutionary learning process driven by the multi-agent collaborative optimization and evaluation learning module; The knowledge retrieval and decision support supply unit is used to respond to requests from other modules, efficiently retrieve and provide the required music theories, style models, creation patterns or learning parameters to support their decision-making process.

[0013] Preferably, the music element transformation and adaptation unit realizes the mapping from the basic music element to the transformed music element through a parameterized transformation function, and the transformation function is: ; in, For the converted musical elements; is a parameterized conversion function; is the basic music element to be converted; is the target transformation parameter vector; is a set of internal parameters of the conversion function.

[0014] Preferably, the dynamic weight parameter mapping unit calculates and generates the dynamic adjustment weight of each mixing parameter, and the calculation method is: ; in, is the dynamic adjustment weight vector of the mixing parameters; is the mapping function; is the context factor vector; is the dynamic weight matrix.

[0015] Preferably, the remix collaborative optimization and instruction generation unit generates the composition instructions by optimizing a parameterized music composition strategy. The goal of the strategy optimization is to maximize a performance objective function defined as: ; in, is the performance objective function of the strategy; For the strategy Next track Take expectations; is the discount factor; At the moment of decision Immediate reward signals obtained; is the length of the trajectory.

[0016] The present invention provides a data processing and production system for dynamic reorganization and automatic mixing of musical elements. It has the following beneficial effects: 1. The present invention achieves semantic-level understanding and dynamic contextual representation of music data by constructing a multi-dimensional music information deep representation system and integrating acoustic feature extraction, structural analysis and emotional semantic modeling technologies. Compared with the limited traditional solutions that only rely on surface audio feature processing, this technical solution helps to solve the problems of incomplete capture of music emotional semantics and ambiguous structural role identification, and provides a more accurate semantic information foundation for subsequent creative links.

[0017] 2. With the help of an intelligent mixing mechanism driven by a dynamic weight matrix, the present invention enables the system to adjust the volume, pan and balance parameters in real time based on contextual factors, forming an adaptive mixing strategy. Compared with traditional fixed-parameter mixing templates, this helps to break through the bottleneck of the disconnection between mixing processing and musical emotional expression, solve the problems of insufficient consistency in cross-style music mixing effects and weak dynamic adaptability, and thus improve the emotional fit of the mixing results.

[0018] 3. The present invention adopts a multi-agent collaborative optimization framework, and realizes the iterative optimization of the creative plan and the evolution of the knowledge base through cross-module information fusion and multi-objective evaluation system. Compared with the traditional single-module independent working mode, this mode helps to solve the problems of poor coordination between the reorganization strategy and mixing parameters and low efficiency of user intention conversion in the creative process, and constructs a closed-loop learning system from evaluation to optimization, promoting the continuous improvement of creative quality.

[0019] 4. Based on parameterized conversion functions and semantic similarity retrieval technology, the present invention realizes the intelligent screening, transformation and structured arrangement of musical elements. Compared with traditional mechanical splicing or regularized reorganization methods, this solution solves the problems of poor tonal adaptability and insufficient style consistency in the element reorganization process, and improves the matching degree between the structural logic and creative intention of the generated music clips, expanding the creative space for automated music creation. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 Schematic diagram of the method flow of the present invention; Figure 2 Schematic diagram of the music information depth representation and semantic understanding module of the present invention; Figure 3 Schematic diagram of the music element dynamic reorganization module of the present invention; Figure 4 Schematic diagram of the intelligent adaptive mixing module of the present invention; Figure 5 Schematic diagram of the multi-agent collaborative optimization and evaluation learning module of the present invention; Figure 6 This is a schematic diagram of the user interaction and collaborative creation interface module of the present invention; Figure 7 Schematic diagram of the knowledge base and learning achievement storage module of the present invention. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0022] Please see the attached Figure 1 -Attached Figure 7 The embodiment of the present invention provides a data processing and production system for dynamic reorganization and automatic mixing of music elements, including: The music information deep representation and semantic understanding module is used to extract multi-dimensional features, analyze structure, and understand emotional semantics of externally input raw music data and user initial intentions, and output context information, context factor vectors, and semantic labels; Specifically, the music information deep representation and semantic understanding module includes: The multimodal music feature extraction unit, the music structure analysis and role recognition unit, the music emotion semantic analysis and dynamic modeling unit, and the context information synthesis and representation unit are connected in sequence through the internal data interface to collaboratively complete the in-depth processing of music information.

[0023] The multimodal music feature extraction unit is used to obtain original music data, which can be audio format files such as WAV and MP3, or symbol format files such as MIDI and MusicXML. The unit extracts acoustic features and symbol features from them.

[0024] For audio data, the unit performs short-time Fourier transform and filter bank analysis to calculate acoustic features such as Mel-frequency cepstral coefficients, chromaticity features, and spectral contrast.

[0025] At the same time, the unit analyzes the envelope and periodicity of the audio signal to extract rhythm information and loudness features.

[0026] If the original music data is in symbolic format, the unit directly parses the file structure and extracts symbolic features such as pitch, duration, dynamics, and instrument type.

[0027] All extracted feature data are organized into feature sets, attached with corresponding timestamps, and passed to subsequent units.

[0028] The music structure analysis and role recognition unit receives feature data or original music data. The unit first automatically segments the music stream and identifies the boundaries of segments with relatively independent musical meanings.

[0029] This segmentation process can be achieved based on the self-similarity matrix analysis of the music features, or by detecting novelty peaks in the feature sequence.

[0030] After segmentation, this unit analyzes the structural hierarchical relationships between the various musical segments. It also analyzes logical relationships such as repetition, contrast, and development, and identifies AABA and rondo structures in the music.

[0031] Furthermore, the unit assigns structural role labels to each identified music clip, such as verse, chorus, introduction, interlude, and ending.

[0032] This process can be completed based on a preset music theory rule library or through a trained machine learning classification model, and the processed structured music information is passed to the next unit.

[0033] The music emotion semantic analysis and dynamic modeling unit receives music feature data and structured music information. Its core function is to identify the emotional characteristics conveyed by the music clips and model the dynamic process of the emotion evolution over time.

[0034] First, the input music features are mapped to a predefined emotional space, which can be a two-dimensional valence-arousal space or a set of discrete emotion categories such as joy, sadness, etc.

[0035] This mapping is achieved through a pre-trained regression model, which can be support vector regression, or through a classification model, such as a deep neural network.

[0036] Subsequently, in order to capture the dynamic changes of musical emotions, a hidden Markov model is used for modeling, and the observed music feature sequence is assumed to be: ; in, is the observed music feature sequence; For the moment Observed musical features.

[0037] The corresponding emotional state sequence is: ; in, is a sequence of affective states; For the moment emotional state.

[0038] Hidden Markov Model The state transition probability matrix A, the observation probability matrix B, and the initial state probability distribution constitute; in, Represents the parameter set of the hidden Markov model; A is the state transition probability matrix; B is the observation probability matrix; is the probability distribution of the initial state.

[0039] The likelihood of a given observation sequence is calculated using the forward algorithm, as follows: ; in, For the model The characteristic sequence is observed The total probability of is a forward variable, indicating that At the end, the system is in an emotional state And the sequence has been observed probability; is the total number of possible affective states.

[0040] The initial values ​​of the forward variables are calculated as follows: ; in, The initial state is probability; In state Observed features probability.

[0041] The most likely emotional state sequence can be decoded through the Viterbi algorithm, which forms the emotional dynamic arc of the music. This emotional dynamic information is passed to the context information synthesis and representation unit.

[0042] The context information synthesis and representation unit is responsible for integrating multi-dimensional information from all the aforementioned units, including original features, structural role information, emotional dynamic information, and possible user preliminary intentions, such as expected style or emotion.

[0043] First, these heterogeneous data are aligned and fused to associate information from different sources onto a common music timeline.

[0044] The fused information is used to quantify and generate three core outputs, namely unified context information, context factor vector, and semantic label.

[0045] This output can be represented as a structured data object in its entirety; ; in, For data objects, It is the comprehensive understanding information of music described in a structured form; It is a low-dimensional dense digital vector that captures the core characteristics of music; It is a set of tags that describe the high-level semantics of music.

[0046] This embodiment performs multi-dimensional processing of original music data through a music information deep representation and semantic understanding module, realizing multimodal feature extraction, structural analysis, emotional dynamic modeling, and contextual comprehensive representation, enabling the system to deeply understand the music content from multiple dimensions and levels.

[0047] A music element dynamic reorganization module is used to perform creative music element selection, arrangement and transformation based on received context information and high-level composition instructions, and output candidate music element reorganization schemes; Specifically, the music element dynamic reorganization module includes: The music element retrieval and screening unit, the music element transformation and adaptation unit, and the music fragment structure generation and arrangement unit are connected in sequence and work together to realize the automatic construction of music fragments.

[0048] The music element retrieval and screening unit is responsible for searching for suitable music elements in the music material library based on the input information. Its input mainly includes the context factor vector and semantic label from the music information deep representation and semantic understanding module and the user's creation instructions. At the same time, each music element in the music material library has its own feature vector and structural attribute information pre-extracted and stored.

[0049] First, the user instructions and semantic tags are parsed to clarify the specific requirements of the target music clip, such as style, emotion, and structure; Then, for each music element in the library, the unit calculates its semantic similarity and structural matching with the target requirement.

[0050] Semantic similarity It can be obtained by calculating the cosine similarity between the target context factor vector and the music element feature vector: ; in, is the semantic similarity; is the context factor vector of the target demand; It is the characteristic vector of an element in the music material library.

[0051] Structural matching is a score obtained by comparing the structural properties of music elements with the structural characteristics of target requirements, and its value range can be normalized to between 0 and 1.

[0052] Then, to comprehensively evaluate the applicability of the music elements, the unit further calculates an overall relevance score, calculated as follows: ; in, Score the overall relevance of the musical elements to the target needs; is the weight coefficient of semantic similarity; is the weight coefficient of structural matching.

[0053] Finally, the music elements are sorted according to the calculated overall relevance score, and several elements with scores higher than a preset threshold or the highest score are selected as candidate sets and passed to the next unit.

[0054] The music element transformation and adaptation unit receives the screened candidate music elements, and aims to adjust these elements to make them conform to the overall music environment of the target music segment, including transformation and adaptation in terms of mode, tonality, rhythm, timbre, etc.

[0055] First, for mode and tonality, the pitch of the element can be transposed to make it consistent with the target tonality; for rhythm, the duration can be scaled or quantized to match the target speed and rhythm; for timbre, the instrument configuration can be adjusted or sound effects can be applied to make the transformed and adapted musical elements more compatible.

[0056] The music fragment structure generation and arrangement unit receives the transformed and adapted music elements, and organizes these music elements into preliminary music fragments with coherence and logic based on the music structure theory, the context structure information provided by the music information deep representation and semantic understanding module, or the arrangement rules specified by the user.

[0057] First, the temporal arrangement of the elements is performed to determine the order and duration of their appearance. Next, the overlapping and combining of the elements is processed to construct a polyphonic texture. Finally, musical rules such as counterpoint and harmony are applied to ensure the harmony and fluidity of the music. Finally, a preliminary, structurally complete musical fragment is output.

[0058] This implementation method realizes the intelligent selection and adaptive arrangement of music elements based on deep semantic understanding through the dynamic reorganization module of music elements, enabling the system to efficiently retrieve and reorganize music elements from the material library according to specific creative needs, and generate preliminary music fragments that meet specific semantic and structural requirements through transformation adaptation and structured arrangement, thus laying a good foundation for subsequent music refinement and mixing processing.

[0059] An intelligent adaptive mixing module receives the music element reorganization scheme and the corresponding context factors, adjusts the mixing parameters in real time through a dynamic weight matrix mechanism, and outputs an audio stream or mixing parameter configuration that has undergone dynamic mixing processing; Specifically, the intelligent adaptive mixing module includes: The reorganization scheme and context receiving unit, the dynamic weight parameter mapping unit, the multi-track audio intelligent mixing processing unit, and the mixing configuration and audio stream generation unit are connected in sequence and work together to realize intelligent adaptive mixing processing of the music element reorganization scheme.

[0060] The reorganization scheme and context receiving unit is responsible for receiving the music element reorganization scheme output by the music element dynamic reorganization module and the corresponding context factors output from the music information deep representation and semantic understanding module.

[0061] The received music element reorganization scheme and the context factors containing quantitative descriptions of the style, emotion, structure, etc. of the current music clip are format checked and preprocessed to provide a standardized data basis for the subsequent dynamic mixing process.

[0062] The core function of the dynamic weight parameter mapping unit is to calculate and generate dynamic adjustment weights for each mixing parameter based on context factors through a dynamic weight matrix mechanism, thereby realizing intelligent mapping of context information to specific mixing strategies.

[0063] Receive the context factor vector passed by the previous unit, and use the dynamic weight matrix that is preset or updated by the multi-agent collaborative optimization and evaluation learning module to calculate a set of dynamic adjustment weight vectors for different mixing parameters of each audio track (such as volume, pan, equalizer key frequency gain, effect send amount, etc.) through a specific mapping function. The calculation method is as follows: ; in, is the dynamic adjustment weight vector of the mixing parameters; is the mapping function; is the context factor vector; is the dynamic weight matrix.

[0064] This dynamically adjusted weight vector will guide subsequent units to make precise and adaptive adjustments to the mixing parameters.

[0065] The multi-track audio intelligent mixing processing unit performs real-time adaptive adjustment of parameters such as volume, pan, balance and effects of each audio track in the music element reorganization scheme based on the dynamic adjustment weights generated by the dynamic weight parameter mapping unit.

[0066] Dynamically adjust weights for each track's fundamental mixing parameters, avoid frequency conflicts through dynamic equalization, optimize the sound field through automated panning, and apply appropriate dynamic effects (compression, limiting) and spatial effects (reverb, delay) based on the musical style and emotion to ensure each track is clear, balanced, and expressive in the mix.

[0067] The mixing configuration and audio stream generation unit is responsible for integrating the mixing parameters of each audio track adjusted by the multi-track audio intelligent mixing processing unit, and based on this, forming a complete mixing parameter configuration file or outputting the final audio stream after dynamic mixing processing.

[0068] First, the final mixing parameters of each audio track (including volume, pan, equalization settings, effect parameters, etc.) are summarized to generate a mixing parameter profile that can be reviewed or reproduced in an external digital audio workstation.

[0069] Subsequently, all processed audio tracks are synthesized into a standard stereo or multi-channel audio stream, and necessary mastering processing (such as bus compression, overall equalization, and loudness normalization) is performed, and finally the finished audio file is output in the user-specified format (such as WAV, MP3).

[0070] This implementation utilizes an intelligent adaptive mixing module, enabling automated, intelligent mixing based on contextual understanding and a dynamic weighting mechanism. This module makes real-time, precise, and adaptive adjustments to each track's mixing parameters based on the specific content of the musical element reorganization scheme and the musical emotion and style indicated by contextual factors. This significantly improves mixing efficiency and enhances the professional listening experience and emotional accuracy of the final music.

[0071] A multi-agent collaborative optimization and evaluation learning module is used to integrate information from each module to perform multi-objective evaluation and optimization on the complete music creation plan, generate high-level creation instructions, coordination instructions and learning signals, and drive knowledge updating; Specifically, the multi-agent collaborative optimization and evaluation learning module includes: The cross-module information fusion and user intention understanding unit, the music creation plan comprehensive evaluation unit, the reorganization mixing collaborative optimization and instruction generation unit, and the system evolution learning and knowledge base feedback unit are connected in sequence and work together to achieve global optimization, evaluation and continuous learning of the entire music creation process.

[0072] The cross-module information fusion and user intention understanding unit is responsible for receiving and integrating data information from other major modules in the system, parsing user intentions, and receiving context factor vectors and semantic labels from the music information deep representation and semantic understanding module; it also receives the current music element reorganization scheme from the music element dynamic reorganization module, which contains the feature vectors of each music element and its information after being processed by a specific parameterized transformation function.

[0073] Receives the current mixing configuration from the smart adaptive mixing module, which reflects the results of applying the dynamically adjusted weights.

[0074] The dynamic adjustment weight is calculated based on the context factor vector and the dynamic weight matrix through a specific mapping function, as follows: ; in, is the dynamic adjustment weight vector of the mixing parameters; is the mapping function; is the context factor vector; is the dynamic weight matrix.

[0075] Receive data from the user interaction and collaborative creation interface module, conduct in-depth analysis of the creative intent input by the user, form a comprehensive understanding of the current creative status, system parameters and user goals, and provide a data basis for subsequent evaluation and optimization.

[0076] The music creation plan comprehensive evaluation unit is used to conduct a comprehensive evaluation of the current complete music creation plan (including the music element reorganization plan and mixing configuration) based on preset multi-dimensional indicators.

[0077] The evaluation is based on indicators such as musicality, innovation, emotional expression accuracy, and user preference. The emotional expression accuracy can be evaluated by calculating the semantic similarity between the target context factor vector and the overall feature vector of the current generated music clip.

[0078] At the same time, the rationality of the music structure and the matching degree between the structural characteristics of the target requirements are evaluated, which is recorded as Then, various evaluation indicators are combined to form a comprehensive performance evaluation score or multi-dimensional reward signal through weighted summation as follows: ; in, Score for comprehensive assessment; Rate the quality of the mix, is the matching degree between structural features; is the semantic similarity; is the semantic similarity weight; is the structural matching weight; is the mixing quality weight.

[0079] This evaluation result is output as the key input for the strategy optimization of the recombinant remix co-optimization and instruction generation unit.

[0080] The recombinant mixing collaborative optimization and instruction generation unit executes the optimization algorithm to adjust the overall creation strategy based on the evaluation results output by the music creation plan comprehensive evaluation unit.

[0081] By optimizing a parameterized music composition strategy to generate composition instructions, the goal of strategy optimization is to maximize a performance objective function defined as: ; in, is the performance objective function of the strategy; For the strategy Next track Take expectations; is the discount factor; At the moment of decision Immediate reward signals obtained; is the length of the trajectory.

[0082] Specific optimization directions include: Adjusting the tendency of element selection in the music element dynamic reorganization module, or the target conversion parameter vector or internal parameters in its music element parameterized conversion function; and adjusting the relevant parameters of the dynamic weight matrix or its mapping function used to calculate the dynamic adjustment weight in the intelligent adaptive mixing module.

[0083] Generate more refined high-level creation instructions for the dynamic reorganization module of music elements, as well as coordination instructions for the intelligent adaptive mixing module and other related execution modules, in order to obtain better evaluation results in the next round of creation.

[0084] The system evolution learning and knowledge base feedback unit is responsible for extracting learning signals and valuable knowledge from successful creation cases and continuous optimization processes.

[0085] Analyze high-performance creation trajectories and their corresponding parameter settings, and use these learning signals to drive the adaptive adjustment and evolution of the system's internal model parameters.

[0086] The extracted valuable experience, successful creation models and optimized system parameters are fed back to the knowledge base and learning outcome storage module for persistent storage and version control, ensuring the system's continuous learning, capability evolution and knowledge accumulation.

[0087] This implementation constructs an intelligent closed-loop feedback system through multi-agent collaborative optimization and evaluation learning modules to dynamically evaluate and continuously optimize the entire process of music creation.

[0088] The user interaction and collaborative creation interface module is used to provide an interface for users to input creative intentions, browse system status, receive system feedback, and conduct iterative interactions; Specifically, the user interaction and collaborative creation interface module includes: The creative intention semantic input and analysis unit, the music creation process status visualization unit, the multi-dimensional parameter interactive control unit, and the user feedback collection and iteration triggering unit are connected through data channels and control signals to jointly complete efficient interaction between users and the system.

[0089] The creative intention semantic input and parsing unit is used to receive the creative intention input by the user through natural language text, sample music input or parameterized control interface, and convert it into structured instructions that the system can understand.

[0090] Natural language input processing: Users can input natural language descriptions such as "compose a piano melody with a sad mood". The system extracts keywords and sentiment tags through semantic parsing models (such as BERT, T5, etc.) and generates a preliminary creation target vector.

[0091] Sample music input analysis: Supports users to upload reference audio in formats such as WAV, MP3, and MIDI. The music information deep representation and semantic understanding module extracts its style, structure, and emotional information and uses it as the target reference vector.

[0092] Parametric control input: Provides adjustable interface sliders or numerical input boxes, allowing users to explicitly set control parameters such as rhythm speed, emotional intensity, style category, harmonic complexity, etc. The system generates a target vector based on the parameters.

[0093] After parsing, the unit will output a creative intention vector and a set of semantic tags in a unified format, which will be passed to the music information deep representation module and the multi-agent optimization module for further processing.

[0094] The music creation process status visualization unit is responsible for visually presenting the system's current processing status, the intermediate results of each module, and the final music output, enhancing the user's control and understanding of the creation process, including but not limited to: Module operation status diagram: displays the real-time progress and data flow path of music representation, element reorganization, and mixing; Structure and emotion visualization: Based on the timeline, it displays the current music structure segmentation (such as ABA) and the dynamic curve of emotion changes (such as joy → anxiety → calm); Audio playback and waveform display: Supports real-time playback of generated music clips, and simultaneously displays audio waveforms, spectrograms, and rhythm grids, making it easy to check details such as timbre, rhythm, and strength; Mixing parameter layer control panel: Displays the volume, pan, EQ, and effect settings of each track in graphical form, allowing users to view the results of automatic system adjustments.

[0095] Through the above methods, users can participate in the system creation process visually throughout the entire process, which helps to timely discover and adjust creation deviations.

[0096] The multi-dimensional parameter interactive control unit provides users with an intuitive and adjustable interactive interface for fine-grained control and intervention of music element reorganization strategies, mixing parameters and emotional expression.

[0097] Music element selection control: allows users to manually add / exclude certain music elements, fine-tune the candidate element sorting results, and intervene in the element splicing logic; Structure control interface: Through the tree structure diagram or drag-and-drop timeline interface, users can adjust the order, length and repetition frequency of the music segments; Mixer Control Interface: Users can fine-tune volume, pan, EQ settings, reverb depth, and other parameters using sliders or numerical input, and can also enable / disable system automatic adjustments; Emotional expression control: Provides a two-dimensional emotional map (such as the valence-arousal space), where users can drag target points to guide the overall emotional trend of the music.

[0098] All control operations will be fed back to the system in real time, affecting the next step of generating results by updating the context factor vector or modifying the strategy weights.

[0099] The user feedback collection and iteration triggering unit is used to collect users' subjective evaluations and modification suggestions on the current generation results, and convert them into iterative optimization signals that can be executed by the system.

[0100] Feedback mechanism: Users can rate the music clips as a whole or in part (e.g., 1-5 stars), provide specific comments (e.g., "too fast tempo," "lack of emotional flow"), or use voice input to provide feedback; Evaluation label generation: After analyzing user feedback, the system generates structured evaluation labels and improvement instructions, such as "reduce speed" and "increase emotional contrast." Iteration trigger: After receiving valid feedback, the system automatically determines whether to trigger the re-creation process or enter the parameter optimization phase, and updates the objective function or strategy weight in the back-end optimization module.

[0101] User feedback can also be stored in the knowledge base for subsequent recommendations, preference modeling, and system adaptive learning.

[0102] This implementation method builds a creative interaction platform that is people-oriented, flexible and controllable, and supports iterative optimization by providing various forms of input analysis, system visualization display, parameterized control and feedback collection methods. It works in conjunction with other modules of the system to significantly enhance the personalized expression and interactive efficiency in the music creation process.

[0103] The knowledge base and learning results storage module is used to store and manage music theory rules, style models, successful creation patterns, user preferences, and various knowledge and parameters learned by the system, and provide decision support for other modules; Specifically, the knowledge base and learning outcomes storage module includes: Music theory and style model storage unit, successful creation mode and user preference recording unit, system learning parameter and evolution knowledge management unit, and knowledge retrieval and decision support supply unit.

[0104] The music theory and style model storage unit is used to persistently store basic music theory rules and style models.

[0105] First, a rule base covering scales, modes, and harmonic progressions is established through a database (such as PostgreSQL). Next, the structured data of harmonic progressions is stored in JSON format for fast query.

[0106] When storing the style model, vector representation is used to support the call of deep learning models so that they can be provided to the generation engine. The style model is indexed using FAISS to achieve efficient data retrieval through emotional and structural features.

[0107] The successful creation pattern and user preference recording unit archives verified effective music element reorganization patterns and user personalized preference data, uses machine learning algorithms to automatically extract patterns from historical creation data, and creates a template library.

[0108] The calculation relationship of the extraction mode can be expressed as: ; in, Create patterns for success; Indicates successful creation of pattern generation function; Represents user historical creation data; For rating or feedback information.

[0109] User preferences are dynamically updated through user behavior logs (e.g., element selection, regeneration frequency), and all extraction patterns and user preferences are stored in vector form for fast retrieval and call.

[0110] The system learning parameter and evolution knowledge management unit is used to store and version control system model parameters and dynamic weight matrix.

[0111] Parameters are saved in the version control system, supporting dynamic adjustment and historical record tracking. During the learning process, the mapping relationship between optimization results and strategies is managed using a graph database to facilitate parameter adjustment and system evolution in subsequent optimization.

[0112] The relationship between dynamic weight parameter mapping is: ; in, Indicates dynamic adjustment of weight vector; is the dynamic weight mapping function; context factor vector; is the dynamic weight matrix.

[0113] This unit is also responsible for summarizing the knowledge generated during the learning process and recording summaries of successful creation cases to achieve adaptive system learning.

[0114] The knowledge retrieval and decision support supply unit implements the knowledge support function for other modules, builds a hybrid retrieval system, combines ElasticSearch and vector databases (such as Milvus) to support the retrieval of music theory, style models and user preferences, implements context-based query processing, returns results through GraphQL, and supports multi-dimensional filtering and sorting.

[0115] This implementation method has storage and management capabilities through the knowledge base and learning outcome storage module for dynamic reorganization and mixing of music elements, while providing support for intelligent decision-making, realizing an innovative and personalized music creation process, and ensuring that it can quickly adapt to user needs and market changes in multi-dimensional music feature evaluation and optimization, providing a highly flexible creation environment.

[0116] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A data processing and production system for dynamic reorganization and automatic mixing of musical elements, characterized by: include: The music information deep representation and semantic understanding module is used to extract multi-dimensional features, analyze structure, and understand emotional semantics of externally input raw music data and user initial intentions, and output context information, context factor vectors, and semantic labels; A music element dynamic reorganization module is used to perform creative music element selection, arrangement and transformation based on received context information and high-level composition instructions, and output candidate music element reorganization schemes; An intelligent adaptive mixing module receives the music element reorganization scheme and the corresponding context factors, adjusts the mixing parameters in real time through a dynamic weight matrix mechanism, and outputs an audio stream or mixing parameter configuration that has undergone dynamic mixing processing; A multi-agent collaborative optimization and evaluation learning module is used to integrate information from each module to perform multi-objective evaluation and optimization on the complete music creation plan, generate high-level creation instructions, coordination instructions and learning signals, and drive knowledge updating; The user interaction and collaborative creation interface module is used to provide an interface for users to input creative intentions, browse system status, receive system feedback, and conduct iterative interactions; The knowledge base and learning results storage module is used to store and manage music theory rules, style models, successful creation patterns, user preferences, and various knowledge and parameters learned by the system, and provide decision support for other modules.

2. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 1 is characterized in that: The music information deep representation and semantic understanding module includes: a multimodal music feature extraction unit, for extracting acoustic features and symbolic features from the input raw music data; Music structure analysis and role identification unit, used to segment music data into segments, analyze the structural relationship between segments, and identify the structural role of each music segment; Music emotion semantic analysis and dynamic modeling unit, used to identify the emotional characteristics contained in music clips and model the dynamic arc of emotion changes over time; The context information integration and representation unit is used to integrate the various features, structural information and emotional semantics extracted by the above units, quantify and generate unified context information, context factor vectors and semantic labels.

3. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 1 is characterized in that: The music element dynamic reorganization module includes: The music element retrieval and screening unit is used to parse user instructions and contextual semantic information, and retrieve and screen candidate music elements that meet the target requirements from the music material library based on the calculated semantic similarity and structural matching degree; The music element transformation and adaptation unit is used to perform parameterized transformation on the selected candidate music elements, including unifying the mode and tonality, adjusting the rhythm and duration, and adapting the timbre to make them conform to the overall musical environment of the target music segment; The music fragment structure generation and arrangement unit is used to arrange the transformed and adapted music elements in time sequence, combine multiple parts and construct harmony according to music structure theory and arrangement rules, so as to generate a preliminary music fragment with a complete structure.

4. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 1 is characterized in that: The intelligent adaptive mixing module includes: A reorganization scheme and context receiving unit is used to receive an input music element reorganization scheme and corresponding context factors to provide a data basis for subsequent mixing processing; A dynamic weight parameter mapping unit, configured to calculate and generate dynamic adjustment weights of various mixing parameters based on context factors through a dynamic weight matrix mechanism; A multi-track audio intelligent mixing processing unit is used to adaptively adjust the volume, pan, balance, and effects of each track in the music element reorganization scheme in real time based on the generated dynamic adjustment weights; The mixing configuration and audio stream generation unit is used to integrate the adjusted mixing parameters of each audio track to form a complete mixing parameter configuration or output the final audio stream after dynamic mixing processing.

5. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 1 is characterized in that: The multi-agent collaborative optimization and evaluation learning module includes: The cross-module information fusion and user intention understanding unit is used to receive and integrate data from the music information deep representation and semantic understanding module, the music element dynamic reorganization module, the intelligent adaptive mixing module, and the user interaction and collaborative creation interface module, and analyze the creative intention proposed by the user; The music creation plan comprehensive evaluation unit is used to comprehensively evaluate the current music element reorganization plan and mixing configuration based on preset multi-dimensional indicators such as musicality, innovation, emotional expression accuracy, and user preferences; A reorganization and remixing collaborative optimization and instruction generation unit, which is used to execute the optimization algorithm based on the evaluation results, adjust the reorganization and remixing strategies, and generate high-level creation instructions for the dynamic reorganization module of the music elements and coordination instructions for each execution module; The system evolution learning and knowledge base feedback unit is used to extract learning signals from successful creation cases and optimization processes, drive the adaptive adjustment of system parameters, and feed back valuable experiences and patterns to the knowledge base and learning results storage module for updating.

6. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 1 is characterized in that: The user interaction and collaborative creation interface module includes: The creative intent semantic input and parsing unit is used to receive the creative intent input by the user through natural language, sample music or parameterized description, and convert it into instructions that the system can understand; A music creation process status visualization unit is used to display the internal status of the system, such as music information representation, element reorganization scheme, mixing parameter configuration, and the final generated music effect in real time; A multi-dimensional parameter interactive control unit allows users to intuitively fine-tune and intervene in the key parameters of musical element selection, reorganization logic, and mixing effects; The user feedback collection and iteration triggering unit is used to receive users' evaluations and modification suggestions on the current generation results, and convert these feedbacks into new iterative instructions to drive the system's re-creation.

7. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 1 is characterized in that: The knowledge base and learning achievement storage module include: Music theory and style model storage unit, used for persistent storage and management of basic music theory rules, harmonic progression paradigms, and characteristic models of different music styles; A successful creation model and user preference recording unit, used to archive proven effective music element reorganization models, mixing strategy templates, and personalized creation preference data for different users; System learning parameter and evolutionary knowledge management unit, used to store and version control system model parameters, dynamic weight matrices, and knowledge generated during the evolutionary learning process driven by the multi-agent collaborative optimization and evaluation learning module; The knowledge retrieval and decision support supply unit is used to respond to requests from other modules, efficiently retrieve and provide the required music theories, style models, creation patterns or learning parameters to support their decision-making process.

8. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 3 is characterized in that: The music element transformation and adaptation unit realizes the mapping from basic music elements to transformed music elements through a parameterized transformation function, and the transformation function is: ; in, For the converted musical elements; is a parameterized conversion function; is the basic music element to be converted; is the target transformation parameter vector; is a set of internal parameters of the conversion function.

9. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 4, characterized in that: The dynamic weight parameter mapping unit calculates and generates the dynamic adjustment weight of each mixing parameter, and the calculation method is: ; in, is the dynamic adjustment weight vector of the mixing parameters; is the mapping function; is the context factor vector; is the dynamic weight matrix.

10. The data processing and production system for dynamic reorganization and automatic mixing of music elements according to claim 5, characterized in that: The remix collaborative optimization and instruction generation unit generates creation instructions by optimizing a parameterized music creation strategy. The goal of the strategy optimization is to maximize a performance objective function defined as: ; in, is the performance objective function of the strategy; For the strategy Next track Take expectations; is the discount factor; At the moment of decision Immediate reward signals obtained; is the length of the trajectory.

Citation Information

Cited By

  • Blind music creation method and electronic equipment

    CN120954363A