A Method for Generating Long-Form Bibliographies of Suzhou Pingtan
By constructing a knowledge graph of Suzhou Pingtan and combining cross-domain transfer enhancement and reinforcement learning, the problem of creating long Suzhou Pingtan stories has been solved, and efficient human-computer collaborative creation has been achieved. The generated content is more in line with the multimodal characteristics of traditional Pingtan and the actual performance requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-03
AI Technical Summary
The creation of long Suzhou Pingtan (storytelling and ballad singing) plays a role in addressing issues such as insufficient dialect and rhythm modeling, mismatch between multimodal representation and alignment, incompatibility between existing technologies and creative needs, and a lack of systematic tools, resulting in poor AI-assisted creation.
We constructed a knowledge graph of Suzhou Pingtan (storytelling and ballad singing in Suzhou dialect), and generated a long-form bibliography that conforms to traditional structure and multimodal characteristics through cross-domain transfer enhancement and reinforcement learning strategies. We also optimized the creation process by combining audience feedback and user interaction.
It lowers the professional creation threshold, improves creation efficiency, and achieves deep integration of Pingtan art and human-computer collaborative creation, resulting in content that better meets actual performance needs.
Smart Images

Figure CN121389997B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the intersection of artificial intelligence technology and the inheritance of traditional culture, and in particular to a method for generating a long-form Suzhou Pingtan (storytelling and ballad singing) repertoire. Background Technology
[0002] Currently, both the creative practice and research on auxiliary technologies for Suzhou Pingtan long-form storytelling suffer from fundamental limitations. At the practical level, creation heavily relies on the individualized experience of senior playwrights, but such talent is facing a severe shortage. At the technological level, existing AI-assisted creation methods are systematically misaligned with the professional and multimodal needs of Suzhou Pingtan long-form storytelling. Specific technological deficiencies include:
[0003] 1. Deficiencies in Dialect and Prosody Modeling: Early rule-based engine methods relied on manually encoded phonological rules, which could not resolve the impromptu, cross-dialectal "local dialect" mixtures in performances. Current deep learning models (such as LSTM, GPT-4, etc.) suffer from training data bias and insufficient modeling of the Wu dialect's voiced and voiceless tone sandhi mechanisms (such as the fundamental frequency depression of voiced initials), resulting in inaccurate dialect prosody in the generated texts. Furthermore, they cannot distinguish the subtle emotions and musical characteristics of related schools such as the Jiang and Yu styles.
[0004] 2. Multimodal Representation and Alignment Mismatch: Pingtan art is a deep fusion of text, music, and performance. Existing digital archives are mostly text scripts, severely lacking key metadata such as stylized action codes, vocal spectrum, and live interaction (gimmicks). This causes the content generated by AI models to deviate from the actual performance; for example, the generated text does not associate "slapping the wooden clapper" with the corresponding "starting gesture," or the virtual human performance appears mechanical due to the lack of this multimodal information.
[0005] 3. Existing patented technologies are incompatible with the creative needs of Pingtan: Currently disclosed related patented technologies, such as the script generation method based on a large model (publication number CN119990078A) or the AIGC short drama processing system (authorization announcement number CN118632049B), are fundamentally incompatible with the creative principles of long-form Pingtan. They either lack dialect adaptability or are limited by the compact narrative logic of short dramas, making them unable to adapt to the traditional structure of long-form Pingtan with its episodic structure and the contextualized adaptation needs of different singing styles. Furthermore, they generally lack cultural training and copyright adaptation for Wu dialect and the musicality of Pingtan, making them unsuitable for direct application.
[0006] 4. Lack of Systematic Tools and Collaborative Mechanisms: Current technological intervention is fragmented, lacking systematic creative tools. Knowledge transfer during the creative process is unstructured, expert experience remains isolated, and human-computer collaborative interaction mechanisms are completely absent. Furthermore, the difficulty in standardizing musical parameters further increases the barriers to technological intervention.
[0007] In summary, the core technical challenges facing the creation of long-form Suzhou Pingtan repertoire can be summarized as: how to construct a systematic technical solution that can deeply integrate the inherent characteristics of Pingtan art (Wu dialect rhythm, singing styles, and performance movements), understand its long narrative structure, and achieve human-machine collaborative innovation. Summary of the Invention
[0008] To address the aforementioned problems, the purpose of this invention is to propose a method for generating long-form Suzhou Pingtan (storytelling and ballad singing) repertoires, thereby enabling the creation of long-form Suzhou Pingtan repertoires.
[0009] The technical solution of this invention is as follows: A method for generating a long-form Suzhou Pingtan (storytelling and ballad singing) repertoire, comprising:
[0010] Based on the constructed Suzhou Pingtan knowledge graph, the system responds to users' creative needs and outputs knowledge graph features.
[0011] Cross-domain transfer enhancement is performed on the features of the knowledge graph to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features;
[0012] By integrating the knowledge graph features and the cross-domain enhancement features, the long-form Suzhou Pingtan repertoire is optimized through a reinforcement learning strategy, and the final result is output.
[0013] Furthermore, this includes updating the Suzhou Pingtan knowledge graph, specifically including:
[0014] Based on audience feedback ratings of the real-time performance, an improvisation event generation mechanism is triggered. This mechanism creates improvisation event nodes to update the Suzhou Pingtan knowledge graph. These improvisation event nodes include timestamps, performance techniques, actors, and bibliographies extracted from the performance context.
[0015] Furthermore, this includes obtaining digital human motion parameters and VR scene instructions based on the output knowledge graph features, driving the digital human to perform in the VR scene, interacting with users to obtain user behavior data streams, dynamically adjusting the knowledge weights in the cognitive engine based on user behavior data, changing the reasoning priority, and optimizing prompt words based on user feedback, performing prompt word self-iteration to guide the cognitive engine and user interaction behavior instructions or templates.
[0016] Furthermore, cross-domain transfer enhancement is performed on the features of the knowledge graph to generate a long-form Suzhou Pingtan bibliography with cross-domain enhancement features, specifically including:
[0017] The knowledge graph features are analyzed, and cross-domain features are extracted from data of other art forms different from Suzhou Pingtan based on the analysis results. The cross-domain features include cross-art text features, cross-art audio features, and cross-art performance features. The cross-art text features include at least Kunqu Opera text features, the cross-art audio features include at least Beijing Drum Music audio features, and the cross-art performance features include at least Kunqu Opera performance features.
[0018] Calculate the similarity between the knowledge graph features and the cross-domain features to generate a cross-domain feature transformation matrix;
[0019] The cross-domain features are converted into cross-domain enhanced features through the cross-domain feature transformation matrix to generate a long list of Suzhou Pingtan stories with cross-domain enhanced features.
[0020] Furthermore, after generating a long Suzhou Pingtan (storytelling and ballad singing) repertoire with cross-domain enhancement features, a traditionality verification score is obtained based on the metrical conformity detection of the generated long Suzhou Pingtan repertoire with cross-domain enhancement features. The Euclidean distance between the text vector of the generated long Suzhou Pingtan repertoire with cross-domain enhancement features and the average semantic vector of the documents in the reference corpus is calculated, and the innovativeness verification score is obtained from the Euclidean distance.
[0021] A comprehensive quality index is calculated based on the traditional verification score and the innovative verification score. A quality index threshold is set. If the comprehensive quality index does not reach the quality index threshold, the cross-domain enhancement features are re-converted to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features until the comprehensive quality index reaches the quality index threshold.
[0022] Furthermore, the generation of long-form Suzhou Pingtan (storytelling and ballad singing) repertoire with cross-domain enhancement features includes:
[0023] Based on the cross-domain feature transformation matrix and the cross-art text features, a literary-enhanced Pingtan text is generated: the hidden state of the Pingtan generator is initialized, the matching degree between the context and each structural unit is calculated through MLP, the selected unit is obtained by weighting and projected onto the Pingtan feature space, the gating vector is calculated based on the hidden state and the selected unit, the migration intensity is controlled by the text migration weight, the cross-art text features and the original Pingtan features are fused by element-wise gating, the hidden state is updated and the literary-enhanced Pingtan text is generated by decoding.
[0024] Based on the cross-domain feature transformation matrix and the cross-art audio features, a prosodic code conforming to the rhyme scheme of Pingtan is generated and aligned with the text semantics and sentiment: a rhythm sub-matrix is extracted from the cross-domain feature transformation matrix to generate a timestamp sequence. A conditional vector is constructed by combining the statistical features of the text semantics and the timestamp sequence. Based on the audio transfer weight cross-domain rhythm features, a prosodic code is generated by a conditional variational autoencoder.
[0025] The text transfer weights and audio transfer weights are adjusted when the cross-domain enhancement features are reconverted.
[0026] Furthermore, the generation of long-form Suzhou Pingtan (storytelling and ballad singing in Suzhou dialect) repertoire with cross-domain enhancement features includes generating performance guidelines adapted to Pingtan based on cross-art performance features:
[0027] The performance features of Kunqu Opera are mapped to the Pingtan space by projection matrix. The style adaptation is optimized by combining style mixing coefficient and Pingtan default parameters. Then, the performance intensity coefficient is predicted based on text semantics and prosodic encoding to obtain the final performance parameters.
[0028] Based on the final performance parameters, a sequence of body movements is generated, the original gesture symbols are mapped to the gesture sequence of the Pingtan system, vocal performance parameters are generated, and the final emotional expression parameters are obtained by fusing the emotional features in the text based on the basic emotional expression.
[0029] Construct a synchronization matrix between body movements and text and calculate the alignment loss. Calculate the multimodal coordination loss of body movements, gestures, vocal performance parameters, and final emotional expression parameters to maintain temporal coordination and stylistic consistency.
[0030] Furthermore, a graph attention network is used to calculate the similarity between the knowledge graph features and the cross-domain features.
[0031] Furthermore, by integrating the knowledge graph features and the cross-domain enhancement features, the long-form Suzhou Pingtan repertoire is optimized through reinforcement learning strategies, specifically including:
[0032] A state space is constructed, which is a fusion encoding of the knowledge graph features, the cross-domain enhancement features, and the creation requirements;
[0033] A reinforcement learning model is trained based on a multi-objective reward function and a proximal policy optimization algorithm. The policy network of the reinforcement learning model selects actions from the mixed action space according to the current state, generates the optimal action sequence using the trained policy network, and optimizes the bibliography.
[0034] The hybrid action space includes discrete actions for content structure adjustment and continuous actions for parameter fine-tuning. The multi-objective reward function balances at least three objectives: tradition, innovation, and user satisfaction, and introduces a dynamic penalty and reward mechanism.
[0035] Furthermore, based on the constructed Suzhou Pingtan knowledge graph, the first intelligent agent responds to the user's creative needs and outputs knowledge graph features;
[0036] The second intelligent agent performs cross-domain transfer enhancement on the features of the knowledge graph to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features;
[0037] The third agent integrates the knowledge graph features and the cross-domain enhancement features, optimizes the long-form Suzhou Pingtan repertoire through a reinforcement learning strategy, and outputs the final result.
[0038] Compared with the prior art, the advantages of this invention are as follows:
[0039] This invention integrates knowledge graphs, cross-domain transfer, and reinforcement learning into a framework for innovative creation of traditional culture, which helps to lower the professional creation threshold and improve the creation efficiency of long Suzhou Pingtan (storytelling and ballad singing) plays. Attached Figure Description
[0040] Figure 1 A flowchart illustrating the method for generating long-form Suzhou Pingtan (storytelling and ballad singing) repertoire according to the present invention.
[0041] Figure 2 A diagram illustrating the knowledge graph of teacher-student relationships.
[0042] Figure 3 A schematic diagram of the workflow of the first intelligent agent according to an embodiment of the present invention.
[0043] Figure 4 A schematic diagram of the workflow of the second intelligent agent in an embodiment of the present invention.
[0044] Figure 5 A schematic diagram of the workflow of the third intelligent agent in an embodiment of the present invention. Detailed Implementation
[0045] To more clearly illustrate the technical solution of the present invention, the present invention will be further described in detail below with reference to embodiments.
[0046] This invention provides a method for generating a long-form Suzhou Pingtan (storytelling and ballad singing) bibliography, which involves multiple intelligent agents collaborating to generate the bibliography. The generation method is as follows: Figure 1 As shown, specifically, it includes: the first intelligent agent, based on the constructed Suzhou Pingtan knowledge graph, responds to the user's creative needs and outputs knowledge graph features by a hybrid cognitive engine; the second intelligent agent performs cross-domain transfer enhancement on the knowledge graph features to generate a long Suzhou Pingtan bibliography with cross-domain enhancement features; the third intelligent agent integrates knowledge graph features and cross-domain enhancement features, optimizes the long Suzhou Pingtan bibliography through reinforcement learning strategies, and outputs the final result.
[0047] Please combine Figure 3 As shown, in one embodiment of the present invention, the specific processing procedure of the first intelligent agent includes the following:
[0048] Step A1: Define the core ontology structure of the Pingtan domain based on the OpenSPG semantic framework.
[0049] We construct a high-level semantic model for the Suzhou Pingtan (storytelling and ballad singing) domain, namely, its core ontology structure. This ontology structure serves as the semantic blueprint for the entire knowledge graph, defining core concepts (entities), concept attributes, and static and dynamic relationships between concepts. The specific definitions are as follows:
[0050] (1) Define the following core entity types:
[0051] The "Bibliography" entity: This entity represents specific performance works of Suzhou Pingtan (a type of storytelling and ballad singing in Suzhou dialect). This entity contains the following key attributes:
[0052] Title: The formal title of the book, such as "The Pearl Tower".
[0053] Dynasty of Creation: The historical period in which the book was created or was mainly popular.
[0054] Chapter structure: Used to describe the compositional structure of a bibliography.
[0055] The "Actor" entity represents a performing artist of Suzhou Pingtan (a traditional Chinese storytelling and ballad singing art). This entity contains the following key attributes:
[0056] Name: The actor's name.
[0057] Genre: The artistic genre to which the actor belongs. This attribute is limited to an enumerated value, such as "Ma Tune" or "Yu Tune".
[0058] Master-disciple relationship: This attribute or relationship is used to establish connections with other entities or relationships that describe the master-disciple lineage in subsequent steps.
[0059] The "Musical Instrument" entity is used to represent the main accompanying musical instruments used in Pingtan performances, such as the sanxian and pipa.
[0060] (2) In order to depict the dynamic, time-dependent performance behaviors in Pingtan art, the following event types are defined:
[0061] "Performance Event": Used to represent a complete Pingtan performance, which can be associated with multiple actors, instruments and performance venues.
[0062] "Improvisational Event": Used to characterize specific improvisational acts performed by an actor during a performance. This event contains the following key attributes:
[0063] Techniques employed: Describe the specific techniques used in improvisation, such as "gimmicks" and "speaking jokes."
[0064] Occurrence timestamp: Precisely records the exact time when the impromptu performance occurred.
[0065] Executor: Establish an association with the executor of the event, i.e., an "actor" entity.
[0066] Related bibliography: Establishes a connection with the specific "bibliographical" entity context in which the improvisation takes place.
[0067] (3) Based on the above definitions, establish a semantic network between entities and events:
[0068] The “bibliography” consists of multiple “chaps”.
[0069] "Actors" belong to a specific "genre".
[0070] An "improvisational performance event" is triggered by a specific "actor" performing a specific "book".
[0071] A "master-apprentice relationship" can be established between "actors" entities, forming a lineage chart.
[0072] Step A2: Based on the defined ontology structure, integrate multi-source data to construct a Suzhou Pingtan knowledge graph.
[0073] Step A2.1: Import structured data such as long-form Suzhou Pingtan repertoire, actors, and musical instruments that have been entered into the public domain, map the imported data to the ontology structure defined in Step 1, and establish semantic relationships between entities (such as actor-repertoire performance relationship).
[0074] Step A2.2: Collect Suzhou Pingtan texts, audio materials, and video materials from the public domain, process unstructured data, form a knowledge graph of master-apprentice relationships, and extract audio features and performance features from support vector indexes.
[0075] The specific implementation process is as follows:
[0076] Step A2.2.1: Based on OpenIE, parse the master-disciple relationship from the text, map the extracted relationship to the actor entity's school and master-disciple relationship attributes, and construct a master-disciple relationship knowledge graph based on ontology structure.
[0077] A specific example is as follows:
[0078] (1) Collect the personal biographies and memoirs of Pingtan artists who have entered the public domain, as well as the publicly published research papers and literary reviews (such as “Passing on the Torch: A Study on the Inheritance Mechanism of Suzhou Pingtan since the Late Qing Dynasty”, “The Construction of the Inheritance System of Suzhou Pingtan in Modern Times - Centered on Guangyu Society” and “A Brief Discussion on the Master-Master Tradition of Suzhou Pingtan”) in electronic or scanned form, convert them into text using OCR technology, remove irrelevant information (such as headers and footers), divide the text into chapters and paragraphs, and standardize punctuation marks and special characters.
[0079] (2) Using Chinese NLP tools such as NLPIR-ICTCLAS or frameworks such as HanLP and Stanford CoreNLP, construct a dictionary of Suzhou Pingtan master-apprentice inheritance, including the names of Pingtan artists and the names of schools.
[0080] (3) Identify typical master-disciple relationship expression patterns, such as: “learned from someone”, “studied by someone”, “disciple of someone”, “someone passed on their skills to someone” or “someone taught someone”, etc., and select the results that conform to the characteristics of master-disciple relationship from the triples extracted by OpenIE.
[0081] (4) Analyze the contextual information of the apprentice-student relationship description, identify auxiliary information such as fellow apprentices, apprenticeship time, and apprenticeship location, and use this information to verify and enrich the extracted relationship.
[0082] (5) Verify the consistency of the description of the same mentor-apprentice relationship in different paragraphs, infer the missing mentor-apprentice relationship by using information such as community organizational structure and timeline, and verify the extraction results manually by domain experts. Construct a feedback mechanism and continuously improve the extraction rules.
[0083] (6) Organize the extracted mentor-apprentice relationships into a knowledge graph. A specific implementation of a mentor-apprentice relationship knowledge graph is as follows: Figure 2 As shown.
[0084] Step A2.2.2: Based on audio processing and feature extraction technology, the MFCC feature vectors of the sanxian and pipa timbres are parsed from the Suzhou Pingtan audio data, and the audio features are associated with the performance events to construct an audio feature database that supports vector retrieval.
[0085] A specific example is as follows:
[0086] (1) Collect high-quality audio materials of representative long Suzhou Pingtan stories by different artists and schools with public authorization or proprietary copyright; divide the audio into shorter segments (e.g., 30 seconds to 2 minutes) according to the chapters or themes of the long stories, use audio processing tools (e.g., librosa, pyAudioAnalysis) to remove background noise, and manually or automatically label the main instruments (sanxian / pipa) in each segment.
[0087] (2) Use the librosa library to extract MFCC features. For each audio segment, set appropriate parameters to extract MFCC coefficients (usually 13-40 coefficients), extract delta and delta-delta features to capture dynamic changes, and normalize the features to ensure the consistency of the feature vectors.
[0088] (3) If the MFCC feature dimension is greater than 128, proceed to step (4); otherwise, proceed to step (5).
[0089] (4) The original MFCC features are processed by dimensionality reduction algorithms such as PCA and t-SNE or deep learning models such as Autoencoder, and compressed to 128 dimensions, while ensuring that the vector after dimensionality reduction can effectively retain the main information of the original features.
[0090] (5) Traditional databases that support vector indexes, such as PostgreSQL+pgvector, record the following data: audio segment ID, original audio path, 128-dimensional feature vector, instrument type (sanxian / pipa), book title, artist information and timestamp.
[0091] Step A2.2.3: Based on computer vision technology, analyze the actors' gestures and expressions from Suzhou Pingtan performance videos, associate visual features with improvisational performance events, and construct a performance feature database that supports multi-dimensional retrieval.
[0092] A specific example is as follows:
[0093] (1) Collect high-quality video materials of representative long stories of Suzhou Pingtan by different artists and schools, either publicly authorized or with their own copyright, unify the video format and resolution, use video processing tools to remove noise and jitter, and extract segments containing the main performance content.
[0094] (2) Detect pixel differences between adjacent frames, identify scene transition points, detect significant changes in actors' movements, and identify key frames; use frameworks such as MediaPipe and OpenPose to detect key points of the hands, track hand movement trajectories, and identify gesture changes; define typical gesture categories in Pingtan (such as orchid finger, sword gesture, palm support, etc.), use CNN or LSTM models to classify gestures, establish a Pingtan gesture dictionary, and record the meaning and application scenarios of each gesture;
[0095] (3) Use frameworks such as FaceMesh and dlib to detect facial key points, extract facial features, define typical expression categories in Pingtan (such as joy, anger, sorrow, happiness, etc.), use CNN or Transformer models to classify expressions, and analyze the meaning and emotional expression of expressions in combination with context.
[0096] (4) Design a relational database (such as MySQL) to store the recognition results. Each record includes: video ID, key frame timestamp, gesture category and confidence level, expression category and confidence level, associated bibliographic content, data index and retrieval, establish time- and content-based indexes, and support multi-dimensional retrieval by gesture, expression, time and other criteria.
[0097] Step A2.3: Establish a unified mapping relationship between "ontology entities and multimodal features" to achieve cross-modal association retrieval capabilities and form a Suzhou Pingtan knowledge graph.
[0098] Step A2.3.1: Based on the structured entity data output in Step 2.1 and the multimodal features output in Step 2.2, construct the entity center index.
[0099] Step A2.3.2: Based on the multimodal feature data output in step A2.2, construct a modal inverse index.
[0100] Step A2.3.3: Construct a unified knowledge graph and cross-modal retrieval system. The specific implementation process is as follows:
[0101] (1) Based on the entity center index established in step A2.3.1 and the modal inverse index constructed in step A2.3.2, deep semantic fusion is performed to construct a unified graph structure knowledge representation. The graph structure contains three types of core nodes: entity nodes inherit from the ontology structure defined in step A2.1, feature nodes are derived from the multimodal features extracted in step A2.2, and relation nodes integrate the semantic relations in step A2.2.1; multiple types of semantic edges such as "has_audio_feature", "has_gesture", "performed_in", "apprentice_of" and "similar_to" are defined and established to form a complete knowledge graph topology.
[0102] (2) Construct a cross-modal unified retrieval engine to realize complex multimodal association query functions. This engine supports composite queries based on graph traversal and can simultaneously utilize the aggregation characteristics of the entity center index and the retrieval capabilities of the modal reverse index to realize a retrieval mechanism that initiates queries from any modal entry point and obtains cross-modal association results.
[0103] In an embodiment of the present invention, the first intelligent agent further performs step A3, updating the Suzhou Pingtan knowledge graph: by capturing real-time events and analyzing the evolution of singing styles, the Suzhou Pingtan knowledge graph is updated. Specifically, this includes the following steps:
[0104] Step A3.1: Capture real-time events, generate impromptu performance event nodes, and update them to the Suzhou Pingtan knowledge graph.
[0105] (1) The intensity of the audience's applause was collected in real time by audio sensors deployed in the storytelling venue. It is the intensity of the i-th applause and the number of laughs. The audience density is obtained by using an array of infrared depth cameras deployed on the roof of the theater to acquire real-time 3D point cloud data of the seating area via TOF (Time-of-Flight) technology. Infrared depth cameras or high-definition cameras deployed in the storytelling venue are used to identify audience facial features in real time and estimate age distribution using face detection and age estimation algorithms (such as convolutional neural network-based models). Record the temporal context of the performance. This includes the time of day, the day of the week, and whether it is a special time period.
[0106] (2) Calculate the sentiment score Among them, the weighting coefficient Using an online prediction model with LSTM ( Dynamic adjustment, i.e. .
[0107] (3) When the sentiment score ( Exceeding the preset threshold ( And the duration meets the threshold. ( When an improvisation event is triggered, the event generation mechanism is activated. The event generation module extracts the current actor and book information from the performance context, combines audio pattern recognition technology to detect improvisation techniques, and creates a complete improvisation event node. This node includes attributes such as timestamp, technique used, performing actor, and associated book, and is updated in real time to the knowledge graph, establishing relationships with relevant entities.
[0108] Step A3.2: Extract hierarchical features from the audio of Suzhou Pingtan singing style, and update the schools and evolutionary relationships in the Suzhou Pingtan knowledge graph from raw data to evolutionary rules.
[0109] Through hierarchical abstraction, hierarchical knowledge of "variant features → musicological patterns → schools and inheritance rules" is extracted from the original singing audio, thereby obtaining the core dimensions (such as melody, rhythm, and embellishment) and evolutionary logic of Pingtan singing variations. The definition of the CoA (Chain of Abstraction) abstraction hierarchy is detailed in Table 1.
[0110] Table 1. CoA Abstraction Levels
[0111]
[0112] Step A3.2.1: Extract acoustic features (L1) from the original audio (L0). The specific implementation process is as follows:
[0113] (1) The original Pingtan singing audio was processed in frames using an audio processing pipeline, with each frame having a length of 10 milliseconds.
[0114] (2) Use the YIN (Yin Fundamental Frequency Estimator) algorithm or the CREPE model (Convolutional Representation for Pitch Estimation is a fundamental frequency extraction method based on deep convolutional neural networks) to extract the fundamental frequency value of each frame and form a melody contour curve. Analyze the audio signal through the autocorrelation function to detect the beat position and identify traditional plate structure (such as one plate with three beats).
[0115] (3) Calculate the trend of the number of beats per minute and the density of notes per unit time.
[0116] (4) Use the music information retrieval tool Essentia to extract features such as MFCC, spectral centroid, and zero-crossing rate, and generate a feature vector for each musical phrase containing a 100-dimensional F0 sequence, a list of rhythmic events, and a 13-dimensional MFCC mean.
[0117] (5) All feature data and metadata (genre, actor, recording time) are stored together as a structured table.
[0118] Step A3.2.2: Map the acoustic features (L1) to the musicological elements (L2).
[0119] Based on a predefined dictionary of Suzhou Pingtan musicological elements (see Table 2), conversion rules from acoustic features to musical concepts were established. For melody analysis, the interval changes between adjacent pitch sequences were calculated, and stepwise (≤2 degrees) and leap (≥3 degrees) patterns were identified. For rhythm analysis, audio segments were classified into slow, medium, or fast tempos based on beat period and note density. For embellishment techniques, rule-based detection algorithms were employed, such as searching for repeated pitch patterns within a short period (100ms) when identifying overlapping embellishments. All mapping rules were validated and optimized using pre-trained machine learning models (random forest, SVM) to ensure a classification accuracy of no less than 90%.
[0120] Table 2 Definitions of Musicological Elements
[0121]
[0122] Step A3.2.3: Cluster genre patterns (L3) from musicological elements (L2). The specific implementation process is as follows:
[0123] (1) Use unsupervised learning algorithms (such as K-means, hierarchical clustering) to cluster musicological features and discover potential variant groups.
[0124] (2) Combine genre labels and train variant classification models using supervised learning (such as SVM, neural networks). Define the feature label set for each genre. Associate the clustering results with actors, eras, and representative arias, and store them as knowledge graph nodes.
[0125] Step A3.2.4: Derive evolutionary rules from variant patterns (L4). The specific implementation process is as follows:
[0126] (1) Multi-dimensional comparative analysis to explore the driving factors of vocal style evolution. By comparing the characteristic differences of different generations of inheritors of the same school, we can analyze the influence of inheritance; by comparing the characteristic changes of singing segments in different historical periods, we can analyze the changes of the times; by detecting the degree of infiltration of elements from other music genres, we can analyze cultural integration.
[0127] (2) Establish a statistical model or graphical model to quantify causal relationships.
[0128] The following is a specific implementation model: ,in, It is variant complexity; It is a branch of algebra (integers, For example, "the first generation successor". "Second-generation successor" (and so on). It is a variable of the times (a binary variable, Indicates modern (after the 2000s). (Indicates tradition (before the 2000s)). Is it the degree of integration of elements from Shanghai Opera or other opera genres? For example, "Shanghai Opera elements account for 30%" ).
[0129] Step A3.2.5: Update the schools and evolutionary relationships in the Suzhou Pingtan knowledge graph. The specific implementation process is as follows:
[0130] (1) For each evolutionary law derived from step 3.2.4, perform the following operations:
[0131] (a) Check whether the same evolutionary pattern already exists in the knowledge graph.
[0132] (b) If not, create an evolutionary pattern entity node in the knowledge graph. Each evolutionary pattern node contains the following core attributes: pattern identifier (a unique ID that identifies the evolutionary pattern, using the naming convention "EVOL_[style name]_[timestamp]"), mathematical model description (storing the specific mathematical expression), set of driving factors (recording key factors affecting the evolution of singing style, including lineage, historical context, cultural integration, etc.), confidence score (the confidence level of the pattern based on statistical significance testing, ranging from 0 to 1), applicable time range (the effective time period of the evolutionary pattern), and data source (the audio samples and feature data on which the pattern is derived).
[0133] (c) If it already exists, update the properties of the existing node and record the updated version.
[0134] Step A4: Deploy a knowledge graph-based hybrid cognitive engine, which responds to user creation needs and outputs knowledge graph features through multimodal intent parsing and hybrid reasoning.
[0135] Step A4.1: The hybrid cognitive engine receives and processes diverse user input, deeply understands query intent, and pays special attention to queries related to genre characteristics, mentor-apprentice relationships, and cross-modal rules.
[0136] (1) Multimodal input interface.
[0137] Supports multiple input methods including natural language text, voice input, and structured queries; voice input is converted to text via an ASR system, preserving paralinguistic information such as intonation and pauses; a query type classifier is built to identify query intent: comparative analysis, trend queries, association discovery, recommendation requests, etc.; the classifier prioritizes the identification of genre features. Master-disciple relationship Cross-modal rules Related intentions.
[0138] (2) Deep semantic parsing.
[0139] The domain-adaptive BERT model is used for named entity recognition to accurately extract entities from the Pingtan domain; semantic role labeling is used to analyze the predicate-argument structure of queries; and a query logic tree is constructed to represent complex nested query relationships.
[0140] (3) Context-aware query enhancement.
[0141] By combining user query history and preferences, implicit query conditions are supplemented; the problem of referential resolution is solved using session context; and a standardized logical representation is generated, which is directly mapped to the knowledge graph. .
[0142] Step A4.2: Perform hybrid reasoning. Combine multiple reasoning modes to extract and deduce deep knowledge from the knowledge graph, generating... .
[0143] (1) Map query and data extraction.
[0144] Generate optimized SPARQL queries, leveraging graph indexes to improve query efficiency. Querying genre entities and their attributes (such as origin, representative works, artistic characteristics), for To query the master-apprentice relationship network (such as the master-apprentice chain, the time of inheritance), for Query cross-modal associations (such as mapping rules between text keywords and audio features); perform multi-hop queries, traverse related entities and relationship networks; extract raw data and build temporary data views.
[0145] (2) Apply a predefined business rule base to perform rule reasoning.
[0146] Rules: Define rules for the inheritance of genre characteristics (e.g., "If genre A influences genre B, then B inherits some characteristics of A").
[0147] Rule: Define the transitivity of the master-apprentice relationship (e.g., "the master's master is also a master-apprentice").
[0148] Rule: Define cross-modal mapping (e.g., "the imagery of 'garden' in the text corresponds to a soft rhythm in the audio").
[0149] Step A4.3: Synthesizing and enhancing cognitive output.
[0150] Transforming reasoning results into cognitively valuable outputs, structured knowledge graph features. It also generates natural language descriptions to support different user types.
[0151] Step A4.3.1: Output a structured JSON object It comprises three core components, each internally divided into hierarchical substructures to ensure data integrity and machine readability:
[0152] (1) (School of thought characteristics): This includes the school of thought's raw data, statistical summaries, insights, and decision support sub-layers.
[0153] (2) (Mentor-Apprentice Relationship): Covering the original network, statistical indicators, pattern recognition, and recommendations for mentor-apprentice relationships.
[0154] (3) (Cross-modal rules): Includes rule definition, applied statistics, anomaly detection, and optimized prediction.
[0155] Step A4.3.3: Encapsulate interaction parameters and output them for digital human motion mapping and VR scene command generation. This includes:
[0156] (1) Action parameter instruction: Parse the self-inference result and use it for the digital human action drive in step A5.1.
[0157] Example fields: action_parameters:
[0158] technique_name (technique name, such as "finger rolling")
[0159] timestamps (lyrics or event timestamps)
[0160] context_entities (context entities extracted from NER, such as technique tags)
[0161] audio_context (audio tempo information, such as BPM value)
[0162] error_cases (error case identifiers, such as "finger flex is too low")
[0163] adjustment_rules (parameter adjustment rules, such as "speed limit increased by 10%").
[0164] (2) VR scene command: used for dynamic updates of VR book scene.
[0165] Example field: vr_commands:
[0166] scene_switch (scene switching command, such as "switch from lobby to backstage")
[0167] event_triggers (real-time event triggers, such as "audience interaction request")
[0168] environment_updates (environment element adjustments, such as lighting and sound parameters).
[0169] (3) Context identifier: used to associate user behavior data.
[0170] Example fields: context_id (inference ID), session_id (session identifier), user_type (user type).
[0171] In some embodiments of the present invention, the hybrid cognitive engine used to output knowledge graph features in step A4 can also be updated, specifically including steps A5 and A6.
[0172] Step A5: Analyze the knowledge graph features output in Step A4 to obtain digital human motion parameters and VR scene commands, enabling knowledge visualization and interaction, and generating real-time user behavior data streams. Specifically, this includes:
[0173] Step A5.1: Construct a digital human with the ability to accurately simulate the movements of Pingtan performance, realize the closed-loop drive of "Pingtan techniques - digital human movement parameters - scene interaction", and support immersive performance in VR storytelling venues.
[0174] Step A 5.1.1: Construct a 3D model of the digital human and bind it to the skeleton.
[0175] (1) Use structured light scanners or multi-view camera arrays to collect high-resolution 3D data of the faces and hands of Pingtan performers (focusing on capturing sensitive areas such as finger joints and wrists); combine texture mapping (such as 4K RGB images) to generate realistic materials for skin and clothing (such as long gowns and cheongsams); define a digital human skeleton system (20-30 key bone nodes) based on human anatomy, focusing on strengthening the hand bones (10-15 subdivided bones, such as the proximal / middle / distal phalanges of the thumb and index finger); set freedom constraints for each bone node (such as finger bones only allowing bending / extension, and wrists allowing rotation ±90°).
[0176] (2) Perform topology optimization on the hand model and bind the muscle deformation weights of actions such as "finger twirling" and "string sweeping"; add fabric simulation to the long gown and water sleeves and set parameters (such as mass 0.5kg / m² and tensile stiffness 80%) to ensure that they hang naturally during the action.
[0177] Step A5.1.2: Collection and annotation of Pingtan technique movement data.
[0178] (1) Deploy optical motion capture systems (such as Vicon) in the storytelling performance area; inertial sensors (such as Xsens): worn on fingers and wrists; force feedback gloves (such as 5DT Data Glove): collect finger bending and contact force;
[0179] (2) Divide the action segments according to the performance process of Pingtan, mark the core techniques corresponding to each segment, and mark the action parameters of the key frames of each technique: finger bending degree, wrist rotation angle and action speed.
[0180] Step A5.1.3: Real-time motion parameter mapping.
[0181] (1) Parse the output of step A4: Use a lightweight parser to extract the technique name, lyrics timestamp, and contextual entities (such as technique tags obtained from the NER results).
[0182] (2) Parameter query and generation: Retrieve the baseline parameters from the parameter library according to the technique name, and make dynamic adjustments in combination with the context provided in step A4 (such as audio rhythm, error cases). For example: if step 4 outputs "finger rolling technique, fast audio rhythm", the upper limit of the movement speed will be automatically adjusted; if step A4 provides an error case (such as "finger rolling bend is too low"), the current parameters will be adjusted immediately to avoid repeating the error.
[0183] (3) Smooth transition processing: For continuous actions (such as the opening and closing of a folding fan to the turning of the fingers), use interpolation algorithms (such as cubic splines) to ensure smooth parameter changes and reduce jumps.
[0184] Step A5.1.4: Real-time calibration and feedback.
[0185] (1) Integrate the verification data of step A4: If step A4 marks the parameter deviation, step A5.1 starts the calibration process, recalculates the parameter mean or invites experts to review (through human-machine interface).
[0186] (2) Use machine learning models to assist: Based on the output of step A4, update the action parameter mapping relationship to ensure semantic consistency.
[0187] Step A5.2: Construct the VR storytelling environment. Convert the output of step A4 into VR scene instructions and process user interactions.
[0188] Step A5.2.1: Dynamically update the scene.
[0189] (1) Parse the event context of step A4. For example, if step A4 detects a change in user interest (inference from the knowledge graph), then generate a scene switching instruction (such as switching from the lobby of the bookstore to the backstage).
[0190] (2) Real-time rendering adjustment: Based on the real-time events provided in step A4 (such as audience interaction requests), update the virtual actor behavior or environmental elements (such as lighting and sound) in the VR Book Theater.
[0191] Step A5.2.2: User interaction processing.
[0192] (1) Multimodal input capture: User input (such as gestures and voice commands) is acquired through full-body motion capture and speech recognition systems.
[0193] (2) Response generation: Based on the reasoning results of step A4, drive the virtual actor to respond. For example, if step A4 recognizes that the user queries "finger rolling technique", the digital human demonstrates the finger rolling action and plays the explanatory audio in sync.
[0194] Step A5.3: Generate user behavior data stream.
[0195] Step A5.3.1 Record the data.
[0196] (1) Record user interaction details: including user actions (such as gesture type, voice content), interaction timestamp, system response (digital human action parameters, VR scene status).
[0197] (2) Associate step A4 context: Associate each data point with the output identifier of step A4 (such as inference ID, error case ID).
[0198] (3) Collect performance indicators: record system latency, action accuracy (such as deviation from standard parameters), and user satisfaction (through implicit feedback such as interaction time).
[0199] Step A5.3.2 Data formatting and output.
[0200] (1) Data standardization: Convert the raw data into JSON format, including fields such as user_action, timestamp, system_response, step4_context, and performance_metrics.
[0201] (2) Real-time streaming output: Send the data stream to step A6 in real time via message queue (such as Kafka) or API to ensure low latency optimization.
[0202] Step A6: Dynamically update the cognitive engine parameters based on the multimodal interaction data stream to form a closed loop that enhances cognition and continuously optimizes it.
[0203] Step A6.1: Data Integration and Preprocessing. Process the user interaction data from Step A5 to provide standardized input for optimization, including data cleaning: removing outliers (such as transient high-latency data) and handling missing values (using interpolation or default values). Feature Extraction: Extract key features from the user interaction data, such as query_count, avg_rating, skip_rate, and laughter_response_rate.
[0204] Step A6.2: Dynamic Adjustment of Knowledge Weights. Based on the integrated data, the weights of nodes in the knowledge graph are adjusted to optimize the reasoning priority in step A4.
[0205] (1) Dynamically adjust knowledge weights based on user interaction data. Calculate the new weights using the following formula:
[0206]
[0207] in, It is the current weight of the node. It represents the number of queries. That is the average rating.
[0208] (2) Weight update rules.
[0209] Execute regularly (e.g., every 100 user interactions or every hour) to avoid frequent fluctuations; limit the weight range to [0, 1] to prevent overflow; set minimum weight thresholds for key nodes (such as high-frequency query nodes) to ensure that basic knowledge is not overly diluted.
[0210] (3) The updated knowledge weight table is synchronized to the cognitive engine in step A4.
[0211] Please combine Figure 4 As shown, the second agent performs cross-domain transfer enhancement on the knowledge graph features to generate a long list of Suzhou Pingtan (storytelling and ballad singing) performances with cross-domain enhancement features. The specific process includes the following:
[0212] Step B1: Initialize the agent and build the feature library
[0213] Step B1.1: Analyze the output of the first agent.
[0214] (1) Receive the output of the first agent .
[0215] (2) Analysis .
[0216] Extracting key information to guide feature extraction: topic keywords (such as from...) Extract =“Suzhou Scenery”, structural requirements (such as “needs fast-paced paragraphs”), style indicators (such as “elegant”, “exhilarating”).
[0217] Step B1.2: Construct a cross-domain feature library. The cross-domain features in the library include cross-art text features, cross-art audio features, and cross-art performance features. In this embodiment, cross-art text features include Kunqu Opera text features, cross-art audio features include Beijing Drum Song audio features, and cross-art performance features include Kunqu Opera performance features.
[0218] based on The thematic and structural information in the text is used to extract features of Kunqu Opera and Beijing Drum Song from external public domain databases.
[0219] Step B1.2.1: Extracting Kunqu Opera Text Features
[0220] (1) According to Using thematic keywords (such as "garden"), relevant Kunqu opera pieces (such as "The Peony Pavilion: Strolling in the Garden") are retrieved from public domain classic Kunqu opera texts, and a Kunqu opera text sequence is constructed. A Bi-GRU-based parser for the ci (lyric) structure extracts the metrical constraint matrix; the specific execution process is as follows:
[0221] Input data: Kunqu Opera text sequence , Indicates the first Each ci (a single character or a fixed phrase) unit, and an annotated dataset, including: tone vectors. , (Level tone) (Oblique tone), rhyme mark (1 indicates the position of the rhyme).
[0222] Calculation process:
[0223] (a) Feature embedding layer, target vector , ,in Indicates the word Perform the embedding operation and output a dense vector; This represents a vector concatenation operation; It is the embedded dimension.
[0224] (b) Bidirectional GRU encoding: Forward GRU hidden state Backward GRU hidden state By concatenating the forward and backward hidden states, a 256-dimensional output vector is obtained. The GRU cell update formula is as follows:
[0225]
[0226] (c) Attention-weighted aggregation, focusing on key metrical positions (such as rhymes, tonal transition points). The sequence... Attention scores for hidden states at each location:
[0227] ,
[0228] Indicates hidden state After weight matrix Perform a linear transformation. This indicates that nonlinearity is introduced through the hyperbolic tangent function. Indicates the use of parameter vectors The transformed vectors are weighted and summed to obtain a scalar score. Hidden state sequence. weighted sum .
[0229] (d) yes The structure vector of each track is adopted. Technology will Clustering A typical ci (lyric poem) structure prototype (matrix row vectors) yields the metrical constraint matrix. , It represents the number of typical ci (lyric) structure prototypes (cluster centers). It is the structural vector dimension (256 dimensions).
[0230] (2) According to Using thematic keywords, relevant Kunqu Opera repertoires are retrieved from public domain classic Kunqu Opera repertoire texts to construct a Kunqu Opera text corpus. Based on a topic model, a literary imagery transfer matrix is mined; the specific execution process is as follows:
[0231] Input data: Kunqu Opera text corpus Target Pingtan theme keywords
[0232] Calculation process:
[0233] (a) Topic-Image Joint Modeling. The document topic distribution is learned simultaneously using the Joint Hidden Dirichlet Allocation (JointLDA) algorithm. Topic-term distribution Theme-Attribute Distribution Three key parameters, in document collection As input, through imagery enhancement factors Control the weight balance of different components.
[0234] (b) Finding the most suitable thematic representation by maximizing the ratio of "thematic relevance / style difference". For imagery in Kunqu Opera... And the target theme in Pingtan Matrix elements The value is the maximized subject. ,Right now
[0235] ,
[0236] The molecular part is an image of Kunqu Opera. On the topic Conditional probability and target topic terms in the topic The product of conditional probabilities; the denominator is the relationship between Kunqu Opera and Pingtan in the theme. The KL divergence between the thematic distributions is used to measure the stylistic differences between the two art forms, thus yielding the imagery transfer matrix. , It is the number of images in Kunqu Opera. It refers to the number of target topics.
[0237] Step B1.2.2: Extract audio features of Beijing-style drum singing.
[0238] according to To meet structural requirements, relevant segments are retrieved from public domain Beijing-style drum music audio. From these segments, Fourier descriptors of the MFCC time sequence matrix are extracted to construct rhythmic templates, and energy abrupt changes in fast-paced segments are annotated to form rhythmic anchor point sequences. The specific execution process is as follows:
[0239] (1) Mel frequency cepstral coefficient matrix , It is the original audio signal. It is a Mel spectrogram calculation function; 5 frequency domain energy characteristics. , Indicates the first The values of each MFCC coefficient across all frames. This represents the Fourier transform, which converts time-domain features to the frequency domain. The L2 norm is used to calculate the energy in the frequency domain representation, resulting in the rhythm template vector. The spectral energy of the first 5 MFCC dimensions is taken as the core rhythmic feature.
[0240] (2) At the time point The short-time energy (summed and averaged over the squares of the signals within the window).
[0241] ,
[0242] This refers to the window size. The second derivative of the energy exceeds a threshold and falls within a specified time interval. The time point within is defined as the anchor point.
[0243] ,
[0244] Second derivative of short-time energy The acceleration representing the change in energy. It is a threshold.
[0245] Step B1.2.3: Extracting Kunqu Opera Performance Features
[0246] (1) From Extracting key themes, performance style requirements, and emotional expression criteria, relevant repertoires are retrieved from public domain Kunqu Opera classic repertoire texts to construct a Kunqu Opera performance resource database.
[0247] (2) Extract body movement features from Kunqu Opera performance videos: Use a posture estimation model to extract the skeleton key point sequence, calculate the joint movement trajectory and speed change, encode continuous action patterns through LSTM network, and identify typical body posture combinations (such as "cloud hands", "lying fish" and "showing off").
[0248] (3) Extract gesture symbol features from Kunqu Opera performance videos: detect hand areas and extract gesture contours, classify traditional Kunqu Opera gestures (finger techniques, palm techniques, fist techniques, etc.), construct a gesture-emotion mapping dictionary, and encode the temporal evolution of gesture sequences.
[0249] (4) Extract vocal performance features from Kunqu Opera performance videos: extract the fundamental frequency trajectory to reflect pitch changes, detect vibrato features (frequency, amplitude, rate), identify ornament patterns (glissando, vibrato, appoggiatura), and analyze breath control patterns: detect the position and duration distribution of the air mouth through energy envelope.
[0250] (5) Extract facial expression features from Kunqu Opera performance videos: identify traditional opera expression paradigms (joy, anger, sorrow, surprise, etc.), quantify the expression intensity change curve over time, and establish an expression-lyric association model.
[0251] (6) Integrate multimodal features (action + vocal cavity + facial expression) to construct an emotion intensity curve: Identify key nodes of emotional transition and establish an emotional performance map.
[0252] (7) By using a multi-layer perceptron feature fusion network, the performance features of Kunqu Opera from different modalities are fused into a unified performance feature vector: .
[0253] Step B2: Align the knowledge graph with cross-domain feature transformation.
[0254] Step B2.1: Construct a shared ontology architecture.
[0255] Based on the cross-artistic features extracted in step B1, a unified knowledge representation framework is constructed. The metrical constraint matrix, imagery transfer matrix, rhythm template, and performance feature vector output from step B1 are standardized to establish a shared semantic space for the three art forms.
[0256] Step B2.2: Based on the graph attention network, align cross-domain features and generate a cross-domain feature transformation matrix.
[0257] The specific implementation process is as follows:
[0258] (1) Constructing a knowledge graph of three art forms , It is a set of nodes, including Pingtan nodes. Kunqu Opera node Big Drum Knot wait; It is a cross-domain relationship edge (such as "tone rule - rhyme pattern"); , , , As the initial feature vector of the corresponding art form node ;
[0259] (2) Heterogeneous graph attention aggregation. Information from neighboring nodes is aggregated through an attention mechanism to update the current node. Hidden state vector in the first layer , It is an activation function (such as ReLU, Sigmoid, etc.). It is a node The set of neighboring nodes, It is a learnable weight matrix. It is a node The initial hidden state vector at layer 0, It is a node To the node The attention weight is calculated using the following formula:
[0260]
[0261] in, It is a learnable projection matrix. It is an attention vector.
[0262] (3) Generate a cross-domain similarity matrix. Node and nodes Similarity score between .
[0263] (4) Generate cross-domain transformation matrix ,in It is a soft alignment function. It is the covariance matrix of art forms.
[0264] Step B3: Transfer enhancement to generate Pingtan bibliographic text, audio, and performance instructions.
[0265] Step B3.1: Based on the transformation matrix generated in step B2 Using the Kunqu Opera text features extracted in step B1, structured literary transfer is achieved to generate Pingtan texts with enhanced literary quality.
[0266] (1) From the metrical constraint matrix of step B1 Extracting the structural units of the ci (a type of Chinese poetry) Using the transformation matrix from step B2 Cross-domain projection is performed on the structural units, and combined with the user's creation instructions, the knowledge graph retrieval results output by the intelligent agent (1) are obtained. and the cross-domain feature transformation matrix generated in step B2.2 Initialize the hidden state of the current commentary generator. ;
[0267] (2) Dynamic unit selection.
[0268] The matching degree between the current generation context and each structural unit is analyzed using a multilayer perceptron. ,in Calculate the matching degree between the current context and each structural unit; calculate the attention weight of each unit based on the matching degree. ,in It is a set The first part of the "introduction, development, transition and conclusion" structure in Chinese Each structural unit is used to achieve dynamic selection; the selected structural unit... The transformation matrix is projected onto the Pingtan feature space.
[0269] (3) Injection condition gating.
[0270] The gating vector is calculated based on the current state and the selected unit to control the migration intensity; the original storytelling features are appropriately preserved by retaining the mask; and element-wise gating is used to achieve smooth feature injection and avoid style conflicts. (Gating vector) ,in It is the hyperbolic tangent activation function, which compresses the value to Interval, weight matrix Used for linear transformations, Indicates concatenation of the currently hidden state. With the selected knowledge vector The updated hidden state ,in Represents element-wise multiplication (Hadamard product), Kunqu Opera transfer weight It is a coefficient that controls the intensity of the migration of Kunqu Opera artistic features to Pingtan works. This indicates that the original hidden state should be retained. This indicates the part that incorporates new knowledge.
[0271] (4) Enhanced literary imagery generates bibliographical texts.
[0272] Based on the target theme From the imagery transfer matrix Extract the corresponding column vector ; through a fully connected layer Mapped to image feature vectors: ,in , Use fusion strength parameters (Default value is 0.3), the image feature vector is fused into the hidden state after gated injection: ;Will Input text decoder to obtain the word probability distribution at the current time step. ,in and These are the weight matrix and bias terms of the output layer. Map the hidden states to a probability distribution on the vocabulary; according to Generate tokens (characters or words) ;Will Proceed to the next time step as a hidden state; after The text of the Pingtan bibliography was obtained after a certain time step. .
[0273] Step B3.2: Generate a prosodic code that conforms to the phonetics of Pingtan and ensures that it is aligned with the semantics and sentiment of the text.
[0274] Step B3.2.1: Decoding of multi-scale timestamp control sequences.
[0275] (1) From the transformation matrix Extracting the rhythm submatrix Calculate the first time interval ,in It is a time scaling factor (default 0.8 to adapt to the speed of Pingtan speech), used to adapt to the speed of Pingtan speech.
[0276] (2) Fine-grained rhythmic adaptation. A fine time step increment ,in Multilayer perceptrons for predicting fine time steps Used to Mapped to a high-dimensional feature space, No. The location encoding of each position (transmitting sequence position information) It is a vector concatenation operation (combining the three input parts into the input vector of the MLP).
[0277] (3) Multi-scale fusion. Total increment at each time step ,in It is the first A “coarse-grained time step increment” (basic component). It is the scaling factor (default is 0.2, used to adjust the overall weight of fine components). For activation function, Represents the adaptive weight matrix ( yes The dimension used for (perform preliminary transformation) It is a text semantic vector. .
[0278] (4) Generate a timestamp sequence by accumulating. ,in Initial time.
[0279] Step B3.2.2: Conditional variational autoencoder (CVAE) generates prosodic codes.
[0280] (1) Constructing condition vectors , among which It calculates the statistical characteristics of a timestamp sequence (such as mean, variance, minimum, maximum, etc.); and encodes latent variables. ,in .
[0281] (2) Latent space sampling: , Calculate the KL divergence loss , used to constrain the distribution of the latent space.
[0282] (3) Constructing cross-domain rhythm feature projection Among them, the big drum migration weight It is the intensity coefficient controlling the transfer of Beijing drum art features to Pingtan works; calculating the decoding condition vector ,in, It is a style embedding vector for Pingtan performance; generating prosodic encoding. .
[0283] (4) Reconstruct training loss and KL divergence loss, .
[0284] Step B3.2.3: Rhythm - Dynamic text alignment.
[0285] (1) Constructing the text sentiment intensity curve
[0286] .
[0287] (2) Calculate the prosody-text alignment loss
[0288] ,
[0289] Discrete approximation
[0290] ,
[0291] in It is the coupling strength coefficient.
[0292] Step B3.3: Based on the performance characteristics of Kunqu Opera, generate performance guidelines adapted to Pingtan.
[0293] Step B3.3.1: Cross-domain projection of performance characteristics and style adaptation.
[0294] (1) Characteristic parameters after projection onto the Pingtan performance space , It is a performance feature projection matrix (realizing the spatial mapping of Kunqu Opera to Pingtan).
[0295] (2) Optimization of Pingtan style constraints: style mixing coefficient Parameters of Pingtan performance after style adaptation ,in , These are the weight matrix and bias term for style adjustment. These are the default performance parameters (style baseline values) for Pingtan (a type of storytelling and ballad singing in Suzhou dialect). This indicates element-wise multiplication (style blending on a dimension-by-dimensional basis).
[0296] (3) Adjusting performance intensity: Performance intensity coefficient (predicted by MLP, controlling the overall performance intensity) Final performance parameters after intensity adjustment .
[0297] Step B3.3.2: Decoding multimodal performance elements.
[0298] (1) Generate body movement sequence: basic body movement parameters (generated by posture decoder) Rhythmic body movements (LSTM incorporates rhythm encoding to ensure a sense of rhythm in the movements) Continuous body movement trajectory (smooth continuous movements obtained through temporal interpolation) ,in It is a body posture decoder (which maps performance parameters to basic movements). It is a long short-term memory network for modeling movement rhythm. It is a time-series interpolation function (generating a continuous sequence of actions based on timestamps);
[0299] (2) Transformation of hand gesture symbols: Original hand gesture symbols extracted from Kunqu Opera Gesture symbols mapped onto the Pingtan system Pingtan gesture sequences (ultimately used to guide the performance) ,in, This represents the gesture extraction function (separating gesture features from Kunqu Opera performance parameters). This represents a text alignment function (that synchronizes gestures with text content and timestamps).
[0300] (3) Generate vocal performance parameters: breath control parameters Articulation and pronunciation parameters Volume dynamic parameters ,in, It is a breath decoder (which integrates performance parameters and rhythm averages to generate breath characteristics). It is a dedicated MLP for articulation parameter prediction. It is a normalization function (mapping the absolute value of the rhythm code to a reasonable volume range).
[0301] (4) Facial expression and emotion mapping: basic emotional expressions based on performance parameters Emotional features extracted from text The final emotional expression parameters after fusion Facial intensity parameters ,in, It is an emotion expression decoder. It is a text sentiment analysis function. It is the emotional fusion weight (controlling the proportion of basic facial expressions and textual emotions). It is a smoothing function (to avoid sudden changes in facial expression intensity). It is a difference function (for calculating the rate of change of rhythm coding).
[0302] Step B3.3.3: Timing synchronization and multimodal coordination.
[0303] (1) Performance-text alignment constraints. Text embedding sequence
[0304] ,
[0305] Text-Action Synchronization Matrix
[0306] ,
[0307] Alignment loss (measures the difference between the synchronization matrix and the identity matrix I)
[0308] ,in It is a text embedding function (encodes a text token into a vector). It is the square of the Frobenius norm (calculated by the sum of squares of the matrix elements). It is an identity matrix.
[0309] (2) Multimodal coordination optimization: Multimodal coordination loss (ensuring coordination and consistency of body posture, rhythm, expression, gestures, etc.).
[0310] ,
[0311] Among them, the coordination weighting coefficient (controls the contribution of each loss term) , , , It is a cross-correlation function. It is a consistency function (measures the degree of matching between a sequence of gestures and facial expressions).
[0312] Step B3.3.4: Integration and generation of performance instructions.
[0313] (1) Structured performance guidance program (including four modules: body posture, gestures, vocal style, and facial expressions):
[0314] ,in
[0315] ,
[0316] ,
[0317] ,
[0318] ,
[0319] These are formatted body posture guidelines, gesture guidelines, vocal guidelines, and facial expression guidelines; , , , These are the formatting functions for the corresponding modules (which convert the parameters into an executable instruction language).
[0320] (2) Quality assessment.
[0321] Overall quality score of the performance instruction program:
[0322] ,
[0323] The quality assessment weight (controlling the impact of each indicator) is included. , It is a style consistency score (measuring the degree of fit with the style of Pingtan). It is an artistic value score (assessed by professional standards, such as aesthetics and expressiveness).
[0324] Step B4: Dual evaluation and self-optimization.
[0325] Step B4.1: Real-time quality assessment.
[0326] Step B4.1.1: Metric conformity test based on expert rule base (traditional verification).
[0327] Based on a pre-defined expert rule base, the system checks the degree to which a work conforms to metrical rules, literary norms, and performance conventions. (Expert rule base) It includes metrical rules such as tonal distribution, rhyme placement, and sentence structure; literary rules such as allusion rules and imagery pairing; and performance rules such as body movements and gestures.
[0328] During the verification process, the system checks each generated work to ensure it does not violate any rules and calculates a score for adherence to tradition. A higher score indicates that the work conforms better to traditional norms and demonstrates stronger artistic inheritance. (Step B3 generates a bibliography of Pingtan storytelling.) Traditional verification scores
[0329] ,
[0330] in Representing a set of rules Size, It is an indicator function that returns 1 if the condition is met, and 0 otherwise. Used for judgment Does it violate the rules? .
[0331] Step B4.1.2: Calculate the semantic space distance between the generated text and the training set (innovation assessment).
[0332] Text of Pingtan (a type of storytelling and ballad singing) to be evaluated semantic vectors Reference Corpus Average semantic vector
[0333] ,
[0334] It refers to the number of documents in the corpus. It is the first in the corpus One document;
[0335] Innovation score ,
[0336] It is the Euclidean distance (L2 norm) between the generated text vector and the average vector of the corpus, and the threshold parameter. Control sensitivity to innovation.
[0337] Step B4.2: Comprehensive quality assessment and optimization triggering.
[0338] .if Proceed to step B4.3. Otherwise, output the bibliography of Pingtan (storytelling and ballad singing). .
[0339] Step B4.3: Self-optimization parameters.
[0340] (1) Update Kunqu Opera migration weights and drum migration weight :
[0341] , It is the learning rate.
[0342] (2) Perform step B3.
[0343] Please combine Figure 5 As shown, the third intelligent agent integrates knowledge graph features and cross-domain enhancement features, optimizes the long-form Suzhou Pingtan repertoire through reinforcement learning strategies, and outputs the final results.
[0344] The architecture of the third-party intelligent agent includes:
[0345] (1) Input layer:
[0346] Output of the first agent: Knowledge graph features ,in Indicates the characteristics of a school of thought. Indicating a master-disciple relationship. Represents cross-modal rules;
[0347] Output of the second agent: Cross-domain augmented narrative ;
[0348] User requirements: .
[0349] (2) Reinforcement learning layer:
[0350] State space:
[0351] Action space: For detailed definitions, please refer to Table 3.
[0352] Table 3 Definition of Action Space Specification Parameters
[0353]
[0354] Reward function: , (Amplify the innovation reward index).
[0355] The execution process of the third-party intelligent agent is as follows:
[0356] Step C1: Construct the state space.
[0357] Step C1.1: Combine the outputs of the first and second agents and the user's requirements into a state vector.
[0358] Integrating the above multimodal features: , , , ;
[0359] Step C1.2: Encode the state into a fixed-dimensional state vector using a neural network. Definition of a state-encoding network: , , State vector dimension: .
[0360] Step C2: Near-end strategy optimization loop.
[0361] Step C2.1 Initialize PPO algorithm parameters.
[0362] Maximum number of iterations Maximum step size per round Discount factor GAE parameters PPO trimming parameters Learning rate .
[0363] Policy network architecture definition: , .
[0364] Step C2.2: Training loop and calculating dynamic rewards. For each episode... arrive Perform the following process:
[0365] Step C2.2.1: Initialize the state (From the output of step 1.2), initialize the episode's cumulative reward. .
[0366] Step C2.2.2: For each step size arrive Perform the following process:
[0367] (1) Action sampling ,in Environmental state transition .
[0368] (2) Calculate dynamic rewards.
[0369] (a) Calculate the base reward based on tradition, innovation and user satisfaction.
[0370] Calculating traditional verification scores and traditional rewards Next, the innovation score is calculated. and innovation awards Then calculate the user satisfaction reward. Finally, the basic reward is received. ,in .
[0371] (b) When the traditionality or the condition is not met, a penalty shall be imposed.
[0372] Penalties for insufficient traditional practices:
[0373] .
[0374] Penalties for insufficient innovation:
[0375] .
[0376] (c) The reward is doubled when both tradition and innovation reach high standards. If Greater than 0.9 and , .
[0377] (d) Final reward ;
[0378] (3) Experience storage and status update.
[0379] (a) Storage experience: ;
[0380] (b) Update status Cumulative rewards ;
[0381] (c) If a termination condition is encountered, terminate the step-size loop early.
[0382] Step C2.3: Update the policy network using the PPO algorithm.
[0383] Step C2.4: Iteration Termination Condition
[0384] Training terminates when one of the following conditions is met:
[0385] (1) Achieve ;
[0386] (2) Policy convergence: nearest The average reward change per episode is less than the threshold. (For example, );
[0387] (3) Manual intervention: The user manually stops the training.
[0388] Step C3: Generate the final bibliography using the optimized strategy.
[0389] (1) Optimize the text: Use the trained PPO policy network ( ),from Begin by selecting actions sequentially until the [number]th action. This process generates a state sequence. After Decode into the final text.
[0390] ;
[0391] (2) Optimize audio: Use an audio synthesizer ( According to the optimized prosodic coding and rhythmic action sequences generated by the PPO policy network Generate the final audio. .
[0392] ;
[0393] (3) Optimize performance guidance: use an adapter Optimized performance features and the performance action sequence generated by the PPO policy network Transformed into specific performance instruction .
[0394]
[0395] (3) The final output is a triple.
[0396] .
[0397] The following example illustrates in detail how this invention is specifically implemented.
[0398] Create a new Suzhou Pingtan (storytelling and ballad singing) long-form book titled "The Love of the Children of the Grand Canal" with the theme of "protection of the Grand Canal cultural heritage".
[0399] User input and interaction (via the multimodal interface of the first agent):
[0400] Users input their creative needs into the system via voice and text: "I hope to create a new long-form Pingtan storytelling piece called 'The Love of the Children of the Grand Canal,' with the theme revolving around the local customs, historical changes, and contemporary protection of the Suzhou section of the Grand Canal. The requirements are to retain the charm of traditional Pingtan while incorporating innovative elements that conform to modern aesthetics."
[0401] First intelligent agent:
[0402] Define the ontology and construct the knowledge graph of Suzhou Pingtan.
[0403] (1) The system calls the predefined Suzhou Pingtan ontology, which includes core concepts such as “book list”, “chapter list”, “role”, “event”, “plot pattern”, “musical tune”, “Sanxian technique”, “Pipa technique”, “singing style”, and “theme” and their interrelationships.
[0404] (2) The knowledge graph engine automatically extracts information from authorized Pingtan text databases (such as "The Pearl Tower" and "The Jade Dragonfly"), historical and cultural databases (such as Suzhou local chronicles and Grand Canal historical and cultural materials), and publicly available data on the Internet. For example, it extracts "Fengqiao", "Hanshan Temple", "Xumen", "Grand Canal", and "Silk" as "location" and "historical event" nodes.
[0405] (3) Construct a structured Suzhou Pingtan knowledge graph with “Grand Canal cultural heritage” as the core and related to history, figures, places and traditional books.
[0406] First-level intelligent agent cognitive engine reasoning and continuous optimization.
[0407] (1) The cognitive engine reasoned based on the Suzhou Pingtan knowledge graph: (Topic: Cultural Heritage Protection) + (Location: Fengqiao) → associated with the classic plot patterns "revisiting the old place" and "things have changed".
[0408] (2) The engine verifies the rationality of the generated content.
[0409] (3) Based on user feedback on the initial generated segments (such as users skipping a certain type of singing segment multiple times), the optimization mechanism will lower the knowledge weight of that type of singing segment style and reduce its frequency of occurrence in subsequent generation.
[0410] Multimodal interaction and data collection.
[0411] (1) The first chapter of the story outline and a song lyrics of "The Love of the Canal Children" were initially generated and performed by digital actors in the VR Hanshan Temple Night Mooring scene.
[0412] (2) Users interact with the digital human through VR devices and express particular satisfaction with the narrative paragraph about the "hardships of the canal transport" (staying for a long time and giving a "thumbs up"). This user behavior data stream is recorded in real time and used to optimize the first intelligent agent's cognitive engine.
[0413] Second intelligent agent:
[0414] Construction of cross-domain feature library.
[0415] (1) Analyze the knowledge graph features output by the first agent (such as the concepts of "farewell", "water scene" and "lyricism") to determine the need for "elegant literary style" and "low rhythm".
[0416] (2) Kunqu text features: It uses the [Zao Luo Pao] tune (four-six parallel prose, gongche notation) and its literary imagery of "colorful flowers" in "The Peony Pavilion: Strolling in the Garden".
[0417] (3) Audio characteristics of Beijing-style drum music: It uses the slow tempo template and accented anchor sequence of "Jiange Wenling" to express the "sorrowful and lingering" emotions.
[0418] (4) Kunqu Opera performance characteristics: It uses the body movements and "water sleeve" gestures of Zhao Kuangyin's "riding a horse" in "Sending Jingniang a Thousand Miles Away", as well as the corresponding tragic vocal style.
[0419] Knowledge graph alignment.
[0420] (1) Establish a shared ontology of "Pingtan-Kunqu-Dagu", with "emotional expression" as the core.
[0421] (2) Using graph attention network (GAT) to calculate the similarity of “Liqingdiao” in Pingtan and “Zaoluopao” in Kunqu and “Jiange Wenling” in Dagu in terms of emotional dimensions such as “sadness” and “gentleness”, a high degree of correlation was found.
[0422] (3) Generate cross-domain feature transformation matrix This matrix defines how to map the textual structure of Kunqu Opera and the rhythmic patterns of the drum into the framework of Pingtan.
[0423] Migration-enhanced generation.
[0424] (1) Literary Infusion: The structure of the Kunqu opera [Zao Luo Pao] lyric pattern is combined with the theme of "scenery on both sides of the canal" to generate lyrics that conform to the rules of Pingtan but are more literary: "Look at the canal with its thousands of miles of shimmering waves, how can it compare to the busy canal transport in the old days..."
[0425] (2) Rhythmic innovation: The slow rhythm and anchor point of the Beijing-style drum ballad "Jiange Wenling" are transformed into the "slow rolling embroidered ball" rhythm code of the Pingtan Sanxian, creating a tragic narrative rhythm that is both traditional and novel.
[0426] (3) Performance guidance: Adapt the Kunqu "Tangma" movement to a virtual performance scheme of a scholar riding a horse along the canal in Pingtan, and simplify the "water sleeve" movement to a gesture "Xuzhiyuanwang" which is more characteristic of Pingtan.
[0427] Evaluate and optimize the loop.
[0428] (1) Verification of traditionality: Compare the generated lyrics with the database of classic books to confirm that the rhyme and tone conform to the Pingtan norms, with a traditionality score of 0.75.
[0429] (2) Innovation assessment: assess its uniqueness in the use of allusions (comparison between the past and present of the canal) and the integration of rhythm, with an innovation score of 0.70.
[0430] (3) Overall quality index Since the preset threshold of 0.80 was not reached, the system automatically and dynamically adjusted the migration weights, slightly increasing the migration weights for the literary aspects of Kunqu Opera. Reduce the audio transfer weight of the bass drum rhythm. Then it is regenerated. After two rounds of iterations... When the value reaches 0.82, the final cross-domain augmented features are output (including optimized text, rhythm, and performance scheme).
[0431] Third intelligent agent:
[0432] Global optimization based on reinforcement learning.
[0433] (1) State Space: The knowledge graphs such as "canal", "farewell" and "contemporary conservation" output by the first agent are embedded and fused with cross-domain enhanced features such as "slow rolling embroidered ball rhythm encoding" and "Kunqu opera lyrics vector" output by the second agent, as well as user preference vector (showing that the user likes "arduous narrative"), to form the current state. .
[0434] (2) Action Space: The PPO policy network is based on Output a set of mixed actions.
[0435] Discrete action: Insert a aria in the style of a dockworker's chant after the lyrical passage.
[0436] Continuous parameters: The intensity (dynamics) of this vocal section is set to 0.8, and the BPM (tempo) is increased by 15%.
[0437] (3) Calculation of reward function:
[0438] Traditional bonus: Since the inserted "work chant" element is derived from real canal history and culture, the traditional bonus is +0.6 after calculation by the S-curve function.
[0439] Innovation Bonus: Combining "workers' chants" with elegant Pingtan singing style is highly innovative, and the bonus is +1.2 after the index is amplified.
[0440] User satisfaction reward: This action aligns with the user's preference for "hard-won narratives", so a reward of +0.5 is given.
[0441] Dynamic Rewards: Due to the good performance in both traditional and innovative aspects, the total reward is doubled, i.e., +4.6 = (0.6 + 1.2 + 0.5) * 2.
[0442] (4) Training and Output:
[0443] The third agent is trained through a large number of similar interactions and eventually learns an optimal policy.
[0444] Applying this strategy, the first draft of "The Love of the Canal Children" output by the agent (2) is globally optimized: the order of each chapter is adjusted to enhance the drama, actions that strengthen the emotions (such as adjusting the BPM) are inserted at key plot points, and finally a structured Pingtan bibliography document is generated.
[0445] This document contains:
[0446] Text content: Complete and optimized chapter descriptions and lyrics.
[0447] Audio configuration: Detailed track usage suggestions, rhythm patterns, and BPM variation markers.
[0448] Performance guidance: A virtual digital human performance solution that incorporates Kunqu Opera movements to complement key singing sections.
[0449] Final output: The system presents the long storytelling performance "The Love of the Canal Children," which blends traditional heritage with modern aesthetics, to users in its entirety through digital human performance and VR scenes, completing the intelligent creation process from scratch.
[0450] Summary of the beneficial effects of this embodiment: This example clearly demonstrates how three intelligent agents can collaborate organically to transform an abstract user request into a concrete, high-quality Suzhou Pingtan (storytelling and ballad singing in Suzhou dialect) long-form storytelling performance that combines tradition and innovation. The entire process demonstrates how clearly defined technical means (knowledge graphs, GAT, PPO, etc.) have solved the technical problems of lowering the professional threshold and improving creative efficiency, achieving significant technical results.
[0451] It should be noted that the specific methods of the above embodiments can form a computer program product. Therefore, the computer program product implemented in this application can be stored on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.).
[0452] Simulation experiments were conducted on the implementation method of the present invention. The hardware environment is detailed in Table 4, and the software environment is detailed in Table 5.
[0453] Table 4 Experimental Hardware Environment
[0454]
[0455] Table 5 Experimental Software Environment
[0456]
[0457] The dataset is derived from publicly available materials from the Suzhou Pingtan Museum and manually scanned publicly published Pingtan repertoire, including works such as "The Golden Fan," "Three Smiles," "The Golden Phoenix," and "The Legend of the White Snake." It includes 20 hours of vocal recordings labeled with genre, emotional, and vocal technique tags, as well as 2 hours of performance videos labeled with the timing of instrument movements and facial expressions. The cross-domain art dataset includes: 50 ci (lyric) structures (Kunqu Opera), 200 rhythmic templates (Beijing Drum Song), and 50 emotional expression segments (Shanghai Opera).
[0458] User requirements (Python):
[0459] def generate_user_requests(num=100):
[0460] themes = ["Legend of the White Snake", "The Butterfly Lovers", "Romance of the Three Kingdoms", "Dream of the Red Chamber"]
[0461] emotions = ["tragic", "poignant", "joyful", "exhilarating"]
[0462] return [{
[0463] "theme": random.choice(themes),
[0464] "emotion": random.choice(emotions),
[0465] "school": f"{random.choice(['Jiang Diao','Li Diao'])} as the main theme + {random.choice(['Yu Diao','Kunqiang'])} as embellishments",
[0466] "duration": random.randint(60, 120)
[0467] } for _ in range(num)]
[0468] The baseline comparison settings are as follows:
[0469] (1) Comparison method:
[0470] Traditional rule engines: sequence generation based on LSTM;
[0471] Single-agent reinforcement learning: without knowledge graph support.
[0472] (2) Evaluation indicators:
[0473] Traditionality: CoA abstract chain verification score (0-1);
[0474] Innovation: Cross-domain cosine similarity (0-1);
[0475] User satisfaction: Simulated user ratings (1-5);
[0476] Generation time: minutes / content;
[0477] Resource consumption: Peak CPU utilization.
[0478] The performance comparison results of the method of the present invention and the comparative methods are shown in Table 6. It was found that, compared with the two comparative methods, the method of the present invention improved the conventionality score by 2.2% and 22.4% respectively, proving that the knowledge graph of agent (1) can effectively maintain the origin of Suzhou Pingtan art; the innovation score improved by 102.4% and 25% respectively, verifying the enhanced effect of cross-domain feature transfer of agent (2); the generation time was reduced by 44.2% and 27.8% respectively, indicating that the multi-agent collaborative reasoning and generation mechanism improved the overall efficiency; the CPU utilization rate was reduced by 23.2% and 19.2% respectively, indicating that multi-agent collaboration has good resource optimization capabilities.
[0479] Table 6 Performance Comparison Experiment Results
[0480]
[0481] To verify the contribution of each core component in the method of this invention, the following comparative experiment was designed:
[0482] (1)-KAG: Knowledge graph of agent (1) removed
[0483] (2)-CD: Remove cross-domain features of agent (2)
[0484] (3)-RL: Reinforcement learning that removes agent (3)
[0485] The results of comparing the traditionality, innovation, and satisfaction of the three methods mentioned above with the method of this invention are shown in Table 7. The lack of knowledge graphs led to a significant decrease in traditionality, indicating that this component plays a key role in the standardization of Pingtan art. The removal of cross-domain features led to a significant decrease in innovation, proving that this component is the main source of innovation in Pingtan bibliography content. The lack of reinforcement learning affected system performance and collaboration efficiency, resulting in a decrease in satisfaction.
[0486] Table 7 Ablation Experiment Results
[0487]
[0488] In summary, the knowledge graph component (first agent) of this invention constructs a structured knowledge system in the field of Suzhou Pingtan, establishes a digital expression of Pingtan art norms through dynamic association of multimodal features, and provides underlying logical support for traditional verification; the cross-domain feature transfer component (second agent) establishes a feature mapping mechanism for multiple art forms, provides an innovative material library, and realizes semantic alignment across art forms; the third agent uses reinforcement learning to achieve collaborative decision-making and resource scheduling among multiple agents, improving the overall efficiency and stability of the system.
Claims
1. A method for generating a long-form Suzhou Pingtan (storytelling and ballad singing) repertoire, characterized in that, include: Based on the constructed Suzhou Pingtan knowledge graph, in response to users' creative needs, the hybrid cognitive engine outputs knowledge graph features; Cross-domain transfer enhancement is performed on the features of the knowledge graph to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features; By integrating the knowledge graph features and the cross-domain enhancement features, the long-form Suzhou Pingtan repertoire is optimized using a reinforcement learning strategy, and the final result is output. Specifically, cross-domain transfer enhancement is performed on the features of the knowledge graph to generate a long-form Suzhou Pingtan (storytelling and ballad singing in Suzhou dialect) bibliography with cross-domain enhancement features, including: The knowledge graph features are analyzed, and cross-domain features are extracted from data of other art forms different from Suzhou Pingtan based on the analysis results. The cross-domain features include cross-art text features, cross-art audio features, and cross-art performance features. The cross-art text features include at least Kunqu Opera text features, the cross-art audio features include at least Beijing Drum Music audio features, and the cross-art performance features include at least Kunqu Opera performance features. Calculate the similarity between the knowledge graph features and the cross-domain features to generate a cross-domain feature transformation matrix; The cross-domain features are converted into cross-domain enhanced features through the cross-domain feature transformation matrix to generate a long list of Suzhou Pingtan stories with cross-domain enhanced features. After generating a long Suzhou Pingtan scripture list with cross-domain enhancement features, the traditionality verification score is obtained by detecting the metrical conformity of the generated long Suzhou Pingtan scripture list with cross-domain enhancement features based on the expert rule base. The Euclidean distance between the text vector of the generated long Suzhou Pingtan scripture list with cross-domain enhancement features and the average semantic vector of the documents in the reference corpus is calculated, and the innovation verification score is obtained from the Euclidean distance. A comprehensive quality index is calculated based on the traditional verification score and the innovative verification score. A quality index threshold is set. If the comprehensive quality index does not reach the quality index threshold, the cross-domain enhancement features are re-converted to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features until the comprehensive quality index reaches the quality index threshold.
2. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, This includes updating the Suzhou Pingtan knowledge graph, specifically including: Based on audience feedback ratings of the live performance, an improvisation event generation mechanism is triggered. This mechanism creates improvisation event nodes to update the Suzhou Pingtan knowledge graph. The improvisation event nodes include timestamps, performance techniques, actors, and bibliographies extracted from the performance context. Based on the audio of Suzhou Pingtan singing style, the hierarchical features of variant features → musicological patterns → schools and inheritance rules are extracted through hierarchical abstraction to update the schools and evolutionary relationships in the Suzhou Pingtan knowledge graph.
3. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, This includes obtaining digital human motion parameters and VR scene instructions based on the output knowledge graph features, driving the digital human to perform in the VR scene, interacting with users to obtain user behavior data streams, dynamically adjusting the knowledge weights in the cognitive engine based on user behavior data, changing the reasoning priority, and optimizing prompt words based on user feedback, performing prompt word self-iteration to guide the cognitive engine and user interaction behavior instructions or templates.
4. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, The generated list of long Suzhou storytelling performances with cross-domain enhancement features includes: Based on the cross-domain feature transformation matrix and the cross-art text features, a literary-enhanced Pingtan text is generated: the hidden state of the Pingtan generator is initialized, the matching degree between the context and each structural unit is calculated through MLP, the selected unit is obtained by weighting and projected onto the Pingtan feature space, the gating vector is calculated based on the hidden state and the selected unit, the migration intensity is controlled by the text migration weight, the cross-art text features and the original Pingtan features are fused by element-wise gating, the hidden state is updated and the literary-enhanced Pingtan text is generated by decoding. Based on the cross-domain feature transformation matrix and the cross-art audio features, a prosodic code conforming to the rhyme scheme of Pingtan is generated and aligned with the text semantics and sentiment: a rhythm sub-matrix is extracted from the cross-domain feature transformation matrix to generate a timestamp sequence. A conditional vector is constructed by combining the statistical features of the text semantics and the timestamp sequence. Based on the audio transfer weight cross-domain rhythm features, a prosodic code is generated by a conditional variational autoencoder. The text transfer weights and audio transfer weights are adjusted when the cross-domain enhancement features are reconverted.
5. The method for generating a long Suzhou storytelling repertoire according to claim 4, characterized in that, The generation of long-form Suzhou Pingtan (storytelling and ballad singing) repertoire with cross-domain enhancement features includes generating performance guidelines adapted to Pingtan based on cross-art performance characteristics: The performance features of Kunqu Opera are mapped to the Pingtan space by projection matrix. The style adaptation is optimized by combining style mixing coefficient and Pingtan default parameters. Then, the performance intensity coefficient is predicted based on text semantics and prosodic encoding to obtain the final performance parameters. Based on the final performance parameters, a sequence of body movements is generated, the original gesture symbols are mapped to the gesture sequence of the Pingtan system, vocal performance parameters are generated, and the final emotional expression parameters are obtained by fusing the emotional features in the text based on the basic emotional expression. Construct a synchronization matrix between body movements and text and calculate the alignment loss. Calculate the multimodal coordination loss of body movements, gestures, vocal performance parameters, and final emotional expression parameters to maintain temporal coordination and stylistic consistency.
6. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, The similarity between the knowledge graph features and cross-domain features is calculated using a graph attention network.
7. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, By integrating the knowledge graph features and the cross-domain enhancement features, the long-form Suzhou Pingtan repertoire is optimized through reinforcement learning strategies, specifically including: A state space is constructed, which is a fusion encoding of the knowledge graph features, the cross-domain enhancement features, and the creation requirements; A reinforcement learning model is trained based on a multi-objective reward function and a proximal policy optimization algorithm. The policy network of the reinforcement learning model selects actions from the mixed action space according to the current state, generates the optimal action sequence using the trained policy network, and optimizes the bibliography. The hybrid action space includes discrete actions for content structure adjustment and continuous actions for parameter fine-tuning. The multi-objective reward function balances at least three objectives: tradition, innovation, and user satisfaction, and introduces a dynamic penalty and reward mechanism.
8. The method for generating a long-form Suzhou Pingtan (storytelling and ballad singing) bibliography according to any one of claims 1 to 7, characterized in that, The first intelligent agent, based on the constructed Suzhou Pingtan knowledge graph, responds to the user's creative needs and outputs knowledge graph features; The second intelligent agent performs cross-domain transfer enhancement on the features of the knowledge graph to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features; The third agent integrates the knowledge graph features and the cross-domain enhancement features, optimizes the long-form Suzhou Pingtan repertoire through a reinforcement learning strategy, and outputs the final result.
Citation Information
Patent Citations
A method and system for processing information of AIGC skit generation
CN118632049B
Intelligent script generation method, system and equipment based on large model
CN119990078A
Three-dimensional warehouse stacker driving motor health state diagnosis method based on knowledge graph
CN120470470A
Suzhou evaluation vocal record genre classification method
CN120596976A