Suzhou evaluation long booklist generation method

By constructing a knowledge graph of Suzhou Pingtan and combining cross-domain transfer and reinforcement learning, the problems of dialect and rhythm modeling, multimodal representation and lack of systematic tools in the creation of long Suzhou Pingtan repertoire have been solved. This has enabled efficient and multimodal AI-assisted creation, which is adapted to traditional structures and cultural characteristics, and improves creation efficiency and quality.

CN121389997AActive Publication Date: 2026-01-23CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511948904.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-01-23
Estimated Expiration
2045-12-23

AI Technical Summary

Technical Problem

The creation of long-form Suzhou Pingtan (storytelling and ballad singing) works faces challenges such as insufficient modeling of dialect and rhythm, mismatch between multimodal representation and alignment, incompatibility between existing technologies and creative needs, and a lack of systematic tools. As a result, AI-assisted creation cannot deeply integrate the intrinsic characteristics of Pingtan art, understand the long narrative structure, and achieve human-machine collaboration.

Method used

We construct a knowledge graph of Suzhou Pingtan (storytelling and ballad singing in Suzhou dialect), and generate a long list of Suzhou Pingtan stories with cross-domain enhancement features through cross-domain transfer enhancement and reinforcement learning strategies. We also optimize the creation process by combining audience feedback and user interaction, and achieve deep integration of multimodal information and human-computer collaboration.

Benefits of technology

It lowers the professional creation threshold, improves the efficiency and quality of Suzhou Pingtan long-form storytelling, and ensures that the generated content is aligned with the multimodal information of the actual performance, adapting to traditional structures and cultural characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121389997A_ABST
    Figure CN121389997A_ABST
Patent Text Reader

Abstract

The invention discloses a Suzhou evaluation long booklist generation method, which comprises the steps of outputting knowledge graph characteristics in response to creation requirements of a user based on a constructed Suzhou evaluation knowledge graph; executing cross-domain migration enhancement on the knowledge graph features to generate a Suzhou evaluation long booklist with cross-domain enhancement features; and fusing the knowledge graph features and the cross-domain enhancement features, optimizing the Suzhou evaluation long booklist through a reinforcement learning strategy, and outputting a final result. According to the method, a traditional cultural innovation creation framework integrating the knowledge graph, the cross-domain migration and the reinforcement learning is fused, so that the professional creation threshold of the Suzhou evaluation long booklist can be reduced, and the creation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the intersection of artificial intelligence technology and traditional cultural heritage, and in particular to a Suzhou Pingtan long book generation method. BACKGROUND

[0002] Currently, the creation practice and auxiliary technology research of Suzhou Pingtan long book are fundamentally limited. In the practice level, the creation is highly dependent on the individual experience of experienced scriptwriters, and such talents are facing a serious fault. In the technical assistance level, the existing AI-assisted creation methods are systematically mismatched with the professionalization and multi-modal requirements of Suzhou Pingtan long book. The specific technical defects are as follows:

[0003] 1. Defects in dialect and prosody modeling: Early methods based on rule engines rely on manually coded phonetic and tonal rules, which cannot analyze the "country talk" mix of improvisation and cross-dialect in performance. Current deep learning models (such as LSTM, GPT-4, etc.) are insufficient in modeling due to training data bias and Wu dialect tone variation mechanisms (such as full-turbulent initial consonant fundamental frequency depression), resulting in inaccurate dialect prosody of generated text, and further unable to distinguish the subtle emotions and musical characteristics of related schools such as Jiang and Yu.

[0004] 2. Multi-modal representation and alignment mismatch: Pingtan art is a deep integration of text, music, and performance. Existing digital archives are mostly text scripts, severely lacking key metadata such as programmed action coding, vocal frequency spectrum, and live interaction (gags). This leads to AI model-generated content that deviates from actual performance, such as generated text not associated with "wooden clapping" and corresponding "rise of the body", or virtual human performance that is obviously mechanical due to the lack of these multi-modal information.

[0005] 3. Existing patent technologies are incompatible with Pingtan creation needs: Currently disclosed related patent technologies, such as big model-based script generation methods (publication number CN119990078A) or AIGC short play processing systems (authorized announcement number CN118632049B), have fundamental incompatibility with the rules of Pingtan long creation. They either lack dialect adaptability or are limited by the compact narrative logic of short plays, making them unable to adapt to the traditional structure of Pingtan long division and the situational adaptation needs of school singing, and generally lack cultural training and copyright adaptation for Wu dialect and Pingtan music, making them not directly applicable.

[0006] 4. Lack of systematic tools and collaborative mechanisms: Current technology intervention is scattered, lacking systematic creation tools. Creation process knowledge is not structured, expert experience is isolated, and human-machine collaborative interaction mechanisms are completely absent. At the same time, the standardization of musicality parameters is difficult, further increasing the barriers to technology intervention.

[0007] In summary, the core technical problem faced by the creation of long book of Suzhou Pingtan can be summarized as: how to build a system that can deeply integrate the artistic characteristics of Pingtan (Wu dialect rhythm, school singing aria, performance posture), understand its long narrative structure, and realize human-machine collaborative innovation. SUMMARY

[0008] To solve the above problems, the purpose of the present application is to propose a Suzhou Pingtan long book generation method to realize the creation of Suzhou Pingtan long book.

[0009] The technical solution of the present application is as follows: a Suzhou Pingtan long book generation method, comprising:

[0010] Based on the constructed Suzhou Pingtan knowledge graph, the knowledge graph features are output in response to the user's creation needs;

[0011] Perform cross-domain transfer enhancement on the knowledge graph features to generate a Suzhou Pingtan long book with cross-domain enhanced features;

[0012] Fuse the knowledge graph features and the cross-domain enhanced features, optimize the Suzhou Pingtan long book through reinforcement learning strategy, and output the final result.

[0013] Further, it includes updating the Suzhou Pingtan knowledge graph, specifically including:

[0014] Based on the feedback score of the audience to the real-time performance, trigger the improvisation performance event generation mechanism, which creates an improvisation performance event node for updating the Suzhou Pingtan knowledge graph, and the improvisation performance event node includes timestamp, performance technique, actor and book based on the performance context.

[0015] Further, it includes obtaining digital human action parameters and VR scene instructions based on the output knowledge graph features, driving the digital human to complete the performance in the VR scene, interacting with the user to obtain user behavior data stream, dynamically adjusting the knowledge weight in the cognitive engine based on the user behavior data, changing the reasoning priority, and optimizing the prompt word based on the user feedback, and the prompt word is self-iterated to guide the instructions or templates of the cognitive engine and user interaction behavior.

[0016] Further, perform cross-domain transfer enhancement on the knowledge graph features to generate a Suzhou Pingtan long book with cross-domain enhanced features, specifically including:

[0017] analyzing the knowledge graph features, extracting cross-domain features from other artistic form data different from Suzhou Pingtan based on the analysis results, the cross-domain features including cross-artistic text features, cross-artistic audio features and cross-artistic performance features, the cross-artistic text features at least including Kunqu text features, the cross-artistic audio features at least including Beijing-rhythm big drum audio features, and the cross-artistic performance features at least including Kunqu performance features;

[0018] calculating the similarity of the knowledge graph features and the cross-domain features to generate a cross-domain feature conversion matrix;

[0019] converting the cross-domain features into cross-domain enhanced features through the cross-domain feature conversion matrix to generate a Suzhou Pingtan long book with cross-domain enhanced features.

[0020] Further, after the generation of the Suzhou Pingtan long book with cross-domain enhanced features, a traditional verification score is obtained based on a metrical compliance degree detection of the expert rule base, and an innovative verification score is obtained by calculating the Euclidean distance between the text vector of the generated Suzhou Pingtan long book with cross-domain enhanced features and the average semantic vector of the reference corpus;

[0021] Based on the traditional verification score and the innovative verification score, a comprehensive quality index is calculated, a quality index threshold is set, and when the comprehensive quality index does not reach the quality index threshold, the cross-domain enhanced features are re-converted to generate a Suzhou Pingtan long book with cross-domain enhanced features, until the comprehensive quality index reaches the quality index threshold.

[0022] Further, the generation of the Suzhou Pingtan long book with cross-domain enhanced features includes:

[0023] Based on the cross-domain feature conversion matrix and the cross-artistic text features, a literature-enhanced Pingtan text is generated: the Pingtan generator hidden state is initialized, the context is calculated by MLP to match each structural unit, the selected unit is weighted and projected to the Pingtan feature space, the gating vector is calculated based on the hidden state and the selected unit, the transfer strength is regulated by the text transfer weight, the cross-artistic text features and the original Pingtan features are fused by element-by-element gating, the hidden state is updated and the literature-enhanced Pingtan text is generated by decoding;

[0024] Based on the cross-domain feature conversion matrix and the cross-artistic audio features, a prosody code conforming to the Pingtan tone is generated and aligned with the text semantics and emotion: a rhythm sub-matrix is extracted from the cross-domain feature conversion matrix to generate a timestamp sequence, a condition vector is constructed by combining the statistical features of the text semantics and the timestamp sequence, and a prosody code is generated by a conditional variational autoencoder based on the cross-domain rhythm feature projection after the audio transfer weight;

[0025] The re-conversion adjusts the text migration weight and the audio migration weight when enhancing the cross-domain features.

[0026] Further, the generating the Suzhou Pingtan long book with cross-domain enhanced features comprises generating performance guidance for adaptive Pingtan based on cross-artistic performance features:

[0027] mapping Kunqu performance features to Pingtan space through a projection matrix, combining style mixing coefficients with Pingtan default parameters to optimize style adaptation, and then predicting performance intensity coefficients according to text semantics and prosody coding to obtain final performance parameters;

[0028] generating body movement sequences based on the final performance parameters, mapping original gesture symbols to gesture sequences in the Pingtan system, generating voice cavity performance parameters, and obtaining final emotional expression parameters by fusing emotional features in the text based on basic emotional expressions;

[0029] constructing a synchronization matrix of body movements and text and calculating an alignment loss, calculating multi-modal coordination loss of body movements, gestures, voice cavity performance parameters, and final emotional expression parameters to maintain timing coordination and consistent style.

[0030] Further, the similarity between the knowledge graph features and the cross-domain features is calculated using a graph attention network.

[0031] Further, the knowledge graph features and the cross-domain enhanced features are fused, and the Suzhou Pingtan long book is optimized through a reinforcement learning strategy, specifically including:

[0032] constructing a state space, which is a fusion coding of the knowledge graph features, the cross-domain enhanced features, and the creative requirements;

[0033] training a reinforcement learning model based on a multi-objective reward function using a proximal policy optimization algorithm, wherein the policy network of the reinforcement learning model selects actions from a hybrid action space according to the current state, generates an optimal action sequence using the trained policy network, and optimizes the Pingtan book;

[0034] The hybrid action space includes discrete actions for content structure adjustment and continuous actions for parameter fine-tuning, and the multi-objective reward function balances at least three objectives of traditionality, innovativeness, and user satisfaction, and introduces a dynamic penalty and reward mechanism.

[0035] Further, a first intelligent agent outputs knowledge graph features based on the constructed Suzhou Pingtan knowledge graph in response to user creative requirements.

[0036] A second intelligent agent performs cross-domain transfer enhancement on the knowledge graph features to generate a Suzhou Pingtan long book with cross-domain enhanced features.

[0037] The third agent fuses the knowledge graph features and the cross-domain enhanced features, optimizes the Suzhou Pingtan long book list through a reinforcement learning strategy, and outputs a final result.

[0038] Compared with the prior art, the present application has the following advantages:

[0039] The present application combines the knowledge graph, cross-domain migration and reinforcement learning traditional culture innovation and creation framework, which helps to reduce the professional creation threshold of the Suzhou Pingtan long book list and improve the creation efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 The flowchart of the Suzhou Pingtan long book list generation method of the present application.

[0041] Figure 2 The schematic diagram of the knowledge graph of the master-teacher relationship.

[0042] Figure 3 The schematic diagram of the first agent workflow of the embodiment of the present application.

[0043] Figure 4 The schematic diagram of the second agent workflow of the embodiment of the present application.

[0044] Figure 5 The schematic diagram of the third agent workflow of the embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to more clearly illustrate the technical solutions of the present application, the present application will be further described in detail below in combination with embodiments.

[0046] The present application provides a Suzhou Pingtan long book list generation method, which is completed by multiple agents cooperating to generate the book list, and the generation method is as shown in Figure 1 The specific process of the first agent includes the following contents:

[0047] Please refer to Figure 3 The specific process of the first agent includes the following contents:

[0048] Step A1, based on the OpenSPG semantic framework, define the core ontology structure of the Pingtan field.

[0049] The upper semantic model of Suzhou Pingtan field is constructed, namely the core ontology structure. As the semantic blueprint of the entire knowledge graph, the ontology structure defines the core concepts (entities), concept attributes, and static and dynamic relationships between concepts in the field. The specific definitions are as follows:

[0050] (1) Define the following core entity types:

[0051] "Book" entity: used to represent the specific performance works of Suzhou Pingtan. This entity contains the following key attributes:

[0052] Name: the official name of the book, for example, "Pearl Tower".

[0053] Creation Dynasty: the historical period when the book was created or mainly popular.

[0054] Chapter structure: used to describe the composition structure of the book.

[0055] "Actor" entity: used to represent the performing artists of Suzhou Pingtan. This entity contains the following key attributes:

[0056] Name: the name of the actor.

[0057] Belonging to the genre: the artistic genre to which the actor belongs, which is limited to an enumeration value, such as "Ma tune" or "Yu tune".

[0058] Master-disciple relationship: this attribute or relationship is used to connect with other entities or relationships that describe the master-disciple lineage in subsequent steps.

[0059] "Musical instrument" entity: used to represent the main accompanying instruments used in Pingtan performances, such as Sanxian, Pipa, etc.

[0060] (2) In order to depict the dynamic and time-varying performance behaviors in Pingtan art, the following event types are defined:

[0061] "Performance event": used to represent a complete Pingtan performance, which can be associated with multiple actors, musical instruments, and performance venues, etc.

[0062] "Improvisation event": used to represent the specific improvisation behavior implemented by the actor during the performance. This event contains the following key attributes:

[0063] Techniques used: describe the specific techniques used in improvisation, such as "gags", "say gags", etc.

[0064] Time stamp: accurately records the time point when the improvisation behavior occurs.

[0065] Performing actor: associated with the performer of the event, namely an "actor" entity.

[0066] Related bibliography: Establishes a connection with the specific "bibliographical" entity context in which the improvisation takes place.

[0067] (3) Based on the above definitions, establish a semantic network between entities and events:

[0068] The “bibliography” consists of multiple “chaps”.

[0069] "Actors" belong to a specific "genre".

[0070] An "improvisational performance event" is triggered by a specific "actor" performing a specific "book".

[0071] A "master-apprentice relationship" can be established between "actors" entities, forming a lineage chart.

[0072] Step A2: Based on the defined ontology structure, integrate multi-source data to construct a Suzhou Pingtan knowledge graph.

[0073] Step A2.1: Import structured data such as long-form Suzhou Pingtan repertoire, actors, and musical instruments that have been entered into the public domain, map the imported data to the ontology structure defined in Step 1, and establish semantic relationships between entities (such as actor-repertoire performance relationship).

[0074] Step A2.2: Collect Suzhou Pingtan texts, audio materials, and video materials from the public domain, process unstructured data, form a knowledge graph of master-apprentice relationships, and extract audio features and performance features from support vector indexes.

[0075] The specific implementation process is as follows:

[0076] Step A2.2.1: Based on OpenIE, parse the master-disciple relationship from the text, map the extracted relationship to the actor entity's school and master-disciple relationship attributes, and construct a master-disciple relationship knowledge graph based on ontology structure.

[0077] A specific example is as follows:

[0078] (1) Collect the personal biographies and memoirs of Pingtan artists who have entered the public domain, as well as the publicly published research papers and literary reviews (such as “Passing on the Torch: A Study on the Inheritance Mechanism of Suzhou Pingtan since the Late Qing Dynasty”, “The Construction of the Inheritance System of Suzhou Pingtan in Modern Times - Centered on Guangyu Society” and “A Brief Discussion on the Master-Master Tradition of Suzhou Pingtan”) in electronic or scanned form, convert them into text using OCR technology, remove irrelevant information (such as headers and footers), divide the text into chapters and paragraphs, and standardize punctuation marks and special characters.

[0079] (2) Using Chinese NLP tools such as NLPIR-ICTCLAS or frameworks such as HanLP and Stanford CoreNLP, construct a dictionary of Suzhou Pingtan master-apprentice inheritance, including the names of Pingtan artists and the names of schools.

[0080] (3) Identify typical master-disciple relationship expression patterns, such as: “learned from someone”, “studied by someone”, “disciple of someone”, “someone passed on their skills to someone” or “someone taught someone”, etc., and select the results that conform to the characteristics of master-disciple relationship from the triples extracted by OpenIE.

[0081] (4) Analyze the contextual information of the apprentice-student relationship description, identify auxiliary information such as fellow apprentices, apprenticeship time, and apprenticeship location, and use this information to verify and enrich the extracted relationship.

[0082] (5) Verify the consistency of the description of the same mentor-apprentice relationship in different paragraphs, infer the missing mentor-apprentice relationship by using information such as community organizational structure and timeline, and verify the extraction results manually by domain experts. Construct a feedback mechanism and continuously improve the extraction rules.

[0083] (6) Organize the extracted mentor-apprentice relationships into a knowledge graph. A specific implementation of a mentor-apprentice relationship knowledge graph is as follows: Figure 2 As shown.

[0084] Step A2.2.2: Based on audio processing and feature extraction technology, the MFCC feature vectors of the sanxian and pipa timbres are parsed from the Suzhou Pingtan audio data, and the audio features are associated with the performance events to construct an audio feature database that supports vector retrieval.

[0085] A specific example is as follows:

[0086] (1) Collect high-quality audio materials of representative long Suzhou Pingtan stories by different artists and schools with public authorization or proprietary copyright; divide the audio into shorter segments (e.g., 30 seconds to 2 minutes) according to the chapters or themes of the long stories, use audio processing tools (e.g., librosa, pyAudioAnalysis) to remove background noise, and manually or automatically label the main instruments (sanxian / pipa) in each segment.

[0087] (2) Use the librosa library to extract MFCC features. For each audio segment, set appropriate parameters to extract MFCC coefficients (usually 13-40 coefficients), extract delta and delta-delta features to capture dynamic changes, and normalize the features to ensure the consistency of the feature vectors.

[0088] (3) If the MFCC feature dimension is greater than 128, proceed to step (4); otherwise, proceed to step (5).

[0089] (4) The original MFCC features are processed by dimensionality reduction algorithms such as PCA and t-SNE or deep learning models such as Autoencoder, and compressed to 128 dimensions, while ensuring that the vector after dimensionality reduction can effectively retain the main information of the original features.

[0090] (5) Traditional databases that support vector indexes, such as PostgreSQL+pgvector, record the following data: audio segment ID, original audio path, 128-dimensional feature vector, instrument type (sanxian / pipa), book title, artist information and timestamp.

[0091] Step A2.2.3: Based on computer vision technology, analyze the actors' gestures and expressions from Suzhou Pingtan performance videos, associate visual features with improvisational performance events, and construct a performance feature database that supports multi-dimensional retrieval.

[0092] A specific example is as follows:

[0093] (1) Collect high-quality video materials of representative long stories of Suzhou Pingtan by different artists and schools, either publicly authorized or with their own copyright, unify the video format and resolution, use video processing tools to remove noise and jitter, and extract segments containing the main performance content.

[0094] (2) Detect pixel differences between adjacent frames, identify scene transition points, detect significant changes in actors' movements, and identify key frames; use frameworks such as MediaPipe and OpenPose to detect key points of the hands, track hand movement trajectories, and identify gesture changes; define typical gesture categories in Pingtan (such as orchid finger, sword gesture, palm support, etc.), use CNN or LSTM models to classify gestures, establish a Pingtan gesture dictionary, and record the meaning and application scenarios of each gesture;

[0095] (3) Use frameworks such as FaceMesh and dlib to detect facial key points, extract facial features, define typical expression categories in Pingtan (such as joy, anger, sorrow, happiness, etc.), use CNN or Transformer models to classify expressions, and analyze the meaning and emotional expression of expressions in combination with context.

[0096] (4) Design a relational database (such as MySQL) to store the recognition results. Each record includes: video ID, key frame timestamp, gesture category and confidence level, expression category and confidence level, associated bibliographic content, data index and retrieval, establish time- and content-based indexes, and support multi-dimensional retrieval by gesture, expression, time and other criteria.

[0097] Step A2.3: Establish a unified mapping relationship between "ontology entities and multimodal features" to achieve cross-modal association retrieval capabilities and form a Suzhou Pingtan knowledge graph.

[0098] Step A2.3.1: Construct entity-centric index based on the structured entity data output from step 2.1 and the multi-modal features output from step 2.2.

[0099] Step A2.3.2: Construct modality reverse index based on the multi-modal feature data output from step A2.2.

[0100] Step A2.3.3: Construct unified knowledge graph and cross-modal retrieval system. The specific implementation process is as follows:

[0101] (1) Based on the entity-centric index established in step A2.3.1 and the modality reverse index constructed in step A2.3.2, perform deep semantic fusion to construct a unified graph structure knowledge representation. The graph structure contains three types of core nodes: entity nodes inherit the ontology structure defined in step A2.1, feature nodes are derived from the multi-modal features extracted in step A2.2, and relationship nodes integrate the semantic relationships in step A2.2.1; define and establish "has_audio_feature", "has_gesture", "performed_in", "apprentice_of" and "similar_to" and other multi-type semantic edges to form a complete knowledge graph topology structure.

[0102] (2) Construct a cross-modal unified retrieval engine to realize complex multi-modal association query function. The engine supports composite query based on graph traversal, can simultaneously utilize the aggregation characteristics of entity-centric index and the retrieval ability of modality reverse index, realizes the retrieval mechanism of initiating query from any modality entrance and obtaining cross-modal association results.

[0103] In the embodiment of the present application, the first intelligent agent also performs step A3, updating of Suzhou Pingtan knowledge graph: by capturing real-time events and analyzing the evolution of aria genres, updating the Suzhou Pingtan knowledge graph. Specifically, the following steps are included:

[0104] Step A3.1: Capture real-time events to generate improvisation event nodes and update to the Suzhou Pingtan knowledge graph.

[0105] (1) Real-time collection of audience applause intensity (A) and laughter frequency (L) through audio sensors deployed in the book hall. is the intensity of the i-th applause ); through the TOF (Time of Flight) technology of the infrared depth camera array deployed at the top of the book hall, the 3D point cloud data of the audience seat is obtained, and the audience density (D) is obtained. ; through the infrared depth camera or high-definition camera deployed in the book hall, the face detection and age estimation algorithm (such as a model based on convolutional neural network) is used to identify the audience facial features in real time and estimate the age distribution (A). ; recording a time context of a performance , including a time of day, a day of the week, and whether it belongs to a special time period.

[0106] (2) Calculate sentiment score . wherein the weight coefficient is dynamically adjusted by an LSTM online prediction model , i.e. .

[0107] (3) When the sentiment score exceeds a preset threshold and the duration meets a threshold ( ), trigger the improvisation performance event generation mechanism. The event generation module extracts the current actor and repertoire information from the performance context, combines audio pattern recognition technology to detect improvisation techniques, and creates a complete improvisation event node. The node contains attributes such as timestamp, adopted technique, performing actor, and associated repertoire, and is updated to the knowledge graph in real time and establishes relationships with related entities.

[0108] Step A3.2: Extract hierarchical features from the audio of the performance, from the original data to the evolution law, and update the genre and evolution relationship in the Suzhou performance knowledge graph.

[0109] Through hierarchical abstraction, the hierarchical knowledge of "variant features → music patterns → genres, inheritance rules" is extracted from the original audio, and the core dimensions (such as melody, rhythm, and vocal embellishment) of the performance variant and the evolution logic are obtained. The CoA (Chain of Abstraction, Abstraction Chain) abstraction level is defined in Table 1.

[0110] Table 1 CoA abstraction level

[0111]

[0112] Step A3.2.1: Extract acoustic features (L1) from the original audio (L0). The specific implementation process is as follows:

[0113] (1) Use an audio processing pipeline to frame the original performance audio, with each frame being 10 milliseconds long.

[0114] (2) Extract the fundamental frequency value of each frame using the YIN (Yin Fundamental Frequency Estimator) algorithm or the CREPE model (Convolutional Representation for Pitch Estimation is a deep convolutional neural network-based pitch extraction method) to form the melody contour curve. Analyze the audio signal through the autocorrelation function to detect the beat position and identify the traditional board structure (such as one board three eyes).

[0115] (3) Calculate the trend of the number of beats per minute and the density of notes per unit time.

[0116] (4) Use the music information retrieval tool Essentia to extract MFCC, spectral centroid, zero-crossing rate, etc. Generate a feature vector containing a 100-dimensional F0 sequence, a rhythm event list, and a 13-dimensional MFCC mean for each musical phrase.

[0117] (5) All feature data and meta-information (genre, performer, recording time) are stored as a structured table.

[0118] Step A3.2.2: Map acoustic features (L1) to musicology elements (L2).

[0119] Based on the pre-defined Suzhou Pingtan musicology element dictionary, see Table 2, establish the conversion rules of acoustic features to music concepts. For melody analysis, calculate the adjacent interval change of the pitch sequence, and identify the step (≤ 2 degrees) and jump (≥ 3 degrees) patterns. For rhythm analysis, classify the audio segments into slow, medium or fast tempo according to the beat period and note density. For the cavity skill, use a rule-based detection algorithm, such as identifying the repetition of the pitch pattern in a short period of time (100ms). All mapping rules are verified and optimized by pre-trained machine learning models (random forest, SVM) to ensure that the classification accuracy is not less than 90%.

[0120] Table 2 Definition of musicology elements

[0121]

[0122] Step A3.2.3: Cluster the genre patterns (L3) from the musicology elements (L2). The specific implementation process is as follows:

[0123] (1) Use unsupervised learning algorithms (such as K-means, hierarchical clustering) to cluster musicology features to discover potential variant groups.

[0124] (2) Combine with genre labels, use supervised learning (such as SVM, neural network) to train variant classification model. Define the feature label set of each genre. Associate the clustering results to actors, ages, representative singing sections, and store as knowledge graph nodes.

[0125] Step A3.2.4: Derive evolution rules (L4) from variant patterns. The specific implementation process is as follows:

[0126] (1) Multi-dimensional comparative analysis to explore the driving factors of singing evolution. Analyze the influence of inheritance by comparing the differences in characteristics of different generations of transmission in the same genre; analyze the changes in the characteristics of singing sections in different historical periods to analyze the changes in the times; analyze cultural integration by detecting the degree of infiltration of other types of elements.

[0127] (2) Establish a statistical model or a graph model to quantify the causal relationship.

[0128] A specific implementation example model is given below: wherein, is the complexity of the variant; is the number of generations of inheritance (integer, such as "first generation of inheritors" , "second generation of inheritors" , and so on); is the era variable (binary variable, representing modern (after 2000s), representing traditional (before 2000s)); is the degree of integration of Shanghai opera or other types of elements ( , for example, "30% of Shanghai opera elements" means ).

[0129] Step A3.2.5: Update the genre and evolution relationship in the Suzhou Pingtan knowledge graph. The specific implementation process is as follows:

[0130] (1) For each evolution rule derived from step 3.2.4, perform the following operations:

[0131] (a) Check if the same evolution rule already exists in the knowledge graph.

[0132] (b) If not, create an evolution rule entity node in the knowledge graph. Each evolution rule node contains the following core attributes: rule identifier (an ID uniquely identifying the evolution rule, following the naming convention "EVOL_[genre name]_[timestamp]"), mathematical model description (storing the specific mathematical expression), driving factor set (recording the key factors affecting the evolution of the chorus, including the number of generations, the background of the times, the degree of cultural integration, etc.), confidence score (the rule reliability based on statistical significance test, with a value range of 0-1), applicable time range (the time period during which the evolution rule is valid), and data source (the audio samples and feature data relied on to derive the rule).

[0133] (c) If it already exists, update the attributes of the existing node and record the updated version.

[0134] Step A4, deploy a hybrid cognitive engine based on the knowledge graph, through multi-modal intent analysis and hybrid reasoning, to respond to the user's creation needs and output knowledge graph features.

[0135] Step A4.1: The hybrid cognitive engine receives and processes diverse user inputs, deeply understands the query intent, and pays special attention to queries related to genre features, master-apprentice relationships, and cross-modal rules.

[0136] (1) Multi-modal input interface.

[0137] Support natural language text, voice input, structured query, and other input methods; voice input is converted to text by the ASR system, retaining intonation, pauses, and other paralinguistic information; establish a query type classifier to identify query intent: comparative analysis, trend query, association discovery, recommendation request, etc.; the classifier prioritizes identifying intents related to genre features , master-apprentice relationships , and cross-modal rules .

[0138] (2) Deep semantic analysis.

[0139] Use a domain-adapted BERT model for named entity recognition to accurately extract field-specific entities; use semantic role labeling to analyze the predicate-argument structure of the query; construct a query logic tree to represent complex nested query relationships.

[0140] (3) Context-aware query enhancement.

[0141] Combine user historical query records and preferences to supplement implicit query conditions; use conversation context to resolve anaphora resolution problems; generate standardized logical form representations that directly map to the of the knowledge graph.

[0142] Step A4.2: Perform mixed reasoning. Extract and derive deep knowledge from the knowledge graph in combination with multiple reasoning modes, generating .

[0143] (1) Graph query and data extraction.

[0144] Generate optimized SPARQL queries to improve query efficiency using graph indexing, for Querying genre entities and their attributes (such as origin era, representative tracks, artistic characteristics), for Querying master-apprentice relationship networks (such as master-disciple chains, inheritance time), for Querying cross-modal associations (such as mapping rules between text keywords and audio features); perform multi-hop queries to traverse related entities and relationship networks; extract raw data and establish temporary data views.

[0145] (2) Apply predefined business rule base for rule-based reasoning.

[0146] Rule: Define genre feature inheritance rules (such as "If genre A influences genre B, then B inherits some characteristics of A");

[0147] Rule: Define the transitivity of master-apprentice relationships (such as "The master's master is also in a master-apprentice relationship");

[0148] Rule: Define cross-modal mapping (such as "The 'garden' imagery in the text corresponds to soft rhythm in the audio").

[0149] Step A4.3: Synthesize and enhance cognitive output.

[0150] Convert reasoning results into cognitive value-added output, structured knowledge graph features , and generate natural language descriptions to support different user types.

[0151] Step A4.3.1: Output a structured JSON object , containing three core components, each component is divided into hierarchical substructures to ensure data integrity and machine readability:

[0152] (1) (Genre characteristics): Including the original data of the genre, statistical summary, insight discovery, and decision support sub-layers.

[0153] (2) (Master-apprentice relationship): Covering the original network of master-apprentice relationships, statistical indicators, pattern recognition, and recommendations.

[0154] (3) (Cross-modal rules): Contains rule definitions, application statistics, anomaly detection, and optimization predictions.

[0155] Step A4.3.3: Interaction parameter packaging, output for digital human action mapping and VR scene instruction generation. Includes:

[0156] (1) Action parameter instructions: Analyze from inference results, used for digital human action driving in step A5.1.

[0157] Example field: action_parameters:

[0158] technique_name (technique name, such as "finger rotation"),

[0159] timestamps (lyrics or event timestamps),

[0160] context_entities (context entities extracted from NER, such as technique labels),

[0161] audio_context (audio rhythm information, such as BPM value),

[0162] error_cases (error case identification, such as "finger rotation curvature too low"),

[0163] adjustment_rules (parameter adjustment rules, such as "speed upper limit increased by 10%").

[0164] (2) VR scene instructions: Used for dynamic updating of VR book scenes.

[0165] Example field: vr_commands:

[0166] scene_switch (scene switching instructions, such as "switch from lobby to backstage"),

[0167] event_triggers (real-time event triggers, such as "audience interaction request"),

[0168] environment_updates (environment element adjustments, such as light, sound parameters).

[0169] (3) Context identification: Used for association of user behavior data.

[0170] Example field: context_id (inference ID), session_id (session identification), user_type (user type).

[0171] In some embodiments of the present application, the hybrid cognitive engine for outputting knowledge graph features in step A4 can also be updated, specifically including steps A5 and A6.

[0172] Step A5, analyze the knowledge graph features output in step A4, obtain digital human action parameters and VR scene instructions, realize knowledge visualization interaction, and generate user behavior data stream in real time. Specifically, it includes:

[0173] Step A5.1: Build a digital human with precise performance action simulation capability, realize the closed-loop driving of "evaluation and performance technique-digital human action parameter-scene interaction", and support immersive performance in VR book scene.

[0174] Step A5.1.1: Build a digital human 3D model and bone binding.

[0175] (1) Use a structured light scanner or a multi-view camera array to collect high-resolution 3D data of the face and hands of the evaluation and performance actor (focus on capturing action-sensitive areas such as finger joints and wrists); combine texture mapping (such as 4K RGB images) to generate realistic materials for skin and clothing (such as long gown and cheongsam); define the digital human skeleton system (20-30 key bone nodes) based on human anatomy, with a focus on strengthening the hand skeleton (10-15 sub-bones such as thumb, index finger proximal / middle / distal phalanges); set the degree of freedom constraints for each bone node (such as allowing only bending / extension for finger bones and rotation ±90° for wrists).

[0176] (2) Topology optimization of the hand model, binding muscle deformation weights for "wheel finger" and "sweeping string" actions; add cloth simulation to long gown and water sleeves, set parameters (such as mass 0.5kg / m², tensile stiffness 80%), to ensure natural drooping during action.

[0177] Step A5.1.2: Collection and annotation of evaluation and performance technique action data.

[0178] (1) Deploy an optical motion capture system (such as Vicon) in the book performance area; inertial sensors (such as Xsens): worn on fingers and wrists; force feedback gloves (such as 5DT Data Glove): collect finger bending degree and contact force;

[0179] (2) Divide the action segments according to the evaluation and performance process, and label the core techniques corresponding to each segment. Label the action parameters for each key frame of each technique: finger bending degree, wrist rotation angle, and action speed.

[0180] Step A5.1.3: Real-time action parameter mapping.

[0181] (1) Parse the output of Step A4: Use lightweight parser extraction techniques to extract technique names, lyrics timestamps, contextual entities (e.g., technique tags from NER results).

[0182] (2) Parameter querying and generation: Retrieve baseline parameters from the parameter library based on the technique name and dynamically adjust them in combination with the context provided by Step A4 (e.g., audio tempo, error cases). For example, if Step 4 outputs "finger rotation technique, fast audio tempo," automatically adjust the upper limit of movement speed; if Step A4 provides error cases (e.g., "finger rotation curvature is too low"), immediately adjust the current parameters to avoid repeating errors.

[0183] (3) Smooth transition processing: For consecutive movements (e.g., fan opening and closing to finger rotation), use interpolation algorithms (e.g., cubic spline) to ensure smooth parameter changes and reduce jumps.

[0184] Step A5.1.4: Real-time calibration and feedback.

[0185] (1) Integrate Step A4 verification data: If Step A4 flags parameter deviations, Step A5.1 initiates the calibration process, recalculates parameter averages, or invites experts to review (through human-machine interfaces).

[0186] (2) Use machine learning model assistance: Based on the output of Step A4, update the action parameter mapping relationship to ensure semantic consistency.

[0187] Step A5.2: Build VR book scene. Convert the output of Step A4 into VR scene instructions and handle user interactions.

[0188] Step A5.2.1: Dynamic update of the scene.

[0189] (1) Analyze the event context of Step A4. For example, if Step A4 detects a change in user interest (from knowledge graph reasoning), generate scene switching instructions (e.g., from the book hall to the backstage).

[0190] (2) Real-time rendering adjustment: Based on real-time events provided by Step A4 (e.g., audience interaction requests), update virtual actor behavior or environmental elements (e.g., lighting, sound) in the VR book scene.

[0191] Step A5.2.2: User interaction processing.

[0192] (1) Multi-modal input capture: Capture user input (e.g., gestures, voice commands) through full-body motion capture and speech recognition systems.

[0193] (2) Response generation: Based on the reasoning results of Step A4, drive virtual actors to respond. For example, if Step A4 identifies the user query "finger rotation technique," the digital person demonstrates the finger rotation action and synchronously plays the explanation audio.

[0194] Step A5.3: Generate User Behavior Data Stream.

[0195] Step A5.3.1 Record Data.

[0196] (1) Record User Interaction Details: Include user actions (like gesture types, voice content), interaction timestamps, system responses (digital human action parameters, VR scene status).

[0197] (2) Associate Step A4 Context: Each data point associates with the output identifier from Step A4 (like inference ID, error case ID).

[0198] (3) Collect Performance Metrics: Record system latency, action accuracy (like deviation from standard parameters), user satisfaction (through implicit feedback like interaction duration).

[0199] Step A5.3.2 Data Formatting and Output.

[0200] (1) Data Standardization: Convert raw data into JSON format, including fields like user_action, timestamp, system_response, step4_context, performance_metrics.

[0201] (2) Real-Time Streaming Output: Send data stream in real-time to Step A6 through message queues (like Kafka) or APIs, ensuring low-latency optimization.

[0202] Step A6: Dynamically Update Cognitive Engine Parameters Based on Multi-Modal Interaction Data Stream, Forming a Closed Loop for Enhanced Cognition and Continuous Optimization.

[0203] Step A6.1: Data Integration and Preprocessing. Process user interaction data from Step A5 to provide standardized inputs for optimization, including data cleaning: remove outliers (like transient high latency data), handle missing values (use interpolation or default values). Feature extraction: extract key features from user interaction data: like query_count (number of queries), avg_rating (average rating), skip_rate (skip rate), laughter_response_rate (laughter response rate).

[0204] Step A6.2: Dynamic Adjustment of Knowledge Weights. Adjust the weights of nodes in the knowledge graph based on integrated data to optimize the inference priority of Step A4.

[0205] (1) Dynamically adjust knowledge weights based on user interaction data. Use the following formula to calculate new weights:

[0206]

[0207] wherein, is the current weight of the node, is the number of queries, is the average rating.

[0208] (2) Weight update rule.

[0209] Regularly executed (e.g., every 100 user interactions or every hour) to avoid frequent fluctuations; the weight range is limited to [0, 1] to prevent overflow; set a minimum weight threshold for key nodes (such as high-frequency query nodes) to ensure that basic knowledge is not diluted too much.

[0210] (3) Update the knowledge weight table and synchronize it to the cognitive engine in step A4.

[0211] Please refer to Figure 4 , the second agent performs cross-domain migration enhancement for knowledge graph features to generate Suzhou Pingtan long book list with cross-domain enhanced features, which includes the following processes:

[0212] Step B1: Initialize the agent and build the feature library

[0213] Step B1.1: Analyze the output results of the first agent

[0214] (1) Receive the output results of the first agent .

[0215] (2) Analyze .

[0216] Extract key information for feature extraction: topic keywords (e.g., extract =“Suzhou Jing” from ), structure requirements (e.g., “need fast paragraph”), style indications (e.g., “elegant”, “ardent”).

[0217] Step B1.2: Build a cross-domain feature library, which includes cross-art text features, cross-art audio features, and cross-art performance features. In this embodiment, the cross-art text features include Kunqu text features, the cross-art audio features include Beijing-rhythm big drum audio features, and the cross-art performance features include Kunqu performance features.

[0218] Based on the theme and structure information in , extract features of Kunqu and Beijing-rhythm big drum from external public domain databases.

[0219] Step B1.2.1: Extract Kunqu text features

[0220] (1) According to the subject keyword (e.g. "garden") in the text, search for relevant pieces (e.g. "You Yuan" in "Mudan Ting") from the public domain classical Kunqu repertoire texts, and construct the Kunqu text sequence. The Bi-GRU-based word card structure parser extracts the prosodic constraint matrix, and the specific implementation process is described as follows:

[0221] Input data: Kunqu text sequence , represents the th word card unit (single character or fixed phrase), and the annotated data set includes: level and tone vectors , (level), tone (1 represents the foot position).

[0222] Calculation process:

[0223] (a) Feature embedding layer, target vector , , where represents the embedding operation on the word , outputting a dense vector; represents the vector concatenation operation; is the embedding dimension.

[0224] (b) Bi-directional GRU encoding: forward GRU hidden state , backward GRU hidden state , and concatenating the forward and backward hidden states to obtain a 256-dimensional output vector , where the GRU unit update formula is as follows:

[0225]

[0226] (c) Attention weighted aggregation, focusing on key prosodic positions (e.g. feet, level and tone transition points). The attention score of the hidden state at the th position in the sequence:

[0227] ,

[0228] represents the hidden state after linear transformation by weight matrix , represents the introduction of nonlinearity by the hyperbolic tangent function, represents the weighted sum of the transformed vectors using the parameter vector , obtaining the scalar score. The weighted sum of the hidden state sequence .

[0229] (d)​​ is the structure vector of one tune, using technology to cluster into a typical word structure prototype (matrix row vector) of , is the number of typical word structure prototypes (the number of cluster centers), is the dimension of the structure vector (256 dimensions).

[0230] (2) According to the theme keywords in , relevant tunes are retrieved from the public domain Kunqu classic tune texts to construct a Kunqu text corpus. Based on the theme model, the literary image migration matrix is mined, and the specific implementation process is as follows:

[0231] Input data: Kunqu text corpus , target review theme keywords

[0232] Calculation process:

[0233] (a) Joint modeling of theme and image. By jointly learning the document theme distribution , theme-word distribution and theme-attribute distribution , three key parameters are simultaneously learned by the JointLDA algorithm, with the document set as input, and the image enhancement factor controls the weight balance of different components.

[0234] (b) Find the most suitable theme representation by maximizing the ratio of "theme relevance / style difference". For the image in Kunqu and the target theme in review, the value of matrix element is the maximized theme , that is

[0235] ,

[0236] The numerator is the product of the conditional probability of Kunqu image under theme and the conditional probability of target theme word under theme ; the denominator is the KL divergence between the theme distribution of Kunqu and review under theme , which is used to measure the style difference between the two art forms, and then the image migration matrix , is the number of Kunqu images, is the number of target themes in review.

[0237] Step B1.2.2: Extracting Jingyun Dagu audio features

[0238] According to the structural requirements in , relevant segments are retrieved from the public domain Jingyun Dagu audio. From the relevant segments, the Fourier descriptor of the MFCC time series matrix is extracted to construct the rhythm template and the energy mutation points of the fast-paced passages to form the rhythm anchor point sequence. The specific implementation process is as follows:

[0239] (1) Mel-frequency cepstral coefficient matrix , is the original audio signal, is the Mel spectrum calculation function; 5 frequency energy features , represent the value of the th MFCC coefficient over all frames, represents the Fourier transform, which converts time-domain features to frequency domain, represents the L2 norm, which calculates the energy of the frequency domain representation. Get the rhythm template vector , take the spectral energy of the first 5 MFCC dimensions as the core rhythm features.

[0240] (2) Short-time energy at time point (sum the square of the signal in the window and take the average)

[0241] ,

[0242] is the window size. The time point where the second derivative of the energy exceeds the threshold and is located in the specified time interval is defined as the anchor point

[0243] ,

[0244] The second derivative of the short-time energy represents the acceleration of energy change, is the threshold.

[0245] Step B1.2.3: Extracting Kunqu performance features

[0246] (1) Extract from : theme keywords, performance style requirements, emotional expression requirements, retrieve relevant works from public domain Kunqu classic works texts, and construct a Kunqu performance resource library.

[0247] (2) Extracting body movement features from Kunqu performance videos: Use pose estimation models to extract skeleton keypoint sequences, calculate joint motion trajectories and velocity changes, encode continuous action patterns through LSTM networks, and identify typical body movement combinations (such as "cloud hands", "lying fish", "bright appearance").

[0248] (3) Extracting gesture symbol features from Kunqu performance videos: Detect hand regions and extract gesture contours, classify traditional Kunqu gestures (finger techniques, palm techniques, fist techniques, etc.), construct a gesture-emotion mapping dictionary, and encode the temporal evolution of gesture sequences.

[0249] (4) Extracting vocal cavity performance features from Kunqu performance videos: Extract the fundamental frequency trajectory to reflect pitch changes, detect tremolo features (frequency, amplitude, rate), identify ornamentation patterns (glissando, tremolo, leaning), and analyze breath control patterns: detect air outlet position and duration distribution through energy envelope.

[0250] (5) Extracting facial expression features from Kunqu performance videos: Identify traditional opera expression paradigms (joy, anger, sadness, surprise, etc.), quantify the time-varying curve of expression intensity, and establish an expression-lyric association model.

[0251] (6) Fusing multi-modal features (action + vocal cavity + expression) to construct an emotion intensity curve: , identify emotion transition key nodes, and establish a performance emotion map.

[0252] (7) Through a multi-layer perceptron (Multi-Layer Perceptron) feature fusion network, fuse Kunqu performance features of different modalities into a unified performance feature vector: .

[0253] Step B2: Align knowledge graph with cross-domain feature conversion.

[0254] Step B2.1: Build a shared ontology architecture.

[0255] Based on the cross-artistic features extracted in step B1, build a unified knowledge representation framework. Standardize the prosodic constraint matrix, image transfer matrix, rhythm template, and performance feature vector output by step B1, and establish a shared semantic space for the three artistic forms.

[0256] Step B2.2: Align cross-domain features based on a graph attention network to generate a cross-domain feature conversion matrix.

[0257] The specific implementation process is as follows:

[0258] (1) Build a knowledge graph for the three artistic forms , is a node set, including the evaluation node , Kunqu node , big drum node , etc. is a cross-domain relationship edge (such as "prosody rules-rhyme patterns"); , , , as the initial feature vector corresponding to the artistic form node ;

[0259] (2) Heterogeneous graph attention aggregation. Aggregate the information of neighbor nodes through attention mechanism, update the current node The hidden state vector of the first layer , is an activation function (such as ReLU, Sigmoid, etc.), and is the neighbor node set of node , is a learnable weight matrix, is the initial hidden state vector of node at layer 0, is the attention weight from node to node , and the calculation formula is as follows:

[0260]

[0261] wherein, is a learnable projection matrix, is an attention vector.

[0262] (3) Generate cross-domain similarity matrix. The similarity score between node and node .

[0263] (4) Generate cross-domain conversion matrix , wherein is a soft alignment function, is an artistic form covariance matrix.

[0264] Step B3: Migrate and enhance to generate a Pingshu text, audio, and performance guide.

[0265] Step B3.1: Based on the conversion matrix generated in step B2 and the Kunqu text features extracted in step B1, implement structured literature migration to generate a literature-enhanced Pingshu text.

[0266] (1) Extract the word card structure unit from the prosody constraint matrix in step B1, and use the conversion matrix The structural units are projected across the domain, combined with user creation instructions, and the knowledge graph retrieval results output by the agent (1) and the cross-domain feature conversion matrix generated in step B2.2 , initialize the hidden state of the current poem generator ;

[0267] (2) Dynamic unit selection.

[0268] The current generation context is analyzed by a multi-layer perception machine to determine the matching degree of each structural unit wherein the matching degree of the current context and each structural unit is calculated; the attention weight of each unit is calculated based on the matching degree wherein is the th structural unit in the set of "beginning, continuation, transition, and conclusion", thereby achieving dynamic selection; the selected structural unit is projected into the poem feature space through the conversion matrix.

[0269] (3) Inject conditional gating.

[0270] A gating vector is calculated based on the current state and the selected unit to control the migration strength; a reservation mask is used to ensure the moderate retention of the original poem features; element-wise gating is used to achieve smooth injection of features, avoiding style conflicts. The gating vector wherein is the hyperbolic tangent activation function, which compresses the value to the interval, and the weight matrix is used for linear transformation, indicates the concatenation of the current hidden state and the selected knowledge vector ; the updated hidden state wherein represents element-wise multiplication (Hadamard product), the quju migration weight is a strength coefficient that controls the migration of quju artistic features to poem works, represents the part of the original hidden state that is reserved, represents the part that integrates new knowledge.

[0271] (4) Literary image enhancement generates bibliography text.

[0272] According to the target theme , the corresponding column vector is extracted from the image migration matrix ; the is mapped to an image feature vector through a fully connected layer: wherein , ; using fusion strength parameter (default value is 0.3) to fuse the image feature vector into the gated-in hidden state: ; fuse into the input text decoder to get the word probability distribution at the current time step , where and are the weight matrix and bias term of the output layer, map the hidden state to a probability distribution over the vocabulary; generate a word token (character or word) according to ; pass to the next time step as the hidden state; after time steps, get the prosody encoded text . .

[0273] Step B3.2: Generate prosody encoding that conforms to the prosody of the opera and ensure alignment with the semantics and emotions of the text.

[0274] Step B3.2.1: Multi-scale timestamp control sequence decoding.

[0275] (1) Extract the rhythm sub-matrix from the conversion matrix , calculate the time interval , where is the time scaling factor (default 0.8 to adapt to the speed of the opera) to adapt to the speed of the opera.

[0276] (2) Fine-grained prosody adaptation. The fine time step increment , where is the multi-layer perceptron used to predict the fine time step, is used to map to a high-dimensional feature space, is the position encoding of the position (passing sequence position information), is the vector concatenation operation (combining the three input parts into an input vector for the MLP);

[0277] (3) Multi-scale fusion. The total increment of the time step , where is the "coarse time step increment" (basic component), is the scaling coefficient (default is 0.2 to adjust the overall weight of the fine component), is the activation function, denotes an adaptive weight matrix is a dimension for the pre-transform of , is a text semantic vector, .

[0278] (4) Generate a time stamp sequence by accumulation where is the initial time.

[0279] Step B3.2.2: Conditioned Variational Autoencoder (CVAE) generates prosody code.

[0280] (1) Construct a conditional vector where is a statistical feature (such as mean, variance, minimum, maximum, etc.) of the time stamp sequence; encode the hidden variable where .

[0281] (2) Hidden space sampling: , , calculate the KL divergence loss to constrain the hidden space distribution.

[0282] (3) Construct cross-domain rhythm feature projection where the bass drum transfer weight is the strength coefficient that controls the transfer of Beijing-style bass drum artistic features to the works of Pianpian performances; calculate the decoding condition vector where, is the Pianpian performance style embedding vector; generate prosody code .

[0283] (4) Reconstruction training loss and KL divergence loss, .

[0284] Step B3.2.3: Prosody-text dynamic alignment.

[0285] (1) Construct a text sentiment intensity curve

[0286] .

[0287] (2) Calculate the prosody-text alignment loss

[0288] ,

[0289] discrete approximation

[0290] ,

[0291] where is the coupling strength coefficient.

[0292] Step B3.3: Generating the performance guide for the adapted Pingsheng performance based on the Kunqu performance features.

[0293] Step B3.3.1: Performance feature cross-domain projection and style adaptation.

[0294] (1) Feature parameters after projection to Pingsheng performance space , is the performance feature projection matrix (realizing the spatial mapping from Kunqu to Pingsheng);

[0295] (2) Pingsheng style constraint optimization: style mixing coefficient , Pingsheng performance parameters after style adaptation , where , are the weight matrix and bias term of style adjustment, respectively, is the default Pingsheng performance parameter (style reference value), denotes element-wise multiplication (style mixing by dimension);

[0296] (3) Adjusting performance intensity: performance intensity coefficient (predicted by MLP, controlling the overall performance intensity) , final performance parameters after intensity adjustment .

[0297] Step B3.3.2: Multi-modal performance element decoding.

[0298] (1) Generating body movement sequences: basic body movement parameters (generated by posture decoder) , body movement with rhythm features (LSTM incorporates rhythm coding to ensure movement rhythm) , continuous body movement trajectory (smooth continuous movement obtained by time series interpolation) , where is the body posture decoder (mapping performance parameters to basic movements), is the long short-term memory network for modeling movement rhythm, is the time series interpolation function (generating continuous action sequence based on timestamp);

[0299] (2) Converting gesture symbols: original gesture symbols extracted from Kunqu , mapped to Pingsheng gesture symbols , Pingsheng gesture sequence (finally used to guide performance) , where, denotes the gesture extraction function (separating gesture features from Kunqu performance parameters), Text alignment function (synchronize gestures with text content, timestamps).

[0300] (3) Generate vocal cavity performance parameters: breath control parameters , articulation parameters , volume dynamic parameters , wherein is a breath decoder (fuses performance parameters and rhythm mean to generate breath features), is an articulation parameter prediction dedicated MLP, is a normalization function (maps rhythm encoding absolute value to a reasonable volume range).

[0301] (4) Expression emotion mapping: basic emotional expression based on performance parameters , emotional features extracted from text , final emotional expression parameters after fusion , expression intensity parameters , wherein is an emotional expression decoder, is a text emotion analysis function, is an emotional fusion weight (controls the proportion of basic expression and text emotion), is a smoothing function (avoids sudden changes in expression intensity), is a difference function (calculates the change rate of rhythm encoding).

[0302] Step B3.3.3: Timing synchronization and multi-modal coordination.

[0303] (1) Performance-text alignment constraint. Text embedding sequence

[0304] ,

[0305] Text-action synchronization matrix

[0306] ,

[0307] Alignment loss (measures the difference between the synchronization matrix and the identity matrix I)

[0308] , wherein is a text embedding function (encodes text tokens into vectors), is the square of the Frobenius norm (calculates the sum of the squares of the matrix elements), is the identity matrix.

[0309] (2) Multi-modal coordination optimization: multi-modal coordination loss (ensures that posture, rhythm, expression, gestures, etc. are consistent).

[0310] ,

[0311] wherein the coordination weight coefficient (controls the contribution of each loss term) , , , is the cross-correlation function, is the consistency function (measures the matching degree of gesture sequence and expression.

[0312] Step B3.3.4: Performance guidance integrated generation.

[0313] (1) Structured performance guidance scheme (contains four modules of body, gesture, tone, and expression):

[0314] wherein

[0315] ,

[0316] ,

[0317] ,

[0318] ,

[0319] are the formatted body guidance, gesture guidance, tone guidance, and expression guidance, respectively; , , , are the formatted functions corresponding to the modules (convert parameters into executable guidance language).

[0320] (2) Quality evaluation.

[0321] The overall quality score of the performance guidance scheme:

[0322] ,

[0323] wherein the quality evaluation weight (controls the influence of each index) , is the style consistency score (measures the degree of fit with the evaluation style), is the artistic value score (evaluated by professional standards, such as aesthetics and expressiveness).

[0324] Step B4: Double evaluation and self-optimization.

[0325] Step B4.1: Real-time quality evaluation.

[0326] Step B4.1.1: Meter compliance detection based on expert rule base (traditional verification).

[0327] Based on a pre-defined expert rule base, the system checks the degree to which a work conforms to metrical rules, literary norms, and performance conventions. (Expert rule base) It includes metrical rules such as tonal distribution, rhyme placement, and sentence structure; literary rules such as allusion rules and imagery pairing; and performance rules such as body movements and gestures.

[0328] During the verification process, the system checks each generated work to ensure it does not violate any rules and calculates a score for adherence to tradition. A higher score indicates that the work conforms better to traditional norms and demonstrates stronger artistic inheritance. (Step B3 generates a bibliography of Pingtan storytelling.) Traditional verification scores

[0329] ,

[0330] in Representing a set of rules Size, It is an indicator function that returns 1 if the condition is met, and 0 otherwise. Used for judgment Does it violate the rules? .

[0331] Step B4.1.2: Calculate the semantic space distance between the generated text and the training set (innovation assessment).

[0332] Text of Pingtan (a type of storytelling and ballad singing) to be evaluated semantic vectors Reference Corpus Average semantic vector

[0333] ,

[0334] It refers to the number of documents in the corpus. It is the first in the corpus One document;

[0335] Innovation score ,

[0336] It is the Euclidean distance (L2 norm) between the generated text vector and the average vector of the corpus, and the threshold parameter. Control sensitivity to innovation.

[0337] Step B4.2: Comprehensive quality assessment and optimization triggering.

[0338] .if Proceed to step B4.3. Otherwise, output the bibliography of Pingtan (storytelling and ballad singing). .

[0339] Step B4.3: Self-optimizing parameters.

[0340] (1) Update Kunqu migration weights and big drum migration weights :

[0341] , is the learning rate.

[0342] (2) Perform step B3.

[0343] Please refer to Figure 5 , the third agent fuses knowledge graph features and cross-domain enhanced features, optimizes Suzhou Piaopi long book through reinforcement learning strategy, and outputs the final result.

[0344] The architecture of the third agent includes:

[0345] (1) Input layer:

[0346] Output of the first agent: knowledge graph features , where represents genre features, represents master-apprentice relationships, represents cross-modal rules;

[0347] Output of the second agent: cross-domain enhanced Piaopi book ;

[0348] User requirements: .

[0349] (2) Reinforcement learning layer:

[0350] State space:

[0351] Action space: , see Table 3 for specific definitions.

[0352] Table 3 Definition of action space specific parameters

[0353]

[0354] Reward function: , (amplification innovation reward index).

[0355] The execution process of the third agent is as follows:

[0356] Step C1: Construct state space.

[0357] Step C1.1: Fuse the outputs of the first and second agents and the user needs into a state vector.

[0358] Fuse the above multi-modal features: , , , ;

[0359] Step C1.2: Encode into a fixed dimension state vector by a neural network. The state encoding network is defined as: , , , the state vector dimension: .

[0360] Step C2: Proximal policy optimization loop.

[0361] Step C2.1 Initialize PPO algorithm parameters.

[0362] Max iteration rounds ; Max steps per round ; Discount factor ; GAE parameter ; PPO clipping parameter ; Learning rate .

[0363] Policy network architecture definition: , .

[0364] Step C2.2: Training loop and computing dynamic rewards. For each episode to , perform the following process:

[0365] Step C2.2.1: Initialize state (from the output of Step 1.2), initialize episode cumulative reward .

[0366] Step C2.2.2: For each step to , perform the following process:

[0367] (1) Action sampling , where , the environment state transitions to .

[0368] (2) Compute dynamic rewards.

[0369] (a) Compute the base reward according to traditionality, innovativeness, and user satisfaction.

[0370] Compute traditionality verification score And traditionality reward Then compute innovation score And innovation reward Then compute user satisfaction reward Finally, get base reward Where .

[0371] (b) Apply penalty when traditionality or dissatisfaction condition is met.

[0372] Traditionality deficiency penalty:

[0373] .

[0374] Innovation deficiency penalty:

[0375] .

[0376] (c) Double reward when both traditionality and innovation meet high standards. If > 0.9 and , .

[0377] (d) Final reward ;

[0378] (3) Experience storage and state update.

[0379] (a) Store experience: ;

[0380] (b) Update state , cumulative reward ;

[0381] (c) If termination condition is met, end step loop early.

[0382] Step C2.3: Update policy network using PPO algorithm.

[0383] Step C2.4: Iteration end condition

[0384] Training terminates when one of the following conditions is met:

[0385] (1) Reach ;

[0386] (2) Policy convergence: the average reward of the last episodes changes less than a threshold (e.g., );

[0387] (3) Manual intervention: The user manually stops the training.

[0388] Step C3: Generate the final bibliography using the optimized strategy.

[0389] (1) Optimize the text: Use the trained PPO policy network ( ),from Begin by selecting actions sequentially until the [number]th action. This process generates a state sequence. After Decode into the final text.

[0390] ;

[0391] (2) Optimize audio: Use an audio synthesizer ( According to the optimized prosodic coding and rhythmic action sequences generated by the PPO policy network Generate the final audio. .

[0392] ;

[0393] (3) Optimize performance guidance: use an adapter Optimized performance features and the performance action sequence generated by the PPO policy network Transformed into specific performance instruction .

[0394]

[0395] (3) The final output is a triplet.

[0396] .

[0397] The following example illustrates in detail how this invention is specifically implemented.

[0398] Create a new Suzhou Pingtan (storytelling and ballad singing) long-form book titled "The Love of the Children of the Grand Canal" with the theme of "protection of the Grand Canal cultural heritage".

[0399] User input and interaction (via the multimodal interface of the first agent):

[0400] Users input their creative needs into the system via voice and text: "I hope to create a new long-form Pingtan storytelling piece called 'The Love of the Children of the Grand Canal,' with the theme revolving around the local customs, historical changes, and contemporary protection of the Suzhou section of the Grand Canal. The requirements are to retain the charm of traditional Pingtan while incorporating innovative elements that conform to modern aesthetics."

[0401] First Agent:

[0402] Define the ontology and build the Suzhou Pingtan knowledge graph.

[0403] (1) The system calls the predefined Suzhou Pingtan ontology, which contains core concepts such as "bibliography", "back eye", "role", "event", "plot mode", "tune", "three-string technique", "pipa technique", "singing style", "theme", and their mutual relationships.

[0404] (2) The knowledge graph engine automatically extracts information from authorized Pingtan text libraries (such as "Pearl Tower" and "Jade Dragonfly"), historical and cultural data libraries (such as Suzhou local chronicles and canal historical and cultural data), and Internet public data. For example, "Fengqiao", "Hanshan Temple", "Xumen", "Cao Yun", and "Silk" are extracted as "location" and "historical event" nodes.

[0405] (3) Build a structured Suzhou Pingtan knowledge graph centered on "Grand Canal Cultural Heritage" and related to history, characters, locations, and traditional books.

[0406] First Agent Cognitive Engine Reasoning and Continuous Optimization.

[0407] (1) The cognitive engine reasons based on the Suzhou Pingtan knowledge graph: (theme: cultural heritage protection) + (location: Fengqiao) → associates the classic plot mode "returning to the old place" and "things are not what they used to be".

[0408] (2) The engine checks the rationality of the generated content.

[0409] (3) According to user feedback on the initial generated segment (such as the user skipping a certain type of singing segment multiple times), the optimization mechanism will lower the knowledge weight of that type of singing style and reduce its frequency in subsequent generation.

[0410] Multi-modal interaction and data collection.

[0411] (1) Preliminary generation of the first story outline and a singing segment of "Canal Children's Love", with digital actors performing in the VR Hanshan Temple night camping scene.

[0412] (2) Users interact with digital people through VR devices and express their satisfaction with the narrative paragraph of "Cao Yun's hardships" (longer dwell time and "like" behavior). This user behavior data stream is recorded in real time and used for optimization of the first agent cognitive engine.

[0413] Second Agent:

[0414] Cross-domain feature library construction.

[0415] (1) Analyze the knowledge graph features output by the first agent (such as concepts like "farewell", "water scene", "lyricism", etc.), and determine the need for "literary subtlety" and "low tempo".

[0416] (2) Kunqu text features: call the

Zao Luopao

[0417] (3) Jingyun big drum audio features: call the slow tempo rhythm template and heavy accent anchor sequence that express "sad and mournful" emotions in "Bell in the Sword Pavilion".

[0418] (4) Kunqu performance features: call the "horseback riding" body posture and "water sleeve" gestures of Zhao Guangyan in "A Thousand Miles of Sending the Beijing Wife", as well as the corresponding sad and grand voice.

[0419] Knowledge graph alignment.

[0420] (1) Establish a shared ontology for "Pingtan-Kunqu-Big Drum", with the core being "emotional expression".

[0421] (2) Use the Graph Attention Network (GAT) to calculate the similarity of "Pingtan's

Liangqing tune

Zao Luopao

[0422] (3) Generate a cross-domain feature conversion matrix , which defines how to map the text structure of Kunqu and the rhythm type of Big Drum into the framework of Pingtan.

[0423] Migration and enhancement generation.

[0424] (1) Literary injection: combine the word pattern structure of Kunqu

Zao Luopao

[0425] (2) Rhythm innovation: convert the slow tempo and anchor points of Jingyun big drum 《Bell in the Sword Pavilion》 into the "slow rolling embroidered ball" rhythm code of Pingtan Sanxian, creating a traditional yet innovative sad narrative rhythm.

[0426] (3) Performance guidance: adapt the Kunqu "horseback riding" body posture to a virtual performance plan for a scholar riding a horse at high speed along the canal in Pingtan, and simplify the "water sleeve" action into a more Pingtan characteristic gesture "virtual finger looking far away".

[0427] Evaluation and optimization cycle.

[0428] (1) Traditional verification: Compare the generated lyrics with the classic book database to confirm that the rhyme and tone meet the evaluation of the opera specifications. Traditional score 0.75.

[0429] (2) Innovation evaluation: Evaluate the uniqueness of the use of classics (comparison of the past and present of the canal) and rhythm fusion. Innovation score 0.70.

[0430] (3) Comprehensive quality index . Since it did not reach the preset threshold of 0.80, the system automatically adjusts the migration weight dynamically, slightly increasing the text migration weight of Kunqu literature , and reducing the audio migration weight of the big drum rhythm , and regenerating. After two rounds of iteration, the value reaches 0.82, and the final cross-domain enhanced features (including optimized text, rhythm, and performance scheme) are output.

[0431] Third intelligent agent:

[0432] Global optimization based on reinforcement learning.

[0433] (1) State space: Embed the knowledge graph of "canal", "farewell", "contemporary protection" output by the first intelligent agent, and the cross-domain enhanced features of "slow roll embroidered ball rhythm coding", "Kunqu singing lyrics vector" output by the second intelligent agent, as well as the user preference vector (showing that the user likes "difficult narrative") to form the current state .

[0434] (2) Action space: PPO policy network outputs a set of mixed actions according to .

[0435] Discrete action: Insert a

docks worker song

[0436] Continuous parameters: The intensity (force) of this singing section is set to 0.8, and the BPM (speed) is increased by 15%.

[0437] (3) Reward function calculation:

[0438] Traditional reward: Since the inserted "song" element is derived from real canal history and culture, the traditional reward is +0.6 after S function calculation.

[0439] Innovation reward: The combination of "worker song" and elegant evaluation of singing lyrics has high innovation, and the reward after exponential amplification is +1.2.

[0440] User satisfaction reward: This action meets the user's "difficult narrative" preference, with a reward of +0.5.

[0441] Dynamic reward: since both traditionality and innovativeness perform well, the total reward is doubled, i.e. +4.6 = (0.6 + 1.2 + 0.5) * 2.

[0442] (4) Training and output:

[0443] The third agent is trained through a large number of similar interactions, and finally learns an optimal strategy.

[0444] Applying this strategy, the first draft of "The Canal Children" output by the agent (2) is globally optimized: adjusting the order of each chapter to enhance the dramatic effect, inserting actions that strengthen emotions at key emotional nodes (such as adjusting BPM), and finally generating a structured document of the book of Pingshu.

[0445] This document contains:

[0446] Text content: complete and optimized chapter-by-chapter speech and lyrics.

[0447] Audio configuration: detailed suggestions for using music, rhythm patterns, and BPM changes.

[0448] Performance guidance: a virtual digital human performance plan that integrates Kunqu body movements for key singing sections.

[0449] Final output: the system presents this long Pingshu book of "The Canal Children" that combines traditional roots and modern aesthetics to users through digital human performance and VR scenes, completing the intelligent creation process from zero to one.

[0450] Summary of the beneficial effects of this embodiment: through this example, it is clearly shown how the three agents work together to transform a user's abstract theme requirement into a specific, high-quality, and innovative long Pingshu book of Suzhou Pingshu. The entire process demonstrates that clear technical means (knowledge graph, GAT, PPO, etc.) solve the technical problem of how to reduce professional barriers and improve creative efficiency, and achieve significant technical effects.

[0451] It should be noted that the specific methods of the above embodiments can form a computer program product, therefore, the computer program product implemented by the present application can be stored on one or more computer usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.).

[0452] The simulation experiment of the method for implementing the application is carried out, the hardware environment is shown in Table 4, and the software environment is shown in Table 5.

[0453] Table 4 Experimental hardware environment

[0454]

[0455] Table 5 Experimental Software Environment

[0456]

[0457] The dataset is derived from publicly available materials from the Suzhou Pingtan Museum and manually scanned publicly published Pingtan repertoire, including works such as "The Golden Fan," "Three Smiles," "The Golden Phoenix," and "The Legend of the White Snake." It includes 20 hours of vocal recordings labeled with genre, emotional, and vocal technique tags, as well as 2 hours of performance videos labeled with the timing of instrument movements and facial expressions. The cross-domain art dataset includes: 50 ci (lyric) structures (Kunqu Opera), 200 rhythmic templates (Beijing Drum Song), and 50 emotional expression segments (Shanghai Opera).

[0458] User requirements (Python):

[0459] def generate_user_requests(num=100):

[0460] themes = ["Legend of the White Snake", "The Butterfly Lovers", "Romance of the Three Kingdoms", "Dream of the Red Chamber"]

[0461] emotions = ["tragic", "poignant", "joyful", "exhilarating"]

[0462] return [{

[0463] "theme": random.choice(themes),

[0464] "emotion": random.choice(emotions),

[0465] "school": f"{random.choice(['Jiang Diao','Li Diao'])} as the main theme + {random.choice(['Yu Diao','Kunqiang'])} as embellishments",

[0466] "duration": random.randint(60, 120)

[0467] } for _ in range(num)]

[0468] The baseline comparison settings are as follows:

[0469] (1) Comparison method:

[0470] Traditional rule engines: sequence generation based on LSTM;

[0471] Single-agent reinforcement learning: without knowledge graph support.

[0472] (2) Evaluation indicators:

[0473] Conventional: CoA abstract chain verification score (0-1);

[0474] Innovative: Cross-domain cosine similarity (0-1);

[0475] User satisfaction: simulated user score (1-5);

[0476] Generation time-consuming: minutes / content;

[0477] Resource consumption: CPU peak occupancy.

[0478] The performance comparison results of the method of the application and the comparative methods are shown in Table 6. It is found that compared with the two comparative methods, the method of the application improves the conventional score by 2.2% and 22.4% respectively, proving that the knowledge graph of the agent (1) can effectively maintain the origin of Suzhou Pingtan art; the innovation increases by 102.4% and 25% respectively, verifying the enhancement effect of cross-domain feature transfer of the agent (2); the generation time-consuming is reduced by 44.2% and 27.8% respectively, indicating that the multi-agent collaborative reasoning and generation mechanism improves the overall efficiency; the CPU occupancy is reduced by 23.2% and 19.2% respectively, indicating that multi-agent collaboration has good resource optimization capability.

[0479] Table 6 Performance comparison experimental results

[0480]

[0481] In order to verify the contribution of each core component in the method of the application, the following comparative experiments are designed:

[0482] (1) -KAG: remove the knowledge graph of the agent (1)

[0483] (2) -CD: remove the cross-domain features of the agent (2)

[0484] (3) -RL: remove the reinforcement learning of the agent (3)

[0485] The conventional, innovative and satisfaction of the above three methods compared with the method of the application are shown in Table 7. The absence of the knowledge graph leads to a significant decrease in the conventional performance, indicating that this component plays a key role in the specification of the inherited Pingtan art; the removal of the cross-domain features leads to a significant decrease in innovation, proving that this component is the main source of Pingtan book content innovation; the absence of reinforcement learning affects the system performance and collaboration efficiency, resulting in a decrease in satisfaction.

[0486] Table 7 Ablation experiment results

[0487]

[0488] In summary, the knowledge graph component (first intelligent agent) of the method of the present application constructs a structured knowledge system in the Suzhou Pingtam field, establishes a digital expression of Pingtam artistic norms through dynamic correlation of multi-modal features, and provides bottom logic support for traditional verification; the cross-domain feature transfer component (second intelligent agent) establishes a feature mapping mechanism for multiple artistic forms, provides an innovative material library, and realizes semantic alignment across artistic forms; the third intelligent agent uses reinforcement learning to realize collaborative decision-making and resource scheduling among multiple intelligent agents, improving the overall efficiency and stability of the system.

Claims

1. A method for generating a long-form Suzhou Pingtan (storytelling and ballad singing) repertoire, characterized in that, include: Based on the constructed Suzhou Pingtan knowledge graph, in response to users' creative needs, the hybrid cognitive engine outputs knowledge graph features; Cross-domain transfer enhancement is performed on the features of the knowledge graph to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features; By integrating the knowledge graph features and the cross-domain enhancement features, the long-form Suzhou Pingtan repertoire is optimized through a reinforcement learning strategy, and the final result is output.

2. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, This includes updating the Suzhou Pingtan knowledge graph, specifically including: Based on audience feedback ratings of the live performance, an improvisation event generation mechanism is triggered. This mechanism creates improvisation event nodes to update the Suzhou Pingtan knowledge graph. The improvisation event nodes include timestamps, performance techniques, actors, and bibliographies extracted from the performance context. Based on the audio of Suzhou Pingtan singing style, the hierarchical features of variant features → musicological patterns → schools and inheritance rules are extracted through hierarchical abstraction to update the schools and evolutionary relationships in the Suzhou Pingtan knowledge graph.

3. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, This includes obtaining digital human motion parameters and VR scene instructions based on the output knowledge graph features, driving the digital human to perform in the VR scene, interacting with users to obtain user behavior data streams, dynamically adjusting the knowledge weights in the cognitive engine based on user behavior data, changing the reasoning priority, and optimizing prompt words based on user feedback, performing prompt word self-iteration to guide the cognitive engine and user interaction behavior instructions or templates.

4. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, Cross-domain transfer enhancement is performed on the features of the knowledge graph to generate a long-form Suzhou Pingtan (storytelling and ballad singing in Suzhou dialect) bibliography with cross-domain enhancement features, specifically including: The knowledge graph features are analyzed, and cross-domain features are extracted from data of other art forms different from Suzhou Pingtan based on the analysis results. The cross-domain features include cross-art text features, cross-art audio features, and cross-art performance features. The cross-art text features include at least Kunqu Opera text features, the cross-art audio features include at least Beijing Drum Music audio features, and the cross-art performance features include at least Kunqu Opera performance features. Calculate the similarity between the knowledge graph features and the cross-domain features to generate a cross-domain feature transformation matrix; The cross-domain features are converted into cross-domain enhanced features through the cross-domain feature transformation matrix to generate a long list of Suzhou Pingtan stories with cross-domain enhanced features.

5. The method for generating a long Suzhou storytelling repertoire according to claim 4, characterized in that, After generating a long-form Suzhou Pingtan scripture with cross-domain enhancement features, the traditionality verification score is obtained by detecting the metrical conformity of the generated long-form Suzhou Pingtan scripture with cross-domain enhancement features based on the expert rule base. The Euclidean distance between the text vector of the generated long-form Suzhou Pingtan scripture with cross-domain enhancement features and the average semantic vector of the documents in the reference corpus is calculated, and the innovation verification score is obtained from the Euclidean distance. A comprehensive quality index is calculated based on the traditional verification score and the innovative verification score. A quality index threshold is set. If the comprehensive quality index does not reach the quality index threshold, the cross-domain enhancement features are re-converted to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features until the comprehensive quality index reaches the quality index threshold.

6. The method for generating a long Suzhou storytelling repertoire according to claim 5, characterized in that, The generated list of long Suzhou storytelling performances with cross-domain enhancement features includes: Based on the cross-domain feature transformation matrix and the cross-art text features, a literary-enhanced Pingtan text is generated: the hidden state of the Pingtan generator is initialized, the matching degree between the context and each structural unit is calculated through MLP, the selected unit is obtained by weighting and projected onto the Pingtan feature space, the gating vector is calculated based on the hidden state and the selected unit, the migration intensity is controlled by the text migration weight, the cross-art text features and the original Pingtan features are fused by element-wise gating, the hidden state is updated and the literary-enhanced Pingtan text is generated by decoding. Based on the cross-domain feature transformation matrix and the cross-art audio features, a prosodic code conforming to the rhyme scheme of Pingtan is generated and aligned with the text semantics and sentiment: a rhythm sub-matrix is ​​extracted from the cross-domain feature transformation matrix to generate a timestamp sequence. A conditional vector is constructed by combining the statistical features of the text semantics and the timestamp sequence. Based on the audio transfer weight cross-domain rhythm features, a prosodic code is generated by a conditional variational autoencoder. The text transfer weights and audio transfer weights are adjusted when the cross-domain enhancement features are reconverted.

7. The method for generating a long Suzhou storytelling repertoire according to claim 6, characterized in that, The generation of long-form Suzhou Pingtan (storytelling and ballad singing) repertoire with cross-domain enhancement features includes generating performance guidelines adapted to Pingtan based on cross-art performance characteristics: The performance features of Kunqu Opera are mapped to the Pingtan space by projection matrix. The style adaptation is optimized by combining style mixing coefficient and Pingtan default parameters. Then, the performance intensity coefficient is predicted based on text semantics and prosodic encoding to obtain the final performance parameters. Based on the final performance parameters, a sequence of body movements is generated, the original gesture symbols are mapped to the gesture sequence of the Pingtan system, vocal performance parameters are generated, and the final emotional expression parameters are obtained by fusing the emotional features in the text based on the basic emotional expression. Construct a synchronization matrix between body movements and text and calculate the alignment loss. Calculate the multimodal coordination loss of body movements, gestures, vocal performance parameters, and final emotional expression parameters to maintain temporal coordination and stylistic consistency.

8. The method for generating a long Suzhou storytelling repertoire according to claim 4, characterized in that, The similarity between the knowledge graph features and cross-domain features is calculated using a graph attention network.

9. The method for generating a long Suzhou storytelling repertoire according to claim 1, characterized in that, By integrating the knowledge graph features and the cross-domain enhancement features, the long-form Suzhou Pingtan repertoire is optimized through reinforcement learning strategies, specifically including: A state space is constructed, which is a fusion encoding of the knowledge graph features, the cross-domain enhancement features, and the creation requirements; A reinforcement learning model is trained based on a multi-objective reward function and a proximal policy optimization algorithm. The policy network of the reinforcement learning model selects actions from the mixed action space according to the current state, generates the optimal action sequence using the trained policy network, and optimizes the bibliography. The hybrid action space includes discrete actions for content structure adjustment and continuous actions for parameter fine-tuning. The multi-objective reward function balances at least three objectives: tradition, innovation, and user satisfaction, and introduces a dynamic penalty and reward mechanism.

10. The method for generating a long-form Suzhou storytelling repertoire according to any one of claims 1 to 9, characterized in that, The first intelligent agent, based on the constructed Suzhou Pingtan knowledge graph, responds to the user's creative needs and outputs knowledge graph features; The second intelligent agent performs cross-domain transfer enhancement on the features of the knowledge graph to generate a long list of Suzhou Pingtan stories with cross-domain enhancement features; The third agent integrates the knowledge graph features and the cross-domain enhancement features, optimizes the long-form Suzhou Pingtan repertoire through a reinforcement learning strategy, and outputs the final result.

Citation Information

Patent Citations

  • A method and system for processing information of AIGC skit generation

    CN118632049B

  • Intelligent script generation method, system and equipment based on large model

    CN119990078A

  • Intelligent creation method and system based on knowledge and data dual drive

    CN117235249A

  • Interdisciplinary scientific research potential assessment method based on dynamic multi-modal knowledge graph

    CN120181409A

  • Three-dimensional warehouse stacker driving motor health state diagnosis method based on knowledge graph

    CN120470470A