Huangmuo paddling inheriting and innovating method based on large model

By adopting a large-scale model-based approach to the inheritance and innovation of Huangmei Opera, the problems of information acquisition difficulties and insufficient semantic analysis in the digital dissemination of Huangmei Opera have been solved. This approach enables efficient and accurate information analysis and content recommendation, thereby improving the dissemination efficiency and innovation capabilities of Huangmei Opera culture.

CN120930743APending Publication Date: 2025-11-11ANQING NORMAL UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511029982.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In the process of digital dissemination of Huangmei Opera, difficulties in information acquisition, insufficient semantic analysis, and lack of artistic innovation have led to inaccurate recommended content and insufficient innovation capabilities.

Method used

This study employs a large-model-based approach to the inheritance and innovation of Huangmei Opera. Through semantic parsing, a multi-module reasoning system, and a credibility scoring mechanism, it achieves efficient parsing and accurate recommendations for user questions. It also combines a multi-dimensional knowledge base for content generation and optimizes feature vectors through self-distillation to improve the accuracy of recommendation results and user satisfaction.

Benefits of technology

It enables rapid and accurate analysis of Huangmei Opera-related issues, generates high-quality recommendation results, improves the efficiency and innovation of Huangmei Opera culture dissemination, and has self-learning capabilities to ensure the accuracy and reliability of recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930743A_ABST
    Figure CN120930743A_ABST
Patent Text Reader

Abstract

The invention discloses a Huangmuo paddling inheriting and innovating method based on a large model, and the method comprises the steps: obtaining the related problem information of Huangmuo paddling, and analyzing the related problem information; semantic understanding and information completion are conducted on the analysis result through a semantic analysis method and a knowledge graph, and a complete feature vector of the problem is obtained; performing classification of different labels on the complete feature vectors to obtain problem feature vectors of different labels; according to labels of the question feature vectors, the question feature vectors are answered and recommended through corresponding large models, and different recommended contents are obtained; integrating and scoring different recommendation contents to obtain an integrated recommendation result and credibility; semantic analysis retouching is conducted on the integrated recommendation result, the retouching integrated recommendation result is output, and the analysis and innovation content of the Huangmuo play is obtained. According to the method, efficient and accurate information acquisition, semantic analysis and innovative content generation can be realized, so that the digital development of the Huangmuo play culture is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and digitalization of traditional culture, and in particular relates to a method for the inheritance and innovation of Huangmei Opera based on a large model. Background Technology

[0002] Huangmei Opera is one of the five major opera genres in China, boasting a long history and rich cultural connotations. With the development of information technology, the digitization of traditional culture has become an important means of cultural inheritance. However, in the current process of digital dissemination of Huangmei Opera, the following problems exist: difficulties in information retrieval, insufficient semantic analysis, and inadequate recommendations of artistic innovations. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention proposes a method for the inheritance and innovation of Huangmei Opera based on a large model. This method enables efficient and accurate information acquisition, semantic parsing, and innovative content generation, thereby promoting the digital development of Huangmei Opera culture and solving the problems existing in the prior art.

[0004] To achieve the above objectives, this invention provides a method for the inheritance and innovation of Huangmei Opera based on a large model, comprising:

[0005] This process involves acquiring relevant questions about Huangmei Opera and analyzing them. Semantic analysis and knowledge graphs are used to understand and complete the analyzed information, resulting in a complete feature vector for each question. This feature vector is then categorized into different labels, yielding different labeled feature vectors. Based on these labels, a corresponding large-scale model is used to provide solutions and recommendations, resulting in various recommended content. These recommendations are then integrated and scored to obtain an integrated recommendation result and its credibility. Finally, the integrated recommendation result is refined through semantic analysis, and the refined result is output, providing an analysis and innovative content related to Huangmei Opera.

[0006] Optionally, the process of parsing relevant question information includes:

[0007] The relevant question information is denoised and segmented, and the segmented question information is semantically understood using the BERT model to obtain the semantic understanding result, i.e., the parsing result.

[0008] Optionally, the process of semantic understanding and information completion of the parsing results includes:

[0009] The parsing results are semantically understood and encoded using the Transformer model to obtain feature vectors. Information is then supplemented into the feature vectors using a knowledge graph of Huangmei Opera-related knowledge. Finally, contextual information extraction and sentiment analysis are performed on the feature vectors using the Coattention mechanism and the BiLSTM+Attention model to obtain the complete feature vector of the question.

[0010] Optionally, the process of classifying the complete feature vector into different labels includes:

[0011] The BERT model is used to perform semantic modeling on the complete feature vector. The TextCNN model is used to extract features of different semantic dimensions from the semantic modeling results. The extracted semantic dimension features are classified using the Softmax function to obtain different labels and corresponding confidence scores. A threshold is applied to the confidence scores. When the confidence score is greater than the threshold, the corresponding label is directly output. Otherwise, the corresponding label is re-obtained through a compensation mechanism.

[0012] Optionally, the process of solving and recommending based on the problem's feature vectors includes:

[0013] The problem feature vectors are solved and recommended using corresponding large models. These large models include a plot analysis model, a historical background model, a character relationship recognition model, and an art innovation generation model. The plot analysis module uses the ChatGLM model, the historical background model uses the LLaMA model, the character relationship recognition model uses BERT embedding + graph attention mechanism, and the art innovation module uses the BLOOMZ model. All of these large models are optimized using corresponding sample data and fine-tuning methods.

[0014] Optional, the credibility assessment process includes:

[0015] The XGBoost regression framework is used to fuse different recommended content features, and the fused feature results are scored in multiple dimensions. The multiple scores are then weighted and fused to obtain the credibility. The multiple scores include semantic relevance score, language fluency score, diversity score, and user preference matching degree.

[0016] Optionally, the integrated recommendation results are semantically refined through a prompt word optimization mechanism. This mechanism includes semantic prompt words and formatted template matching. The prompt words are unified through formatted template matching. The integrated recommendation results are input into a large model with prompt words that include generation style, extraction of key entities and target requirements, and language refinement, resulting in refined integrated recommendation results.

[0017] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for the inheritance and innovation of Huangmei Opera based on a large model.

[0018] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the aforementioned method for the inheritance and innovation of Huangmei Opera based on a large model.

[0019] Compared with the prior art, the present invention has the following advantages and technical effects:

[0020] Through large-scale model analysis and a multi-module reasoning system, the system achieves rapid and accurate analysis of Huangmei Opera-related issues and generates high-quality recommendation results, effectively improving the dissemination efficiency of Huangmei Opera culture. Combined with an art innovation module, the system integrates traditional elements with modern aesthetic needs, generating innovative content that aligns with contemporary characteristics and promoting the diversified development of Huangmei Opera content. Simultaneously, the system employs a prompt word optimization mechanism and user feedback loop to continuously optimize the format and accuracy of recommended content, providing users with a more personalized and intelligent interactive experience. Furthermore, the system possesses self-distillation optimization capabilities, continuously improving the parsing accuracy and reasoning effect of feature vectors, thereby constantly optimizing recommendation results, enhancing the system's self-learning ability, and ensuring more accurate and reliable recommendation results. Attached Figure Description

[0021] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0022] Figure 1 The flowchart of the Huangmei Opera inheritance and innovation large model system and method is shown in the embodiment of the present invention. Detailed Implementation

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0025] This invention relates to the fields of artificial intelligence and digitalization of traditional culture, and in particular to a method for the inheritance and innovation of Huangmei Opera based on a large model. Based on natural language processing, deep learning, semantic understanding and multiple reasoning models, it realizes accurate recommendation, in-depth analysis and artistic innovation of Huangmei Opera-related content, and enhances the intelligent dissemination and innovation capabilities of Huangmei Opera culture.

[0026] The following explanation addresses the problems existing in the digital dissemination of Huangmei Opera:

[0027] Difficulty in obtaining information: Users need to manually search extensively to obtain information about Huangmei Opera plots, character backgrounds, singing styles, etc., lacking an efficient information integration mechanism. This will lead to insufficient basic information acquisition capabilities in the innovation process of Huangmei Opera, reducing innovation capabilities based on data.

[0028] Insufficient semantic parsing: Existing recommendation systems are mostly based on keyword matching, which cannot accurately understand the complex questions raised by users, resulting in inaccurate parsing results. In the process of Huangmei Opera mentioned above, further association and parsing based on relevant information content is required. Existing word matching methods cannot meet the innovation requirements of Huangmei Opera, thereby reducing the innovation capability of Huangmei Opera.

[0029] Insufficient artistic innovation: The recommendation process for Huangmei Opera content lacks an innovative mechanism and fails to generate creative content by incorporating modern artistic concepts. As a result, effective results cannot be generated during the innovation process of Huangmei Opera, and the recommended content fails to meet the requirements for innovation in Huangmei Opera.

[0030] To address the aforementioned issues, this invention provides a method for the inheritance and innovation of Huangmei Opera based on a large-scale model. The method includes: receiving Huangmei Opera-related questions input by users; inputting user questions into a large-scale model for semantic parsing and information completion to obtain the complete feature vector of the questions; classifying the questions using a tag classification model and pushing them to the corresponding reasoning large-scale model; combining a multi-dimensional knowledge base for feature matching, and inferring and outputting Huangmei Opera content recommendations including categories such as plot analysis, historical background, and singing style, along with a credibility score. The output content can effectively guide Huangmei Opera to align with current information, enhancing its innovative capabilities. The reasoning large-scale model is pre-trained with a large amount of Huangmei Opera-related knowledge base data and fine-tuned through a deep learning model, enabling it to accurately analyze user needs and improve the accuracy of Huangmei Opera-related content recommendations and the quality of innovative content generation. The system directly outputs high-credibility results based on the score, or further optimizes the feature vector through a self-distillation process when the score is insufficient, until a preset standard is reached. This invention effectively enhances users' understanding of Huangmei Opera culture and promotes the digital dissemination and innovative development of Huangmei Opera art.

[0031] The purpose of this invention is to provide a method for the inheritance and innovation of Huangmei Opera based on a large-scale model. By combining semantic parsing, large-scale model reasoning, a multi-module reasoning system, and a credibility scoring mechanism, the system efficiently analyzes Huangmei Opera-related questions input by users and accurately matches corresponding large-scale reasoning models based on a multi-dimensional knowledge base, outputting content recommendations in multiple categories, including plot analysis, historical background, and singing style. Simultaneously, the system continuously optimizes feature vectors through a self-distillation mechanism, improving the accuracy of recommendation results and user satisfaction. Furthermore, the system also possesses a content generation mechanism based on artistic innovation, injecting new vitality into the inheritance and development of Huangmei Opera culture in the digital age.

[0032] The technical solution adopted in this invention is as follows:

[0033] A method for the inheritance and innovation of Huangmei Opera based on a large model:

[0034] The semantic parsing module obtains relevant questions about Huangmei Opera from the input and performs semantic parsing on these questions.

[0035] The feature generation module uses a Transformer-based architecture combined with relevant knowledge graphs to further perform deep semantic understanding and information completion on the semantic parsing results mentioned above. It also expands the feature vector of the question by using knowledge graphs of different knowledge types of Huangmei Opera.

[0036] The label classification module uses the label classification model to classify the complete feature vectors of the above questions to obtain different types of questions. These types of questions include different levels of questions such as plot analysis, historical background, character relationship recognition, and artistic innovation generation, providing separate question inputs for the subsequent large model inference module.

[0037] The large model reasoning module includes large models at different levels, including but not limited to plot analysis, historical background, character relationship recognition, and artistic innovation generation models. These large models at different levels are used to answer and analyze the questions categorized above to generate corresponding recommended content.

[0038] The recommendation and innovation module integrates the output results of different levels of the large model in the above-mentioned large model reasoning module to integrate the recommendation results. The integrated recommendation results include plot analysis, historical background, character relationship analysis or innovative script fragments, and add a credibility score to the recommendation results.

[0039] The refinement module further unifies and formats the integrated recommendation results, and refines their semantics to improve the logical and semantic naturalness of the integrated recommendation results.

[0040] The verification module evaluates the semantically refined integrated recommendation results from multiple dimensions to determine whether they meet the output requirements. If the verification results do not meet the requirements, the results are passed to the self-distillation optimization module.

[0041] The self-distillation optimization module is used to perform self-distillation optimization on the complete feature vector of the problem, thereby improving the accuracy and richness of feature representation;

[0042] The iterative control module is used to re-input the self-distilled and optimized feature vectors into the semantic polishing module, inference module, and recommendation module of the system for iterative processing until the verification module confirms that the output criteria are met.

[0043] In conjunction with the above system, the present invention provides a method for the inheritance and innovation of Huangmei Opera based on a large model, which includes the following:

[0044] A method for the inheritance and innovation of Huangmei Opera based on a large model includes:

[0045] S1. Receive user input of questions related to Huangmei Opera and perform semantic parsing on the input content; the related questions include questions about the script, background, character relationships, and directions for innovation in Huangmei Opera.

[0046] S2. Input the user's question into a large model based on the Transformer architecture and pre-trained with the Huangmei Opera knowledge graph to perform deep semantic understanding and information completion on the question. At the same time, expand the graph content such as Huangmei Opera script, Huangmei Opera background and Huangmei Opera characters to obtain the complete feature vector of the question.

[0047] S3. Input the complete feature vector of the problem into the label classification model to automatically classify the problem;

[0048] S4. Based on the classification results, the questions are pushed to the specially fine-tuned large model reasoning modules. The reasoning modules include, but are not limited to: plot analysis large model, historical background large model, character relationship recognition large model, and artistic innovation generation large model. Each module is specially fine-tuned and trained based on Huangmei Opera knowledge graph and corpus to ensure that it has accurate recommendation and content generation capabilities in its respective professional field.

[0049] S5. Each module outputs corresponding content recommendations or innovative results. The system integrates the outputs of each module to generate structured Huangmei Opera content integration recommendation results. The content includes plot analysis, historical background, character relationship analysis or innovative script excerpts, and adds a credibility score to the recommendation results.

[0050] S6. The integrated results are formatted and semantically refined through a prompt word optimization mechanism to ensure that the recommended content is logically clear and the language is natural, thereby improving the user's reading and comprehension experience.

[0051] S7. When the credibility score of the output result is higher than the preset threshold, the answer is provided to the user directly; if the score is lower than the preset threshold, the system automatically returns to step S2, optimizes the feature vector and semantic completion effect through the self-distillation process, and outputs the result after generating a high credibility result.

[0052] As some embodiments, the user interaction and semantic parsing in step S1 include:

[0053] Receive Huangmei Opera-related questions from users via text input;

[0054] The received information is preprocessed, including noise reduction, word segmentation, and semantic annotation, to improve the accuracy of subsequent semantic parsing.

[0055] As some embodiments, the large model feature encoding and information completion in step S2 include:

[0056] The user-input question is fed into a large model pre-trained based on the Transformer architecture and combined with the Huangmei Opera knowledge graph.

[0057] This large model includes a dataset expansion module, a context association module, and a fine-grained sentiment analysis module. It uses deep learning to perform semantic parsing and information completion on the questions to obtain the complete feature vector of the questions.

[0058] As some embodiments, the multiple large model inferences and problem classifications in step S3 include:

[0059] A label classification model based on a deep neural network is used to perform preliminary classification on the complete feature vector;

[0060] Based on the classification results, questions are automatically pushed to the corresponding reasoning model, achieving intelligent question triage.

[0061] As some embodiments, the inference module corresponding to step S4 includes:

[0062] The plot analysis model includes a plot structure analysis unit and a character relationship extraction unit.

[0063] A large-scale historical background model, which includes historical event matching units and document comparison units;

[0064] A large-scale model for role relationship recognition, which includes a role attribute extraction unit and a relationship graph construction unit;

[0065] A large-scale model for generating artistic innovations includes a style fusion unit and a content transformation unit.

[0066] Among them, each model combines multi-dimensional knowledge base data for feature matching and deep reasoning to generate targeted Huangmei Opera content recommendations.

[0067] As some embodiments, the recommendation results and credibility scores in step S5 include:

[0068] Based on the output of each reasoning model, a multi-parameter weighted algorithm is used to calculate and generate Huangmei Opera recommendation results that include categories such as plot analysis, historical background, and singing style.

[0069] A credibility score is added to the recommendation results to ensure the accuracy and usefulness of the output.

[0070] As some embodiments, the prompt word optimization and content integration in step S6 includes:

[0071] The prompt word optimization mechanism integrates and formats the parsing results output by the large inference model. This mechanism includes:

[0072] Intelligent prompt generation unit: Generates optimized prompts based on contextual information and user history interaction data;

[0073] Formatted template matching unit: Matches the parsed results with the preset display template to ensure that the output format is standardized;

[0074] User feedback loop: Collect user feedback in real time and dynamically adjust prompts and output formats to continuously improve the system's interactive experience and service quality.

[0075] As some embodiments, the self-distillation optimization and high-confidence result generation in step S7 include:

[0076] When the credibility score of the output result is higher than the preset threshold, the system directly provides the answer to the user;

[0077] If the score is lower than the preset threshold, the system will automatically return to step S2 to perform a self-distillation process, optimize the feature vector and semantic completion effect, and output the result only after generating a highly reliable result.

[0078] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method for the inheritance and innovation of Huangmei Opera based on a large model.

[0079] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the aforementioned method for the inheritance and innovation of Huangmei Opera based on a large model.

[0080] The technical details of the above technical solutions will be described in detail below with reference to the accompanying drawings and relevant embodiments:

[0081] The following is combined with Figure 1 The present invention will be further described as follows:

[0082] Step 1: Receive questions related to Huangmei Opera and perform semantic analysis;

[0083] Users can input questions through text input boxes, and the system listens for input events and captures text data. Examples of input questions include: "What is the reason for the Seven Fairies descending to earth in 'The Fairy Couple'?", "What stages has the origin and development of Huangmei Opera gone through?", and "Please generate an innovative Huangmei Opera script dialogue." First, in the preprocessing stage, the input text is initially denoised using regular expressions to remove irrelevant characters and abnormal formatting information, and word segmentation is optimized using BERT-Tokenizer to preserve semantic integrity. Subsequently, a finely tuned BERT is used for semantic understanding, while an entity recognition mechanism based on the Huangmei Opera knowledge graph is introduced: by encoding node attributes (such as repertoire, characters, region, historical background, etc.) in the knowledge graph into embedding vectors during the pre-training stage, and then extending the embeddings through cross-layer attention fusion in the Transformer architecture, knowledge supplementation and context enhancement are achieved.

[0084] During training, the system filters data based on large-scale Huangmei Opera scripts, encyclopedia entries, and question-and-answer data. Noisy data is cleaned using keyword matching and manual review to ensure corpus quality. The annotation method primarily uses question-and-answer pairs, annotating question intent, key information entities, and corresponding node attributes to improve the model's understanding of deep semantics and feature extraction capabilities. Finally, combining the deep encoding of the Transformer architecture and the semantic completion mechanism of the knowledge graph, the system generates a complete and rich question feature vector through steps such as entity recognition, entity alignment and disambiguation (optimized by injecting Huangmei Opera domain dictionaries), attribute retrieval (expanding the Huangmei Opera attribute set based on SPARQL syntax), semantic expansion (combined with BERT context embedding), feature fusion encoding (hierarchical attention mechanism), and feature weight allocation (dynamic weighting using graph attention networks). This provides a foundation for accurate reasoning and recommendation in subsequent modules.

[0085] Step 2: Complete and expand the information and transform it into a feature vector;

[0086] The system employs a Transformer-based deep semantic parsing module, using a 6-layer encoder and 6-layer decoder structure. Each layer contains 8 self-attention heads, the hidden layer dimension is set to 512, and the feedforward network dimension is 2048. The input question is transformed into a feature vector, and a Seq2Seq model is used to enhance parsing capabilities, ensuring the completeness of the question's context. For information completion, the system expands the dataset using a knowledge graph to fill in missing information, and utilizes a Coattention mechanism to mine contextual information from historical questions to enhance contextual relevance. Furthermore, the system uses a BiLSTM+Attention model for sentiment analysis to identify users' emotional tendencies, thereby optimizing recommended content and improving user experience.

[0087] Step 3: Multi-module system reasoning for label classification;

[0088] The system is trained using a large amount of labeled corpus, sourced from Huangmei Opera script databases, Huangmei Opera knowledge encyclopedias, and interactive question-and-answer platforms. The complete feature vector constructed in step two is input into the label classification model to perform semantic label recognition on questions. The classification model employs a deep architecture combining BERT and TextCNN. First, BERT is used for contextual semantic modeling, then TextCNN is used to extract features from different semantic dimensions, and finally, Softmax is used to output multi-label classification probabilities. The system sets a confidence threshold of 0.75, a value determined through multiple rounds of experimental validation. After multiple parameter tuning tests within the 0.70–0.80 range, it was found that setting it to 0.75 achieves the best balance between accuracy and recall, ensuring the reliability of label prediction results without being overly conservative and missing valid questions. When the model's classification confidence exceeds this threshold, the system directly pushes the question to the corresponding inference module. If the confidence level is low, the system activates a compensation mechanism: First, it calls the built-in rule engine, a classification system built on expert rules, to match and identify the input question using a keyword-label mapping database. The labels include categories such as plot analysis, historical background, character relationship identification, and artistic innovation generation. If the rule matching fails to cover the current input, the system further calls the semantic similarity recall module. This module compares the BERT vector representation with the historical corpus to retrieve Top-K similar questions and their corresponding labels, thus determining the most likely category of the current input. This combined judgment mechanism ensures that the classification results have semantic accuracy and comprehensive coverage, effectively improving the model's ability to handle complex, ambiguous, or cold-start problems.

[0089] Step 4: Inference using multiple large models;

[0090] The system inputs questions into corresponding specialized large-scale model inference modules based on classification results. Each module is specifically fine-tuned based on an open-source pre-trained language model. The plot analysis module uses the ChatGLM model as its base, collecting hundreds of Huangmei Opera scripts and paragraph analysis texts as training corpora. It employs a lightweight training technique with efficient LoRA parameters (rank set to 8, learning rate 1e-4) to enable it to understand the plot structure, plot development, and character conflicts of Huangmei Opera. In the historical background module, the LLaMA model is used to perform text alignment fine-tuning in conjunction with Huangmei Opera historical materials, regional background, and historical context. Scalable training objectives are also set, including enhancing factual consistency, improving historical background inference ability, and fine-grained geographical and cultural association modeling, to improve its ability to generate accurate historical annotations. The character relationship recognition module introduces graph modeling capabilities, constructing a character interaction network through BERT embedding + graph attention mechanism and training it through a graph neural network (GAT). The art innovation module uses the BLOOMZ model, fine-tuned in conjunction with modern dialogue, short video texts, and online language corpora, while introducing prefix-tuning for style-oriented control.

[0091] Step 5: Integrate the content and add a credibility score.

[0092] The content generated by each reasoning module is uniformly aggregated into the results fusion module for structured integration and credibility scoring. The system first performs preliminary clustering of the output content by type. Then, through content deduplication and information conflict resolution mechanisms, it merges identical or similar answers and performs confidence-weighted fusion on differing content to generate a unified answer representation. After integration, the system outputs each aggregation result in a structured format according to a predefined structure. Finally, the credibility scoring model, based on the XGBoost regression framework, integrates multiple feature dimensions to quantitatively evaluate and rank the quality and relevance of the answers, providing user-oriented recommendation results.

[0093] 1. Semantic Relevance Score (Rs): The cosine similarity between the content and the user's input question is calculated using BERT vector representation.

[0094]

[0095] Among them User question feature vector Outputs a vector of content for the module. · represents the dot product operation.

[0096] 2. Language fluency score (R) l ): Use the perplexity of the generated language as an inverse indicator of language fluency:

[0097]

[0098] Where x i For the generated language text, a lower Perplexity value indicates more natural and fluent language. In the formula, x... <i This represents all the context content preceding the i-th token, which is the result of the language model predicting the current position x. i The historical information on which it is based. This contextual sequence, as a conditional input, determines the probability of the generated word, thus affecting the perplexity score.

[0099] Standardize by taking the reciprocal, so that the higher the score, the better.

[0100] 3. Diversity score (R d ): Evaluating the diversity of generated text using Distinct-n:

[0101]

[0102] This represents the number of unique n1 segments appearing in the generated text. In other words, it represents the number of n-grams after deduplication. This represents the total number of n2 values ​​in the generated text, including duplicates. Where n1 = 2 and n2 = 3.

[0103] 4. User preference matching degree (R) u ): Generate a preference vector based on the user's past behavior, and calculate the similarity with the current content topic vector:

[0104]

[0105] in This represents the user preference vector. If the user is anonymous, the system uses the default user preference model instead; if the user is a long-term user, an interest profile can be adaptively constructed based on ratings, favorites, and click history.

[0106] 5. The final credibility score is calculated using a weighted fusion method:

[0107] Score final =αR s +βR l +γR d +δR u

[0108] Where α+β+γ+δ=1, the weight parameters can be optimized through historical data.

[0109] Score final Content with a value >0.75 is judged as "high-quality recommended content" and proceeds to the next step of prompt word optimization and output process; otherwise, the push is temporarily suspended, pending further correction or distillation optimization.

[0110] Step Six: Prompt Word Optimization Mechanism

[0111] The integrated results are formatted and semantically refined through a prompt word optimization mechanism. This mechanism comprises three modules: an intelligent prompt generation unit, a formatted template matching unit, and a user feedback loop.

[0112] 1. Intelligent prompt generation unit

[0113] The system generates guiding semantic prompts based on the user's original input and the model's intermediate inference results. final Its basic structure is as follows:

[0114] Prompt final =Style prefix +Entity context +Goal guide

[0115] Style prefix Control the language used to generate the style, such as "Please rewrite it using modern teenage language" or "Please perform it in a style that mimics traditional lyrics."

[0116] Enity context Key entities (characters, events, historical background, etc.) extracted from the previous steps, such as "Young Master Liu Zhiyuan" and "the background is set in the late Qing Dynasty and early Republic of China".

[0117] Goal guide The target requirements for the results are, for example, "output a piece of lyrics" or "generate a piece of dialogue".

[0118] 2. Formatted template matching module

[0119] For different generation goals (such as plot analysis, innovative singing segments, character analysis, etc.), the system has designed different output structures and format templates. The initial Prompt templates are automatically generated by an independent generation model based on the training corpus content, and the system can generate approximately 1 million initial Prompt samples at a time. Subsequently, through the evaluation of classification and generation capabilities after training, a weighted scoring formula is introduced to comprehensively evaluate the Prompt templates corresponding to different labels. The weighted formula comprehensively considers multiple dimensions such as template generation quality, classification accuracy, and user feedback to generate different Prompt template evaluation scores, realizing automatic screening and optimization of Prompt templates. Through weighted scoring, the system can effectively identify the optimal template type and merge similar templates and encapsulate them, ultimately forming an efficient and accurate Prompt-Template set. In the inference stage, the system automatically matches the most suitable template according to the question type label and encapsulates it into a standardized Prompt input, improving the response accuracy and generation quality of various model modules.

[0120] 3. User feedback enhancement module

[0121] The system allows users to rate or annotate generated content, and collects data to construct a Prompt-effect mapping relationship. This data is then used to train a prompt generation strategy model, which automatically learns, through reinforcement learning, which prompt patterns are more likely to produce high-quality results.

[0122] Step 7: Output Results

[0123] The system directly outputs a high-confidence result based on the score. If the score is below a preset threshold, the system automatically returns to the semantic parsing and information completion steps, performing a self-distillation process. During self-distillation, the system uses the currently generated feature vector as a basis, taking the model-predicted soft labels as new pseudo-labels, and re-supervises the feature vector generation and semantic completion process. Specifically, it employs joint optimization using Feature Vector Consistency Loss (MSE Loss) and Semantic Preservation Loss (KL Divergence Loss) to ensure that the new feature representation closely approximates the original prediction while further improving the fit to the semantic context of the input question. After each iteration, the system re-evaluates the confidence score. If the score still does not reach the set threshold, it continues to fine-tune the feature vector and the completed content, with the number of iterations limited to a maximum of 5 to prevent overfitting or training oscillations. When the confidence score of the generated result exceeds 0.75, the system stops the self-distillation process and outputs the final high-confidence content. This step, through the confidence scoring mechanism, ensures the high quality and credibility of the final output result. For results with insufficient ratings, the system will further optimize them until they reach the preset standards, thereby providing users with the highest quality content recommendations.

[0124] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for the inheritance and innovation of Huangmei Opera based on a large model, characterized in that, include: The process involves acquiring relevant question information about Huangmei Opera, parsing this information, and then using semantic analysis and knowledge graphs to perform semantic understanding and information completion on the parsed results, ultimately obtaining the complete feature vector of the question. The complete feature vector is used to classify different labels, resulting in problem feature vectors for different labels; based on... The labels of the problem feature vectors are used to answer and recommend different recommended content through the corresponding large model. The different recommended content is integrated and scored to obtain the integrated recommendation result and credibility. The integrated recommendation result is semantically analyzed and polished, and the polished integrated recommendation result is output to obtain the analysis and innovative content of Huangmei Opera.

2. The method according to claim 1, characterized in that, The process of analyzing relevant problem information includes: The relevant question information is denoised and segmented, and the segmented question information is semantically understood using the BERT model to obtain the semantic understanding result, i.e., the parsing result.

3. The method according to claim 1, characterized in that, The process of semantic understanding and information completion of the parsing results includes: The parsing results are semantically understood and encoded using the Transformer model to obtain feature vectors. Information is then supplemented into the feature vectors using a knowledge graph of Huangmei Opera-related knowledge. Finally, contextual information extraction and sentiment analysis are performed on the feature vectors using the Coattention mechanism and the BiLSTM+Attention model to obtain the complete feature vector of the question.

4. The method according to claim 1, characterized in that, The process of classifying a complete feature vector into different labels includes: The BERT model is used to perform semantic modeling on the complete feature vector. The TextCNN model is used to extract features of different semantic dimensions from the semantic modeling results. The extracted semantic dimension features are classified using the Softmax function to obtain different labels and corresponding confidence scores. A threshold is applied to the confidence scores. When the confidence score is greater than the threshold, the corresponding label is directly output. Otherwise, the corresponding label is re-obtained through a compensation mechanism.

5. The method according to claim 1, characterized in that, The process of solving and recommending based on the feature vectors of a problem includes: The problem feature vectors are solved and recommended using corresponding large models. These large models include a plot analysis model, a historical background model, a character relationship recognition model, and an art innovation generation model. The plot analysis module uses the ChatGLM model, the historical background model uses the LLaMA model, the character relationship recognition model uses BERT embedding + graph attention mechanism, and the art innovation module uses the BLOOMZ model. All of these large models are optimized using corresponding sample data and fine-tuning methods.

6. The method according to claim 1, characterized in that, The process of obtaining credibility includes: The XGBoost regression framework is used to fuse different recommended content features, and the fused feature results are scored in multiple dimensions. The multiple scores are then weighted and fused to obtain the credibility. The multiple scores include semantic relevance score, language fluency score, diversity score, and user preference matching degree.

7. The method according to claim 1, characterized in that, The integrated recommendation results are semantically refined through a prompt word optimization mechanism. This mechanism includes semantic prompt words and formatted template matching. The prompt words are standardized through formatted template matching. The integrated recommendation results are then input into a large model with prompt words that include generation style, extraction of key entities and target requirements, and language refinement, resulting in the refined integrated recommendation results.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the Huangmei Opera inheritance and innovation method based on a large model as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the Huangmei Opera inheritance and innovation method based on a large model as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Commodity recommendation method and commodity recommendation device based on comment text sentiment analysis

    CN110517121A

  • Recommended content processing method and device, electronic equipment and readable storage medium

    CN111625710A

  • BERT improved model-based text sentiment analysis method

    CN114781392A

  • Retrieval method and system for film and television extensive search, electronic equipment and storage medium

    CN119415726A

  • Knowledge question-answering method and system based on knowledge graph of hydropower station

    CN119961408A