A macro-micro fusion method and system for optimizing copywriting language
Through the copywriting language optimization method of macro-micro fusion, combined with field classification, structured inspector and large model technology, the problem of insufficient copywriting optimization in the existing technology is solved, multi-dimensional optimization of copywriting is achieved, and the quality and efficiency of copywriting is improved.
Patent Information
- Application Number
- CN202510518888.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing AI-based copywriting optimization methods have shortcomings in macro structure, field professionalism and multi-dimensional integration optimization, resulting in the optimized copywriting that may lack logic, professionalism and overall quality improvement.
The copywriting language optimization method of macro-micro-fusion is adopted, and through intelligent macro-optimization and micro-modification mechanism, the global structure is rewritten and refined by using field classification, structured inspector, multi-level optimization framework, sentence pattern knowledge base and large models to achieve multi-dimensional and systematic optimization of copywriting.
It significantly improves the structural rationality, logical rigor, language accuracy, professionalism and expressiveness of copywriting, meets the needs of high-standard copywriting in different industries and achieves a significant improvement in the efficiency and quality of copywriting creation.
Smart Images

Figure CN120046587B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and natural language processing, and particularly relates to a macro-micro integrated copywriting language optimization method and system. Background Art
[0002] With the continuous development of artificial intelligence (AI) and natural language processing (NLP) technologies, the intelligence and automation of copywriting creation have gradually become a hot topic in the industry. Modern enterprises and institutions, especially in professional fields such as technology, finance, healthcare, and education, have a higher demand for high-quality and high-efficiency copywriting creation. In these fields, the structured expression, logical rigor, and language accuracy of copywriting directly affect the effective dissemination of information, the establishment of a professional image, and the trust of the audience. To improve the quality of copywriting, many organizations have begun to attempt to use AI technology to assist in copywriting creation, including using large-scale pre-trained language models (such as GPT, BERT, etc.) to generate and optimize copywriting content.
[0003] However, the existing AI-based copywriting optimization methods and systems still have several deficiencies in practical applications, mainly reflected in the following aspects:
[0004] (1) Insufficient macro-structure optimization: Currently, many copywriting optimization technologies mainly focus on the polishing and improvement of the language level, such as grammar correction, vocabulary replacement, and improvement of fluency. However, they often neglect the optimization of the overall structure and core content of the copywriting. Although large models can improve language expression to a certain extent, at the macro level, in terms of the logic, paragraph organization, information hierarchy, support of arguments and evidence, clarity of information transmission, etc. of the copywriting, the existing technologies still cannot meet the high standards of the industry requirements, resulting in the optimized copywriting may be "similar in form" but "scattered in spirit".
[0005] (2) Lack of domain expertise: Most of the existing copywriting generation and optimization tools are trained based on general language models, lacking in-depth calibration for specific industries and the injection of professional knowledge. This makes the generated copywriting significantly lacking in professionalism, accuracy of term usage, industry-conventional expressions, and compliance with specific norms. Especially for fields such as technology, finance, and healthcare, which have extremely high requirements for copywriting quality, the copywriting generated by general models is difficult to meet their rigor and professionalism. In addition, the term processing ability of the existing systems is limited, and it is often unable to accurately control the unique expression styles and nuances of the industry, resulting in the generated copywriting being difficult to meet the industry-recognized standards.
[0006] (3) Single optimization dimension and lack of integration: Current copywriting optimization methods are often limited to single-level improvements, such as grammar checking or synonym replacement, and fail to achieve comprehensive and systematic optimization of multiple dimensions such as copywriting structure, content, logic, language, and style. The improvement of copywriting quality is the result of the synergy of multiple factors. Existing methods usually lack an integrated framework and cannot conduct unified and coordinated evaluation and optimization in different dimensions, resulting in scattered optimization effects and difficulty in achieving a breakthrough improvement in the overall quality of copywriting.
[0007] Therefore, there is an urgent need for a copywriting optimization method and system that can integrate macro-structure planning and micro-language refinement, take into account both field expertise and expression accuracy, and achieve multi-dimensional, systematic, and intelligent copywriting optimization to overcome the limitations of existing technologies. Summary of the invention
[0008] The main purpose of the present invention is to solve the deficiencies of the existing copywriting optimization methods in the macro structure, field expertise and multi-dimensional integrated optimization, and to provide a macro-micro integrated copywriting language optimization method and system. The present invention aims to systematically improve the structural rationality, logical rigor, language accuracy, professionalism and expressiveness of copywriting creation through the combination of intelligent macro optimization and micro polishing mechanism, thereby greatly improving the efficiency and quality of copywriting creation and meeting the urgent needs of different industries for high-standard copywriting.
[0009] In order to achieve the above-mentioned object, the first aspect of the present invention provides a macro-micro fusion copywriting language optimization method, the method comprising the following steps:
[0010] S1: Classify the texts to be optimized and determine their industry and topic tags;
[0011] S2: Select a set of structured checkers corresponding to the industry field, determine the highlight coverage of the structured checker set relative to the excellent texts in the same field of the copy to be optimized, and if the coverage is lower than a preset threshold, generate text inspection content based on the big model and update the structured checker set;
[0012] S3: loading the updated structured checker set according to the topic tag, and generating a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the structured checker set;
[0013] S4: using the big model to rewrite the global structure of the text to be optimized according to the multi-level optimization framework and the optimization thinking chain, and generate a primary optimized text;
[0014] S5: By performing syntactic analysis and semantic annotation on excellent texts, break down the sentences of the excellent texts into fixed templates and variable parameters to extract specific sentence structures in the industry field, and construct a reusable sentence pattern knowledge base including phrase level, sentence pattern level, and paragraph level;
[0015] S6: Perform sentence segmentation on the primary optimized text to generate a sentence sequence with context markers;
[0016] S7: Adopt a vectorization technique based on pre-trained language model embedding to calculate the semantic similarity between each sentence in the sentence sequence and the sentence pattern templates in the reusable sentence pattern knowledge base, and match the sentence pattern template with the highest similarity for each sentence;
[0017] S8: Based on the matched sentence pattern template and the sentence sequence with context markers, use a large model to perform refined rewriting on the sentences in the sentence sequence to generate the final optimized sentences.
[0018] As an optional implementation manner of the first aspect of the present application, the determination of the highlight coverage of the structured checker set relative to the excellent texts in the same field as the text to be optimized in step S2 includes: inputting the excellent text into the structured checker set, counting the number of highlight features recognized by the structured checker set in the excellent text, and calculating the ratio of the number to the total number of highlight features in the excellent text as the coverage; if the coverage is lower than a preset threshold, then by comparing the differences between the excellent text and the text to be optimized, a new checker is obtained based on a large model to generate text inspection content, and the new checker includes function description, detection dimension, evaluation criteria, and improvement suggestions.
[0019] As an optional implementation manner of the first aspect of the present application, in step S3: the multi-level optimization framework includes a structure layer, a semantic layer, and a style layer, where: the structure layer is used to optimize the logical process and information organization of the text, the semantic layer is used to ensure the accuracy of terms and the integrity of content, and the style layer is used to adjust the expression style to conform to industry norms; the optimization thinking chain includes a text analysis step and a text optimization step, where: the text analysis step includes identifying the keywords and core ideas of the text to be optimized, and the text optimization step includes structure adjustment, content supplementation, and term unification.
[0020] As an alternative implementation of the first aspect of the present application, in step S5, syntactic analysis and semantic annotation are performed on excellent texts, including: using NLP tools for dependency syntactic analysis and named entity recognition to identify fixed templates and variable parameters in sentences; abstracting the fixed templates into template structures containing placeholders and recording the filling examples of the variable parameters; the reusable sentence pattern knowledge base is indexed according to fields, topics, and sentence pattern functions and includes templates at the phrase level, sentence pattern level, and paragraph level.
[0021] As an alternative implementation of the first aspect of the present application, in step S6, generating a sentence sequence with context markers includes: using a sentence segmentation tool to segment the primary optimized text to obtain a list of sentence sequences; traversing the list, for the i-th sentence, recording the content of its previous sentence and the next sentence as context markers.
[0022] As an alternative implementation of the first aspect of the present application, in step S7, the vectorization technology based on pre-trained language model embedding includes: converting each sentence in the sentence sequence and the sentence pattern templates in the sentence pattern knowledge base into vector representations of a fixed dimension; calculating the cosine similarity between the vector of the sentence and the template vector, and selecting the template with the highest similarity as the matching result.
[0023] As an alternative implementation of the first aspect of the present application, step S8 includes: for each sentence in the sentence sequence and its best-matched sentence pattern template, as well as the recorded context information and example sentences and excellent vocabulary obtained from the sentence pattern template, forming a polishing instruction for the large model; inputting the original sentence, the matching template, the context, the examples, the vocabulary, and the polishing instruction into the large model to perform refined rewriting on the original sentence, and the rewriting includes sentence pattern adjustment, word replacement, and expression optimization; recombining all the sentences in the sentence sequence that have been refined and rewritten by the large model in order to form the final optimized copywriting.
[0024] In a second aspect, an embodiment of the present application provides a macro-micro fusion copywriting language optimization system, and the system includes:
[0025] A field classification module for classifying the copywriting to be optimized to determine its industry field and theme label;
[0026] An inspector management module for selecting a set of structured inspectors corresponding to the industry field, judging the highlight coverage of the set of structured inspectors with respect to excellent texts in the same field as the copywriting to be optimized, and if the coverage is lower than a preset threshold, generating text inspection content based on a large model and updating the set of structured inspectors;
[0027] A macro-optimization framework generation module, configured to load the updated set of structured checkers according to the topic tags, and generate a multi-level optimization framework and an optimization thought chain based on the rule constraints of the set of structured checkers;
[0028] A macro-rewriting module, configured to globally restructure the text to be optimized by using a large model according to the multi-level optimization framework and the optimization thought chain, and generate a primary optimized text;
[0029] A sentence pattern knowledge base construction module, configured to disassemble the sentences of the excellent text into fixed templates and variable parameters by performing syntactic analysis and semantic annotation on the excellent text, so as to extract industry-specific sentence structures, and construct a reusable sentence pattern knowledge base including phrase level, sentence pattern level and paragraph level;
[0030] A sentence processing and matching module, configured to perform sentence segmentation on the primary optimized text to generate a sentence sequence with context markers, and calculate the semantic similarity between each sentence in the sentence sequence and the sentence pattern templates in the reusable sentence pattern knowledge base by using a vectorization technique based on pre-trained language model embeddings, and match the sentence pattern template with the highest similarity for each sentence;
[0031] A micro-polishing module, configured to perform fine-grained rewriting on the sentences in the sentence sequence by using a large model based on the matched sentence pattern templates and the sentence sequence with context markers, and generate a final optimized sentence.
[0032] In a third aspect, an embodiment of the present application provides an electronic device, where the electronic device includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0033] In a fourth aspect, an embodiment of the present application provides a readable storage medium, where a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] (1) Macro-micro integration and comprehensive optimization: The present invention innovatively combines macro-structural optimization and micro-language refinement. Through macro-rewriting based on a multi-level framework and thought chain of domain checkers, the logic, structure and content integrity of the text are ensured; through template matching and fine-grained rewriting based on the sentence pattern knowledge base, the professionalism, accuracy and expressiveness of the language are improved. This macro-micro integration strategy overcomes the limitation of the prior art that only focuses on a single level, and realizes the comprehensive and in-depth optimization of the text quality.
[0036] (2) Domain Adaptation with Strong Professionalism: By introducing domain classification, domain-specific checkers, and sentence pattern knowledge bases, the present invention achieves adaptive optimization of copywriting in different industry domains. The checkers and sentence pattern libraries can be dynamically updated, continuously learning and absorbing excellent expression paradigms and evaluation criteria in specific domains to ensure that the optimized copywriting highly conforms to domain requirements in terms of term usage, expression style, professional norms, etc., solving the problem of insufficient professionalism in existing general models.
[0037] (3) Intelligent and Efficient with High Degree of Automation: The present invention deeply integrates the capabilities of large models and achieves a high degree of automation in multiple links such as checker generation, macro rewriting, sentence pattern extraction, and micro polishing. Especially through CoT-guided macro rewriting and micro polishing based on template matching, the need for manual intervention and subjectivity is greatly reduced, improving the efficiency and consistency of the optimization process and making the production of high-quality copywriting faster and more standardized.
[0038] (4) Continuous Learning and Ability Evolution: By dynamically updating the checker set and continuously expanding the sentence pattern knowledge base (especially based on excellent texts provided by users), the methods and systems of the present invention have the ability of continuous learning and self-improvement, can continuously adapt to new domain requirements and language development trends, and maintain the leading edge of optimization effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of a macro-micro fusion copywriting language optimization method provided by an embodiment of the present invention;
[0040] Figure 2 is a schematic structural diagram of a macro-micro fusion copywriting language optimization system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.
[0042] The terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the description and claims means at least one of the connected objects. The character " / ", generally represents an "or" relationship between the associated objects before and after. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0043] Embodiment 1
[0044] Please refer to Figure 1 , which is a flowchart of a macro-micro fusion copywriting language optimization method proposed in the first embodiment of the present invention. This method can be executed by a computer program, for example, running on a server or a terminal device. The specific steps are as follows:
[0045] Step S1: Classify the copywriting to be optimized to determine its industry field and topic tags.
[0046] Specifically, receive the copywriting to be optimized uploaded or input by the user. Use a pre-trained text classification model (which can be a model fine-tuned based on BERT, RoBERTa, etc., or the classification ability of large language models such as the GPT series) to analyze the text. The model outputs the predefined industry field category to which the text belongs (such as scientific research, finance, medical, education, etc.). At the same time, keyword extraction algorithms (such as TF-IDF, TextRank) or the summary / topic generation ability of large models can be used to generate a set of topic tags reflecting the core content of the copywriting.
[0047] Exemplarily, the user inputs the copywriting to be optimized: "Artificial intelligence has many applications in the medical field. For example, doctors can use AI to assist in diagnosing diseases. AI technology can improve the accuracy of diagnosis." The system analyzes through a large model and determines that it belongs to the "medical" field, and generates topic tags such as "artificial intelligence", "medical diagnosis", "AI application". If the user also provides excellent text: "In recent years, the application of deep learning technology in medical image analysis has made remarkable progress. Through convolutional neural networks (CNNs), doctors can more accurately identify and diagnose diseases, such as cancer and cardiovascular diseases. The application of this technology not only improves the accuracy of diagnosis, but also greatly shortens the diagnosis time." The system will also classify it into the "medical" field and generate topic tags such as "deep learning", "medical imaging", "CNN", "disease diagnosis".
[0048] Step S2: Select the structured checker set corresponding to the industry field, and judge the highlight coverage of the structured checker set with respect to the excellent texts in the same field as the text to be optimized. If the coverage is lower than the preset threshold, generate text inspection content based on the large model and update the structured checker set.
[0049] Ⅰ. Selection: According to the "medical" field determined in Step S1, retrieve all the checkers marked as the "medical" field from the stored checker knowledge base to form an initial checker set.
[0050] Ⅱ. Coverage evaluation: Use the checkers in this initial set to analyze the excellent texts in the "medical" field provided by the user. Each checker attempts to identify the features it focuses on (such as research detail and frontier). Count the number of highlight features in the excellent texts that are successfully identified or evaluated as "meeting the standards" by at least one checker, and calculate the proportion of them in the total highlight features of the excellent texts (which can be manually marked or automatically discovered by comparing with ordinary texts) to obtain the coverage.
[0051] Ⅲ. Judgment and update: Compare the calculated coverage with the preset threshold (such as 80%).
[0052] If the coverage >= threshold (Case 1): It means the existing checkers are good enough and no update is needed. Directly use the currently selected checker set.
[0053] If the coverage < threshold (Case 2): It means the existing checkers fail to capture some key features of the excellent texts. At this time, by comparing the excellent texts with the ordinary texts (provided in Step S1), find the differences (i.e., the highlights of the excellent texts). For example, it is found that the excellent text clearly mentions the specific technology CNN and its application effects, while the ordinary text description is general. Based on this difference, construct a specific Prompt, such as: "Please generate a checker to evaluate whether the AI technology (such as algorithm name) used in the medical field texts is specifically described and its specific role and advantages in diagnosis. The checker should include function description, detection dimension, evaluation criteria, and improvement suggestions." Input this Prompt into the large model (such as GPT-4) to obtain a new checker to generate text inspection content, such as: "Interdisciplinary and innovative integration expression checker" (as shown in the example, or more precisely, "Specificity of AI technology application checker"). Add the generated text inspection content to the checker set of the "medical" field to complete the update.
[0054] As shown in the provided example, no update is needed in Case 1; in Case 2, a new "Interdisciplinary and innovative integration expression checker" is generated and added to the set.
[0055] Step S3: Load the updated set of structured checkers according to the subject tag, and generate a multi-level optimization framework and an optimization thought chain based on the rule constraints of the set of structured checkers.
[0056] I. Loading: According to the subject tag (such as "Application of Deep Learning in Medical Image Analysis") and the field (medical) determined in Step S1, further screen out the checkers most relevant to this subject from the set of checkers in the "medical" field obtained in Step S2 (or use all field checkers).
[0057] II. Generating a multi-level optimization framework: Traverse the loaded set of checkers. According to the "function description" and "detection dimension" of each checker, classify it into predefined optimization levels (structural level, semantic level, style level). For example, the "Research Content Detail Evaluation Checker" may mainly belong to the semantic level, and the "Academic Frontier Expansion Checker" may span the semantic level and the style level. Summarize the requirements of all checkers to form a structured optimization framework containing multi-level checkpoints.
[0058] III. Generating an optimization thought chain (CoT): Based on the multi-level optimization framework, transform the optimization process into a series of specific, logically sequenced operation instructions or thinking steps. This is usually done by a dedicated CoT generation module (which can be rule-based or generated by a large model). It will analyze the initial structure and content of the text to be optimized, and combine the checker rules to generate a guiding optimization path.
[0059] Exemplarily (continuing the example in Step S4): Loaded checkers for the "Path Planning Algorithm Test" subject (which may include general logic checkers and domain-specific checkers). Framework construction: Structural level (check the consistency of introduction - body - conclusion, paragraph logic), semantic level (check the unity of terms, the accuracy and completeness of the DRT-PP method description, the matching of experimental results and conclusions), style level (check the conciseness of expression, the compliance with the scientific research paper style). CoT generation: As shown in the provided example, it includes text analysis (keywords, core ideas) and text optimization steps (clarify the theme, simplify the logic, unify the terms, remove redundancy, supplement details, add summaries, etc.).
[0060] Step S4: Use a large model to globally restructure the text to be optimized according to the multi-level optimization framework and the optimization thought chain, and generate a primary optimized text.
[0061] Specifically, the text to be optimized, the multi-level optimization framework generated in step S3, and the optimization chain of thought (CoT) are provided as input to the big model (such as GPT-4). The big model is instructed through a carefully designed prompt: "Please structurally rewrite the original text provided according to the following optimization framework and chain of thought, focusing on optimizing the logical flow, information organization, and core content expression, and generate a version with a clearer structure and more rigorous logic." The big model modifies the original text one by one or as a whole according to the steps of CoT, and performs operations such as chapter reorganization, paragraph adjustment, content addition and deletion, and argument reinforcement. The output result is the primary optimized text.
[0062] As shown in the example provided, the original text is input and the CoT is generated, and the optimized sample text is output by the large model. Among them, the blue mark reflects the structural adjustment (such as focusing on the challenges in advance), the red mark reflects the content addition (such as supplementing the DRT-PP core mechanism and dynamic adjustment function), and the green mark reflects the vocabulary replacement (improving professionalism).
[0063] Step S5: By performing syntactic analysis and semantic annotation on the excellent text, the sentences of the excellent text are disassembled into fixed templates and variable parameters to extract industry-specific sentence structures and construct a reusable sentence pattern knowledge base including phrase level, sentence level and paragraph level.
[0064] Ⅰ. Analyze excellent texts: Select a collection of excellent texts in the same field as the text to be optimized. Use NLP tools (such as spaCy, Stanford CoreNLP, or the built-in capabilities of large models) to perform dependency syntactic analysis and named entity recognition / semantic role labeling on the sentences in each excellent text.
[0065] Ⅱ. Disassembly and extraction: Based on the analysis results, identify the core predicate structure, subject, object, adverbial and other components of the sentence and their relationships. Identify the fixed parts (such as specific verbs, prepositional phrases, conjunctions) and variable parts (usually noun phrases, concepts, specific data, etc., corresponding to semantic roles or entity types) in the sentence. Abstract the fixed parts into templates with placeholders (such as [something]), and use the variable parts as fill-in examples or parameter types.
[0066] III. Build a knowledge base: The extracted sentence pattern templates (including "sentence patterns" i.e. fixed templates, sample sentences, "excellent vocabulary" lists, etc.) are stored in a structured manner. Index by field, theme, sentence function (such as definition, comparison, explanation of reasons, description of results, etc.) to form a reusable sentence pattern knowledge base. This reusable sentence pattern knowledge base can contain patterns at the phrase level (such as "based on the method of..."), sentence level (such as the long sentence in the example), and even paragraph level (such as the standard summary structure, introduction paragraph template).
[0067] As shown in the provided example, analyze the excellent sentence "Based on the multipath transmission protocol...", extract the fixed template "Based on [so and so], break through the [so and so] paradigm, and design the [so and so] theory, mechanism, and mode.", and record the corresponding variable parameter filling content and excellent vocabulary. What is finally stored in the knowledge base is the structured template record. The processing object is the excellent text.
[0068] Step S6: Perform sentence segmentation on the primary optimized text to generate a sentence sequence with context markers.
[0069] Specifically, use a standard sentence segmentation tool (such as a rule-based or machine learning-based sentence breaker) to process the primary optimized text generated in step S4. Obtain a list (sequence) of sentences. Traverse this list. For the i-th sentence, record the content of its previous sentence (i - 1, if i > 0) and the next sentence (i + 1, if i is not the last sentence) as its context marker. This context marker is used to maintain the coherence between sentences in the refined rewriting in step S8.
[0070] As shown in the provided example, the primary optimized text is segmented into 4 sentences, and each sentence is marked with its "previous text" and "next text" (marked as "no previous text" or "no next text" if there is none).
[0071] Step S7: Adopt a vectorization technique based on pre-trained language model embeddings to calculate the semantic similarity between each sentence in the sentence sequence and the sentence templates in the reusable sentence pattern knowledge base, and match the sentence template with the highest similarity for each sentence.
[0072] Ⅰ. Vectorize the sentence: For each sentence in the sentence sequence generated in step S6, use the pre-trained Sentence-BERT model (or other models that can generate high-quality sentence embeddings) to encode it into a vector with a fixed dimension.
[0073] Ⅱ. Vectorize the template: From the sentence pattern knowledge base constructed in step S5, take out the sentence templates related to the current copywriting field. For each template, also encode the core "sentence pattern" text (fixed template part) into a vector using the same Sentence-BERT model.
[0074] Ⅲ. Calculate the similarity: For the vector of each sentence to be optimized, calculate the cosine similarity between it and all relevant sentence template vectors in the knowledge base.
[0075] Ⅳ. Match the best template: For each sentence to be optimized, select the sentence template with the highest cosine similarity to its vector as the best matching result.
[0076] As shown in the provided example, the sentence "The experimental results show..." is vectorized by BERT. Then, the cosine similarity between this vector and the vectors of all medical domain templates in the knowledge base (such as "To verify the effectiveness of [so-and-so]...") is calculated. If the similarity with the "To verify..." template is found to be the highest (0.98), then this template is associated with the sentence.
[0077] Step S8: Based on the matched sentence pattern template and the sentence sequence with context markers, use a large model to refine and rewrite the sentences in the sentence sequence to generate the final optimized sentences.
[0078] Ⅰ. Generate a polishing control signal: For each sentence in the sentence sequence, its best-matched sentence pattern template in step S7, the context information recorded in step S6, as well as the example sentences and excellent words obtained from the sentence pattern template, integrate this information to form a polishing instruction (Prompt) for the large model. For example, the instruction may include: "Please use the structure of the following sentence pattern template (template structure), refer to the expression of the template example (template example), and preferably use these excellent words (excellent words), combined with the context (above text, below text), to polish and rewrite this original sentence (original sentence content) to make it more professional, concise, and better expressed, while maintaining the original meaning."
[0079] Ⅱ. Refined rewriting by the large model: Input the original sentence, the matched template, the context, the examples, the words, and the above polishing instruction Prompt into the large model. The large model makes detailed modifications to the original sentence according to these constraints, which may involve sentence pattern adjustment, word replacement, expression optimization, etc. The processing object is each sentence in the sentence sequence generated in step S6.
[0080] Ⅲ. Combine the final text: Recombine all the sentences in the sentence sequence that have been refined and rewritten by the large model in the original order to form the final optimized copywriting.
[0081] As shown in the provided example:
[0082] Original sentence 1: "The experimental results show...", combined with the matched "To verify..." template, template example, excellent words (effectiveness, significant, diversity) and context, generate a Prompt. The large model outputs the polished text: "To verify the effectiveness of the DRT-PP method..." (which reflects the learning of template structure, example expression, and vocabulary application).
[0083] Original sentence 2: "Regarding the above problems...", combined with the matched "To balance..." template, template example, excellent words (innovatively, efficiently, contradiction) and context, generate a Prompt. The large model outputs the polished text: "Regarding the existing problems in the current field of autonomous driving..." (which also reflects the integrated application of various aspects of information).
[0084] Example 2
[0085] Please refer to Figure 2 , which shows a schematic structural diagram of a macro-micro fusion copywriting language optimization system proposed in the second embodiment of this application. This system can be deployed on a server and provide services to users through an API or a web interface. The system includes the following key modules:
[0086] Domain classification module 100, which is used to classify the copywriting to be optimized, and determine its industry domain and theme tags;
[0087] Inspector management module 200, which is used to select a set of structured inspectors corresponding to the industry domain, judge the highlight coverage of the set of structured inspectors relative to excellent texts in the same domain as the copywriting to be optimized. If the coverage is lower than a preset threshold, text inspection content is generated based on a large model and the set of structured inspectors is updated;
[0088] Macro optimization framework generation module 300, which is used to load the updated set of structured inspectors according to the theme tags, and generate a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the set of structured inspectors;
[0089] Macro rewriting module 400, which is used to globally restructure the copywriting to be optimized by using a large model according to the multi-level optimization framework and the optimization thinking chain, and generate a primary optimized text;
[0090] Sentence pattern knowledge base construction module 500, which is used to disassemble the sentences of the excellent texts into fixed templates and variable parameters by performing syntactic analysis and semantic annotation on the excellent texts, so as to extract industry domain-specific sentence patterns, and construct a reusable sentence pattern knowledge base including phrase level, sentence pattern level and paragraph level;
[0091] Sentence processing and matching module 600, which is used to segment the primary optimized text into sentences to generate a sentence sequence with context markers, and calculate the semantic similarity between each sentence in the sentence sequence and the sentence pattern templates in the reusable sentence pattern knowledge base by using a vectorization technology based on pre-trained language model embeddings, and match the sentence pattern template with the highest similarity for each sentence;
[0092] Micro refinement module 700, which is used to finely rewrite the sentences in the sentence sequence by using a large model based on the matched sentence pattern templates and the sentence sequence with context markers, and generate the final optimized sentences.
[0093] Example of the system working process:
[0094] The user submits the text to be optimized and optional reference text through the user input module -> The domain classification module determines the domain and theme -> Macro modification module: Obtain the corresponding domain checker from the storage and management module -> (If necessary) Call the large model algorithm engine to update the checker -> Generate an optimization framework and CoT -> Call the large model algorithm engine for macro rewriting to obtain a primary optimized text -> Micro polishing module: Perform sentence segmentation and context marking -> Vectorize the sentences and obtain sentence pattern templates from the storage and management module for matching -> Call the large model algorithm engine to perform refined rewriting in combination with templates, context, etc. -> Output the final optimized text to the user.
[0095] A macro-micro fusion copywriting language optimization system in an embodiment of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), etc. The embodiments of the present application do not make specific limitations.
[0096] A macro-micro fusion copywriting language optimization system in an embodiment of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.
[0097] A macro-micro fusion copywriting language optimization system provided in an embodiment of the present application can implement Figure 1 each process implemented in a method embodiment of a macro-micro fusion copywriting language optimization method. To avoid repetition, it will not be elaborated here.
[0098] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a macro-micro fusion copywriting language optimization method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0099] An embodiment of the present application further provides a readable storage medium, on which a program or instructions are stored. When the program or instructions are executed by a processor, each process of the above embodiment of the method for optimizing copywriting language with macro-micro fusion is implemented, and the same technical effects can be achieved. To avoid repetition, it will not be elaborated here.
[0100] Wherein, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.
[0101] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0102] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0103] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A macro-micro integration copywriting language optimization method, characterized in that: The method comprises the following steps: S1: Classify the texts to be optimized and determine their industry and topic tags; S2: Select a set of structured checkers corresponding to the industry field, determine the highlight coverage of the structured checker set relative to the excellent text in the same field of the copy to be optimized, and if the coverage is lower than a preset threshold, generate text inspection content based on the big model and update the structured checker set, specifically including: input the excellent text into the structured checker set, count the number of highlight features in the excellent text recognized by the structured checker set, and calculate the ratio of the number to the total number of highlight features in the excellent text as the coverage; if the coverage is lower than the preset threshold, obtain a new checker based on the big model by comparing the difference between the excellent text and the copy to be optimized, so as to generate text inspection content, and the new checker includes function description, detection dimension, evaluation standard and improvement suggestion; S3: loading the updated structured checker set according to the topic tag, and generating a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the structured checker set; S4: using the big model to rewrite the global structure of the text to be optimized according to the multi-level optimization framework and the optimization thinking chain, and generate a primary optimized text; S5: By performing syntactic analysis and semantic annotation on excellent texts, the sentences of the excellent texts are disassembled into fixed templates and variable parameters to extract industry-specific sentence structures and build a reusable sentence pattern knowledge base including phrase level, sentence level and paragraph level; S6: Segment the primary optimized text into sentences to generate a sentence sequence with context tags; S7: using a vectorization technology based on pre-trained language model embedding, calculating the semantic similarity between each sentence in the sentence sequence and the sentence template in the reusable sentence pattern knowledge base, and matching each sentence with the sentence template with the highest similarity; S8: Based on the matched sentence template and the sentence sequence with context tags, the sentences in the sentence sequence are refined and rewritten using the large model to generate a final optimized sentence.
2. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: In step S3: The multi-level optimization framework includes a structure layer, a semantic layer, and a style layer, wherein: The structural layer is used to optimize the logical flow and information organization of the copy. The semantic layer is used to ensure terminology accuracy and content integrity. The style layer is used to adjust the expression style to comply with industry standards; The optimization thinking chain includes a text analysis step and a text optimization step, wherein: The text analysis step includes identifying the keywords and core ideas of the text to be optimized, The text optimization steps include structural adjustment, content supplementation and terminology unification.
3. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: The step S5 includes: performing syntactic analysis and semantic annotation on the excellent text. Use NLP tools to perform dependency parsing and named entity recognition to identify fixed templates and variable parameters in sentences; Abstracting the fixed template into a template structure containing placeholders, and recording a filling example of the variable parameters; The reusable sentence pattern knowledge base is indexed by domain, theme and sentence function, and contains templates at phrase level, sentence level and paragraph level.
4. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: Generating a sentence sequence with context tags in step S6 includes: Using a sentence segmentation tool, segment the primary optimized text into sentences to obtain a list of sentence sequences; Traverse the list, and for the i-th sentence, record the contents of its previous sentence and next sentence as context tags.
5. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: The vectorization technology based on pre-trained language model embedding adopted in step S7 includes: Convert each sentence in the sentence sequence and the sentence pattern template in the sentence pattern knowledge base into a vector representation of a fixed dimension respectively; The cosine similarity between the sentence vector and the template vector is calculated, and the template with the highest similarity is selected as the matching result.
6. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: The step S8 comprises: For each sentence in the sentence sequence and its matching best sentence template, as well as the recorded context information and sample sentences and excellent vocabulary obtained from the sentence template, form polishing instructions for the large model; Inputting the original sentence, matching template, context, examples, vocabulary and polishing instructions into the big model, and rewriting the original sentence in a refined manner, wherein the rewriting includes sentence structure adjustment, word replacement and expression optimization; All the sentences in the sentence sequence that have been refined and rewritten by the large model are recombined in order to form the final optimized copy.
7. A macro-micro integration copywriting language optimization system, characterized in that: The system comprises: The field classification module is used to classify the optimized copywriting into fields and determine its industry field and subject tags; The checker management module is used to select a set of structured checkers corresponding to the industry field, determine the highlight coverage of the structured checker set relative to the excellent text in the same field of the copy to be optimized, and if the coverage is lower than a preset threshold, generate text inspection content based on the big model and update the structured checker set, specifically including: inputting the excellent text into the structured checker set, counting the number of highlight features in the excellent text recognized by the structured checker set, and calculating the ratio of the number to the total number of highlight features in the excellent text as the coverage; if the coverage is lower than the preset threshold, by comparing the difference between the excellent text and the copy to be optimized, obtain a new checker based on the big model to generate text inspection content, the new checker includes function description, detection dimension, evaluation standard and improvement suggestion; A macro optimization framework generation module, used for loading the updated structured checker set according to the topic tag, and generating a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the structured checker set; A macro rewriting module is used to use a macro model to rewrite the global structure of the text to be optimized according to the multi-level optimization framework and the optimization thinking chain to generate a primary optimized text; A sentence pattern knowledge base construction module is used to perform syntactic analysis and semantic annotation on excellent texts, decompose the sentences of the excellent texts into fixed templates and variable parameters, extract industry-specific sentence structures, and construct a reusable sentence pattern knowledge base including phrase level, sentence level and paragraph level; A sentence processing and matching module, which is used to segment the primary optimized text into sentences, generate a sentence sequence with context tags, use a vectorization technology based on pre-trained language model embedding, calculate the semantic similarity between each sentence in the sentence sequence and the sentence template in the reusable sentence pattern knowledge base, and match each sentence with the sentence template with the highest similarity; The micro-polishing module is used to perform fine rewriting on the sentences in the sentence sequence based on the matched sentence template and the sentence sequence with context markers using a large model to generate a final optimized sentence.
8. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a macro-micro fusion copywriting language optimization method as described in any one of claims 1 to 6 are implemented.
9. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the macro-micro fusion copywriting language optimization method as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method and system for generating bionic hierarchical memory fusion document of large electric semantic model
CN119474347A