Macro-micro fused copywriting language optimization method and system

Through the copywriting language optimization method of macro-micro-fusion, combined with field classification, structured inspector and sentence pattern knowledge base, the problem of insufficient copywriting optimization in the existing technology is solved, and the structural rationality, professionalism and language accuracy of copywriting is improved, meeting the industry needs of high-standard copywriting.

CN120046587AActive Publication Date: 2025-05-27JIANGXI NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510518888.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing AI-based copywriting optimization methods have shortcomings in macro structure, field professionalism and multi-dimensional integration optimization, resulting in the optimized copywriting that may lack logic, professionalism and overall quality improvement.

Method used

The copywriting language optimization method of macro-micro-fusion is adopted, and through the combination of intelligent macro-optimization and micro-modification mechanism, the field classification, structured inspector, multi-level optimization framework and sentence pattern knowledge base are used to realize the global structure and language rewriting of copywriting.

Benefits of technology

It significantly improves the structural rationality, logical rigor, language accuracy, professionalism and expressiveness of copywriting, meets the needs of high-standard copywriting in different industries and achieves a significant improvement in the efficiency and quality of copywriting creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046587A_ABST
    Figure CN120046587A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of natural language processing, and discloses a macro-micro fusion copywriting language optimization method and system, and the method comprises the steps: determining an industry field and a theme tag of a to-be-optimized copywriting; selecting a corresponding structured checker set, and if it is detected that the bright spot feature coverage of the same-field excellent texts is lower than a preset threshold value, generating a text check content set by using a large model and updating the text check content set; loading the updated novel checker set according to the topic tag, and generating a multi-level optimization framework and a thinking chain to perform global structure rewriting to obtain a primary optimization text; syntactic analysis and semantic annotation are carried out on excellent texts in the same field, and a specific sentence pattern structure in the industry field is extracted and stored in a reusable sentence pattern knowledge base; segmenting the primary optimization text into a sentence sequence with a context, and matching the sentence sequence with the highest similarity template in the reusable sentence pattern knowledge base; and on the basis of refined rewriting of the template and the sentence sequence, the final optimized sentence is generated, and the document quality and creation efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and natural language processing, and particularly relates to a method and system for optimizing copywriting language with macro-micro integration. Background Art

[0002] With the continuous development of artificial intelligence (AI) and natural language processing (NLP) technologies, the intelligence and automation of copywriting creation have gradually become a hot topic in the industry. Modern enterprises and institutions, especially in professional fields such as technology, finance, healthcare, and education, have higher demands for high-quality and high-efficiency copywriting creation. In these fields, the structured expression, logical rigor, and language precision of copywriting directly affect the effective dissemination of information, the establishment of a professional image, and the trust of the audience. To improve the quality of copywriting, many organizations have begun to attempt to use AI technology to assist in copywriting creation, including using large-scale pre-trained language models (such as GPT, BERT, etc.) to generate and optimize copywriting content.

[0003] However, the existing AI-based copywriting optimization methods and systems still have several deficiencies in practical applications, mainly reflected in the following aspects: (1) Insufficient optimization of the macro structure: Currently, many copywriting optimization technologies mainly focus on the polishing and improvement of the language level, such as grammar correction, vocabulary replacement, and improvement of fluency. However, they often neglect the optimization of the overall structure and core content of the copywriting. Although large models can improve language expression to a certain extent, at the macro level, in terms of the logic, paragraph organization, information hierarchy, support of arguments and evidence, clarity of information transmission, etc. of the copywriting, the existing technologies still cannot meet the high standards of the industry, resulting in the optimized copywriting may be "similar in form" but "scattered in spirit".

[0004] (2) Lack of domain professionalism: Most of the existing copywriting generation and optimization tools are trained based on general language models and lack in-depth tuning for specific industries and the injection of professional knowledge. This makes the generated copywriting significantly lack in professionalism, accuracy of term usage, industry-conventional expressions, and compliance with specific norms. Especially for fields such as technology, finance, and healthcare, which have extremely high requirements for the quality of copywriting, the copywriting generated by general models is difficult to meet their rigor and professionalism. In addition, the term processing ability of the existing systems is limited, and it is often unable to precisely control the unique expression styles and nuances of the industry, resulting in the generated copywriting being difficult to meet the industry-recognized standards.

[0005] (3) Single optimization dimension and lack of integration: Current copywriting optimization methods are often limited to single-level improvements, such as grammar checking or synonym replacement, and fail to achieve comprehensive and systematic optimization of multiple dimensions such as copywriting structure, content, logic, language, and style. The improvement of copywriting quality is the result of the synergy of multiple factors. Existing methods usually lack an integrated framework and cannot conduct unified and coordinated evaluation and optimization in different dimensions, resulting in scattered optimization effects and difficulty in achieving a breakthrough improvement in the overall quality of copywriting.

[0006] Therefore, there is an urgent need for a copywriting optimization method and system that can integrate macro-structure planning and micro-language refinement, take into account both field expertise and expression accuracy, and achieve multi-dimensional, systematic, and intelligent copywriting optimization to overcome the limitations of existing technologies. Summary of the invention

[0007] The main purpose of the present invention is to solve the deficiencies of the existing copywriting optimization methods in the macro structure, field expertise and multi-dimensional integrated optimization, and to provide a macro-micro integrated copywriting language optimization method and system. The present invention aims to systematically improve the structural rationality, logical rigor, language accuracy, professionalism and expressiveness of copywriting creation through the combination of intelligent macro optimization and micro polishing mechanism, thereby greatly improving the efficiency and quality of copywriting creation and meeting the urgent needs of different industries for high-standard copywriting.

[0008] In order to achieve the above-mentioned object, the first aspect of the present invention provides a macro-micro fusion copywriting language optimization method, the method comprising the following steps: S1: Classify the texts to be optimized and determine their industry and topic tags; S2: Select a set of structured checkers corresponding to the industry field, determine the highlight coverage of the structured checker set relative to the excellent texts in the same field of the copy to be optimized, and if the coverage is lower than a preset threshold, generate text inspection content based on the big model and update the structured checker set; S3: loading the updated structured checker set according to the topic tag, and generating a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the structured checker set; S4: using the big model to rewrite the global structure of the text to be optimized according to the multi-level optimization framework and the optimization thinking chain, and generate a primary optimized text; S5: By performing syntactic analysis and semantic annotation on excellent texts, the sentences of the excellent texts are disassembled into fixed templates and variable parameters to extract industry-specific sentence structures and build a reusable sentence pattern knowledge base including phrase level, sentence level and paragraph level; S6: Perform sentence segmentation on the primary optimized text to generate a sentence sequence with context markers; S7: Use a vectorization technique based on pre-trained language model embeddings to calculate the semantic similarity between each sentence in the sentence sequence and the sentence patterns in the reusable sentence pattern knowledge base, and match the sentence pattern template with the highest similarity for each sentence; S8: Based on the matched sentence pattern template and the sentence sequence with context markers, use a large model to perform refined rewriting on the sentences in the sentence sequence to generate the final optimized sentences.

[0009] As an optional implementation manner of the first aspect of this application, the determination of the highlight coverage of the structured checker set relative to the excellent text in the same field as the text to be optimized in step S2 includes: inputting the excellent text into the structured checker set, counting the number of highlight features recognized by the structured checker set in the excellent text, and calculating the ratio of the number to the total number of highlight features in the excellent text as the coverage; if the coverage is lower than the preset threshold, then by comparing the differences between the excellent text and the text to be optimized, a new checker is obtained based on the large model to generate text inspection content, and the new checker includes function descriptions, detection dimensions, evaluation criteria, and improvement suggestions.

[0010] As an optional implementation manner of the first aspect of this application, in step S3: the multi-level optimization framework includes a structure layer, a semantic layer, and a style layer, where: the structure layer is used to optimize the logical flow and information organization of the text, the semantic layer is used to ensure the accuracy of terms and the integrity of content, and the style layer is used to adjust the expression style to conform to industry norms; the optimization thinking chain includes a text analysis step and a text optimization step, where: the text analysis step includes identifying the keywords and core ideas of the text to be optimized, and the text optimization step includes structure adjustment, content supplementation, and term unification.

[0011] As an optional implementation manner of the first aspect of this application, in step S5, the syntactic analysis and semantic annotation of the excellent text include: using NLP tools to perform dependency syntactic analysis and named entity recognition to identify fixed templates and variable parameters in the sentence; abstracting the fixed template into a template structure containing placeholders and recording the filling examples of the variable parameters; the reusable sentence pattern knowledge base is indexed according to fields, topics, and sentence pattern functions, and includes templates at the phrase level, sentence pattern level, and paragraph level.

[0012] As an alternative implementation of the first aspect of the present application, generating the sentence sequence with context tags in step S6 includes: using a sentence segmentation tool to segment the primary optimized text to obtain a list of sentence sequences; traversing the list, and for the i-th sentence, recording the content of its previous sentence and the next sentence as context tags.

[0013] As an alternative implementation of the first aspect of the present application, the vectorization technique based on pre-trained language model embedding in step S7 includes: converting each sentence in the sentence sequence and the sentence patterns in the sentence pattern knowledge base into vector representations of a fixed dimension respectively; calculating the cosine similarity between the vector of the sentence and the template vector, and selecting the template with the highest similarity as the matching result.

[0014] As an alternative implementation of the first aspect of the present application, step S8 includes: for each sentence in the sentence sequence and its best-matched sentence pattern template, as well as the recorded context information and the example sentences and excellent words obtained from the sentence pattern template, forming a polishing instruction for the large model; inputting the original sentence, the matching template, the context, the examples, the words, and the polishing instruction into the large model to perform refined rewriting on the original sentence, and the rewriting includes sentence pattern adjustment, word replacement, and expression optimization; recombining all the sentences in the sentence sequence that have been refined and rewritten by the large model in order to form the final optimized text.

[0015] In a second aspect, an embodiment of the present application provides a macro-micro integrated copywriting language optimization system, and the system includes: A field classification module, which is used to classify the copywriting to be optimized, and determine its industry field and theme tags; An inspector management module, which is used to select a set of structured inspectors corresponding to the industry field, and judge the highlight coverage of the set of structured inspectors relative to the excellent texts in the same field as the copywriting to be optimized. If the coverage is lower than a preset threshold, text inspection content is generated based on the large model and the set of structured inspectors is updated; A macro-optimization framework generation module, which is used to load the updated set of structured inspectors according to the theme tags, and generate a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the set of structured inspectors; A macro-rewriting module, which is used to globally restructure the copywriting to be optimized by using the large model according to the multi-level optimization framework and the optimization thinking chain to generate a primary optimized text; The sentence pattern knowledge base construction module is used to disassemble the sentences of the excellent text into fixed templates and variable parameters by performing syntactic analysis and semantic annotation on the excellent text, so as to extract specific sentence structure of the industry field and construct a reusable sentence pattern knowledge base including phrase level, sentence pattern level and paragraph level; The sentence processing and matching module is used to perform sentence segmentation on the primary optimized text to generate a sentence sequence with context markers, and adopt a vectorization technology based on pre-trained language model embedding to calculate the semantic similarity between each sentence in the sentence sequence and the sentence pattern template in the reusable sentence pattern knowledge base, and match the sentence pattern template with the highest similarity for each sentence; The micro-polishing module is used to perform refined rewriting on the sentences in the sentence sequence by using a large model based on the matched sentence pattern template and the sentence sequence with context markers to generate the final optimized sentences.

[0016] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0017] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0018] Compared with the prior art, the present invention has the following beneficial effects: (1) Macro-micro integration and comprehensive optimization: The present invention innovatively combines macro-structural optimization and micro-language refinement. Through macro-rewriting based on a multi-level framework and chain of thought of the domain checker, the logic, structure and content integrity of the copywriting are ensured; through template matching and refined rewriting based on the sentence pattern knowledge base, the professionalism, accuracy and expressiveness of the language are improved. This macro-micro integration strategy overcomes the limitation of the prior art that only focuses on a single level and realizes the comprehensive and in-depth optimization of the copywriting quality.

[0019] (2) Domain adaptability and strong professionalism: The present invention realizes the adaptive optimization of copywriting in different industry fields by introducing domain classification, domain-specific checkers and sentence pattern knowledge bases. The checker and sentence pattern library can be dynamically updated, continuously learning and absorbing excellent expression paradigms and evaluation criteria in specific fields, ensuring that the optimized copywriting highly meets the domain requirements in terms of term usage, expression style, professional norms, etc., and solving the problem of insufficient professionalism of existing general models.

[0020] (3) Intelligent and efficient, with a high degree of automation: The present invention deeply integrates the capabilities of large models and achieves a high degree of automation in multiple links such as checker generation, macro rewriting, sentence pattern extraction, and micro polishing. In particular, through CoT-guided macro rewriting and template matching-based micro polishing, the need for manual intervention and subjectivity are greatly reduced, the efficiency and consistency of the optimization process are improved, and the production of high-quality copywriting is faster and more standardized.

[0021] (4) Continuous learning and capability evolution: By dynamically updating the checker set and continuously expanding the sentence knowledge base (especially based on excellent texts provided by users), the method and system of the present invention have the ability of continuous learning and self-improvement, can constantly adapt to new field requirements and language development trends, and maintain the leading optimization effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flow chart of a macro-micro fusion copywriting language optimization method provided by an embodiment of the present invention; Figure 2 It is a structural schematic diagram of a macro-micro integrated copywriting language optimization system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the front and back associated objects are a kind of "or" relationship. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically limited.

[0025] Example 1 See also Figure 1 , is a flow chart of a macro-micro fusion copywriting language optimization method proposed in the first embodiment of the present invention. The method can be executed by a computer program, for example, running on a server or terminal device. The specific steps are as follows: Step S1: Classify the text to be optimized into different fields and determine its industry field and topic tags.

[0026] Specifically, receive the copywriting to be optimized uploaded or input by the user. Analyze the text using a pre-trained text classification model (which can be a model fine-tuned based on BERT, RoBERTa, etc., or the classification ability of large language models such as the GPT series). The model outputs the predefined industry domain category to which the text belongs (such as scientific research, finance, healthcare, education, etc.). At the same time, a keyword extraction algorithm (such as TF-IDF, TextRank) or the summary / topic generation ability of a large model can be used to generate a set of topic tags reflecting the core content of the copywriting.

[0027] Exemplarily, the user inputs the copywriting to be optimized: "Artificial intelligence has many applications in the healthcare field. For example, doctors can use AI to assist in diagnosing diseases. AI technology can improve the accuracy of diagnosis." The system analyzes through a large model and determines that it belongs to the "healthcare" field, and generates topic tags such as "artificial intelligence", "medical diagnosis", and "AI applications". If the user also provides excellent text: "In recent years, the application of deep learning technology in medical image analysis has made remarkable progress. Through convolutional neural networks (CNNs), doctors can more accurately identify and diagnose diseases such as cancer and cardiovascular diseases. The application of this technology not only improves the accuracy of diagnosis but also greatly shortens the diagnosis time." The system will also classify it into the "healthcare" field and generate topic tags such as "deep learning", "medical imaging", "CNN", and "disease diagnosis".

[0028] Step S2: Select the set of structured checkers corresponding to the industry domain, and judge the highlight coverage of the set of structured checkers with respect to the excellent text in the same domain as the copywriting to be optimized. If the coverage is lower than the preset threshold, generate text inspection content based on a large model and update the set of structured checkers.

[0029] Ⅰ. Selection: According to the "healthcare" field determined in step S1, retrieve all the checkers marked as the "healthcare" field from the stored checker knowledge base to form an initial set of checkers.

[0030] Ⅱ. Coverage evaluation: Use the checkers in this initial set to analyze the excellent text in the "healthcare" field provided by the user. Each checker attempts to identify the features it focuses on (such as research detail, frontier). Count the number of highlight features in the excellent text that are successfully identified or evaluated as "meeting the standard" by at least one checker, and calculate the proportion of it to the total highlight features of the excellent text (which can be manually marked or automatically discovered by comparing with ordinary text) to obtain the coverage.

[0031] Ⅲ. Judgment and update: Compare the calculated coverage with the preset threshold (such as 80%).

[0032] If the coverage >= threshold (Case 1): It indicates that the existing checker is good enough and does not need to be updated. Directly use the currently selected checker set.

[0033] If the coverage < threshold (Case 2): It means that the existing checker fails to capture some key features of excellent texts. At this time, by comparing excellent texts with ordinary texts (provided in Step S1), find the differences (i.e., the highlights of excellent texts). For example, it is found that the excellent text clearly mentions the specific technology CNN and its application effects, while the description in the ordinary text is general. Based on this difference, construct a specific Prompt, such as: "Please generate a checker to evaluate whether the AI technology (such as algorithm name) used in medical copywriting is specifically described and its specific role and advantages in diagnosis. The checker should include function description, detection dimensions, evaluation criteria, and improvement suggestions." Input this Prompt into a large model (such as GPT-4) to obtain a new checker to generate text inspection content, such as: "Interdisciplinary and Innovative Integration Expression Checker" (as shown in the example, or more precisely, "Specificity of AI Technology Application Checker"). Add the generated text inspection content to the checker set in the "medical" field to complete the update.

[0034] As shown in the provided example, Case 1 does not require an update; in Case 2, a new "Interdisciplinary and Innovative Integration Expression Checker" is generated and added to the set.

[0035] Step S3: Load the updated structured checker set according to the theme label, and generate a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the structured checker set.

[0036] Ⅰ. Loading: According to the theme label (such as "Application of Deep Learning in Medical Image Analysis") and field (medical) determined in Step S1, further screen out the checkers most relevant to this theme from the (possibly updated) "medical" field checker set obtained in Step S2 (or use all field checkers).

[0037] Ⅱ. Generating a multi-level optimization framework: Traverse the loaded checker set. According to the "function description" and "detection dimensions" of each checker, classify it into predefined optimization levels (structural layer, semantic layer, style layer). For example, the "Research Content Detail Evaluation Checker" may mainly belong to the semantic layer, and the "Academic Frontier Expansion Checker" may span the semantic layer and the style layer. Summarize the requirements of all checkers to form a structured optimization framework containing multi-level checkpoints.

[0038] Ⅲ. Generate Optimized Chain of Thought (CoT): Based on a multi-level optimization framework, transform the optimization process into a series of specific and logically sequenced operation instructions or thinking steps. This is usually accomplished by a dedicated CoT generation module (which can be rule-based or generated by a large model). It analyzes the initial structure and content of the text to be optimized, and combines with the checker rules to generate a guiding optimization path.

[0039] Exemplarily (continuing the example of step S4): Loaded a checker for the topic of "Path Planning Algorithm Test" (which may include a general logic checker and a domain-specific checker). Framework construction: Structural layer (check the consistency of introduction-body-conclusion, paragraph logic), semantic layer (check the unity of terms, the accuracy and completeness of DRT-PP method description, the matching of experimental results and conclusions), style layer (check the conciseness of expression, the compliance with the scientific research paper style). CoT generation: As shown in the provided example, it includes text analysis (keywords, core ideas) and text optimization steps (clarify the theme, simplify the logic, unify the terms, remove redundancy, supplement details, add summaries, etc.).

[0040] Step S4: Use a large model to globally restructure the text to be optimized according to the multi-level optimization framework and the optimized Chain of Thought, and generate a primary optimized text.

[0041] Specifically, take the text to be optimized, the multi-level optimization framework generated in step S3, and the optimized Chain of Thought (CoT) as inputs, and provide them to a large model (such as GPT-4). Through a carefully designed Prompt, instruct the large model: "Please structurally rewrite the provided original text according to the following optimization framework and chain of thought, focusing on optimizing the logical process, information organization, and expression of the core content, and generate a version with a clearer structure and more rigorous logic." The large model modifies the original text one by one or as a whole according to the steps of the CoT, performing operations such as chapter reorganization, paragraph adjustment, content addition and deletion, and argument strengthening. The output result is the primary optimized text.

[0042] As shown in the provided example, for the input original text and the generated CoT, the optimized example text output by the large model. Among them, the blue markings reflect the structural adjustment (such as presenting the challenges in advance and concentrating the elaboration), the red markings reflect the content supplementation (such as supplementing the core mechanism of DRT-PP and the role of dynamic adjustment), and the green markings reflect the vocabulary replacement (enhancing professionalism).

[0043] Step S5: By performing syntactic analysis and semantic annotation on excellent texts, disassemble the sentences of the excellent texts into fixed templates and variable parameters to extract industry domain-specific sentence structures, and construct a reusable sentence pattern knowledge base including phrase level, sentence pattern level, and paragraph level.

[0044] Ⅰ. Analyze excellent texts: Select a collection of excellent texts in the same field as the text to be optimized. Use NLP tools (such as spaCy, Stanford CoreNLP, or the built-in capabilities of large models) to perform dependency syntactic analysis and named entity recognition / semantic role labeling on the sentences in each excellent text.

[0045] Ⅱ. Disassembly and extraction: Based on the analysis results, identify the core predicate structure, subject, object, adverbial and other components of the sentence and their relationships. Identify the fixed parts (such as specific verbs, prepositional phrases, conjunctions) and variable parts (usually noun phrases, concepts, specific data, etc., corresponding to semantic roles or entity types) in the sentence. Abstract the fixed parts into templates with placeholders (such as [something]), and use the variable parts as fill-in examples or parameter types.

[0046] III. Build a knowledge base: The extracted sentence pattern templates (including "sentence patterns" i.e. fixed templates, sample sentences, "excellent vocabulary" lists, etc.) are stored in a structured manner. Index by field, theme, sentence function (such as definition, comparison, explanation of reasons, description of results, etc.) to form a reusable sentence pattern knowledge base. This reusable sentence pattern knowledge base can contain patterns at the phrase level (such as "based on the method of..."), sentence level (such as the long sentence in the example), and even paragraph level (such as the standard summary structure, introduction paragraph template).

[0047] For example, the excellent sentence "Based on the multi-path transmission protocol..." is analyzed, and the fixed template "Based on [something], breaking through the [something] paradigm, designing [something] theory, mechanism and model" is extracted, and the corresponding variable parameter filling content and excellent vocabulary are recorded. What is finally stored in the knowledge base is the structured template record. The processing object is the excellent text.

[0048] Step S6: Segment the primary optimized text into sentences to generate a sentence sequence with context tags.

[0049] Specifically, a standard sentence segmentation tool (such as a rule-based or machine learning sentence segmenter) is used to process the primary optimized text generated in step S4. A list (sequence) of sentences is obtained. This list is traversed, and for the i-th sentence, the content of its previous sentence (i-1, if i>0) and the next sentence (i+1, if i is not the last sentence) is recorded as its context tag, which is used to maintain the coherence between sentences in the refined rewriting of step S8.

[0050] As shown in the example provided, the primary optimized text is divided into 4 sentences, and each sentence is marked with its "previous context" and "next context" (if there is no previous context, it is marked as "no previous context" or "no subsequent context").

[0051] Step S7: Using the vectorization technique based on pre-trained language model embedding, calculate the semantic similarity between each sentence in the sentence sequence and the sentence pattern templates in the reusable sentence pattern knowledge base, and match the sentence pattern template with the highest similarity for each sentence.

[0052] Ⅰ. Vectorize the sentence: For each sentence in the sentence sequence generated in Step S6, use the pre-trained Sentence-BERT model (or other models that can generate high-quality sentence embeddings) to encode it into a vector with a fixed dimension.

[0053] Ⅱ. Vectorize the template: From the sentence pattern knowledge base constructed in Step S5, extract the sentence pattern templates related to the current copywriting field. For each template, also encode the core "sentence pattern" text (fixed template part) into a vector using the same Sentence-BERT model.

[0054] Ⅲ. Calculate the similarity: For the vector of each sentence to be optimized, calculate the cosine similarity between it and the vectors of all relevant sentence pattern templates in the knowledge base.

[0055] Ⅳ. Match the best template: For each sentence to be optimized, select the sentence pattern template with the highest cosine similarity to its vector as the best matching result.

[0056] As shown in the provided example, the sentence "The experimental results show..." is vectorized by BERT. Then, calculate the cosine similarity between this vector and the vectors of all medical field templates in the knowledge base (such as "To verify the effectiveness of [so-and-so]..."). It is found that the similarity with the "To verify..." template is the highest (0.98), so this template is associated with this sentence.

[0057] Step S8: Based on the matched sentence pattern template and the sentence sequence with context markers, use a large model to refine and rewrite the sentences in the sentence sequence to generate the final optimized sentences.

[0058] Ⅰ. Generate the polishing control signal: For each sentence in the sentence sequence, its best-matched sentence pattern template in Step S7, the context information recorded in Step S6, as well as the example sentences and excellent words obtained from the sentence pattern template, integrate this information to form a polishing instruction (Prompt) for the large model. For example, the instruction may include: "Please use the structure of the following sentence pattern template (template structure), refer to the expression method of the template example (template example), and give priority to using these excellent words (excellent words), combined with the context (above text, below text), to polish and rewrite this original sentence (original sentence content) to make it more professional, concise, and better expressed while maintaining the original meaning." Ⅱ. Fine-tuning and rewriting by large models: Input the original sentence, the matching template, context, examples, vocabulary, and the above polishing instructions (Prompt) into the large model. Based on these constraints, the large model makes meticulous modifications to the original sentence, which may involve sentence pattern adjustment, word replacement, and optimization of expression methods, etc. The processing object is each sentence in the sentence sequence generated in step S6.

[0059] Ⅲ. Combining the final text: Recombine all the sentences in the sentence sequence that have been fine-tuned and rewritten by the large model in their original order to form the final optimized copywriting.

[0060] As shown in the provided example: Original sentence 1: "The experimental results show...", combined with the matching template of "To verify...", template examples, excellent vocabulary (effectiveness, significance, diversity), and context, generate Prompt. The large model outputs the polished text: "To verify the effectiveness of the DRT-PP method...". (This reflects the learning of template structure, example expression, and vocabulary application).

[0061] Original sentence 2: "Regarding the above problems...", combined with the matching template of "To balance...", template examples, excellent vocabulary (innovatively, efficiently, contradiction), and context, generate Prompt. The large model outputs the polished text: "Regarding the existing problems in the current field of autonomous driving...". (This also reflects the integrated application of various aspects of information).

[0062] Embodiment 2 Please refer to Figure 2 , which shows a schematic structural diagram of a macro-micro fusion copywriting language optimization system proposed in the second embodiment of the present application. This system can be deployed on a server and provide services to users through an API or a web interface. The system includes the following key modules: Domain classification module 100, used to classify the copywriting to be optimized, and determine its industry domain and theme tags; Inspector management module 200, used to select a set of structured inspectors corresponding to the industry domain, judge the highlight coverage of the set of structured inspectors relative to excellent texts in the same domain as the copywriting to be optimized. If the coverage is lower than a preset threshold, generate text inspection content based on the large model and update the set of structured inspectors; Macro optimization framework generation module 300, used to load the updated set of structured inspectors according to the theme tags, and generate a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the set of structured inspectors; Macro rewriting module 400, used to globally rewrite the structure of the copywriting to be optimized by using the large model according to the multi-level optimization framework and the optimization thinking chain, and generate a primary optimized text; The sentence pattern knowledge base construction module 500 is used to disassemble the sentences of the excellent text into fixed templates and variable parameters by performing syntactic analysis and semantic annotation on the excellent text, so as to extract the specific sentence structure of the industry field and construct a reusable sentence pattern knowledge base including phrase level, sentence pattern level and paragraph level; The sentence processing and matching module 600 is used to perform sentence segmentation on the primary optimized text to generate a sentence sequence with context markers, and adopt the vectorization technology based on the pre-trained language model embedding to calculate the semantic similarity between each sentence in the sentence sequence and the sentence pattern template in the reusable sentence pattern knowledge base, and match the sentence pattern template with the highest similarity for each sentence; The micro refinement module 700 is used to perform fine-grained rewriting on the sentences in the sentence sequence by using a large model based on the matched sentence pattern template and the sentence sequence with context markers to generate the final optimized sentences.

[0063] Example of the system working process: The user submits the text to be optimized and the optional reference text through the user input module -> the domain classification module determines the domain and theme -> the macro modification module: obtains the corresponding domain checker from the storage and management module -> (if necessary) calls the large model algorithm engine to update the checker -> generates the optimization framework and CoT -> calls the large model algorithm engine to perform macro rewriting to obtain the primary optimized text -> the micro refinement module: performs sentence segmentation and context marking -> performs sentence vectorization, and obtains the sentence pattern template from the storage and management module for matching -> calls the large model algorithm engine to perform fine-grained rewriting in combination with the template, context, etc. -> outputs the final optimized text to the user.

[0064] A macro-micro fusion copywriting language optimization system in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a Network Attached Storage (NAS), a personal computer (PC), etc., which are not specifically limited in the embodiments of the present application.

[0065] An example of a macro-micro fusion copywriting language optimization system in this application can be a device with an operating system. The operating system can be the Android operating system, the iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of this application.

[0066] The macro-micro fusion copywriting language optimization system provided by the embodiments of this application can implement Figure 1 each process implemented by a macro-micro fusion copywriting language optimization method in the method embodiments. To avoid repetition, it will not be elaborated here.

[0067] Optionally, the embodiments of this application also provide an electronic device, including a processor, a memory, a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above-mentioned macro-micro fusion copywriting language optimization method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0068] The embodiments of this application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, it implements each process of the above-mentioned macro-micro fusion copywriting language optimization method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0069] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc.

[0070] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of this application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0071] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0072] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. A macro-micro integration copywriting language optimization method, characterized in that: The method comprises the following steps: S1: Classify the texts to be optimized and determine their industry and topic tags; S2: Select a set of structured checkers corresponding to the industry field, and determine the highlight coverage of the structured checker set relative to the excellent texts in the same field of the copy to be optimized. If the coverage is lower than a preset threshold, generate text inspection content based on the big model and update the structured checker set; S3: loading the updated structured checker set according to the topic tag, and generating a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the structured checker set; S4: using the big model to rewrite the global structure of the text to be optimized according to the multi-level optimization framework and the optimization thinking chain, and generate a primary optimized text; S5: By performing syntactic analysis and semantic annotation on excellent texts, the sentences of the excellent texts are disassembled into fixed templates and variable parameters to extract industry-specific sentence structures and build a reusable sentence pattern knowledge base including phrase level, sentence level and paragraph level; S6: Segment the primary optimized text into sentences to generate a sentence sequence with context tags; S7: using a vectorization technology based on pre-trained language model embedding, calculating the semantic similarity between each sentence in the sentence sequence and the sentence template in the reusable sentence pattern knowledge base, and matching each sentence with the sentence template with the highest similarity; S8: Based on the matched sentence template and the sentence sequence with context tags, the sentences in the sentence sequence are refined and rewritten using the large model to generate a final optimized sentence.

2. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: The step S2 of determining the highlight coverage of the structured checker set relative to the excellent text in the same field of the copy to be optimized includes: Input the excellent text into the structured checker set, count the number of highlight features in the excellent text that are recognized by the structured checker set, and calculate the ratio of the number to the total number of highlight features in the excellent text as the coverage; If the coverage is lower than a preset threshold, a new checker is obtained based on the big model by comparing the differences between the excellent text and the text to be optimized to generate text inspection content. The new checker includes function description, detection dimensions, evaluation criteria and improvement suggestions.

3. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: In step S3: The multi-level optimization framework includes a structure layer, a semantic layer, and a style layer, wherein: The structural layer is used to optimize the logical flow and information organization of the copy. The semantic layer is used to ensure terminology accuracy and content integrity. The style layer is used to adjust the expression style to comply with industry standards; The optimization thinking chain includes a text analysis step and a text optimization step, wherein: The text analysis step includes identifying the keywords and core ideas of the text to be optimized, The text optimization steps include structural adjustment, content supplementation and terminology unification.

4. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: The step S5 includes: performing syntactic analysis and semantic annotation on the excellent text. Use NLP tools to perform dependency parsing and named entity recognition to identify fixed templates and variable parameters in sentences; Abstracting the fixed template into a template structure containing placeholders, and recording a filling example of the variable parameters; The reusable sentence pattern knowledge base is indexed by domain, theme and sentence function, and contains templates at phrase level, sentence level and paragraph level.

5. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: Generating a sentence sequence with context tags in step S6 includes: Using a sentence segmentation tool, segment the primary optimized text into sentences to obtain a list of sentence sequences; Traverse the list, and for the i-th sentence, record the contents of its previous sentence and next sentence as context tags.

6. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: The vectorization technology based on pre-trained language model embedding adopted in step S7 includes: Convert each sentence in the sentence sequence and the sentence pattern template in the sentence pattern knowledge base into a vector representation of a fixed dimension respectively; The cosine similarity between the sentence vector and the template vector is calculated, and the template with the highest similarity is selected as the matching result.

7. The method for optimizing copywriting language by macro-micro integration according to claim 1, characterized in that: The step S8 comprises: For each sentence in the sentence sequence and its matching best sentence template, as well as the recorded context information and sample sentences and excellent vocabulary obtained from the sentence template, form polishing instructions for the large model; Inputting the original sentence, matching template, context, examples, vocabulary and polishing instructions into the big model, and rewriting the original sentence in a refined manner, wherein the rewriting includes sentence structure adjustment, word replacement and expression optimization; All the sentences in the sentence sequence that have been refined and rewritten by the large model are recombined in order to form the final optimized copy.

8. A macro-micro integration copywriting language optimization system, characterized in that: The system comprises: The field classification module is used to classify the optimized copywriting into fields and determine its industry field and subject tags; The checker management module is used to select a set of structured checkers corresponding to the industry field, determine the highlight coverage of the structured checker set relative to the excellent texts in the same field of the copy to be optimized, and if the coverage is lower than a preset threshold, generate text check content based on the big model and update the structured checker set; A macro optimization framework generation module, used for loading the updated structured checker set according to the topic tag, and generating a multi-level optimization framework and an optimization thinking chain based on the rule constraints of the structured checker set; A macro rewriting module is used to use a macro model to rewrite the global structure of the text to be optimized according to the multi-level optimization framework and the optimization thinking chain to generate a primary optimized text; A sentence pattern knowledge base construction module is used to perform syntactic analysis and semantic annotation on excellent texts, decompose the sentences of the excellent texts into fixed templates and variable parameters, extract industry-specific sentence structures, and construct a reusable sentence pattern knowledge base including phrase level, sentence level and paragraph level; A sentence processing and matching module, which is used to segment the primary optimized text into sentences, generate a sentence sequence with context tags, use a vectorization technology based on pre-trained language model embedding, calculate the semantic similarity between each sentence in the sentence sequence and the sentence template in the reusable sentence pattern knowledge base, and match each sentence with the sentence template with the highest similarity; The micro-polishing module is used to perform fine rewriting on the sentences in the sentence sequence based on the matched sentence template and the sentence sequence with context markers using a large model to generate a final optimized sentence.

9. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a macro-micro fusion copywriting language optimization method as described in any one of claims 1 to 7 are implemented.

10. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the macro-micro fusion copywriting language optimization method as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Method and system for generating bionic hierarchical memory fusion document of large electric semantic model

    CN119474347A