Multi-modal large model-based copywriting editing method and device and storage medium
By generating the copywriting theme and background through a multimodal large model and determining the copywriting construction plan, the problems of single copywriting format and writing method in the existing technology are solved, and personalized design of copywriting and improved user experience are achieved.
Patent Information
- Application Number
- CN202510718128.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The copywriting generated by existing technologies has a single format and writing style, and the expression method does not conform to the convention, resulting in a reduced user experience.
A multimodal large model is used to generate the copy theme and background by receiving external prompt data, determine the copy construction plan, including the copy expression method and form, and output the copy construction plan to achieve personalized design.
It realizes the personalized design of copywriting, meets the diverse needs of users, and improves the differentiation of copywriting and user experience.
Smart Images

Figure CN120706392A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of copywriting processing, and in particular to a copywriting editing method, device and storage medium based on a multimodal large model. Background Art
[0002] With the development of technology, artificial intelligence technology has been widely used in various copywriting to improve people's work efficiency.
[0003] The common practice of existing solutions is to form a corresponding template for a specific genre, then use the template and the corresponding copy as a training set, and train the model to fill the corresponding copy into the template.
[0004] However, the copy generated by the above method has a single format and writing style, and often the expression does not conform to conventional expressions, thereby reducing the user experience. Summary of the Invention
[0005] In view of the above-mentioned solution, the present application aims to propose a copy editing method, device and storage medium based on a multimodal large model to solve at least one of the above-mentioned technical problems.
[0006] In a first aspect, one or more embodiments of this specification provide a copywriting editing method based on a multimodal large model, which pre-builds the multimodal large model, including:
[0007] receiving first prompt data input from an external source;
[0008] Generate a corresponding copy theme and copy background according to the first prompt data;
[0009] receiving second prompt data input from an external source;
[0010] Obtaining a copywriting construction plan according to the copywriting theme, the copywriting background, and the second prompt data, wherein the copywriting construction plan includes a copywriting expression method and a copywriting expression form; and
[0011] According to the copywriting construction plan and the preset template, the target copywriting is obtained.
[0012] Furthermore, the copywriting expression includes: pictures; the second prompt data includes prompt words;
[0013] Get a copywriting solution, including:
[0014] determining first target data that needs to be expressed in a graphic from the second prompt data;
[0015] Classifying the first target data to obtain a type corresponding to the first target data;
[0016] determining a text expression corresponding to the first target data according to the type and the prompt word corresponding to the first target data;
[0017] According to the type, the prompt word corresponding to the first target data, and the text expression form, a graph corresponding to the first target data is obtained.
[0018] Furthermore, the types include: text-to-image conversion;
[0019] Obtaining a graph corresponding to the first target data includes:
[0020] According to the expression form of the text, determining the drawing reference information in a preset picture-text database;
[0021] A graph corresponding to the first target data is drawn according to the mapping reference information and the text expression form.
[0022] Furthermore, the copywriting expression includes: text; the second prompt data includes prompt words;
[0023] Get a copywriting solution, including:
[0024] determining second target data that needs to be expressed in text according to the second prompt data;
[0025] extracting keywords from the second target data;
[0026] determining a text expression corresponding to the second target data according to the keyword and the prompt word corresponding to the second target data;
[0027] extracting corresponding text from a preset text pair database according to the keyword and the prompt word corresponding to the second target data;
[0028] The text corresponding to the second target data is determined according to the copywriting expression corresponding to the second target data and the extracted text.
[0029] Furthermore, the second target data is a picture, and the method further includes:
[0030] Recognizing text in the second target data;
[0031] Extracting keywords based on the recognized text;
[0032] Determining reference information from a preset image-text pair database according to the keyword and the prompt word corresponding to the second target data;
[0033] The text corresponding to the second target data is determined according to the reference information.
[0034] In a second aspect, an embodiment of the present application provides a copy editing device based on a multimodal large model, comprising:
[0035] A first receiving module receives first prompt data input from an external source;
[0036] A generation module, generating a corresponding copy theme and copy background according to the first prompt data;
[0037] A second receiving module receives second prompt data input from the outside;
[0038] The data processing module inputs the copy theme, the copy background and the second prompt data into a preset multimodal large model to obtain a copy construction plan, which includes a copy expression method and a copy expression form; and obtains the target copy according to the copy construction plan and the preset template.
[0039] Furthermore, the copywriting expression includes: pictures; the second prompt data includes prompt words;
[0040] The data processing module is used to determine the first target data that needs to be expressed in a graphic from the second prompt data; classify the first target data to obtain the type corresponding to the first target data; determine the text expression form corresponding to the first target data based on the type and the prompt word corresponding to the first target data; and obtain the graphic corresponding to the first target data based on the type, the prompt word corresponding to the first target data, and the text expression form.
[0041] Furthermore, the types include: text-to-image conversion;
[0042] The data processing module is used to determine mapping reference information in a preset image-text pair database according to the text expression form; and draw a graph corresponding to the first target data according to the mapping reference information and the text expression form.
[0043] Furthermore, the copywriting expression includes: text; the second prompt data includes prompt words;
[0044] The data processing module is used to determine second target data that needs to be expressed in text based on the second prompt data; extract keywords from the second target data; determine the text expression corresponding to the second target data based on the keywords and the prompt words corresponding to the second target data; extract corresponding text from a preset text pair database based on the keywords and the prompt words corresponding to the second target data; and determine the text corresponding to the second target data based on the text expression corresponding to the second target data and the extracted text.
[0045] In a third aspect, an embodiment of the present application provides a storage medium for storing computer-executable instructions, which, when executed, implement the steps of the multimodal large model-based copy editing method described in any one of the first aspects.
[0046] Compared with the existing technology, this application can at least achieve the following technical effects:
[0047] Copywriting typically consists of two parts: content and format. Relatively speaking, content is more complex because it involves multiple forms of expression. To better present the content, this application utilizes a large, multimodal model with strong adaptability, wide application scope, strong scalability, and a high degree of automation. To reflect the diversity of copywriting, this application does not directly output the copywriting, but instead outputs a copywriting construction plan. In this application, the copywriting construction plan includes the copywriting expression method and the copywriting expression form. The copywriting expression methods include: images and text. The copywriting expression form of images includes but is not limited to color, size, brightness, coordinate axes, visualization chart type, and data source selection. The textual copywriting expression form includes but is not limited to style, writing style, theme, paragraph format, and word count. In other words, the copywriting construction plan is a plan for generating the copywriting based on the copywriting expression form and the copywriting expression method. This allows users to obtain detailed design ideas for the copywriting. Based on this detailed idea, users can easily modify the copywriting construction plan to achieve personalized report design. In addition, since the present application can select the corresponding copywriting expression method and copywriting expression form according to the needs of the user, the copywriting generated by the present application can also meet personalized needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0049] Figure 1 A flowchart of a copy editing method based on a multimodal large model provided for one or more embodiments of this specification;
[0050] Figure 2 A schematic structural diagram of a copy editing device based on a multimodal large model provided in one or more embodiments of this specification. DETAILED DESCRIPTION
[0051] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.
[0052] Existing technologies rely on templates and utilize artificial intelligence to generate copy. This approach is essentially a form-filling approach. For example, consider hospital examination reports. Regardless of the type of illness, most people perceive these reports as uniform. This is because hospital examination reports often simply modify the information in the form. This means that this "form-filling" approach to copywriting is only suitable for this type of scenario. For copywriting that relies heavily on expressiveness, this approach struggles to meet client needs.
[0053] Take a technical report, for example. Before a report is completed, its evaluation metrics must be defined. These metrics often evolve with policy changes, technological advancements, and market demands. Obviously, the impact of changes in evaluation metrics on the entire report cannot be simply addressed by changing the data.
[0054] Fundamentally, these reports are based on expressive ability: using language and text within a specific genre to articulate perspectives, opinions, or express thoughts and feelings. Therefore, before writing a document like a scientific report, careful planning is required, including overall layout, paragraph division, choice of expression, and data selection. However, the "fill-in-the-form" approach lacks this planning, resulting in a uniform appearance and even a sense of incoherence between sentences and paragraphs.
[0055] In order to solve the above problems, the present application provides a copy editing method based on a multimodal large model, such as Figure 1 As shown, the following steps are included:
[0056] Step 1: Receive first prompt data input from the outside.
[0057] In the embodiment of the present application, taking a technical report as an example, the first prompt data includes: keywords input by the user, scientific research data, and survey data.
[0058] Step 2: Generate a corresponding copy theme and copy background based on the first prompt data.
[0059] In the embodiment of the present application, taking a technical report as an example, the purpose and scope of the report (theme and background of the report) can be determined based on the first prompt data.
[0060] Step 3: Receive the second prompt data input from the outside.
[0061] In the embodiment of the present application, taking a technical report as an example, after the user determines the purpose and scope of the report, the user prompts the model to continue generating the report by asking questions. Detailed data can also be provided for the model to refer to.
[0062] Step 4: Obtain a copywriting construction plan based on the copywriting theme, the copywriting background and the second prompt data.
[0063] In an embodiment of the present application, the copywriting construction scheme includes a copywriting expression method and a copywriting expression form. Among them, the copywriting expression method includes: pictures and text, and the copywriting expression form of the picture includes but is not limited to color, size, brightness, coordinate axis, visualization chart type and data source selection, etc. The text copywriting expression form includes but is not limited to style, style, theme, paragraph format and word count, etc. Through the text copywriting expression form and copywriting expression method, the specific details of the design report are realized to realize a personalized design technical report. Users can modify the copywriting construction scheme according to their own preferences, such as paragraph format, picture size, etc., to further realize a personalized design technical report.
[0064] For example, when a technology report presents a certain technological development trend, the copywriting plan is as follows:
[0065] A line graph is used, with the X-axis representing time, scaled in years, and the Y-axis representing investment amounts, scaled in 100 million yuan. The graph is accompanied by explanatory and argumentative texts, written in a rigorous style.
[0066] It can be seen that personalized settings of technical reports can be achieved through the above method.
[0067] In this embodiment, the language model used in this application can be IXC-2.5 (Intern Language Model - XComposer2.5). This model is a general-purpose large visual language model that supports long context input and output. It excels in a variety of text-image understanding and synthesis applications, achieving GPT-4V-level capabilities using only a 7B LLM backend. After training on 24KB of interleaved image-text context, it can be seamlessly extended to 96KB of context using rotational encoding extrapolation.
[0068] The IXC-2.5 model structure is:
[0069] (1) Visual Encoder: IXC-2.5 uses a pre-trained visual encoder to extract high-quality features from images. Furthermore, when used with the Partial LoRA (Partial Low-Rank Adaptation) module, this lightweight visual model achieves better results and is more efficient.
[0070] (2) Multimodal large model: A fine-tuned version of InternLM2-Chat is selected. InternLM2-Chat is a version of InternLM2 that has undergone SFT (supervised fine-tuning) and RLHF (reinforcement learning with human feedback) and is optimized for conversational interaction. Therefore, InternLM2-Chat has excellent multilingual capabilities, command following, empathetic chat, and tool calling capabilities.
[0071] (3) Partial Low-Rank Adaptation: Since the alignment of individual models in a large multimodal model has not been well solved, we hope that model alignment can enrich the capabilities of a large multimodal model while preserving its original capabilities. Currently proposed alignment methods all align different models in the same way or treat them as separate entities, which will lose the inherent properties between models or increase the alignment cost.
[0072] Since the multimodal large model is a general model, it is necessary to fine-tune the multimodal large model. The specific process is as follows:
[0073] A large number of technical reports are selected as samples. There are two ways to select samples:
[0074] Proportional data collection: For example, consider sampling 20,000 reports from a company's technical report database. The sampled data is drawn based on the proportion of each technical category in the overall database. This approach preserves the true distribution of the data, as class imbalance is natural in these scenarios. However, it also presents a disadvantage: minority classes may be difficult for the model to effectively learn due to insufficient samples, resulting in poor predictive performance for these classes. Furthermore, the model may be biased towards the majority class, resulting in seemingly good overall performance, but low recall and F1 scores for minority classes.
[0075] Balanced data collection: 2k samples of data were collected for each category, categorized by technology. The advantage of this approach is that a balanced dataset helps the model better learn the characteristics of each category, as each category has an equal number of samples. This can improve the recognition rate of the minority class, as the model is no longer dominated by the majority class. However, oversampling can lead to overfitting, as the model may over-learn repeated samples from the minority class. Undersampling can cause information loss, especially when samples removed from the majority class contain important information. It also changes the original distribution of the data, making it unsuitable for applications where class imbalance is a natural characteristic.
[0076] Since the expression and drawings of technical reports are unfamiliar to the multimodal large model and are difficult to learn, in addition to adding field-specific content for training, it is also necessary to add general data sets. The ratio of general data to technical reports is 1:5. A certain proportion of general data sets is added according to the amount of original data collected.
[0077] The training process based on the above training samples is:
[0078] First, based on the InternLM-XComposer-2.5 model, we trained the pure technical report dataset using the default fine-tuning parameters, setting a higher LR (Learning Rate) (such as 1e-5) and a longer epoch (such as 10 rounds). The trained model is first tested on the training set, that is, the same data used for training should be used for testing. Testing on the training set is very important. It can eliminate issues with the dataset quality to a certain extent, eliminate the need to worry about overfitting, and ensure that the underlying framework is in good working order. This is because no matter how poor the data quality is, no matter how poor the generalization is on the test set, there should be good learning effects on the training set. At the same time, ablation experiments can be performed on the first-stage description data and the second-stage instruction data of the technical report data.
[0079] After ensuring that there are no problems on the training set, we then combined the curves on the validation set to roughly confirm most of the hyperparameters, including the training rounds (Epochs), technical report data ratio, learning rate (LR, Learning Rate), batch size, text length, number of MoE experts, parallel configuration (TP PP DP), etc.
[0080] The parameters involved in the above training process are as follows:
[0081] 1. Training Round (Epoch)
[0082] Definition: 1 epoch means that the model completes a complete traversal of the entire training dataset.
[0083] Function: The number of epochs determines the total number of times the model is trained. You should choose an appropriate number of epochs based on the size of your dataset and the complexity of your model.
[0084] For example, if the dataset has 10,000 samples and the batch_size is 100, then one epoch requires 100 iterations (10,000 / 100).
[0085] 2. Data ratio
[0086] Definition: The ratio of technical report data to other types of data in the training data.
[0087] Function: If the task involves technical report-related fields, the proportion of technical report data will affect the domain adaptability and performance of the model.
[0088] For example, if technical reports account for 70% of the training data and other data account for 30%, the model will be more likely to learn features related to technical reports.
[0089] 3. Learning Rate (LR)
[0090] Definition: A hyperparameter that controls the step size for updating model parameters.
[0091] Function: The learning rate determines the magnitude of parameter adjustments during each gradient descent. A learning rate that is too large leads to unstable training, while a learning rate that is too small leads to slow convergence.
[0092] Examples: Common learning rate values are 0.001, 0.0001, etc.
[0093] 4.Batch Size
[0094] Definition: The number of samples fed into the model at each iteration.
[0095] Effect: Batch size affects training speed and memory usage. A larger batch size can improve training efficiency but requires more memory; a smaller batch size may cause training instability.
[0096] Example: A batch size of 32 means that 32 samples are used in each iteration to calculate gradients and update parameters.
[0097] 5. Text Length
[0098] Definition: The maximum text length of the input model (in tokens).
[0099] Purpose: The length of the text determines the maximum capacity of the model to process input data. Text lengths that are too short will truncate the input, while text lengths that are too long will waste computing resources.
[0100] Example: A typical text length for the BERT model is 512 tokens.
[0101] 6. Number of MoE experts
[0102] Definition: MoE (Mixture of Experts) is a model architecture that contains multiple "expert" sub-models, each of which is responsible for processing a specific type of input.
[0103] Role: The number of experts determines the complexity and expressiveness of the MoE model. More experts can improve model performance, but also increase computational costs.
[0104] Example: A MoE model might contain 8 experts, each of which is a small neural network.
[0105] 7. Parallel configuration (TP, PP, DP)
[0106] Definition: Parallel configuration refers to how to distribute computing tasks to multiple devices (such as GPUs) in distributed training.
[0107] TP (Tensor Parallelism): Split a single layer of the model onto multiple devices.
[0108] PP (Pipeline Parallelism): Distribute different layers of the model to multiple devices and execute them in a pipeline manner.
[0109] DP (Data Parallelism): Split the data across multiple devices, with each device running a complete copy of the model.
[0110] Function: Parallel configuration can speed up the training process and support larger models and data.
[0111] For example, when training on 8 GPUs, DP can be used to split the data into 8 parts, with each GPU processing a portion of the data.
[0112] Step 5: According to the copywriting scheme and preset template, the target copywriting is obtained.
[0113] In the embodiment of the present application, the template is the main part of the technical report, and you can refer to the chapters in the table of contents of the technical report. The copywriting construction plan will be presented in the form of a technical report.
[0114] From this, we can see that the way to write a report in this application is not based on a template, but it can plan the details of the report, thereby achieving personalized construction of the report.
[0115] In the embodiment of the present application, the second prompt data includes prompt words. When the text expression is a picture, the process of constructing the corresponding text construction plan is as follows:
[0116] A1. Determine first target data that needs to be expressed in a graphic from the second prompt data.
[0117] In the embodiment of the present application, the user can indicate in the prompt which parts need to be represented in the form of a graph. For example, the following prompt can be given: When analyzing a certain data, it needs to be presented in the form of a graph, and the type of graph can be selected according to the content of the display.
[0118] A2. Classify the first target data to obtain a type corresponding to the first target data.
[0119] In the embodiments of this application, the types of graphs include: data graphs, structure graphs, flow charts, and data tables. Data graphs and data tables are primarily used to visually display statistical information attributes (temporality, quantity, etc.), and are graphical structures that play a key role in knowledge mining and the intuitive and vivid perception of information. Structure graphs are primarily used to display product structures, technical routes, and material structures. Flow charts are primarily used to display processes or logical structures.
[0120] A3. Determine the text expression form corresponding to the first target data according to the type and the prompt word corresponding to the first target data.
[0121] In this application example, using a line graph in a technical report as an example, the textual representation includes: line color, coordinate scale, coordinate maximum value, and the physical quantity corresponding to the coordinate axis. If the user believes that a large model cannot properly handle a data analysis process, they can constrain the model using prompts. For example, a prompt might say, "Trend analysis charts are all displayed as line graphs."
[0122] A4. Obtain a graph corresponding to the first target data according to the type, the prompt word corresponding to the first target data, and the text expression form.
[0123] In an embodiment of the present application, in a technical report, when introducing a certain technology, a simple diagram needs to be attached after the text so that readers can better understand the meaning of the report. In this case, text-to-image conversion is required. Therefore, in order to better solve the problem of this scenario, the present application specifically designs the type corresponding to the first target data to include text-to-image conversion. The specific process of text-to-image conversion is as follows:
[0124] According to the expression form of the text, mapping reference information is determined in a preset image-text database, and a graph corresponding to the first target data is drawn according to the mapping reference information and the expression form of the text.
[0125] It should be noted that the large model itself has the ability to convert text into images. Although the model has been fine-tuned using technical reports, there may still be significant deviations between images and text. To minimize this deviation, based on user needs, the large model's drawing actions are restricted in the form of prompt words and image-text pairs, thus achieving image-text matching.
[0126] In the embodiment of the present application, the copywriting expression method includes: text; the second prompt data includes prompt words; the copywriting process in the form of text includes the following steps:
[0127] B1. Determine second target data that needs to be expressed in text based on the second prompt data.
[0128] In the embodiment of the present application, the user can indicate in the prompt words which parts need to be expressed in the form of text. For example, the following prompt words are given: Analyze the changing trend of relevant data in the form of text.
[0129] B2. Extract keywords from the second target data.
[0130] In the embodiment of the present application, the extracted keywords include the object, content, style and writing style that the text is to describe.
[0131] B3. Determine the copywriting expression form corresponding to the second target data based on the keyword and the prompt word corresponding to the second target data.
[0132] In the embodiments of this application, the textual expression forms include, but are not limited to, style, tone, theme, paragraph format, and word count. Taking the review section of a technical report as an example, the textual expression forms are: using expository writing, following the style of a master's or doctoral dissertation, written in paragraphs in the order of publication of the documents, and with a word count of no less than 5,000 words.
[0133] B4. Extract corresponding text from a preset text pair database based on the keyword and the prompt word corresponding to the second target data.
[0134] In an embodiment of the present application, the training samples of the multimodal large model can be text pairs, and the text pairs can be obtained using the multimodal large model. For example, a text-based training sample is sent to the multimodal large model, and the multimodal large model will split the training sample into text pairs in the form of questions and answers.
[0135] B5. Determine the text corresponding to the second target data based on the copywriting expression corresponding to the second target data and the extracted text.
[0136] In the embodiments of this application, when introducing a certain technology in a technical report, it is necessary to attach a paragraph of text after the figure so that readers can better understand the meaning of the figure. In this case, it is necessary to perform image-to-text conversion. The specific process of image-to-text conversion is as follows:
[0137] When the second target data is a picture, the text in the second target data is identified; keywords are extracted based on the identified text; reference information is determined from a preset image-text pair database based on the keywords and the prompt words corresponding to the second target data; and the text corresponding to the second target data is determined based on the reference information.
[0138] It should be noted that although this application involves image-text conversion and text-image conversion, these two processes share a training sample during training.
[0139] The embodiment of the present application provides a copy editing device based on a multimodal large model, such as Figure 2 Shown, including:
[0140] The first receiving module 201 receives first prompt data input from the outside;
[0141] A generation module 202 generates a corresponding copy theme and copy background according to the first prompt data;
[0142] The second receiving module 203 receives the second prompt data input from the outside;
[0143] The data processing module 204 inputs the copy theme, the copy background and the second prompt data into a preset multimodal large model to obtain a copy construction plan, which includes a copy expression method and a copy expression form; and obtains the target copy based on the copy construction plan and the preset template.
[0144] In the embodiment of the present application, the copywriting expression includes: a picture; the second prompt data includes a prompt word;
[0145] The data processing module 204 is used to determine the first target data that needs to be expressed in a graphic from the second prompt data; classify the first target data to obtain the type corresponding to the first target data; determine the text expression form corresponding to the first target data based on the type and the prompt word corresponding to the first target data; and obtain the graphic corresponding to the first target data based on the type, the prompt word corresponding to the first target data, and the text expression form.
[0146] In the embodiment of the present application, the types include: text-to-image conversion;
[0147] The data processing module 204 is configured to determine drawing reference information in a preset image-text pair database according to the text expression form; and draw a graph corresponding to the first target data according to the drawing reference information and the text expression form.
[0148] In the embodiment of the present application, the copywriting expression includes: text; the second prompt data includes prompt words;
[0149] The data processing module 204 is used to determine the second target data that needs to be expressed in text based on the second prompt data; extract keywords from the second target data; determine the text expression corresponding to the second target data based on the keywords and the prompt words corresponding to the second target data; extract corresponding text from a preset text pair database based on the keywords and the prompt words corresponding to the second target data; and determine the text corresponding to the second target data based on the text expression corresponding to the second target data and the extracted text.
[0150] The technical solution of this application is not only applicable to generating technical reports, but also to generating other text data, such as news, academic, and advertising. Specifically, when training a multimodal large model, it is necessary to convert the training samples from technical reports to corresponding text data. The textual expression of technical reports and the textual expression of images are also applicable to other text types. For example, news usually requires certain requirements for the format of images, so the language large model needs to design the color, size, and brightness of the images to ensure that the images meet the requirements. Many images in academic texts are used to demonstrate technical effects, so text-to-image conversion is also involved. During the text-to-image conversion process, the large model also needs to select coordinate axes, visualization chart types, and data sources to better express the technical effects. The above-mentioned types of text require explanations of concepts, nouns, and legal provisions, so it is also necessary to select the style, tone, theme, paragraph format, and word count. When writing advertising texts, users will also upload relevant images, so text-to-image conversion is also required.
[0151] Unlike technical reports, graphics and text for other text data can be generated separately or simultaneously. Therefore, the corresponding copywriting construction solutions are structurally different and can be divided into the following scenarios:
[0152] Scenario 1
[0153] The user first determines the type of news report, such as a feature story, and then provides the corresponding topic name and core idea. Based on the user's description and system prompts, the system generates a text-only copywriting plan based on a multimodal large model, including core factual information, emotional resonance points, character traits, and narrative structure. Based on this content, the system then generates a corresponding graphic construction plan for the feature story.
[0154] Scenario 2
[0155] The user first identifies the technical field covered by the academic text and then provides relevant content, such as the topic and the main idea of the text. Based on the user's description and system prompts, the system directly generates a complete copywriting plan based on a multimodal macro model. For example, the full text is divided into an introduction, body, and conclusion. The introduction uses a concise style to introduce the background and the problem. The body adopts a precise and rigorous style, with the main paragraphs using a topic sentence and supporting sentences. The conclusion adopts a precise and concise style.
[0156] Scenario 3
[0157] The user first determines the specific category of the ad and then provides the desired theme. Based on the user's description and system prompts, the system generates a copywriting plan for the ad text based on a multimodal large model, including the title, slogan, body slogan, and specific requirements for related graphics. After reviewing the copywriting plan, the user may feel the need for further supplementary plans. One or more relevant schematics or sketches are uploaded to the system. Based on preset prompts and the user-submitted schematics, the system converts the image into text and updates the copywriting plan for the title, slogan, body slogan, or related graphics.
[0158] When the user uploads a video, the multimodal large model can call the corresponding video analysis software to decompose the video frame by frame to obtain the pictures corresponding to the supplementary solution.
[0159] An embodiment of the present application provides a storage medium for storing computer-executable instructions, characterized in that when the computer-executable instructions are executed, the steps of the copy editing method based on a multimodal large model described in any one of the embodiments are implemented.
[0160] It should be noted that the embodiment of the storage medium in this specification and the embodiment of the blockchain-based service provision method in this specification are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding copy editing method based on the multimodal large model mentioned above, and the repeated parts will not be repeated.
[0161] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0162] In the 1930s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly done using "logic compiler" software. This is similar to the software compiler used when developing programs. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0163] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code format, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules that implement the method and structures within the hardware component.
[0164] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0165] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0166] Those skilled in the art will appreciate that one or more embodiments of this specification may be provided as a method, system, or computer program product. Thus, one or more embodiments of this specification may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0167] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0168] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0170] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0171] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0172] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0173] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0174] One or more embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0175] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0176] The foregoing description is merely an example of the present invention and is not intended to limit the present invention. Persons skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims herein.
Claims
1. A copywriting editing method based on a multimodal large model, characterized in that: Pre-built multimodal large models, including: receiving first prompt data input from an external source; Generate a corresponding copy theme and copy background according to the first prompt data; receiving second prompt data input from an external source; Obtaining a copywriting construction plan according to the copywriting theme, the copywriting background, and the second prompt data, wherein the copywriting construction plan includes a copywriting expression method and a copywriting expression form; and According to the copywriting construction plan and the preset template, the target copywriting is obtained.
2. The method according to claim 1, characterized in that The copywriting expression includes: pictures; the second prompt data includes prompt words; Get a copywriting solution, including: determining first target data that needs to be expressed in a graphic from the second prompt data; Classifying the first target data to obtain a type corresponding to the first target data; determining a text expression corresponding to the first target data according to the type and the prompt word corresponding to the first target data; According to the type, the prompt word corresponding to the first target data, and the text expression form, a graph corresponding to the first target data is obtained.
3. The method according to claim 2, characterized in that The types include: text-to-image conversion; Obtaining a graph corresponding to the first target data includes: According to the expression form of the text, determining the drawing reference information in a preset picture-text database; A graph corresponding to the first target data is drawn according to the mapping reference information and the text expression form.
4. The method according to claim 1, wherein The copywriting expression includes: text; the second prompt data includes prompt words; Get a copywriting solution, including: determining second target data that needs to be expressed in text according to the second prompt data; extracting keywords from the second target data; determining a text expression corresponding to the second target data according to the keyword and the prompt word corresponding to the second target data; extracting corresponding text from a preset text pair database according to the keyword and the prompt word corresponding to the second target data; The text corresponding to the second target data is determined according to the copywriting expression corresponding to the second target data and the extracted text.
5. The method according to claim 4, characterized in that The second target data is a picture, and the method further includes: Recognizing text in the second target data; Extracting keywords based on the recognized text; Determining reference information from a preset image-text pair database according to the keyword and the prompt word corresponding to the second target data; The text corresponding to the second target data is determined according to the reference information.
6. A copy editing device based on a multimodal large model, characterized in that: include: A first receiving module receives first prompt data input from an external source; A generation module, generating a corresponding copy theme and copy background according to the first prompt data; A second receiving module receives second prompt data input from the outside; The data processing module inputs the copy theme, the copy background and the second prompt data into a preset multimodal large model to obtain a copy construction plan, which includes a copy expression method and a copy expression form; and obtains the target copy according to the copy construction plan and the preset template.
7. The device according to claim 6, characterized in that The copywriting expression includes: pictures; the second prompt data includes prompt words; The data processing module is used to determine the first target data that needs to be expressed in a graphic from the second prompt data; classify the first target data to obtain the type corresponding to the first target data; determine the text expression form corresponding to the first target data based on the type and the prompt word corresponding to the first target data; and obtain the graphic corresponding to the first target data based on the type, the prompt word corresponding to the first target data, and the text expression form.
8. The device according to claim 7, characterized in that The types include: text-to-image conversion; The data processing module is used to determine mapping reference information in a preset image-text pair database according to the text expression form; and draw a graph corresponding to the first target data according to the mapping reference information and the text expression form.
9. The device according to claim 6, characterized in that The copywriting expression includes: text; the second prompt data includes prompt words; The data processing module is used to determine second target data that needs to be expressed in text based on the second prompt data; extract keywords from the second target data; determine the text expression corresponding to the second target data based on the keywords and the prompt words corresponding to the second target data; extract corresponding text from a preset text pair database based on the keywords and the prompt words corresponding to the second target data; and determine the text corresponding to the second target data based on the text expression corresponding to the second target data and the extracted text.
10. A storage medium for storing computer-executable instructions, characterized in that: When executed, the computer executable instructions implement the steps of the copy editing method based on a multimodal large model according to any one of claims 1 to 5.