A method, system and electronic device for generating a document oriented to user intent

By providing users with function keys and predefined templates, AI can better understand user tasks, solve the problem that the existing technology of manuscript generation cannot meet the real needs of users, and achieve more accurate and reliable manuscript generation.

CN117494725BActive Publication Date: 2025-06-17ZHEJIANG FANGZHENG PRINTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311583443.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-06-17
Estimated Expiration
2043-11-24

AI Technical Summary

Technical Problem

The automatically generated documents in the prior art cannot accurately meet the real or deep needs of users, resulting in a poor user experience.

Method used

By providing users with selectable function keys and combining predefined templates, rules and algorithms, AI helps better understand users' tasks and generate more accurate documents.

Benefits of technology

It improves the accuracy of AI in understanding user intentions, reduces misunderstandings and errors caused by unclear language expression, and provides more accurate and reliable document generation services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117494725B_ABST
    Figure CN117494725B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence, and provides an AI manuscript generation method, system, device, and electronic device for intelligent emergence of cloud native and distributed infrastructure. The manuscript generation method includes: determining the server type of the user, where the server type includes a public domain server and a private domain server; providing corresponding function keys for the user based on the server type of the user; collecting the function keys selected by the user and the prompt words input; and inputting the function keys and prompt words into a manuscript generation model corresponding to the server type for manuscript generation. It is used to solve the problem that the automatically generated manuscripts in the prior art cannot truly meet the real needs or deep-level needs of users. By providing function keys for the user to select, combining the function keys selected by the user and the prompt words manually input, the industry characteristics and private data processing of the user can be more accurately known, so as to provide more accurate manuscript generation services for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, system and electronic device for generating a document oriented to user intent. Background Art

[0002] Document generation technology is an application based on computer science and natural language processing technology. It can automatically generate text content, usually according to user input, requirements or specified topics. Such a system uses advanced algorithms and models to understand natural language. By learning a large amount of text data, it has autonomously "understood" some basic laws and patterns of language, showing an "intelligence", an intelligence that emerges autonomously from countless data and experience samples and generates a document that conforms to grammar, semantics and context.

[0003] When using a related document generation model for automatic document generation, the quality of the model writing depends on the data of the model itself and the understanding of natural language on the one hand, and on whether the prompt content (prompt) we input is accurate and specific on the other hand. Users need to provide some keywords to express their needs, but users often cannot clearly express their true needs or deep needs through simple descriptions, resulting in the automatically generated document may not be what the user really needs, making the user experience poor. Summary of the Invention

[0004] The present invention provides a method, system and electronic device for generating a document oriented to user intent, to solve the defect that the automatically generated document in the prior art cannot truly meet the true needs or deep needs of users, resulting in poor user experience. By providing users with function keys that can be selected and converted into a specific format of input, and using predefined templates, rules and algorithms to process, enabling AI to better understand the task and give corresponding answers. Maximize the ability of AI to accurately understand the task, reduce misunderstandings and errors caused by unclear language expression, enable it to accurately and reliably execute specific tasks, and combine the function keys selected by the user and the prompt words manually input to perform thinking combination reasoning for document generation, so as to provide more accurate document generation services for customers.

[0005] The present invention provides a method for generating a document oriented to user intent, including:

[0006] Determine the server type of the user, where the server type includes a public domain server and a private domain server;

[0007] Based on the server type of the user, provide corresponding function keys for the user;

[0008] Collect the function keys selected by the user and the prompt words input.

[0009] Input the function keys and prompt words into the manuscript generation model corresponding to the server type for manuscript generation. The manuscript generation model includes a public domain manuscript generation model and a private domain manuscript generation model. When the user is a public domain server, use the public domain manuscript generation model for manuscript generation. When the user is a private domain server, use the private domain manuscript generation model for manuscript generation.

[0010] According to the manuscript generation method for user intent provided by the present invention, the private domain manuscript generation model is established in the following manner:

[0011] Based on a large number of private domain manuscripts, establish a corpus model library;

[0012] Based on the corpus model library, further establish a private domain manuscript style library to obtain the private domain manuscript generation model;

[0013] Provide a learning framework for the private domain manuscript generation model, and the private domain manuscript generation model performs self-learning and updating under the learning framework;

[0014] Automatically assign weights to each part of the manuscript.

[0015] According to the manuscript generation method for user intent provided by the present invention, when the user's server type is a private domain server, determine the field where the private domain server is located, and the function keys include function keys related to the field.

[0016] When the user's server type is a public domain server, the function keys include function keys for different fields.

[0017] According to the manuscript generation method for user intent provided by the present invention, the manuscript generation method further includes:

[0018] Based on each generated manuscript, perform reinforcement training on the public domain manuscript generation model;

[0019] Train the private domain manuscript generation model based on the public domain manuscript generation model.

[0020] According to the manuscript generation method for user intent provided by the present invention, training the private domain manuscript generation model based on the public domain manuscript generation model includes:

[0021] Determine the semantic data and syntactic data in the public domain manuscript generation model, and determine the semantic data and syntactic data in the private domain manuscript generation model;

[0022] If the semantic data and syntactic data in the public domain manuscript generation model are inconsistent with the semantic data and syntactic data in the private domain manuscript generation model, supplement the semantic data and syntactic data in the public domain manuscript generation model to the private domain manuscript generation model.

[0023] The method for generating a document oriented to user intent provided by the present invention further includes:

[0024] Adjust the usage priorities of semantic data and syntactic data in the supplemented private-domain document generation model.

[0025] According to the method for generating a document oriented to user intent provided by the present invention, input the function keys and prompt words into the document generation model corresponding to the server type for document generation, including:

[0026] Automatically generate the body of the document through the document generation model;

[0027] Automatically generate the title of the document based on the body of the document.

[0028] According to the method for generating a document oriented to user intent provided by the present invention, automatically generate the title of the document, including:

[0029] Calculate the correlation weight values of each paragraph of the body of the document;

[0030] Generate a set of key paragraphs based on the correlation weight values between paragraphs;

[0031] Determine the function keys included in the document based on the set of key paragraphs;

[0032] Automatically generate the title of the document based on the function keys included in the document.

[0033] The present invention also provides a document generation device oriented to user intent, including:

[0034] A determination module for determining the server type of the user, where the server type includes a public-domain server and a private-domain server;

[0035] A function key selection module for providing corresponding function keys for the user based on the server type of the user;

[0036] An acquisition module for acquiring the function keys selected by the user and the input prompt words;

[0037] A generation module is used to input function keys and prompting words into a manuscript generation model corresponding to the server type. The manuscript generation model performs thinking combination reasoning for manuscript generation based on function keys and prompting words. The manuscript generation model includes a public-domain manuscript generation model and a private-domain manuscript generation model. When the user is a public-domain server, the public-domain manuscript generation model is used for manuscript generation. When the user is a private-domain server, the private-domain manuscript generation model takes "industry knowledge" as the core feature. On the basis of a pre-trained model, a large-scale knowledge graph is further integrated to mine a large amount of industry-specific data and knowledge existing in industry application scenarios, and then combined with the knowledge of industry experts to conduct integrated learning from large-scale knowledge and massive data, and internalize the knowledge into the model parameters to perform manuscript generation.

[0038] The present invention also provides a manuscript generation system for user intent, including:

[0039] A control module is used to execute any one of the above manuscript generation methods;

[0040] A server is used for manuscript storage. The server includes a public-domain server and a private-domain server;

[0041] A login module is electrically connected to the server and is used to verify the user's identity and determine the user's server type.

[0042] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements any one of the above manuscript generation methods for user intent.

[0043] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above manuscript generation methods for user intent.

[0044] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements any one of the above manuscript generation methods for user intent.

[0045] In the solution of this application, a method, system, and device for generating a document oriented to user intent with intelligent emergence of cloud native and distributed infrastructure are provided. By integrating a public domain server and a private domain server, a flexible and secure document generation experience is provided for users. The public domain server is responsible for providing public document generation storage services, while the private domain server focuses on providing private document generation services, creating a differentiated service level for users. This design not only takes into account the privacy and security needs of users, but also allows users to have more choices and control during the document generation process, making the entire system more elastic and practical. On the other hand, function keys can be provided for users to choose from. By combining the function keys selected by users and the prompt words manually input, the true needs of users can be more accurately understood, so as to provide more accurate document generation services for customers. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 is a schematic flowchart of the method for generating a document oriented to user intent provided by an embodiment of the present invention;

[0048] Figure 2 is a schematic structural diagram of the device for generating a document oriented to user intent provided by an embodiment of the present invention;

[0049] Figure 3 is a schematic structural diagram of the system for generating a document oriented to user intent provided by an embodiment of the present invention;

[0050] Figure 4 is a schematic structural diagram of the electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0052] Figure 1 is a schematic flowchart of the method for generating a document oriented to user intent provided by an embodiment of the present invention.

[0053] Such asFigure 1 As shown in the figure, this embodiment provides a method for generating a document oriented to user intent, including:

[0054] Step 101, determine the server type of the user, where the server type includes a public domain server and a private domain server;

[0055] Step 102, based on the server type of the user, provide the corresponding function keys for the user;

[0056] Step 103, collect the function keys selected by the user and the input prompt words;

[0057] Step 104, input the function keys and prompt words into the document generation model corresponding to the server type for document generation. The document generation model includes a public domain document generation model and a private domain document generation model. When the user is a public domain server, the public domain document generation model is used for document generation. When the user is a private domain server, it covers more complex text analysis and generation tasks, and the private domain document generation model is used for document generation.

[0058] In implementation, in the document generation model, the grammar and semantic rules of natural language can be learned by means of pre-training, and the semantics of the user can be understood according to the function keys selected by the user and the input tube detection. Thinking and learning ability, and analyze the value of participating keywords in the document for SEO ranking, and generate a document that meets industry optimization;

[0059] The order of the training data of the private domain server enables the model to master as many knowledge points and skills as possible, reduce the waste of parameters, and thus achieve the "lightweight" of scenario application. The public domain server restricts the prompt words, embeds a large amount of domain-specific knowledge and fine-tuning technology, so that the model can only be based on a certain type of questions for a certain type of identity. For other types of questions, the robot will inform the user that it does not understand the relevant content. This usage can effectively restrict the input of the user, reduce many unnecessary risks, but it also requires a lot of effort to train an excellent scenario.

[0060] In practical applications, the function keys in this embodiment can actually include keywords and the corresponding semantic explanations of the keywords. The semantic explanations can be oriented to the user, so that the user can more clearly know the semantics of the function key based on the semantic explanation. In addition, the semantic explanations can be oriented to the instructions of the document generation model in the subsequent process, so that the document generation model can more clearly know the needs of the user and generate a more accurate document.

[0061] In this embodiment, by setting up a private domain server and establishing a dedicated private domain document generation model for the private domain server, the private domain server can process private data according to industry characteristics, which is more in line with the extended thinking direction of industry characteristics; in addition, the processing effect of the large language model can be further improved, and any thoughts can be preferentially combined together to extract documents that conform to industry characteristics.

[0062] In this embodiment, the industry characteristics are aggregated by first presenting the reasoning process and then obtaining the final answer, and the documents within the industry are summarized, similar to adding an intermediate reasoning process, so that the model performance of the model in complex reasoning problems such as common sense reasoning and mathematical problems can be significantly improved.

[0063] In an exemplary embodiment, the private domain document generation model is established in the following manner:

[0064] Based on a large number of private domain documents, a corpus model library is established;

[0065] Based on the corpus model library, a private domain document style library is further established to obtain the private domain document generation model;

[0066] A learning framework is provided for the private domain document generation model, and an objective function is designed based on feature extraction, similarity measurement, iterative optimization algorithms, etc., and the private domain document generation model performs self-learning and updating under the learning framework;

[0067] Weights are automatically assigned to each part of the document.

[0068] In practical applications, the private domain documents of each customer can be learned first, and the AI is used to establish a corpus model library for the writing style of this document; in the second step, the documents that the user thinks are good can be focused on "feeding" to the AI, with the aim of establishing a document model to absorb the private domain document style library; in the third step, the AI can be allowed to further learn and change its own answers.

[0069] In implementation, the provided learning framework can be the structure of the document. Specifically, a message generally consists of four parts: a title, a lead, a body, and an ending. Beginning part of the document: Introduce the theme. Content part: Output value. Ending part: Summarize or sublimate the theme.

[0070] In practical applications, common leads include: the lead, the first paragraph of the message, or the first sentence, which mainly uses concise words to summarize the most important content in the text and hint at the theme of the article; the lead grasps the core of the event and attracts readers to continue reading; the following is an example to illustrate the common lead structure:

[0071] The flashback structure, also known as the "inverted pyramid" structure, that is, it comes straight to the point and has strong generalization. The most important fact or result is placed in the lead for statement;

[0072] The chronological style, also known as the "pyramid" structure or the chronological structure, is written in the order of the time when the events occur, which is convenient for a clear introduction and a general description of each stage of the event development;

[0073] The suspense style sets a mystery at the beginning, making readers eager to know the development and result of the event. Then the mystery is solved in the main body or at the end;

[0074] The visual style is written in an image-based and three-dimensional way, using typical details and vivid pictures to reflect and report news facts;

[0075] The prose style has no fixed form, and its content is extensive but not miscellaneous, similar to the free and lively prose structure. For example, at the beginning of the news, it briefly depicts the relevant scene, situation, atmosphere, color, or spontaneously expresses the personal experience and feeling; or triggers and mobilizes the readers' association to arouse their interest; or sets a suspense, etc. Then, it presents the news facts rhythmically and comprehensively.

[0076] In practical applications, the conclusion can also be called the ending. A well-written ending can deepen the theme, enhance the readers' understanding of the whole text, and increase the readability and appeal of the news. The following are examples of common conclusions:

[0077] The summary style summarizes the content of the news at the end, points out the significance of the news facts, or makes a generalization, once again points out the theme, and finds the common significance of many facts, leaving a whole impression on the readers. In terms of writing methods, it can either simply summarize the news facts or make appropriate comments on the facts;

[0078] The inspiring style makes a "finishing touch" at the end, making readers have a sense of "sudden enlightenment and awakening". This ending method often uses a lyrical tone or subtly implies the theme through the characters' language, deepens the theme, and helps readers deepen their understanding of the essence of the news;

[0079] The analytical style uses the analytical description method at the end, making people feel as if they are on the scene, seeing the people, and hearing the sounds, so as to enhance the appeal of the news. It is mainly used for some news themes that are mainly applicable to analytical descriptions and strive to reproduce the scene or rare wonders. Using this method at the end can enhance the credibility of the news and the vividness of the scene;

[0080] The prospective style is that after narrating the news facts, at the end of the news work, the author expresses hopes, suggestions, or warnings for the news facts or the characters in the news, but this expression should be euphemistic and appropriate, and easy for readers to accept;

[0081] Metaphorical, that is, using a metaphor in common parlance. Chinese is extremely rich, and the category of metaphors is very broad. For example, "the imperial sword" is used to metaphorize the instructions of superiors; "the setting sun" is used to metaphorize old age; "the spring silkworm" is used to metaphorize people's diligence; "the black gauze cap" is used to metaphorize people's official positions; "the sun" is used to metaphorize "light", etc. Using a metaphor as the ending is vivid and intuitive, enabling readers to draw inferences from one instance and triggering their associations and imaginations, thereby achieving the purpose of enhancing the allure of the report;

[0082] Suspenseful, which refers to the tense mood of readers regarding the development of news events and the fates of characters. This way of ending sets a "question" for readers at the end, and then quickly provides an answer, or sets a question to guide readers to think and find the answers themselves;

[0083] Natural, also known as "the inverted pyramid ending", that is, the factual materials are arranged in the order of decreasing importance. After all the factual elements of the news are presented, the full text already has a "natural progression" momentum. When a relatively complete concept or impression has been given to readers without the need for an additional ending paragraph or conclusion, the text can end. Most "inverted pyramid" dynamic news adopts this natural ending.

[0084] This embodiment also gives examples of the opening methods of some manuscripts, such as:

[0085] 1. Starting with a question

[0086] Starting with a question stimulates readers' thinking and curiosity by asking them questions, attracting them to read the answers given in the article. It is a very common opening method.

[0087] 2. Starting with a dialogue

[0088] Starting with a dialogue means presenting the events related to the article in the form of a dialogue at the beginning of the article, which has a sense of immersion.

[0089] 3. The ever - useful famous quotes

[0090] This method has always been used the most. Famous quotes have an authoritative effect, increasing the credibility of the article. Such an opening form is used more in knowledge - based, opinion - based, and experience - sharing articles.

[0091] 4. Quoting user comments

[0092] This is a highly interactive way of expression. Starting an article with user comments, basically the content of the article also revolves around solving users' problems.

[0093] 5. Being good at using current affairs hotspots

[0094] The beginning of a current affairs hot topic refers to opening with a reference to current hot events, hot figures, hot topics, popular film and television works, etc.

[0095] 6. Self-introduction

[0096] Self-introduction means that the author, from the first-person perspective, introduces to the reader "what kind of person they are", "what they have done", "what they want to do", "what they are doing", etc.

[0097] This embodiment also gives examples of the ending methods of some manuscripts, such as:

[0098] 1. Summarize the full text

[0099] This is a very common ending template that summarizes the theme or core idea of the article.

[0100] 2. End with a famous quote to add the finishing touch

[0101] Similar to the method of starting with a famous quote, it also ends with a famous quote that has been passed down for many years, which can also produce catchy lines.

[0102] 3. Issue an appeal

[0103] Issuing an appeal is a common way to trigger interaction between the reader and the author. For example, when soliciting an activity or expressing a view, the author hopes to see the reader's opinions.

[0104] 4. End with a parallelism

[0105] Parallelism is a very useful and versatile sentence pattern because its function is to align and rhyme, making it catchy.

[0106] In an exemplary embodiment, when the user's server type is a private domain server, the domain where the private domain server is located is determined, and the function keys include function keys related to the domain.

[0107] When the user's server type is a public domain server, the function keys include function keys for different domains.

[0108] Exemplarily, based on the function keys and prompts for combined thinking and reasoning, the language model's ability to perform multi-step reasoning enables it to go beyond simple pattern matching. The function keys + some large models are deployed to the edge side. Undoubtedly, behind the function keys (keywords) is your best source of inspiration. It allows users to have both data security and the advanced productivity of large models. Moving towards true understanding in manuscript generation.

[0109] In an exemplary embodiment, the manuscript generation model includes a public domain manuscript generation model and a private domain manuscript generation model;

[0110] The method for generating manuscripts oriented to user intentions further includes:

[0111] Based on each generated manuscript, the public-domain manuscript generation model is intensively trained; a hybrid training system can be used to combine self-supervised learning and supervised learning, which enables it to more deeply understand the text semantics and structure, and generate more accurate, fluent and logical content. Discover patterns from massive data and string the data together according to rules to form articles and content written by humans; in the intensive training session, by limiting the prompt words, embedding a large amount of domain-specific knowledge and fine-tuning techniques, the influence of malicious user input of wrong values can be reduced.

[0112] Train the private-domain manuscript generation model based on the public-domain manuscript generation model, so as to solve the industry characteristics and refine the manuscript of the private data set that conforms to the industry characteristics.

[0113] In an exemplary embodiment, training the private-domain manuscript generation model based on the public-domain manuscript generation model includes:

[0114] Determine the semantic data and syntactic data in the public-domain manuscript generation model, and determine the semantic data and syntactic data in the private-domain manuscript generation model;

[0115] If the semantic data and syntactic data in the public-domain manuscript generation model are inconsistent with the semantic data and syntactic data in the private-domain manuscript generation model, supplement the semantic data and syntactic data in the public-domain manuscript generation model to the private-domain manuscript generation model.

[0116] In an exemplary embodiment, it further includes:

[0117] For the supplemented private-domain manuscript generation model, build a model with private data. Since the service is independently deployed and private data will not be used for large model training, private data can be effectively protected. Generalize and make the experience of professional talents universal; allow users to create their own model versions according to specific needs. This service gives users more control, enabling them to customize the use priority of private data processing, semantic data and syntactic data according to their specific needs.

[0118] In an exemplary embodiment, inputting keywords and prompt words into the manuscript generation model corresponding to the server type for manuscript generation includes:

[0119] Automatically generate the main body of the manuscript through the manuscript generation model;

[0120] Automatically generate the title of the manuscript based on the main body of the manuscript.

[0121] There is a real demand from users to extract key information from a large amount of data, and it is a universal demand. For those with niche needs, it is not special or personalized enough. "Information screening" + "Content precipitation".

[0122] In an exemplary embodiment, automatically generating a title for a document, including:

[0123] Calculating the correlation weight value of each paragraph in the body of the document;

[0124] Generating a set of key paragraphs based on the correlation weight values between paragraphs;

[0125] Determining the keywords included in the document based on the set of key paragraphs;

[0126] Automatically generating a title for the document based on the keywords included in the document.

[0127] In practical applications, calculating the correlation weight value of each paragraph in the body of the document, including:

[0128] Splitting each paragraph into several complete sentences, and calculating the weight ratio of each sentence in the paragraph where it is located through the following formula (1);

[0129]

[0130] Among them, c represents a constant, d l represents the word frequency of the l-th word in the sentence, M represents the number of sentences in the paragraph, e represents the exponent, and L represents the number of words in the sentence.

[0131] According to the weight ratio of each sentence in the paragraph where it is located, screening the effective sentences of each paragraph, and using the weight ratios corresponding to all the effective sentences as the weight sequence of the paragraph;

[0132] Calculating the correlation weight value of each paragraph according to the weight sequence of each paragraph.

[0133] In implementation, screening the effective sentences of each paragraph, including:

[0134] Sorting the weight ratios of all the sentences in the paragraph from large to small, and taking the sentences corresponding to the first weight ratios in the sorting as the effective sentences; where max(·) represents the maximum value operation, A represents the number of paragraphs in the news document, int(·) represents the rounding operation, and ε represents a very small value.

[0135] In implementation, calculating the correlation weight value of each paragraph according to the weight sequence of each paragraph, including:

[0136] Calculating the correlation weight value f of the paragraph through the following formula (2):

[0137]

[0138] Among them, K n+1 represents the weight ratio of the (n + 1)-th effective sentence in the paragraph, and K n represents the weight ratio of the n-th effective sentence in the paragraph, N represents the number of effective sentences in the paragraph, and K max represents the maximum weight ratio of the paragraph, and K min represents the minimum weight ratio of the paragraph.

[0139] In an exemplary embodiment, the specific method for generating the set of key paragraphs is as follows: According to the associated weight values of each paragraph, calculate the association threshold of the news manuscript through the following formula (3), eliminate the paragraphs corresponding to the associated weight values less than the association threshold, and use the remaining paragraphs as the key paragraphs with title keywords to generate the set of key paragraphs:

[0140]

[0141] Among them, A represents the number of paragraphs in the news manuscript, and f a represents the associated weight value of the a-th paragraph.

[0142] In an exemplary embodiment, based on the keywords included in the manuscript, automatically generate the title of the manuscript, including:

[0143] Use the TextRank algorithm to extract the keywords of each key paragraph in the set of key paragraphs to generate a keyword set;

[0144] Extract the word vectors of each keyword in the keyword set and generate a keyword matrix;

[0145] Perform rounding operations on the values of the keyword matrix as the title length of the news manuscript;

[0146] Within the title length of the news manuscript, generate the title of the news manuscript based on the keywords of each paragraph using a support vector machine.

[0147] Among them, the specific method for generating the keyword matrix is as follows:

[0148] Use the number of paragraphs as the number of rows of the keyword matrix, and use the maximum value of the number of keywords in all paragraphs as the number of columns of the keyword matrix; Calculate the matrix label values of each keyword according to the word vectors of each keyword, arrange the matrix label values of each keyword in each paragraph from small to large as the elements of each row of the keyword matrix, and fill the empty spaces with 1 to generate the keyword matrix.

[0149] In implementation, the keyword matrix generates label values through the following formula (4):

[0150]

[0151] Among them, g represents the character length of the keyword, and X represents the word vector of the keyword.

[0152] In this embodiment, when generating a title, a support vector machine can be used to classify and predict the probability of the next word or phrase appearing. Specifically, an n-gram language model can be used to represent the text sequence, and the next word or phrase can be predicted by maximum likelihood estimation or conditional probability; then the title is generated based on the prediction result.

[0153] Based on the word vector of the keyword, the keyword matrix corresponding to the entire news manuscript is generated. The value of the keyword matrix generated can be used to constrain the length of the news manuscript title to prevent the news manuscript title from being too long and causing the key points to be unclear. In addition, in order to prevent the value of the keyword matrix from being an integer, the value of the keyword matrix needs to be rounded.

[0154] The document generation method oriented to user intention provided by this embodiment is described below in conjunction with specific embodiments.

[0155] Take news releases as an example. First, select the news genre settings (four categories: general, news, communication, new media), and according to the template: select the type, structure, introduction (five categories: flashback, chronological, suspense, visual, prose), ending (six categories: summary, inspiration, analysis, metaphor, suspense, natural), including five elements: time, place, people, events and causes. (Title + introduction (reveal the report theme with the language of news characters) - body (explain, deepen and concretize the introduction) - ending (emphasize and deepen the theme)) to obtain the content of the manuscript. The news usually consists of five parts: title, news head, introduction, theme, background, and ending. Before the introduction of the news, there are often words such as "This newspaper", "This station news", "This newspaper (×× agency) × place × month × day", which is the news head. The main content is "At 19:00 on the evening of October 14, Professor X personally gave a special lecture on SIT. Professor X started this lecture from many aspects. First, he briefly explained to the students what SIT is and why the university started such an innovative project for college students. Professor X emphasized that college students do not need to have all the professional knowledge, but should learn by doing, explore by themselves, and innovate by themselves. SIT aims to cultivate college students' potential for knowledge application, analysis and problem solving, and teamwork. When talking about how to get started with the SIT project, Professor X said... After more than an hour, the lecture ended smoothly, and the students benefited a lot. After the lecture, Professor X said in an interview: We are not cultivating nerds, but talents with all-round development. All students can participate in SIT or subject competitions to improve their comprehensive quality potential and innovation potential.

[0156] First, the manuscript content is divided into two paragraphs, namely "At 19:00 on the evening of October 14, Professor X... analyzed the potential for problem solving and teamwork" and "When it comes to how to get started with the SIT project... improve one's own comprehensive quality potential and innovation potential"; generate correlation weight values ​​for the two paragraphs.

[0157] Then, according to the association weight values ​​of the two paragraphs, it is determined that both paragraphs have title function keys, and the function key set is specifically "SIT training", "SIT introduction" and "lecture".

[0158] Finally, the title of the news article generated by support vector machine is "Lecture on SIT Training and SIT Introduction".

[0159] Figure 2 It is a structural schematic diagram of a document generation device oriented to user intention provided by an embodiment of the present invention.

[0160] like Figure 2 As shown, the document generation device oriented to user intention provided in this embodiment includes:

[0161] A determination module 201 is used to determine a server type of a user, where the server type includes a public domain server and a private domain server;

[0162] A function key selection module 202, for providing a user with corresponding function keys based on the user's server type;

[0163] The collection module 203 is used to collect the function keys selected by the user and the prompt words input;

[0164] Generation module 204 is used to input function keys and prompt words into a document generation model corresponding to the server type. The document generation model generates a document based on thought combination reasoning based on the function keys and prompt words. The document generation model includes a public domain document generation model and a private domain document generation model. When the user is a public domain server, the public domain document generation model is used to generate the document. When the user is a private domain server, the private domain document generation model is used to generate the document.

[0165] Function keys + some large models are decentralized to the end side. The function keys (keywords) are undoubtedly your best source of inspiration. It allows users to have both data security and advanced productivity of large models. It relies on self-learning capabilities to intelligently reorganize massive amounts of knowledge; and it communicates and interacts with people based on a massive corpus database to complete the task of writing text generation.

[0166] The following is a description of a document generation system oriented toward user intention provided by the present invention. The document generation system oriented toward user intention described below and the document generation method oriented toward user intention described above can be referenced to each other.

[0167] Figure 3 It is a schematic structural diagram of a document generation system for user intent provided by an embodiment of the present invention.

[0168] As Figure 3 shown, the document generation system for user intent includes:

[0169] A control module for executing the document generation method of any of the above embodiments;

[0170] A server for storing documents, and the server includes a public domain server and a private domain server;

[0171] A login module electrically connected to the server for authenticating the user's identity and determining the user's server type.

[0172] In practical applications, the document generation system provided by this embodiment has the following functions: language understanding, text generation, content expansion, text polishing, plot sorting, resuming after discontinuation, logical reasoning, title generation, typo correction, creativity stimulation, thinking control, and mathematical ability, etc.

[0173] In implementation, the public domain server is used for providing public document generation storage services, and the private domain server is used for providing private document generation services;

[0174] It further includes a document generation module. The document generation module is used for generating documents, which is the core "creation" of generative AI. By learning elements from data, it can generate brand-new and original documents. It can not only implement the analysis, judgment, and decision-making functions of traditional AI, but also implement creative functions that traditional AI cannot achieve.

[0175] In implementation, the document generation module includes a main text generation unit and a title generation unit. The main text generation unit is electrically connected to the upload module, the verification unit, the public domain server, and the private domain server respectively. When the verification unit determines that the user to be served is allowed to use the private domain server, the main text generation unit generates documents based on the document model of the private domain server for the text information uploaded by the user to be served; the main text generation unit is also used when the verification unit determines that the user to be served is only allowed to use the public domain server, and the main text generation unit generates documents based on the document model of the public domain server for the text information uploaded by the user to be served;

[0176] The title generation unit is electrically connected to the body text generation unit, the private domain server, and the public domain server. When the body text generation unit generates a manuscript based on the manuscript model of the private domain server for the text information uploaded by the user, the title generation unit generates a title for the manuscript generated by the body text generation unit according to the title model in the private domain server; the title generation unit is also used when the body text generation unit generates a manuscript for the text information uploaded by the user according to the manuscript model of the public domain server, and the title generation unit generates a title for the manuscript generated by the body text generation unit according to the title model in the public domain server.

[0177] In the manuscript generation module, by setting the title keywords + function keys and network inference to enhance the semantic understanding ability, it is possible to perform fine-grained control of the user's personal ideas, apply advanced idea transformation, and combine the most promising ideas in the ongoing inference into new ideas.

[0178] In an exemplary embodiment, the login module includes:

[0179] A login unit for obtaining the user's identity information;

[0180] A verification unit is electrically connected to the login module and the distributed server respectively. The verification unit is used to judge whether to allow the user to log in according to the relationship between the user's identity information and the retained information in the distributed server. Among them,

[0181] When the user's identity information is consistent with the retained information in the distributed server, the verification unit judges that the user fails the identity verification and does not allow the user to log in;

[0182] When the user's identity information is consistent with the retained information in the distributed server, the verification unit judges that the user passes the identity verification and determines whether the public domain server or the private domain server is used by the user according to the relationship between the user's identity information and the retained information stored in the public domain server and several private domain servers in the distributed server.

[0183] In an exemplary embodiment, the verification unit is also used to obtain the retained information stored in several private domain servers and judge whether the public domain server or the private domain server is used by the user according to the relationship between the user's identity information and the retained information stored in several private domain servers. Among them,

[0184] When the user's identity information is consistent with the retained information stored in several private domain servers, the verification unit allows the user to use the private domain server whose stored retained information is consistent with the identity information of the user to be used;

[0185] When the identity information of the user to be used is inconsistent with the retained information stored in several private domain servers, the verification unit only allows the user to be used to use the public domain server.

[0186] In an exemplary embodiment, the text generation unit includes a calling subunit, which is electrically connected to the verification unit. The calling subunit is configured to call the manuscript model stored in the corresponding server according to the result of the verification unit's judgment on whether the user to be used is a public domain server or a private domain server. Among them,

[0187] When the verification unit determines that the user to be used is only allowed to use the public domain server, the calling subunit selects to call the manuscript model in the public domain server;

[0188] When the verification unit determines that the user to be used is allowed to use the private domain server, the calling subunit selects to call the manuscript model in the private domain server;

[0189] An analysis subunit, which is electrically connected to the upload module. The analysis subunit is configured to analyze the semantic data and syntactic data in the text information uploaded by the user to be used;

[0190] A text generation subunit, which is electrically connected to the analysis subunit and the calling subunit. When the calling subunit selects to call the manuscript model in the public domain server, the text generation subunit is configured to match the semantic data and syntactic data in the text information uploaded by the user to be used with the semantic data and syntactic data in the manuscript model in the public domain server, and generate a manuscript according to the word data of the semantic data and syntactic data in the manuscript model in the public domain server that match; The text generation subunit is further configured to, when the calling subunit selects to call the manuscript model in the private domain server, the text generation subunit matches the semantic data and syntactic data in the text information uploaded by the user to be used with the semantic data and syntactic data in the manuscript model in the private domain server, and generates a manuscript according to the word data of the semantic data and syntactic data in the manuscript model in the private domain server that match.

[0191] In implementation, the text generation unit further includes:

[0192] A text instruction subunit, which is provided in several numbers, and the several text instruction subunits are electrically connected to the text generation subunit, so that when the text generation subunit generates text according to the word data of the semantic data and syntactic data in the text model in the private domain server or the public domain server, the text instruction subunit is configured to control the text generation subunit.

[0193] In an exemplary embodiment, a model training unit is further included. The model training unit includes:

[0194] A model parameter acquisition subunit, which is electrically connected to the text generation subunit and the public domain server respectively. The model parameter acquisition subunit is configured to acquire the technical parameters of reinforcement learning for the manuscript model based on human feedback and the manuscript model parameters after pre-training of the preference model;

[0195] The model training subunit is electrically connected to the model parameter acquisition subunit and the private domain server. The model training subunit is used to learn the basic language rules and patterns in the text data according to the technical parameters for reinforcement learning of the manuscript model based on human feedback and the manuscript model parameters after pre-training of the preference model obtained by the model parameter acquisition subunit, and to adjust the manuscript model in the private domain server according to the learned model parameters.

[0196] In implementation, the model training subunit is also used to obtain the real-time word data of semantic data and syntactic data in the manuscript model of the public domain server. The model training subunit is also used to determine whether to adjust the word data of semantic data and syntactic data in the manuscript model of the private domain server according to the relationship between the real-time word data of semantic data and syntactic data in the manuscript model of the public domain server and the word data of semantic data and syntactic data in the manuscript model of the private domain server. Among them,

[0197] When the word data of the real-time word data of semantic data and syntactic data in the manuscript model of the public domain server is consistent with the word data of the semantic data and syntactic data in the manuscript model of the private domain server, the model training subunit does not adjust the word data of the semantic data and syntactic data in the manuscript model of the private domain server;

[0198] When the word data of the real-time word data of semantic data and syntactic data in the manuscript model of the public domain server is inconsistent with the word data of the semantic data and syntactic data in the manuscript model of the private domain server, the model training subunit supplements the extra word data of the real-time word data of semantic data and syntactic data in the manuscript model of the public domain server into the word data of the semantic data and syntactic data in the manuscript model of the private domain server.

[0199] In practical applications, the model training subunit is also used to obtain the real-time word data of semantic data and syntactic data in the manuscript model generated by the main text generation subunit. The model training subunit is also used to adjust the word data of the semantic data and syntactic data in the manuscript model of the supplemented private domain server according to the relationship between the real-time word data of semantic data and syntactic data in the manuscript model generated by the main text generation subunit and the word data of the semantic data and syntactic data in the manuscript model of the supplemented private domain server. Among them,

[0200] When the real-time word data of semantic data and syntactic data in the manuscript model generated by the main text generation subunit is consistent with the word data of semantic data and syntactic data in the manuscript model in the supplemented private domain server, the model training subunit is further configured to adjust the priority usage ranking of each word in the word data of semantic data and syntactic data in the manuscript model in the supplemented private domain server according to the usage ranking of each word in the real-time word data of semantic data and syntactic data in the manuscript model generated by the main text generation subunit;

[0201] When the real-time word data of semantic data and syntactic data in the manuscript model generated by the main text generation subunit is inconsistent with the word data of semantic data and syntactic data in the manuscript model in the supplemented private domain server, the model training subunit is further configured to perform secondary supplementation according to the missing word data in the word data of semantic data and syntactic data in the manuscript model in the supplemented private domain server, and perform intelligent integration on a large amount of word data according to the self-learning ability. And according to the usage ranking of each word in the real-time word data of semantic data and syntactic data in the manuscript model generated by the main text generation subunit, adjust the priority usage ranking of each word in the word data of semantic data and syntactic data in the manuscript model in the supplemented private domain server after the secondary supplementation.

[0202] In the exemplary embodiment, the title generation unit includes:

[0203] A title model acquisition subunit, electrically connected to the manuscript generation subunit. When the manuscript generation subunit generates a manuscript based on the manuscript model of the private domain server for the text information uploaded by the user to be processed, the title model acquisition subunit acquires the title model in the private domain server; when the manuscript generation subunit generates a manuscript based on the manuscript model of the public domain server for the text information uploaded by the user to be processed, the title model acquisition subunit acquires the title model in the public domain server;

[0204] A title generation subunit, electrically connected to the title model acquisition subunit. When the title model acquisition subunit acquires the title model in the private domain server, the title generation subunit is configured to generate a title for the manuscript generated by the manuscript generation subunit according to the title model in the private domain server; the title generation subunit is further configured to, when the title model acquisition subunit acquires the title model in the public domain server, generate a title for the manuscript generated by the manuscript generation subunit according to the title model in the public domain server.

[0205] The specific implementation method of the manuscript generation system provided in this embodiment can be implemented with reference to the above embodiments, and will not be elaborated here.

[0206] Figure 4Illustrates a schematic diagram of the physical structure of an electronic device, as follows Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call logical instructions in the memory 430 to execute a document generation method, which includes:

[0207] Determine the server type of the user. The server type includes a public domain server and a private domain server;

[0208] Based on the server type of the user, provide a corresponding function key for the user;

[0209] Collect the function key selected by the user and the input prompt word;

[0210] Input the function key and the prompt word into a document generation model corresponding to the server type for document generation. The document generation model includes a public domain document generation model and a private domain document generation model. When the user is a public domain server, use the public domain document generation model for document generation. When the user is a private domain server, use the private domain document generation model for document generation.

[0211] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.

[0212] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the document generation method provided by the above-mentioned various methods. The method includes:

[0213] Determine the server type of the user, where the server type includes a public domain server and a private domain server;

[0214] Based on the user's server type, provide the corresponding function keys for the user;

[0215] Collect the function keys selected by the user and the input prompt words;

[0216] Input the function keys and prompt words into the document generation model corresponding to the server type for document generation. The document generation model includes a public domain document generation model and a private domain document generation model. When the user is a public domain server, use the public domain document generation model for document generation. When the user is a private domain server, use the private domain document generation model for document generation.

[0217] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is used to execute the document generation method provided by the above-mentioned various methods. The method includes:

[0218] Determine the server type of the user, where the server type includes a public domain server and a private domain server;

[0219] Based on the user's server type, provide the corresponding function keys for the user;

[0220] Collect the function keys selected by the user and the input prompt words;

[0221] Input the function keys and prompt words into the document generation model corresponding to the server type for document generation. The document generation model includes a public domain document generation model and a private domain document generation model. When the user is a public domain server, use the public domain document generation model for document generation. When the user is a private domain server, use the private domain document generation model for document generation.

[0222] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0223] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a document oriented to user intent, characterized in that, Including: Determine the server type of the user, where the server type includes a public domain server and a private domain server; Based on the server type of the user, provide corresponding function keys for the user; Collect the function keys selected by the user and the input prompt words; Input the function keys and the prompt words into the manuscript generation model corresponding to the server type for manuscript generation. The manuscript generation model includes a public domain manuscript generation model and a private domain manuscript generation model. When the user is a public domain server, use the public domain manuscript generation model for manuscript generation. When the user is a private domain server, use the private domain manuscript generation model for manuscript generation; Based on each generated manuscript, perform reinforcement training on the public domain manuscript generation model; Train the private domain manuscript generation model based on the public domain manuscript generation model; The training of the private domain manuscript generation model based on the public domain manuscript generation model includes: Determine the semantic data and syntactic data in the public domain manuscript generation model, and determine the semantic data and syntactic data in the private domain manuscript generation model; If the semantic data and syntactic data in the public domain manuscript generation model are inconsistent with the semantic data and syntactic data in the private domain manuscript generation model, supplement the semantic data and syntactic data in the public domain manuscript generation model to the private domain manuscript generation model; Adjust the usage priority of the semantic data and syntactic data in the supplemented private domain manuscript generation model.

2. The method for generating a document oriented to user intent according to claim 1, characterized in that, The private domain manuscript generation model is established in the following manner: Based on a large number of private domain manuscripts, establish a corpus model library; Further establish a private domain manuscript style library based on the corpus model library to obtain the private domain manuscript generation model; Provide a learning framework for the private domain manuscript generation model, and the private domain manuscript generation model performs self-learning and updating under the learning framework; Automatically assign weights to each part of the manuscript.

3. The method for generating a document oriented to user intent according to claim 1, characterized in that, When the server type of the user is a private domain server, determine the field where the private domain server is located, and the function keys include function keys related to the field; When the server type of the user is a public domain server, the function keys include function keys in different fields.

4. The method for generating a document oriented to user intent according to claim 1, characterized in that, The inputting the function keys and the prompt words into the manuscript generation model corresponding to the server type for manuscript generation includes: Through the manuscript generation model, automatically generate the body text of the manuscript; Based on the body text of the manuscript, automatically generate the title of the manuscript.

5. The method for generating a document oriented to user intent according to claim 4, characterized in that, The automatically generating the title of the manuscript includes: Calculate the associated weight values of each paragraph of the body text of the manuscript; Based on the associated weight values between the paragraphs, generate a set of key paragraphs; Based on the set of key paragraphs, determine the function keys included in the manuscript; Based on the function keys included in the manuscript, automatically generate the title of the manuscript.

6. A device for generating a document oriented to user intent, applied to the method for generating a document oriented to user intent according to any one of claims 1-5, characterized in that, Including: A determination module for determining the server type of the user, where the server type includes a public domain server and a private domain server; A function key selection module for providing corresponding function keys for the user based on the server type of the user; A collection module for collecting the function keys selected by the user and the input prompt words; A generation module, configured to input the function key and the prompt word into a document generation model corresponding to the server type. The document generation model performs thinking combination reasoning for document generation based on the function key and the prompt word. The document generation model includes a public domain document generation model and a private domain document generation model. When the user is a public domain server, the public domain document generation model is used for document generation. When the user is a private domain server, the private domain document generation model is used for document generation.

7. A manuscript generation system oriented to user intent, characterized in that Comprising: A control module, configured to execute the document generation method for user intent according to any one of claims 1-5; A server, configured to store documents. The server includes a public domain server and a private domain server; A login module, electrically connected to the server, configured to verify the identity of the user and determine the server type of the user.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the manuscript generation method oriented to user intent according to any one of claims 1-5 above.

Citation Information

Patent Citations

  • Powerpoint generation method and device, computer readable storage medium and server

    CN114398883A

  • Fine adjustment method and device for large model in banking business, equipment and storage medium

    CN116579402A