Local position recognition model training method and device, local position recognition model reasoning method and device, and local position recognition model equipment
By training a local position recognition model to identify and modify the local positions of text generated by a large generative model, the problem of poor user experience in existing technologies is solved, and efficient local text modification and computing resource conservation are achieved.
Patent Information
- Application Number
- CN202510830548.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
AI Technical Summary
When users need to make partial modifications to text generated by a large generative model, existing technologies usually require the entire text to be regenerated or manually performed, resulting in a poor user experience.
A training method for a local location recognition model is provided. By obtaining a target sample set and a modification request, a target training sample is constructed, and an initial local location recognition model is trained to recognize and modify local locations in an article.
It simplifies the model training process, improves location recognition efficiency and user experience, reduces computing resource requirements, and enables efficient local text modification.
Smart Images

Figure CN120744490A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to the technical fields of artificial intelligence, large models, and big data. Background Art
[0002] With the development of large model technology, generative large language model (G-LLM) text generation technology is becoming increasingly mature. In real-world scenarios, the following situations often occur: users need to make partial modifications to the generated text. In this case, existing technologies usually regenerate the entire text, or the user manually determines the local location to be modified and then calls the large model to regenerate the content. Obviously, this process reduces the user experience. Summary of the Invention
[0003] The present disclosure provides a training method, an inference method, and an apparatus and device for a local location recognition model.
[0004] According to one aspect of the present disclosure, a method for training a local position recognition model is provided, comprising:
[0005] Acquire a target sample set, wherein the target sample set includes: a plurality of target sample articles, and a target position indicated by each of a plurality of preset quantifiers in each target sample article;
[0006] Determine a plurality of first modification requests; wherein the first modification requests include a target quantifier, and the first modification requests are used to instruct to modify the text content at the corresponding position through the included target quantifier;
[0007] Based on the target positions indicated by the preset quantifiers in the target sample articles, determining a target position that can match the target quantifier included in the first modification request;
[0008] Based on the target sample article, the first modification request, and the target position matching the target quantifier contained in the first modification request, a target training sample is constructed to train an initial local position recognition model based on the target training sample, and a target local position recognition model is obtained.
[0009] According to another aspect of the present disclosure, a training device for a local position recognition model is provided, comprising:
[0010] A sample determination unit is configured to obtain a target sample set, wherein the target sample set includes: a plurality of target sample articles, and a target position indicated by each of a plurality of preset quantifiers in each target sample article; determine a plurality of first modification requests; wherein the first modification request includes a target quantifier, and the first modification request is configured to instruct to modify text content at a corresponding position using the target quantifier included therein; and determine, based on the target position indicated by each of the preset quantifiers in each target sample article, a target position that can match the target quantifier included in the first modification request;
[0011] The model training unit is used to construct a target training sample based on the target sample article, the first modification request, and the target position matching the target quantifier contained in the first modification request, so as to train the initial local position recognition model based on the target training sample and obtain the target local position recognition model.
[0012] According to another aspect of the present disclosure, a local position identification method is provided, comprising:
[0013] Obtaining a target modification request for an initial article, wherein the target modification request is used to indicate the required modification content through a quantifier; wherein the initial article is generated by calling a large model;
[0014] Inputting the initial article and the target modification request into a target local position recognition model to obtain the specific position of the modification content requested by the target modification request in the initial article;
[0015] Among them, the model parameter amount of the target local position recognition model is smaller than the model parameter amount of the large model.
[0016] According to another aspect of the present disclosure, there is provided a local position identification device, comprising:
[0017] a request acquisition unit, configured to acquire a target modification request for an initial article, wherein the target modification request is used to indicate the required modification content by using a quantifier; wherein the initial article is generated by calling a large model;
[0018] A model inference unit is configured to input the initial article and the target modification request into a target local position recognition model to obtain a specific position of the modification content requested by the target modification request in the initial article;
[0019] Among them, the model parameter amount of the target local position recognition model is smaller than the model parameter amount of the large model.
[0020] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0021] at least one processor; and
[0022] a memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.
[0024] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0025] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.
[0026] In this way, the disclosed solution can perform model training on the initial local position recognition model based on the target training sample of "target sample article-first modification request-target position matching the target quantifier contained in the first modification request", and then obtain a model that can effectively identify local positions in the article. The training process is simple and efficient, providing strong support for the efficient implementation of local modification of article content, and at the same time, laying the foundation for improving user experience.
[0027] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0029] Figure 1 This is a schematic flow chart of a method for training a local position recognition model according to an embodiment of the present application. Figure 1 ;
[0030] Figure 2 This is a schematic flow chart of a method for training a local position recognition model according to an embodiment of the present application. Figure 2 ;
[0031] Figure 3 This is a schematic flow chart of a method for training a local position recognition model according to an embodiment of the present application. Figure 3 ;
[0032] Figure 4 is a schematic flow chart of a local position identification method according to an embodiment of the present application;
[0033] Figure 5 1 is a schematic structural diagram of a training device for a local position recognition model according to an embodiment of the present application;
[0034] Figure 6 is a structural diagram of a local position identification device according to an embodiment of the present application;
[0035] Figure 7 It is a block diagram of an electronic device used to implement the training method or reasoning method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0036] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0037] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.
[0038] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0039] The disclosed solution provides a training method for a local position recognition model. The training method can first determine a modification request, then match the determined modification request with the target position indicated by each preset quantifier in each target sample article. In other words, the target position corresponding to the modification request is determined, and then the target position that can match the quantifier contained in the modification request is determined. Finally, according to the triple of "target sample article (which can be recorded as article)-modification request (which can be recorded as query)-target position (position) matching the quantifier contained in the modification request" (which can be recorded as<article,query,position> ) to construct a target training sample, so as to train the initial local position recognition model based on the target training sample, and then obtain the target local position recognition model. In the above process, since the disclosed solution can first obtain a modification request, and then use the quantifier contained in the modification request to match the target position that matches the quantifier, compared with the method of "first determining the article and then obtaining a modification request for the article", when preparing the same amount of training samples, the data processing effect of the disclosed solution is higher, thereby effectively improving the efficiency of model training, and thus laying the foundation for improving the model's position recognition ability in text content.
[0040] Specifically, Figure 1 This is a schematic flow chart of a method for training a local position recognition model according to an embodiment of the present application. Figure 1 The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0041] Furthermore, the method includes at least part of the following contents. Figure 1 Shown, including:
[0042] Step S101: Acquire a target sample set.
[0043] Here, the target sample set includes: a plurality of target sample articles, and a target position indicated by each preset quantifier in each target sample article among a plurality of preset quantifiers.
[0044] Furthermore, in one example, the preset quantifiers may be conventional quantifiers used in an article, such as "chapter" quantifiers (e.g., Chapter 1, Chapter 1, ..., Chapter 5, etc.), "section" quantifiers (e.g., Section 1, Section 2, ..., Section 5, etc.), "paragraph" quantifiers (e.g., Paragraph 1, Paragraph 2, ..., Paragraph 5, etc.), etc. Alternatively, in another example, the preset quantifiers may be user-defined quantifiers. The present disclosure does not impose any restrictions on the setting of preset quantifiers. In other words, the present disclosure may set preset quantifiers based on different usage scenarios, thereby improving the flexibility and applicability of the present disclosure.
[0045] Furthermore, in one example, the target position indicated by the preset quantifier in the target sample article may be specifically: the position of the text content corresponding to the preset quantifier in the target sample article; further, the target position may specifically include the starting position (which can be recorded as start) and the ending position (which can be recorded as end) of the text content corresponding to the preset quantifier. In this case, the target position can also be recorded as article[start:end].
[0046] Step S102: Determine multiple first modification requests.
[0047] Here, the first modification request includes a target quantifier, and the first modification request is used to instruct to modify the text content at the corresponding position through the included target quantifier.
[0048] Furthermore, in one example, the multiple first modification requests may specifically include at least one of the following types of requests: modification requests (e.g., "modify the first paragraph," "polish the second paragraph," etc.), deletion requests (e.g., "delete the third paragraph," etc.), and supplement requests (e.g., "add a paragraph after the third paragraph"). The disclosed solution does not impose specific restrictions on modification requests, and they can be adaptively adjusted based on actual scenario requirements.
[0049] Moreover, since the multiple first modification requests determined by the disclosed solution have many request types and rich data volume, when training a model based on these requests, the generalization ability of the model can be effectively improved.
[0050] Furthermore, in one example, the target quantifier is one of the plurality of preset quantifiers; this facilitates subsequent matching to obtain the target position corresponding to the first modification request.
[0051] Step S103: Based on the target positions indicated by the preset quantifiers in the target sample articles, a target position that can match the target quantifier included in the first modification request is determined.
[0052] That is to say, in this step, the target position to which the target quantifier in the first modification request can correspond can be effectively determined; it should be noted that there are one or more target positions that can match the first modification request. For example, the first modification request is specifically "modify the first chapter". At this time, the target position that can match the first modification request is the position where the first chapter is located. Based on this, the target position that matches the "modify the first chapter" can specifically include: the position where the first chapter of the target sample article 1 is located, the position where the first chapter of the target sample article 2 is located, the position where the first chapter of the target sample article 3 is located, etc. In other words, the target sample articles that have "Chapter 1" in the target sample set all have the target position corresponding to "Modify the First Chapter", that is, they can all be used as the target position matched by "Modify the First Chapter". In this way, it provides strong support for the subsequent rapid acquisition of training triples.
[0053] Step S104: construct a target training sample based on the target sample article, the first modification request, and the target position matching the target quantifier contained in the first modification request, so as to train an initial local position recognition model based on the target training sample and obtain a target local position recognition model.
[0054] In a specific example, after constructing the target training sample, the target sample article and the first modification request in the target training sample can be input into the initial local position recognition model to obtain the position of the model output, and then the loss value is calculated based on the target position in the target training sample that matches the target quantifier contained in the first modification request and the position output by the model to fine-tune some adjustable parameters in the initial local position recognition model to obtain a target local position recognition model that meets the training requirements.
[0055] In this way, the disclosed solution can perform model training on the initial local position recognition model based on the target training sample of "target sample article-first modification request-target position matching the target quantifier contained in the first modification request", and then obtain a model that can effectively identify local positions in the article. The training process is simple and efficient, providing strong support for the efficient implementation of local modification of article content, and at the same time, laying the foundation for improving user experience.
[0056] Furthermore, since the disclosed solution can first obtain a modification request and then use the quantifier contained in the modification request to match the target position that matches the quantifier, compared to the method of "first determining the article and then obtaining a modification request for the article", the processing resources required for "matching" in the disclosed solution are lower. Moreover, when the same amount of training samples need to be prepared, the processing efficiency of the disclosed solution is higher, and the data quality of the samples is also higher, thereby effectively improving the training efficiency and training effect of the model.
[0057] Furthermore, since the disclosed solution first obtains the target position that matches the target quantifier contained in the first modification request based on the target quantifier contained in the first modification request during the process of constructing the target training sample, compared with other sample construction methods (for example, calling a large model and obtaining an "article with a position annotation" based on the "article-modification request"), the disclosed solution does not need to rely on the article (that is, there is no need to determine the article first and then obtain the modification request based on the article), and can obtain high-quality target training samples at low cost, which is more controllable and provides strong support for the subsequent improvement of the generalization ability of the model.
[0058] Here, in actual applications, given preset quantifiers, a small number of modification requests, and specific locations that match the quantifiers contained in the modification requests, the disclosed solution can automatically iterate a local location recognition model that meets user needs, thereby enabling rapid adaptation in vertical fields.
[0059] Furthermore, in a specific example, the target local location recognition model can output the specific location of the modification content requested by the target modification request in the initial article based on the target modification request and the initial article generated by the large model. Here, the target modification request is used to indicate the required modification content through quantifiers.
[0060] Furthermore, the model parameter amount of the target local position recognition model is smaller than that of the large model, thus providing strong support for efficiently identifying the local position and laying the foundation for improving user experience.
[0061] It should be noted that since the model parameters of the target local position recognition model in the disclosed solution are smaller than the model parameters of the large model, in other words, compared with the large model, the target local position recognition model of the disclosed solution can be understood as a "lightweight model". At this time, the disclosed solution locates the position of the text content based on the "lightweight model", and the computing resources required are less. In other words, compared with the solution of calling the large model to locate the position of the text content, the disclosed solution has lower cost and higher efficiency for position recognition while ensuring the position recognition capability.
[0062] In a specific example, the target position indicated by the preset quantifier in the target sample article is represented by a regular expression. In this case, the target position can be recorded as article_start_end_regrex. Here, the amount of data occupied by the regular expression is smaller than the amount of data of the text content indicated by the preset quantifier.
[0063] That is to say, in the process of training the initial local position recognition model, the target position indicated by the preset quantifier input to the initial local position recognition model in the target sample article is a target position based on a regular expression. In this way, compared with the method of directly inputting the text content indicated by the preset quantifier into the model for training, the disclosed solution reduces the model's processing capacity for input data. In other words, the training resources required for model training in the disclosed solution are updated, for example, the number of tokens required is smaller, and the training efficiency is higher. In this way, the subsequent efficient acquisition of local positions provides strong support, and thus also lays the foundation for improving user experience.
[0064] Figure 2 This is a schematic flow chart of a method for training a local position recognition model according to an embodiment of the present application. Figure 2 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 1 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.
[0065] Furthermore, the method includes at least part of the following contents. Figure 2 Shown, including:
[0066] Step S201: Based on a plurality of preset quantifiers, each of the N1 first initial articles is disassembled to obtain a first position indicated by each preset quantifier in each first initial article.
[0067] Step S202: Based on each first initial article and the first position indicated by each preset quantifier in each first initial article, and calling the large model for sample amplification, N2 second initial articles and the second position indicated by each preset quantifier in each second initial article are obtained.
[0068] Here, N1 and N2 are both positive integers, and N2 is greater than N1. In other words, the disclosed solution can generalize samples based on the first initial articles that have been annotated, thereby enriching the sample size, thereby providing strong support for subsequent improvements in the model's training efficiency, training effect, and generalization ability.
[0069] Step S203: constructing the target sample set based on each first initial article and the first position indicated by each preset quantifier in each first initial article, and each second initial article and the second position indicated by each preset quantifier in each second initial article.
[0070] Here, the target sample set includes: a plurality of target sample articles, and a target position indicated by each preset quantifier in each target sample article among a plurality of preset quantifiers.
[0071] Furthermore, the multiple target sample articles include the first initial article mentioned above, and a second initial article obtained by generalizing the annotation result of the first initial article.
[0072] For example, in one example, first, the first initial article is disassembled according to multiple preset quantifiers, such as common quantifiers such as chapter and section, to obtain the position indicated by each chapter (such as the first chapter, the second chapter, etc.) in the first initial article, the position indicated by each section (such as the first section, the second section, etc.) in the first initial article, etc.; secondly, the large model is called, and based on the annotation results described above, that is, "the first initial article-the position indicated by each preset quantifier in the first initial article", the reasoning ability of the large model (such as few-shot learning) is called to perform sample amplification to obtain new samples, and then construct a target sample set with rich sample size and sufficient sample diversity.
[0073] Furthermore, in one example, after calling the large model sample to amplify the new sample, a rule strategy can be preset to verify the new sample generated by the large model (that is, each second initial article, and the second position indicated by each preset quantifier in each second initial article) to screen out samples that meet the requirements, thereby effectively improving the data quality of the sample.
[0074] Furthermore, in one example, the "first position" obtained in step S201 can be expressed by a regular expression. In this way, compared with the solution that directly relies on specific content for sample generalization, the disclosed solution requires fewer resources when calling a large model for sample generalization, for example, fewer tokens are required and the efficiency is higher.
[0075] Step S204: Determine multiple first modification requests.
[0076] Here, the first modification request includes a target quantifier, and the first modification request is used to instruct to modify the text content at the corresponding position through the included target quantifier.
[0077] Here, for relevant examples of the first modification request, please refer to the above statement and will not be repeated here.
[0078] Step S205: Based on the target positions indicated by the preset quantifiers in the target sample articles, a target position that can match the target quantifier included in the first modification request is determined.
[0079] Step S206: construct a target training sample based on the target sample article, the first modification request, and the target position matching the target quantifier included in the first modification request, so as to train an initial local position recognition model based on the target training sample and obtain a target local position recognition model.
[0080] Here, for relevant instructions on the first modification request, please refer to the above instructions and will not be repeated here.
[0081] In this way, the disclosed solution can call a large model and perform sample amplification based on a small number of pre-obtained samples of "each first initial article, the first position indicated by each preset quantifier in each first initial article" to construct a target sample set with a large sample size and sufficient sample diversity. In this way, the manpower cost and time cost required to construct the target sample set are effectively saved. Moreover, the target training samples obtained based on the target sample set and multiple first modification requests have higher data quality, which further effectively improves the training efficiency and training efficiency of the model, thereby laying the foundation for improving user experience.
[0082] Furthermore, in a specific example, sample amplification can be performed in the following manner to obtain N2 second initial articles and the second position indicated by each preset quantifier in each second initial article. Specifically, the above-described method of performing sample amplification based on each first initial article and the first position indicated by each preset quantifier in each first initial article, and calling the large model to obtain N2 second initial articles and the second position indicated by each preset quantifier in each second initial article (for example, step S202) can specifically include:
[0083] Step S202 - 1 : Based on each first initial article and the first position indicated by each preset quantifier in each first initial article, determine N3 (N3 is a positive integer less than or equal to N1) sample examples.
[0084] Here, N3 is a positive integer less than or equal to N1, or, in one example, N3=N1. In this case, each first initial article may correspond to a specific sample example.
[0085] Step S202-2: construct a target prompt word based on N3 sample examples.
[0086] Step S202 - 3 : Input the target prompt word into the large model to perform sample amplification to obtain N2 second initial articles and the second position indicated by each preset quantifier in each second initial article.
[0087] In this way, the disclosed solution provides a specific solution for calling a large model for sample amplification, that is, based on N3 sample examples, constructing target prompt words, and then calling the large model, and combining the target prompt words to perform sample amplification. The above process is simple, practical and has a low threshold for use. In this way, it is possible to quickly obtain sample data with rich data volume and sufficient diversity. Moreover, since the "first position" in the sample examples in the disclosed solution can be represented by a regular expression, compared to the solution that directly relies on the content of the article for sample generalization, the disclosed solution requires fewer tokens when calling a large model and a small number of sample examples for sample amplification, and is more efficient, thereby effectively improving the training efficiency and training efficiency of the model, thereby laying the foundation for improving user experience.
[0088] Figure 3 This is a schematic flow chart of a method for training a local position recognition model according to an embodiment of the present application. Figure 3 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understood that the above Figure 1 and Figure 2 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.
[0089] Furthermore, the method includes at least part of the following contents. Figure 3 Shown, including:
[0090] Step S301: Acquire a target sample set.
[0091] Here, the target sample set includes: a plurality of target sample articles, and a target position indicated by each preset quantifier in each target sample article among a plurality of preset quantifiers.
[0092] Here, for relevant content about the target sample set, please refer to the above examples and will not be repeated here.
[0093] Step S302: Determine M1 initial modification requests.
[0094] Here, the initial modification request includes an initial quantifier, and the initial modification request is used to instruct to modify the text content at the corresponding position through the included initial quantifier.
[0095] Step S303: Based on the M1 initial modification requests, the large model is called to perform sample amplification to obtain M2 new modification requests.
[0096] Here, M1 and M2 are both positive integers, and M2 is greater than M1; in other words, the disclosed solution can generalize samples based on the initial modification requests that have been labeled, thereby enriching the sample size, thereby providing strong support for subsequent improvements in the model's training efficiency, training effects, and generalization capabilities.
[0097] Step S304: Based on the M1 initial modification requests and the M2 newly added modification requests, the multiple first modification requests are obtained.
[0098] For example, the initial modification request and the newly added modification request may be directly used as the first modification request, thereby obtaining multiple first modification requests.
[0099] Here, the first modification request includes a target quantifier, and the first modification request is used to instruct to modify the text content at the corresponding position through the included target quantifier.
[0100] For example, in one example, a small number of initial modification requests can be generated based on rules. A large model can then be invoked and, based on this small number of initial modification requests, sample amplification can be performed to generate a large number of new modification requests. This can then be used to generate multiple first modification requests with a large volume and sufficient diversity, based on this small number of initial modification requests and the amplified new modification requests. This effectively increases the richness of modification requests and provides strong support for subsequent improvements in model training efficiency, training effectiveness, and generalization capabilities.
[0101] Step S305: Based on the target positions indicated by the preset quantifiers in the target sample articles, a target position that can match the target quantifier included in the first modification request is determined.
[0102] Step S306: construct a target training sample based on the target sample article, the first modification request, and the target position matching the target quantifier contained in the first modification request, so as to train an initial local position recognition model based on the target training sample and obtain a target local position recognition model.
[0103] In this way, the disclosed solution can call upon a large model and generate multiple new modification requests based on a small number of initial modification requests, thereby enriching the sample size of modification requests. Compared to methods based on manually generated samples, the modification requests generated by calling upon a large model are more diverse and have a richer data volume, thus providing strong support for subsequent improvements in model training efficiency, training effectiveness, and generalization capabilities.
[0104] Moreover, compared with the method of "first determining the article and then obtaining a modification request for the article", the disclosed solution can efficiently obtain multiple first modification requests without pre-determining the article. In this way, the number of occupied tokens is effectively saved, and the processing efficiency of the large model is improved, thereby further providing strong support for the subsequent improvement of the model's training efficiency, training effect and generalization ability.
[0105] The disclosed solution also provides a local position recognition method, specifically, Figure 4 This is a schematic flow chart of a method for identifying a local location according to an embodiment of the present application. The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0106] Furthermore, the method includes at least part of the following contents. Figure 4 Shown, including:
[0107] Step S401: Obtain a target modification request for an initial article.
[0108] Here, the target modification request is used to indicate the required modification content through quantifiers. Further, the initial article is generated by calling a large model.
[0109] Step S402: Input the initial article and the target modification request into a target local position recognition model to obtain the specific position of the modification content requested by the target modification request in the initial article.
[0110] Here, the model parameter amount of the target local position recognition model is smaller than the model parameter amount of the large model.
[0111] Here, in one example, the target local position recognition model is trained using the training method described above.
[0112] In this way, the disclosed solution can quickly obtain the specific location of the modification content requested by the target modification request in the initial article based on the target local position recognition model and combined with the target modification request. The model parameter amount of the target local position recognition model used in the above process is smaller than (or even much smaller than) the model parameter amount of the large model. In this way, the computing resources required for position recognition based on the target local position recognition model are less. In other words, on the basis of ensuring the location recognition capability of the model, the cost required for location recognition is lower and more efficient, thereby providing strong support for the subsequent call of the large model for local modification, and at the same time, laying the foundation for improving user experience.
[0113] Furthermore, in a specific example, after obtaining the specific location of the modification content requested by the target modification request in the initial article, the method further includes:
[0114] Step S403: Based on the specific location of the modification content requested by the target modification request in the initial article, a location identifier for the requested modification content is generated.
[0115] Step S404: Based on the location identifier, determine the text content targeted by the target modification request from the initial article.
[0116] Step S405: input the determined text content targeted by the target modification request and the target modification request into the large model to obtain updated text content.
[0117] Step S406: Based on the location identifier, the updated text content is updated into the initial article to obtain an updated target article.
[0118] That is to say, after obtaining the specific position of the modified content requested by the target modification request in the initial article, a position identifier can be generated based on the specific position, and then the text content requested to be modified by the target modification request can be determined from the initial article based on the position identifier; the text content requested to be modified and the target modification request are then input into the big model to call the big model to modify the text content requested to be modified, and obtain the updated text content; finally, based on the position identifier, the text content requested to be modified in the initial article is replaced with the updated text content, and the updated target article is obtained.
[0119] In this way, the disclosed solution provides a detailed solution that utilizes the target local location recognition model and the large model to collaboratively implement local content modification. This solution not only implements the modification of local content and improves the user experience, but also effectively reduces resource usage and reduces the cost of local modification.
[0120] Furthermore, since the disclosed solution can first locate the specific position of the text content to be modified in the target modification request based on the target local position recognition model (compared to the large model, the target local position recognition model can be understood as a "lightweight model"), and after obtaining the text content to be modified at the specific position, the large model is called to complete the modification of the text content to be modified. Therefore, compared with the solution of directly calling the large model to perform local modifications on the initial article, the disclosed solution requires less processing resources, occupies fewer tokens during model inference, and is more efficient.
[0121] The disclosed solution also provides a training device for a local position recognition model, such as Figure 5 Shown, including:
[0122] The sample determination unit 501 is configured to obtain a target sample set, wherein the target sample set includes: a plurality of target sample articles, and a target position indicated by each of a plurality of preset quantifiers in each target sample article; determine a plurality of first modification requests; wherein the first modification request includes a target quantifier, and the first modification request is used to instruct to modify the text content at a corresponding position using the target quantifier included therein; and determine a target position that can match the target quantifier included in the first modification request based on the target position indicated by each preset quantifier in each target sample article;
[0123] The model training unit 502 is used to construct a target training sample based on the target sample article, the first modification request, and the target position matching the target quantifier contained in the first modification request, so as to train the initial local position recognition model based on the target training sample and obtain the target local position recognition model.
[0124] In a specific example of the disclosed solution, the target local location recognition model can output the specific location of the modification content requested by the target modification request in the initial article based on the target modification request and the initial article generated by the large model;
[0125] The target modification request is used to indicate the required modification content through a quantifier; the model parameter quantity of the target local position recognition model is less than the model parameter quantity of the large model.
[0126] In a specific example of the disclosed solution, the target position indicated by the preset quantifier in the target sample article is represented by a regularized expression;
[0127] The amount of data occupied by the regular expression is smaller than the amount of data of the text content indicated by the preset quantifier.
[0128] In a specific example of the disclosed solution, the sample determination unit is further configured to:
[0129] Based on a plurality of preset quantifiers, each of the N1 first initial articles is decomposed to obtain a first position indicated by each preset quantifier in each first initial article;
[0130] Based on each first initial article and the first position indicated by each preset quantifier in each first initial article, and calling the large model to perform sample amplification, N2 second initial articles and the second position indicated by each preset quantifier in each second initial article are obtained; wherein N1 and N2 are both positive integers, and N2 is greater than N1;
[0131] The target sample set is constructed based on each first initial article and the first position indicated by each preset quantifier in each first initial article, and each second initial article and the second position indicated by each preset quantifier in each second initial article.
[0132] In a specific example of the present disclosure, the sample determination unit is specifically configured to include:
[0133] Based on each first initial article and the first position indicated by each preset quantifier in each first initial article, determining N3 sample examples; N3 is a positive integer less than or equal to N1;
[0134] Based on N3 sample examples, construct the target prompt word;
[0135] The target prompt word is input into the large model to perform sample amplification to obtain N2 second initial articles and the second position indicated by each preset quantifier in each second initial article.
[0136] In a specific example of the present disclosure, the sample determination unit is specifically configured to:
[0137] Determine M1 initial modification requests; the initial modification requests include initial quantifiers, and the initial modification requests are used to indicate modification of text content at corresponding positions through the included initial quantifiers;
[0138] Based on M1 initial modification requests, the large model is called to perform sample amplification, resulting in M2 new modification requests. M1 and M2 are both positive integers, and M2 is greater than M1.
[0139] Based on the M1 initial modification requests and the M2 newly added modification requests, the multiple first modification requests are obtained.
[0140] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0141] The present disclosure also provides a local position identification device, such as Figure 6 Shown, including:
[0142] The request acquisition unit 601 is used to acquire a target modification request for an initial article, wherein the target modification request is used to indicate the required modification content through a quantifier; wherein the initial article is generated by calling a large model;
[0143] The model inference unit 602 is configured to input the initial article and the target modification request into a target local location recognition model to obtain the specific location of the modification content requested by the target modification request in the initial article;
[0144] Among them, the model parameter amount of the target local position recognition model is smaller than the model parameter amount of the large model.
[0145] In a specific example of the disclosed solution, the model inference unit is also used to
[0146] generating a location identifier for the requested modification content based on the specific location of the modification content requested by the target modification request in the initial article;
[0147] Based on the position identifier, determining the text content targeted by the target modification request from the initial article;
[0148] Inputting the determined text content targeted by the target modification request and the target modification request into the large model to obtain updated text content;
[0149] Based on the position identifier, the updated text content is updated into the initial article to obtain an updated target article.
[0150] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0151] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0152] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0153] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0154] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0155] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0156] The computing unit 701 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as training methods or reasoning methods. For example, in some embodiments, the training method or reasoning method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the training method or reasoning method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the training method or reasoning method by any other appropriate means (e.g., by means of firmware).
[0157] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0158] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0159] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0160] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0161] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0162] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0163] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0164] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A training method for a local position recognition model, comprising: Acquire a target sample set, wherein the target sample set includes: a plurality of target sample articles, and a target position indicated by each of a plurality of preset quantifiers in each target sample article; Determine a plurality of first modification requests; wherein the first modification requests include a target quantifier, and the first modification requests are used to instruct to modify the text content at the corresponding position through the included target quantifier; Based on the target positions indicated by the preset quantifiers in the target sample articles, determining a target position that can match the target quantifier included in the first modification request; Based on the target sample article, the first modification request, and the target position matching the target quantifier contained in the first modification request, a target training sample is constructed to train an initial local position recognition model based on the target training sample, and a target local position recognition model is obtained.
2. The method according to claim 1, wherein The target local position recognition model can output the specific position of the modification content requested by the target modification request in the initial article based on the target modification request and the initial article generated by the large model; The target modification request is used to indicate the required modification content through a quantifier; the model parameter quantity of the target local position recognition model is smaller than the model parameter quantity of the large model.
3. The method according to claim 1 or 2, wherein The target position indicated by the preset quantifier in the target sample article is represented by a regularized expression; The amount of data occupied by the regular expression is smaller than the amount of data of the text content indicated by the preset quantifier.
4. The method according to any one of claims 1 to 3, further comprising: Based on a plurality of preset quantifiers, each of the N1 first initial articles is decomposed to obtain a first position indicated by each preset quantifier in each first initial article; Based on each first initial article and the first position indicated by each preset quantifier in each first initial article, and calling the large model to perform sample amplification, N2 second initial articles and the second position indicated by each preset quantifier in each second initial article are obtained; wherein N1 and N2 are both positive integers, and N2 is greater than N1; The target sample set is constructed based on each first initial article and the first position indicated by each preset quantifier in each first initial article, and each second initial article and the second position indicated by each preset quantifier in each second initial article.
5. The method according to claim 4, wherein The method of amplifying the sample based on each first initial article and the first position indicated by each preset quantifier in each first initial article and calling the large model to obtain N2 second initial articles and the second position indicated by each preset quantifier in each second initial article includes: Based on each first initial article and the first position indicated by each preset quantifier in each first initial article, determining N3 sample examples; N3 is a positive integer less than or equal to N1; Based on N3 sample examples, construct the target prompt word; The target prompt word is input into the large model to perform sample amplification to obtain N2 second initial articles and the second position indicated by each preset quantifier in each second initial article.
6. The method according to any one of claims 1 to 3, wherein: The determining of the plurality of first modification requests comprises: Determine M1 initial modification requests; the initial modification requests include initial quantifiers, and the initial modification requests are used to indicate modification of text content at corresponding positions through the included initial quantifiers; Based on M1 initial modification requests, the large model is called to perform sample amplification, resulting in M2 new modification requests. M1 and M2 are both positive integers, and M2 is greater than M1. Based on the M1 initial modification requests and the M2 newly added modification requests, the multiple first modification requests are obtained.
7. A local position recognition method, comprising: Obtaining a target modification request for an initial article, wherein the target modification request is used to indicate the required modification content through a quantifier; wherein the initial article is generated by calling a large model; Inputting the initial article and the target modification request into a target local position recognition model to obtain the specific position of the modification content requested by the target modification request in the initial article; Among them, the model parameter amount of the target local position recognition model is smaller than the model parameter amount of the large model.
8. The method according to claim 7, further comprising: generating a location identifier for the requested modification content based on the specific location of the modification content requested by the target modification request in the initial article; Based on the position identifier, determining the text content targeted by the target modification request from the initial article; Inputting the determined text content targeted by the target modification request and the target modification request into the large model to obtain updated text content; Based on the position identifier, the updated text content is updated into the initial article to obtain an updated target article.
9. A training device for a local position recognition model, comprising: A sample determination unit is configured to obtain a target sample set, wherein the target sample set includes: a plurality of target sample articles, and a target position indicated by each of a plurality of preset quantifiers in each target sample article; determine a plurality of first modification requests; wherein the first modification request includes a target quantifier, and the first modification request is configured to instruct to modify text content at a corresponding position using the target quantifier included therein; and determine, based on the target position indicated by each of the preset quantifiers in each target sample article, a target position that can match the target quantifier included in the first modification request; The model training unit is used to construct a target training sample based on the target sample article, the first modification request, and the target position matching the target quantifier contained in the first modification request, so as to train the initial local position recognition model based on the target training sample and obtain the target local position recognition model.
10. A local position recognition device, comprising: a request acquisition unit, configured to acquire a target modification request for an initial article, wherein the target modification request is used to indicate the required modification content by using a quantifier; wherein the initial article is generated by calling a large model; A model inference unit is configured to input the initial article and the target modification request into a target local position recognition model to obtain a specific position of the modification content requested by the target modification request in the initial article; Among them, the model parameter amount of the target local position recognition model is smaller than the model parameter amount of the large model.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Text processing, model training and speech synthesis method, device and system and medium
CN114822491A
Code processing model training method and code task processing method
CN118939245A
Method and apparatus for training pre-trained knowledge model, and electronic device
US20210248498A1
Method and apparatus for processing text, electronic device, and storage medium
US20250173501A1