Data processing method, electronic device, storage medium and computer program product

By performing targeted enhancement training on a large language model and generating a target text generation model, the problems of low processing efficiency and poor accuracy in existing technologies are solved, and efficient and accurate text transcription processing is achieved, especially for multimedia content processing in educational and learning scenarios.

CN120705306APending Publication Date: 2025-09-26ALIBABA (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410325219.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing large language models have problems with low processing efficiency, poor accuracy and adaptability when transcribing target files in specific fields. Especially in educational and learning scenarios such as meetings, interviews, and training, it is difficult to meet the needs of efficient and accurate multimedia digital content processing.

Method used

By obtaining sample generation instructions and sample data sets, the initial text generation model is trained with targeted reinforcement using preset conditional constraints to generate a target text generation model, which is then used to transcribe the target file to ensure that the generated text meets specific scenario and condition requirements.

Benefits of technology

It achieves efficient and accurate text transcription of target files, improves processing efficiency and adaptability, meets the text generation needs of specific fields, and improves the accuracy and adaptability of file transcription.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705306A_ABST
    Figure CN120705306A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, electronic equipment, a storage medium and a computer program product, and relates to the technical field of large models and computers. The method comprises the steps that a sample generation instruction and a sample data set are obtained, the sample generation instruction is used for determining a to-be-generated text type, and the sample data set covers multiple types of labels; according to a preset condition constraint mode, the sample generation instruction and the sample data set are adopted to conduct directional enhancement training on the initial text generation model, a target text generation model is generated, and the preset condition constraint mode is used for generating model input instructions matched with file content of a to-be-transcribed target file from multiple different dimensions; and performing text transcription processing on the target file by adopting the target text generation model to obtain a target text. According to the method and the device, the technical problems of low processing efficiency and poor accuracy and adaptability when a file processing mode provided in the related technology is used for transferring the target file are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of large model technology and computer technology, and specifically to a data processing method, electronic equipment, storage medium and computer program product. Background Art

[0002] With the advent of the digital age, the production and consumption of digital content have exploded. In educational and learning scenarios such as meetings, interviews, and training, efficiently and accurately processing multimedia digital content has become a major challenge. After the introduction of Large Language Models (LLMs), the use of LLMs has greatly facilitated the processing of multimedia digital content. However, although LLMs perform well in understanding and generating text, they mainly rely on a wide range of text data during the pre-training phase, which may lead to certain deviations when transcribing target files in specific fields. Because the default generation logic of LLMs tends to be more general and flexible rather than optimized for specific tasks, the fine-tuning of general LLMs instructions often cannot meet the precision requirements of product levels, resulting in low processing efficiency, poor accuracy, and poor adaptability in actual applications.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide a data processing method, electronic device, storage medium and computer program product to at least solve the technical problems of low processing efficiency, poor accuracy and adaptability when transcribing target files in the file processing method provided in the related art.

[0005] According to one aspect of an embodiment of the present application, a data processing method is provided, including: obtaining a sample generation instruction and a sample data set, wherein the sample generation instruction is used to determine the type of text to be generated, and the sample data set covers multiple types of labels; according to a preset condition constraint method, the sample generation instruction and the sample data set are used to perform targeted reinforcement training on an initial text generation model to generate a target text generation model, wherein the preset condition constraint method is used to generate model input instructions that are adapted to the file content of a target file to be transcribed from multiple different dimensions; and the target text generation model is used to perform text transcription processing on the target file to obtain a target text.

[0006] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: obtaining a target file to be transcribed; using a target text generation model to perform text transcription processing on the target file to obtain a target text; wherein the target text generation model is obtained after performing targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0007] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: obtaining a conference video file to be transcribed; using a conference text generation model to perform text transcription processing on the conference video file to obtain a conference summary; wherein the conference text generation model is obtained after targeted reinforcement training according to a conference condition constraint method, and the conference condition constraint method is used to generate model input instructions that are adapted to the file content of the conference video file from multiple different dimensions.

[0008] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: obtaining a file processing request through a first application programming interface; returning a file processing response through a second application programming interface; wherein the request data carried in the file processing request includes: a target file to be transcribed, and the response data carried in the file processing response includes: a target text, the target text is obtained by performing text transcription processing on the target file using a target text generation model, and the target text generation model is obtained by performing targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0009] According to another aspect of an embodiment of the present application, a data processing method is also provided, including: obtaining a currently input file processing dialogue request; returning a file processing dialogue reply in response to the file processing dialogue request; wherein the request data carried in the file processing dialogue request includes: a target file to be transcribed, and the information carried in the file processing dialogue reply includes: a target text, the target text is obtained by performing text transcription processing on the target file using a target text generation model, the target text generation model is obtained by performing targeted reinforcement training according to a preset condition constraint method, the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions; and the target text is displayed in a graphical user interface.

[0010] According to another aspect of an embodiment of the present application, an electronic device is further provided, comprising: a memory storing an executable program; and a processor for running the program, wherein the program executes the data processing method described in any one of the embodiments of the present application when running.

[0011] According to another aspect of an embodiment of the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the data processing method described in any one of the embodiments of the present application.

[0012] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program, which implements the data processing method described in any one of the embodiments of the present application when executed by a processor.

[0013] In an embodiment of the present application, by obtaining sample generation instructions and sample data sets, and then according to preset conditional constraints, using sample generation instructions and sample data sets to perform targeted reinforcement training on the initial text generation model, a target text generation model is generated, and finally the target text generation model is used to perform text transcription processing on the target file to obtain the target text, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related art when transcribing the target file.

[0014] It is easy to notice that the above general description and the following detailed description are merely for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0016] Figure 1 This is a schematic diagram of an application scenario of a data processing method according to Example 1 of the present application;

[0017] Figure 2 is a flow chart of a data processing method according to Example 1 of the present application;

[0018] Figure 3 This is a word cloud corresponding to a sample data set according to Example 1 of the present application;

[0019] Figure 4 is a schematic diagram of a constraint evolution process according to Example 1 of the present application;

[0020] Figure 5 is a schematic diagram of a data processing method according to Example 1 of the present application;

[0021] Figure 6 is a flow chart of a data processing method according to Example 2 of the present application;

[0022] Figure 7 is a flow chart of a data processing method according to Example 3 of the present application;

[0023] Figure 8 is a flow chart of a data processing method according to Example 4 of the present application;

[0024] Figure 9 is a flow chart of a data processing method according to Example 5 of the present application;

[0025] Figure 10 is a structural block diagram of a file processing device according to embodiment 6 of the present application;

[0026] Figure 11 is a structural block diagram of another file processing device according to embodiment 6 of the present application;

[0027] Figure 12 is a structural block diagram of another file processing device according to embodiment 6 of the present application;

[0028] Figure 13 is a structural block diagram of another file processing device according to embodiment 6 of the present application;

[0029] Figure 14 is a structural block diagram of another file processing device according to embodiment 6 of the present application;

[0030] Figure 15 This is a structural block diagram of a computer terminal according to Example 7 of the present application. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] The technical solution provided in this application is mainly implemented using large-scale model technology. The large model here refers to a deep learning model with large-scale model parameters, which can usually contain hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. The large model can also be called a cornerstone model / foundation model (Foundation Model). The large model is pre-trained by large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as large-scale language models (LLMs) and multi-modal pre-training models.

[0034] It should be noted that when the large model is actually used, the pre-training model can be fine-tuned by a small amount of samples so that the large model can be applied to different tasks. For example, the large model can be widely used in fields such as natural language processing (NLP), computer vision, speech processing, etc., and can be specifically applied to computer vision tasks such as visual question answering (VQA), image description (IC), and image generation. It can also be widely used in natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. Therefore, the main application scenarios of the large model include but are not limited to digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc. In the embodiment of the present application, data processing is performed by generating a target text model in educational learning scenarios such as meetings, interviews, and training as an example for explanation.

[0035] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:

[0036] Large language models are a type of model in the field of artificial intelligence specifically designed to understand and generate natural language text. By learning from large amounts of text data, these models can grasp the structure, grammar, vocabulary, and context of a language, enabling them to perform various language tasks such as text generation, summarization, and question answering.

[0037] Summary: The process of extracting key information from the original text and summarizing it. The purpose of a summary is to provide a short, concise version that retains the main content and important details of the original text while omitting any minor or unnecessary information.

[0038] Instruction fine-tuning: Instruction fine-tuning is a technique in Natural Language Processing (NLP) used to improve the performance of large-scale pre-trained language models. This technique typically involves further training or adjusting the model using specific instructions or tasks based on the large-scale pre-trained model so that the model can better understand and execute these instructions or tasks.

[0039] Conditional generation: Conditional generation is a natural language processing and machine learning technique used to impose specific conditions or rules when generating text or other types of data. This technique is often applied to large language models. The goal of conditional generation is to ensure that the generated content is not only high-quality and conforms to linguistic standards, but also meets specific requirements or restrictions, such as the content's topic, style, structure, or adherence to specific ethical and social standards.

[0040] With the advent of the digital age, the production and consumption of digital content has exploded. Efficiently and accurately processing multimedia digital content has become a major challenge in educational and learning scenarios such as meetings, interviews, and training. The introduction of LLMs has greatly facilitated the processing of multimedia digital content. However, LLMs primarily use large-scale text data during pre-training, which can lead to some discrepancies in their scope.

[0041] Specifically, the open-source language model (WizardLM) in the related art trains complex instruction data using an unsupervised learning method (Evol-Instruct) based on evolutionary strategies. This aims to improve the complexity and diversity of instruction data by evolving it, thereby enabling the open-source language model to better handle complex instructions. The development of WizardLM focuses on enhancing the model's instruction-following capabilities, increasing the complexity and coverage of instructions to improve the model's performance and generalization capabilities.

[0042] However, WizardLM suffers from two shortcomings: First, it cannot generate targeted enhancements for the audio and video domain. During its evolution, Evol-Instruct only considers the diversity and complexity of the instruction itself, without considering the input source knowledge that the instruction relies on, such as transcribed text records in the audio and video domain. Second, Evol-Instruct cannot achieve targeted enhancements. During its evolution, it generates instructions that are completely unrelated to the target instruction (such as the summary instruction), and does not explicitly model the relationship between the instruction and the constraints, making it difficult to judge the difficulty of the evolved instructions. These two shortcomings make WizardLM difficult to directly apply to scenarios where targeted enhancement conditional constraint summaries are generated.

[0043] In summary, the file processing methods provided in the related art have technical problems of low processing efficiency, poor accuracy and adaptability when transcribing the target file. Currently, no effective solutions have been proposed to the above problems.

[0044] Example 1

[0045] According to an embodiment of the present application, a data processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0046] Considering the huge number of model parameters of large models and the limited computing resources of mobile terminals, the above data processing method provided in the embodiment of the present application can be applied to Figure 1 The application scenarios shown are not limited to this. Figure 1 In the illustrated application scenario, the large model is deployed on a server 10. The server 10 can be connected to one or more client devices 20 via a local area network, a wide area network, the Internet, or other types of data networks. The client devices 20 herein may include, but are not limited to, smartphones, tablet computers, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. The client devices 20 can interact with users via a graphical user interface to access the large model and thereby implement the methods provided in the embodiments of the present application.

[0047] In an embodiment of the present application, a system composed of a client device and a server can perform the following steps: the server obtains a currently input file processing dialogue request from the client device, and then returns a file processing dialogue reply to the client device in response to the file processing dialogue request; wherein, the request data carried in the file processing dialogue request includes: the target file to be transcribed, and the information carried in the file processing dialogue reply includes: the target text, the target text is obtained by using a target text generation model to perform text transcription processing on the target file, and the target text generation model is obtained after targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions, and finally the target text is displayed in the graphical user interface of the client device.

[0048] It should be noted that, when the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiments of the present application can be carried out in the client device.

[0049] Under the above operating environment, this application provides Figure 2 The data processing method shown. Figure 2 This is a flow chart of a data processing method according to Example 1 of the present application. Figure 2 As shown, the method may include the following steps:

[0050] Step S21: obtaining a sample generation instruction and a sample data set, wherein the sample generation instruction is used to determine the type of text to be generated, and the sample data set covers multiple types of labels;

[0051] Step S22: performing targeted reinforcement training on the initial text generation model using the sample generation instructions and the sample data set according to a preset condition constraint method to generate a target text generation model, wherein the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file to be transcribed from multiple different dimensions;

[0052] Step S23: Using the target text generation model to perform text transcription processing on the target file to obtain the target text.

[0053] The sample generation instructions may be a set of instructions or rules for generating a summary, which are used to determine the type, format, content, and other characteristics of the text to be generated. The sample generation instructions may include, but are not limited to, requirements such as text length, language style, and subject matter, so as to generate a text sample that meets specific requirements.

[0054] The above-mentioned sample data set can be a set of existing text samples that cover multiple types of labels and can be used to train models, test algorithms, or analyze text features. The sample data set can include various types of transcribed texts, as well as corresponding labels or classification information. For example, the goal of summarization is to automatically extract and generate a concise, coherent summary that retains the main information of the original content from long speeches, meeting minutes, or any other form of text. Therefore, the sample data set also includes the source data content of the speech, meeting minutes, etc. of the summary target.

[0055] Furthermore, training the initial text generation model using sample generation instructions and a sample dataset can improve its generation capabilities and accuracy, enabling it to better generate target text. During training, the initial text generation model learns the language patterns, vocabulary usage, and syntactic structure in the sample dataset, enabling it to more accurately generate content that meets the requirements of the target text. Through continuous iterative training, the generation performance of the initial text generation model will gradually improve, achieving higher generation quality and accuracy.

[0056] The target file to be transcribed can be audio or video content selected by an artificial intelligence (AI) assistant. The audio or video content can include, but is not limited to, multimedia digital content associated with scenarios such as meeting minutes, course lectures, media interviews, and professional training, which needs to be converted into text for transcription. For example, the target file to be transcribed can be at least the following: audio files, video files, meeting minutes, interview recordings, etc. The transcription process can help convert the multimedia content into text for easy reading, editing, and searching.

[0057] The target text generation model can be a pre-trained large language model with spoken summary capabilities. Compared to general natural language generation models, the target text generation model focuses on generating text output for a specific domain or task based on a given input and context. During training, the target text generation model learns a wealth of linguistic knowledge and contextual information, enabling it to understand and generate natural language text. The target text generation model can generate concise and accurate summaries or abstracts by understanding and analyzing the spoken information in the target document and performing targeted reinforcement training based on pre-set conditional constraints. Based on its understanding of spoken expressions and context, the target text generation model, leveraging the generative capabilities of the language model, can transform spoken information into more formal, structured text output to achieve spoken summary functionality. This provides a more professional and automated text generation service, helping users summarize and transform spoken information more efficiently.

[0058] Text transcription refers to converting text content from speech or images into an editable text format. This process typically involves speech recognition or optical character recognition technology. For example, when a target file needs to be transcribed, it is input into a trained target text generation model. The target text generation model will then generate appropriate target text based on pre-set constraints. The resulting target text can be used in various application scenarios, such as speech recognition and natural language processing. This enables the target text generation model to summarize audio and video content strictly according to product-defined scenarios, conditions, and formats. For example, a scenario might be: Summarize a meeting minutes based on a given meeting conversation; a condition might be: Summarize the conversation into a summary of no more than 300 words; and a format might be: Summarize the conversation into a summary and a title, output in JSON format. The "title" field in the JSON corresponds to the title of the summary, and the "summary" field corresponds to the summary of the summary. By introducing a conditional constraint enhancement mechanism, the target text generation model after targeted enhancement is able to comply with specific rules and constraints when generating text. It can not only generate summaries and abstracts based on audio and video content, but also ensure that these generated abstracts are more in line with the specific conditions and forms defined by the product.

[0059] Based on the above steps S21 to S23, by obtaining sample generation instructions and sample data sets, and then according to the preset condition constraints, the sample generation instructions and sample data sets are used to perform targeted enhancement training on the initial text generation model to generate a target text generation model, and finally the target text generation model is used to perform text transcription processing on the target file to obtain the target text, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file during file transcription, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the relevant technology when transcribing the target file.

[0060] The following is a step-by-step introduction to the data processing method in the embodiment of the present application.

[0061] In an optional embodiment, in step S21, obtaining a sample generation instruction includes:

[0062] Step S211, obtaining a sample generation task, wherein the sample generation task is a text generation task corresponding to the type of text to be generated;

[0063] Step S212: Determine a sample generation instruction based on the task seed instruction corresponding to the sample generation task.

[0064] Specifically, the above-mentioned sample generation task can be a preset number of summary generation tasks, and the above-mentioned task seed instruction is an instruction to guide the model to generate a specific type of sample, which may include a description of the task, requirements and restrictions, as well as constraints and guidance on the expected results of the generated samples. Task seed instructions can help the initial text generation model better understand the requirements of the task and generate samples that meet expectations. In sample-based generation tasks, task seed instructions can include requirements and guidance on the content, format, quantity, quality, etc. of the generated samples. Through task seed instructions, the model can more accurately generate samples that meet the task requirements.

[0065] For example, 13 summary generation tasks are obtained, which may include but are not limited to summarizing Chinese and English documents, conversations, keywords, summarizing by topic, summarizing by role, summarizing into a mind map, etc., where a mind map is a series of related ideas, concepts or information presented in a graphical way to help people better understand and remember these contents. Mind maps usually use elements such as branches, keywords, colors and icons to express information, which can deepen the understanding of knowledge and improve memory and learning efficiency. Each summary generation task is set with 10 task seed instructions, each task seed instruction contains 1 restriction condition, and a total of 130 task seed instructions. For example, the task seed instruction can be "Please summarize the given conversation content in an ordered list."

[0066] Based on the above optional embodiment, by obtaining the sample generation task, and then determining the sample generation instruction based on the task seed instruction corresponding to the sample generation task, the corresponding task seed instruction can be accurately generated according to the requirements of the sample generation task, thereby determining the sample generation instruction, thereby improving the completion quality of the file processing task and saving manpower and time costs.

[0067] In an optional embodiment, in step S21, obtaining a sample data set includes:

[0068] Step S213: labeling the multi-source data set using the text labeling model to obtain a labeling result;

[0069] Step S214: normalize the labeling results to obtain multiple types of labels;

[0070] Step S215 : selecting a sample data set from the multi-source data set based on multiple types of labels.

[0071] Specifically, in order to help text tagging models (such as large language models or other artificial intelligence models suitable for labeling) adapt to the audio and video field, it is necessary to select as rich and diverse source data as possible to cover typical situations in audio and video scenes. In order to evaluate the diversity of source data, the multi-source data set is first labeled using a large language model to obtain labeling results, such as education labels. At the same time, similar labels in the labeling results are normalized in a heuristic rule manner to obtain multiple types of labels. For example, after normalization of the labels "education", "educational", and "educationally", they can be uniformly normalized to the label "education".

[0072] After labeling each source data in a multi-source data set, a greedy algorithm can be used to select a sample data set from the multi-source data set. For example, each time, source data with more labels and less overlap with the selected labels is selected, so that a sample data set covering all labels can be selected. The idea of ​​the greedy algorithm is to select data that can cover more uncovered labels each time until all labels are covered. Suppose there are the following source data sets: Data 1: Labels A and B; Data 2: Labels B and C; Data 3: Labels C and D; Data 4: Labels D and E; Data 5: Labels E and F. First, select the data that can cover more labels, that is, Data 1, because it contains labels A and B. Then, remove the already covered label B. Next, select Data 3 because it contains labels C and D. Then remove the already covered labels C and D. Finally, select Data 5 because it contains labels E and F. The sample data sets that can thus cover all labels are Data 1, Data 3, and Data 5.

[0073] In practical applications, we can select more than 3,200 Chinese-English transcription results from the accumulated more than 27,000 Chinese-English transcription results as the sample data set, which accounts for only 11.7% and covers the total 4,086 tags. On average, each Chinese-English transcription result contains 3.61 tags, while also reflecting the diversity and complexity of the Chinese-English transcription results. Figure 3 is a word cloud diagram corresponding to a sample data set according to Example 1 of the present application, such as Figure 3As shown in FIG, the corresponding word cloud diagram in the sample data set may include the following tags: entertainment, language, learning, personal, confirmation, technology, marketing strategy, education, dialogue, consumer service, introduction, data analysis, personal growth, etc.

[0074] Based on the above optional embodiment, by using a large language model to label the multi-source data set, a labeling result is obtained, and then the labeling result is normalized to obtain multiple types of labels. Finally, a sample data set is selected from the multi-source data set based on the multiple types of labels, so that the multi-source data set can be labeled using a large language model and the labeling result is normalized, thereby realizing automatic data labeling and the acquisition of multiple types of labels, providing effective technical support and basic data for subsequent data analysis and mining.

[0075] In an optional embodiment, in step S22, the initial text generation model is subjected to targeted enhancement training using the sample generation instruction and the sample data set to generate the target text generation model, including:

[0076] Step S221, constraining and evolving the sample generation instruction according to a preset condition constraint method to obtain a constrained evolution instruction;

[0077] Step S222: Use the constrained evolution instruction and the sample data set to perform targeted enhancement training on the initial text generation model to generate a target text generation model.

[0078] The above-mentioned preset condition constraint method is a given constraint condition. Combined with the given constraint condition, the sample generation instruction is directed to enhance the instruction complex constraint, and then the constraint evolution instruction is obtained.

[0079] Exemplarily, the constraints can be used to limit the length or complexity of the sample generation instructions so that the generated samples are more in line with the actual situation. When performing directional enhancement of complex constraints on instructions, an upper limit on the length of the sample generation instructions is set to ensure that the generated sample generation instructions are not too long, thereby avoiding the generation of overly complex samples. By setting the types and number of operators and operators allowed to appear in the instructions, the complexity of the sample generation instructions is limited, so that the generated samples are more in line with the actual situation. According to the actual application scenarios and needs, the generated samples are directionally enhanced, for example, corresponding constraints are added for specific application fields to ensure that the generated samples meet the actual needs. In this way, samples that are more in line with the actual situation are effectively generated, and then the constrained evolution instructions are obtained. Furthermore, the constrained evolution instructions and the sample data set are used to train the initial text generation model to generate the target text generation model.

[0080] Based on the above optional embodiments, by constraining the evolution of sample generation instructions in accordance with preset condition constraints, constrained evolution instructions are obtained, and then the constrained evolution instructions and the sample data set are used to perform targeted enhancement training on the initial text generation model to generate a target text generation model. The constrained evolution instructions can help the initial text generation model better understand and learn the characteristics and rules in the sample data set, thereby generating a target text that better meets the preset conditions, thereby further improving the accuracy, fluency and diversity of the target text generation model, thereby effectively improving the quality and diversity of the generated text.

[0081] In an optional embodiment, in step S221, the sample generation instruction is constrained and evolved according to a preset condition constraint method, and the obtained constrained evolution instruction includes:

[0082] Step S2211: determining a target evolution direction according to a preset condition constraint method, wherein the target evolution direction includes: a replacement evolution direction and an addition evolution direction. The replacement evolution direction is used to replace the initial constraint conditions of the sample generation instruction, and the addition evolution direction is used to add new constraint conditions based on the initial constraint conditions.

[0083] Step S2212: Constrained evolution is performed on the sample generation instruction based on the target evolution direction to obtain a constrained evolution instruction.

[0084] The target evolution directions mentioned above include replacement and increment. Each sample generation instruction undergoes two evolutionary changes to increase diversity. The replacement evolution direction uses a given sample generation instruction and a corresponding list of constraints to replace existing conditions with less common conditions, while maintaining the type and difficulty of the task. The increment evolution direction uses a given sample generation instruction and a corresponding list of constraints, combined with the given source knowledge, to add constraints, maintaining the same task type but increasing the difficulty.

[0085] Figure 4 is a schematic diagram of a constraint evolution process according to Example 1 of the present application, such as Figure 4 As shown, the sample generation instruction is "please summarize the given conversation content in an ordered list". Based on the replacement evolution direction, the sample generation instruction is constrained and evolved, and the constrained evolution instruction is "please summarize the given conversation content in an unordered list". Based on the addition evolution direction, the sample generation instruction is constrained and evolved, and the constrained evolution instruction is "please summarize the given conversation content in an ordered list, and the number of element data in the list shall not exceed 5".

[0086] Based on the above optional embodiments, the target evolution direction is determined in accordance with preset condition constraints, and then the sample generation instructions are constrained and evolved based on the target evolution direction to obtain constrained evolution instructions. This allows the target evolution direction to be determined in accordance with preset condition constraints, thereby realizing constrained evolution of instructions during the sample generation process, so that the generated samples are more in line with the preset conditions and meet specific target requirements.

[0087] In an optional embodiment, in step S322, the initial text generation model is subjected to directed enhancement training using the constrained evolution instruction and the sample data set to generate the target text generation model, including:

[0088] Step S3221, using the constrained evolution instruction and the sample data set to train the initial text generation model, and obtain the training result corresponding to the constrained evolution instruction;

[0089] Step S3222: continue to iteratively train the initial text generation model based on the constrained evolution instruction and the training results corresponding to the constrained evolution instruction until all sample data in the sample data set are used up, thereby generating a target text generation model.

[0090] Based on the above optional embodiment, the initial text generation model is trained by using the constrained evolution instruction and the sample data set to obtain the training results corresponding to the constrained evolution instruction, and then the initial text generation model is iteratively trained based on the constrained evolution instruction and the training results corresponding to the constrained evolution instruction until all the sample data in the sample data set are used up, and the target text generation model is generated, thereby improving the accuracy and quality of the generated text. Since the constrained evolution instruction can guide the model to generate text that better meets the requirements, and the sample data set can provide more training data, the model can learn the laws and characteristics of text generation more accurately. The target text generation model obtained by training through the above process can better meet actual needs and has higher technical effects.

[0091] In an optional embodiment, in step S3222, the initial text generation model is iteratively trained based on the constrained evolution instruction and the training results until all sample data in the sample data set are used up, and generating the target text generation model includes:

[0092] The initial text generation model is continuously iteratively trained based on the constrained evolution instructions and the training results, and the current round training results obtained from each round of training are sequentially obtained;

[0093] Determine the target loss based on the current round of training results and the actual results corresponding to the current round of training results;

[0094] The target loss is used to fine-tune the model parameters of the intermediate text generation model to obtain the text generation model to be used in the next round of training, until all the sample data in the sample data set are used up, and the target text generation model is generated, where the intermediate text generation model is the model obtained by the initial text generation model after completing the current round of training.

[0095] For example, 13 summary generation tasks are obtained, and each summary generation task is correspondingly set with 10 task seed instructions. Each task seed instruction contains 1 constraint, for a total of 130 task seed instructions. Each evolution is based on the newly generated instructions of the previous iteration, and the original instructions will not be iterated again. Combining the above 130 seed instructions and the selected 3200+ diverse source data, 23,000 data are generated after constrained evolution, and the average number of constraints per data increases to 7.3. After using the 23,000 data generated after the above constrained evolution, the model parameters of the intermediate text generation model can be fine-tuned and optimized to generate the target text generation model. Among them, the iteration termination condition is that the source data is used up, and finally a specific large model of oral summarization ability is trained, that is, the target text generation model is obtained. After inputting an audio file or video file into the target text generation model, the target text generation model can perform text transcription processing on it to obtain a spoken summary of the audio file or video file.

[0096] Based on the above optional embodiment, the initial text generation model is iteratively trained based on the constrained evolution instruction and the training results, and the current round training results obtained in each round of training are obtained in turn. Then, the target loss is determined based on the current round training results and the actual results corresponding to the current round training results. Finally, the target loss is used to fine-tune the model parameters of the intermediate text generation model to obtain the text generation model to be used in the next round of training, until all the sample data in the sample data set are used up, and the target text generation model is generated, thereby improving the performance of the target text generation model so that it can better generate text content that meets expectations, and continuously optimize the model parameters, thereby improving the accuracy and usability of the model.

[0097] In an optional embodiment, the multiple different dimensions include at least some of the following dimensions: scenario dimension, condition dimension, and format dimension, wherein the scenario dimension is used to determine the usage scenario of the target text, the condition dimension is used to determine the generation condition of the target text, and the format dimension is used to determine the usage format of the target text.

[0098] The usage scenarios of the above-mentioned target text include but are not limited to meeting minutes, course lectures, media interviews, professional training, etc. The generation conditions include but are not limited to word count, language, style, etc. The usage formats include but are not limited to text format, lightweight data exchange format (JSON), etc.

[0099] Figure 5 is a schematic diagram of a data processing method according to Example 1 of the present application, such as Figure 5As shown, first, diverse source data is selected to cover typical situations in audio and video scenarios. Then, constraints are evolved for summarization instructions. Finally, training data is updated, and the model is fine-tuned and evaluated. As a result, this embodiment of the application can better meet specific functional requirements for language models, such as the ability to summarize spoken dialogue information.

[0100] The data processing method in the embodiment of the present application has the following advantages compared to the data processing method improved in the related art, which are mainly reflected in the following two aspects:

[0101] First, for the spoken summary scenario, the embodiment of the present application adds a step of selecting diverse source data. In the models in the related art, the selection of data sources often focuses on breadth rather than diversity, which causes the model to perform poorly when processing specific types of summary tasks due to the lack of sufficient sample diversity. In order to solve this problem, a strategy for selecting diverse source data is introduced in the data preparation stage, which collects and screens data specifically for different types of summary scenarios (such as speeches, interviews, and interviews). This not only enhances the model's adaptability to different summary tasks, but also helps to improve the accuracy and relevance of the summary.

[0102] Secondly, the evolution process is focused on replacing and adding constraints. Traditional model evolution often lacks a clear goal, which can lead to over-dispersion of model capabilities across different tasks, impacting the efficiency of specific tasks. To overcome this challenge, we ensure that the evolution process is strictly focused on improving summary capabilities by precisely replacing and purposefully adding constraints. This not only prevents the model's capabilities from over-evolving into non-target tasks, but also enables more effective targeted optimization for summary tasks.

[0103] Through the above two improvements, the performance of the data processing method in the embodiment of the present application in the summary scenario has been significantly improved. Specifically, the addition of the step of selecting diverse source data enables the model to more flexibly adjust its own processing strategy when facing different types of summary tasks, thereby improving the quality and coverage of the summary. At the same time, focusing on the evolutionary process and replacing and adding conditional constraints makes the directional enhancement of the model's summary ability more effective, avoids ability drift, and ensures that the model's ability to summarize spoken dialogue information in the target text generation model is better.

[0104] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0105] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0106] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0107] Example 2

[0108] According to an embodiment of the present application, a data processing method is also provided. Figure 6 This is a flow chart of a data processing method according to Example 2 of the present application. Figure 6 As shown, the method may include the following steps:

[0109] Step S61, obtaining the target file to be transcribed;

[0110] Step S62, using the target text generation model to perform text transcription processing on the target file to obtain the target text; wherein, the target text generation model is obtained after targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0111] The target file to be transcribed can be audio or video content selected by an AI assistant. This audio or video content can include, but is not limited to, multimedia digital content associated with scenarios such as meeting minutes, course lectures, media interviews, and professional training, which needs to be converted into text for transcription. For example, the target file to be transcribed can be at least the following: audio files, video files, meeting minutes, interview recordings, etc. The transcription process can help convert the multimedia content into text for easier reading, editing, and searching.

[0112] The target text generation model can be a pre-trained large language model with spoken summary capabilities. Compared to general natural language generation models, the target text generation model focuses on generating text output for a specific domain or task based on a given input and context. During training, the target text generation model learns a wealth of linguistic knowledge and contextual information, enabling it to understand and generate natural language text. The target text generation model can generate concise and accurate summaries or abstracts by understanding and analyzing the spoken information in the target document and performing targeted reinforcement training based on pre-set conditional constraints. Based on its understanding of spoken expressions and context, the target text generation model, leveraging the generative capabilities of the language model, can transform spoken information into more formal, structured text output to achieve spoken summary functionality. This provides a more professional and automated text generation service, helping users summarize and transform spoken information more efficiently.

[0113] For example, when the target file needs to be transcribed, the target file is input into the trained target text generation model. The target text generation model will generate an appropriate target text according to preset conditional constraints. The obtained target text can be used in various application scenarios, such as speech recognition, natural language processing, etc.

[0114] Based on the above steps S61 to S62, by obtaining the target file to be transcribed, and then using the target text generation model to perform text transcription processing on the target file, the target text is obtained, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file during file transcription, and thus solving the technical problems of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the relevant technology when transcribing the target file.

[0115] In an optional embodiment, the data processing method in the embodiment of the present application further includes:

[0116] Step S63, in response to the editing operation performed on the target text, obtaining an edited text;

[0117] Step S64, determining a target loss based on the target text and the edited text;

[0118] Step S65: fine-tune the model parameters of the target text generation model using the target loss to obtain an adjusted text generation model.

[0119] Based on the above optional embodiment, by responding to an edit operation performed on the target text, an edited text is obtained, and then a target loss is determined based on the target text and the edited text. Finally, the target loss is used to fine-tune the model parameters of the target text generation model to obtain an adjusted text generation model, thereby improving the generation quality and accuracy of the text generation model. By performing edit operations on the target text and calculating the target loss, the model parameters can be adjusted more accurately, so that the text generated by the text generation model better meets the expected requirements, thereby improving the effectiveness and performance of the text generation model.

[0120] In an optional embodiment, the data processing method in the embodiment of the present application also includes: obtaining sample generation instructions and a sample data set, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data set covers multiple types of labels; using the sample generation instructions and the sample data set to train the initial text generation model to generate a target text generation model.

[0121] The sample generation instructions may be a set of instructions or rules for generating a summary, which are used to determine the type, format, content, and other characteristics of the text to be generated. The sample generation instructions may include, but are not limited to, requirements such as text length, language style, and subject matter, so as to generate a text sample that meets specific requirements.

[0122] The above-mentioned sample data set can be a set of existing text samples that cover multiple types of labels and can be used to train models, test algorithms, or analyze text features. The sample data set can include various types of transcribed texts, as well as corresponding labels or classification information. For example, the goal of summarization is to automatically extract and generate a concise, coherent summary that retains the main information of the original content from long speeches, meeting minutes, or any other form of text. Therefore, the sample data set also includes the source data content of the speech, meeting minutes, etc. of the summary target.

[0123] Furthermore, training the initial text generation model using sample generation instructions and a sample dataset can improve its generation capabilities and accuracy, enabling it to better generate target text. During training, the initial text generation model learns the language patterns, vocabulary usage, and syntactic structure in the sample dataset, enabling it to more accurately generate content that meets the requirements of the target text. Through continuous iterative training, the generation performance of the initial text generation model will gradually improve, achieving higher generation quality and accuracy.

[0124] Based on the above optional embodiments, by obtaining sample generation instructions and sample data sets, and then using the sample generation instructions and sample data sets to train the initial text generation model to generate a target text generation model, it can help improve the quality and accuracy of the initial text generation model, so that it can better understand and generate natural language text. The trained target text generation model can generate more fluent, accurate and logical text content, thereby improving the practicality and reliability of text transcription.

[0125] For the parts not described in detail in the above embodiments of the present application, please refer to the relevant description of Example 1 and will not be repeated here.

[0126] Example 3

[0127] According to an embodiment of the present application, a data processing method is also provided. Figure 7 This is a flow chart of a data processing method according to Example 3 of the present application. Figure 7 As shown, the method may include the following steps:

[0128] Step S71, obtaining the conference video file to be transcribed;

[0129] Step S72: Use a conference text generation model to perform text transcription processing on the conference video file to obtain a conference summary; wherein, the conference text generation model is obtained after targeted reinforcement training according to a conference condition constraint method, and the conference condition constraint method is used to generate model input instructions that are adapted to the file content of the conference video file from multiple different dimensions.

[0130] Based on the above steps S71 to S72, by obtaining the conference video file to be transcribed, and then using the conference text generation model to perform text transcription on the conference video file, a conference summary is obtained, thereby achieving the purpose of efficiently and accurately transcribing the target file into text, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problems of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the relevant technology when transcribing the target file.

[0131] In an optional embodiment, the above method may further include the following steps:

[0132] Step S73, in response to the editing operation performed on the conference summary, obtaining an edited summary;

[0133] Step S74, determining the target loss based on the meeting summary and the edited summary;

[0134] Step S75 , fine-tuning the model parameters of the conference text generation model using the target loss to obtain an adjusted conference text generation model.

[0135] In an optional embodiment, a feedback mechanism can be set up for the conference summary obtained after the conference text generation model performs text transcription processing on the conference video file. That is, the user can perform a second edit on the conference summary so as to rewrite the relevant parts of the conference summary that involve problems such as inaccurate descriptions or irregular formats to obtain an edited summary. Therefore, the target loss is further determined based on the conference summary and the edited summary, and the target loss is used to fine-tune the model parameters of the conference text generation model to obtain an adjusted conference text generation model. This process is actually equivalent to a new round of iterative training of the above-mentioned conference text generation model, so that the summary generated subsequently by the adjusted conference text generation model can better meet the user's expectations.

[0136] For the parts not described in detail in the above embodiments of the present application, please refer to the relevant description of Example 1 and will not be repeated here.

[0137] Example 4

[0138] According to an embodiment of the present application, a data processing method is also provided. Figure 8 This is a flow chart of a data processing method according to Example 4 of the present application. Figure 8 As shown, the method may include the following steps:

[0139] Step S81, obtaining a file processing request through a first application programming interface;

[0140] Step S82, returning a file processing response through a second application programming interface; wherein, the request data carried in the file processing request includes: a target file to be transcribed, and the response data carried in the file processing response includes: a target text, the target text is obtained by performing text transcription processing on the target file using a target text generation model, and the target text generation model is obtained by performing targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0141] Based on the above steps S81 to S82, a file processing request is obtained through the first application programming interface, and then a file processing response is returned through the second application programming interface, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file during file transcription, and thus solving the technical problems of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the relevant technology when transcribing the target file.

[0142] The first application programming interface and the second application programming interface can be the same application programming interface or different application programming interfaces. In an optional embodiment, the interface parameters in the first application programming interface and the second application programming interface can include, but are not limited to: an interface global identifier, an interface signature key, an interface timestamp, an interface request identifier, a system call credential identifier, etc. The first application programming interface can use GET or POST as an interface request method to obtain a file processing request. The second application programming interface can use JSON format to feedback a file processing response.

[0143] For the parts not described in detail in the above embodiments of the present application, please refer to the relevant description of Example 1 and will not be repeated here.

[0144] Example 5

[0145] According to an embodiment of the present application, a data processing method is also provided. Figure 9 This is a flow chart of a data processing method according to Example 5 of the present application. Figure 9 As shown, the method may include the following steps:

[0146] Step S91, obtaining the currently input file processing session request;

[0147] Step S92: Returning a file processing dialogue reply in response to the file processing dialogue request; wherein the request data carried in the file processing dialogue request includes: a target file to be transcribed; and the information carried in the file processing dialogue reply includes: a target text, which is obtained by transcribing the target file using a target text generation model, which is obtained by performing targeted reinforcement training according to a preset conditional constraint method, wherein the preset conditional constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions;

[0148] Step S93: Display the target text in the graphical user interface.

[0149] Based on the above steps S91 to S93, by obtaining the currently input file processing dialogue request, and then responding to the file processing dialogue request, returning the file processing dialogue reply, and finally displaying the target text in the graphical user interface, the purpose of efficiently and accurately transcribing the text of the target file is achieved, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problems of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the relevant technology when transcribing the target file.

[0150] For the parts not described in detail in the above embodiments of the present application, please refer to the relevant description of Example 1 and will not be repeated here.

[0151] Example 6

[0152] According to an embodiment of the present application, a file processing device for implementing the above data processing method is also provided. Figure 10 This is a structural block diagram of a file processing device according to embodiment 6 of the present application. Figure 10 As shown, the device includes:

[0153] Acquisition module 1001 is used to acquire a sample generation instruction and a sample data set, wherein the sample generation instruction is used to determine the type of text to be generated, and the sample data set covers multiple types of labels;

[0154] A training module 1002 is configured to perform targeted reinforcement training on the initial text generation model using sample generation instructions and a sample data set in accordance with a preset conditional constraint method to generate a target text generation model, wherein the preset conditional constraint method is configured to generate model input instructions adapted to the file content of the target file to be transcribed from multiple different dimensions;

[0155] The processing module 1003 is used to perform text transcription processing on the target file using the target text generation model to obtain the target text.

[0156] Optionally, the acquisition module 1001 is further configured to acquire a sample generation task, wherein the sample generation task is a text generation task corresponding to the type of text to be generated; and determine the sample generation instruction based on the task seed instruction corresponding to the sample generation task.

[0157] Optionally, the training module 1002 is also used to: constrain evolution of the sample generation instruction according to preset condition constraints to obtain constrained evolution instructions; use the constrained evolution instructions and the sample data set to perform targeted enhancement training on the initial text generation model to generate a target text generation model.

[0158] Optionally, the training module 1002 is also used to: determine the target evolution direction according to a preset condition constraint method, wherein the target evolution direction includes: replacing the evolution direction and adding the evolution direction. The replacing evolution direction is used to replace the initial constraint conditions of the sample generation instruction, and the adding evolution direction is used to add new constraint conditions based on the initial constraint conditions; based on the target evolution direction, the sample generation instruction is constrained to evolve to obtain a constrained evolution instruction.

[0159] Optionally, the training module 1002 is also used to: train the initial text generation model using the constrained evolution instruction and the sample data set to obtain the training results corresponding to the constrained evolution instruction; continue to iteratively train the initial text generation model based on the constrained evolution instruction and the training results corresponding to the constrained evolution instruction until all the sample data in the sample data set are used up, and generate the target text generation model.

[0160] Optionally, the training module 1002 is also used to: continue to iteratively train the initial text generation model based on the constrained evolution instructions and the training results, and obtain the current round training results obtained in each round of training in turn; determine the target loss based on the current round training results and the actual results corresponding to the current round training results; use the target loss to fine-tune the model parameters of the intermediate text generation model to obtain the text generation model to be used in the next round of training, until all the sample data in the sample data set are used up, and generate the target text generation model, wherein the intermediate text generation model is the model obtained by the initial text generation model after completing the current round of training.

[0161] Optionally, the multiple different dimensions include at least some of the following dimensions: scenario dimension, condition dimension, and format dimension, wherein the scenario dimension is used to determine the usage scenario of the target text, the condition dimension is used to determine the generation condition of the target text, and the format dimension is used to determine the usage format of the target text.

[0162] It should be noted that the acquisition module 1001, training module 1002, and processing module 1003 correspond to steps S21 to S23 in Example 1. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The modules can also be run in a computer terminal as part of the device.

[0163] By adopting the embodiment of the present application, by obtaining sample generation instructions and sample data sets, and then according to the preset condition constraints, using the sample generation instructions and sample data sets to perform targeted reinforcement training on the initial text generation model, a target text generation model is generated, and finally the target text generation model is used to perform text transcription processing on the target file to obtain the target text, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related art when transcribing the target file.

[0164] Figure 11 This is a structural block diagram of another file processing device according to embodiment 6 of the present application. Figure 11 As shown, the device includes:

[0165] An acquisition module 1101 is used to acquire a target file to be transcribed;

[0166] The processing module 1102 is configured to perform text transcription processing on the target file using the target text generation model to obtain the target text;

[0167] Among them, the target text generation model is obtained after targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0168] Optionally, the data processing device also includes: a response module 1103, used to respond to the editing operation performed on the target text to obtain the edited text; a determination module 1104, used to determine the target loss based on the target text and the edited text; a fine-tuning module 1105, used to use the target loss to fine-tune the model parameters of the target text generation model to obtain an adjusted text generation model.

[0169] Optionally, the acquisition module 1101 is also used to obtain sample generation instructions and sample data sets, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data sets cover multiple types of labels; the file processing device also includes: a training module 11106, which is used to use the sample generation instructions and the sample data sets to train the initial text generation model to generate a target text generation model.

[0170] It should be noted that the acquisition module 1101 and the processing module 1102 correspond to steps S61 to S62 in Example 2. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 2. It should be noted that the modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The modules can also be run in a computer terminal as part of a device.

[0171] By adopting the embodiment of the present application, by obtaining the target file to be transcribed, and then using the target text generation model to perform text transcription processing on the target file, the target text is obtained, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problems of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related art when transcribing the target file.

[0172] Figure 12 This is a structural block diagram of another file processing device according to embodiment 6 of the present application. Figure 12 As shown, the device includes:

[0173] An acquisition module 1201 is used to acquire a conference video file to be transcribed;

[0174] Processing module 1202 is used to perform text transcription processing on the conference video file using the conference text generation model to obtain a conference summary;

[0175] Among them, the conference text generation model is obtained after targeted reinforcement training according to the conference condition constraint method. The conference condition constraint method is used to generate model input instructions that are adapted to the file content of the conference video file from multiple different dimensions.

[0176] Optionally, the data processing device also includes: a response module 1203, used to respond to the editing operation performed on the conference summary to obtain the edited summary; a determination module 1204, used to determine the target loss based on the conference summary and the edited summary; a fine-tuning module 1205, used to use the target loss to fine-tune the model parameters of the conference text generation model to obtain the adjusted conference text generation model.

[0177] It should be noted that the above-mentioned acquisition module 1201 and processing module 1202 correspond to steps S71 to S72 in Example 3. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 3. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The above-mentioned modules can also be run in a computer terminal as part of the device.

[0178] By adopting the embodiment of the present application, by obtaining the conference video file to be transcribed, and then using the conference text generation model to perform text transcription processing on the conference video file, a conference summary is obtained, thereby achieving the purpose of efficiently and accurately transcribing the target file into text, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related technology when transcribing the target file.

[0179] Figure 13 This is a structural block diagram of another file processing device according to embodiment 6 of the present application. Figure 13 As shown, the device includes:

[0180] An acquisition module 1301 is configured to acquire a file processing request through a first application programming interface;

[0181] Return module 1302, configured to return a file processing response via a second application programming interface;

[0182] Among them, the request data carried in the file processing request includes: the target file to be transcribed, and the response data carried in the file processing response includes: the target text. The target text is obtained by transcribing the target file using a target text generation model. The target text generation model is obtained after targeted reinforcement training according to a preset condition constraint method. The preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0183] It should be noted that the acquisition module 1301 and the return module 1302 correspond to steps S81 to S82 in Example 4. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 4. It should be noted that the modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The modules can also be part of a device and run in a computer terminal.

[0184] By adopting the embodiment of the present application, a file processing request is obtained through the first application programming interface, and a file processing response is returned through the second application programming interface, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related technology when transcribing the target file.

[0185] Figure 14 This is a structural block diagram of another file processing device according to embodiment 6 of the present application. Figure 14 As shown, the device includes:

[0186] An acquisition module 1401 is used to acquire a currently input file processing session request;

[0187] A return module 1402 is configured to return a file processing dialogue reply in response to the file processing dialogue request;

[0188] The request data carried in the file processing dialogue request includes the target file to be transcribed, and the information carried in the file processing dialogue reply includes the target text, which is obtained by transcribing the target file using a target text generation model. The target text generation model is obtained by performing targeted reinforcement training according to a preset condition constraint method. The preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0189] The display module 1403 is used to display the target text in the graphical user interface.

[0190] It should be noted that the acquisition module 1401, return module 1402, and display module 1403 correspond to steps S91 to S93 in Example 5. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 5. It should be noted that the modules or units can be hardware components or software components stored in a memory and processed by one or more processors. The modules can also be part of a device and run in a computer terminal.

[0191] By adopting the embodiment of the present application, by obtaining the currently input file processing dialogue request, responding to the file processing dialogue request, returning the file processing dialogue reply, and finally displaying the target text in the graphical user interface, the purpose of efficiently and accurately transcribing the text of the target file is achieved, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related technology when transcribing the target file.

[0192] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0193] Example 7

[0194] The embodiment of the present application can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.

[0195] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of a computer network.

[0196] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the data processing method: obtaining sample generation instructions and sample data sets, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data sets cover multiple types of labels; according to the preset condition constraint method, the sample generation instructions and the sample data sets are used to perform targeted reinforcement training on the initial text generation model to generate a target text generation model, wherein the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file to be transcribed from multiple different dimensions; and the target text generation model is used to perform text transcription processing on the target file to obtain the target text.

[0197] Optionally, Figure 15 1 is a block diagram of a computer terminal according to Embodiment 7 of the present application. As shown in the figure, the computer terminal may include: one or more (only one is shown in the figure) processors 152, a memory 154, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display.

[0198] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the data processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned data processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0199] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain sample generation instructions and sample data sets, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data sets cover multiple types of labels; according to the preset condition constraint method, the sample generation instructions and the sample data set are used to perform targeted reinforcement training on the initial text generation model to generate a target text generation model, wherein the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file to be transcribed from multiple different dimensions; and the target text generation model is used to perform text transcription processing on the target file to obtain the target text.

[0200] Optionally, the processor may further execute program code of the following steps: obtaining a sample generation task, wherein the sample generation task is a text generation task corresponding to the type of text to be generated; and determining a sample generation instruction based on a task seed instruction corresponding to the sample generation task.

[0201] Optionally, the processor may also execute the program code of the following steps: constraining the evolution of the sample generation instruction according to preset conditional constraints to obtain a constrained evolution instruction; using the constrained evolution instruction and the sample data set to perform targeted enhancement training on the initial text generation model to generate a target text generation model.

[0202] Optionally, the above-mentioned processor can also execute the program code of the following steps: determine the target evolution direction according to the preset condition constraint method, wherein the target evolution direction includes: replacing the evolution direction and adding the evolution direction, the replacing evolution direction is used to replace the initial constraint conditions of the sample generation instruction, and the adding evolution direction is used to add new constraint conditions based on the initial constraint conditions; constrain the evolution of the sample generation instruction based on the target evolution direction to obtain a constrained evolution instruction.

[0203] Optionally, the above-mentioned processor can also execute the program code of the following steps: use the constrained evolution instruction and the sample data set to train the initial text generation model to obtain the training results corresponding to the constrained evolution instruction; continue to iteratively train the initial text generation model based on the constrained evolution instruction and the training results corresponding to the constrained evolution instruction until all the sample data in the sample data set are used up, and generate the target text generation model.

[0204] Optionally, the processor may also execute the following program code: continuing to iteratively train the initial text generation model based on the constrained evolution instructions and the training results, and sequentially obtaining the current round training results obtained from each round of training; determining the target loss based on the current round training results and the actual results corresponding to the current round training results; using the target loss to fine-tune the model parameters of the intermediate text generation model to obtain the text generation model to be used in the next round of training, until all the sample data in the sample data set are used up, and generating the target text generation model, wherein the intermediate text generation model is the model obtained by the initial text generation model after completing the current round of training.

[0205] Optionally, the multiple different dimensions include at least some of the following dimensions: scenario dimension, condition dimension, and format dimension, wherein the scenario dimension is used to determine the usage scenario of the target text, the condition dimension is used to determine the generation condition of the target text, and the format dimension is used to determine the usage format of the target text.

[0206] Optionally, the processor may also execute the program code of the following steps: obtaining a target file to be transcribed; performing text transcription processing on the target file using a target text generation model to obtain a target text; wherein the target text generation model is obtained after performing targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0207] Optionally, the processor may also execute the program code of the following steps: responding to an editing operation performed on the target text to obtain an edited text; determining a target loss based on the target text and the edited text; and using the target loss to fine-tune the model parameters of the target text generation model to obtain an adjusted text generation model.

[0208] Optionally, the above-mentioned processor can also execute the program code of the following steps: obtaining sample generation instructions and sample data sets, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data sets cover multiple types of labels; using the sample generation instructions and sample data sets to train the initial text generation model to generate a target text generation model.

[0209] Optionally, the processor may also execute the program code of the following steps: obtaining a conference video file to be transcribed; performing text transcription processing on the conference video file using a conference text generation model to obtain a conference summary; wherein the conference text generation model is obtained after targeted reinforcement training according to a conference condition constraint method, and the conference condition constraint method is used to generate model input instructions that are adapted to the file content of the conference video file from multiple different dimensions.

[0210] Optionally, the processor may also execute the program code of the following steps: responding to an editing operation performed on the conference summary to obtain an edited summary; determining a target loss based on the conference summary and the edited summary; and using the target loss to fine-tune the model parameters of the conference text generation model to obtain an adjusted conference text generation model.

[0211] Optionally, the above-mentioned processor can also execute the program code of the following steps: obtain a file processing request through a first application programming interface; return a file processing response through a second application programming interface; wherein, the request data carried in the file processing request includes: the target file to be transcribed, and the response data carried in the file processing response includes: the target text, the target text is obtained by using a target text generation model to perform text transcription processing on the target file, and the target text generation model is obtained after targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0212] Optionally, the processor may also execute the following program code: obtaining a currently input file processing dialogue request; returning a file processing dialogue reply in response to the file processing dialogue request; wherein the request data carried in the file processing dialogue request includes: the target file to be transcribed, and the information carried in the file processing dialogue reply includes: the target text, the target text is obtained by performing text transcription processing on the target file using a target text generation model, the target text generation model is obtained by performing targeted reinforcement training according to a preset condition constraint method, the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions; and the target text is displayed in a graphical user interface.

[0213] By adopting the embodiment of the present application, by obtaining sample generation instructions and sample data sets, and then according to the preset condition constraints, using the sample generation instructions and sample data sets to perform targeted reinforcement training on the initial text generation model, a target text generation model is generated, and finally the target text generation model is used to perform text transcription processing on the target file to obtain the target text, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related art when transcribing the target file.

[0214] It can be understood by those skilled in the art that Figure 15 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 15 It does not limit the structure of the above electronic device. For example, the computer terminal may also include Figure 15 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 15 Different configurations shown.

[0215] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0216] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0217] Example 8

[0218] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the data processing method provided in the first embodiment.

[0219] Optionally, in this embodiment, the above-mentioned storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.

[0220] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: obtaining sample generation instructions and a sample data set, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data set covers multiple types of labels; according to a preset condition constraint method, the sample generation instructions and the sample data set are used to perform targeted enhancement training on the initial text generation model to generate a target text generation model, wherein the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file to be transcribed from multiple different dimensions; and the target text generation model is used to perform text transcription processing on the target file to obtain the target text.

[0221] Optionally, the storage medium can also be configured to store program code for executing the following steps: obtaining a sample generation task, wherein the sample generation task is a text generation task corresponding to the type of text to be generated; and determining a sample generation instruction based on a task seed instruction corresponding to the sample generation task.

[0222] Optionally, the storage medium can also be configured to store program code for executing the following steps: constraining the evolution of the sample generation instructions in accordance with preset conditional constraints to obtain constrained evolution instructions; using the constrained evolution instructions and the sample data set to perform targeted enhancement training on the initial text generation model to generate a target text generation model.

[0223] Optionally, the storage medium can also be configured to store program code for executing the following steps: determining the target evolution direction in accordance with preset conditional constraints, wherein the target evolution direction includes: replacing the evolution direction and adding the evolution direction, the replacing evolution direction is used to replace the initial constraint conditions of the sample generation instruction, and the adding evolution direction is used to add new constraint conditions based on the initial constraint conditions; constraining the evolution of the sample generation instruction based on the target evolution direction to obtain a constrained evolution instruction.

[0224] Optionally, the storage medium can also be configured to store program code for executing the following steps: training the initial text generation model using constrained evolution instructions and a sample data set to obtain training results corresponding to the constrained evolution instructions; continuing to iteratively train the initial text generation model based on the constrained evolution instructions and the training results corresponding to the constrained evolution instructions until all the sample data in the sample data set have been used up, thereby generating a target text generation model.

[0225] Optionally, the storage medium can also be configured to store program code for executing the following steps: continuing to iteratively train the initial text generation model based on the constrained evolution instructions and training results, and obtaining the current round training results obtained in each round of training in turn; determining the target loss based on the current round training results and the actual results corresponding to the current round training results; using the target loss to fine-tune the model parameters of the intermediate text generation model to obtain the text generation model to be used in the next round of training, until all the sample data in the sample data set are used up, and generating the target text generation model, wherein the intermediate text generation model is the model obtained by the initial text generation model after completing the current round of training.

[0226] Optionally, the multiple different dimensions include at least some of the following dimensions: scenario dimension, condition dimension, and format dimension, wherein the scenario dimension is used to determine the usage scenario of the target text, the condition dimension is used to determine the generation condition of the target text, and the format dimension is used to determine the usage format of the target text.

[0227] Optionally, the storage medium can also be configured to store program code for executing the following steps: obtaining a target file to be transcribed; performing text transcription processing on the target file using a target text generation model to obtain a target text; wherein the target text generation model is obtained after targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0228] Optionally, the storage medium can also be configured to store program code for executing the following steps: responding to an editing operation performed on a target text to obtain an edited text; determining a target loss based on the target text and the edited text; and using the target loss to fine-tune the model parameters of a target text generation model to obtain an adjusted text generation model.

[0229] Optionally, the storage medium can also be configured to store program code for executing the following steps: obtaining sample generation instructions and a sample data set, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data set covers multiple types of labels; using the sample generation instructions and the sample data set to train the initial text generation model to generate a target text generation model.

[0230] Optionally, the storage medium can also be configured to store program code for executing the following steps: obtaining a conference video file to be transcribed; using a conference text generation model to perform text transcription processing on the conference video file to obtain a conference summary; wherein the conference text generation model is obtained after targeted reinforcement training according to a conference condition constraint method, and the conference condition constraint method is used to generate model input instructions that are adapted to the file content of the conference video file from multiple different dimensions.

[0231] Optionally, the storage medium can also be configured to store program code for performing the following steps: responding to an editing operation performed on the conference summary to obtain an edited summary; determining a target loss based on the conference summary and the edited summary; and using the target loss to fine-tune the model parameters of the conference text generation model to obtain an adjusted conference text generation model.

[0232] Optionally, the storage medium can also be configured to store program code for executing the following steps: obtaining a file processing request through a first application programming interface; returning a file processing response through a second application programming interface; wherein the request data carried in the file processing request includes: the target file to be transcribed, and the response data carried in the file processing response includes: the target text, the target text is obtained by performing text transcription processing on the target file using a target text generation model, and the target text generation model is obtained after performing targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0233] Optionally, the storage medium can also be configured to store program code for executing the following steps: obtaining a currently input file processing dialogue request; returning a file processing dialogue reply in response to the file processing dialogue request; wherein the request data carried in the file processing dialogue request includes: the target file to be transcribed, and the information carried in the file processing dialogue reply includes: the target text, the target text is obtained by transcribing the target file using a target text generation model, the target text generation model is obtained by performing targeted reinforcement training according to a preset condition constraint method, the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions; and the target text is displayed in a graphical user interface.

[0234] By adopting the embodiment of the present application, by obtaining sample generation instructions and sample data sets, and then according to the preset condition constraints, using the sample generation instructions and sample data sets to perform targeted reinforcement training on the initial text generation model, a target text generation model is generated, and finally the target text generation model is used to perform text transcription processing on the target file to obtain the target text, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related art when transcribing the target file.

[0235] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0236] Example 9

[0237] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and when the computer program is executed by a processor, the method provided in the embodiment is implemented.

[0238] Optionally, the computer program included in the above-mentioned computer program product is executed by the processor to perform the following steps: obtain sample generation instructions and a sample data set, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data set covers multiple types of labels; according to a preset condition constraint method, the sample generation instructions and the sample data set are used to perform targeted reinforcement training on the initial text generation model to generate a target text generation model, wherein the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file to be transcribed from multiple different dimensions; and the target text generation model is used to perform text transcription processing on the target file to obtain the target text.

[0239] Optionally, the computer program included in the above-mentioned computer program product is used by the processor to execute the following steps: obtaining a sample generation task, wherein the sample generation task is a text generation task corresponding to the type of text to be generated; and determining a sample generation instruction based on a task seed instruction corresponding to the sample generation task.

[0240] Optionally, the computer program included in the above-mentioned computer program product is executed by the processor to perform the following steps: constraining evolution of sample generation instructions in accordance with preset conditional constraints to obtain constrained evolution instructions; using the constrained evolution instructions and the sample data set to perform targeted reinforcement training on the initial text generation model to generate a target text generation model.

[0241] Optionally, the computer program included in the above-mentioned computer program product is executed by the processor to perform the following steps: determine the target evolution direction according to the preset condition constraint method, wherein the target evolution direction includes: replacing the evolution direction and adding the evolution direction, the replacing evolution direction is used to replace the initial constraint conditions of the sample generation instruction, and the adding evolution direction is used to add new constraint conditions based on the initial constraint conditions; constrain the evolution of the sample generation instruction based on the target evolution direction to obtain a constrained evolution instruction.

[0242] Optionally, the computer program included in the above-mentioned computer program product is used by the processor to execute the following steps: using constrained evolution instructions and a sample data set to train the initial text generation model to obtain training results corresponding to the constrained evolution instructions; continuing to iteratively train the initial text generation model based on the constrained evolution instructions and the training results corresponding to the constrained evolution instructions until all the sample data in the sample data set have been used up, thereby generating a target text generation model.

[0243] Optionally, the computer program included in the above-mentioned computer program product is used by the processor to execute the following steps: continue to iteratively train the initial text generation model based on the constrained evolution instruction and the training results, and obtain the current round training results obtained in each round of training in turn; determine the target loss based on the current round training results and the actual results corresponding to the current round training results; use the target loss to fine-tune the model parameters of the intermediate text generation model to obtain the text generation model to be used in the next round of training, until all the sample data in the sample data set are used up, and generate the target text generation model, wherein the intermediate text generation model is the model obtained by the initial text generation model after completing the current round of training.

[0244] Optionally, the multiple different dimensions include at least some of the following dimensions: scenario dimension, condition dimension, and format dimension, wherein the scenario dimension is used to determine the usage scenario of the target text, the condition dimension is used to determine the generation condition of the target text, and the format dimension is used to determine the usage format of the target text.

[0245] Optionally, the computer program included in the above-mentioned computer program product is used by the processor to execute the following steps: obtaining a target file to be transcribed; using a target text generation model to perform text transcription processing on the target file to obtain a target text; wherein the target text generation model is obtained after targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0246] Optionally, the computer program included in the above-mentioned computer program product is used by the processor to execute the following steps: responding to an editing operation performed on the target text to obtain an edited text; determining a target loss based on the target text and the edited text; and using the target loss to fine-tune the model parameters of the target text generation model to obtain an adjusted text generation model.

[0247] Optionally, the computer program included in the above-mentioned computer program product is executed by the processor to perform the following steps: obtain sample generation instructions and a sample data set, wherein the sample generation instructions are used to determine the type of text to be generated, and the sample data set covers multiple types of labels; use the sample generation instructions and the sample data set to train the initial text generation model to generate a target text generation model.

[0248] Optionally, the computer program included in the above-mentioned computer program product is used by the processor to execute the following steps: obtaining a conference video file to be transcribed; using a conference text generation model to perform text transcription processing on the conference video file to obtain a conference summary; wherein the conference text generation model is obtained after targeted reinforcement training according to a conference condition constraint method, and the conference condition constraint method is used to generate model input instructions that are adapted to the file content of the conference video file from multiple different dimensions.

[0249] Optionally, the computer program included in the above-mentioned computer program product is used by the processor to execute the following steps: responding to an editing operation performed on the conference summary to obtain an edited summary; determining a target loss based on the conference summary and the edited summary; and using the target loss to fine-tune the model parameters of the conference text generation model to obtain an adjusted conference text generation model.

[0250] Optionally, the computer program included in the above-mentioned computer program product is executed by the processor to perform the following steps: obtain a file processing request through a first application programming interface; return a file processing response through a second application programming interface; wherein the request data carried in the file processing request includes: the target file to be transcribed, and the response data carried in the file processing response includes: the target text, the target text is obtained by using a target text generation model to perform text transcription processing on the target file, and the target text generation model is obtained after performing targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

[0251] Optionally, the computer program included in the above-mentioned computer program product is executed by the processor to perform the following steps: obtain the currently input file processing dialogue request; return a file processing dialogue reply in response to the file processing dialogue request; wherein the request data carried in the file processing dialogue request includes: the target file to be transcribed, and the information carried in the file processing dialogue reply includes: the target text, the target text is obtained by using a target text generation model to perform text transcription processing on the target file, and the target text generation model is obtained after targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions; and display the target text in a graphical user interface.

[0252] By adopting the embodiment of the present application, by obtaining sample generation instructions and sample data sets, and then according to the preset condition constraints, using the sample generation instructions and sample data sets to perform targeted reinforcement training on the initial text generation model, a target text generation model is generated, and finally the target text generation model is used to perform text transcription processing on the target file to obtain the target text, thereby achieving the purpose of efficiently and accurately transcribing the text of the target file, thereby achieving the technical effect of improving the processing efficiency, accuracy and adaptability of the target file when transcribing the file, and thus solving the technical problem of low processing efficiency, poor accuracy and adaptability of the file processing method provided in the related art when transcribing the target file.

[0253] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.

[0254] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0255] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0256] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0257] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0258] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0259] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0260] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data processing method, characterized in that: include: Obtaining a sample generation instruction and a sample data set, wherein the sample generation instruction is used to determine the type of text to be generated, and the sample data set covers multiple types of labels; According to a preset condition constraint method, the sample generation instruction and the sample data set are used to perform targeted reinforcement training on the initial text generation model to generate a target text generation model, wherein the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file to be transcribed from multiple different dimensions; The target text generation model is used to perform text transcription processing on the target file to obtain the target text.

2. The data processing method according to claim 1, wherein: Obtaining the sample generation instruction includes: Obtaining a sample generation task, wherein the sample generation task is a text generation task corresponding to the type of text to be generated; The sample generation instruction is determined based on a task seed instruction corresponding to the sample generation task.

3. The data processing method according to claim 1, wherein: Performing targeted enhancement training on the initial text generation model using the sample generation instruction and the sample data set to generate the target text generation model includes: Performing constraint evolution on the sample generation instruction according to the preset condition constraint method to obtain a constraint evolution instruction; The constrained evolution instruction and the sample data set are used to perform directed enhancement training on the initial text generation model to generate the target text generation model.

4. The data processing method according to claim 3, wherein: Constraining and evolving the sample generation instruction according to the preset condition constraint method to obtain the constraint evolution instruction includes: Determining a target evolution direction according to the preset condition constraint method, wherein the target evolution direction includes: a replacement evolution direction and an addition evolution direction, wherein the replacement evolution direction is used to replace the initial constraint condition of the sample generation instruction, and the addition evolution direction is used to add a new constraint condition based on the initial constraint condition; Constrained evolution is performed on the sample generation instruction based on the target evolution direction to obtain the constrained evolution instruction.

5. The data processing method according to claim 3, wherein: Performing directed enhancement training on the initial text generation model using the constrained evolution instruction and the sample data set to generate the target text generation model includes: Using the constrained evolution instruction and the sample data set to train the initial text generation model, and obtain a training result corresponding to the constrained evolution instruction; The initial text generation model is iteratively trained based on the constrained evolution instruction and the training result corresponding to the constrained evolution instruction until all the sample data in the sample data set are used up, thereby generating the target text generation model.

6. The data processing method according to claim 5, characterized in that: Continuing to iteratively train the initial text generation model based on the constrained evolution instruction and the training result until all sample data in the sample data set are used up, generating the target text generation model includes: Continuing to iteratively train the initial text generation model based on the constrained evolution instruction and the training result, and sequentially obtaining the current round training results obtained from each round of training; Determining a target loss based on the current round training result and a true result corresponding to the current round training result; The target loss is used to fine-tune the model parameters of the intermediate text generation model to obtain the text generation model to be used in the next round of training, until all the sample data in the sample data set are used up, and the target text generation model is generated, wherein the intermediate text generation model is the model obtained by the initial text generation model after completing the current round of training.

7. The data processing method according to claim 1, wherein: The multiple different dimensions include at least some of the following dimensions: scenario dimension, condition dimension, and format dimension, wherein the scenario dimension is used to determine the usage scenario of the target text, the condition dimension is used to determine the generation condition of the target text, and the format dimension is used to determine the usage format of the target text.

8. A data processing method, characterized in that: include: Get the target file to be transcribed; Using a target text generation model to perform text transcription processing on the target file to obtain a target text; The target text generation model is obtained after targeted enhancement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

9. The data processing method according to claim 8, characterized in that: The data processing method further includes: In response to the editing operation performed on the target text, an edited text is obtained; determining a target loss based on the target text and the edited text; The target loss is used to fine-tune the model parameters of the target text generation model to obtain an adjusted text generation model.

10. The data processing method according to claim 8, characterized in that: The data processing method further includes: Obtaining a sample generation instruction and a sample data set, wherein the sample generation instruction is used to determine the type of text to be generated, and the sample data set covers multiple types of labels; The sample generation instruction and the sample data set are used to train an initial text generation model to generate the target text generation model.

11. A data processing method, characterized in that: include: Obtain the conference video file to be transcribed; Using a conference text generation model to perform text transcription on the conference video file to obtain a conference summary; Among them, the conference text generation model is obtained after targeted reinforcement training according to the conference condition constraint method, and the conference condition constraint method is used to generate model input instructions that are adapted to the file content of the conference video file from multiple different dimensions.

12. The data processing method according to claim 11, characterized in that: The data processing method further includes: In response to an editing operation performed on the conference summary, obtaining an edited summary; determining target loss based on the meeting summary and the edited summary; The target loss is used to fine-tune the model parameters of the conference text generation model to obtain an adjusted conference text generation model.

13. A data processing method, characterized in that: include: obtaining a file processing request through a first application programming interface; returning a file processing response via a second application programming interface; Among them, the request data carried in the file processing request includes: the target file to be transcribed, and the response data carried in the file processing response includes: the target text, which is obtained by using a target text generation model to perform text transcription processing on the target file, and the target text generation model is obtained after performing targeted reinforcement training according to a preset condition constraint method, and the preset condition constraint method is used to generate model input instructions that are adapted to the file content of the target file from multiple different dimensions.

14. An electronic device, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the data processing method according to any one of claims 1 to 13 when running.

15. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the data processing method according to any one of claims 1 to 13.

16. A computer program product, characterized in that The computer program comprises a computer program which, when executed by a processor, implements the data processing method according to any one of claims 1 to 13.