Dialogue text parsing methods, devices, electronic devices, and computer-readable media

By annotating and manually correcting the dialogue text information set of goods circulation, a training dataset is generated to train the base large model in a targeted manner. This solves the problem of uncontrollable parsing results of general large language models, and achieves high-quality dialogue text information extraction and reduces bandwidth resource waste.

CN122133848APending Publication Date: 2026-06-02HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2026-04-27
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, when dialogue text is parsed by directly calling a general large language model, the generated results are arbitrary and uncontrollable, resulting in chaotic parsed information formats, easy omission of key information and errors, and increased waste of bandwidth resources.

Method used

By acquiring a set of dialogue text information about the flow of goods, labeling and manually reviewing and correcting it, training and testing datasets are generated. A large-scale base model is then trained to obtain a dialogue text information extraction model, which outputs parsed information that meets the requirements.

Benefits of technology

It improves the accuracy and reliability of dialogue text information extraction, reduces the possibility of regenerating and sending parsed information, and reduces the waste of broadband resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133848A_ABST
    Figure CN122133848A_ABST
Patent Text Reader

Abstract

This disclosure presents embodiments of a dialogue text parsing method, apparatus, electronic device, and computer-readable medium. One specific implementation of the method includes: acquiring a set of dialogue text information related to the flow of goods and inputting it into a pre-trained dialogue analysis model to obtain a labeled dataset; sending the dialogue text information set and the labeled dataset to a preset manual processing terminal to obtain a calibration dataset; generating a training dataset and a test dataset; training a preset base model to obtain a trained dialogue text information extraction model; acquiring dialogue text information to be processed and inputting it into the dialogue text information extraction model to obtain extracted information; parsing the extracted information to obtain outputtable parsed information; and sending the dialogue text information to be processed and the outputtable parsed information to a preset downstream business system. This implementation reduces the waste of broadband resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically to methods, apparatus, electronic devices, and computer-readable media for parsing dialogue text. Background Technology

[0002] With the development of applications such as intelligent customer service and telephone sales quality inspection, the automated analysis and evaluation of dialogue content between customer service representatives and customers has become a key requirement. Dialogue text parsing is a technology for structured understanding of dialogue text. Currently, the common approach to parsing dialogue text information is to directly call a general large language model for text analysis using prompt words.

[0003] However, when parsing dialogue text using the above method, the following technical problems often arise: By directly calling a general-purpose language model using prompt words for text analysis, the lack of targeted training for the model leads to unpredictable and arbitrary generation of results. This results in chaotic output parsed information formats, the potential for missing key information, or the creation of "illusory" information. Consequently, the parsed information sent to downstream business systems contains errors. These erroneous parsed information may be meaningless in the downstream systems, but still needs to be transmitted over the network. This increases the likelihood of regenerating and resending the dialogue text parsing information, leading to a waste of bandwidth resources.

[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention

[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0006] Some embodiments of this disclosure provide methods, apparatuses, electronic devices, and computer-readable media for parsing dialogue text to address one or more of the technical problems mentioned in the background section above.

[0007] In a first aspect, some embodiments of this disclosure provide a dialogue text parsing method, which includes: acquiring a set of dialogue text information on the flow of goods; inputting the set of dialogue text information on the flow of goods into a pre-trained dialogue analysis model to obtain a labeled dataset, wherein each labeled data in the labeled dataset includes a dialogue summary, tag text, keyword information, and flow result information; sending the set of dialogue text information on the flow of goods and the labeled dataset to a preset manual processing terminal for manual review and correction processing to obtain a corrected dataset, wherein each corrected data in the corrected dataset corresponds to a labeled data in the labeled dataset; generating a training dataset and a test dataset based on the corrected dataset and the set of dialogue text information on the flow of goods; training a preset base model based on the training dataset and the test dataset to obtain a trained dialogue text information extraction model; acquiring dialogue text information to be processed; inputting the dialogue text information to be processed into the dialogue text information extraction model to obtain extracted information; parsing the extracted information to obtain outputtable parsing information; and sending the dialogue text information to be processed and the outputtable parsing information to a preset downstream business system.

[0008] Secondly, some embodiments of this disclosure provide a dialogue text parsing apparatus, comprising: a first acquisition unit configured to acquire a set of dialogue text information on the flow of goods; a first input unit configured to input the set of dialogue text information on the flow of goods into a pre-trained dialogue analysis model to obtain a labeled dataset, wherein each labeled data in the labeled dataset includes a dialogue summary, tag text, keyword information, and flow result information; and a sending unit configured to send the set of dialogue text information on the flow of goods and the labeled dataset to a preset manual processing terminal for manual review and correction processing to obtain a corrected dataset, wherein each corrected data in the corrected dataset is consistent with the data in the labeled dataset. A labeled data correspondence; a generation unit is configured to generate a training dataset and a test dataset based on the aforementioned calibration dataset and the aforementioned item flow dialogue text information set; a training unit is configured to train a preset base large model based on the aforementioned training dataset and the aforementioned test dataset to obtain a trained dialogue text information extraction model; a second acquisition unit is configured to acquire dialogue text information to be processed; a second input unit is configured to input the aforementioned dialogue text information to be processed into the aforementioned dialogue text information extraction model to obtain extracted information; a sending unit is configured to parse the aforementioned extracted information to obtain output parsing information and send the aforementioned dialogue text information to be processed and the aforementioned output parsing information to a preset downstream business system.

[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0011] The above-described embodiments of this disclosure have the following beneficial effects: the dialogue text parsing method of some embodiments of this disclosure reduces the waste of broadband resources. Specifically, the reason for the waste of broadband resources is that by directly calling a general large language model for text analysis using prompt words, the generation results are arbitrary and uncontrollable due to the lack of targeted training of the model. This leads to a chaotic format of the output parsed information, which is prone to omitting key information or generating "illusionary" information. As a result, there are errors in the parsed information sent to downstream business systems. The erroneous parsed information may not have practical significance in the downstream system, but it still needs to be transmitted through the network, which increases the possibility of regenerating and sending dialogue text parsing information, thus wasting bandwidth resources. Based on this, the dialogue text parsing method of some embodiments of this disclosure first obtains a set of item flow dialogue text information. Then, the above-mentioned item flow dialogue text information set is input into a pre-trained dialogue analysis model to obtain a labeled dataset, wherein each labeled data in the above-mentioned labeled dataset includes a dialogue summary, tag text, keyword information, and flow result information. Then, the aforementioned set of dialogue text information about the item flow and the aforementioned labeled dataset are sent to a pre-set manual processing terminal for manual review and correction, resulting in a corrected dataset. Each corrected data point in the corrected dataset corresponds to a labeled data point in the labeled dataset. This yields a manually corrected dataset that meets the requirements. Next, based on the corrected dataset and the aforementioned set of dialogue text information about the item flow, training and testing datasets are generated. Then, based on the training and testing datasets, a pre-set large-scale model is trained to obtain a trained dialogue text extraction model. Thus, the training and testing datasets generated from the manually corrected data can be used to specifically train the model, resulting in a dialogue text extraction model that extracts information from dialogue text, outputting extracted information that better meets the parsing requirements. Then, the dialogue text information to be processed is obtained. Then, the dialogue text information to be processed is input into the dialogue text extraction model to obtain extracted information. This extracted information is then used for subsequent parsing processing. Finally, the extracted information is parsed to obtain outputtable parsing information, and the dialogue text information to be processed and the outputtable parsing information are sent to the preset downstream business system. Thus, outputtable parsing information reflecting the dialogue summary, tag text, keyword information, and flow result information in the dialogue text information to be processed can be obtained and sent to the preset downstream business system. This is because the process involves first performing dialogue analysis on the item flow dialogue text information set to obtain an annotated dataset, and then manually reviewing and correcting the annotated data to obtain a corrected dataset that meets the requirements.Furthermore, training and correction datasets generated from targeted and corrected data of the dialogue text information of the goods flow are used to train the dialogue of the pre-set base large model. This reduces the interference of cluttered information on model learning, improves the accuracy and reliability of extracting key information, and thus provides downstream business systems with high-quality parsed information that can be directly processed and applied. This reduces the possibility of regenerating and sending dialogue text parsing information, thereby reducing the waste of broadband resources. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 This is a flowchart of some embodiments of the dialogue text parsing method according to the present disclosure; Figure 2 This is a schematic diagram of the structure of some embodiments of the dialogue text parsing apparatus according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0016] Figure 1 A flow 100 of some embodiments of a dialogue text parsing method according to the present disclosure is shown. The dialogue text parsing method includes the following steps: Step 101: Obtain the set of dialogue text information for item transfer.

[0017] In some embodiments, the executing entity of the dialogue text parsing method (e.g., a computing device) can obtain the item flow dialogue text information set via a wired or wireless connection. Each item flow dialogue text piece in the aforementioned item flow dialogue text information set can be text generated after voice-to-text conversion of the call content. In practice, the executing entity can obtain the item flow dialogue text information set from a storage device. This storage device can be a hard disk drive, solid-state drive, or the like connected to the executing entity.

[0018] It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future wireless connection methods.

[0019] Step 102: Input the item flow dialogue text information set into the pre-trained dialogue analysis model to obtain the labeled dataset.

[0020] In some embodiments, the aforementioned executing entity can input the aforementioned item flow dialogue text information set into a pre-trained dialogue analysis model to obtain a labeled dataset. Each labeled data in the labeled dataset includes a dialogue summary, tag text, keyword information, and flow result information. Each labeled data in the labeled dataset corresponds to one item flow dialogue text in the aforementioned item flow dialogue text information set. The aforementioned dialogue analysis model can be a general large model that takes item flow dialogue text information as input and labeled data as output (such as Qwen2.5-32B-Instruct). The aforementioned dialogue summary can be text content representing the item flow process in the item flow dialogue text information (for example, the aforementioned dialogue summary could be: (Agent: Hello, Mr. Wang! I am customer service from the credit card center… We have now launched an X customer: What are the application requirements?… Goodbye!)). The aforementioned tag text can be at least one tag selected by the dialogue analysis model from a preset tag list based on the item flow dialogue text information. Each tag in the aforementioned preset tag list can be pre-set words representing key points of the conversation of the person receiving the call (e.g., "high cost", "not needed"). The aforementioned keyword information can be at least one text containing key points of the dialogue content selected by the dialogue analysis model from the dialogue text information of the item flow. The aforementioned flow result information can be information indicating the result of the dialogue, which can be "yes" or "no" (for example, the aforementioned flow result information can be "yes" for a successful credit card sales dialogue and "no" for an unsuccessful sales dialogue).

[0021] Step 103: Send the item flow dialogue text information set and the labeled dataset to the preset manual processing terminal for manual review and correction to obtain the corrected dataset.

[0022] In some embodiments, the executing entity may send the item flow dialogue text information set and the labeled dataset to a preset manual processing terminal for manual review and correction to obtain a corrected dataset. Each corrected data point in the corrected dataset corresponds to a labeled data point in the labeled dataset. In practice, firstly, the executing entity may send the labeled dataset to the preset manual processing terminal. Then, the preset manual processing terminal randomly sends the item flow dialogue text information set and the labeled dataset to at least three preset manual processing sub-terminals. Next, each of the at least three preset manual processing sub-terminals, for each labeled data point in the labeled dataset, performs manual verification and correction of the tag text and keyword information in the corrected data based on the item flow dialogue text information corresponding to the labeled data (the content can be added, removed, or replaced). Then, the verified and corrected labeled data set is sent to the preset manual processing terminal. The preset manual processing terminal determines the received at least three sets of verified and corrected labeled data sets as the corrected dataset and sends the corrected dataset to the executing entity. (As an example, the label text in the labeled data can be {"low credit limit"}, and the keyword information can be {"cash flow"}. The label text in the labeled data after verification and correction by the preset manual processing terminal can be {"low credit limit"}, and the keyword information can be {"convenient repayment"}. The preset manual processing terminal can be a pre-set computing device that receives the aforementioned goods flow dialogue text information set and the aforementioned labeled dataset, and randomly sends the aforementioned goods flow dialogue text information set and the aforementioned labeled dataset to the preset manual processing terminal. The preset manual processing terminal can be a pre-set computing device for manual verification and correction.)

[0023] Step 104: Based on the calibration dataset and the item flow dialogue text information set, generate the training dataset and the test dataset.

[0024] In some embodiments, the executing entity may generate a training dataset and a test dataset based on the calibration dataset and the item flow dialogue text information set.

[0025] In some optional implementations of certain embodiments, the aforementioned executing entity may generate a training dataset and a test dataset based on the aforementioned calibration dataset and the aforementioned item flow dialogue text information set through the following steps: The first step is to obtain preset instruction words. These preset instruction words can be a pre-defined natural language prompt used to invoke a large model to analyze the dialogue text (for example, the preset instruction words could be "Analyze the call between the customer and customer service representatives, output a summary, the customer's reasons for rejection, key marketing points from customer service, and the original text, and determine whether the result is successful."). In practice, the executing entity can obtain the preset instruction words from a storage device.

[0026] The second step is to perform the following operations on each correction data in the above correction dataset: The first sub-step involves generating a target summary key-value pair based on the dialogue summary included in the aforementioned correction data. In practice, the executing entity can use the dialogue summary included in the aforementioned correction data as the value and "Summary:" as the key to obtain a fixed key-value pair format of "Summary: Dialogue Summary" as the target summary key-value pair.

[0027] The second sub-step involves generating target tag key-value pairs based on the tag text included in the aforementioned correction data. In practice, the executing entity can use the tag text included in the aforementioned correction data as the value and "tag:" as the key to obtain a fixed key-value pair format of "tag:tag text" as the target tag key-value pair.

[0028] The third sub-step involves generating target keyword information key-value pairs based on the keyword information included in the aforementioned correction data. In practice, the executing entity can use the keyword information included in the aforementioned correction data as the value and "keyword:" as the key to obtain a fixed key-value pair format of "keyword: keyword information" as the target keyword information key-value pair.

[0029] The fourth sub-step involves generating target result information key-value pairs based on the flow result information included in the aforementioned correction data. In practice, the executing entity can use the flow result information included in the aforementioned correction data as the value and "Result:" as the key to obtain a fixed key-value pair format of "Result: Flow Result Information" as the target result information key-value pair.

[0030] The fifth sub-step generates target output information based on the aforementioned target summary key-value pairs, target tag key-value pairs, target keyword information key-value pairs, and target result information key-value pairs. The target output information can be the standard result obtained after extracting information from the dialogue text information of the item flow (for example, the target output information could be: Summary: "Customer service representatives marketed a large credit card installment plan to a customer. The customer stated that the installment fees were high and that they did not need the money. The customer service representative emphasized that applying for the plan would enjoy a limited-time interest rate discount and alleviate financial pressure. Ultimately, the customer agreed to apply." Tags: "High fees", "Not needed". Keywords: "Preferential interest rate, very cost-effective". Flow result: "Yes")). In practice, the executing entity can concatenate the aforementioned target summary key-value pairs, target tag key-value pairs, target keyword information key-value pairs, and target result information key-value pairs to obtain a complete text content as the target output information.

[0031] The sixth sub-step involves determining the item flow dialogue text information corresponding to the above-mentioned correction data from the above-mentioned item flow dialogue text information set as the input information.

[0032] The seventh sub-step involves determining the aforementioned preset instruction words, the aforementioned target output information, and the aforementioned input information as a ternary information group.

[0033] The third step is to define the obtained ternary information groups as ternary information group sets.

[0034] The fourth step involves generating training and testing datasets based on the aforementioned ternary information set. In practice, the executing entity can randomly divide the ternary information set into two parts according to a preset ratio, and then designate these two parts as the training and testing datasets, respectively. This preset ratio can be information representing the division ratio between the training and testing datasets. For example, the preset ratio could be 3:1.

[0035] Step 105: Based on the training dataset and the test dataset, train the preset base model to obtain the trained dialogue text information extraction model.

[0036] In some embodiments, the execution entity may train a preset base model based on the training dataset and the test dataset to obtain a trained dialogue text information extraction model.

[0037] In some optional implementations of certain embodiments, the aforementioned execution entity may train a preset base model based on the aforementioned training dataset and the aforementioned test dataset through the following steps to obtain a trained dialogue text information extraction model: The first step involves reconstructing the training dataset to obtain summary training data sets, label training data sets, keyword training data sets, and flow result training data sets. In practice, for each training data set, the execution entity first retains only the target summary key-value pairs from the target output information. Then, the preset instruction words are replaced with preset summary extraction instruction words. Next, the training data after replacement is determined as the summary training data. Then, the obtained summary training data is divided into summary training data sets according to preset rules, resulting in summary training sets. Then, the execution entity can reconstruct the training dataset to obtain the label training data sets, keyword training data sets, and flow result training sets using the same method as obtaining the summary training data sets. The preset summary extraction instruction words can be a preset natural language prompt used to call the large model to analyze the dialogue text. For example, "Analyze the call between customers and customer service representatives and output a summary." The preset rules can group four pieces of information together.

[0038] The second step is to obtain the preset parameters. These preset parameters include the matrix rank and the learning rate. The matrix rank is a numerical value representing the rank of the subsequently constructed parameter tuning matrix. The learning rate is a numerical value representing the step size for updating the model parameters during training. This step size can be the magnitude by which the model adjusts its parameters in each training epoch.

[0039] The third step is to generate a parameter tuning matrix based on the aforementioned preset parameters. This parameter tuning matrix can be the product of the second parameter matrix and the first parameter matrix. The first parameter matrix can be a matrix with rows equal to the matrix rank and columns equal to the hidden layer dimension of the model. For example, the hidden layer dimension of Qwen2.5-7B is 4096. All parameters in the first parameter matrix are 0. The second parameter matrix can also be a matrix with rows equal to the hidden layer dimension and columns equal to the matrix rank. All parameters in the second parameter matrix can be randomly generated values ​​conforming to a normal distribution. In practice, firstly, the execution entity can construct the first parameter matrix and the second parameter matrix. Then, the product of the second parameter matrix and the first parameter matrix is ​​determined as the parameter tuning matrix.

[0040] The fourth step involves integrating the aforementioned parameter tuning matrix with the pre-defined base model, and then designating the integrated pre-defined base model as the base training sub-model. This pre-defined base model can be a general-purpose large model (such as the Qwen2.5-Omni-7B large model). In practice, firstly, the execution entity can freeze all parameters of the pre-defined base model to ensure that none of its parameters are modified during training. Next, the parameter tuning matrix is ​​attached to the attention layer of the pre-defined base model, forming a combination of "base model + LoRA adapter". Finally, the attached pre-defined base model is designated as the base training sub-model.

[0041] Step 5: Based on the aforementioned summary training dataset, train the aforementioned base training sub-model to obtain the summary extraction model. In practice, firstly, based on the aforementioned summary training dataset, perform the following training steps on the aforementioned base training sub-model: First, input at least one summary training dataset from the summary training dataset into the base training sub-model to obtain the output information corresponding to each summary training dataset in the at least one summary training dataset. Next, compare the output information corresponding to each summary training dataset in the at least one summary training dataset with the target output information in the corresponding summary training dataset. Then, determine whether the base training sub-model has reached the preset optimization objective based on the comparison result. Then, in response to determining that the base training sub-model has reached the aforementioned optimization objective, use the base training sub-model as the trained summary extraction model. In response to determining that the base training sub-model has not reached the aforementioned optimization objective, adjust the parameters in the parameter tuning matrix, and use unused summary training datasets, and use the adjusted base training sub-model as the base training sub-model to perform the aforementioned training steps again. The optimization objective can be that the cosine similarity between the output information and the target output information in the corresponding summary training data set is greater than a preset optimization value. This optimization value can be a pre-set value used as a condition for obtaining the minimum cosine similarity.

[0042] Step 6: Based on the aforementioned label training set, train the aforementioned base training sub-model to obtain the label extraction model. It should be noted that the method used to train the aforementioned base training sub-model to obtain the label extraction model is the same as the method used to train the aforementioned base training sub-model to obtain the summary extraction model based on the aforementioned summary training data set.

[0043] Step 7: Based on the aforementioned keyword training set, train the aforementioned base training sub-model to obtain the keyword extraction model. It should be noted that the method used to train the aforementioned base training sub-model to obtain the keyword extraction model is the same as the method used to train the aforementioned base training sub-model to obtain the summary extraction model based on the aforementioned summary training data set.

[0044] Step 8: Based on the training set of the aforementioned flow results, train the aforementioned base training sub-model to obtain the flow result extraction model. It should be noted that the method used to train the aforementioned base training sub-model based on the aforementioned flow result training set to obtain the flow result extraction model is the same as the method used to train the aforementioned base training sub-model based on the aforementioned summary training data set to obtain the summary extraction model.

[0045] Step nine involves integrating the above-mentioned summary extraction model, tag extraction model, keyword extraction model, and flow result extraction model with the above-mentioned preset base model to obtain the base training model. In practice, firstly, the execution entity can determine the parameter tuning matrix of the above-mentioned summary extraction model as the first sub-task matrix, the parameter tuning matrix of the above-mentioned tag extraction model as the second sub-task matrix, the parameter tuning matrix of the above-mentioned keyword extraction model as the third sub-task matrix, and the parameter tuning matrix of the above-mentioned flow result extraction model as the fourth sub-task matrix. Then, the parameters of the above-mentioned summary matrix, tag matrix, keyword matrix, flow result matrix, and preset base model are linearly added together. Finally, the matrix obtained after linear addition is used as the base training model. It should be noted that when the above-mentioned base training model is trained, it will call the corresponding sub-task matrix according to the input prompt. As an example, if the input prompt is "analyze content, output summary", the base training model will only call the first sub-task matrix when performing analysis.

[0046] Step 10: Based on the aforementioned training dataset, generate a training data set. In practice, the executing entity can divide the training data in the aforementioned training dataset into at least one training data set according to a preset rule. Then, the at least one training data set is determined as the training data set. The preset rule can be to randomly divide the training data into groups according to a preset number. The preset number can be 4.

[0047] Step 11: Based on the above training data set and the above test dataset, train the above-mentioned base training model to obtain the dialogue text information extraction model.

[0048] In some optional implementations of certain embodiments, the aforementioned execution entity may train the aforementioned base training model based on the aforementioned training data set and the aforementioned test dataset through the following steps to obtain the dialogue text information extraction model: The first step is to train a large model on the base based on the training dataset, performing the following training steps: The first sub-step involves randomly selecting a training data set from the training data set set, identifying a training data set as the target training data set, and deleting the aforementioned training data set from the training data set set.

[0049] The second sub-step involves segmenting and filling each target training data point in the target training data set to obtain a sequence of word indexes. In practice, firstly, the execution entity can load the corresponding word segmenter from the aforementioned base training model. Then, each target training data point in the target training data set is input into the word segmenter to obtain a sequence of word indexes. Each word index sequence corresponds to one target training data point from the target training data set. The word index sequence can be a sequence composed of word indices arranged according to the order of their corresponding words in the target training data. The word index can be the numerical value representing the word for each semantic unit after all sentences in the target training data have been segmented into the smallest semantic units. The aforementioned base training model contains a vocabulary. The vocabulary includes each smallest semantic unit and the word index corresponding to each smallest semantic unit.

[0050] The third sub-step involves inputting the various lexical index sequences into the base training model to obtain various probability distribution sequences. In practice, the aforementioned execution entity can input the various lexical index sequences into the aforementioned base training model. For each probability distribution sequence, the model predicts the probability of the lexical index corresponding to the next smallest semantic unit at each position in the sequence, and outputs the probability sequence as the probability distribution sequence.

[0051] The fourth sub-step involves generating an average loss value based on each probability distribution sequence and each word index sequence. In practice, firstly, for each probability distribution sequence, the execution entity can generate the cross-entropy loss between the probability distribution sequence and the corresponding word index sequence. Then, the execution entity can use the average of the obtained cross-entropy losses as the average loss value.

[0052] The fifth sub-step involves updating the parameters of the large-scale training model on the pedestal based on the average loss value and the learning rate in the preset parameters. In practice, the aforementioned execution entity can utilize the backpropagation algorithm to obtain the difference between the predicted and the true values ​​based on the average loss value. This allows the generation of the gradient value for each trainable parameter in the model. Next, the execution entity can use a preset optimizer to update the parameters of the large-scale training model on the pedestal according to a preset update formula. This preset optimizer can be stochastic gradient descent (SGD). The preset update formula can be... Represented as: Among them, the above This can be the updated parameter value, as mentioned above. This can be the parameter value before the update, as mentioned above. It can be the learning rate, as mentioned above. It can be a gradient value.

[0053] The sixth sub-step involves determining the average loss value as the historical loss value and storing the historical loss value in a preset historical loss value queue, thereby updating the preset historical loss value queue. The preset historical loss value queue is a pre-set queue of fixed length used to store historical loss values.

[0054] The seventh sub-step, in response to determining that a historical loss value enqueue operation exists in the preset historical loss value queue and that the number of historical loss values ​​in the preset historical loss value queue after the enqueue operation is equal to the length of the preset historical loss value queue, generates a historical loss representation value based on each historical loss value in the preset historical loss value queue. In practice, the aforementioned executing entity can determine the variance of each of the aforementioned historical loss values ​​as the historical loss representation value.

[0055] The eighth sub-step, in response to determining that the historical loss representation value meets a preset completion condition or the training data set is empty, updates and verifies the base-trained large model with updated parameters based on the test dataset to obtain the dialogue text information extraction model. The preset completion condition can be that the historical loss representation value is less than a preset loss value. The preset loss value can be pre-set, limiting the maximum value that the historical loss representation value can reach.

[0056] It should be noted that, in response to the determination that the historical loss representation value does not meet the preset completion condition and the updated training data set is not empty, the training steps are performed again on the base training large model with updated parameters based on the training data set.

[0057] In some optional implementations of certain embodiments, the aforementioned execution entity can obtain a dialogue text information extraction model by performing update and verification processing on the base-trained large model with updated parameters based on a test dataset through the following steps: The first step, in response to determining that the historical loss representation value meets the preset completion condition or the training data set is empty, involves word segmentation and padding of each test data point in the test dataset to obtain each test word index sequence. The resulting base training model, after updating the parameters with these test word index sequences, then obtains each predicted probability distribution sequence. It should be noted that the method used for word segmentation and padding of each test data point in the test dataset to obtain each test word index sequence, and then inputting these test word index sequences into the base training model to obtain each predicted probability distribution sequence, is the same as the method used for word segmentation and padding of each target training data point in the target training dataset to obtain each word index sequence, and then inputting these word index sequences into the base training model to obtain each probability distribution sequence.

[0058] The second step involves generating test loss values ​​based on each predicted probability distribution sequence and each test term index sequence. In practice, firstly, the execution entity can generate the cross-entropy loss between each predicted probability distribution sequence and its corresponding test term index sequence. Then, the execution entity can use the average of the obtained cross-entropy losses as the average loss value.

[0059] The third step involves, in response to the determination that the test loss value is greater than the preset loss value, generating a new training data set based on the training dataset and training a large model on the updated base model using this training data set, then repeating the above training steps. The preset loss value can be a pre-defined value representing the maximum test loss value. In practice, the executing entity can randomly divide the training dataset into two parts according to a preset ratio, designating these two parts as the training dataset and the test dataset, respectively.

[0060] The fourth step is to determine the test loss value as less than or equal to the preset loss value and then use the updated base training model as the dialogue text information extraction model.

[0061] In some embodiments, a method for training a pre-defined base model based on training and test datasets to obtain a trained dialogue text information extraction model involves first training the base training sub-models four times using a reconstructed training dataset (summary training dataset, tag training dataset, keyword training dataset, and flow result training dataset), resulting in more targeted sub-models for each part of the content. Then, using the training dataset, the pre-defined base model, obtained by integrating the four sub-models with the pre-defined base training model, is trained. This avoids the parameters of each sub-task matrix from interfering with each other when extracting information from complete communication text. A pre-defined historical loss value queue is used to obtain the variance of the loss values ​​over a certain period to determine whether the model has reached a stable state and avoid overfitting. This results in a dialogue text information extraction model with balanced performance when calling each sub-task matrix (i.e., summary extraction, tag extraction, keyword extraction, and flow result extraction).

[0062] Step 106: Obtain the dialogue text information to be processed.

[0063] In some embodiments, the executing entity may obtain the dialogue text information to be processed. In practice, the executing entity may obtain the dialogue text information to be processed from a storage device. The dialogue text information to be processed may be text information representing the content of a dialogue that needs to be processed.

[0064] Step 107: Input the dialogue text information to be processed into the dialogue text information extraction model to obtain the extracted information.

[0065] In some embodiments, the executing entity may input the dialogue text information to be processed into the dialogue text information extraction model to obtain extracted information. The extracted information may be the information output by the dialogue text information extraction model, including dialogue summary, tag text, keyword information, and flow result information.

[0066] Step 108: The extracted information is parsed to obtain output parsable information, and the dialogue text information to be processed and the output parsable information are sent to the preset downstream business system.

[0067] In some embodiments, the executing entity may parse the extracted information to obtain outputtable parsing information and send the dialogue text information to be processed and the outputtable parsing information to a preset downstream business system. The preset downstream business system may be a pre-configured computing device for receiving outputtable parsing information. The parsable information may represent dialogue summaries, tag texts, keyword information, and flow result information in the dialogue text information to be processed.

[0068] In some optional implementations of certain embodiments, the aforementioned execution entity may perform parsing processing on the extracted information through the following steps to obtain outputtable parsing information and send the dialogue text information to be processed and the outputtable parsing information to a preset downstream business system: The first step is to input the extracted information into a preset information parser to obtain a text information dictionary. This preset information parser can be a pre-defined regular expression parser. The text information dictionary can be a structured table output by the preset information parser. For example, if the extracted information text is: "Summary: Customer service recommended a product, customer declined. Tag: Not needed at the moment." the resulting text information dictionary would have the following keys: "Summary": value "Customer service recommended a product, customer declined" and "Tag": value "Not needed at the moment".

[0069] The second step is to obtain a preset historical text information dictionary sequence. In practice, the aforementioned execution entity can obtain the preset historical text information dictionary sequence from a storage device. This preset historical text information dictionary sequence can be a sequence composed of previously obtained text information dictionaries arranged in chronological order of their generation. Each historical text information dictionary in the preset historical text information dictionary sequence corresponds to the dialogue text information to be processed.

[0070] The third step is to add the aforementioned text information dictionary to the aforementioned preset historical text information dictionary sequence to update the preset historical text information dictionary sequence. In practice, the executing entity can add the aforementioned text information dictionary to the last position of the aforementioned preset historical text information dictionary sequence. The aforementioned preset historical text information dictionary sequence has a fixed length. Each historical text information dictionary in the aforementioned preset historical text information dictionary sequence exists for a fixed period of time.

[0071] The fourth step involves generating dictionary judgment information based on the aforementioned text information dictionary. In practice, the executing entity can compare the contents of the text information dictionary according to a preset format and output dictionary judgment information based on the comparison results. This dictionary judgment information can be a set of multiple sub-judgment information. These sub-judgment information can be information that judges whether the output content of each part of the text information dictionary—dialogue summary, tag text, keyword information, and flow result information—is consistent with a preset format. For example, the judgment information could be "Summary: Yes", "Tag Text: Yes", "Keywords: No", or "Flow Result Information: Yes". The preset format can be information representing pre-defined specifications for each part of the content. For example, Summary: must be a paragraph (non-empty string), Tags: must be a list, Keywords: must be a dictionary (containing key points and the original text), Flow Result Information: can only be "Yes" or "No".

[0072] Fifth, in response to determining that the above dictionary judgment information meets the preset conditions, the above text information dictionary is determined as output parsable information. The preset conditions can be that all entries in the dictionary judgment information are "yes".

[0073] Step 6: In response to the determination that the dictionary judgment information does not meet the preset conditions, a prompt message is generated, and the dialogue text information to be processed and the text information dictionary are sent to the preset manual processing terminal. The preset manual processing terminal then reviews and repairs the text information dictionary to obtain a repaired text information dictionary, and determines that the repaired text information dictionary is output and can be parsed. In practice, firstly, after receiving the dialogue text information to be processed and the text information dictionary, the preset manual processing terminal manually modifies and repairs the text information dictionary based on the dialogue text information to be processed (manually adding, deleting, or replacing missing or erroneous parts) to obtain a repaired text information dictionary. Then, after the manual modification and repair of the text information dictionary is completed, the preset manual processing terminal sends the repaired text information dictionary to the executing entity.

[0074] Step 7: Send the above-mentioned output parsable information to the preset downstream business system.

[0075] In addressing the technical problems mentioned above by adopting technical solutions, the following technical issues often arise in the application scenario: extracting information from high-intent customers during telemarketing. Because telemarketing generates a large amount of dialogue information within a certain timeframe, and communication with high-intent customers results in lengthy conversations, processing large volumes of complex and numerous long dialogues can easily lead to errors in the parsed text dictionary. To ensure the accuracy of the parsed information, it is common practice to send all information obtained from a single extraction (including both correct and incorrect content) to a pre-set manual processing end for correction. This also results in a large amount of correct information that does not require correction being sent to the pre-set manual processing end, wasting bandwidth resources. Considering the following requirements for this application scenario: parsing large volumes of long dialogues and reducing bandwidth waste, we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity may further perform the following steps to parse the extracted information, obtain outputtable parsed information, and send the unprocessed dialogue text information and the outputtable parsed information to a preset downstream business system: The first step, in response to determining that the number of text characters in the above-mentioned dialogue text information to be processed has reached a preset number, is to perform the following steps on the above-mentioned extracted information: The first sub-step involves inputting the extracted information into a preset extracted information parser to obtain a text information dictionary.

[0076] The second sub-step is to obtain a preset dictionary sequence of historical text information.

[0077] The third sub-step involves adding the aforementioned text information dictionary to the aforementioned preset historical text information dictionary sequence to update the aforementioned preset historical text information dictionary sequence.

[0078] The fourth sub-step involves generating dictionary determination information based on the aforementioned text information dictionary.

[0079] The fifth sub-step involves determining that the dictionary judgment information meets the preset conditions, identifying the text information dictionary as output parsable information, and sending the output parsable information to the preset downstream business system.

[0080] The sixth sub-step involves, in response to the determination that the dictionary judgment information does not meet the preset conditions, re-extracting the dialogue text information to be processed based on the dictionary judgment information to obtain re-extracted information. In practice, firstly, the executing entity may, in response to determining that the information corresponding to "no" in the dictionary judgment information (i.e., at least one of summary, tag, keyword, or flow result information), re-input the dialogue text information to be processed into the dialogue text information extraction model, and call the sub-task matrix corresponding to the above information to extract information from the dialogue text information to be processed, thus obtaining re-extracted information. As an example, if the judgment information corresponding to the keyword in the dictionary judgment information is "no", after re-inputting the dialogue text information to be processed into the dialogue text information extraction model, only the third sub-task matrix is ​​called to extract information, resulting in extracted information containing only keyword information as re-extracted information.

[0081] The seventh sub-step involves inputting the re-extracted information into a preset extraction information parser to obtain a re-extracted text information dictionary. This re-extracted text information dictionary can be a structured table obtained after parsing the re-extracted information.

[0082] The eighth sub-step generates duplicate dictionary determination information based on the aforementioned re-extracted text information dictionary. In practice, the executing entity can compare the contents of the aforementioned text information dictionary according to a preset format and output dictionary determination information based on the comparison results.

[0083] The ninth sub-step, in response to determining that the aforementioned re-extracted text information meets preset conditions, generates outputtable parsing information based on the aforementioned re-extracted text information dictionary and the aforementioned text information dictionary, and sends the outputtable parsing information to a preset downstream business system. In practice, the aforementioned executing entity can replace the corresponding content in the aforementioned text information dictionary with the obtained re-extracted text information dictionary, obtaining the replaced text information dictionary as the outputtable parsing information. As an example, the aforementioned re-extracted text information dictionary could be "Keywords: Credit limit released immediately after installment for emergency use," and the aforementioned text information dictionary could be "Summary: Customer service recommends installment products to customers via telephone. Customers say they don't need them at the moment. Customer service emphasizes the preferential handling fees and flexible use of funds during the promotion period. Ultimately, the customer agrees to process the bill installment. Tags: Low credit limit, worried about forgetting to repay. Keyword: null. Flow result: Yes." The corresponding content of the re-extracted text information dictionary in the aforementioned text information dictionary is "Keyword: null." The text information dictionary, after replacing "keyword: null" in the original text information dictionary with the extracted text information dictionary, can be: "Summary: A customer recommends an installment product to another customer. The customer states that they do not currently need it. The customer service representative emphasizes the preferential handling fees and flexible use of funds during the promotional period. The customer agrees to proceed. Tags: Low credit limit. Keywords: Convenient for emergencies after installment payment. Flow result: Yes." Here, "null" can be a character representing an empty content.

[0084] The tenth sub-step involves, in response to the determination that the aforementioned re-dictionary judgment information does not meet the preset conditions, sending the aforementioned dialogue text information to be processed and the aforementioned re-extracted text information dictionary to a preset manual processing terminal. The preset manual processing terminal then reviews and repairs the re-parsed information to obtain the processed re-parsed information. In practice, after receiving the aforementioned dialogue text information to be processed and the aforementioned re-extracted text information dictionary, the preset manual processing terminal manually modifies and repairs the re-extracted text information dictionary based on the aforementioned dialogue text information (manually adding, deleting, or replacing any missing or erroneous parts). After the manual modification and repair of the re-extracted text information dictionary is completed, the corresponding content in the aforementioned text information dictionary is replaced with the repaired re-extracted text information dictionary, resulting in a replaced text information dictionary as the processed re-parsed information. The preset manual processing terminal then sends the processed re-parsed information to the aforementioned execution entity.

[0085] The eleventh sub-step involves determining the re-parsed information after the above processing as output parsing information and sending the parsing information to a preset downstream business system.

[0086] The above-mentioned technical solution and related content, as an inventive point of this disclosure, solve the technical problem of "waste of bandwidth resources." Factors leading to bandwidth resource waste often include: During telemarketing, a large amount of dialogue information is generated within a certain timeframe; simultaneously, communication with high-intent customers generates lengthy dialogue content. When processing a large volume of long dialogue texts, the complexity and quantity of the dialogue content easily lead to errors in the parsed text dictionary. To ensure the accuracy of the parsed information, all information obtained after a single information extraction process (including correct and incorrect content) is typically sent to a preset manual processing terminal for correction. This also sends a large amount of correct information that does not require correction to the preset manual processing terminal, resulting in bandwidth resource waste. Solving these factors can reduce bandwidth resource waste. To achieve this effect, firstly, in response to determining that the number of text characters in the above-mentioned dialogue text information to be processed reaches a preset number, the following steps are performed on the extracted information: First, the extracted information is input into a preset extracted information parser to obtain a text information dictionary. Thus, a text information dictionary parsed from the extracted information can be obtained. Then, a preset historical text information dictionary sequence is acquired. Next, the aforementioned text information dictionary is added to the aforementioned preset historical text information dictionary sequence to update the preset historical text information dictionary sequence. This yields an updated preset historical text information dictionary sequence. Then, based on the aforementioned text information dictionary, dictionary determination information is generated. This yields dictionary determination information indicating whether the generated text information dictionary is compliant. Next, in response to determining that the aforementioned dictionary determination information meets preset conditions, the aforementioned text information dictionary is determined to be output parsable information, and the output parsable information is sent to a preset downstream business system. This yields output parsable information that can be extracted and output to the preset downstream business system. Then, in response to determining that the aforementioned dictionary determination information does not meet preset conditions, the aforementioned dialogue text information to be processed is re-extracted based on the aforementioned dictionary determination information to obtain re-extracted information. This allows for re-extraction of the content of the dialogue text information to be processed when preset conditions are not met, resulting in re-extracted information. Next, the re-extracted information is input to a preset extraction information parser to obtain a re-extracted text information dictionary. This yields a text information dictionary parsed from the re-extracted information. Then, based on the aforementioned re-extracted text information dictionary, re-dictionary determination information is generated. This provides re-dictionary determination information indicating whether the generated re-extracted text information dictionary is compliant. Next, in response to determining that the aforementioned re-dictionary determination information meets preset conditions, outputtable parsing information is generated based on the aforementioned re-extracted text information dictionary and the aforementioned text information dictionary, and this outputtable parsing information is sent to a preset downstream business system. This provides outputtable parsing information that can be extracted and output to the preset downstream business system.Next, in response to the determination that the aforementioned re-dictionary judgment information does not meet the preset conditions, the aforementioned dialogue text information to be processed and the aforementioned re-extracted text information dictionary are sent to a preset manual processing terminal for review and repair of the re-parsed information, resulting in processed re-parsed information. Thus, only the erroneous portion of the information can be transmitted to the preset manual processing terminal, obtaining re-parsed information after manual review of the erroneous portion. Finally, the processed re-parsed information is determined to be output parsed information and sent to a preset downstream business system. Thus, output parsed information that can be output to the preset downstream business system after review and processing can be obtained. Because when there is a lot of content in the dialogue text information to be processed, the text information dictionary that does not meet the preset conditions is re-extracted in combination with the dialogue text information to be processed. For the part that still does not meet the preset conditions after re-extraction, it is sent to the preset manual processing terminal for review and repair. This avoids sending all the generated text information dictionary content to the preset manual processing terminal when the dictionary judgment information represents the text information dictionary content incorrectly. When extracting a large amount of long dialogue text, the amount of data transmitted to the preset manual processing terminal for repair is reduced, and the waste of bandwidth resources is reduced.

[0087] Optionally, after adding the aforementioned text information dictionary to the aforementioned preset historical text information dictionary sequence, the method may further include: The first step is to obtain a preset reward formula and a preset historical text information dictionary sequence. The preset reward formula can be a pre-set formula used to generate the global reward value for subsequent text information. The global reward value can be a numerical value representing the quality of a single extraction result.

[0088] The second step involves performing the following steps on the aforementioned preset historical text information dictionary sequence, based on the preset reward formula: The first sub-step is to determine the last historical text information dictionary in the above-mentioned preset historical text information dictionary sequence as the reference text information dictionary.

[0089] The second sub-step generates a global reward value for the text information based on the aforementioned preset reward formula and the aforementioned reference text information dictionary. In practice, firstly, the executing entity can retrieve the dialogue text information to be processed corresponding to the reference text information dictionary from the storage device. Then, the executing entity can input the aforementioned reference text information dictionary and the corresponding dialogue text information to be processed into a pre-set natural language inference model to obtain a semantic reward value. Next, a completeness check is performed on the content in the aforementioned reference text information dictionary (for example, calculating the string length of the "summary" field. If the number of words is within the preset ideal range, a high score is obtained; if it is too short or too long, points are deducted proportionally. Calculating the number of elements in the "tag" list. A reward is obtained if the number is within a reasonable range (e.g., 1-3); points are deducted if there are too many or the list is empty). This yields a complete reward value. Next, a procedural check of deterministic rules is performed on the content in the aforementioned reference text information dictionary (e.g., whether key markers such as "Summary:" and "Tags" appear in the text in sequence. If missing, the reward for that item is 0 points or a negative score. In the parsed dictionary, whether the value corresponding to each key is of a valid type (e.g., whether "Tags" is a list, "Keywords" is a dictionary). If it is a default empty value, points are deducted). This yields the rule reward value. Finally, the aforementioned rule reward value, the aforementioned complete reward value, and the aforementioned semantic reward value are substituted into a preset reward formula to obtain the global text information reward value, which serves as the global text information reward value. The aforementioned rule reward value can be a numerical value representing the compliance of the content generated by each subtask in the reference text information dictionary, and can include the rule reward value for the first subtask, the rule reward value for the second subtask, and the rule reward value for the third subtask. The aforementioned rule reward values ​​for the first, second, and third subtasks can represent the rule reward values ​​for different subtask parts in the reference text information dictionary. The aforementioned complete reward value can be a numerical value representing the compliance of the length of each part in the reference text information dictionary, and can include the complete reward value of the first subtask, the complete reward value of the second subtask, and the complete reward value of the third subtask. The aforementioned complete reward values ​​of the first, second, and third subtasks can represent the complete reward values ​​for different subtask parts in the reference text information dictionary. The aforementioned semantic reward value can be a numerical value representing the authenticity of the content in the reference text information dictionary, and can include the semantic reward values ​​of the first, second, and third subtasks. The aforementioned semantic reward values ​​of the first, second, and third subtasks can represent the semantic reward values ​​for different subtask parts in the reference text information dictionary. The aforementioned global text information reward value can be a numerical value representing the comprehensive situation of the execution results of each subtask of the model.The pre-set natural language inference model mentioned above can be a pre-trained language model that takes a reference text information dictionary and the corresponding dialogue text information to be processed as input and a semantic reward value as output (e.g., RoBERTa-Large-MNLI, DeBERTa). The pre-set reward formula mentioned above can be expressed as: .in, It can be used to assign a global information reward value to text information. It can represent the first The value is obtained by weighting the rule reward value, complete reward value, and semantic reward value of each subtask. It can represent the first Sub-tasks. This can be set to the number of subtasks. , , It can be an adjustable hyperparameter, and satisfies... . Can the first The rule reward value for each sub-task (i.e., the first sub-task) (Sub-task rule reward value). It can be the first The complete reward value of each subtask (i.e., the first) (Complete reward value for sub-tasks). It can be the first The semantic reward value of the first subtask (i.e., the first subtask) (Subtask semantic reward value). As an example, the reference text information dictionary and the corresponding dialogue text information to be processed can be first determined as a target input information group. Next, at least one sample target input information group and the sample semantic reward value corresponding to each sample target input information group in the at least one sample target input information group are obtained. Then, each sample target input information group in the at least one sample target input information group is used as input, and the sample semantic reward value corresponding to each sample target input information group in the at least one sample target input information group is used as the expected output to train a natural language inference model.

[0090] The third sub-step involves optimizing the parameter strategy of the aforementioned dialogue text extraction model based on the global reward value information of the text information, resulting in a fine-tuned dialogue text extraction model. In practice, the executing entity can invoke the PPO (Proximal Policy Optimization) algorithm to optimize the parameter strategy of the dialogue text extraction model using the global reward value information of the text information, thereby obtaining a fine-tuned dialogue text extraction model.

[0091] The third step is to define the fine-tuned dialogue text extraction model as the dialogue text extraction model.

[0092] In addressing the aforementioned technical challenges by employing technical solutions, the application scenario of real-time monitoring of financial telemarketing often presents further technical issues: financial telemarketing is strictly regulated, and any use of inappropriate scripts could result in penalties, license revocation, or even suspension of operations. Furthermore, unresolved issues can drastically increase the risk of customer churn or escalating complaints, leading to a poor user experience. Considering the specific requirements of this application scenario—short processing time and timely feedback on problematic calls—we have decided to adopt the following solution: In some optional implementations of certain embodiments, after parsing the extracted information to obtain outputtable parsable information and sending the dialogue text information to be processed and the outputtable parsable information to a preset downstream business system, the method may further include: The first step is to obtain the historical dialogue quality score sequence and the corresponding dialogue speech data for the dialogue text information to be processed. The historical dialogue quality score sequence can be obtained by arranging the various dialogue quality scores in chronological order of their generation. The dialogue quality scores can be numerical values ​​representing the quality of the dialogue process. The dialogue speech data can be the original recording file or audio stream of the dialogue text information to be processed. In practice, the executing entity can obtain the historical dialogue quality score sequence and the corresponding dialogue speech data for the dialogue text information to be processed from a storage device.

[0093] The second step is to perform dialogue sentiment analysis on the above-mentioned dialogue voice data to obtain sentiment data.

[0094] The third step is to generate compliance data based on the aforementioned dialogue text information to be processed. In practice, the executing entity can input the dialogue text information to be processed into a pre-trained compliance analysis model to obtain compliance data. The pre-trained compliance analysis model can be a neural network model (e.g., Legal-BERT, RoBERTa) that takes the dialogue text information to be processed as input and the compliance data as output. The compliance data can be a numerical value representing the degree to which the content of the dialogue text information to be processed conforms to preset requirements. The preset requirements can include both non-compliant and necessary statements in the dialogue. As an example, at least one sample of dialogue text information to be processed and sample compliance data corresponding to each sample of dialogue text information to be processed can be obtained first. Then, each sample of dialogue text information to be processed in the at least one sample of dialogue text information to be processed is used as input, and the sample compliance data corresponding to each sample of dialogue text information to be processed in the at least one sample of dialogue text information to be processed is used as the expected output to train the compliance analysis model.

[0095] The fourth step involves generating a dialogue quality score based on the aforementioned outputtable parsing information, sentiment data, and compliance data. In practice, the executing entity can input the outputtable parsing information, sentiment data, and compliance data into a pre-trained dialogue quality scoring model to obtain the dialogue quality score. The dialogue quality score can be a numerical value representing the overall quality of the dialogue. The dialogue quality scoring model can be a pre-trained linear weighted sum model that takes the outputtable parsing information, sentiment data, and compliance data as input and the dialogue quality score as output; for example, a fixed-weight linear fusion model or a learned-weight linear fusion model. As an example, the outputtable parsing information, sentiment data, and compliance data can first be defined as an input information group. Next, at least one sample input information group and the sample dialogue quality score corresponding to each sample input information group in the at least one sample input information group are obtained. Then, each sample input information group in the at least one sample input information group is used as input, and the sample dialogue quality score corresponding to each sample input information group in the at least one sample input information group is used as the expected output to train the dialogue quality scoring model.

[0096] The fifth step involves generating a core indicator trend chart based on the aforementioned historical dialogue quality score sequence and the aforementioned dialogue quality scores, and then outputting and displaying the core indicator trend chart. In practice, firstly, the executing entity can use a visualization library to create a line graph showing the score trend of each historical dialogue quality score in the aforementioned historical dialogue quality score sequence and the aforementioned dialogue quality scores, serving as the core indicator trend chart. Then, the executing entity can output the core indicator trend chart to a preset display screen. This preset display screen can be the electronic screen of a pre-defined target display device.

[0097] The sixth step is to add the aforementioned dialogue quality scores to the aforementioned historical dialogue quality score sequence to update the historical dialogue quality score sequence. In practice, the executing entity can add the aforementioned dialogue quality scores to the aforementioned historical dialogue quality score sequence in chronological order of their generation time.

[0098] Step 7: In response to determining that the dialogue score does not meet a preset quality condition, feedback information is generated and sent to a preset receiving end. The preset quality condition may be that the dialogue score is greater than a preset value. The feedback information may indicate that the dialogue text to be processed has a problem that needs to be addressed, such as, "The score is below the threshold; please process it promptly." The preset receiving end may be a pre-set computing device used to display the feedback information.

[0099] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem mentioned in the background: "Financial telemarketing is subject to strict regulation. If illegal scripts are used, it may lead to penalties such as company penalties or license revocation. Furthermore, if dialogue issues are not handled promptly, the risk of customer churn or escalating complaints increases dramatically, resulting in a poor user experience." Solving these factors achieves the effect of short processing time and timely feedback on problematic calls. To achieve this effect, firstly, the historical dialogue quality score sequence and the dialogue voice data corresponding to the above-mentioned dialogue text information to be processed are obtained. Then, dialogue voice data is processed using dialogue sentiment analysis to obtain sentiment data. This provides data representing the other party's emotions during the dialogue, used to generate a dialogue quality score. Next, based on the above-mentioned dialogue text information to be processed, compliance data is generated. This provides compliance data indicating whether the dialogue was conducted as required and whether illegal scripts were used. Then, based on the above-mentioned dialogue text information to be processed, the above-mentioned outputtable parsing information, the above-mentioned sentiment data, and the above-mentioned compliance data, a dialogue quality score is generated. This provides a numerical value representing the quality level of the dialogue process. Then, based on the aforementioned historical dialogue quality score sequence and the aforementioned dialogue quality scores, a core indicator trend chart is generated, and the core indicator trend chart is output and displayed. This provides a core indicator trend chart that visually shows changes in dialogue quality scores. Next, the aforementioned dialogue quality score sequence is added to the aforementioned historical dialogue quality score sequence to update the historical dialogue quality score sequence. This results in a historical dialogue quality score sequence containing the latest dialogue quality scores. Finally, in response to the determination that the aforementioned dialogue score does not meet the preset quality conditions, feedback information is generated and sent to the preset receiving end. This allows for timely generation of feedback information (indicating the presence of inappropriate language or customer agitation, etc.) when the dialogue quality score does not reach the preset value, reminding for prompt handling of the dialogue. Furthermore, by performing sentiment analysis on the dialogue voice data, data representing customer emotions is obtained. By analyzing the dialogue text to be processed, compliance data is obtained, used to determine whether inappropriate language has occurred, allowing for timely handling of dialogues with inappropriate language. Finally, by processing the dialogue text information to be processed, outputtable parsable information, sentiment data, and compliance data, a dialogue quality score representing the call quality is obtained, characterizing the quality of this dialogue. Furthermore, by incorporating trend charts of core metrics, the dialogue quality over a period of time is visualized, allowing for a more intuitive view of changes in dialogue quality. Feedback is also provided for dialogue quality scores that do not meet preset values, facilitating timely identification of dialogue problems and shortening processing time.

[0100] In addressing the technical problems mentioned above, and considering the application scenario—telemarketing requires extensive voice analysis—the following technical issues arise: When analyzing large amounts of audio information from conversations, using a multimodal large-scale model for sentiment analysis can lead to exponentially increasing computational costs due to the large number of parameters, resulting in a heavy computational burden and excessive consumption of computing resources. Given the specific requirements of this application scenario—suitability for high concurrency and high availability in production environments—we have decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned executing entity may perform dialogue emotion analysis processing on the aforementioned dialogue voice data through the following steps to obtain emotion data: The first step is to perform activity separation processing on the aforementioned dialogue speech data to obtain a sequence of speech frame data to be processed. In practice, firstly, the executing entity can use VAD (Voice Activity Detection) to remove silent segments from the dialogue speech data, retaining only the "sounding" segments. Then, the dialogue speech data is divided according to preset standards to obtain a sequence of speech frame data to be processed. This sequence can be composed of individual speech frame data arranged in the order they appear in the dialogue speech data. Each speech frame data segment can be a single speech frame fragment that needs processing. The preset standards can be a frame length of 20ms and a frame shift of 0.01 seconds.

[0101] The second step involves processing each speech frame in the above sequence of speech frame data as follows: The first sub-step involves extracting Mel-frequency cepstral coefficient features from the aforementioned speech frame data to be processed, thereby obtaining a Mel-frequency cepstral coefficient feature vector. In practice, the executing entity can extract Mel-frequency cepstral coefficient features from the aforementioned speech frame data to be processed, thus obtaining a Mel-frequency cepstral coefficient feature vector.

[0102] The second sub-step involves generating speech domain feature data based on the aforementioned speech frame data to be processed. In practice, firstly, the executing entity can use the cepstral method to obtain the fundamental frequency of the speech frame data to be processed. Then, the square root sum of the sampling points in the speech frame data is used as the speech energy value. Next, the average frequency of the speech frame data is used as the spectral centroid. Finally, the fundamental frequency, the speech energy value, and the spectral centroid are combined into a vector as the speech domain feature data.

[0103] The third sub-step involves generating a speech feature vector based on the aforementioned Mel-frequency cepstral coefficient feature vector and the aforementioned speech domain feature data. In practice, the executing entity can concatenate the aforementioned Mel-frequency cepstral coefficient feature vector and the aforementioned speech domain feature data, using the concatenated vector as the speech feature vector.

[0104] The fourth sub-step involves inputting the aforementioned speech feature vectors into a pre-trained emotion data model to obtain speech frame emotion data. This emotion data model can be a single neural network model (e.g., an FNN model or a DNN model) that takes speech feature vectors as input and outputs speech frame emotion data. The speech frame emotion data can represent the degree of emotion (higher values ​​for better mood) within the processed speech frame data. For example, 100-80 indicates happiness, 79-65 indicates calmness, and 64 and below indicates anger. As an example, at least one sample speech feature vector and the corresponding sample speech frame emotion data for each of these vectors can be obtained first. Then, each of the at least one sample speech feature vectors is used as input, and the corresponding sample speech frame emotion data is used as the desired output to train the emotion data model.

[0105] The third step is to generate a sequence of voice frame emotion data based on the obtained emotion data of each voice frame. In practice, the aforementioned execution entity can arrange the emotion data of each voice frame according to the order of the corresponding voice frame data to be processed in the sequence of voice frame data to be processed, thus obtaining the sequence of voice frame emotion data.

[0106] The fourth step involves generating an emotion curve and emotion index data based on the aforementioned speech frame emotion data sequence. In practice, firstly, the executing entity can extract the emotion intensity value of each speech frame emotion data from the speech frame emotion data sequence (for example, mapping the label "anger" to -1, "calm" to 0, and "happy" to +1). Then, Gaussian filtering is applied to process each emotion intensity value. Next, the executing entity can use a visualization library to create a line graph of the processed emotion intensity values ​​as the emotion curve. Then, the maximum and minimum values ​​among the aforementioned emotion intensity values ​​are determined as the emotion peak. Next, the average value of the aforementioned emotion intensity values ​​is determined as the overall emotion value. Next, the standard deviation of the aforementioned emotion intensity values ​​is determined as the emotion fluctuation value. Finally, the aforementioned emotion peak, overall emotion value, and emotion fluctuation value are determined as the emotion index data.

[0107] The fifth step is to identify the above-mentioned sentiment curve and sentiment index data as sentiment data.

[0108] The above-described technical solution and its related content, as an inventive point of this disclosure, solve the technical problem mentioned in the background art: "When analyzing a large amount of audio information from dialogues, using a multimodal large model for sentiment analysis leads to an exponential increase in computational cost due to the large number of parameters in the multimodal large model, resulting in a heavy computational burden and thus a large consumption of computational resources." Solving these factors can reduce the consumption of computational resources. To achieve this effect, firstly, the dialogue speech data is subjected to activity separation processing to obtain a sequence of speech frame data to be processed. This removes noise from the dialogue speech data, resulting in a sequence of speech frame data to be processed, which serves as the processing unit. Next, each speech frame data in the sequence is processed as follows: First, Mel-frequency cepstral coefficient feature extraction is performed on the speech frame data to obtain a Mel-frequency cepstral coefficient feature vector. This yields a Mel-frequency cepstral coefficient feature vector that meets the requirements of the input sentiment data model and can represent the spectral characteristics of the speech frame data to be processed. Then, based on the speech frame data to be processed, speech domain feature data is generated. Therefore, speech domain feature data reflecting the fundamental frequency, average frequency, and other speech frame characteristics in the frequency domain of the speech frame data to be processed can be obtained. Then, based on the aforementioned Mel-frequency cepstral coefficient feature vector and the aforementioned speech domain feature data, a speech feature vector is generated. Thus, the Mel-frequency cepstral coefficient feature vector and the speech domain feature data can be concatenated to obtain the speech feature vector, which is then input into the emotion data model to generate emotion data. Next, the aforementioned speech feature vector is input into a pre-trained emotion data model to obtain speech frame emotion data. Thus, speech frame emotion data representing the emotional state of the speakers in the speech frame data to be processed can be obtained. Next, based on the obtained speech frame emotion data, a speech frame emotion data sequence is generated. Then, based on the aforementioned speech frame emotion data sequence, an emotion curve and emotion index data are generated. Thus, an emotion curve and emotion index data representing the emotional changes of the speakers during the dialogue can be obtained. Finally, the aforementioned emotion curve and emotion index data are identified as the emotion data. Thus, emotion data used to generate a dialogue quality score can be obtained. This is also because the extraction of features such as Mel frequency cepstral coefficients and fundamental frequency are lightweight signal processing operations, requiring less computational resources during processing. At the same time, the pre-trained sentiment data model used (essentially a single neural network model) is small in scale and has fast inference, which reduces the overall computational resources consumed in a single analysis.

[0109] The above-described embodiments of this disclosure have the following beneficial effects: the dialogue text parsing method of some embodiments of this disclosure reduces the waste of broadband resources. Specifically, the reason for the waste of broadband resources is that by directly calling a general large language model for text analysis using prompt words, the generation results are arbitrary and uncontrollable due to the lack of targeted training of the model. This leads to a chaotic format of the output parsed information, which is prone to omitting key information or generating "illusionary" information. As a result, there are errors in the parsed information sent to downstream business systems. The erroneous parsed information may not have practical significance in the downstream system, but it still needs to be transmitted through the network, which increases the possibility of regenerating and sending dialogue text parsing information, thus wasting bandwidth resources. Based on this, the dialogue text parsing method of some embodiments of this disclosure first obtains a set of item flow dialogue text information. Then, the above-mentioned item flow dialogue text information set is input into a pre-trained dialogue analysis model to obtain a labeled dataset, wherein each labeled data in the above-mentioned labeled dataset includes a dialogue summary, tag text, keyword information, and flow result information. Then, the aforementioned set of dialogue text information about the item flow and the aforementioned labeled dataset are sent to a pre-set manual processing terminal for manual review and correction, resulting in a corrected dataset. Each corrected data point in the corrected dataset corresponds to a labeled data point in the labeled dataset. This yields a manually corrected dataset that meets the requirements. Next, based on the corrected dataset and the aforementioned set of dialogue text information about the item flow, training and testing datasets are generated. Then, based on the training and testing datasets, a pre-set large-scale model is trained to obtain a trained dialogue text extraction model. Thus, the training and testing datasets generated from the manually corrected data can be used to specifically train the model, resulting in a dialogue text extraction model that extracts information from dialogue text, outputting extracted information that better meets the parsing requirements. Then, the dialogue text information to be processed is obtained. Then, the dialogue text information to be processed is input into the dialogue text extraction model to obtain extracted information. This extracted information is then used for subsequent parsing processing. Finally, the extracted information is parsed to obtain outputtable parsing information, and the dialogue text information to be processed and the outputtable parsing information are sent to the preset downstream business system. Thus, outputtable parsing information reflecting the dialogue summary, tag text, keyword information, and flow result information in the dialogue text information to be processed can be obtained and sent to the preset downstream business system. This is because the process involves first performing dialogue analysis on the item flow dialogue text information set to obtain an annotated dataset, and then manually reviewing and correcting the annotated data to obtain a corrected dataset that meets the requirements.Furthermore, training and correction datasets generated from targeted and corrected data of the dialogue text information of the goods flow are used to train the dialogue of the pre-set base large model. This reduces the interference of cluttered information on model learning, improves the accuracy and reliability of extracting key information, and thus provides downstream business systems with high-quality parsed information that can be directly processed and applied. This reduces the possibility of regenerating and sending dialogue text parsing information, thereby reducing the waste of broadband resources.

[0110] Further reference Figure 2 As an implementation of the methods shown in the figures, this disclosure provides some embodiments of a dialog text parsing apparatus, which are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0111] like Figure 2 As shown, the dialogue text parsing device 200 in some embodiments includes: a first acquisition unit 201, a first input unit 202, a first sending unit 203, a generation unit 204, a training unit 205, a second acquisition unit 206, a second input unit 207, and a second sending unit 208. The first acquisition unit 201 is configured to acquire a set of dialogue text information related to the flow of goods; the first input unit 202 is configured to input the set of dialogue text information related to the flow of goods into a pre-trained dialogue analysis model to obtain a labeled dataset, wherein each labeled data in the labeled dataset includes a dialogue summary, tag text, keyword information, and flow result information; the first sending unit 203 is configured to send the set of dialogue text information related to the flow of goods and the labeled dataset to a preset manual processing terminal for manual review and correction to obtain a corrected dataset, wherein each corrected data in the corrected dataset corresponds to a labeled data in the labeled dataset; the generation unit 204 is configured to... The system is configured to generate a training dataset and a test dataset based on the aforementioned calibration dataset and the aforementioned item flow dialogue text information set; the training unit 205 is configured to train a preset base large model based on the aforementioned training dataset and the aforementioned test dataset to obtain a trained dialogue text information extraction model; the second acquisition unit 206 is configured to acquire dialogue text information to be processed; the second input unit 207 is configured to input the aforementioned dialogue text information to be processed into the aforementioned dialogue text information extraction model to obtain extracted information; the second sending unit 208 is configured to parse the aforementioned extracted information to obtain output parsing information and send the aforementioned dialogue text information to be processed and the aforementioned output parsing information to a preset downstream business system.

[0112] It is understandable that the units described in the device 200 are related to the reference. Figure 1The steps in the method described above correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.

[0113] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0114] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0115] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.

[0116] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of technical features, but should also cover other technical solutions formed by arbitrary combinations of technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for parsing dialogue text, comprising: Obtain the item transfer dialogue text information set; The set of dialogue text information about the item flow is input into a pre-trained dialogue analysis model to obtain a labeled dataset. Each labeled data in the labeled dataset includes a dialogue summary, label text, keyword information, and flow result information. The item flow dialogue text information set and the labeled dataset are sent to a preset manual processing terminal for manual review and correction to obtain a corrected dataset, wherein each corrected data in the corrected dataset corresponds to a labeled data in the labeled dataset; Based on the calibration dataset and the item flow dialogue text information set, a training dataset and a test dataset are generated; Based on the training dataset and the test dataset, the preset base model is trained to obtain a trained dialogue text information extraction model. Obtain the text information of the dialogue to be processed; The dialogue text information to be processed is input into the dialogue text information extraction model to obtain the extracted information; The extracted information is parsed to obtain output parsable information, and the dialogue text information to be processed and the output parsable information are sent to a preset downstream business system.

2. The method according to claim 1, wherein, Each corrected data point in the corrected dataset includes a dialogue summary, tagged text, and keyword information flow results. Each tagged data point in the labeled dataset corresponds to an item flow dialogue text in the item flow dialogue text information set. Based on the corrected dataset and the item flow dialogue text information set, a training dataset and a test dataset are generated, including: Retrieve preset command words; Perform the following operations on each correction data point in the correction dataset: Based on the dialogue summary included in the correction data, generate target summary key-value pairs; Based on the label text included in the correction data, generate target label key-value pairs; Based on the keyword information included in the correction data, generate key-value pairs of target keyword information; Based on the flow result information included in the correction data, generate target result information key-value pairs; Based on the target summary key-value pairs, the target tag key-value pairs, the target keyword information key-value pairs, and the target result information key-value pairs, target output information is generated; The item flow dialogue text information in the set of item flow dialogue text information that corresponds to the correction data is determined as the input information; The preset instruction word, the target output information, and the input information are determined as a ternary information group; Each obtained ternary information group is defined as a ternary information group set; Based on the aforementioned ternary information set, training and testing datasets are generated.

3. The method according to claim 1, wherein, The step of training a pre-defined large-scale model based on the training dataset and the test dataset to obtain a trained dialogue text information extraction model includes: The training dataset is reconstructed to obtain a summary training dataset, a tag training dataset, a keyword training dataset, and a flow result training dataset. Obtain preset parameters, wherein the preset parameters include the matrix rank and the learning rate; Based on the preset parameters, a parameter tuning matrix is ​​generated; The parameter tuning matrix is ​​integrated with the preset base model, and the integrated preset base model is determined as the base training sub-model. Based on the aforementioned summary training dataset, the base training sub-model is trained to obtain the summary extraction model. Based on the aforementioned tag training set, the base training sub-model is trained to obtain the tag extraction model; Based on the keyword training set, the base training sub-model is trained to obtain the keyword extraction model; Based on the training set of the circulation results, the training sub-model of the base is trained to obtain the circulation result extraction model; The abstract extraction model, the tag extraction model, the keyword extraction model, the flow result extraction model, and the preset base model are integrated to obtain the base training model. Based on the training dataset, a training data set is generated; Based on the training data set and the test dataset, the base training model is trained to obtain the dialogue text information extraction model.

4. The method according to claim 3, wherein, The step of training the large-scale model on the base based on the training data set and the test dataset to obtain the dialogue text information extraction model includes: Based on the training dataset, a large model is trained on the base by performing the following training steps: Randomly select a training data set from the training data set set as the target training data set, and delete the training data set from the training data set set; The target training data in the target training data group is segmented and filled with words to obtain the word index sequence of each word; Input each word index sequence into the base to train the large model and obtain each probability distribution sequence; Based on each probability distribution sequence and each word index sequence, an average loss value is generated; Based on the average loss value and the learning rate in the preset parameters, the parameters of the large model trained on the base are updated. The average loss value is determined as the historical loss value, and the historical loss value is stored in a preset historical loss value queue for updating the preset historical loss value queue. In response to the determination that there is a historical loss value enqueue operation in the preset historical loss value queue and the number of historical loss values ​​in the preset historical loss value queue after the enqueue operation is equal to the length of the preset historical loss value queue, a historical loss characterization value is generated based on each historical loss value in the preset historical loss value queue. In response to the determination that the historical loss representation value meets the preset completion condition or the training data set is empty, the base training large model with updated parameters is updated and verified based on the test dataset to obtain the dialogue text information extraction model.

5. The method according to claim 4, wherein, The process of updating and validating the large-scale training model on the base with updated parameters based on the test dataset yields a dialogue text information extraction model, including: In response to determining that the historical loss representation value meets the preset completion condition or the training data set is empty, the word segmentation and filling process is performed on each test data in the test dataset to obtain each test word index sequence, and the base training large model after updating the input parameters of each obtained test word index sequence is used to obtain each prediction probability distribution sequence. Based on each predicted probability distribution sequence and each test word index sequence, a test loss value is generated; In response to the determination that the test loss value is greater than the preset loss value, a new training data set is generated based on the training dataset, and a large model is trained on the base with updated parameters based on the training data set, and the training steps are executed again. In response to the determination that the test loss value is less than or equal to the preset loss value, the base training large model with updated parameters is determined as the dialogue text information extraction model.

6. The method according to claim 1, wherein, The step of parsing the extracted information to obtain outputtable parsable information and sending the dialogue text to be processed and the outputtable parsable information to a preset downstream business system includes: The extracted information is input into a preset extracted information parser to obtain a text information dictionary; Retrieve a preset dictionary sequence of historical text information; The text information dictionary is added to the preset historical text information dictionary sequence to update the preset historical text information dictionary sequence; Based on the text information dictionary, dictionary determination information is generated; In response to determining that the dictionary determination information meets the preset conditions, the text information dictionary is determined to be output parsable information; In response to determining that the dictionary judgment information does not meet the preset conditions, a prompt message is generated and the dialogue text information to be processed and the text information dictionary are sent to a preset manual processing terminal, so that the preset manual processing terminal can review and repair the text information dictionary to obtain the repaired text information dictionary, and determine the repaired text information dictionary as output parsable information; The outputtable parsable information is sent to a preset downstream business system.

7. The method according to claim 6, wherein, After adding the text information dictionary to the preset historical text information dictionary sequence, the method further includes: Obtain the preset reward formula and the preset historical text information dictionary sequence; Based on the preset reward formula, the following steps are performed on the preset historical text information dictionary sequence: The last historical text information dictionary in the preset historical text information dictionary sequence is determined as the reference text information dictionary; Based on the preset reward formula and the reference text information dictionary, global reward value information for text information is generated. Based on the global reward value of the text information, the parameter strategy of the dialogue text information extraction model is optimized to obtain the fine-tuned dialogue text information extraction model. The fine-tuned dialogue text extraction model was determined as the dialogue text extraction model.

8. A dialog text parsing device, comprising: The first acquisition unit is configured to acquire a set of dialogue text information about item flow. The first input unit is configured to input the item flow dialogue text information set into a pre-trained dialogue analysis model to obtain a labeled dataset, wherein each labeled data in the labeled dataset includes a dialogue summary, label text, keyword information and flow result information; The first sending unit is configured to send the item flow dialogue text information set and the labeled dataset to a preset manual processing terminal for manual review and correction to obtain a corrected dataset, wherein each corrected data in the corrected dataset corresponds to a labeled data in the labeled dataset. The generation unit is configured to generate a training dataset and a test dataset based on the calibration dataset and the item flow dialogue text information set; The training unit is configured to train a preset base model based on the training dataset and the test dataset to obtain a trained dialogue text information extraction model. The second acquisition unit is configured to acquire the dialogue text information to be processed. The second input unit is configured to input the dialogue text information to be processed into the dialogue text information extraction model to obtain extracted information. The second sending unit is configured to parse the extracted information to obtain output parsable information and send the dialogue text information to be processed and the output parsable information to a preset downstream business system.

9. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 7.

10. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 7.