Mediation phrase generation method and device based on large model, equipment and medium

By generating mediation scripts using a large model and utilizing improved speech separation and text conversion technologies, combined with scene information to select script templates, the problem of authenticity caused by fixed scripts in AI mediation has been solved, and more logical and realistic dialogue generation has been achieved.

CN120877733BActive Publication Date: 2025-12-05BEIJING HUAYU JIUPIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511383665.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-12-05
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

In existing AI mediation processes, the use of fixed scripts causes users to perceive the communicator as a robot, generating negative feelings and affecting the authenticity of the communication.

Method used

A large model is used to generate mediation scripts. By acquiring speech separation models and online mediation recording data, the target speech separation model is improved, effective recording data is selected, mediation dialogue text is generated, mediation script templates are constructed, and scripts are selected and generated based on scenario information.

Benefits of technology

It improves the realism of AI calls, making the generated mediation scripts more closely resemble the current conversation, logical, and enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877733B_ABST
    Figure CN120877733B_ABST
Patent Text Reader

Abstract

The application relates to a mediation speech generation method and device based on a large model, equipment and a medium, applied to the technical field of model training, which comprises the following steps: acquiring a voice separation model and online mediation recording data; improving and training the voice separation model to generate a target voice separation model; performing screening processing on the online mediation recording data to obtain effective recording mediation data; performing data processing on the effective recording mediation data based on the target voice separation model and a preset text conversion model to generate mediation dialogue text; generating a mediation speech template based on the mediation dialogue text and preset special knowledge data; in response to a template use instruction, acquiring scene information of the template use instruction; selecting the mediation speech template based on the scene information to obtain a target speech template; and generating mediation speech based on the target speech template. The application has the effect of improving the authenticity of AI calls.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model training, and in particular to a method, apparatus, device and medium for generating mediation scripts based on a large model. Background Technology

[0002] In recent years, with the continuous development of AI, this process is increasingly being handled by outbound call robots, making AI-driven mediation an inevitable trend. To achieve AI-driven mediation, the following points need to be considered: first, understanding the current situation of the mediation event; second, understanding user information in real time and providing corresponding responses; third, guiding the user's conversation to focus on mediation-related topics and avoiding idle chatter; and fourth, providing solutions based on the current situation, historical chat logs, and the demands of both parties.

[0003] However, existing AI calls all use pre-set fixed scripts for communication. This communication method can make people immediately perceive that the communicator is a robot, thus generating a negative feeling and ending the conversation. How to improve the authenticity of AI calls has become an important problem to be solved. Summary of the Invention

[0004] To improve the realism of AI calls, this application provides a method, apparatus, device, and medium for generating mediation scripts based on a large model.

[0005] Firstly, this application provides a method for generating mediation scripts based on a large model, employing the following technical solution:

[0006] A method for generating mediation scripts based on a large model, comprising:

[0007] Acquire speech separation model and online mediation recording data;

[0008] The speech separation model is improved and trained to generate a target speech separation model;

[0009] The online mediation recording data is filtered to obtain valid mediation recording data;

[0010] Based on the target speech separation model and the preset text conversion model, the effective recorded mediation data is processed to generate mediation dialogue text.

[0011] Generate a mediation script template based on the mediation dialogue text and preset special knowledge data;

[0012] In response to a template usage instruction, obtain the scenario information of the template usage instruction;

[0013] Based on the scenario information, the mediation script template is selected to obtain the target script template;

[0014] Mediation scripts are generated based on the target script template.

[0015] By adopting the above technical solution, online mediation recording data is collected, and the speech separation model that needs improvement is determined. The speech separation model is improved and trained to obtain a target speech separation model with more accurate speech separation effect. The obtained online mediation recording data is filtered to obtain usable effective recording mediation data. The target speech separation model and a preset text conversion model are used to perform speech separation and text conversion processing on the effective recording mediation data to obtain the mediation dialogue text. The mediation dialogue text and special knowledge data are used to construct a mediation script template. During AI calls, the scene information of the call is collected, and a template is selected from the constructed mediation script template according to the scene information. The obtained target script template is used to generate the mediation script, making the generated mediation script closer to the current dialogue situation and more logical, thereby improving the realism of AI calls.

[0016] Optionally, the step of improving and training the speech separation model to generate the target speech separation model includes:

[0017] Dynamic convolutional enhancement modules are inserted between encoder layers to dynamically capture speech harmonic structures using learnable offsets.

[0018] Set up a hierarchical feature fusion mechanism to connect the lower and higher layers;

[0019] A dual-branch role separation network was constructed, the voice recognition module was optimized, and voice activity detection was added.

[0020] Optionally, the step of processing the effective recorded mediation data based on the target speech separation model and the preset text conversion model to generate mediation dialogue text includes:

[0021] Based on the target speech separation model, the effective recording mediation data is separated to generate dialogue roles and corresponding dialogue speech information;

[0022] Based on the preset text conversion model, the dialogue voice information is processed to generate text data.

[0023] The text data is subjected to terminology correction processing to generate processed text data;

[0024] The processed text data is bound to the corresponding dialogue roles to generate mediation dialogue text.

[0025] Optionally, generating a mediation script template based on the mediation dialogue text and preset special knowledge data includes:

[0026] Identify sensitive data in the mediation dialogue text;

[0027] Based on a preset desensitization strategy and the sensitive data, the mediation dialogue text is dynamically desensitized to generate desensitized dialogue text.

[0028] The special knowledge data is standardized to generate a special knowledge graph;

[0029] Obtain dialogue behavior tags and template construction logic;

[0030] Add the dialogue behavior tags to the desensitized dialogue text to generate tag-desensitized dialogue text;

[0031] Based on the anonymized dialogue text with the aforementioned tags and the template construction logic, a mediation script template is constructed.

[0032] Optionally, the step of selecting the mediation script template based on the scenario information to obtain the target script template includes:

[0033] Obtain the template type of the mediation script template;

[0034] The scene information is matched with the template type to generate a matching result;

[0035] Based on the matching results, a target dialogue template is generated by selecting from the mediation dialogue templates.

[0036] Optionally, generating mediation scripts based on the target script template includes:

[0037] Determine whether the matching result is empty;

[0038] If the matching result is not empty, then obtain the speech information in the target speech template;

[0039] Generate mediation scripts based on the aforementioned script information;

[0040] If the matching result is empty, then a mediation script is generated based on the scenario information and the preset large model.

[0041] Optionally, after generating the mediation script template based on the mediation dialogue text and preset special knowledge data, the method further includes:

[0042] Obtain updated information on the evaluation system, evaluation dimensions, and specific data.

[0043] The mediation script template is evaluated based on the evaluation system and the evaluation dimensions to generate an evaluation result.

[0044] Obtain the mediation script generated based on the mediation script template;

[0045] The mediation script template is iteratively optimized based on the special data update information, the evaluation results, and the mediation script.

[0046] Secondly, this application provides a mediation script generation device based on a large model, which adopts the following technical solution:

[0047] A mediation script generation device based on a large model, comprising:

[0048] The model data acquisition module is used to acquire speech separation model and online mediation recording data;

[0049] A separation model generation module is used to improve and train the speech separation model to generate a target speech separation model.

[0050] The effective data filtering module is used to filter the online mediation recording data to obtain effective mediation recording data;

[0051] The dialogue text generation module is used to process the effective recorded mediation data based on the target speech separation model and the preset text conversion model to generate mediation dialogue text.

[0052] The dialogue template generation module is used to generate a mediation dialogue template based on the mediation dialogue text and preset special knowledge data;

[0053] The scene information acquisition module is used to acquire scene information of the template usage instruction in response to the template usage instruction;

[0054] The target template selection module is used to select the mediation script template based on the scenario information to obtain the target script template;

[0055] The mediation script generation module is used to generate mediation scripts based on the target script template.

[0056] By adopting the above technical solution, online mediation recording data is collected, and the speech separation model that needs improvement is determined. The speech separation model is improved and trained to obtain a target speech separation model with more accurate speech separation effect. The obtained online mediation recording data is filtered to obtain usable effective recording mediation data. The target speech separation model and a preset text conversion model are used to perform speech separation and text conversion processing on the effective recording mediation data to obtain the mediation dialogue text. The mediation dialogue text and special knowledge data are used to construct a mediation script template. During AI calls, the scene information of the call is collected, and a template is selected from the constructed mediation script template according to the scene information. The obtained target script template is used to generate the mediation script, making the generated mediation script closer to the current dialogue situation and more logical, thereby improving the realism of AI calls.

[0057] Thirdly, this application provides an electronic device that adopts the following technical solution:

[0058] An electronic device includes a processor coupled to a memory;

[0059] The processor is configured to execute a computer program stored in the memory, such that the electronic device executes the computer program of the large-model-based mediation script generation method as described in any of the first aspects.

[0060] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution:

[0061] A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing the mediation script generation method based on a large model as described in any of the first aspects. Attached Figure Description

[0062] Figure 1 This is a flowchart illustrating a mediation script generation method based on a large model, as provided in an embodiment of this application.

[0063] Figure 2 This is a structural block diagram of a mediation script generation device based on a large model provided in an embodiment of this application.

[0064] Figure 3 This is a structural block diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0065] The present application will be further described in detail below with reference to the accompanying drawings.

[0066] This application provides a method for generating mediation scripts based on a large model. This method can be executed by an electronic device, which can be a server or a terminal device. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smartphone, tablet computer, desktop computer, etc., but is not limited to these.

[0067] Figure 1 This is a flowchart illustrating a mediation script generation method based on a large model, provided as an embodiment of this application.

[0068] like Figure 1 As shown, the main process of this method is described below (steps S101 to S108):

[0069] Step S101: Obtain the speech separation model and online mediation recording data.

[0070] In this embodiment, the voice separation model is a model that separates the recording, dividing it into multiple separate audio files according to the number of participants. For example, in a call recording, if two people are involved, the voice separation model will extract the individual conversations of each person separately, resulting in two separate audio files containing only one person's voice. The online mediation recording data refers to the mediation recording data generated during historical call mediation processes.

[0071] Step S102: Improve and train the speech separation model to generate the target speech separation model.

[0072] For step S102, a dynamic convolutional enhancement module is inserted between encoder layers to dynamically capture the speech harmonic structure through learnable offsets; a hierarchical feature fusion mechanism is set to connect the bottom layer and the top layer; a dual-branch role separation network is constructed to optimize the voice recognition module and increase speech activity detection.

[0073] In this embodiment, the speech separation model uses the MossFormer2 model. Since noise and speech have local correlation in the time-frequency domain, such as the short-term continuity of burst noise, and the previous speech separation model is difficult to take into account both local features and global context, the speech separation model is improved and trained. In low signal-to-noise ratio scenarios, speech clarity is enhanced by time-frequency masking generative adversarial network, that is, overlapping separation and word ordering after separation, as well as short word recognition, thereby improving the MOS score.

[0074] The specific improvements consist of three parts: adding a dynamic convolution enhancement module, hierarchical feature fusion, and constructing a dual-branch role separation network. The dynamic convolution enhancement module inserts deformable convolutions between encoder layers, dynamically capturing speech harmonic structures, such as the continuity of the fundamental frequency F0, using learnable offsets. The convolution kernels are set to 3×3, with 8 offsets per group, covering a 20ms time window. Hierarchical feature fusion connects the bottom and top layers. The bottom layer includes Layers 1-4, used to process 10ms frame-level features, focusing on eliminating transient impulse noise. The top layer includes Layers... Sections 5-8 are used to model a 500ms long-term context and suppress steady-state background noise such as air conditioner sounds. A dual-branch role separation network is constructed, consisting of two branches. Branch 1 builds an ECAPA-TDNN speaker recognition module based on timbre features to improve accuracy. This involves optimizing the ECAPA-TDNN structure, performing multi-scale feature aggregation, introducing cross-frame skip connections between TDNN layers, fusing contextual information from 1 / 3 / 7 frames to improve the robustness of timbre features, enhancing channel attention, and adding a Squeeze-Excitation module before the statistical pooling layer to dynamically strengthen key frequency band features. Branch 2 constructs a speech activity detection network based on energy thresholds to solve the problem of separating overlapping speech. Based on these processes, independent audio tracks for different roles, such as mediators and users, can be output, reducing the separation error rate. The extracted 192-dimensional speaker embedding vector is used as the conditional input to the separation network. Separation is guided by feature splicing or cross-attention mechanisms, and a decoupling strategy for overlapping regions is added. Voiceprint contrast constraints are added, and a constraint to maximize cross-track voiceprint differences is added to the loss function. Based on voiceprint features, the typical fundamental frequency range of the target speaker is predicted, such as 100-150Hz for males and 200-300Hz for females. Harmonic component matching is performed in overlapping regions, and an identity-aware speech separation architecture is implemented. The role embedding conditional input includes strong identity binding and end-to-end optimization. Strong identity binding uses the extracted 192-dimensional high-recognition voiceprint features as decoding conditions to ensure that the mediator's multiple speeches are always classified as the same track and that the identity is not lost when the party is emotionally agitated. End-to-end optimization directly applies the timbre contrast loss function to the separation process and sets the cosine distance between the output tracks of different roles to be greater than or equal to 0.65, thereby improving the accuracy of separation and reducing noise interference.

[0075] Step S103: Filter and process the online mediation recording data to obtain valid mediation recording data.

[0076] In this embodiment, some of the obtained online mediation recording data may be unusable or invalid. Therefore, the obtained online mediation recording data is filtered to select valid recordings that have passed, been successfully mediated, or failed. A voice quality assessment model is constructed, and low-quality recordings are filtered based on three indicators: signal-to-noise ratio, voice continuity, and effective dialogue ratio. Samples with a pass rate are selected and the selected samples are used as valid recording mediation data.

[0077] Step S104: Based on the target speech separation model and the preset text conversion model, the effective recorded mediation data is processed to generate mediation dialogue text.

[0078] For step S104, the effective recorded mediation data is separated based on the target speech separation model to generate dialogue roles and corresponding dialogue speech information; the dialogue speech information is processed into text based on a preset text conversion model to generate text data; the text data is processed for terminology correction to generate processed text data; and the processed text data is bound to the corresponding dialogue roles to generate mediation dialogue text.

[0079] In this embodiment, the constructed target speech separation model is used to separate roles and dialogue data in valid recording mediation data. This determines the number of dialogue roles present in a single valid recording mediation dataset and separates the overall speech according to these roles, obtaining the dialogue speech information corresponding to each role. The target speech separation model yields dialogue roles and dialogue speech information with higher correlation, resulting in more accurate recognition and segmentation. After obtaining the dialogue speech information, a preset text-to-text conversion model is used to convert the audio data into text data, i.e., the dialogue speech information undergoes text-to-text conversion processing to obtain the corresponding text data. This data may contain out-of-order text or useless interjections. Therefore, error correction processing is performed on the text data. A regular expression mapping table containing common errors in special terms is constructed. For example, the error text is: Debt transfer, the regular expression pattern is Debt [transfer] transfer, and the correction result is Debt transfer. Redundant words are detected, and a blacklist of redundant words is constructed to achieve collaborative optimization of interjection filtering, removing redundant words such as "uh" and "ah", retaining the core semantics, and removing documents unrelated to financial mediation, such as chat records. Core content such as dispute descriptions and agreement texts are retained to obtain the final processed text data. The processed text data is then bound to the corresponding dialogue roles to obtain the mediation dialogue text.

[0080] Step S105: Generate a mediation script template based on the mediation dialogue text and preset special knowledge data.

[0081] For step S105, identify sensitive data in the mediation dialogue text; perform dynamic desensitization of sensitive information on the mediation dialogue text based on the preset desensitization strategy and sensitive data to generate desensitized dialogue text; standardize special knowledge data to generate a special knowledge graph; obtain dialogue behavior tags and template construction logic; add dialogue behavior tags to the desensitized dialogue text to generate tagged desensitized dialogue text; and construct a mediation script template based on the tagged desensitized dialogue text and template construction logic.

[0082] In this embodiment, to hide users' personal information, the mediation dialogue text needs to be anonymized. A three-level anonymization strategy is designed, namely a preset anonymization strategy. The regular expression matching layer quickly identifies explicit sensitive information such as phone numbers; the semantic reasoning layer detects implicit personal information based on a context-sensitive entity recognition model of a large model, such as "Zhang Mingliang owes 500,000" → "Mr. Zhang owes yuan"; the adversarial substitution layer uses a large model to generate semantically consistent virtual substitute content to avoid the text being reversible after anonymization. The preset anonymization strategy is used to complete the dynamic anonymization operation of sensitive information in the dialogue text, resulting in an anonymized dialogue text. Then, a special knowledge graph is constructed using special knowledge data. Clauses are extracted from multiple relevant official documents using entity extraction, and a "dispute type-mandatory clause-reference case" triple relationship is established. Based on the BERT-CRF model, the text clauses are automatically aligned with the FLKG. Warnings are triggered for dialogue segments lacking mandatory clauses, such as interest rate disputes that do not cite Article 680 of the official document. A special knowledge graph for handling pending warnings is constructed. In addition, it is necessary to add dialogue behavior tags, that is, to add at least one behavior tag to a dialogue to standardize the definition of the current dialogue. Dialogue behavior tags include requesting mediation, presenting and examining evidence, special clarification, agreement negotiation, emotional soothing, reaching an agreement, etc., which need to be set in advance according to the actual situation. Then, according to the corresponding meaning, they are added to each sentence of the dialogue to obtain the tag-desensitized dialogue text. Then, according to the template construction logic, the dialogue data in the tag-desensitized dialogue text is added to the corresponding logical positions to construct the mediation dialogue template. For example, in the interest rate dispute scenario, the three-part structure of "explanation of special terms → calculation demonstration → solution confirmation" is constructed to generate 8 major categories and 26 subcategories of standard dialogue templates, specifically: 1. Identity confirmation; 1.1 Customer denial; 1.2 Customer request for rescheduling; 1.3 Confirmation time; 1.4 Customer questions identity; 1.5 Call not answered by the person in question; 1.6 Strong resistance; 2. Confirmation of debt authenticity; 2.1 Acknowledgment of debt; 3. Dispute resolution process; 3.1 Factual disputes; 3.2 Procedural disputes; 3.3 Acceptance of debt authenticity; 4. Repayment ability assessment; 4.1 Verification process; 4.2 Income covers debt; 4.3 Insufficient income; 4.4 Assets available; 4.5 No assets available; 4.6 Request for delayed mediation; 5. Negotiation of repayment plan; 5.1 Standard plan; 5.2 Adjustment of down payment; 6. Notification of special consequences; 6.1 Explanation of litigation process; 6.2 Warning of property preservation; 6.3 Simultaneous litigation process; 7. Agreement signing and execution; 7.1 Agreement generation; 7.2 After signing; 7.3 Not signed; 8. Case closure and archiving; 8.1 Withdrawal of case after agreement execution; 8.2 Handling of non-execution; 8.3 End of process. It should be noted that the constructed adjustment script templates are templates corresponding to different response content. When generating dialogues, multiple templates need to be recombined in order for use.

[0083] Step S106: In response to the template usage instruction, obtain the scenario information of the template usage instruction.

[0084] In this embodiment, after all the construction is completed, in response to the template usage instruction, it is triggered at the start of the dialogue and the current scene information is collected. The scene information is the current dialogue stage and dialogue content.

[0085] Step S107: Select a mediation script template based on the scenario information to obtain the target script template.

[0086] For step S107, obtain the template type of the mediation script template; match the scenario information with the template type to generate a matching result; select from the mediation script templates based on the matching result to generate the target script template.

[0087] In this embodiment, the mediation dialogue template is selected using scene information. First, the template type corresponding to each template in the dialogue template is determined. It is necessary to ensure that the template type matches the scene information, so that the corresponding template is returned as the matching result. The template in the returned template matching result is used as the target speech template. In some special cases, the matching result may be empty, and the corresponding target speech template will also be empty.

[0088] Step S108: Generate mediation script based on the target script template.

[0089] For step S108, determine whether the matching result is empty; if the matching result is not empty, obtain the script information in the target script template; generate mediation script based on the script information; if the matching result is empty, generate mediation script based on the scenario information and the preset large model.

[0090] In this embodiment, when generating mediation scripts, it is necessary to first determine whether the matching result is empty. If the matching result is empty, it means that the target script template is also empty and there is no available template to provide the script. The script needs to be generated using a preset large model based on the current scenario information to obtain the final mediation script. If the matching result is not empty, it means that there is a target script template and the script content in the target script template will be used directly as the mediation script.

[0091] In this embodiment, the evaluation system, evaluation dimensions, and special data update information are obtained; the mediation script template is evaluated based on the evaluation system and evaluation dimensions to generate evaluation results; the mediation script generated based on the mediation script template is obtained; and the mediation script template is iteratively optimized based on the special data update information, evaluation results, and mediation script.

[0092] An evaluation system and evaluation dimensions were established, and these were used to evaluate the mediation script templates. This included evaluations of specific data compliance, communication effectiveness, and dynamic evaluation using the entropy weight method. The specific data compliance evaluation included calculating the coverage rate of mandatory clauses and the accuracy of specific information citations. Communication effectiveness included dispute resolution rate (percentage of dialogues reaching an agreement), average number of dialogue rounds (ideally <15 rounds), and user experience. User experience included user emotional stability (voice feature analysis) and acceptance rate of the mediation outcome. The dynamic evaluation using the entropy weight method included incorporating Monte Carlo simulations to evaluate the model's robustness in extreme scenarios. The mediation scripts generated using the templates, along with updated specific data information, were obtained. The evaluation results, the updated scripts, and the updated specific data information were then used to optimize and iterate the mediation script templates, generating templates that better suit actual needs and current information.

[0093] Figure 2 The present invention relates to a structural block diagram of a mediation script generation device 200 based on a large model, provided in the embodiments of the application.

[0094] like Figure 2 As shown, the mediation script generation device 200 based on a large model mainly includes:

[0095] The model data acquisition module 201 is used to acquire speech separation model and online mediation recording data;

[0096] The speech separation model generation module 202 is used to improve and train the speech separation model to generate the target speech separation model;

[0097] The effective data filtering module 203 is used to filter and process online mediation recording data to obtain effective mediation recording data;

[0098] The dialogue text generation module 204 is used to process effective recorded mediation data based on the target speech separation model and the preset text conversion model to generate mediation dialogue text.

[0099] The dialogue template generation module 205 is used to generate mediation dialogue templates based on mediation dialogue text and preset special knowledge data.

[0100] The scene information acquisition module 206 is used to acquire scene information of the template usage instruction in response to the template usage instruction;

[0101] The target template selection module 207 is used to select a mediation script template based on scenario information to obtain the target script template;

[0102] Mediation script generation module 208 is used to generate mediation scripts based on the target script template.

[0103] As an optional implementation of this embodiment, the separation model generation module 202 is specifically used to insert a dynamic convolution enhancement module between encoder layers, dynamically capture speech harmonic structure through learnable offsets; set up a hierarchical feature fusion mechanism to connect the bottom layer and the top layer; construct a dual-branch role separation network, optimize the voice recognition module, and increase voice activity detection.

[0104] As an optional implementation of this embodiment, the dialogue text generation module 204 is specifically used to separate and process the effective recorded mediation data based on the target speech separation model to generate dialogue roles and corresponding dialogue speech information; to perform text conversion processing on the dialogue speech information based on a preset text conversion model to generate text data; to perform terminology correction processing on the text data to generate processed text data; and to bind the processed text data with the corresponding dialogue roles to generate mediation dialogue text.

[0105] As an optional implementation of this embodiment, the script template generation module 205 is specifically used to determine sensitive data in the mediation dialogue text; perform dynamic desensitization of sensitive information on the mediation dialogue text based on a preset desensitization strategy and sensitive data to generate desensitized dialogue text; perform standardization processing on special knowledge data to generate a special knowledge graph; obtain dialogue behavior tags and template construction logic; add dialogue behavior tags to the desensitized dialogue text to generate tagged desensitized dialogue text; and construct a mediation script template based on the tagged desensitized dialogue text and template construction logic.

[0106] As an optional implementation of this embodiment, the target template selection module 207 is specifically used to obtain the template type of the mediation script template; match the usage scenario information with the template type to generate a matching result; and select from the mediation script templates based on the matching result to generate a target script template.

[0107] As an optional implementation of this embodiment, the mediation script generation module 208 is specifically used to determine whether the matching result is empty; if the matching result is not empty, the script information in the target script template is obtained; the mediation script is generated based on the script information; if the matching result is empty, the mediation script is generated based on the scenario information and the preset large model.

[0108] As an optional implementation of this embodiment, the mediation script generation device 200 based on a large model further includes:

[0109] The assessment update acquisition module is used to acquire update information on the assessment system, assessment dimensions, and special data.

[0110] The evaluation result generation module is used to evaluate the mediation script template based on the evaluation system and evaluation dimensions, and generate evaluation results.

[0111] The mediation script acquisition module is used to acquire mediation scripts generated based on mediation script templates;

[0112] The template iteration and optimization module is used to iteratively optimize the mediation script template based on special data update information, evaluation results, and mediation scripts.

[0113] In one example, the module in any of the above devices may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), or one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs), or a combination of at least two of these integrated circuit forms.

[0114] For example, when modules in a device can be implemented via a processing element scheduler, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling programs. Alternatively, these modules can be integrated together as a system-on-a-chip (SOC).

[0115] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0116] Figure 3 This is a structural block diagram of the electronic device 300 provided in an embodiment of this application.

[0117] like Figure 3 As shown, the electronic device 300 includes a processor 301 and a memory 302, and may further include one or more of an information input / output (I / O) interface 303, a communication component 304, and a communication bus 305.

[0118] The processor 301 controls the overall operation of the electronic device 300 to complete all or part of the steps of the large-model-based mediation script generation method described above. The memory 302 stores various types of data to support the operation of the electronic device 300. This data may include, for example, instructions for any application or method operating on the electronic device 300, as well as application-related data. The memory 302 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as one or more of Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0119] I / O interface 303 provides an interface between processor 301 and other interface modules, such as keyboards, mice, and buttons. These buttons can be virtual or physical. Communication component 304 is used for wired or wireless communication between electronic device 300 and other devices. Wireless communication includes Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof. Therefore, the corresponding communication component 304 may include a Wi-Fi component, a Bluetooth component, and an NFC component.

[0120] The electronic device 300 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the mediation script generation method based on a large model given in the above embodiments.

[0121] The communication bus 305 may include a path for transmitting information between the aforementioned components. The communication bus 305 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 305 may be divided into an address bus, a data bus, a control bus, etc.

[0122] Electronic device 300 may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers, and may also be servers.

[0123] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for generating mediation scripts based on a large model.

[0124] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0125] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0126] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions claimed in this application.

Claims

1. A mediation dialogue generation method based on a large model, characterized by, The method comprises the following steps: obtain a voice separation model and online mediation recording data; improve the training of the voice separation model to generate a target voice separation model; screen the online mediation recording data to obtain valid recording mediation data; perform data processing on the valid recording mediation data based on the target voice separation model and a preset text conversion model to generate mediation dialogue text; generate a mediation dialogue template based on the mediation dialogue text and preset special knowledge data; comprising: determine sensitive data in the mediation dialogue text; perform a sensitive information dynamic desensitization operation on the mediation dialogue text based on a preset desensitization strategy and the sensitive data to generate a desensitized dialogue text; standardize the special knowledge data to generate a special knowledge graph; obtain a dialogue behavior label and a template construction logic; add the dialogue behavior label to the desensitized dialogue text to generate a labeled desensitized dialogue text; construct a mediation dialogue template based on the labeled desensitized dialogue text and the template construction logic; in response to a template usage instruction, obtain scenario information of the template usage instruction; select the mediation dialogue template based on the scenario information to obtain a target dialogue template; generate a mediation dialogue based on the target dialogue template.

2. The method of claim 1, wherein, The method comprises the following steps: insert a dynamic convolution enhancement module between the encoder layers to dynamically capture the harmonic structure of the voice through a learnable offset; set a hierarchical feature fusion mechanism to connect the bottom layer and the high layer; construct a double-branch role separation network to optimize the sound recognition module and increase the voice activity detection.

3. The method of claim 1, wherein, The method comprises the following steps: perform separation processing on the valid recording mediation data based on the target voice separation model to generate dialogue roles and dialogue voice information corresponding to the dialogue roles; perform text conversion processing on the dialogue voice information based on the preset text conversion model to generate text data; perform term error correction processing on the text data to generate processed text data; bind the processed text data to the corresponding dialogue roles to generate mediation dialogue text.

4. The method of claim 1, wherein, The method comprises the following steps: obtain the template type of the mediation dialogue template; match the scenario information with the template type to generate a matching result; select the mediation dialogue template based on the matching result to generate a target dialogue template.

5. The method of claim 4, wherein, The method comprises the following steps: determine whether the matching result is empty; if the matching result is not empty, obtain dialogue information in the target dialogue template; generate a mediation dialogue based on the dialogue information; if the matching result is empty, generate a mediation dialogue based on the scenario information and a preset large model.

6. The method of claim 1, wherein, After the mediation dialogue template is generated based on the mediation dialogue text and the preset special knowledge data, the method further comprises the following steps: obtain an evaluation system, an evaluation dimension, and special data update information; The mediation speech template is evaluated based on the evaluation system and the evaluation dimension, and an evaluation result is generated; Obtaining mediation speech generated based on the mediation speech template; Based on the special data update information, the evaluation result and the mediation speech, the mediation speech template is iteratively optimized. 7.A mediation phrase generation apparatus based on a large model, characterized by Comprising: The model data acquisition module is used for acquiring a speech separation model and online mediation recording data; The separation model generation module is used for improving training of the speech separation model to generate a target speech separation model; The effective data screening module is used for screening the online mediation recording data to obtain effective recording mediation data; The dialogue text generation module is used for data processing of the effective recording mediation data based on the target speech separation model and a preset text conversion model to generate mediation dialogue text; The speech template generation module is used for generating a mediation speech template based on the mediation dialogue text and preset special knowledge data; comprising: Determine the sensitive data in the mediation dialogue text; Based on a preset desensitization strategy and the sensitive data, the mediation dialogue text is dynamically desensitized to generate a desensitized dialogue text; The special knowledge data is standardized to generate a special knowledge graph; Obtain the dialogue behavior label and the template construction logic; The dialogue behavior label is added to the desensitized dialogue text to generate a labeled desensitized dialogue text; Based on the labeled desensitized dialogue text and the template construction logic, a mediation speech template is constructed; The scene information acquisition module is used for acquiring scene information of the template use instruction in response to the template use instruction; The target template selection module is used for selecting the mediation speech template based on the scene information to obtain a target speech template; The mediation speech generation module is used for generating mediation speech based on the target speech template.

8. An electronic device, comprising: Comprising a processor coupled with a memory; The processor is used to execute a computer program stored in the memory, so that the electronic device executes the method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Comprising a computer program or instructions, when the computer program or instructions run on a computer, make the computer execute the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • AI mediation system based on large data model

    CN119128106A

  • Voice-based role separation method and device, equipment and medium

    CN119724196A