Auxiliary mediation method and system based on large model and storage medium
Through the large-model-based auxiliary mediation method, the dispute exchange voices are analyzed, mediation suggestions and risk levels are generated, and the problem of inefficient mediation in the existing technology is solved, and a more efficient and accurate mediation process is achieved.
Patent Information
- Application Number
- CN202510205523.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Mediation efficiency in the prior art is low, mediators are seriously insufficient, and quality assurance is difficult to achieve, resulting in low mediation efficiency.
The auxiliary mediation method based on the big model is adopted to obtain the voice of dispute communication, conduct speaker analysis and voice recognition, generate the text of the dispute speaker, and input it into the pre-trained big model for mediation analysis, generate mediation suggestions and risk levels, and send mediation prompts and risk control alarms to the mediator.
It reduces the decision-making cost of mediators, improves mediation efficiency, enables mediation to achieve results faster, reduces the work burden of mediators, and improves the efficiency and accuracy of the mediation process.
Smart Images

Figure CN120015035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a large model-based auxiliary mediation method, system and storage medium. Background Art
[0002] In the public security and judicial systems, mediation is an effective means to resolve conflicts and disputes among residents in the jurisdiction and to eliminate risks at the subtlest level. In order to improve the utilization rate of judicial resources, the issue of mediation efficiency has received more and more attention.
[0003] In the existing mediation process, mediation is generally conducted manually by mediators, resulting in a serious shortage of mediators, the quality of mediators cannot be guaranteed, and the efficiency of mediation is reduced. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide an auxiliary mediation method, system and storage medium based on a large model to solve the problem of low mediation efficiency in the prior art.
[0005] The embodiment of the present invention is implemented as follows: an auxiliary mediation method based on a large model, the method comprising:
[0006] Acquire dispute communication voice, and perform speaker analysis on the dispute communication voice to obtain a speaker analysis result;
[0007] Segmenting the dispute communication speech by speaker according to the speaker analysis result to obtain speaker speech, and performing speech recognition on the speaker speech to obtain dispute speaker text;
[0008] Input the dispute speaker text into the pre-trained large model for mediation analysis to obtain mediation suggestions and risk levels, and send mediation prompts to the mediator based on the mediation suggestions;
[0009] If the risk level is greater than the level threshold, a dispute risk control alarm is sent to the mediator.
[0010] Preferably, before inputting the dispute speaker text into the pre-trained large model for mediation analysis, the method further includes:
[0011] Obtaining a sample of a dispute speaker, and inputting the sample of the dispute speaker into the large model for word segmentation to obtain a sample word segmentation;
[0012] Extracting features from the sample word segments to obtain context sample features, and decoding the context sample features to obtain decoded sample features;
[0013] Performing mediation prediction according to the decoded sample features to obtain a sample prediction suggestion and a sample prediction grade, and determining a model loss according to the sample prediction suggestion and the sample prediction grade;
[0014] The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
[0015] Preferably, feature extraction is performed on the sample word segmentation to obtain context sample features, including:
[0016] Performing word embedding processing on the sample word segmentation to obtain a word vector, and performing self-attention mechanism calculation on the word vector to obtain a self-attention vector;
[0017] Performing a multi-head attention mechanism calculation on the self-attention vector to obtain a multi-head attention vector, and normalizing the multi-head attention vector to obtain a normalized feature;
[0018] Feedforward processing is performed on the normalized features to obtain feedforward features, and residual processing is performed on the feedforward features to obtain the context sample features.
[0019] Preferably, performing speaker analysis on the dispute communication voice to obtain a speaker analysis result includes:
[0020] Performing speech segmentation on the dispute communication speech to obtain communication segmentation speech, and obtaining Mel frequency cepstrum coefficients, fundamental frequency and formant frequency of the communication segmentation speech;
[0021] Combining the Mel frequency cepstral coefficient, the fundamental frequency and the formant frequency to obtain a segmented speech feature, and performing vector conversion on the segmented speech feature to obtain a segmented feature vector;
[0022] Classify the communication segmented speech according to the segmentation feature vector to obtain a segmented speech set, and obtain a face image of the disputer corresponding to the dispute communication speech;
[0023] A pronunciation time set is determined according to the face image, and a voice correspondence relationship between the disputer and the segmented voice set is determined according to the pronunciation time set to obtain the speaker analysis result.
[0024] Preferably, the communication segmented speech is classified according to the segmentation feature vector to obtain a segmented speech set, including:
[0025] Calculating the similarities between adjacent segmentation feature vectors to obtain a first vector similarity;
[0026] If the first vector similarity is greater than a first similarity threshold, merging the communication segmented speech corresponding to the segmented feature vector to obtain segmented merged speech;
[0027] Calculating an average vector of the segmentation feature vectors in the segmented and merged speech to obtain a segmentation average vector, and calculating similarities between different segmentation average vectors to obtain a second vector similarity;
[0028] If the second vector similarity is greater than a second similarity threshold, the different segmented and merged speech corresponding to the second vector similarity are divided into sets to obtain the segmented speech set.
[0029] Preferably, determining the pronunciation time set according to the face image includes:
[0030] Performing feature point recognition on the face image to obtain face feature points, and determining a lip image based on the face feature points;
[0031] Obtaining the opening distance between the upper lip and the lower lip in the lip image, and calculating the distance difference of the opening distance between adjacent lip images;
[0032] If the distance difference is greater than the difference threshold, the adjacent lip image corresponding to the distance difference is determined as a speaking image, and the image acquisition time of the speaking image is obtained to obtain the pronunciation acquisition time;
[0033] The pronunciation collection time is stored in a preset set to obtain the pronunciation time set.
[0034] Preferably, after sending a mediation reminder to the mediator according to the mediation suggestion, the method further includes:
[0035] Obtaining mediation feedback for the mediation suggestion, and generating parameter fine-tuning samples according to the mediation feedback, the dispute speaker text and the mediation suggestion;
[0036] The parameters of the pre-trained large model are adjusted according to the parameter fine-tuning samples.
[0037] Another object of an embodiment of the present invention is to provide an auxiliary mediation system based on a large model, the system comprising:
[0038] A speaker analysis module is used to obtain dispute communication voice, and perform speaker analysis on the dispute communication voice to obtain a speaker analysis result;
[0039] A speech recognition module, used to perform speaker segmentation on the dispute communication speech according to the speaker analysis result to obtain speaker speech, and perform speech recognition on the speaker speech to obtain the dispute speaker text;
[0040] The mediation prompt module is used to input the dispute speaker text into the pre-trained large model for mediation analysis, obtain mediation suggestions and risk levels, and send mediation prompts to the mediator based on the mediation suggestions; if the risk level is greater than the level threshold, a dispute risk control alarm is sent to the mediator.
[0041] Preferably, the mediation prompt module is also used for:
[0042] Obtaining a sample of a dispute speaker, and inputting the sample of the dispute speaker into the large model for word segmentation to obtain a sample word segmentation;
[0043] Extracting features from the sample word segments to obtain context sample features, and decoding the context sample features to obtain decoded sample features;
[0044] Performing mediation prediction according to the decoded sample features to obtain a sample prediction suggestion and a sample prediction grade, and determining a model loss according to the sample prediction suggestion and the sample prediction grade;
[0045] The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
[0046] The embodiments of the present invention can effectively determine the speaker voices corresponding to each speaker in the dispute communication voice by performing speaker analysis on the dispute communication voice, can effectively convert the speaker voices into dispute speaker text by performing voice recognition on the speaker voices, can automatically generate mediation suggestions based on the big model by inputting the dispute speaker text into a pre-trained big model for mediation analysis, and can send mediation prompts to the mediator through the mediation suggestions, thereby reducing the decision-making cost of the mediator, enabling mediation to achieve results more quickly, and improving mediation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flow chart of an auxiliary mediation method based on a large model provided in a first embodiment of the present invention;
[0048] Figure 2 is a schematic diagram of the structure of an auxiliary mediation system based on a large model provided in a second embodiment of the present invention;
[0049] Figure 3 It is a schematic diagram of specific implementation steps of the auxiliary mediation system based on a large model provided in the second embodiment of the present invention;
[0050] Figure 4 It is a schematic diagram of the structure of a terminal device provided in the third embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0052] In order to illustrate the technical solution of the present invention, a specific embodiment is provided below for illustration.
[0053] Embodiment 1
[0054] See also Figure 1 , is a flow chart of an auxiliary mediation method based on a large model provided in a first embodiment of the present invention. The auxiliary mediation method based on a large model can be applied to any device or system. The auxiliary mediation method based on a large model includes the following steps:
[0055] Step S10, obtaining dispute communication voice, and performing speaker analysis on the dispute communication voice to obtain a speaker analysis result;
[0056] Optionally, performing speaker analysis on the dispute communication voice to obtain a speaker analysis result includes:
[0057] Segment the dispute communication speech to obtain the communication segmentation speech, and obtain the Mel-frequency cepstral coefficients, fundamental frequency and formant frequency of the communication segmentation speech; wherein, the communication segmentation speech is passed through a Mel filter group to obtain the energy of different Mel frequency bands, and a discrete cosine transform is performed to obtain the Mel-frequency cepstral coefficients, the Mel-frequency cepstral coefficients can effectively reflect the spectral envelope characteristics of the speech, the fundamental frequency reflects the frequency of the vocal cord vibration, and the fundamental frequency range of different speakers is often different, and the formant frequency is related to the shape and size of the vocal tract and is also speaker-specific;
[0058] The Mel frequency cepstral coefficient, the pitch frequency and the formant frequency are combined to obtain a segmented speech feature, and the segmented speech feature is vector-converted to obtain a segmented feature vector; wherein, by combining the Mel frequency cepstral coefficient, the pitch frequency and the formant frequency, the segmented speech feature reflecting the speech representation of the speaker can be effectively obtained;
[0059] Classify the communication segmented speech according to the segmentation feature vector to obtain a segmented speech set, and obtain a face image of the disputer corresponding to the dispute communication speech;
[0060] A pronunciation moment set is determined according to the facial image, and a voice correspondence relationship between the disputer and the segmented speech set is determined according to the pronunciation moment set to obtain the speaker analysis result; wherein, the voice correspondence relationship between the disputer and the segmented speech set is determined based on the matching relationship between the pronunciation collection moment in the pronunciation moment set and the communication segmented speech in the segmented speech set.
[0061] Furthermore, the communication segmented speech is classified according to the segmentation feature vector to obtain a segmented speech set, including:
[0062] Calculating the similarity between adjacent segmentation feature vectors to obtain a first vector similarity; wherein the distance between adjacent segmentation feature vectors can be calculated using a Euclidean distance formula to obtain the first vector similarity;
[0063] If the first vector similarity is greater than a first similarity threshold, the communication segmented speech corresponding to the segmentation feature vector is voice-merged to obtain a segmented merged speech; wherein the first similarity threshold can be set according to demand, and if the first vector similarity is greater than the first similarity threshold, it is determined that the communication segmented speech corresponding to the segmentation feature vector is the speech emitted by the same speaker, and therefore, the corresponding communication segmented speech is voice-merged to obtain a segmented merged speech;
[0064] Calculating an average vector of the segmentation feature vectors in the segmented and merged speech to obtain a segmentation average vector, and calculating similarities between different segmentation average vectors to obtain a second vector similarity;
[0065] If the second vector similarity is greater than the second similarity threshold, the different segmented and merged speech corresponding to the second vector similarity are divided into groups to obtain the segmented speech set; wherein, the second similarity threshold can be set according to demand, and if the second vector similarity is greater than the second similarity threshold, it is determined that the different segmented and merged speech corresponding to the second vector similarity are speech uttered by the same speaker, and therefore, the different segmented and merged speech corresponding to the second vector similarity are divided into groups to obtain a segmented speech set.
[0066] Furthermore, determining a pronunciation time set according to the face image includes:
[0067] Performing feature point recognition on the face image to obtain face feature points, and determining the lip image based on the face feature points; wherein, by performing feature point recognition on the face image, the face feature points in the face image can be effectively extracted, the lip position can be determined based on the face feature points, and the lip image can be determined based on the lip position;
[0068] Obtaining the opening distance between the upper lip and the lower lip in the lip image, and calculating the distance difference of the opening distance between adjacent lip images;
[0069] If the distance difference is greater than the difference threshold, the adjacent lip images corresponding to the distance difference are determined as speaking images, the image acquisition time of the speaking image is obtained, the pronunciation acquisition time is obtained, and the pronunciation acquisition time is stored in a preset set to obtain the pronunciation time set, wherein the difference threshold can be set according to demand. If the distance difference is greater than the difference threshold, it is determined that the adjacent lip images correspond to the speaker speaking, therefore, the image acquisition time of the speaking image is obtained to obtain the pronunciation acquisition time.
[0070] Step S20, performing speaker segmentation on the dispute communication speech according to the speaker analysis result to obtain speaker speech, and performing speech recognition on the speaker speech to obtain the dispute speaker text;
[0071] Among them, based on the speaker analysis result, the speaker of the dispute communication voice is distinguished to obtain the speaker voice, and the speaker voice is recognized to obtain the dispute speaker text corresponding to each speaker.
[0072] Step S30, inputting the dispute speaker text into the pre-trained large model for mediation analysis, obtaining mediation suggestions and risk levels, and sending mediation prompts to the mediator according to the mediation suggestions;
[0073] Among them, mediation analysis is performed by inputting the text of the dispute speaker into a pre-trained big model, and mediation suggestions are automatically generated based on the big model. Mediation prompts are sent to the mediator through the mediation suggestions, thereby reducing the decision-making cost of the mediator.
[0074] Optionally, before inputting the dispute speaker text into the pre-trained large model for mediation analysis, the method further includes:
[0075] Obtain a sample of dispute speakers, and input the sample of dispute speakers into the large model for word segmentation to obtain sample word segmentation; wherein, collect mediation-related data from various channels, such as real mediation case documents, legal mediation document libraries, case collections of professional mediation agencies, etc., to ensure that the corpus covers various types of mediation cases, such as civil disputes, labor disputes, neighborhood conflicts, etc. Clean up the collected corpus and remove duplicate, erroneous, and irregularly formatted data. For example, correct typos, fill in missing information, unify date formats, etc. Annotate the corpus to clarify the information of different parts, such as the basic situation of the case, the mediation process, the mediation results, etc. For risk assessment-related information, mark the key factors affecting the risk assessment and the corresponding risk level to obtain a sample of dispute speakers. Evaluate and select suitable pre-trained large models, such as the GPT series, BERT, etc., based on factors such as computing resources and model performance. Consider the model's versatility, parameter count, and fine-tuning adaptability in natural language processing tasks;
[0076] Performing feature extraction on the sample segmentation to obtain context sample features, and performing feature decoding on the context sample features to obtain decoded sample features; wherein, by performing feature extraction on the sample segmentation, context sample features containing context representation are obtained;
[0077] Mediation prediction is performed based on the decoded sample features to obtain sample prediction suggestions and sample prediction levels, and the model loss is determined based on the sample prediction suggestions and the sample prediction levels; wherein, based on the decoder in the large model, the mediation suggestions are outputted in a word-by-word generation manner, and in this step, the decoded sample features are mapped to the dimensions of risk assessment through one or more fully connected layers. For example, if the risk is divided into three levels: high, medium, and low, the large model will eventually output a three-dimensional vector, and the value of each dimension represents the probability of the corresponding risk level, and the sample prediction level is obtained. The cross-entropy loss is generally used to measure the difference between the generated sample prediction suggestions and the true mediation suggestions. For each word unit generated, the large model will predict a word unit probability distribution, and the cross-entropy loss will calculate the difference between the predicted distribution and the true word unit (that is, the word unit in the annotated correct mediation suggestion). By minimizing the cross-entropy loss, the model will gradually adjust the parameters to make the generated mediation suggestions closer to the actual situation; for risk assessment, the commonly used loss function is the cross-entropy loss or the mean square error loss. Taking the classification problem as an example, the cross-entropy loss between the risk level probability distribution output by the model and the true annotated risk level will guide the large model to adjust the parameters and improve the accuracy of risk assessment;
[0078] The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model; wherein, after the calculation of the model loss is completed, the gradient of the loss to the large model parameters is calculated by the back propagation algorithm, and the gradient indicates how a small change in the parameter will affect the loss value, indicating the direction of the parameter update, and the optimizer (such as Adam, SGD, etc.) is used to update the model parameters according to the calculated gradient.
[0079] Furthermore, feature extraction is performed on the sample word segmentation to obtain context sample features, including:
[0080] The sample word segmentation is processed by word embedding to obtain a word vector, and the word vector is calculated by a self-attention mechanism to obtain a self-attention vector; wherein each word vector calculates the degree of association with other word vectors in the text through the self-attention mechanism, thereby capturing the long-distance dependency relationship in the text. For example, when analyzing the statements of the parties in a mediation case, the large model can focus on key information, such as the focus of the dispute, the demands of the parties, etc., through the self-attention mechanism;
[0081] A multi-head attention mechanism is performed on the self-attention vector to obtain a multi-head attention vector, and the multi-head attention vector is normalized to obtain a normalized feature, wherein the multi-head attention mechanism is performed on the self-attention vector based on the multi-layer Transformer block in the large model, and the output of the self-attention mechanism passes through the multi-layer Transformer block. The multi-layer Transformer block calculates the attention of multiple heads in parallel, and each head focuses on different aspects of the input sequence, thereby capturing richer semantic information and dependencies. Specifically, a weighted context representation is generated for each position by calculating the similarity between the query, key, and value;
[0082] The normalized features are subjected to feedforward processing to obtain feedforward features, and the feedforward features are subjected to residual processing to obtain the context sample features; wherein, the output of the multi-head attention mechanism is subjected to layer normalization operation to normalize the data in order to accelerate the convergence of large models and reduce the problem of gradient vanishing or exploding, and the normalized results are input into a feedforward neural network, which usually includes two fully connected layers to perform further feature extraction and transformation on the data, and after the output of the feedforward neural network, a residual connection is performed to add the input of the current layer to the output of the feedforward neural network to help the model be better trained, solve the problem of gradient vanishing and degradation, and enable the model to learn long-term dependencies more easily.
[0083] Furthermore, after sending a mediation prompt to the mediator according to the mediation suggestion, it also includes: obtaining mediation feedback for the mediation suggestion, generating parameter fine-tuning samples according to the mediation feedback, the dispute speaker text and the mediation suggestion, and performing parameter adjustment on the pre-trained large model according to the parameter fine-tuning samples, wherein, based on the mediator's mediation feedback on the mediation suggestion, parameter fine-tuning samples can be effectively generated, and the pre-trained large model is adjusted by the parameter fine-tuning samples, so as to achieve the effect of real-time and continuous model optimization, thereby improving the model quality of the pre-trained large model.
[0084] Step S40: If the risk level is greater than the level threshold, a dispute risk control alarm is issued to the mediator;
[0085] Among them, the level threshold can be set according to needs. If the risk level is greater than the level threshold, it is determined that the current dispute is more serious and needs to be transferred to the public security department (to deal with more serious crimes) for processing. Therefore, a dispute risk control alarm is sent to the mediator to prompt the mediator to transfer the case.
[0086] In this embodiment, by performing speaker analysis on the dispute communication voice, the speaker voice corresponding to each speaker in the dispute communication voice can be effectively determined, and by performing voice recognition on the speaker voice, the speaker voice can be effectively converted into dispute speaker text. By inputting the dispute speaker text into a pre-trained large model for mediation analysis, mediation suggestions are automatically generated in a large model-based manner, and mediation prompts are sent to the mediator through the mediation suggestions, thereby reducing the decision-making cost of the mediator, enabling mediation to achieve results more quickly, and improving mediation efficiency.
[0087] Embodiment 2
[0088] See also Figure 2 , is a schematic diagram of the structure of a large model-based auxiliary mediation system 100 provided in a second embodiment of the present invention, including:
[0089] The speaker analysis module 10 is used to obtain dispute communication speech, and perform speaker analysis on the dispute communication speech to obtain a speaker analysis result.
[0090] Optionally, the speaker analysis module 10 is further used to: perform speech segmentation on the dispute communication speech to obtain communication segmentation speech, and obtain Mel frequency cepstral coefficients, fundamental frequency and formant frequency of the communication segmentation speech;
[0091] Combining the Mel frequency cepstral coefficient, the fundamental frequency and the formant frequency to obtain a segmented speech feature, and performing vector conversion on the segmented speech feature to obtain a segmented feature vector;
[0092] Classify the communication segmented speech according to the segmentation feature vector to obtain a segmented speech set, and obtain a face image of the disputer corresponding to the dispute communication speech;
[0093] A pronunciation time set is determined according to the face image, and a voice correspondence relationship between the disputer and the segmented voice set is determined according to the pronunciation time set to obtain the speaker analysis result.
[0094] Furthermore, the speaker analysis module 10 is further used to: calculate the similarity between adjacent segmented feature vectors to obtain a first vector similarity;
[0095] If the first vector similarity is greater than a first similarity threshold, merging the communication segmented speech corresponding to the segmented feature vector to obtain segmented merged speech;
[0096] Calculating an average vector of the segmentation feature vectors in the segmented and merged speech to obtain a segmentation average vector, and calculating similarities between different segmentation average vectors to obtain a second vector similarity;
[0097] If the second vector similarity is greater than a second similarity threshold, the different segmented and merged speech corresponding to the second vector similarity are divided into sets to obtain the segmented speech set.
[0098] Furthermore, the speaker analysis module 10 is also used to: perform feature point recognition on the face image to obtain face feature points, and determine the lip image according to the face feature points;
[0099] Obtaining the opening distance between the upper lip and the lower lip in the lip image, and calculating the distance difference of the opening distance between adjacent lip images;
[0100] If the distance difference is greater than the difference threshold, the adjacent lip image corresponding to the distance difference is determined as a speaking image, and the image acquisition time of the speaking image is obtained to obtain the pronunciation acquisition time;
[0101] The pronunciation collection time is stored in a preset set to obtain the pronunciation time set.
[0102] The speech recognition module 11 is used to perform speaker segmentation on the dispute communication speech according to the speaker analysis result to obtain speaker speech, and perform speech recognition on the speaker speech to obtain the dispute speaker text.
[0103] The mediation prompt module 12 is used to input the dispute speaker text into the pre-trained large model for mediation analysis, obtain mediation suggestions and risk levels, and send mediation prompts to the mediator based on the mediation suggestions; if the risk level is greater than the level threshold, a dispute risk control alarm is sent to the mediator.
[0104] Optionally, the mediation prompt module 12 is further used to: obtain a sample of a dispute speaker, and input the sample of the dispute speaker into the large model for word segmentation to obtain a sample word segmentation;
[0105] Extracting features from the sample word segments to obtain context sample features, and decoding the context sample features to obtain decoded sample features;
[0106] Performing mediation prediction according to the decoded sample features to obtain a sample prediction suggestion and a sample prediction grade, and determining a model loss according to the sample prediction suggestion and the sample prediction grade;
[0107] The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
[0108] Furthermore, the mediation prompt module 12 is further used to: perform word embedding processing on the sample word segmentation to obtain a word vector, and perform self-attention mechanism calculation on the word vector to obtain a self-attention vector;
[0109] Performing a multi-head attention mechanism calculation on the self-attention vector to obtain a multi-head attention vector, and normalizing the multi-head attention vector to obtain a normalized feature;
[0110] Feedforward processing is performed on the normalized features to obtain feedforward features, and residual processing is performed on the feedforward features to obtain the context sample features.
[0111] Preferably, the mediation prompt module 12 is further used to: obtain mediation feedback for the mediation suggestion, and generate a parameter fine-tuning sample according to the mediation feedback, the dispute speaker text and the mediation suggestion;
[0112] The parameters of the pre-trained large model are adjusted according to the parameter fine-tuning samples.
[0113] See also Figure 3 The specific implementation steps of the auxiliary mediation system 100 based on the large model include:
[0114] 1. During the mediation process, the auxiliary mediation system 100 based on the large model is started to collect sound in real time;
[0115] 2. Use TTS speech recognition technology to recognize the recorded content into text and distinguish different speakers based on their timbre;
[0116] 3. Use mediation corpus to fine-tune the big model so that it has professional knowledge in the field of mediation, can generate mediation suggestions in real time based on the mediation content, and conduct risk assessment on the case;
[0117] 4. During the mediation process, the big model processes the dialogue content during the mediation process, generates mediation suggestions in real time, and assesses the risk level of the current case;
[0118] 5. The police can conduct more professional mediation between the two parties of the case based on the mediation suggestions; and pay attention to the risk assessment level in a timely manner. When the big model assesses that the case is serious, a risk control alarm will be issued. After confirmation, the police can promptly transfer the case to the public security department (to deal with more serious crimes);
[0119] 6. During the mediation process, the system updates the mediation suggestions and risk levels in real time;
[0120] 7. Regularly use mediation data to further fine-tune the big model and enhance the big model’s professional capabilities and risk assessment capabilities in the mediation field.
[0121] Through systematic risk assessment and intelligent mediation suggestions, the decision-making costs of mediators are reduced, and mediation can be achieved more quickly, thereby improving mediation efficiency. The large model generates mediation suggestions based on the context of the conversation, which is more objective than the traditional way of mediators providing mediation suggestions based on personal experience. The system needs to be evaluated before it is put into use to ensure that the product launched has a mediation accuracy rate above the average. As an assistant to the mediator, it can point out errors to the mediator and reduce the probability of system or personnel making mistakes.
[0122] In this embodiment, by performing speaker analysis on the dispute communication voice, the speaker voice corresponding to each speaker in the dispute communication voice can be effectively determined, and by performing voice recognition on the speaker voice, the speaker voice can be effectively converted into dispute speaker text. By inputting the dispute speaker text into a pre-trained large model for mediation analysis, mediation suggestions are automatically generated in a large model-based manner, and mediation prompts are sent to the mediator through the mediation suggestions, thereby reducing the decision-making cost of the mediator, enabling mediation to achieve results more quickly, and improving mediation efficiency.
[0123] Embodiment 3
[0124] Figure 4 2 is a block diagram of a terminal device 2 provided in the third embodiment of the present application. Figure 4As shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program of the auxiliary mediation method based on a large model. When the processor 20 executes the computer program 22, the steps in each embodiment of the auxiliary mediation method based on a large model are implemented.
[0125] Exemplarily, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.
[0126] The processor 20 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0127] The memory 21 may be an internal storage unit of the terminal device 2, such as a hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 2. Further, the memory 21 may also include both an internal storage unit and an external storage device of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is to be output.
[0128] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional unit.
[0129] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunication signals.
[0130] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. An auxiliary mediation method based on a large model, characterized in that: The method comprises: Acquire dispute communication voice, and perform speaker analysis on the dispute communication voice to obtain a speaker analysis result; Segmenting the dispute communication speech by speaker according to the speaker analysis result to obtain speaker speech, and performing speech recognition on the speaker speech to obtain dispute speaker text; Input the dispute speaker text into the pre-trained large model for mediation analysis to obtain mediation suggestions and risk levels, and send mediation prompts to the mediator based on the mediation suggestions; If the risk level is greater than the level threshold, a dispute risk control alarm is sent to the mediator.
2. The large model-based assisted mediation method according to claim 1, characterized in that: Before inputting the dispute speaker text into the pre-trained large model for mediation analysis, it also includes: Obtaining a sample of a dispute speaker, and inputting the sample of the dispute speaker into the large model for word segmentation to obtain a sample word segmentation; Extracting features from the sample word segments to obtain context sample features, and decoding the context sample features to obtain decoded sample features; Performing mediation prediction according to the decoded sample features to obtain a sample prediction suggestion and a sample prediction grade, and determining a model loss according to the sample prediction suggestion and the sample prediction grade; The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
3. The large model-based auxiliary mediation method according to claim 2, characterized in that: Perform feature extraction on the sample word segmentation to obtain context sample features, including: Performing word embedding processing on the sample word segmentation to obtain a word vector, and performing self-attention mechanism calculation on the word vector to obtain a self-attention vector; Performing a multi-head attention mechanism calculation on the self-attention vector to obtain a multi-head attention vector, and normalizing the multi-head attention vector to obtain a normalized feature; Feedforward processing is performed on the normalized features to obtain feedforward features, and residual processing is performed on the feedforward features to obtain the context sample features.
4. The large model-based auxiliary mediation method according to claim 1, characterized in that: Perform speaker analysis on the dispute communication voice to obtain speaker analysis results, including: Performing speech segmentation on the dispute communication speech to obtain communication segmentation speech, and obtaining Mel frequency cepstral coefficients, fundamental frequency and formant frequency of the communication segmentation speech; Combining the Mel frequency cepstral coefficient, the fundamental frequency and the formant frequency to obtain a segmented speech feature, and performing vector conversion on the segmented speech feature to obtain a segmented feature vector; Classify the communication segmented speech according to the segmentation feature vector to obtain a segmented speech set, and obtain a face image of the disputer corresponding to the dispute communication speech; A pronunciation time set is determined according to the face image, and a voice correspondence relationship between the disputer and the segmented voice set is determined according to the pronunciation time set to obtain the speaker analysis result.
5. The large model-based auxiliary mediation method according to claim 4, characterized in that: The communication segmented speech is classified according to the segmentation feature vector to obtain a segmented speech set, including: Calculating the similarities between adjacent segmentation feature vectors to obtain a first vector similarity; If the first vector similarity is greater than a first similarity threshold, merging the communication segmented speech corresponding to the segmented feature vector to obtain segmented merged speech; Calculating an average vector of the segmentation feature vectors in the segmented and merged speech to obtain a segmentation average vector, and calculating similarities between different segmentation average vectors to obtain a second vector similarity; If the second vector similarity is greater than a second similarity threshold, the different segmented and merged speech corresponding to the second vector similarity are divided into sets to obtain the segmented speech set.
6. The large model-based auxiliary mediation method according to claim 4, characterized in that: Determining a pronunciation time set according to the face image includes: Performing feature point recognition on the face image to obtain face feature points, and determining a lip image based on the face feature points; Obtaining the opening distance between the upper lip and the lower lip in the lip image, and calculating the distance difference of the opening distance between adjacent lip images; If the distance difference is greater than the difference threshold, the adjacent lip image corresponding to the distance difference is determined as a speaking image, and the image acquisition time of the speaking image is obtained to obtain the pronunciation acquisition time; The pronunciation collection time is stored in a preset set to obtain the pronunciation time set.
7. The large model-based assisted mediation method according to claim 1, characterized in that: After sending a mediation reminder to the mediator according to the mediation proposal, it also includes: Obtaining mediation feedback for the mediation suggestion, and generating parameter fine-tuning samples according to the mediation feedback, the dispute speaker text and the mediation suggestion; The parameters of the pre-trained large model are adjusted according to the parameter fine-tuning samples.
8. An auxiliary mediation system based on a large model, characterized in that: The system comprises: A speaker analysis module is used to obtain dispute communication voice, and perform speaker analysis on the dispute communication voice to obtain a speaker analysis result; A speech recognition module, used to perform speaker segmentation on the dispute communication speech according to the speaker analysis result to obtain speaker speech, and perform speech recognition on the speaker speech to obtain the dispute speaker text; The mediation prompt module is used to input the dispute speaker text into the pre-trained large model for mediation analysis, obtain mediation suggestions and risk levels, and send mediation prompts to the mediator based on the mediation suggestions; if the risk level is greater than the level threshold, a dispute risk control alarm is sent to the mediator.
9. The large model-based auxiliary mediation system according to claim 8, characterized in that: The mediation prompt module is also used for: Obtaining a sample of a dispute speaker, and inputting the sample of the dispute speaker into the large model for word segmentation to obtain a sample word segmentation; Extracting features from the sample word segments to obtain context sample features, and decoding the context sample features to obtain decoded sample features; Performing mediation prediction according to the decoded sample features to obtain a sample prediction suggestion and a sample prediction grade, and determining a model loss according to the sample prediction suggestion and the sample prediction grade; The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Automatic collection method for multivariate social contradictory dispute information
CN117056510A
Mediation strategy output method, device and equipment based on intelligent interactive questions and answers
CN118262725A
Voice content compliance intelligent judgment and risk grading system based on AI large model
CN118609567A
Upper welding rods inspection apparatus and inspection method using the same
KR1020240138741A
KR20220117802A