A digital mediation method and system fusing multi-modal sentiment analysis
By using an AI-based mediation system to conduct multimodal sentiment analysis, the problems of low efficiency and low strategy adaptability in traditional mediation methods have been solved, realizing intelligent dispute mediation and improving mediation efficiency and strategy adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN FANRAN INFORMATION TECH CO LTD
- Filing Date
- 2026-03-11
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional mediation methods suffer from low efficiency and poor adaptability of mediation strategies when faced with dynamically changing emotions, making it difficult to achieve real-time analysis and dynamic response to multi-dimensional emotional states.
An AI-powered mediation system, comprising a large-scale mediation model system, an AI mediation engine system, and an AI interaction system, is employed to collect multimodal data and perform sentiment analysis. This generates mediation robot templates and parameter sets, and intelligent mediation is achieved through sentiment analysis annotation and mediation template matching.
It improves mediation efficiency, enhances the adaptability of mediation strategies, better addresses the dynamic emotional changes in disputes, and generates effective mediation outcome reports.
Smart Images

Figure CN122133803A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction technology, and in particular to a digital intelligence mediation method and system that integrates multimodal emotion analysis. Background Technology
[0002] With the rapid development of society and the economy, the number of civil disputes, commercial disputes, and neighborhood conflicts has continued to increase. Dispute mediation, as a core means of efficiently resolving conflicts, reducing judicial costs, and maintaining social stability, is increasingly in demand in the fields of judicial assistance and social governance. The introduction of intelligent mediation technology into the field of mediation provides a new, intelligent, and efficient approach to dispute mediation, and has significant application value.
[0003] Traditional mediation methods primarily rely on human mediators who resolve disputes through face-to-face communication and experience-based judgment. While this approach can resolve conflicts, it faces challenges in handling the ever-increasing number of disputes. The dynamic nature of the parties' emotional states during mediation, coupled with the limited experience of human mediators, makes it difficult to achieve real-time analysis and dynamic response to multi-dimensional emotional states. This results in shortcomings in mediation efficiency and professionalism. Consequently, current dispute mediation processes suffer from low efficiency and limited adaptability of mediation strategies. Summary of the Invention
[0004] This invention provides a digital intelligent mediation method and system that integrates multimodal sentiment analysis, with the main purpose of solving the problems of low mediation efficiency and low adaptability of mediation strategies in the current dispute mediation process.
[0005] To achieve the above objectives, this invention provides a digital mediation method integrating multimodal sentiment analysis, comprising: An AI mediation system is acquired, comprising: a large-scale mediation model system, an AI mediation engine system, and an AI interaction system; The AI mediation engine system and the large-scale mediation model system in the AI mediation system are used to perform mediation matching, and the AI mediation robot template and AI mediation parameter set are obtained. Initial interaction is performed using an AI interaction system to obtain multimodal data; Based on the AI mediation engine system, sentiment analysis and annotation are performed on multimodal data to obtain a set of element annotations; Based on the feature annotation set and AI mediation parameter set, the AI mediation engine system is used to perform mediation template matching to obtain multimodal interaction information; Based on multimodal interaction information, an AI interaction system is used to mediate the interaction and obtain feedback information; If the feedback information does not meet the preset termination condition, the feedback information is treated as multimodal data, and the process of performing sentiment analysis and annotation on the multimodal data based on the AI mediation engine system is returned until the feedback information meets the termination condition. If the feedback information meets the termination conditions, a mediation result report is generated, and the AI mediation is completed.
[0006] Optionally, the process of using the AI mediation engine system and the large-scale mediation model system in the AI mediation system for mediation matching to obtain an AI mediation robot template and an AI mediation parameter set includes: The AI mediation initiation subsystem is determined from the AI mediation engine system; Obtain information on mediated cases; The AI mediation initiation subsystem is used to analyze the case characteristics of the mediation cases to obtain a case feature set; The AI mediation library subsystem was determined from the aforementioned large-scale mediation model system; Based on the characteristics of the mediation object and the AI mediation library subsystem, a mediation robot is matched to obtain an AI mediation robot template; Based on the case feature set, mediation parameters are set to obtain the AI mediation parameter set.
[0007] Optionally, the AI mediation initiation subsystem is used to perform case feature analysis on the mediation case information to obtain a case feature set, including: Based on the AI-based mediation initiation subsystem, the mediation case information is preprocessed to obtain standard case information; The standard case information is split to obtain multiple case terms; For each of the multiple case terms, perform the following operation: Lexical analysis is performed based on the vocabulary of the case to obtain lexical analysis results, wherein the lexical analysis results include: keywords or non-keywords; If the vocabulary analysis results are keywords, then the case vocabulary will be marked as reserved vocabulary; If the vocabulary analysis result is a non-keyword, then the case vocabulary will be marked as removed vocabulary; By summarizing the removed and retained words, a vocabulary tag set is obtained; Based on the aforementioned vocabulary tag set, vocabulary filtering is performed on the vocabulary of multiple cases to obtain an effective vocabulary set; Semantic features are extracted from the effective vocabulary set to obtain multiple semantic feature vectors; Based on the multiple semantic feature vectors and the AI mediation library subsystem, semantic feature matching is performed to obtain a case feature set.
[0008] Optionally, the step of setting mediation parameters based on the case feature set to obtain an AI mediation parameter set includes: Extract case type, cause of action, and mediation object characteristics from the case feature set; The complexity of the case is calculated based on the characteristics of the mediation subjects and the preset calculation standards. Based on the case type, cause of action, and complexity, the AI mediation library subsystem is used to configure basic parameters to obtain a set of basic parameters. Obtain the start time of mediation; The termination time is determined based on the mediation start time and the preset maximum mediation duration. Based on the aforementioned case types and causes of action, a set of risk keywords was determined; Based on the aforementioned set of basic parameters, mediation start time, termination time, risk keyword set, and preset human intervention conditions, an AI mediation parameter set is constructed.
[0009] Optionally, the initial interaction using the AI interaction system to obtain multimodal data includes: The AI digital human system was identified from the AI interaction system. Based on the AI mediation robot template, the AI digital human system is configured with robot settings to obtain a mediation digital human; The AI interaction subsystem was identified from the AI interaction system; Based on the case feature set, obtain the initial dialogue plan; The initial dialogue plan is input into the mediation digital human using the AI interaction subsystem, and the initial mediation interaction is performed using the mediation digital human to obtain multimodal data, which includes video data, audio data and text data.
[0010] Optionally, the AI-based mediation engine system performs sentiment analysis and annotation on multimodal data to obtain a feature annotation set, including: Based on the AI interaction subsystem, facial expression recognition is performed on the video data in the multimodal data to obtain video text elements; The AI interaction subsystem is used to perform speech feature analysis on audio data to obtain audio text elements. The AI interaction subsystem is used to perform keyword analysis on text data to obtain message text elements; The AI intelligent annotation subsystem was identified from the AI mediation engine system; The video text elements, audio text elements, and message text elements are transmitted to the AI intelligent annotation subsystem to obtain an object element set. Feature annotation is performed based on the AI mediation library subsystem and the object feature set to obtain the feature annotation set.
[0011] Optionally, the AI-based interaction subsystem performs facial expression recognition on the video data in the multimodal data to obtain video text elements, including: Based on the AI interaction subsystem, the video data in the multimodal data is processed by frame segmentation to obtain a video image sequence, wherein the video image sequence contains multiple video images; For each video image in the video image sequence, the following operation is performed: Target face detection is performed on the video image to obtain the face image region; Facial key point detection is performed based on the face image region to obtain the coordinates of multiple facial key points; Facial expression status is obtained by using the coordinates of multiple facial key points for expression recognition. The facial expression states are summarized to obtain a facial expression sequence; Extract the set of termination expressions from the facial expression sequence; Based on the set of terminating expressions, determine the object's final emotion; Based on the final emotion and facial expression sequence of the object, video text elements are constructed.
[0012] Optionally, the AI interaction subsystem is used to perform speech feature analysis on the audio data to obtain audio text elements, including: The audio data is segmented using a preset time interval to obtain multiple audio segments; For each of the plurality of audio segments, the following operation is performed: Audio features are extracted from the audio segment to obtain multiple Mel frequency cepstral coefficient feature vectors; Multiple acoustic probability distributions are calculated using multiple Mel frequency cepstral coefficient eigenvectors and a pre-constructed acoustic model; Phoneme decoding is performed on the multiple acoustic probability distributions to obtain the optimal phoneme state sequence; The optimal phoneme state sequence is transformed using the AI interaction subsystem to obtain audio-recognized text; The audio-recognized text is labeled with keywords to obtain an audio keyword set; Extract the fundamental frequency and energy characteristics of the audio segment; Based on the fundamental frequency characteristics and energy characteristics, the audio pitch characteristics are determined; Based on the aforementioned audio intonation features and audio keyword set, audio emotion features are constructed; By summarizing the aforementioned audio emotional features, audio text elements are obtained.
[0013] Optionally, the step of using an AI mediation engine system to perform mediation template matching based on the feature annotation set and AI mediation parameter set to obtain multimodal interaction information includes: The AI mediation dialogue subsystem was identified from the AI mediation engine system. Based on the element annotation set, the AI mediation dialogue subsystem is used to perform feature fusion to obtain the mediation feature vector; Based on the AI mediation library subsystem, similarity retrieval is performed using mediation feature vectors to obtain the target mediation dialogue template; Based on the AI mediation parameter set, obtain the mediation constraints; Historical interaction content is obtained from the multimodal data; Based on the mediation constraints, historical interaction content, and target mediation dialogue template, obtain multimodal interaction information.
[0014] To achieve the above objectives, the present invention also provides a digital intelligence mediation system integrating multimodal sentiment analysis, comprising: The mediation preparation module is used to acquire the AI mediation system, wherein the AI mediation system includes: a large-scale mediation model system, an AI mediation engine system, and an AI interaction system; The emotion labeling module is used to perform mediation matching using the AI mediation engine system and mediation big model system in the AI mediation system, to obtain AI mediation robot templates and AI mediation parameter sets, and to perform initial interaction using the AI interaction system to obtain multimodal data; The mediation interaction execution module is used to perform sentiment analysis and annotation on multimodal data based on the AI mediation engine system to obtain an element annotation set. Based on the element annotation set and the AI mediation parameter set, the AI mediation engine system is used to perform mediation template matching to obtain multimodal interaction information. The mediation process control module is used to conduct mediation interaction using an AI interaction system based on multimodal interaction information, obtain feedback information, and if the feedback information does not meet the preset termination conditions, the feedback information is treated as multimodal data and the process is returned to the above-mentioned step of performing sentiment analysis and annotation on the multimodal data based on the AI mediation engine system until the feedback information meets the termination conditions. If the feedback information meets the termination conditions, a mediation result report is generated, and the AI mediation is completed.
[0015] To address the above problems, the present invention also provides an electronic device, the electronic device comprising: A memory that stores at least one instruction; and a processor that executes the instructions stored in the memory to implement the digital mediation method for integrating multimodal sentiment analysis as described above.
[0016] To address the aforementioned issues, the present invention also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor in an electronic device to implement the aforementioned digital mediation method integrating multimodal sentiment analysis.
[0017] To address the problems described in the background art, this invention provides an AI mediation system. This AI mediation system comprises a large-scale mediation model system, an AI mediation engine system, and an AI interaction system. Through this step, the invention clarifies the technical architecture of the mediation task. It utilizes the AI mediation engine system and the large-scale mediation model system within the AI mediation system for mediation matching, obtaining an AI mediation robot template and an AI mediation parameter set. The AI interaction system is then used for initial interaction, yielding multimodal data. This invention uses an AI digital human within the AI interaction system to conduct initial interaction with the mediation subject, collecting multimodal data including video, audio, and text data. This provides data support for comprehensively capturing the real state of the mediation subject, enhancing the adaptability of the mediation strategy. The AI mediation engine system performs sentiment analysis and annotation on multimodal data to obtain a feature annotation set. Based on the feature annotation set and the AI mediation parameter set, the system performs mediation template matching to obtain multimodal interaction information. Based on this information, an AI interaction system is used for mediation interaction to obtain feedback information. If the feedback information does not meet a preset termination condition, it is treated as multimodal data, and the process returns to the previous steps of sentiment analysis and annotation on the multimodal data using the AI mediation engine system, until the feedback information meets the termination condition. If the feedback information does meet the termination condition, a mediation result report is generated, completing the AI mediation. This invention improves the efficiency of dispute mediation by using an intelligent mediation system. Therefore, this invention can solve the problems of low mediation efficiency and low adaptability of mediation strategies in current dispute mediation processes. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a digital mediation method integrating multimodal sentiment analysis provided in an embodiment of the present invention. Figure 2 This is a functional block diagram of a digital intelligence mediation system integrating multimodal emotion analysis provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device that implements the digital intelligence mediation method that integrates multimodal emotion analysis, according to an embodiment of the present invention.
[0019] Explanation of reference numerals in the attached figures: 10. Electronic device; 11. Processor; 12. Memory; 13. Bus.
[0020] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0022] This application provides a digital mediation method integrating multimodal sentiment analysis. The executing entity of the digital mediation method integrating multimodal sentiment analysis includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application embodiment: a server, a terminal, etc. In other words, the digital mediation method integrating multimodal sentiment analysis can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0023] Reference Figure 1 The diagram shown is a flowchart illustrating a digital mediation method integrating multimodal sentiment analysis according to an embodiment of the present invention. In this embodiment, the digital mediation method integrating multimodal sentiment analysis includes: S1. Obtain an AI mediation system, wherein the AI mediation system includes: a large-scale mediation model system, an AI mediation engine system, and an AI interaction system.
[0024] It should be noted that, in this embodiment, the dialogue and interaction targets in all mediation processes are the mediation objects, which refer to the parties accepting mediation in a dispute case. The AI mediation system refers to an integrated system used to automate or assist in the mediation of dispute cases, which includes a large-scale mediation model system, an AI mediation engine system, and an AI interaction system. The aforementioned mediation model system refers to the knowledge base and model training platform within the AI mediation system. It includes: an AI mediation library subsystem (explained in subsequent implementation steps), an AI mediation training subsystem, and an AI mediation modeling subsystem. The AI mediation modeling subsystem is the functional module responsible for constructing the algorithm model within the large mediation model. It employs a multi-core algorithm structure, integrating various professional domain algorithms (such as LR-Linear Regression, Log-LR, LDA-Linear Discriminant Analysis, Cost-sensitive learning, DAG, Adaboost, GBDT-Gradient Boosting Decision Tree, XGBoost-eXtreme Gradient Boosting, and AGNES-Agglomerative...). The AI mediation training subsystem refers to the functional module in the large-scale mediation model system that constructs AI mediator templates. It can train AI mediator templates (i.e., templates recording relevant data such as the dispute cases mediated, mediation scripts, facial expressions, and gestures) by calling data generated by mediators during dispute resolution stored in the AI mediation library subsystem and the mediation algorithm model generated by the AI mediation modeling subsystem. The AI mediation engine system refers to the execution center of the AI mediation system, responsible for the logical control and decision-making execution of the mediation process. Its specific functions include: calling AI mediator templates from the large-scale mediation model system and configuring AI parameters based on the characteristics of the dispute cases; analyzing multimodal information from the parties involved during the mediation process and generating mediation scripts and interaction strategies based on the analysis results. The AI interaction system refers to the execution module that enables information interaction between the AI mediation system and the mediation subject. It can collect video, audio and text data of the mediation subject, send the collected video, audio and text data to the AI mediation engine system, receive the mediation scripts and interaction strategies fed back by the AI mediation engine system, and output them to the mediation subject to complete the mediation dialogue process.
[0025] S2. Utilize the AI mediation engine system and mediation model system in the AI mediation system to perform mediation matching, and obtain the AI mediation robot template and AI mediation parameter set.
[0026] In detail, the AI mediation engine system and the large-scale mediation model system in the AI mediation system are used for mediation matching to obtain the AI mediation robot template and the AI mediation parameter set, including: The AI mediation initiation subsystem is determined from the AI mediation engine system; Obtain information on mediated cases; The AI mediation initiation subsystem is used to analyze the case characteristics of the mediation cases to obtain a case feature set; The AI mediation library subsystem was determined from the aforementioned large-scale mediation model system; Based on the characteristics of the mediation object and the AI mediation library subsystem, a mediation robot is matched to obtain an AI mediation robot template; Based on the case feature set, mediation parameters are set to obtain the AI mediation parameter set.
[0027] It should be noted that the AI mediation initiation subsystem refers to the module in the AI mediation engine system responsible for receiving information on mediation cases. It analyzes the case information to provide data for subsequently obtaining AI mediation robot templates and AI mediation parameter sets. The mediation case information refers to relevant information about the dispute case, including: case number, cause of action (such as loan disputes, contract disputes, neighborhood conflicts, etc.), and basic information of the parties involved (such as name, age, occupation, place of residence, etc.). The AI mediation database subsystem refers to the module in the large-scale mediation model system responsible for data storage. Its stored content includes: AI mediator templates generated by the AI mediation training subsystem, and historical dispute case information (i.e., relevant information from past dispute cases). The mediation robot matching refers to the process of selecting the mediator template that best matches the case feature set from the AI mediation database subsystem. Specifically, the process involves: converting the case feature set into vector form to obtain case feature vectors; calculating the similarity between these case feature vectors and the mediator feature vectors (i.e., the case feature vectors of the disputes the mediator is good at handling) that match the templates of each mediator in the AI mediation database subsystem; obtaining multiple case feature similarities; and selecting the mediator template with the highest similarity. The formula for calculating the similarity is as follows: ,in, Indicates the similarity of case characteristics. Represents the case feature vector. Represents the mediator's feature vector. This represents the dot product of the case feature vector and the mediator feature vector. The term "modulus" refers to the length of the vector. The method for converting the case feature set into vector form is existing technology and will not be elaborated upon here. The AI mediation robot template refers to the template of the mediator selected after matching with the mediation robot, which can be used for robot setup in subsequent embodiments.
[0028] Specifically, the AI mediation initiation subsystem is used to perform case feature analysis on mediation case information to obtain a case feature set, including: Based on the AI-based mediation initiation subsystem, the mediation case information is preprocessed to obtain standard case information; The standard case information is split to obtain multiple case terms; For each of the multiple case terms, perform the following operation: Lexical analysis is performed based on the vocabulary of the case to obtain lexical analysis results, wherein the lexical analysis results include: keywords or non-keywords; If the vocabulary analysis results are keywords, then the case vocabulary will be marked as reserved vocabulary; If the vocabulary analysis result is a non-keyword, then the case vocabulary will be marked as removed vocabulary; By summarizing the removed and retained words, a vocabulary tag set is obtained; Based on the aforementioned vocabulary tag set, vocabulary filtering is performed on the vocabulary of multiple cases to obtain an effective vocabulary set; Semantic features are extracted from the effective vocabulary set to obtain multiple semantic feature vectors; Based on the multiple semantic feature vectors and the AI mediation library subsystem, semantic feature matching is performed to obtain a case feature set.
[0029] It is understood that the data preprocessing refers to the process of standardizing the data format of mediation case information. For example, different formats of date information (such as "2025.12.01", "December 1, 2025", "12 / 01 / 2025") are uniformly converted into the standard "year-month-date" format, and different formats of case numbers (such as "2025-001", "001-2025") are uniformly standardized into the standard "year-serial number" encoding format. The standard case information refers to the mediation case information after data preprocessing. The information splitting refers to the operation of using a word segmentation algorithm (such as the forward maximum matching method) to split continuous standard text data into multiple independent words according to semantic logic. The case vocabulary refers to the independent words obtained after information splitting. The vocabulary analysis refers to the process of filtering case vocabulary based on the case keywords stored in the AI mediation database subsystem. The lexical analysis results refer to the results obtained after performing lexical analysis on the case vocabulary, including keywords and non-keywords. Keywords refer to case vocabulary that are considered case keywords and can represent the characteristics of the dispute, such as: case type terms (e.g., contract, loan, tort, etc.), case cause terms (e.g., repayment, compensation, arrears, etc.), dispute element terms (e.g., breach of contract, overdue payment, damage, etc.), and mediation subject characteristic terms (e.g., age, occupation, etc.). Non-keywords refer to case vocabulary that is not considered case keywords.
[0030] It should be understood that the "retained vocabulary" refers to the labeling of case words that are keywords in the vocabulary analysis results. These retained words represent the characteristics of the dispute case and need to be retained. The "removed vocabulary" refers to the labeling of case words that are not keywords in the vocabulary analysis results. The "vocabulary label set" refers to the collection of all case words and their corresponding labels forming a combined structure (combination form: case word - label). The "vocabulary filtering based on the vocabulary label set to obtain an effective vocabulary set" means: selecting case words marked as retained words from multiple case words; the set of these retained case words is the effective vocabulary set. The "semantic feature extraction" refers to the process of converting each effective word in the effective vocabulary set into a word vector using a pre-trained word embedding method (such as Word2Vec). The method of converting effective words in the text form into word vectors using pre-trained word embedding is existing technology and will not be elaborated here. The "semantic feature vector" refers to the word vector obtained after semantic feature extraction of the effective words. The semantic feature matching refers to grouping multiple semantic feature vectors using a clustering algorithm (such as K-means clustering). This involves automatically grouping the semantic feature vectors based on their distance in the semantic space, forming multiple semantic feature vector groups. The semantic feature vectors within each group are close to each other in the semantic space, meaning that the words they represent are semantically closely related and may jointly describe a certain aspect or sub-topic of the case, such as jointly describing the loan facts or jointly describing the consequences of harm. The method of grouping multiple semantic feature vectors using a clustering algorithm is existing technology and will not be elaborated here. Then, the similarity of the semantic feature vector groups with the feature vectors of various case features stored in the AI mediation database subsystem is calculated, and the case features with the highest similarity are selected. For example, if a semantic feature vector set for a dispute case is {loan, lending, repayment}, and the Word2Vec algorithm has a dimension of 5, after semantic feature extraction, three semantic feature vectors are obtained: Vloan = {0.85, 0.12, 0.78, 0.05, 0.91}, Voverdue = {0.72, 0.88, 0.15, 0.90, 0.23}, and Vrepayment = {0.80, 0.10, 0.75, 0.08, 0.93}. Taking the average of the three semantic feature vectors yields a combined feature vector of {0.79, 0.37, 0.56, 0.34, 0.69}. The similarity calculation between this combined feature vector and the feature vectors of various case features in the AI mediation database subsystem shows that the combined feature vector has the highest similarity to the feature vector of the loan dispute case. Therefore, the loan dispute is considered a case feature of the dispute case. The case feature set refers to the set of selected case features.
[0031] Furthermore, the process of setting mediation parameters based on the case feature set to obtain an AI mediation parameter set includes: Extract case type, cause of action, and mediation object characteristics from the case feature set; The complexity of the case is calculated based on the characteristics of the mediation subjects and the preset calculation standards. Based on the case type, cause of action, and complexity, the AI mediation library subsystem is used to configure basic parameters to obtain a set of basic parameters. Obtain the start time of mediation; The termination time is determined based on the mediation start time and the preset maximum mediation duration. Based on the aforementioned case types and causes of action, a set of risk keywords was determined; Based on the aforementioned set of basic parameters, mediation start time, termination time, risk keyword set, and preset human intervention conditions, an AI mediation parameter set is constructed.
[0032] It should be explained that the "case type" refers to the type of dispute, such as loan disputes, contract disputes, labor disputes, neighborhood conflicts, etc. The case type determines the applicable legal provisions and mediation strategies (i.e., the purpose of mediation, such as resolving conflicts, repaying loans, etc.). The "case cause of action" refers to the description of the points of contention in the dispute, such as overdue repayment, breach of contract compensation, wage arrears, etc. The "characteristics of the mediation subjects" refers to the set of relevant descriptions of the mediation subjects participating in the mediation, including: the age, occupation, education level, past mediation records (whether there is previous experience in mediation), number of legal relationships, number of mediation subjects, etc. The "calculation standard" refers to a pre-set standard for calculating the complexity of the case. Optionally, in this embodiment, the number of legal relationships and the number of mediation subjects in the dispute are used as the calculation standard. The number of legal relationships refers to the total number of different types of legal relationships involved in the same dispute case. For example, a case with only one legal relationship (such as a loan relationship) has a legal relationship count of 1; a case with two legal relationships (such as a loan relationship and a guarantee relationship) has a legal relationship count of 2; and a case with three or more legal relationships (such as a loan relationship + a guarantee relationship + a mortgage relationship) has a legal relationship count of 3. The number of mediation participants is a quantitative indicator determined by the number of participants in the mediation. For example, if the number of participants is 2 or less, the number of participants is set to 2; if the number of participants is greater than or equal to 3 and less than or equal to 5, the number of participants is set to 5; and if the number of participants is greater than 5, the number of participants is set to 10. The case complexity is a quantitative indicator representing the degree of complexity of the dispute case after evaluation according to calculation standards. Calculating the case complexity based on the characteristics of the mediation participants and the preset calculation standards means: calculating the case complexity according to the calculation standards using a calculation formula, which is: ,in, Indicates the complexity of the case. Indicates the number of legal relationships. This indicates the number of parties involved in the mediation. A weighting coefficient representing the number of pre-defined legal relationships. This represents the weighting coefficient for the pre-defined number of mediation targets. The process of configuring basic parameters using the AI mediation database subsystem based on the case type, cause of action, and complexity to obtain the basic parameter set involves: determining the mediation strategy type based on the case type and cause of action (e.g., if the case type and cause of action are loan disputes and overdue payments, the determined mediation strategy type is: loan disputes - overdue payments); searching for a suitable mediation strategy template in the AI mediation database subsystem based on the determined mediation strategy type and case complexity (the AI mediation database subsystem stores historical mediation strategies for various disputes; using the mediation strategy type and case complexity as standards, similarity matching is performed in the AI mediation database subsystem, and the mediation strategy with the highest similarity is the suitable mediation strategy template; the similarity matching method is the same as the similarity calculation method in semantic feature matching, and will not be elaborated here); and configuring the basic parameters of the mediation digital human (which will be explained in subsequent implementation steps) based on the mediation strategy template (such as single speaking time limit, pause interval, legal knowledge involved in the dispute, etc.). The set of basic parameters of the mediation robot is the basic parameter set. In addition, the basic parameter set also includes the number of mediation robots, which is set according to the number of mediation objects, with one mediation robot corresponding to each mediation object.
[0033] It is understood that the mediation start time refers to the specific time at which mediation of the dispute case begins. Optionally, the mediation start time can be set according to the requirements of the parties involved in the mediation. For example, the parties involved in the mediation may request that the mediation start time be 3 PM on January 17, 2026. The mediation duration limit refers to the maximum duration of mediation that is preset by the user. When the duration reaches the mediation duration limit, the mediation process will automatically terminate. Optionally, the mediation duration limit can be set according to the user's needs. For example, the user may set the mediation duration limit to 3 hours. The termination time refers to the latest end time of the mediation process calculated based on the mediation start time and the mediation duration limit. The determination of the risk keyword set based on the case type and the cause of action refers to: determining the specific type of the case based on the case type and the cause of action, and identifying the risk words that may cause conflict, misunderstanding, or mediation failure corresponding to the specific type of case in the AI mediation library subsystem. The set of risk words constitutes the risk parameter set. The aforementioned human intervention conditions refer to pre-set triggering conditions that require human mediators to intervene in the mediation process. Examples include: the mediator actively requests intervention, risk keywords appearing more frequently in the dialogue than a threshold (e.g., 10 times), and the termination time being reached. Constructing the AI mediation parameter set based on the aforementioned basic parameter set, mediation start time, termination time, risk keyword set, and pre-set human intervention conditions means: summarizing the basic parameter set, mediation start time, termination time, risk keyword set, and human intervention conditions into a single set; this set constitutes the AI mediation parameter set.
[0034] S3. Use the AI interaction system to perform initial interaction and obtain multimodal data.
[0035] In detail, the initial interaction using the AI interaction system to obtain multimodal data includes: The AI digital human system was identified from the AI interaction system. Based on the AI mediation robot template, the AI digital human system is configured with robot settings to obtain a mediation digital human; The AI interaction subsystem was identified from the AI interaction system; Based on the case feature set, obtain the initial dialogue plan; The initial dialogue plan is input into the mediation digital human using the AI interaction subsystem, and the initial mediation interaction is performed using the mediation digital human to obtain multimodal data, which includes video data, audio data and text data.
[0036] It should be understood that the AI digital human system refers to the module in the AI interaction system responsible for generating and driving the virtual mediator image, which can generate a digital character based on input text or instructions. The robot setup refers to the process by which the AI digital human system configures the parameters of the digital character according to the AI mediation robot template. The configured parameters include: appearance (gender, age, clothing, etc.), voice timbre, tone of voice, etc. The mediation digital human refers to the digital character generated by the AI digital human system based on the AI mediation robot template. The AI interaction subsystem refers to the system in the AI interaction system responsible for collecting the interaction data of the mediation subjects during the mediation process and scheduling the AI digital human subsystem. It can receive multimodal interaction information output by the AI mediation dialogue subsystem (which will be explained in the steps of subsequent embodiments) and input it into the AI digital human subsystem to complete the dialogue with the mediation subjects. The initial dialogue scheme refers to the first round of dialogue content (i.e., opening remarks) retrieved from the AI mediation database subsystem based on the case feature set before formally starting the dialogue with the mediation subject. The retrieval method is as follows: according to the case type and case cause in the case feature set, select the opening dialogue template with the same case type and case cause in the AI mediation database subsystem, and replace and fill the template information in the opening dialogue template with the specific information of the dispute case (such as the name of the mediation subject, the time of the dispute, etc.). For example, if the opening dialogue template is: "Hello, this is my party," and the name of the mediation subject is Li Si, then replace "my party" with "Li Si."
[0037] It should be noted that the initial mediation interaction refers to the dialogue initiated by the digital mediation human and conducted with the mediation subject; the content of this dialogue constitutes the initial dialogue plan. The multimodal data refers to the information collected by the AI interaction subsystem from the mediation subject after the initial mediation interaction, including video data, audio data, and text data. The video data refers to the video footage of the mediation subject during the interaction, collected by the AI interaction subsystem. The audio data refers to the voice data collected by the AI interaction subsystem, including the voice content, tone, speed, and volume of the mediation subject during the interaction. The text data refers to the text information entered by the mediation subject during the interaction, collected by the AI interaction subsystem. By collecting multimodal data, the emotions and mediation needs of the mediation subject can be determined, providing data support for subsequent mediation processes.
[0038] S4. Based on the AI mediation engine system, perform sentiment analysis and annotation on multimodal data to obtain a set of element annotations.
[0039] Specifically, the AI-based mediation engine system performs sentiment analysis and annotation on multimodal data to obtain a feature annotation set, including: Based on the AI interaction subsystem, facial expression recognition is performed on the video data in the multimodal data to obtain video text elements; The AI interaction subsystem is used to perform speech feature analysis on audio data to obtain audio text elements. The AI interaction subsystem is used to perform keyword analysis on text data to obtain message text elements; The AI intelligent annotation subsystem was identified from the AI mediation engine system; The video text elements, audio text elements, and message text elements are transmitted to the AI intelligent annotation subsystem to obtain an object element set. Feature annotation is performed based on the AI mediation library subsystem and the object feature set to obtain the feature annotation set.
[0040] It is understood that the keyword analysis refers to the process of extracting object keywords related to the mediation subject from text data. The method for extracting object keywords related to the mediation subject from text data is consistent with the method for lexical analysis of case vocabulary, and will not be elaborated here. The object keywords refer to words or phrases directly related to the mediation subject or their dispute case, extracted from the text data of the mediation subject during the keyword analysis process. Examples include: the mediation subject's identity (e.g., self-employed individual, migrant worker, etc.), the mediation subject's living conditions (e.g., income, daily consumption, etc.), and the mediation subject's emotions (e.g., anxiety, depression, etc.). The message text elements refer to the set of all object keywords obtained after keyword analysis of the text data. The AI intelligent annotation subsystem refers to the module in the AI mediation engine system responsible for receiving video text elements, audio text elements, and message text elements analyzed by the AI interaction subsystem, and annotating the mediation subject in multiple dimensions (e.g., living conditions, personality, repayment ability, willingness to mediate, etc.) based on these three elements. The object element set refers to the set formed by the AI intelligent annotation subsystem after summarizing the received video text elements, audio text elements, and message text elements. The element annotation refers to the process of annotating mediation subjects based on the AI mediation database subsystem and the subject element set. The specific steps are as follows: Video text elements, audio text elements, and message text elements are combined and analyzed; the video text elements, audio text elements, and message text elements are time-aligned; the AI mediation database subsystem is retrieved for each annotation dimension (such as living conditions, personality, repayment ability, and willingness to mediate); for each annotation dimension, video text elements, audio text elements, and message text elements are simultaneously retrieved; if the keywords matched by the content of the video text elements, audio text elements, and message text elements at the same time are consistent, then the keyword is valid. The matching method is the same as the method for lexical analysis of case vocabulary, and will not be elaborated here. When keywords in each dimension (such as living conditions, personality, repayment ability, and willingness to mediate) of the AI mediation database subsystem are matched among the three elements, the mediation subject is annotated according to the matching results (i.e., the keyword is used as the label for the mediation subject in that dimension). The element annotation set refers to the set of labels obtained after element annotation.
[0041] Furthermore, the AI-based interaction subsystem performs facial expression recognition on the video data in the multimodal data to obtain video text elements, including: Based on the AI interaction subsystem, the video data in the multimodal data is processed by frame segmentation to obtain a video image sequence, wherein the video image sequence contains multiple video images; For each video image in the video image sequence, the following operation is performed: Target face detection is performed on the video image to obtain the face image region; Facial key point detection is performed based on the face image region to obtain the coordinates of multiple facial key points; Facial expression status is obtained by using the coordinates of multiple facial key points for expression recognition. The facial expression states are summarized to obtain a facial expression sequence; Extract the set of termination expressions from the facial expression sequence; Based on the set of terminating expressions, determine the object's final emotion; Based on the final emotion and facial expression sequence of the object, video text elements are constructed.
[0042] It should be understood that the frame segmentation process refers to the process of dividing video data into multiple independent images according to a pre-set time interval (e.g., 0.02 seconds). The video image refers to the image obtained after frame segmentation. The video image sequence refers to the sequence formed by arranging all video images in chronological order. The target face detection refers to extracting the facial image of the mediation subject using a face detection model. The method for extracting the facial image of the mediation subject using a face detection model is existing technology and will not be described in detail here. Optionally, MTCNN can be used as the face detection model. The face image region refers to the facial image of the mediation subject extracted using the face detection model. The facial key point detection refers to the process of locating the facial features and points with significant features on the facial contour (e.g., corners of the eyes, tip of the nose, corners of the mouth, chin contour points, etc.) using a key point localization model. The method for locating the facial features and points with significant features on the facial contour using a key point localization algorithm is existing technology and will not be described in detail here. Optionally, the RetinaFace algorithm can be used as the key point localization algorithm. The facial keypoint coordinates refer to the two-dimensional pixel coordinates of facial keypoints (i.e., points with significant features on the facial contours) within the face image region after facial keypoint detection. These coordinates can be represented by (x, y), where x represents the pixel position of the facial keypoint along the horizontal (width) direction within the face image region, and y represents the pixel position of the facial keypoint along the vertical (height) direction within the face image region. It should be noted that the origin of the coordinates is the top-left corner of the face image region.
[0043] It should be explained that the aforementioned expression recognition refers to the process of using a vector composed of multiple facial key point coordinates (the method of converting multiple facial key point coordinates into vector form is existing technology and will not be elaborated here) as input, and then using an expression classification model (such as a model built on the Transformer architecture, which uses the Transformer's self-attention mechanism to capture the correlation features and spatial dependencies between facial key point coordinates, thereby achieving expression classification) to match the facial state corresponding to a face image to the corresponding expression category (such as peaceful, anxious, resistant, excited, irritable). The method of matching the facial state corresponding to a face image to the corresponding expression category using an expression classification model is existing technology and will not be elaborated here. The facial expression state refers to the expression category obtained after expression recognition. The facial expression sequence refers to the sequence formed by arranging all facial expression states in chronological order. The terminating expression set refers to the set formed by the last part of facial expression states extracted from the facial expression sequence; optionally, the number of the last part can be set to 5. The final emotion of the object refers to the emotion of the mediation object after the initial mediation interaction, as determined by the set of terminating facial expressions. Optionally, it can be determined based on the facial expression state with the most occurrences in the set of terminating facial expressions. If the number of facial expression states is the same, then the last facial expression state in time is taken as the object's final emotion. The video text elements refer to the set formed by summarizing the object's final emotion and the sequence of facial expressions.
[0044] In detail, the AI interaction subsystem is used to perform speech feature analysis on the audio data to obtain audio text elements, including: The audio data is segmented using a preset time interval to obtain multiple audio segments; For each of the plurality of audio segments, the following operation is performed: Audio features are extracted from the audio segment to obtain multiple Mel frequency cepstral coefficient feature vectors; Multiple acoustic probability distributions are calculated using multiple Mel frequency cepstral coefficient eigenvectors and a pre-constructed acoustic model; Phoneme decoding is performed on the multiple acoustic probability distributions to obtain the optimal phoneme state sequence; The optimal phoneme state sequence is transformed using the AI interaction subsystem to obtain audio-recognized text; The audio-recognized text is labeled with keywords to obtain an audio keyword set; Extract the fundamental frequency and energy characteristics of the audio segment; Based on the fundamental frequency characteristics and energy characteristics, the audio pitch characteristics are determined; Based on the aforementioned audio intonation features and audio keyword set, audio emotion features are constructed; By summarizing the aforementioned audio emotional features, audio text elements are obtained.
[0045] It is understood that segmenting audio data using a preset time interval to obtain multiple audio segments means: dividing the audio data according to a fixed duration to obtain multiple audio frames. An audio segment is an audio fragment consisting of N consecutive audio frames. Optionally, the duration is 20ms, the frame shift of the audio fragment is 10ms, the number of N is 5, and the segment shift is 2 frames. For example, a 1s effective speech segment, after segmentation, yields 99 audio frames, constituting 48 audio segments. Audio feature extraction refers to the process of extracting quantized parameters that characterize the acoustic properties of an audio segment using an audio feature extraction algorithm and converting these quantized parameters into vectors. Optionally, the Mel-frequency cepstral coefficient method is used as the audio feature extraction algorithm. The Mel-frequency cepstral coefficient feature vector refers to the vector obtained after audio feature extraction. The acoustic model refers to a model that can output the acoustic probability distribution of the input Mel-frequency cepstral coefficient feature vector. Optionally, a hidden Markov model is used as the acoustic model. The multiple acoustic probability distributions refer to the set of probability values for all possible phonemes corresponding to a single audio segment. Each probability value is in the interval [0, 1], and the sum of the probability values of all phonemes is 1. For example, after the multiple Mel-frequency cepstral coefficient feature vectors of an audio segment are calculated by the acoustic model, the probability corresponding to the phoneme b is 0.62, the probability corresponding to the phoneme a is 0.32, and the sum of the probabilities corresponding to other phonemes is 0.06. This set of probability values constitutes the acoustic probability distribution of the speech frame. The phoneme decoding refers to the process of using a preset decoding algorithm to select the phoneme combination with the highest probability value from multiple acoustic probability distributions. Optionally, the Viterbi algorithm can be used as the decoding algorithm. The method of using the Viterbi algorithm to select the phoneme combination with the highest probability value from multiple acoustic probability distributions is existing technology and will not be described in detail here. The optimal phoneme state sequence refers to the sequence formed by arranging the phoneme combinations with the highest probability values obtained after phoneme decoding in chronological order.
[0046] It should be noted that the transformation of the optimal phoneme state sequence using the AI interaction subsystem to obtain audio recognition text means that the AI interaction subsystem maps the phoneme combinations in the optimal phoneme state sequence to the corresponding text based on a pre-defined correspondence between phonemes (or phoneme combinations) and corresponding Chinese characters (or words). The resulting text is the audio recognition text. Keyword annotation refers to the process of extracting object keywords related to the mediation object from the audio recognition text. The method for extracting object keywords related to the mediation object from the audio recognition text is consistent with the method of keyword analysis of text data using the AI interaction subsystem, and will not be elaborated further here. The audio keyword set refers to the set of object keywords extracted from the audio recognition text. The fundamental frequency feature refers to the numerical feature extracted by the audio signal processing algorithm that characterizes the change of speech pitch over time. The fundamental frequency feature reflects the pitch change of the mediation object and can be used to judge the pitch, emotional excitement, etc. Optionally, a fundamental tone detection algorithm can be used as the audio signal processing algorithm. The energy feature refers to the numerical feature representing the change of speech amplitude over time, extracted by an energy feature algorithm. The energy feature reflects the loudness of the subject's speech and can be used to judge emotional intensity, tone strength, etc. Optionally, a short-time energy calculation algorithm is used as the energy feature algorithm. The audio intonation feature refers to a descriptive label derived from the changing trends of the fundamental frequency feature and energy feature through pre-defined quantitative judgment rules. For example, when both the fundamental frequency feature and energy feature show an upward trend (fundamental frequency change rate...), the energy feature is... 15Hz / s, and energy fluctuation amplitude 60), labeled as emotionally agitated, when both fundamental frequency characteristics and energy characteristics show a downward trend (fundamental frequency change rate). -10Hz / s, and energy fluctuation amplitude 25) The label is "depressed mood". If it does not fall into either of the above two categories, the label will be confirmed as a pre-defined label (e.g., "peaceful mood"). The fundamental frequency change rate refers to a quantitative indicator representing the relative rate of change of pitch, and the calculation formula is: ,in, Indicates the rate of change of the fundamental frequency. The fundamental frequency that indicates the starting frame of an audio segment. The fundamental frequency of the last frame of an audio segment. This indicates the duration of the audio segment. The energy fluctuation amplitude refers to a quantitative indicator representing the degree of change in speech amplitude, calculated using the following formula: ,in, Indicates the amplitude of energy fluctuations. Indicates the first The short-time energy value of the frame, This represents the average short-time energy value of all audio frames within an audio segment. This indicates the number of audio frames within an audio segment. The audio sentiment features refer to the set formed by combining audio intonation features with the set of audio keywords. The audio text elements refer to the set formed by combining all audio sentiment features.
[0047] S5. Based on the feature annotation set and AI mediation parameter set, the AI mediation engine system is used to perform mediation template matching to obtain multimodal interaction information.
[0048] Specifically, based on the feature annotation set and the AI mediation parameter set, the AI mediation engine system performs mediation template matching to obtain multimodal interaction information, including: The AI mediation dialogue subsystem was identified from the AI mediation engine system. Based on the element annotation set, the AI mediation dialogue subsystem is used to perform feature fusion to obtain the mediation feature vector; Based on the AI mediation library subsystem, similarity retrieval is performed using mediation feature vectors to obtain the target mediation dialogue template; Based on the AI mediation parameter set, obtain the mediation constraints; Historical interaction content is obtained from the multimodal data; Based on the mediation constraints, historical interaction content, and target mediation dialogue template, obtain multimodal interaction information.
[0049] It should be explained that the AI mediation dialogue subsystem refers to the module in the AI mediation engine system used to generate dialogue content with the mediation subject. The process of using the AI mediation dialogue subsystem to perform feature fusion based on the element annotation set to obtain the mediation feature vector means that the AI mediation dialogue subsystem extracts various tags from the element annotation set and transforms these tags into a unified vector using feature encoding technology (such as one-hot encoding). This vector is the mediation feature vector. The method of transforming various tags into a unified vector using feature encoding technology is existing technology and will not be elaborated here. The similarity retrieval refers to the process of using the mediation feature vector as input and calculating the similarity with the feature vectors of multiple mediation dialogue templates (the mediation dialogue templates refer to templates of dialogue with the mediation subject) stored in the AI mediation library subsystem, and selecting the mediation dialogue template with the highest similarity. The method of calculating the similarity with the feature vectors of multiple mediation dialogue templates stored in the AI mediation library subsystem is consistent with the method of similarity calculation in semantic feature matching and will not be elaborated here. The target mediation dialogue template refers to the selected mediation dialogue template with the highest similarity. The mediation constraints refer to the restrictions determined from the AI mediation parameter set that must be followed in this mediation process, including but not limited to: duration constraints (i.e., the mediation duration cannot exceed the upper limit), risk word constraints (i.e., a set of risky keywords to be avoided), and human intervention constraints (i.e., conditions for human intervention to be avoided). These constraints ensure the legality and compliance of the mediation process and the safety of the mediation subjects. The historical interaction content refers to all information generated during previous interactions between the mediation digital human and the mediation subjects, obtained from multimodal data, including video data, audio data, text data, and their corresponding analysis results (such as emotional state, keywords, tone features, etc.), used to maintain the coherence and contextual consistency of the dialogue.The process of obtaining multimodal interaction information based on the mediation constraints, historical interaction content, and target mediation dialogue template refers to: using the target mediation dialogue template as a framework, replacing and filling the template information of the target mediation dialogue template with specific information from the historical interaction content (i.e., the general script in the target mediation dialogue template). The method of replacing and filling the template information of the target mediation dialogue template with specific information from the historical interaction content is the same as the method of replacing and filling the template information in the opening dialogue template with specific information from the dispute case. For example, if the specific information extracted from the historical interaction content is "Mediation subject name: Zhang San, Dispute type: Loan dispute, Zhang San's claim: Request Li Si to repay the loan of 20,000 yuan", then the template information can be: [Party name], about you and [Opponent's title]. The original text discusses a loan dispute involving RMB 20,000 between a person named Zhang San and another person named Li Si. It then describes the process of revising the dialogue based on the mediation constraints. The revision steps involve analyzing the dialogue using a sensitive word filtering algorithm (such as the Aho-Corasick algorithm) to detect and replace risky words (the latter being risky keywords in the mediation constraints; this method is existing technology and will not be elaborated upon here). Finally, it describes retrieving corresponding dialogue parameters (voice, facial expressions, etc.) from the target mediation dialogue template based on the revised dialogue. The resulting set of revised dialogue content and parameters constitutes the multimodal interaction information.
[0050] S6. Based on multimodal interaction information, use an AI interaction system to conduct mediation and interaction, and obtain feedback information.
[0051] It should be explained that the so-called mediation interaction based on multimodal interaction information and obtaining feedback information through an AI interaction system means that the AI interaction subsystem in the AI interaction system receives multimodal information from the AI mediation dialogue subsystem and sends the multimodal information to the AI digital human system. The AI digital human system drives the AI digital human to communicate and dialogue with the mediation object based on the dialogue content and dialogue parameters in the multimodal information. The response information of the mediation object to the AI digital human (i.e., the information collected from the mediation object after interacting with the AI digital human, including video data, audio data and text data) is the feedback information.
[0052] S7. If the feedback information does not meet the preset termination conditions, the feedback information is treated as multimodal data, and the process of performing sentiment analysis and annotation on the multimodal data based on the AI mediation engine system is returned until the feedback information meets the termination conditions. If the feedback information meets the termination conditions, a mediation result report is generated, and the AI mediation is completed.
[0053] It should be understood that the termination conditions refer to pre-set criteria for determining whether the mediation process needs to be ended, specifically the following situations: 1. The feedback information indicates acceptance of the mediation proposal. 2. The feedback information indicates rejection of the mediation proposal and no longer accepting mediation. 3. The termination time is reached. When the termination conditions are met, the mediation can be considered complete.
[0054] It should be noted that the AI mediation engine system has an AI intelligent monitoring subsystem (i.e., the AI mediation engine system is responsible for monitoring the mediation process). When the AI intelligent monitoring subsystem detects an abnormal situation, it will terminate the mediation or transfer it to manual mediation. The abnormal situation includes the following: 1. Multiple occurrences of risky keywords (e.g., more than 5 times) in the feedback information. 2. The mediation subject requests to transfer to manual mediation. 3. The audio emotional characteristics of the mediation subject's multimodal data for multiple consecutive rounds (e.g., 3 rounds or more) and the subject's final emotional state are all negative emotional characteristics (e.g., anxiety, agitation, irritability, etc.).
[0055] It is understood that the mediation result report refers to the collection of all data generated during the mediation process (such as multimodal data, element annotation sets, feedback information, etc.), which will be sent to the AI mediation library subsystem for storage, so as to facilitate subsequent improvement of the AI mediation system based on the data.
[0056] To address the problems described in the background art, this invention provides an AI mediation system. This AI mediation system comprises a large-scale mediation model system, an AI mediation engine system, and an AI interaction system. Through this step, the invention clarifies the technical architecture of the mediation task. It utilizes the AI mediation engine system and the large-scale mediation model system within the AI mediation system for mediation matching, obtaining an AI mediation robot template and an AI mediation parameter set. The AI interaction system is then used for initial interaction, yielding multimodal data. This invention uses an AI digital human within the AI interaction system to conduct initial interaction with the mediation subject, collecting multimodal data including video, audio, and text data. This provides data support for comprehensively capturing the real state of the mediation subject, enhancing the adaptability of the mediation strategy. The AI mediation engine system performs sentiment analysis and annotation on multimodal data to obtain a feature annotation set. Based on the feature annotation set and the AI mediation parameter set, the system performs mediation template matching to obtain multimodal interaction information. Based on this information, an AI interaction system is used for mediation interaction to obtain feedback information. If the feedback information does not meet a preset termination condition, it is treated as multimodal data, and the process returns to the previous steps of sentiment analysis and annotation on the multimodal data using the AI mediation engine system, until the feedback information meets the termination condition. If the feedback information does meet the termination condition, a mediation result report is generated, completing the AI mediation. This invention improves the efficiency of dispute mediation by using an intelligent mediation system. Therefore, this invention can solve the problems of low mediation efficiency and low adaptability of mediation strategies in current dispute mediation processes.
[0057] like Figure 2 The diagram shown is a functional block diagram of a digital intelligence mediation system that integrates multimodal emotion analysis, provided in an embodiment of the present invention.
[0058] The digital mediation system 100 integrating multimodal sentiment analysis described in this invention can be installed in an electronic device. Depending on the functions implemented, the digital mediation system 100 may include a mediation preparation module 101, a sentiment labeling module 102, a mediation interaction execution module 103, and a mediation process control module 104. The module described in this invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0059] The mediation preparation module 101 is used to acquire an AI mediation system, wherein the AI mediation system includes: a large-scale mediation model system, an AI mediation engine system, and an AI interaction system; The emotion labeling module 102 is used to perform mediation matching using the AI mediation engine system and the mediation big model system in the AI mediation system, to obtain the AI mediation robot template and the AI mediation parameter set, and to perform initial interaction using the AI interaction system to obtain multimodal data. The mediation interaction execution module 103 is used to perform sentiment analysis and annotation on multimodal data based on the AI mediation engine system to obtain an element annotation set. Based on the element annotation set and the AI mediation parameter set, the AI mediation engine system is used to perform mediation template matching to obtain multimodal interaction information. The mediation process control module 104 is used to conduct mediation interaction using an AI interaction system based on multimodal interaction information to obtain feedback information. If the feedback information does not meet the preset termination conditions, the feedback information is treated as multimodal data and the process is returned to the above-mentioned step of performing sentiment analysis and annotation on the multimodal data based on the AI mediation engine system until the feedback information meets the termination conditions. If the feedback information meets the termination conditions, a mediation result report is generated to complete the AI mediation.
[0060] In detail, the modules in the digital mediation system 100 that integrates multimodal emotion analysis described in this embodiment of the invention employ the same methods as described above when in use. Figure 1 The method uses the same technical means as the digital intelligence mediation method that integrates multimodal sentiment analysis described in the article, and can produce the same technical effect, so it will not be elaborated here.
[0061] like Figure 3 The diagram shown is a structural schematic of an electronic device that implements a digital intelligence mediation method that integrates multimodal emotion analysis, according to an embodiment of the present invention.
[0062] The electronic device 1 may include a processor 10, a memory 11 and a bus 12, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a digital intelligence mediation method program that integrates multimodal emotion analysis.
[0063] The memory 11 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as a portable hard drive. In other embodiments, the memory 11 can be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 1. Furthermore, the memory 11 includes both internal storage units and external storage devices of the electronic device 1. The memory 11 can be used not only to store application software and various types of data installed on the electronic device 1, such as the code of a digital mediation method program integrating multimodal sentiment analysis, but also to temporarily store data that has been output or will be output.
[0064] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., a digital mediation method program integrating multimodal emotion analysis) and calls data stored in the memory 11 to perform various functions of the electronic device 1 and process data.
[0065] The bus 12 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is configured to realize the connection and communication between the memory 11 and at least one processor 10, etc.
[0066] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0067] For example, although not shown, the electronic device 1 may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0068] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the electronic device 1 and other electronic devices.
[0069] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device 1 and to display a visual user interface.
[0070] The digital mediation method program that integrates multimodal emotion analysis, stored in the memory 11 of the electronic device 1, is a combination of multiple instructions. When run in the processor 10, it can achieve the following: An AI mediation system is acquired, comprising: a large-scale mediation model system, an AI mediation engine system, and an AI interaction system; The AI mediation engine system and the large-scale mediation model system in the AI mediation system are used to perform mediation matching, and the AI mediation robot template and AI mediation parameter set are obtained. Initial interaction is performed using an AI interaction system to obtain multimodal data; Based on the AI mediation engine system, sentiment analysis and annotation are performed on multimodal data to obtain a set of element annotations; Based on the feature annotation set and AI mediation parameter set, the AI mediation engine system is used to perform mediation template matching to obtain multimodal interaction information; Based on multimodal interaction information, an AI interaction system is used to mediate the interaction and obtain feedback information; If the feedback information does not meet the preset termination condition, the feedback information is treated as multimodal data, and the process of performing sentiment analysis and annotation on the multimodal data based on the AI mediation engine system is returned until the feedback information meets the termination condition. If the feedback information meets the termination conditions, a mediation result report is generated, and the AI mediation is completed.
[0071] Specifically, the processor 10's implementation method for the above instructions can be found in [reference needed]. Figures 1 to 3 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0072] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0073] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following: An AI mediation system is acquired, comprising: a large-scale mediation model system, an AI mediation engine system, and an AI interaction system; The AI mediation engine system and the large-scale mediation model system in the AI mediation system are used to perform mediation matching, and the AI mediation robot template and AI mediation parameter set are obtained. Initial interaction is performed using an AI interaction system to obtain multimodal data; Based on the AI mediation engine system, sentiment analysis and annotation are performed on multimodal data to obtain a set of element annotations; Based on the feature annotation set and AI mediation parameter set, the AI mediation engine system is used to perform mediation template matching to obtain multimodal interaction information; Based on multimodal interaction information, an AI interaction system is used to mediate the interaction and obtain feedback information; If the feedback information does not meet the preset termination condition, the feedback information is treated as multimodal data, and the process of performing sentiment analysis and annotation on the multimodal data based on the AI mediation engine system is returned until the feedback information meets the termination condition. If the feedback information meets the termination conditions, a mediation result report is generated, and the AI mediation is completed.
[0074] In the embodiments provided by this invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and actual implementations may have other classification methods.
[0075] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0076] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0077] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A digital intelligence mediation method integrating multimodal sentiment analysis, characterized in that, The method includes: An AI mediation system is acquired, comprising: a large-scale mediation model system, an AI mediation engine system, and an AI interaction system; The AI mediation engine system and the large-scale mediation model system in the AI mediation system are used to perform mediation matching, and the AI mediation robot template and AI mediation parameter set are obtained. Initial interaction is performed using an AI interaction system to obtain multimodal data; Based on the AI mediation engine system, sentiment analysis and annotation are performed on multimodal data to obtain a set of element annotations; Based on the feature annotation set and AI mediation parameter set, the AI mediation engine system is used to perform mediation template matching to obtain multimodal interaction information; Based on multimodal interaction information, an AI interaction system is used to mediate the interaction and obtain feedback information; If the feedback information does not meet the preset termination condition, the feedback information is treated as multimodal data, and the process of performing sentiment analysis and annotation on the multimodal data based on the AI mediation engine system is returned until the feedback information meets the termination condition. If the feedback information meets the termination conditions, a mediation result report is generated, and the AI mediation is completed.
2. The digital mediation method integrating multimodal sentiment analysis as described in claim 1, characterized in that, The AI mediation system utilizes its AI mediation engine and large-scale mediation model to perform mediation matching, resulting in an AI mediation robot template and a set of AI mediation parameters, including: The AI mediation initiation subsystem is determined from the AI mediation engine system; Obtain information on mediated cases; The AI mediation initiation subsystem is used to analyze the case characteristics of the mediation cases to obtain a case feature set; The AI mediation library subsystem was determined from the aforementioned large-scale mediation model system; Based on the characteristics of the mediation object and the AI mediation library subsystem, a mediation robot is matched to obtain an AI mediation robot template; Based on the case feature set, mediation parameters are set to obtain the AI mediation parameter set.
3. The digital mediation method integrating multimodal sentiment analysis as described in claim 2, characterized in that, The AI mediation initiation subsystem is used to analyze the case characteristics of the mediation cases to obtain a case feature set, including: Based on the AI-based mediation initiation subsystem, the mediation case information is preprocessed to obtain standard case information; The standard case information is split to obtain multiple case terms; For each of the multiple case terms, perform the following operation: Lexical analysis is performed based on the vocabulary of the case to obtain lexical analysis results, wherein the lexical analysis results include: keywords or non-keywords; If the vocabulary analysis results are keywords, then the case vocabulary will be marked as reserved vocabulary; If the vocabulary analysis result is a non-keyword, then the case vocabulary will be marked as removed vocabulary; By summarizing the removed and retained words, a vocabulary tag set is obtained; Based on the aforementioned vocabulary tag set, vocabulary filtering is performed on the vocabulary of multiple cases to obtain an effective vocabulary set; Semantic features are extracted from the effective vocabulary set to obtain multiple semantic feature vectors; Based on the multiple semantic feature vectors and the AI mediation library subsystem, semantic feature matching is performed to obtain a case feature set.
4. The digital mediation method integrating multimodal sentiment analysis as described in claim 3, characterized in that, The process of setting mediation parameters based on the case feature set to obtain an AI mediation parameter set includes: Extract case type, cause of action, and mediation object characteristics from the case feature set; The complexity of the case is calculated based on the characteristics of the mediation subjects and the preset calculation standards. Based on the case type, cause of action, and complexity, the AI mediation library subsystem is used to configure basic parameters to obtain a set of basic parameters. Obtain the start time of mediation; The termination time is determined based on the mediation start time and the preset maximum mediation duration. Based on the aforementioned case types and causes of action, a set of risk keywords was determined; Based on the aforementioned set of basic parameters, mediation start time, termination time, risk keyword set, and preset human intervention conditions, an AI mediation parameter set is constructed.
5. The digital mediation method integrating multimodal sentiment analysis as described in claim 4, characterized in that, The initial interaction using the AI interaction system to obtain multimodal data includes: The AI digital human system was identified from the AI interaction system. Based on the AI mediation robot template, the AI digital human system is configured with robot settings to obtain a mediation digital human; The AI interaction subsystem was identified from the AI interaction system; Based on the case feature set, obtain the initial dialogue plan; The initial dialogue plan is input into the mediation digital human using the AI interaction subsystem, and the initial mediation interaction is performed using the mediation digital human to obtain multimodal data, which includes video data, audio data and text data.
6. The digital mediation method integrating multimodal sentiment analysis as described in claim 5, characterized in that, The AI-based mediation engine system performs sentiment analysis and annotation on multimodal data to obtain a feature annotation set, including: Based on the AI interaction subsystem, facial expression recognition is performed on the video data in the multimodal data to obtain video text elements; The AI interaction subsystem is used to perform speech feature analysis on audio data to obtain audio text elements. The AI interaction subsystem is used to perform keyword analysis on text data to obtain message text elements; The AI intelligent annotation subsystem was identified from the AI mediation engine system; The video text elements, audio text elements, and message text elements are transmitted to the AI intelligent annotation subsystem to obtain an object element set. Feature annotation is performed based on the AI mediation library subsystem and the object feature set to obtain the feature annotation set.
7. The digital mediation method integrating multimodal sentiment analysis as described in claim 6, characterized in that, The AI-based interactive subsystem performs facial expression recognition on the video data in the multimodal data to obtain video text elements, including: Based on the AI interaction subsystem, the video data in the multimodal data is processed by frame segmentation to obtain a video image sequence, wherein the video image sequence contains multiple video images; For each video image in the video image sequence, the following operation is performed: Target face detection is performed on the video image to obtain the face image region; Facial key point detection is performed based on the face image region to obtain the coordinates of multiple facial key points; Facial expression status is obtained by using the coordinates of multiple facial key points for expression recognition. The facial expression states are summarized to obtain a facial expression sequence; Extract the set of termination expressions from the facial expression sequence; Based on the set of terminating expressions, determine the object's final emotion; Based on the final emotion and facial expression sequence of the object, video text elements are constructed.
8. The digital mediation method integrating multimodal sentiment analysis as described in claim 7, characterized in that, The AI interaction subsystem is used to perform speech feature analysis on audio data to obtain audio text elements, including: The audio data is segmented using a preset time interval to obtain multiple audio segments; For each of the plurality of audio segments, the following operation is performed: Audio features are extracted from the audio segment to obtain multiple Mel frequency cepstral coefficient feature vectors; Multiple acoustic probability distributions are calculated using multiple Mel frequency cepstral coefficient eigenvectors and a pre-constructed acoustic model; Phoneme decoding is performed on the multiple acoustic probability distributions to obtain the optimal phoneme state sequence; The optimal phoneme state sequence is transformed using the AI interaction subsystem to obtain audio-recognized text; The audio-recognized text is labeled with keywords to obtain an audio keyword set; Extract the fundamental frequency and energy characteristics of the audio segment; Based on the fundamental frequency characteristics and energy characteristics, the audio pitch characteristics are determined; Based on the aforementioned audio intonation features and audio keyword set, audio emotion features are constructed; By summarizing the aforementioned audio emotional features, audio text elements are obtained.
9. The digital mediation method integrating multimodal sentiment analysis as described in claim 8, characterized in that, The method, based on the feature annotation set and AI mediation parameter set, utilizes the AI mediation engine system to perform mediation template matching, obtaining multimodal interaction information, including: The AI mediation dialogue subsystem was identified from the AI mediation engine system. Based on the element annotation set, the AI mediation dialogue subsystem is used to perform feature fusion to obtain the mediation feature vector; Based on the AI mediation library subsystem, similarity retrieval is performed using mediation feature vectors to obtain the target mediation dialogue template; Based on the AI mediation parameter set, obtain the mediation constraints; Historical interaction content is obtained from the multimodal data; Based on the mediation constraints, historical interaction content, and target mediation dialogue template, obtain multimodal interaction information.
10. A digital mediation system integrating multimodal sentiment analysis, characterized in that, The system includes: The mediation preparation module is used to acquire the AI mediation system, wherein the AI mediation system includes: a large-scale mediation model system, an AI mediation engine system, and an AI interaction system; The emotion labeling module is used to perform mediation matching using the AI mediation engine system and mediation big model system in the AI mediation system, to obtain AI mediation robot templates and AI mediation parameter sets, and to perform initial interaction using the AI interaction system to obtain multimodal data; The mediation interaction execution module is used to perform sentiment analysis and annotation on multimodal data based on the AI mediation engine system to obtain an element annotation set. Based on the element annotation set and the AI mediation parameter set, the AI mediation engine system is used to perform mediation template matching to obtain multimodal interaction information. The mediation process control module is used to conduct mediation interaction using an AI interaction system based on multimodal interaction information, obtain feedback information, and if the feedback information does not meet the preset termination conditions, the feedback information is treated as multimodal data and the process is returned to the above-mentioned step of performing sentiment analysis and annotation on the multimodal data based on the AI mediation engine system until the feedback information meets the termination conditions. If the feedback information meets the termination conditions, a mediation result report is generated, and the AI mediation is completed.