Multi-modal digital intelligent evaluation method and system for foreign talk ability

By constructing an analysis model of diplomatic discourse capability through multimodal data processing and deep learning technology, the problem of inconsistent and incomplete evaluation of diplomatic discourse capability has been solved, enabling real-time, quantitative, and scientific assessment of diplomatic discourse capability and improving the objectivity and practicality of the evaluation.

CN120974138APending Publication Date: 2025-11-18ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511073064.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

The existing technology lacks a unified standard for evaluating diplomatic discourse competence, the evaluation methods are not comprehensive enough, and it is difficult to quantify, resulting in a lack of scientific rigor and objectivity in the assessment of diplomatic personnel's discourse competence.

Method used

By employing multimodal data processing techniques, combined with Transformer architecture, graph neural networks (GNN), and multi-task learning, a diplomatic discourse capability analysis model is constructed. Through quantifying evaluation indicators of diplomatic discourse capability, scientific and comprehensive evaluation results are generated.

Benefits of technology

It enables real-time, quantitative, and scientific assessment of diplomatic discourse capabilities, enhancing the objectivity and practicality of the evaluation, and is applicable to scenarios such as diplomatic negotiations, media releases, and diplomat training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974138A_ABST
    Figure CN120974138A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode digital intelligence evaluation method and system for foreign exchange speech ability, and belongs to the technical field of foreign exchange data processing. Multi-modal utterance data (videos, audios and texts) of utterances of foreign traffic personnel are acquired, a multi-modal utterance data processing analysis module is used for data synchronization, format conversion and standardization, a foreign traffic utterance ability analysis model is constructed in combination with a Transform architecture, a graph neural network (GNN) and multi-task learning, foreign traffic utterance ability evaluation indexes are embedded, and foreign traffic utterance ability evaluation is realized. And finally, generating a foreign talk ability evaluation result through algorithm analysis. The method can quantitatively and scientifically evaluate the discourse ability of the foreign exchange personnel in real time, solves the problems that a traditional evaluation mode is high in subjectivity and difficult to quantify, has the advantages of being efficient, accurate and comprehensive, is suitable for various scenes such as foreign exchange negotiation, media release and foreign exchange training, and remarkably improves objectivity and practicability of foreign exchange discourse ability evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of diplomatic data processing technology, and more specifically to a multimodal, digitalized evaluation method and system for diplomatic discourse capabilities. Background Technology

[0002] "Diplomatic discourse competence" refers to the comprehensive ability of diplomatic actors to convey foreign policy positions, safeguard national interests, shape their international diplomatic image, promote the construction, translation, and dissemination of diplomatic discourse, and ultimately gain diplomatic discourse power through language and paralanguage in international exchanges. This concept is closely related to constructing diplomatic behavior, interpreting diplomatic events, and expounding diplomatic concepts, and is an important indicator of diplomatic soft power and comprehensive national strength.

[0003] Currently, with the deepening of globalization and the increasing frequency of international exchanges, the importance of diplomatic discourse competence is becoming increasingly prominent. Diplomats need to possess proficient diplomatic communication skills and abilities to ensure the smooth implementation of national foreign policy and the effective dissemination of the country's image. However, the evaluation of diplomatic discourse competence still faces some challenges, such as inconsistent evaluation standards, incomplete evaluation methods, and difficulties in quantifying the evaluation process.

[0004] Therefore, how to propose a multimodal and digital evaluation method and system for diplomatic discourse competence, and achieve a scientific and reasonable quantitative evaluation of diplomatic discourse competence, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a multimodal and digital evaluation method and system for diplomatic discourse competence, which evaluates the discourse competence of diplomats in a quantitative, scientific and comprehensive manner, thereby helping to improve the efficiency of diplomatic activities.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] On the one hand, this invention discloses a multimodal, digitally-driven evaluation method for diplomatic discourse competence, comprising the following steps:

[0008] Acquire multimodal discourse data from diplomats and diplomatic missions, process the multimodal discourse data, and obtain a diplomatic discourse dataset;

[0009] Construct and train a diplomatic discourse capability analysis model, embed diplomatic discourse capability evaluation indicators into the trained diplomatic discourse capability analysis model, and obtain a diplomatic discourse capability evaluation model.

[0010] The diplomatic discourse dataset is input into the diplomatic discourse competence evaluation model to obtain the diplomatic discourse competence evaluation results.

[0011] Preferably, the multimodal discourse data includes video data, audio data, and text data;

[0012] Processing the multimodal discourse data includes:

[0013] The video data, audio data, and text data are synchronized based on timestamps.

[0014] The synchronized multimodal discourse data is format-converted and standardized to obtain preprocessed data;

[0015] The preprocessed data is processed using data augmentation techniques to construct a diplomatic discourse dataset.

[0016] Preferably, the synchronization processing of the video data, the audio data, and the text data based on timestamps includes:

[0017] A timestamp-based multimodal alignment technique is employed to ensure the temporal consistency of the multimodal speech data by performing preliminary synchronization processing on the video data, audio data, and text data.

[0018] When timestamps are lost or deviated, secondary calibration is performed using specific markers in the audio data or keyframes in the video data as reference points to ensure the synchronization of multimodal speech data.

[0019] Preferably, the diplomatic discourse capability analysis model is constructed by combining the Transformer architecture, GNN, and multi-task learning;

[0020] The diplomatic discourse capability analysis model is trained based on a diplomatic knowledge base.

[0021] The construction of the diplomatic knowledge base includes:

[0022] The diplomatic knowledge base is updated in real time using a combination of scheduled tasks and event-driven methods.

[0023] A cross-review mechanism is introduced to annotate the updated data in the diplomatic knowledge base;

[0024] After data labeling is completed, data cleaning and standardization are performed, and the standardized diplomatic knowledge base data is processed using data augmentation technology.

[0025] Preferably, the diplomatic discourse competence evaluation indicators are embedded into the trained diplomatic discourse competence analysis model to obtain the diplomatic discourse competence evaluation model, including:

[0026] Establish the aforementioned evaluation indicators for diplomatic discourse capabilities;

[0027] Formalize the aforementioned indicators for evaluating diplomatic discourse capabilities;

[0028] By using multi-task learning, the formalized evaluation indicators of diplomatic discourse competence are embedded into the trained diplomatic discourse competence analysis model to obtain the diplomatic discourse competence evaluation model.

[0029] Preferably, the evaluation indicators for diplomatic discourse competence are formalized, including:

[0030] The evaluation indicators of diplomatic discourse capability are quantified, and the evaluation indicators are classified based on scoring rules and logical expressions.

[0031] A diplomatic discourse capability evaluation model is constructed using machine learning algorithms. The quantitative and classification results of the diplomatic discourse capability evaluation indicators are then optimized using the model to obtain the prediction results for each diplomatic discourse capability evaluation indicator.

[0032] The predicted results of each diplomatic discourse capability evaluation indicator are weighted and summed to obtain the diplomatic discourse capability score.

[0033] Preferably, the total loss function of the diplomatic discourse evaluation model is Loss total as follows:

[0034] Loss total =α·Loss 分析 +β·Loss 评价 ;

[0035] In the formula, α and β are the weighting parameters for adjusting the importance of the analysis task and the evaluation task, respectively, and Loss 分析 Loss is the loss of the diplomatic discourse capability analysis model. 评价 The loss of the diplomatic discourse capability evaluation model;

[0036] Loss of the diplomatic discourse power analysis model 分析 And the loss of the diplomatic discourse capability evaluation model. 评价 Both are weighted cross-entropy losses.

[0037] Preferably, a multimodal and digital evaluation method for diplomatic discourse capability also includes:

[0038] Generate a diplomatic discourse capability assessment report and visualize the assessment results, enabling personalized data display and interactive exploration functions.

[0039] On the other hand, this invention also discloses a multimodal, digitalized evaluation system for diplomatic discourse competence, used to implement the aforementioned evaluation method for diplomatic discourse competence, comprising:

[0040] The data processing module is used to acquire multimodal discourse data from diplomats and diplomatic institutions, process the multimodal discourse data, and obtain a diplomatic discourse dataset.

[0041] The model building module is used to build and train the diplomatic discourse capability analysis model, embed the diplomatic discourse capability evaluation index into the trained diplomatic discourse capability analysis model, and obtain the diplomatic discourse capability evaluation model.

[0042] The results output module is used to input the diplomatic discourse dataset into the diplomatic discourse competence evaluation model to obtain the diplomatic discourse competence evaluation results.

[0043] Preferably, a multimodal, digitalized evaluation system for diplomatic discourse capabilities also includes:

[0044] The visualization module is used to generate a diplomatic discourse capability assessment report and to visualize the assessment results, enabling personalized data display and interactive exploration functions.

[0045] As can be seen from the above technical solution, compared with the prior art, this invention discloses a multimodal, digitally intelligent evaluation method and system for diplomatic discourse competence. It acquires multimodal discourse data (video, audio, and text) from diplomatic personnel, utilizes a multimodal discourse data processing and analysis module for data synchronization, format conversion, and standardization, and combines Transformer architecture, Graph Neural Networks (GNNs), and multi-task learning to construct a diplomatic discourse competence analysis model. It embeds diplomatic discourse competence evaluation indicators and finally generates diplomatic discourse competence evaluation results through algorithmic analysis. This invention can evaluate the discourse competence of diplomatic personnel in real time, quantitatively, and scientifically, solving the problems of strong subjectivity and difficulty in quantification in traditional evaluation methods. It has the advantages of high efficiency, accuracy, and comprehensiveness, and is applicable to various scenarios such as diplomatic negotiations, media releases, and diplomatic training, significantly improving the objectivity and practicality of diplomatic discourse competence evaluation. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0047] Figure 1 A flowchart of the method provided by the present invention;

[0048] Figure 2 The system architecture diagram provided for this invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] In this invention, the supporting hardware facilities are a fundamental component of the intelligent analysis and evaluation system for diplomatic discourse capabilities, providing necessary physical support for the system's data acquisition and real-time analysis. These hardware devices are responsible for efficiently and accurately acquiring multimodal discourse data (such as video, audio, and text) from diplomatic negotiations, meetings, and media scenarios, laying the foundation for subsequent data processing and analysis. The appropriate configuration of the hardware facilities directly affects the system's data quality, processing efficiency, and analytical accuracy. The supporting hardware facilities include video acquisition equipment, audio acquisition equipment, data processing servers, and thermal imaging equipment.

[0051] Video capture equipment: ① High-definition cameras: Used for real-time recording of video data in diplomatic meetings or negotiation scenarios, featuring high resolution, low latency, and autofocus. They support multi-angle shooting to ensure comprehensive capture of diplomats' facial expressions, gestures, and other non-verbal behaviors. ② Multi-camera systems: In large-scale conference scenarios, multi-camera systems can be used for scene switching and multi-view capture, enabling richer data input.

[0052] Audio Acquisition Equipment: ① High-sensitivity microphone: Used for audio data acquisition, equipped with noise reduction capabilities, enabling clear pickup of diplomats' voices even in noisy environments. Supports multi-channel audio input to capture the voices of different speakers. ② Audio processor: Used for real-time audio processing, capable of filtering background noise, enhancing speech clarity, and providing high-quality audio input for subsequent speech analysis. ③ Voiceprint technology: Utilizes Mel-frequency cepstral coefficients (MFCC) and linear predictive cepstral coefficients (LPCC) to extract representative features from speech signals. After preprocessing, voiceprint features are extracted to construct a recognition model.

[0053] Data Processing Servers: High-performance computing servers are used to process large amounts of video, audio, and text data. These servers are equipped with GPUs to accelerate the inference process of deep learning models. The servers need to have large-capacity storage, fast read speeds, and the ability to process multimodal speech data and perform complex analysis tasks. A distributed computing architecture is adopted, leveraging frameworks such as Hadoop and Spark to efficiently process large-scale data. A distributed database system is used to store the data, and data preprocessing and filtering are performed at the acquisition end. The servers collaborate with the cloud to reduce data transmission volume and latency, and improve response speed.

[0054] Thermal imaging equipment: Introducing thermal imaging sensors, using infrared detectors to capture infrared radiation emitted by objects, and generating thermal maps reflecting temperature distribution through signal processing, providing raw data for subsequent analysis of object temperature characteristics, temperature-based behavior analysis, etc.

[0055] On the one hand, embodiments of the present invention disclose a multimodal, digitally-driven evaluation method for diplomatic discourse competence, such as... Figure 1 As shown, it includes the following steps:

[0056] S1. Acquire multimodal discourse data from diplomatic personnel (including career diplomats, multilateral diplomats, non-governmental diplomats, scholars from diplomatic think tanks, etc.) and diplomatic institutions; process the multimodal discourse data to obtain a diplomatic discourse dataset, including:

[0057] S11. Synchronize video data, audio data, and text data based on timestamps, including:

[0058] A timestamp-based multimodal alignment technique is employed to ensure temporal consistency across video frames, audio signals, and text transcription data through synchronized processing. When timestamps are lost or deviated due to the complexity of diplomatic scenarios, specific markers in the audio signal (such as pitch at a specific frequency) or keyframes in the video (such as specific actions or expressions of diplomats) can be used as reference points for secondary calibration to ensure data synchronization.

[0059] In this embodiment, video data requires frame extraction and face recognition, audio data requires speech recognition and sentiment analysis, and text data requires semantic parsing and keyword extraction.

[0060] In video processing, this embodiment utilizes computer vision technologies, such as OpenCV or deep learning models (such as YOLO or MTCNN), for image processing and feature extraction. It also employs object detection algorithms like optical flow in OpenCV or YOLOv5 to track diplomats in real-time, separating them from the background and ensuring the focus remains on the individuals while minimizing background interference. These technologies can identify nonverbal behaviors such as facial expressions and gestures, providing multi-layered input for subsequent evaluation of diplomatic discourse competence. In audio processing, a speech recognition system is used (considering network dependence and privacy concerns, an open-source, localized speech recognition system, such as Kaldi or Mozilla's DeepSpeech, is employed. These systems can run on local servers without relying on external networks and use encryption to protect data security when converting audio to text), combined with emotion recognition and intonation analysis tools (such as OpenSmile) to capture speaker emotions and tone changes. These technologies not only enhance the understanding of discourse content but also provide crucial support for nonverbal communication analysis in diplomatic scenarios.

[0061] S12. Perform format conversion and standardization on the synchronized multimodal discourse data to obtain preprocessed data.

[0062] After initial synchronization and processing, multimodal discourse data requires format conversion and standardization to ensure smooth input into the diplomatic discourse competence analysis and evaluation model. Video and audio data need to be converted into feature vectors or sequence data, while text data needs to be embedded to generate numerical forms easily readable by the model. Encoder-decoder architectures from deep learning (such as LSTM and BERT) are used for feature extraction and embedding generation of different modalities. Furthermore, standardization is crucial to ensure that different data sources have the same scale and format when input into the analysis model, reducing noise interference with the results. To achieve this conversion, pre-trained models (such as ResNet and VGG for video feature extraction and BERT for text embedding) and a custom processing pipeline are used to convert multimodal discourse data into a unified input format. Especially when processing audio data, spectral analysis and speech embedding techniques are employed to convert the raw sound wave signal into a spectrogram or speech feature vector, providing a foundation for subsequent semantic understanding and sentiment analysis.

[0063] Specifically, the video data is scaled to a fixed size (224×224), and pixel values ​​are normalized to [0,1] to ensure consistency in format and numerical range across different videos. For audio data, the sampling rate is unified to 16kHz, and the audio amplitude is normalized to [-1,1] to guarantee consistency in audio features. For text data, stemming is performed to simplify vocabulary, stop words are removed, and L2 normalization is applied after embedding the text to ensure consistent vector scale and provide standard input for the analysis model. To further improve processing efficiency, dimensionality reduction techniques such as Principal Component Analysis (PCA) and Linear Discriminant Analysis (LDA) are used to reduce the dimensionality of the converted and standardized video, audio, and text feature vectors. This reduces data complexity while preserving key information, thereby improving the system's processing speed and efficiency.

[0064] S13. Use data augmentation techniques to process preprocessed data and construct a diplomatic discourse dataset.

[0065] Due to the complex data acquisition environment in diplomatic scenarios, different devices may produce variations in data quality. Therefore, data augmentation and optimization are crucial steps in improving the quality of model input. Data augmentation techniques (such as data expansion, random noise injection, and data smoothing) can enhance the robustness of multimodal discourse data, making it more adaptable to complex diplomatic scenarios. Simultaneously, denoising algorithms (such as filters and autoencoders) are employed to optimize audio and video data, eliminating background noise and unnecessary interference to ensure data purity and reliability.

[0066] Furthermore, after the multimodal discourse data has been integrated, processed, transformed, and optimized, the results need to be seamlessly integrated into the diplomatic discourse capability analysis and evaluation model. This process requires consideration of data interface design to ensure smooth data flow into the model for subsequent in-depth analysis and evaluation. To achieve efficient collaboration between modules, this embodiment utilizes compatible API interfaces, data pipelines, or real-time data stream processing frameworks (such as Apache Kafka) to introduce an asynchronous processing mechanism, constructing a stable data transmission channel to ensure the system can respond in real time to dynamic data inputs in diplomatic scenarios and smoothly transition during system upgrades.

[0067] S2. Construct and train a diplomatic discourse capability analysis model, embed diplomatic discourse capability evaluation indicators into the trained diplomatic discourse capability analysis model, and obtain a diplomatic discourse capability evaluation model.

[0068] S21. Construct a diplomatic discourse capability analysis model by combining Transformer architecture, GNN and multi-task learning.

[0069] Transformer models use self-attention to model global dependencies in the input text, enabling them to capture complex relationships and deep context in language. Representative models of this type include BERT (Bidirectional Encoder Representations from Transformers) and GPT (Generative Pre-trained Transformer). BERT, through bidirectional encoding—understanding text simultaneously from left to right and right to left—significantly improves its ability to understand context, making it suitable for tasks such as text classification, relation extraction, and named entity recognition. Its bidirectional nature allows it to accurately capture subtle differences and multi-layered semantic relationships in diplomatic texts. GPT, on the other hand, focuses on generative tasks, excelling in language generation, dialogue understanding, and discourse coherence analysis. It is suitable for predicting potential next moves or linguistic trends in diplomatic speeches, providing reasonable predictions for complex diplomatic scenarios.

[0070] To further enhance the model's ability to understand multi-layered relationships in diplomatic events, this embodiment combines Graph Neural Networks (GNNs) to handle entity relationships and network structures within events. GNNs model relationships between different entities, such as countries, diplomats, and agreements, through graph structures (nodes and edges). Information is propagated in the graph via message passing, enabling the model to capture complex interactions and hierarchical relationships between entities. In a diplomatic negotiation scenario, GNNs can connect elements such as diplomats, national positions, and historical events, providing a more in-depth analysis of the event's background. GNNs also support multi-hop reasoning, enabling layer-by-layer deduction based on the relationships between different nodes in the graph structure, revealing the implicit logical chains behind diplomatic language, and enhancing the model's insight into complex international events.

[0071] Multi-task learning is another key technique for building efficient analytical models. By sharing some of the model's parameters, multi-task learning allows the model to handle multiple related tasks simultaneously, such as text understanding, sentiment analysis, discourse style detection, and contextual analysis. This approach can fully utilize resources even with relatively limited data, and knowledge sharing between different tasks can improve the model's overall performance. When dealing with diplomatic language, the model can simultaneously perform semantic understanding and sentiment analysis, thereby comprehensively assessing the rationality and emotional tone of a speech and accurately judging the potential direction of diplomatic strategy. Furthermore, by coordinating the learning objectives of different tasks, multi-task learning enables the model to more comprehensively understand contextual changes, underlying positions, and strategic intentions in speeches, providing multi-dimensional data support for assessing diplomats' communication skills and strategy execution capabilities.

[0072] By combining the Transformer architecture, GNN, and multi-task learning, the analytical model demonstrates strong adaptability and understanding in diplomatic event analysis. It not only handles the complexity of language but also dynamically integrates multi-layered information relationships, providing scientific support for the assessment of real-time diplomatic discourse capabilities. This integrated model design ensures that the system possesses comprehensive language analysis and event correlation capabilities when facing complex international situations, contributing to more accurate decision-making and assessment.

[0073] This embodiment constructs a diplomatic knowledge graph within the Transformer architecture, embedding a diplomatic knowledge module. Key diplomatic information is used as nodes, and semantic relationships (such as causal relationships, temporal order, and logical connections) are used as edges. Graph features are added to the input text vocabulary to help the model accurately identify and process key information. Simultaneously, the attention mechanism is optimized by adding a weight adjustment module. The words in the input text are mapped to a low-dimensional space via a linear layer and then dot-producted with the embedding vectors of key information in the knowledge graph to obtain a relevance score. The attention weights are adjusted based on this score, paying more attention to the context when key information appears. Furthermore, pruning techniques are employed to lightweight the model. During model training, an L1 regularization term is added to the loss function, causing unimportant parameters to gradually approach zero, and connections with zero parameter values ​​are removed, reducing the number of model parameters and improving real-time evaluation capabilities.

[0074] S22. Training a diplomatic discourse capability analysis model based on a diplomatic knowledge base.

[0075] Building a diplomatic knowledge base includes:

[0076] S221. A combination of scheduled tasks and event-driven methods is used to update the diplomatic knowledge base in real time.

[0077] For scheduled tasks, the system utilizes libraries like Python's APScheduler to retrieve new data from data sources at fixed intervals (e.g., hourly, daily). Simultaneously, it leverages event-driven mechanisms (monitoring specific news sources) to immediately trigger the data retrieval process upon the release of new diplomatic events or policies. This utilizes appropriate web crawling techniques or API services (if available) to promptly obtain the latest reports according to the website's structure rules. The system regularly acquires the latest international news and political event data in this way, continuously updating the knowledge base to ensure the analysis model possesses up-to-date international contextual knowledge. This data should cover different languages, cultural backgrounds, and historical periods to enrich the model's training corpus.

[0078] S222. Introduce a cross-checking mechanism to annotate the data in the updated diplomatic knowledge base.

[0079] The collected data is meticulously labeled, with tags including event type, discourse style, diplomatic strategy, context, and sentiment. To ensure high-quality labeling, a combination of manual and semi-automated labeling tools can be used to improve efficiency and accuracy. A cross-review mechanism is also introduced, assigning labeling tasks to multiple labelers for independent completion. After labeling, consistency among labelers is assessed using a consistency check algorithm (such as the Kappa statistic). If the consistency index falls below a set threshold (e.g., 0.8), the data is marked as disputed and subject to arbitration.

[0080] S223. After completing data labeling, perform data cleaning and standardization, and then process the standardized diplomatic knowledge base data through data augmentation techniques.

[0081] The collected corpus undergoes cleaning and enhancement processes. Operations such as removing irrelevant symbols, standardizing terminology, and adjusting sentence structure ensure the consistency and accuracy of the training data. Data augmentation techniques, such as text reconstruction and translation enhancement, increase the diversity of the corpus, helping the model better handle speeches in different diplomatic scenarios.

[0082] The training process of the diplomatic discourse capability analysis model includes four key steps: data preprocessing, model pretraining, fine-tuning training, and adaptive learning.

[0083] 1) Data Preprocessing and Augmentation: The data in the diplomatic knowledge base is cleaned, standardized, and augmented to ensure high-quality and consistent input data for the model. First, natural language processing techniques are used for data cleaning to remove redundant information and noise, such as duplicate text, invalid symbols, or formatting errors. Next, the data is standardized, including language unification, tense adjustment, and terminology standardization, to ensure compatibility between data from different sources. Finally, data augmentation is performed, expanding the diversity of the data through methods such as data restructuring, semantic transformation, and translation augmentation, thereby improving the model's ability to understand different diplomatic discourses.

[0084] 2) Model pre-training: Deep learning using a large-scale knowledge base.

[0085] (1) Objectives and procedures of pre-training

[0086] Pre-training is the initial stage of deep learning model training, aiming to enable the model to master the basic structure, contextual relationships, and deep semantics of language through large-scale unsupervised learning. This stage does not require manually labeled data, but instead utilizes rich text resources in the diplomatic knowledge base (such as diplomatic documents, news reports, official statements, etc.) for training. The core objective of pre-training is to equip the model with broad language comprehension capabilities, enabling it to identify key elements and multi-layered relationships in diplomatic discourse, laying the foundation for subsequent fine-tuning for specific tasks.

[0087] (2) Transformer architecture: the core of self-attention mechanism

[0088] Pre-trained models are typically based on the Transformer architecture, such as BERT and GPT. Transformers excel in natural language processing due to their self-attention mechanism, which models global dependencies within sentences. This self-attention mechanism allows the model to consider the relationships between other words in the sentence while processing each word, thus capturing contextual semantic information. Compared to traditional recurrent neural networks (RNNs), the parallel processing capabilities of Transformers significantly improve training efficiency and model performance.

[0089] BERT (Bidirectional Encoder Representations from Transformers) employs bidirectional encoding (simultaneous left-to-right and right-to-left encoding), enabling the model to fully understand contextual information during pre-training. It is trained using a Masked Language Model (MLM) and a NextSentence Prediction (NSP) task: MLM randomly masks parts of the words, allowing the model to predict the masked content, training its semantic understanding ability; NSP, on the other hand, allows the model to determine whether two sentences are consecutive, learning the logical relationships between sentences. This enables BERT to accurately capture subtle semantic differences and complex contexts in diplomatic texts. GPT (Generative Pre-trained Transformer), with generation as its core task, uses a unidirectional decoder structure, learning the contextual relationships of sentence generation by predicting the next word. It excels in language generation and dialogue understanding, and can be used to analyze the coherence and logic of diplomatic speeches. In diplomatic scenarios, GPT can predict the potential next strategies of diplomats, providing reasonable linguistic trend analysis for complex dialogues and negotiations.

[0090] (3) Data processing and training details

[0091] The first step is corpus selection. The pre-training phase requires a massive amount of high-quality corpus. Dynamic knowledge bases offer a rich variety of data sources, including news articles, historical diplomatic records, national policy documents, and academic papers. To improve the model's adaptability to multicultural and multilingual diplomatic texts, the corpus needs to cover the language habits and expression styles of different countries and regions.

[0092] Next comes data cleaning and augmentation. The collected corpus undergoes cleaning and augmentation, including removing irrelevant symbols, standardizing terminology and sentence structure, to ensure the consistency and accuracy of the training data. Data augmentation techniques such as text reconstruction and translation enhancement can increase the diversity of the corpus, helping the model better handle speeches in different diplomatic scenarios.

[0093] Finally, there's batch training and hyperparameter tuning. Pre-training is typically done in batches to avoid memory shortages caused by massive amounts of data. During training, the settings of hyperparameters (such as learning rate, batch size, number of attention heads, etc.) directly affect model performance and need to be optimized through experiments to ensure the model performs best in its specific domain.

[0094] 3) Fine-tuning training

[0095] (1) Purpose and significance of fine-tuning

[0096] Fine-tuning training is a supervised learning phase that applies a pre-trained model to a specific application scenario. Its core objective is to optimize the model using labeled data, enabling it to perform exceptionally well in real-world diplomatic tasks. During fine-tuning, the model needs to not only understand the surface meaning of the text but also delve deeper to uncover key information such as the speaker's emotional attitude, stance, and strategic intent.

[0097] (2) Task Design and Multi-Task Learning

[0098] Fine-tuning tasks include classification tasks (such as determining the stance of a statement), sequence labeling tasks (such as identifying entities and relationships in a statement), and generation tasks (such as predicting the next step in a statement). The task design should be closely aligned with the needs of diplomatic discourse competence analysis, ensuring the model can achieve multi-dimensional understanding in complex scenarios. During fine-tuning, multi-task learning is employed, allowing the model to simultaneously handle multiple related tasks within the same framework, such as text understanding, sentiment analysis, discourse style detection, and context analysis. Multi-task learning, by sharing some model parameters, enables complementary and transferable knowledge between different tasks, thereby enhancing the model's overall analytical capabilities. For example, when analyzing diplomatic statements, the model can simultaneously process sentiment and context tasks to determine the rational and emotional components of a statement and predict potential strategic moves.

[0099] (3) Refine task training and strategy optimization

[0100] First, supervised learning and data annotation. Fine-tuning training relies on high-quality labeled data. The labeled content needs to cover important information in diplomatic statements, such as event type, emotional attitude, and contextual changes. Semi-automated annotation tools combined with manual annotation can be used to ensure data accuracy and diversity. During annotation, attention should be paid to diplomatic-specific terminology and strategic expressions to ensure the professionalism of the data. Second, adaptive adjustment and strategy optimization. During training, the training strategy is adaptively adjusted based on the model's performance in various tasks, such as adjusting the weights of the loss function to balance the learning effects across different tasks. Optimization strategies also include adversarial training, which improves the model's coping ability by generating virtual data of complex diplomatic scenarios, enabling it to flexibly adjust its analysis methods when facing emerging topics. Third, combining graph neural networks (GNNs). To enhance the model's understanding of deep relationships in diplomatic events, GNNs can be introduced into fine-tuning training. Through graph structure modeling, entities (such as countries, individuals, agreements, etc.) and relationships in diplomatic events are represented as nodes and edges. GNNs use message passing mechanisms to integrate information from these nodes, helping the model identify and deduce complex international relationship networks.

[0101] 4) Adaptive learning and incremental updates:

[0102] To enable the analysis model to possess knowledge expansion and autonomous learning capabilities, and to automatically expand relevant knowledge domains based on the content of diplomats' speeches, thereby enhancing the system's understanding of emerging international affairs and topics, various technical means can be employed, including the introduction of a dynamic knowledge base, adaptive learning mechanisms, natural language understanding and reasoning technologies, and user feedback and proactive learning mechanisms. First, a dynamic knowledge base is established. This base possesses real-time update and automatic expansion capabilities, capturing the latest information on international affairs, diplomatic trends, and policy changes. Through automated data crawling and natural language processing technologies, the knowledge base can acquire data from authoritative international political databases, diplomatic archives, news media, government documents, and academic resources, constructing a multi-layered, cross-domain knowledge system. Simultaneously, knowledge graph technology is used to store information on diplomats, countries, events, and agreements in a graph structure, forming a highly interconnected knowledge network. Through semantic association and knowledge mapping capabilities, the knowledge base maps the content of diplomats' speeches to relevant fields, individuals, and events, enabling the system to understand the background, intentions, and complex international relations involved in the diplomats' speeches, thereby identifying and learning emerging topics. Secondly, adaptive learning mechanisms enable the model to continuously optimize and expand its knowledge in practical applications. Through incremental learning, the model can gradually absorb new information without losing existing knowledge, improving its analytical capabilities for current international political events. Combined with contextual reinforcement learning, the model can self-adjust according to the context of diplomatic statements, integrating relevant knowledge through automatic retrieval, allowing it to gradually adapt to and master new content when facing new domains. Adversarial training techniques enhance the model's ability to cope with challenging content and improve its generalization level by generating complex diplomatic scenarios and emerging topics. To further enhance the system's understanding and reasoning capabilities regarding complex international affairs in diplomatic statements, the integration of Natural Language Understanding (NLU) and machine reasoning technologies is particularly important. Deep semantic analysis technology can help the system semantically understand diplomatic statements, extract core content and key themes, and enhance context awareness and reasoning capabilities when combined with contextual understanding models (such as BERT and GPT). Causal reasoning and knowledge deduction techniques can help the system deduce hidden logical relationships, causal chains, and potential diplomatic motivations from diplomatic statements, thereby achieving automatic knowledge expansion. Finally, user feedback and proactive learning mechanisms further enhance the model's knowledge expansion capabilities. Through user feedback, the system can receive suggestions from diplomatic experts or users regarding the analysis results and adjust the model's analysis methods accordingly, achieving closed-loop optimization. Simultaneously, proactive learning technology allows the system to identify content in statements that is difficult to judge or not yet covered in the knowledge base, and proactively learn and update it. This mechanism enables the model to quickly adapt to emerging topics, enhancing its ability to cope with complex diplomatic situations. Through these measures, the analysis model can continuously expand its knowledge system in real-time operation and improve its understanding of emerging international affairs and topics.This intelligent expansion and self-learning capability enables the system to remain at the forefront in complex diplomatic environments, providing more accurate analysis and evaluation for diplomats' capability assessment and decision support.

[0103] S23. Embed the evaluation indicators of diplomatic discourse competence into the trained diplomatic discourse competence analysis model to obtain the diplomatic discourse competence evaluation model, including:

[0104] S231. Set evaluation indicators for diplomatic discourse capabilities.

[0105] Evaluation indicators are set from multiple dimensions, and the evaluation dimensions covered in the indicator system are shown in Table 1.

[0106] Table 1 Evaluation Index System for Diplomatic Discourse Capability

[0107]

[0108]

[0109] S232. Formalize the evaluation indicators of diplomatic discourse competence, including:

[0110] S2321. Quantify the evaluation indicators of various diplomatic discourse capabilities, and classify the evaluation indicators of various diplomatic discourse capabilities based on scoring rules and logical expressions.

[0111] Design mathematical expressions to transform each evaluation indicator into a quantifiable score function. For example, language expression ability can be quantified through sub-indicators such as the accuracy of diplomatic language, the artistry of diplomatic rhetoric, and the persuasiveness of diplomatic language. The score function can be defined as:

[0112] S-language expression ability = a × accuracy of diplomatic language + b × additional rhetorical artistry + c × persuasiveness of diplomatic language;

[0113] Here, a, b, and c represent the weights of each sub-indicator, set based on the AHP results. The scoring function converts the performance of each evaluation dimension in the text into specific numerical values, facilitating comprehensive calculation by the model.

[0114] To facilitate the model's understanding and processing of evaluation indicators, clear classification criteria and labeling systems need to be established. For example, for intercultural communication competence, classification criteria (such as high, medium, and low) can be established, with clear scoring rules and logical expressions for each category. When analyzing diplomatic discourse competence, the model can classify texts according to these criteria and assign corresponding scores.

[0115] For some complex characteristics of diplomatic language (such as political equivalence and aesthetic consistency), formalization can be achieved through logical rules and decision trees. For example, the model can handle political equivalence by identifying specific sentence patterns or keywords.

[0116] Political equivalence refers to the ability of diplomatic discourse to convey its original meaning and maintain balance across different cultural and political contexts. It is a crucial indicator of whether diplomatic language accurately conveys policy intentions without error. The model can assess the political equivalence of discourse through methods such as language pattern recognition, keyword identification, and contextual analysis. The logical expression for the political equivalence scoring is as follows:

[0117]

[0118] S2322. Construct a diplomatic discourse evaluation model using machine learning algorithms, optimize the quantitative and classification results of diplomatic discourse capability evaluation indicators using the diplomatic discourse capability evaluation model, and obtain the prediction results of each diplomatic discourse capability evaluation indicator.

[0119] Machine learning algorithms (such as linear regression, support vector machines (SVM), and neural networks) can be used to optimize scoring mechanisms and enhance the credibility of the output. For example, linear regression can adjust the weight coefficients in the scoring function to make the final score more consistent with the actual evaluation results; SVM can be used to optimize data classification and accurately distinguish the strengths and weaknesses of different diplomatic discourse performances; and neural networks can learn complex evaluation patterns during training, automatically adjust the scoring mechanism, and improve the comprehensive analytical level of diplomatic discourse competence.

[0120] S2323. The predicted results of each diplomatic discourse competence evaluation indicator are weighted and summed to obtain the diplomatic discourse competence score.

[0121] The prediction results of each sub-task (such as diplomatic discourse construction ability, diplomatic discourse translation ability, etc.) are weighted and synthesized to obtain the final diplomatic discourse ability score, as shown in the following formula:

[0122] Total score = W1 × S 外交话语构建能力 +W2×S 外交话语翻译能力 +…+W n ×S 其他能力

[0123] Where W1, W2, ..., W n The model assigns weights to each competency score. Through the synthesized overall score, it can assess the discourse performance of diplomats in real-world scenarios in real time, providing a scientific basis for diplomatic decision-making.

[0124] Among them, the weights W1, W2, ..., W of each ability score are... nFirst, the Analytic Hierarchy Process (AHP) is used, where experts score evaluation indicators based on their experience and knowledge to determine subjective weights. Second, the entropy method is used to calculate objective weights, determining the weights based on the numerical variations of each indicator's data (such as standard deviation or variance). The greater the data dispersion, the greater the variation of the indicator under different conditions, and the higher its objective weight. Then, the subjective weights obtained from AHP and the objective weights obtained from entropy are combined, and the comprehensive weight W(comprehensive) is calculated using the formula: W(AHP) + (1-λ)*W(entropy method). λ can be adjusted according to the actual situation (range 0-1). If expert opinions are trusted more, λ can be larger; if the data itself is valued more, λ can be smaller. Finally, the weights are dynamically adjusted according to different diplomatic scenarios. In the early stages of training, when data is scarce, λ can be set large to rely on expert opinions; as data increases, λ is reduced to allow the data to play a greater role. Under different task scenarios, the weights of related indicators can be adjusted to provide a more reliable basis for subsequent evaluations.

[0125] The above technologies and methods ensure that the various evaluation indicators of diplomatic discourse competence can be effectively integrated with the analytical model, enabling the model to possess accurate and real-time evaluation capabilities. This integrated design not only enhances the breadth of the model's application but also ensures the reliability and accuracy of the evaluation results.

[0126] S233. By using multi-task learning, the formalized evaluation indicators of diplomatic discourse competence are embedded into the trained diplomatic discourse competence analysis model to obtain the diplomatic discourse competence evaluation model.

[0127] Multi-task learning is a technique that trains multiple related tasks simultaneously, improving the learning performance of each task by sharing underlying model parameters. In this system, diplomatic event analysis and discourse competence assessment are interrelated tasks, making multi-task learning suitable. The core of multi-task learning lies in: 1. Shared representation: By sharing the underlying neural network structure, the model allows diplomatic event analysis and discourse competence assessment tasks to share most of the parameters, thereby improving performance through mutual reinforcement. 2. Task-specific evaluation head: In a multi-task model, a dedicated output layer is customized for each task. For example, a dedicated evaluation head can be designed for discourse competence assessment to score discourse performance. This head can be a fully connected layer or a more complex neural network module, such as an LSTM or convolutional neural network, to capture specific features of the evaluation metrics.

[0128] The total loss function of the diplomatic discourse competence evaluation model. total as follows:

[0129] Loss total =α·Loss 分析 +β·Loss评价 ;

[0130] In the formula, α and β are the weighting parameters for adjusting the importance of the analysis task and the evaluation task, respectively, and Loss 分析 Loss for the diplomatic discourse power analysis model 评价 The loss of the diplomatic discourse capability evaluation model;

[0131] Loss in the analysis model of diplomatic discourse power 分析 Loss of the diplomatic discourse competence assessment model 评价 Both are weighted cross-entropy losses, and the formulas are as follows:

[0132]

[0133] In the formula, n is the number of samples, k is the number of categories, and y ij It is the true label of the j-th class of the i-th sample. It is the corresponding predicted probability, w j This refers to category weights. Due to class imbalance in diplomatic discourse evaluation, important categories (such as "high equivalence") have fewer samples but require more precise predictions. The weight (wj) can be determined by combining expert knowledge, prior information, and experimental adjustments, selecting the optimal combination on the validation set. In implementation, the categories and weights are first determined, the weighted loss for each sample is calculated and summed, and then an optimization algorithm (such as stochastic gradient descent) is used to minimize this loss to update the model parameters, ensuring accurate evaluation of important indicators (such as political equivalence) under multi-task learning.

[0134] The S3. Diplomatic Discourse Competence Evaluation Model analyzes and evaluates the diplomatic discourse dataset to generate evaluation results. The results combine qualitative descriptions with quantitative indicators, categorizing them into four levels: unsatisfactory, satisfactory, good, and excellent. The quantitative evaluation criteria for each level are shown in Table 2.

[0135] Table 2 Quantitative Evaluation Standards at the Sequence Level

[0136] excellent 90 points (inclusive) - 100 points good 80 points (inclusive) - 90 points (exclusive) qualified 60 points (inclusive) - 80 points (exclusive) Unqualified 0-60 points (excluding 60 points)

[0137] S4. Generate a diplomatic discourse capability assessment report and visualize the assessment results, enabling personalized data display and interactive exploration functions.

[0138] Develop a visualization module, i.e., a user interface, to display evaluation results in the form of charts, scores, and text feedback, enabling users to intuitively understand the diplomats' performance. Furthermore, the feedback mechanism allows the system to provide suggestions and improvement plans based on the diplomats' real-time performance, which is of great significance for improving diplomats' communication skills in practical work. Specifically:

[0139] Real-time monitoring and data display:

[0140] ① Data Visualization Interface: Visually displays the current analysis progress and diplomat performance scores through charts, real-time video playback, and audio waveform graphs. ② Real-time Feedback and Analysis Report Generation: Users can view the scores for various indicators in real time on the interface and generate detailed evaluation reports for diplomat performance assessment and capacity building. ③ Personalized Data Display Options: Provides personalized data display options for different users, meeting diverse needs. ④ Interactive Exploration Function: Adds interactive exploration functionality, allowing users to delve deeper into specific data points and enhance their data analysis capabilities.

[0141] Operating System and Compatibility:

[0142] ① User Guide and Tips: The interface is simple and intuitive, featuring user guides and real-time tips for quick and easy operation. ② Report Export and Sharing: Evaluation reports can be exported to various formats (such as PDF and Excel) and support one-click sharing for easy dissemination and application. ③ Cross-Platform Compatibility: Compatibility with different operating systems (such as Windows, macOS, and Linux) and devices (such as PCs, tablets, and mobile phones) is considered to ensure smooth system operation in various environments.

[0143] On the other hand, this invention also discloses a multimodal, digitalized evaluation system for diplomatic discourse competence, with reference to Figure 2 This system is used to implement the aforementioned methods for evaluating diplomatic discourse competence, including:

[0144] The data processing module is used to acquire multimodal discourse data from diplomats, process the multimodal discourse data, and obtain a diplomatic discourse dataset.

[0145] The model building module is used to build and train the diplomatic discourse capability analysis model, embed the diplomatic discourse capability evaluation index into the trained diplomatic discourse capability analysis model, and obtain the diplomatic discourse capability evaluation model.

[0146] The results output module is used to input the diplomatic discourse dataset into the diplomatic discourse competence evaluation model to obtain the diplomatic discourse competence evaluation results.

[0147] The visualization module is used to generate a report on the assessment of diplomatic discourse capabilities and to visualize the assessment results, enabling personalized data display and interactive exploration functions.

[0148] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0149] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal, digitalized evaluation method for diplomatic discourse competence, characterized in that, Includes the following steps: Acquire multimodal discourse data from diplomats and diplomatic missions, process the multimodal discourse data, and obtain a diplomatic discourse dataset; Construct and train a diplomatic discourse capability analysis model, embed diplomatic discourse capability evaluation indicators into the trained diplomatic discourse capability analysis model, and obtain a diplomatic discourse capability evaluation model. The diplomatic discourse dataset is input into the diplomatic discourse competence evaluation model to obtain the diplomatic discourse competence evaluation results.

2. The multimodal, digitalized evaluation method for diplomatic discourse competence according to claim 1, characterized in that, The multimodal discourse data includes video data, audio data, and text data; Processing the multimodal discourse data includes: The video data, audio data, and text data are synchronized based on timestamps. The synchronized multimodal discourse data is format-converted and standardized to obtain preprocessed data; The preprocessed data is processed using data augmentation techniques to construct a diplomatic discourse dataset.

3. The multimodal, digitalized evaluation method for diplomatic discourse competence according to claim 2, characterized in that, Synchronization processing of the video data, audio data, and text data based on timestamps includes: A timestamp-based multimodal alignment technique is employed to ensure the temporal consistency of the multimodal speech data by performing preliminary synchronization processing on the video data, audio data, and text data. When timestamps are lost or deviated, secondary calibration is performed using specific markers in the audio data or keyframes in the video data as reference points to ensure the synchronization of multimodal speech data.

4. The multimodal, digitalized evaluation method for diplomatic discourse competence according to claim 1, characterized in that, The diplomatic discourse capability analysis model is constructed by combining the Transformer architecture, GNN, and multi-task learning. The diplomatic discourse capability analysis model is trained based on a diplomatic knowledge base. The construction of the diplomatic knowledge base includes: The diplomatic knowledge base is updated in real time using a combination of scheduled tasks and event-driven methods. A cross-review mechanism is introduced to annotate the updated data in the diplomatic knowledge base; After data labeling is completed, data cleaning and standardization are performed, and the standardized diplomatic knowledge base data is processed using data augmentation technology.

5. The multimodal, digitalized evaluation method for diplomatic discourse competence according to claim 1, characterized in that, By embedding the evaluation indicators of diplomatic discourse competence into the trained diplomatic discourse competence analysis model, the evaluation model of diplomatic discourse competence is obtained, including: Establish the aforementioned evaluation indicators for diplomatic discourse capabilities; Formalize the aforementioned indicators for evaluating diplomatic discourse capabilities; By using multi-task learning, the formalized evaluation indicators of diplomatic discourse competence are embedded into the trained diplomatic discourse competence analysis model to obtain the diplomatic discourse competence evaluation model.

6. The multimodal, digitalized evaluation method for diplomatic discourse competence according to claim 5, characterized in that, The evaluation indicators for diplomatic discourse competence are formalized, including: The evaluation indicators of diplomatic discourse capability are quantified, and the evaluation indicators are classified based on scoring rules and logical expressions. A diplomatic discourse evaluation model is constructed using machine learning algorithms. The quantitative and classification results of the diplomatic discourse capability evaluation indicators are then optimized using the model to obtain the prediction results for each indicator. The predicted results of each diplomatic discourse capability evaluation indicator are weighted and summed to obtain the diplomatic discourse capability score.

7. The multimodal, digitalized evaluation method for diplomatic discourse competence according to claim 6, characterized in that, The total loss function of the diplomatic discourse capability evaluation model is Loss total as follows: Loss total =α·Loss 分析 +β·Loss 评价 ; In the formula, α and β are the weighting parameters for adjusting the importance of the analysis task and the evaluation task, respectively, and Loss 分析 Loss is the loss of the diplomatic discourse capability analysis model. 评价 The loss of the diplomatic discourse capability evaluation model; Loss of the diplomatic discourse power analysis model 分析 And the loss of the diplomatic discourse capability evaluation model. 评价 Both are weighted cross-entropy losses.

8. The multimodal, digitalized evaluation method for diplomatic discourse competence according to claim 1, characterized in that, Also includes: Generate a diplomatic discourse capability assessment report and visualize the assessment results, enabling personalized data display and interactive exploration functions.

9. A multimodal, digitalized evaluation system for diplomatic discourse competence, characterized in that, include: The data processing module is used to acquire multimodal discourse data from diplomats and diplomatic institutions, process the multimodal discourse data, and obtain a diplomatic discourse dataset. The model building module is used to build and train the diplomatic discourse capability analysis model, embed the diplomatic discourse capability evaluation index into the trained diplomatic discourse capability analysis model, and obtain the diplomatic discourse capability evaluation model. The results output module is used to input the diplomatic discourse dataset into the diplomatic discourse competence evaluation model to obtain the diplomatic discourse competence evaluation results.

10. A multimodal, digitalized evaluation system for diplomatic discourse competence according to claim 9, characterized in that, Also includes: The visualization module is used to generate a diplomatic discourse capability assessment report and to visualize the assessment results, enabling personalized data display and interactive exploration functions.