A method for generating digital transaction outbound call strategies based on multimodal large model recognition

CN122675486APending Publication Date: 2026-09-01SHANGHAI JIASU E-COMMERCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610847101.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0003]目前,仅依赖单模态数据,无法对客户语音流、经授权的情绪面部表情图像及交易操作界面截图等多源异构数据进行实时融合与深度理解,导致难以捕捉客户微表情中的犹豫信号、界面操作中的决策延迟等隐性特征,造成客户意向判断滞后且误判率较高;同时,现有策略生成多基于固定话术库或简单决策树规则,缺乏对历史外呼策略知识图谱的结构化建模与图神经网络推理,无法在交互过程中动态识别异议触发点、成交临界点等关键决策节点,更无法通过反事实推理引擎在主流策略失效时主动生成备选话术序列,导致策略僵化且缺乏因果层面的纠错能力

Benefits of technology

1.本发明中,通过多模态数据采集模块同时获取客户语音流、经授权的情绪面部表情图像及交易操作界面截图,并利用跨模态注意力机制与门控融合网络对视觉、语音、文本特征进行加权融合与动态组合,实现了对客户微表情中的犹豫信号、语音中的情绪波动及界面操作中的决策延迟等多源异构信息的实时感知与深度对齐,解决了传统单模态分析因信息盲区导致的客户意向判断滞后与误判率高的问题,提升了外呼交互中客户真实意图的捕捉精度与响应速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122675486A_ABST
    Figure CN122675486A_ABST
Patent Text Reader

Abstract

This invention discloses a method for generating digital transaction outbound calling strategies based on multimodal large model recognition, relating to the fields of intelligent outbound calling and digital transaction technology. The method includes the following steps: acquiring real-time interaction data during the outbound calling process; extracting and fusing features from the real-time interaction data to generate a unified multimodal interaction feature vector; extracting key decision nodes and generating a strategy requirement feature vector; evaluating the strategy requirement feature vector in multiple dimensions and outputting an original strategy score; performing dynamic Bayesian calibration based on customer historical transaction information and an expert strategy library to generate a strategy confidence score; constructing a personalized outbound calling strategy based on the strategy confidence score; and dynamically optimizing model parameters and strategy generation thresholds. This invention achieves the generation and adaptive optimization of outbound calling strategies, solving the problems of traditional outbound calling strategies relying on fixed rules, lacking personalization, and being unable to dynamically adjust, thereby improving transaction conversion rates and customer experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent outbound calling and digital transaction technology, and in particular to a method for generating digital transaction outbound calling strategies based on multimodal large model recognition. Background Technology

[0002] Digital transaction outbound call strategy generation based on multimodal large model recognition refers to the use of large-scale deep learning models that can simultaneously process multiple types of information such as voice, text, and images to comprehensively identify and understand customer status and scenario characteristics during real-time outbound call interactions, and automatically generate or recommend suitable outbound call strategies accordingly.

[0003] Currently, relying solely on single-modal data makes it impossible to perform real-time fusion and deep understanding of multi-source heterogeneous data such as customer voice streams, authorized emotional facial expression images, and transaction operation interface screenshots. This makes it difficult to capture implicit features such as hesitation signals in customer micro-expressions and decision delays in interface operations, resulting in delayed judgment of customer intentions and a high misjudgment rate. At the same time, existing strategy generation is mostly based on fixed script libraries or simple decision tree rules, lacking structured modeling and graph neural network reasoning of historical outbound call strategy knowledge graphs. It is unable to dynamically identify key decision nodes such as objection trigger points and transaction thresholds during the interaction process, and it is also unable to proactively generate alternative script sequences when mainstream strategies fail through counterfactual reasoning engines. This results in rigid strategies and a lack of causal error correction capabilities.

[0004] Therefore, a digital transaction outbound call strategy generation method based on multimodal large model recognition is proposed to solve the above problems. Summary of the Invention

[0005] The main objective of this invention is to provide a method for generating digital transaction outbound call strategies based on multimodal large model recognition, so as to solve the problems mentioned in the background above.

[0006] To achieve the above objectives, the technical solution adopted by this invention is as follows: a method for generating digital transaction outbound call strategies based on multimodal large model recognition, comprising the following steps: S1. Acquire real-time interactive data during the digital transaction outbound call process through the multimodal data acquisition module; S2. Use a multimodal large model to extract and fuse features from real-time interactive data to generate a unified multimodal interactive feature vector; S3. Based on the historical outbound call strategy knowledge graph, extract key decision nodes from the multimodal interaction feature vector through graph neural network and generate strategy requirement feature vector; S4. Use a deep reinforcement learning model to evaluate the strategy requirement feature vector in multiple dimensions, including customer intent, risk level and conversion potential, and output the original strategy score. S5. Combine customer historical transaction information and expert strategy database to perform dynamic Bayesian calibration on strategy scores and generate strategy confidence scores. S6. Construct personalized outbound calling strategies based on strategy confidence, including script templates, recommended timing, and alternative solutions for handling objections; S7. Based on the execution feedback of personalized outbound calling strategies, the parameters of the multimodal large model and the strategy generation threshold are dynamically optimized through an online learning mechanism.

[0007] Preferably, the acquisition of real-time interactive data in S1 includes the following steps: S11. Acquire customer speech stream through microphone array and real-time speech recognition engine, extract acoustic features including pitch, speech rate and energy and semantic segments, and generate text dialogue records simultaneously. S12. Capture customer-authorized emotional facial expression images using a camera and facial motion coding system, and extract facial muscle movement units and micro-expression temporal features. S13. Capture screenshots of the transaction operation interface through the screen recording interface and optical character recognition module, identify the state of interface controls, operation focus position and input box content, and pack the voice stream, text and image into multimodal data frames according to timestamp alignment.

[0008] Preferably, the generation of the multimodal interaction feature vector in S2 includes the following steps: S21. Construct a multimodal large model, in which the visual encoder uses a lightweight residual network to process emotional facial expression images, the speech encoder uses a time-delay convolutional network to process speech streams, and the text encoder uses a Transformer to process dialogue records. S22. Visual, speech, and text features are weighted and fused through a cross-modal attention mechanism to generate multimodal interaction feature vectors and align them to the same embedding space. S23. A gated fusion network is used to dynamically combine the interaction feature vectors to output a fixed-dimensional multimodal interaction feature vector. The gate weights are calculated in real time by the current interaction context, and the inference latency of the gated network is controlled within 200 milliseconds.

[0009] Preferably, the generation of the strategy requirement feature vector in S3 includes the following steps: S31. Based on the semi-automatic annotation of historical call logs, and through active learning to reduce the amount of manual annotation, a knowledge graph of historical outbound call strategies is constructed. Nodes represent outbound call actions, customer responses and transaction stages, and edges represent outbound call scenario relationships and strategy dependency constraints. The success conversion rate and risk coefficient of nodes are also annotated. S32. Use graph attention network to align multimodal interaction feature vectors with knowledge graph node embeddings, and use node activation propagation algorithm to identify key decision nodes in the current interaction, including objection trigger points, hesitation signals and transaction thresholds. S33. Extract strategy requirement parameters based on key decision nodes, including customer urgency, anti-interference ability and preference offset, and concatenate them into a strategy requirement feature vector.

[0010] Preferably, the output of the original policy score in S4 includes the following steps: S41. Design a deep Q-network as the evaluation model, where the state space is the strategy demand feature vector, the action space is the candidate strategy action, and the reward function integrates customer intention gain, risk aversion term and conversion potential prediction. S42. Input the strategy demand feature vector into the deep Q network, calculate the Q value of each candidate action, and output the customer intention score, risk level score and conversion potential score through the soft maximization function. S43. The network parameters are updated by sampling training samples from historical interaction records using a priority experience replay mechanism, with the goal of minimizing temporal difference error, and the weighted fusion of the original policy scores is output.

[0011] Preferably, generating policy confidence in step S5 includes the following steps: S51. Establish an expert strategy library to store the best strategies labeled by experts in different trading scenarios and their corresponding rating records, and associate them with scenario complexity coefficients and customer type weights to form a prior distribution. S52. Retrieve customer historical transaction information and extract recent transaction frequency, average response time and strategy adoption rate as dynamic prior parameters; S53. Use Bayesian linear regression to probabilistically fuse the policy scores with the prior distribution in the expert policy library, and adjust the policy confidence using dynamic prior parameters as regularization factors.

[0012] Preferably, the probabilistic fusion in S53 includes the following steps: S531. Parameterize the prior distribution in the expert strategy library into a normal distribution, with its mean being the weighted average of expert scores and its variance being calculated jointly by the scenario complexity coefficient and the customer type weight. S532. Using the strategy score as the observed value, calculate the posterior distribution using the Bayesian linear regression formula: The observation noise variance is dynamically estimated from the recent transaction frequency and average response time in the dynamic prior parameters; S533. Take the mean of the posterior distribution as the adjusted policy confidence, and use the policy adoption rate in the dynamic prior parameters as the regularization factor to scale and correct the confidence before outputting it.

[0013] Preferably, the personalized outbound calling strategy constructed in step S6 includes the following steps: S61. Based on strategy confidence, a decision tree classification model is used to identify the strategy type of the current interaction scenario, including closing, appeasing, and information confirmation, and to match the corresponding script template library. S62. Analyze the pause patterns, speech rate changes and available emotional features in the customer's voice stream through a time-series prediction model, including the timing of voice emotions and facial expressions, estimate the best timing for script delivery and waiting intervals, and generate recommendation timing tags. S63. Generate alternative solutions using a counterfactual reasoning engine. The counterfactual reasoning engine is built on a structural causal model. When the customer's intentions do not meet expectations after the main strategy is executed, it automatically switches to the objection response sub-strategy and outputs the alternative solutions in order of confidence. The structural causal model learns from intervention records in historical dialogue data.

[0014] Preferably, the step S63, which uses a counterfactual reasoning engine to generate alternative solutions, includes the following steps: S631. Construct a structural causal model, in which variables include customer emotions, speech rate, operation delay, agent script type and transaction result. Learn the causal graph structure and conditional probability distribution through intervention records in historical dialogue data. S632. When the customer's intention does not meet expectations after the main strategy is executed, the current observation status is used as the fact input. The agent's script type variable is changed through the counterfactual reasoning algorithm, while keeping other exogenous variables unchanged, and the alternative script sequence is calculated in reverse. S633. Sort the generated alternative solutions in descending order of the probability of transaction in the counterfactual results, and select the top three objection response sub-strategies with the highest confidence to output.

[0015] Preferably, the step S7, which dynamically optimizes model parameters and thresholds through an online learning mechanism, includes the following steps: S71. Construct an online learning pipeline. After each outbound call, the execution feedback is used as a reward signal. The execution feedback includes whether the customer has made a purchase, the call duration, and the resolution of objections. Update the parameters of the deep reinforcement learning model through the policy gradient algorithm. S72. Set an adaptive threshold adjuster to statistically analyze the confidence distribution of the comprehensive strategy for the most recent N outbound calls using a sliding window. When the mean and variance of the distribution exceed the control limits, the strategy generation threshold is automatically adjusted. The threshold includes the intention threshold and the risk tolerance threshold. S73. The updated model parameters and thresholds are synchronously pushed to the multimodal large model and deep reinforcement learning evaluation module.

[0016] The present invention has the following beneficial effects: 1. In this invention, a multimodal data acquisition module simultaneously acquires customer voice streams, authorized emotional facial expression images, and screenshots of the transaction operation interface. By utilizing a cross-modal attention mechanism and a gating fusion network to perform weighted fusion and dynamic combination of visual, voice, and text features, real-time perception and deep alignment of multi-source heterogeneous information such as hesitation signals in customer micro-expressions, emotional fluctuations in voice, and decision delays in interface operations are achieved. This solves the problem of delayed customer intention judgment and high misjudgment rate caused by information blind spots in traditional single-modal analysis, and improves the accuracy and response speed of capturing the true intention of customers in outbound call interactions.

[0017] 2. In this invention, a graph attention network based on a historical outbound call strategy knowledge graph is used to identify key decision nodes such as objection trigger points and transaction thresholds. A deep Q-network is combined to evaluate the strategy requirement feature vector from multiple dimensions, including customer intent, risk level, and conversion potential. At the same time, a counterfactual reasoning engine is introduced to dynamically generate alternative dialogue sequences based on a structural causal model when the main strategy fails. This realizes a complete reasoning chain from feature perception to strategy evaluation to causal error correction. It solves the problems of rigid strategy and weak error correction ability caused by the lack of structured knowledge guidance and causal reasoning ability in existing decision trees or fixed rule methods, and enhances the adaptability and robustness of the outbound call system to complex transaction scenarios.

[0018] 3. In this invention, the strategy score is probabilistically fused with the prior distribution in the expert strategy library through dynamic Bayesian calibration. The confidence level is adjusted using recent transaction frequency, average response time, and strategy adoption rate from the customer's historical transaction information as dynamic prior parameters. Simultaneously, based on execution feedback, the model parameters and strategy generation threshold are optimized online through a strategy gradient algorithm and an adaptive threshold adjuster, forming a closed-loop adaptive optimization mechanism of execution, feedback, and update. This solves the problems of low personalization and insufficient continuous optimization capability caused by the inability of traditional static strategy models to dynamically adjust according to individual customer characteristics and real-time feedback. This improves transaction conversion rate, shortens call duration, and reduces risk exposure. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a digital transaction outbound call strategy generation method based on multimodal large model recognition according to the present invention. Detailed Implementation

[0020] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.

[0021] Example 1, please refer to Figure 1 As shown: A method for generating digital transaction outbound call strategies based on multimodal large model recognition, comprising the following steps: Multimodal perception and fusion: S1. Acquire real-time interactive data during the digital transaction outbound call process through the multimodal data acquisition module; S2. Use a multimodal large model to extract and fuse features from real-time interactive data to generate a unified multimodal interactive feature vector; Knowledge Reasoning and Assessment: S3. Based on the historical outbound call strategy knowledge graph, extract key decision nodes from the multimodal interaction feature vector through graph neural network and generate strategy requirement feature vector; S4. Use a deep reinforcement learning model to evaluate the strategy requirement feature vector in multiple dimensions, including customer intent, risk level and conversion potential, and output the original strategy score. Policy generation and adaptive optimization: S5. Combine customer historical transaction information and expert strategy database to perform dynamic Bayesian calibration on strategy scores and generate strategy confidence scores. S6. Construct personalized outbound calling strategies based on strategy confidence, including script templates, recommended timing, and alternative solutions for handling objections; S7. Based on the execution feedback of personalized outbound calling strategies, the parameters of the multimodal large model and the strategy generation threshold are dynamically optimized through an online learning mechanism.

[0022] The acquisition of real-time interactive data in S1 includes the following steps: S11. Acquire customer speech streams through microphone arrays and real-time speech recognition engines, extract acoustic features including pitch, speech rate and energy, and semantic segments, and simultaneously generate text dialogue records. In specific implementation: The system uses a single microphone built into the client terminal to collect speech signals. After transmission over the network, the server runs a deep learning-based single-channel speech enhancement model (such as RNNoise) for noise suppression. The enhanced speech stream is then input into a real-time speech recognition engine. This engine, based on an end-to-end Conformer model, extracts 80-dimensional log-Mel spectrum features with a frame length of 25 milliseconds and a frame shift of 10 milliseconds. It outputs the corresponding text frame by frame and concatenates the text dialogue records according to sentence boundaries. Simultaneously, it extracts three types of acoustic features: Pitch: Fundamental frequency extracted using autocorrelation method Calculate the autocorrelation function of the speech frame. ,in For the first speech signal One sampling point, Let the delay points be the number of points to make Maximum and of The value, and its corresponding frequency, is the fundamental frequency. .

[0023] Speech rate: Calculated by dividing the total number of recognized words by the duration of the speech, in words per second.

[0024] Energy: Short-time root-mean-square energy is used. ,in This represents the number of intra-frame sampling points. For the first frame The amplitude of each sampling point.

[0025] Semantic fragments: Extracting pre-defined keywords (such as interest rate, term, refund) and their dependencies from the identified text to form structured semantic units.

[0026] S12. Capture customer-authorized emotional facial expression images using a camera and facial motion coding system, extract facial muscle movement units and micro-expression temporal features. Specifically: If the client terminal supports and authorizes video capture, an RGB camera with a resolution of at least 1280×720 and a frame rate of 30fps is used to continuously capture facial images of the client. After face detection and alignment, the facial region is cropped. Based on the Facial Action Coding System (FACS), a pre-trained convolutional neural network is used to locate 68 facial key points and identify the activation intensity of 12 core motion units (such as AU1 inner eyebrow lifting, AU4 frowning, and AU12 corner of the mouth lifting), with intensity values ​​mapped to the 0–5 range. Micro-expression temporal features are captured by calculating the optical flow field between adjacent frames, and the optical flow constraint equation is solved using the Lucas-Kanade optical flow method. ; in , The images are respectively in , Spatial gradient of direction, For time gradient, , Each pixel is located at , The velocity components in the direction are statistically analyzed for the velocity vectors in the neighborhood of each key point. The start and end times, maximum displacement amplitude, and duration of micro-expressions are extracted, ultimately forming the facial motion unit vector and micro-expression temporal sequence for each frame.

[0027] S13. Capture screenshots of the transaction operation interface through the screen recording interface and optical character recognition module, identify the state of interface controls, the focus position of the operation, and the content of the input box, and package the audio stream, text, and images into a multimodal data frame according to the timestamp. In specific implementation: User operation events and interface states are collected through a client-authorized front-end SDK (such as a web JavaScript SDK or mobile SDK). The DOM element tree or control tree of the current page is obtained at a sampling frequency of 5 times per second. The state of controls such as buttons, input boxes, and checkboxes (such as enabled / disabled, selected / unselected, focus state) is identified. The coordinates or identifier of the currently active control are obtained by listening to the page focus event. The text entered by the user is obtained by listening to the input event. If the client additionally authorizes screen recording permission, the interface screenshot can be obtained at 5 frames per second through the operating system screen capture interface (such as Windows DWM or Android Media Projection). The text area in the screenshot is identified by a lightweight OCR engine (CRNN+CTC). Alignment operation: Based on the system clock, each frame of the speech stream, each sentence of the text dialogue, each frame of the image, and each record captured by the interface are timestamped at the millisecond level. Data with inconsistent sampling rates are synchronized by nearest neighbor or linear interpolation to form multimodal data frames sorted by time. Each frame contains: timestamp, speech feature vector, text paragraph, facial motion unit set, interface control state, and focus coordinates.

[0028] The generation of multimodal interaction feature vectors in S2 includes the following steps: S21. Construct a multimodal large model, in which the visual encoder uses a lightweight residual network to process emotional facial expression images, the speech encoder uses a time-delayed convolutional network to process speech streams, and the text encoder uses a Transformer to process dialogue records. In specific implementation: The visual encoder uses ResNet-18, with a cropped and aligned 96×96 pixel grayscale face image as input. After passing through 5 residual blocks, it outputs a 512-dimensional visual feature vector. The speech encoder employs a temporal delay convolutional network (TDNN). The input consists of 80-dimensional log-Mel spectrum features per frame, stacked with three layers of time delay convolutional layers. The context windows for each layer are [-2, 2], [-1, 1], and [0, 0], respectively. The output is a 512-dimensional speech feature vector. The text encoder employs a 6-layer Transformer architecture. The input is a segmented dialogue text sequence (maximum length 128). Through positional encoding and multi-head self-attention, it outputs a 512-dimensional text feature vector. The three encoders process their respective modal data in parallel.

[0029] S22. Visual, speech, and text features are weighted and fused using a cross-modal attention mechanism to generate a multimodal interaction feature vector, which is then aligned to the same embedding space. Specifically: Text features As a query, visual features and speech features Using these as keys and values ​​respectively, we calculate cross-modal attention for text-visual and text-speech interactions; taking text-visual attention as an example, we calculate the attention weight matrix. ,in , It is a learnable linear projection matrix. The projection dimension is 64. Indicates transpose, outputs the weighted visual context vector. , Similarly, the speech context vector is obtained by using the projection matrix. Then, the original text features are concatenated with the two context vectors and mapped to the same embedding space through a linear layer to generate an aligned multimodal interaction feature vector. ,in This indicates vector concatenation, with an output dimension of 512.

[0030] S23. A gated fusion network is used to dynamically combine the interaction feature vectors, outputting a fixed-dimensional multimodal interaction feature vector. The gate weights are calculated in real time by the current interaction context, and the inference latency of the gated network is controlled within 200 milliseconds. In specific implementation: The gated fusion network receives the multimodal interaction feature vector output in step S22. The current interaction context features (including customer sentiment scores in the last 3 frames, speech rate changes in the last 5 seconds, and screen dwell time) are concatenated and input into a two-layer fully connected network. The first layer outputs 128 dimensions, and the second layer outputs 3 dimensions. The gate weights for the three modalities are obtained through the Softmax function. Naturally satisfied If the visual modality is missing due to customer non-authorization or device incompatibility, the visual encoder output will be... Set to zero vector, only the text-to-speech part is calculated in cross-modal attention, and the gating weights are set to... and to , Renormalization: ; ; Final multimodal interaction feature vector (Use when missing) , Alternative , ),in ,a, The outputs are the original feature vectors from the three encoders in step S21. All network operations use GPU-accelerated inference. The processing time per frame is monitored by a CUDA event timer, and in actual tests, it is controlled within 200 milliseconds. As input for the next stage.

[0031] Example 2: The generation of the strategy requirement feature vector in S3 includes the following steps: S31. Based on semi-automatic annotation of historical call logs, and through active learning to reduce manual annotation, a historical outbound call strategy knowledge graph is constructed. Nodes represent outbound call actions, customer responses, and transaction stages, while edges represent outbound call scenario relationships and strategy dependency constraints. The success conversion rate and risk coefficient of each node are also annotated. In specific implementation: Call logs (including agent dialogue sequences, customer responses, and transaction results) are exported from the historical outbound call system. A rule engine is used to automatically extract three types of candidate nodes: outbound action nodes (e.g., introducing products, closing contracts), customer reaction nodes (e.g., silence, questioning, agreement), and transaction stage nodes (e.g., opening remarks, objection handling, closing). The active learning process involves randomly selecting 10% of the logs, having experts annotate the nodes and edges, and training an initial graph neural network (e.g., GCN or GAT). Uncertainty scores (entropy) are calculated for the remaining unannotated samples. ,in Predicting the category to which a sample belongs for the model The probability of success is calculated; the top 5% of samples with the highest entropy are selected and labeled by experts, and the process is iterated for 3 rounds until the model converges; the final knowledge graph is a directed graph. ,node With attribute: Success conversion rate (The final transaction percentage after the action corresponding to this node), risk factor (Based on historical complaint rates normalized to 0-1), edge With constraint types (such as must precede, can be skipped, mutually exclusive).

[0032] S32. Use a graph attention network to align the multimodal interaction feature vectors with the knowledge graph node embeddings, and use a node activation propagation algorithm to identify key decision nodes in the current interaction, including objection trigger points, hesitation signals, and transaction thresholds. In specific implementation: Initial embedding of knowledge graph nodes Input a graph attention network (GAT), and for node i, compute the attention coefficients of its neighbor node j. ,in For attention weight vectors, It is a linear transformation matrix. This indicates splicing; after softmax normalization, we get... Update node embedding , The activation function is used; two layers of GAT are stacked to obtain the final node embedding, which is then used to generate the multimodal interaction feature vector. Linear projection to 128 dimensions is performed, and cosine similarity is calculated with node embeddings. The five nodes with the highest similarity are selected as seed nodes, and node activation propagation is performed: starting from the seed nodes, the propagation proceeds along the edges according to weights. Breadth-first propagation, decay factor Node activation score ; in Seed node similarity, The shortest path length. The score is calculated as the product of the weights of the edges on the path. Nodes with scores exceeding the threshold of 0.7 are identified as key decision nodes: objection trigger points (customer questioning reaction nodes), hesitation signals (long silence or repeated confirmation nodes), and closing thresholds (closing action nodes).

[0033] S33. Extract strategy requirement parameters based on key decision nodes, including customer urgency, anti-interference ability, and preference offset, and concatenate them into a strategy requirement feature vector. In specific implementation: Three types of parameters are extracted from the identified key decision nodes and their real-time interaction data: Customer urgency ,in The number of times the customer urged them to use certain phrases in the last 30 seconds. This represents the number of seconds remaining in the average transaction time window. Prevent division by zero.

[0034] Anti-interference capability ,in The time interval (in seconds) from the point of objection triggering to the disappearance of the hesitation signal. This represents the normalized standard deviation of the emotional characteristics (vocal energy or facial AU intensity) during this period.

[0035] Preference offset ,in Attribute embedding for the current key decision node (generated by GAT). The preferences exhibited by this customer in historical transactions are embedded (obtained by averaging the customer's historical paths in the knowledge graph); the above three parameters are concatenated with the node type one-hot encoding: the node type is divided into four categories: objection trigger point, hesitation signal, transaction threshold point and others, encoded as a 4-dimensional vector, finally yielding a 7-dimensional strategy demand feature vector s. The output is sent to the deep reinforcement learning module.

[0036] The steps involved in outputting the original policy score in S4 are as follows: S41. Design a deep Q-network as the evaluation model, where the state space is the policy demand feature vector, the action space is the candidate policy actions, and the reward function integrates customer intention gain, risk aversion term, and conversion potential prediction. In specific implementation: Deep Q-networks employ a three-layer fully connected network, with the input layer receiving the feature vector required by the strategy. The hidden layers are 128-dimensional and 64-dimensional respectively, the activation function is ReLU, and the output layer dimension is equal to the number of candidate actions. The actions include product introduction, price inquiry, closing the contract, resolving objections, and follow-up. The state space is a 7-dimensional vector output from step S33. The reward function is defined as: ; in , and These are the customer intent scores calculated by S42 before and after the action was performed; Risk coefficient corresponding to the current action (taken from knowledge graph node attributes) ); This represents the probability of a transaction being completed under similar conditions based on historical statistics; a balance coefficient is used. , .

[0037] S42. Input the strategy demand feature vector into a deep Q-network, calculate the Q-value of each candidate action, and output the customer intention score, risk level score, and conversion potential score through a soft maximization function. In specific implementation: Will Inputting the data into a deep Q-network yields each action. of value Three mutually exclusive subsets of actions are predefined: Action Set for Enhancing Intention ; High-risk action set ; High conversion potential action set .

[0038] The three scores are calculated as follows: Customer Intent Score: , which is the sum of the Softmax probability mass of the intended action, with a value range of (0, 1), and a higher value indicates a stronger customer intention.

[0039] Risk level score: Similarly, it is the sum of the probability quality of high-risk actions, with a value range of (0, 1). The higher the value, the greater the risk of the current strategy.

[0040] Conversion potential score: , which is the maximum Softmax probability in the high conversion potential action set, with a value range of (0, 1), representing the expected conversion potential of the optimal action.

[0041] All three scores output scalar values ​​between 0 and 1, which are used by the reward function and subsequent scoring.

[0042] S43. Employ a priority experience replay mechanism to sample training samples from historical interaction records, update network parameters with the goal of minimizing temporal difference error, and output the weighted fusion of the original policy scores. Specifically, in implementation: Maintain an experience replay buffer and store tuples after each interaction. Prioritize experience playback based on the absolute value of timing difference error. Assign sampling probabilities, where the discount factor , For the target network parameters, calculate the mean squared error loss based on a probability sampling batch (32 samples): ; Update the main network parameters using the Adam optimizer, with a soft update rate every 100 steps. Copy the parameters of the main network to the target network. After training converges, analyze the current state. Output the original policy score: ; Among them, weight , , , The risk score is converted into a safety score, and the final score is a scalar between 0 and 1, which is used as the original policy score output.

[0043] Example 3: Generating policy confidence in S5 includes the following steps: S51. Establish an expert strategy library to store the best strategies labeled by experts and their corresponding rating records under different trading scenarios, and associate them with scenario complexity coefficients and customer type weights to form a prior distribution. In specific implementation: Ten senior call center experts with over 5 years of experience and a historical prediction accuracy rate exceeding 80% were selected to annotate strategies for 100 typical trading scenarios (categorized by transaction amount, product type, and customer age group). For each scenario, the best strategy was selected using the Delphi method (multiple rounds of anonymous voting), and a score from 0 to 100 was assigned. Two coefficients were also assigned to each scenario. Scene complexity coefficient The higher the complexity, the better, as the number of decision branches and the density of product terminology are normalized. The larger.

[0044] Customer type weight : Set based on customers' historical spending power and credit rating, high-value customers Relatively large.

[0045] Treating expert ratings as priors that follow a normal distribution, calculate the weighted mean. ,in For the number of experts, For the first The authority weight of each expert (based on their historical prediction accuracy normalized to 0.5-1.5). The prior variance is calculated as the expert's strategy score for the scenario. ; The second term in the formula (Pick This makes the prior distribution of complex scenarios or high-weight customers more dispersed (increasing variance), and the final prior distribution of each scenario is represented as follows: Stored in the expert strategy database.

[0046] S52. Retrieve customer historical transaction information and extract recent transaction frequency, average response time, and strategy adoption rate as dynamic prior parameters. In specific implementation: Retrieve the customer's transaction history for the past 90 days from the customer relationship management system, defining the following three dynamic prior parameters: Recent trading frequency (times / day), of which This represents the total number of transactions within 90 days.

[0047] Average response time (seconds), in the formula This refers to the number of seconds between the customer's question and the agent's response during each outbound call. This represents the total number of interaction rounds in the last 90 days.

[0048] Strategy adoption rate ,in The number of times a customer accepts a seat suggestion. This represents the total number of outbound calls (range 0-1).

[0049] The dynamic prior parameter vector is This is used for estimation of observation noise variance and regularization scaling in subsequent Bayesian calibration.

[0050] S53. Use Bayesian linear regression to probabilistically fuse the policy scores with the prior distribution in the expert policy library, and adjust the policy confidence using dynamic prior parameters as regularization factors. In specific implementation: S531. Parameterize the prior distribution in the expert strategy base into a normal distribution, with its mean being the weighted average of expert scores, and its variance being calculated jointly by the scenario complexity coefficient and the customer type weight. In specific implementation: For the current transaction scenario, match the closest scenario template based on scenario characteristics (transaction amount, product type, customer age group) to obtain the corresponding prior distribution. ,in and The calculation method is the same as that of S51 (weighted mean plus variance of complexity term). This prior distribution reflects the prior knowledge of domain experts about the strategy score in the current scenario.

[0051] S532. Using the strategy score as the observed value, calculate the posterior distribution using the Bayesian linear regression formula: The observation noise variance is dynamically estimated from the recent transaction frequency and average response time in the dynamic prior parameters. In practice: The original policy score output in step S4 is recorded as... (Scalar between 0 and 1), observation noise variance Recent transaction frequency from dynamic prior parameters and average response time Dynamic estimation: ; In the formula The hyperbolic tangent function slightly reduces the noise variance for high-frequency trading clients (because their behavior is more predictable). Ensure that the variance is not less than 0.01.

[0052] Calculate the mean of the posterior distribution using Bayesian linear regression (normal-normal conjugate): ; in The prior mean, For the prior variance, Score the original strategy. To observe the noise variance, the posterior standard deviation is: ; Take the posterior mean As a preliminary strategy confidence level.

[0053] S533. Take the mean of the posterior distribution as the adjusted policy confidence score, and use the policy adoption rate in the dynamic prior parameters as a regularization factor to scale and correct the confidence score before outputting it. In specific implementation: Policy adoption rate in dynamic prior parameters As a regularization factor, combined with recent trading frequency Calculate the scaling factor: ; In the formula Using the natural logarithm to smooth out the impact of trading frequency, the final overall strategy confidence level is: ; By pruning, ensure the output is in the 0-1 range. The output is used in step S6 to build a personalized outbound calling strategy.

[0054] Building a personalized outbound calling strategy in S6 includes the following steps: S61. Based on strategy confidence, a decision tree classification model is used to identify the strategy type of the current interaction scenario, including closing, reassurance, and information confirmation, and to match the corresponding script template library. In specific implementation: The overall policy confidence level output by S5 Step S42 outputs the customer intention score and risk level score Using these features as input, a CART decision tree classification model (depth of 4, splitting criterion: Gini impurity) is constructed. , (The proportion of samples belonging to the i-th class of strategies), example of classification rules: if and Then it is judged as a sales-boosting type; if or If the result is positive, it is classified as a reassurance type; otherwise, it is classified as an information confirmation type. Based on the identified strategy type, the corresponding template is matched from the script template library (each script includes an opening sentence, a core sentence, a closing sentence, and replaceable variables such as product name and amount), and the variables are filled in according to the customer's historical preferences (such as product categories that appear in past purchase records) to output personalized scripts.

[0055] S62. Analyze the pause patterns, speech rate changes, and available emotional features in the customer's voice stream using a time-series prediction model, including the timing of vocal emotions and facial expressions, to estimate the optimal timing for script delivery and waiting intervals, and generate recommendation timing tags. In specific implementation: A Bidirectional Long Short-Term Memory (BiLSTM) network is used as the temporal prediction model. The input consists of multimodal features from each frame within the last 10 seconds: speech rate. (Words per second, taken from S11), Short-time energy (Taken from S11), mean activation intensity of facial motor units (Only when video is available; if the customer has not authorized video, then...) Setting to 0), probability of voice emotion classification (Extracted from acoustic features using a speech emotion classifier); the model outputs two values: optimal timing for broadcasting. (Number of seconds until the current time) and waiting interval (Number of seconds the agent should wait); Specific calculation: BiLSTM extracts the timing hidden state. The output broadcast confidence sequence is obtained through a fully connected layer. Take the time corresponding to the peak as The formula for estimating the waiting interval is: ; in The second is the base interval. Emotional similarity is defined as the current fused emotional value. (Based on facial and voice data, range 0-1) How close to a neutral mood score of 0.5: ; when When it approaches 0.5 Approaching 1, wait for the interval to decrease; if the visual modality is missing, The system outputs only from the voice emotion model and ultimately generates recommendation timing tags (such as immediate broadcast, wait 1.5 seconds, etc.) for use by the outbound calling system.

[0056] S63. Utilize a counterfactual reasoning engine to generate alternative solutions. The counterfactual reasoning engine is built upon a structural causal model. When the customer's intentions do not meet expectations after the main strategy is executed, it automatically switches to an objection response sub-strategy and outputs alternative solutions sorted by confidence level. The structural causal model learns from intervention records in historical dialogue data. In specific implementation: The steps involved in generating alternative solutions using the counterfactual reasoning engine in S63 are as follows: S631. Construct a structural causal model, where variables include customer emotion, speech rate, operational delay, agent script type, and transaction outcome. Learn the causal graph structure and conditional probability distribution through intervention records in historical dialogue data. In specific implementation: Based on the time-series characteristics output by S62 and the strategy type determined by S61, a structural causal model (SCM) containing five endogenous variables is constructed: customer sentiment. (Integrating voice and facial expressions, in real time), speech rate (Words / second, current frame), operation latency (Seconds are defined as the duration of the customer's response to the previous dialogue in the current interaction, i.e., the interval from when the agent ends asking a question to when the customer begins to answer), Agent dialogue type (Corresponding to 9 actions) and results of changes in intention (Continuous scoring, compared with customer intent score in S42) (Alignment), the exogenous variable is customer personality tendency. and time-of-day interference Utilizing natural intervention records in historical dialogue data (the actual types of dialogue used by agents) Learning causal structures: Greedy Equivalence Search (GES) and BIC scoring are used to obtain a directed acyclic graph, and relationships are determined as follows: , , For example, assuming a linear Gaussian model for continuous variables, such as: ; Parameters are estimated using maximum likelihood estimation. , , , For discrete variables Historical frequency was statistically analyzed, and the distribution of exogenous variables was estimated using residual analysis. If facial expressions were unavailable, then... The causal relationship remains unchanged even when output by the speech model alone.

[0057] S632. When the customer's intention does not meet expectations after the main strategy is executed, the current observed state is used as the factual input. The agent's script type variable is changed through a counterfactual reasoning algorithm, while keeping other exogenous variables unchanged. The alternative script sequence is calculated in reverse. In specific implementation: After the main strategy is implemented, changes in customer intent will be monitored in real time. ,like If the expected result is not achieved, then the current observed state will be taken as a fact: , , Actual sales pitch Intended outcome Counterfactual steps: First, infer the exogenous variable from the factual state. , (For example, obtaining residuals by solving the structure equations) That is to be regarded as , (the embodiment of) , Leave the dialogue type variable unchanged and change it to another candidate value. Update the affected variables in causal order: according to Counterfactual delay ,in This is the factual residual.

[0058] according to Counterfactual sentiment .

[0059] Update transaction results .

[0060] Calculate change in intention For all After calculation, select The action, according to Arrange in descending order to generate a sequence of candidate dialogues (each action corresponds to a pre-stored dialogue template).

[0061] S633. Sort the generated multiple alternative solutions in descending order of the probability of transaction in the counterfactual results, and select the top three objection response sub-strategies with the highest confidence levels for output. In specific implementation: For each alternative action Calculate the probability of a transaction. (or use directly) The variance was estimated using the MCDropout method: during inference, the same input was subjected to 10 random forward propagations (with the Dropout layer retained) to obtain 10 transaction results. Calculate the sample variance The confidence score is then: ; (make sure Normalized to [0, 1]), sorted in descending order of confidence, the top three are taken as the sub-strategies for objection response; each scheme includes the type of speech, recommended content and confidence score (e.g.: Option 1: provide an instant discount (confidence 0.92); Option 2: transfer to a senior agent (confidence 0.85); Option 3: change the explanation angle (confidence 0.76)), for use by agents or automated execution systems.

[0062] In S7, dynamically optimizing model parameters and thresholds through an online learning mechanism includes the following steps: S71. Construct an online learning pipeline. After each outbound call, the execution feedback is used as a reward signal. The execution feedback includes whether the customer made a purchase, the call duration, and the resolution of objections. Update the parameters of the deep reinforcement learning model through a policy gradient algorithm. In specific implementation: Receive execution feedback after each outbound call: Transaction confirmation flag Call duration (seconds), objection resolution mark The comprehensive reward signal is: ; A reward of 10 is given for a successful transaction, and 2 is given for resolving an objection. After 120 seconds, 0.01 will be deducted per second.

[0063] Adopting an Actor-Critic architecture: Actor network Output action probabilities, Critic network Output state value, sharing the feature extraction layer of the S41 deep Q network, for a single outbound call trajectory Calculate cumulative discount return Discount factor Advantage function Actor parameter update gradient: ; Minimize Critic parameter updates The Adam optimizer is used with a learning rate of 0.001; gradient updates are performed after each outbound call to achieve real-time online learning, and the updated parameters... , Stored in the parameter server.

[0064] S72. Set an adaptive threshold adjuster to statistically analyze the confidence distribution of the comprehensive strategy for the most recent N outbound calls using a sliding window. When the mean or variance of the distribution exceeds the control limits, automatically adjust the strategy-generated threshold. The thresholds include the intention threshold and the risk tolerance threshold. In specific implementation: Maintenance size is A sliding window records the overall strategy confidence level of the most recent N outbound calls. (From S5) and transaction results Calculate the built-in confidence mean in the window. and standard deviation Control limits: , , .

[0065] Intention threshold (Initial 0.6, used for S61 decision tree judgment to promote orders): When And in the last 20 transactions with a transaction rate > 0.8, ;when And when the transaction rate is less than 0.3, .

[0066] Risk tolerance threshold (Initial 0.7, used by S42 to judge high-risk actions): When hour, (Tighten); when And when the transaction rate is stable (transaction rate fluctuation <0.1 for 20 consecutive times), (Relaxed), the adjusted threshold is written to the configuration center.

[0067] S73. Synchronously push the updated model parameters and thresholds to the multimodal large model and deep reinforcement learning evaluation module. Specifically: After each outbound call is completed, the parameters updated in S71 are sent via Remote Procedure Call (RPC). and the threshold adjusted by S72 Asynchronous push to all online service nodes.

[0068] The deep reinforcement learning evaluation module (S4) receives new parameters, which are directly used for action selection and value evaluation in subsequent states.

[0069] Multimodal large model (S2) receive threshold updates: utilizing Adjust the gating fusion weights (when) (Increase text modal weights at the same time); utilize Control the use of visual features (if) (Then the visual features are forced to be set to zero). In addition, every 1,000 outbound calls, the cumulative feedback data is used to efficiently fine-tune the encoder parameters in the multimodal large model (PEFT, only 5% of the parameters are updated) to adapt to changes in customer behavior distribution. The push delay is controlled within 50 milliseconds to ensure that online services are not interrupted. This forms a closed-loop adaptive optimization system of execution, feedback, update and then execution again.

[0070] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for generating digital transaction outbound call strategies based on multimodal large model recognition, characterized in that, Includes the following steps: S1. Acquire real-time interactive data during the digital transaction outbound call process through the multimodal data acquisition module; S2. Use a multimodal large model to extract and fuse features from real-time interactive data to generate a unified multimodal interactive feature vector; S3. Based on the historical outbound call strategy knowledge graph, extract key decision nodes from the multimodal interaction feature vector through graph neural network and generate strategy requirement feature vector; S4. Use a deep reinforcement learning model to evaluate the strategy requirement feature vector in multiple dimensions, including customer intent, risk level and conversion potential, and output the original strategy score. S5. Combine customer historical transaction information and expert strategy database to perform dynamic Bayesian calibration on strategy scores and generate strategy confidence scores. S6. Construct personalized outbound calling strategies based on strategy confidence, including script templates, recommended timing, and alternative solutions for handling objections; S7. Based on the execution feedback of personalized outbound calling strategies, the parameters of the multimodal large model and the strategy generation threshold are dynamically optimized through an online learning mechanism.

2. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 1, characterized in that, The acquisition of real-time interactive data in S1 includes the following steps: S11. Acquire customer speech stream through microphone array and real-time speech recognition engine, extract acoustic features including pitch, speech rate and energy and semantic segments, and generate text dialogue records simultaneously. S12. Capture customer-authorized emotional facial expression images using a camera and facial motion coding system, and extract facial muscle movement units and micro-expression temporal features. S13. Capture screenshots of the transaction operation interface through the screen recording interface and optical character recognition module, identify the state of interface controls, operation focus position and input box content, and pack the voice stream, text and image into multimodal data frames according to timestamp alignment.

3. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 1, characterized in that, The generation of the multimodal interaction feature vector in S2 includes the following steps: S21. Construct a multimodal large model, in which the visual encoder uses a lightweight residual network to process emotional facial expression images, the speech encoder uses a time-delay convolutional network to process speech streams, and the text encoder uses a Transformer to process dialogue records. S22. Visual, speech, and text features are weighted and fused through a cross-modal attention mechanism to generate multimodal interaction feature vectors and align them to the same embedding space. S23. A gated fusion network is used to dynamically combine the interaction feature vectors to output a fixed-dimensional multimodal interaction feature vector. The gate weights are calculated in real time by the current interaction context, and the inference latency of the gated network is controlled within 200 milliseconds.

4. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 1, characterized in that, The generation of the strategy requirement feature vector in S3 includes the following steps: S31. Based on the semi-automatic annotation of historical call logs, and through active learning to reduce the amount of manual annotation, a knowledge graph of historical outbound call strategies is constructed. Nodes represent outbound call actions, customer responses and transaction stages, and edges represent outbound call scenario relationships and strategy dependency constraints. The success conversion rate and risk coefficient of nodes are also annotated. S32. Use graph attention network to align multimodal interaction feature vectors with knowledge graph node embeddings, and use node activation propagation algorithm to identify key decision nodes in the current interaction, including objection trigger points, hesitation signals and transaction thresholds. S33. Extract strategy requirement parameters based on key decision nodes, including customer urgency, anti-interference ability and preference offset, and concatenate them into a strategy requirement feature vector.

5. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 1, characterized in that, The output of the original policy score in S4 includes the following steps: S41. Design a deep Q-network as the evaluation model, where the state space is the strategy demand feature vector, the action space is the candidate strategy action, and the reward function integrates customer intention gain, risk aversion term and conversion potential prediction. S42. Input the strategy demand feature vector into the deep Q network, calculate the Q value of each candidate action, and output the customer intention score, risk level score and conversion potential score through the soft maximization function. S43. The network parameters are updated by sampling training samples from historical interaction records using a priority experience replay mechanism, with the goal of minimizing temporal difference error, and the weighted fusion of the original policy scores is output.

6. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 1, characterized in that, The generation of policy confidence in S5 includes the following steps: S51. Establish an expert strategy library to store the best strategies labeled by experts in different trading scenarios and their corresponding rating records, and associate them with scenario complexity coefficients and customer type weights to form a prior distribution. S52. Retrieve customer historical transaction information and extract recent transaction frequency, average response time and strategy adoption rate as dynamic prior parameters; S53. Use Bayesian linear regression to probabilistically fuse the policy scores with the prior distribution in the expert policy library, and adjust the policy confidence using dynamic prior parameters as regularization factors.

7. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 6, characterized in that, The probabilistic fusion in S53 includes the following steps: S531. Parameterize the prior distribution in the expert strategy library into a normal distribution, with its mean being the weighted average of expert scores and its variance being calculated jointly by the scenario complexity coefficient and the customer type weight. S532. Using the strategy score as the observed value, calculate the posterior distribution using the Bayesian linear regression formula: The observation noise variance is dynamically estimated from the recent transaction frequency and average response time in the dynamic prior parameters; S533. Take the mean of the posterior distribution as the adjusted policy confidence, and use the policy adoption rate in the dynamic prior parameters as the regularization factor to scale and correct the confidence before outputting it.

8. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 1, characterized in that, The process of constructing a personalized outbound calling strategy in S6 includes the following steps: S61. Based on strategy confidence, a decision tree classification model is used to identify the strategy type of the current interaction scenario, including closing, appeasing, and information confirmation, and to match the corresponding script template library. S62. Analyze the pause patterns, speech rate changes and available emotional features in the customer's voice stream through a time-series prediction model, including the timing of voice emotions and facial expressions, estimate the best timing for script delivery and waiting intervals, and generate recommendation timing tags. S63. Generate alternative solutions using a counterfactual reasoning engine. The counterfactual reasoning engine is built on a structural causal model. When the customer's intentions do not meet expectations after the main strategy is executed, it automatically switches to the objection response sub-strategy and outputs the alternative solutions in order of confidence. The structural causal model learns from intervention records in historical dialogue data.

9. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 8, characterized in that, The process of generating alternative solutions using a counterfactual reasoning engine in S63 includes the following steps: S631. Construct a structural causal model, in which variables include customer emotions, speech rate, operation delay, agent script type and transaction result. Learn the causal graph structure and conditional probability distribution through intervention records in historical dialogue data. S632. When the customer's intention does not meet expectations after the main strategy is executed, the current observation status is used as the fact input. The agent's script type variable is changed through the counterfactual reasoning algorithm, while keeping other exogenous variables unchanged, and the alternative script sequence is calculated in reverse. S633. Sort the generated alternative solutions in descending order of the probability of transaction in the counterfactual results, and select the top three objection response sub-strategies with the highest confidence to output.

10. The method for generating digital transaction outbound call strategies based on multimodal large model recognition according to claim 1, characterized in that, The S7 step of dynamically optimizing model parameters and thresholds through an online learning mechanism includes the following steps: S71. Construct an online learning pipeline. After each outbound call, the execution feedback is used as a reward signal. The execution feedback includes whether the customer has made a purchase, the call duration, and the resolution of objections. Update the parameters of the deep reinforcement learning model through the policy gradient algorithm. S72. Set an adaptive threshold adjuster to statistically analyze the confidence distribution of the comprehensive strategy for the most recent N outbound calls using a sliding window. When the mean and variance of the distribution exceed the control limits, the strategy generation threshold is automatically adjusted. The threshold includes the intention threshold and the risk tolerance threshold. S73. The updated model parameters and thresholds are synchronously pushed to the multimodal large model and deep reinforcement learning evaluation module.