Intelligent interaction method and system based on modal perception and edge collaboration

By employing multimodal affective computing and edge collaboration architecture, the system addresses the shortcomings of traditional intelligent customer service systems in terms of accuracy of emotional interaction, response efficiency, and cross-language adaptability, thereby achieving efficient and secure intelligent interaction.

CN120892145APending Publication Date: 2025-11-04湖北消费金融股份有限公司
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510817214.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Traditional intelligent customer service systems lack multimodal emotional integration capabilities, have low response efficiency, cannot adapt to dynamic business scenarios, and have insufficient cross-language support.

Method used

Employing a multimodal emotion computing and edge collaboration architecture, a comprehensive perception index is generated through a multimodal perception unit. Combined with a dynamic scene adaptation network and a federated learning scheduling algorithm, the accuracy of emotion interaction, response efficiency, and resource scheduling are optimized.

Benefits of technology

It improves the accuracy of emotional interaction, reduces response latency, supports multilingual scenarios, and ensures efficient use of edge resources and privacy security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892145A_ABST
    Figure CN120892145A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent interaction method and system based on modal perception and edge collaboration. The method comprises the following steps: obtaining interaction request information of a target user; generating modal sensing information from the interaction request information based on a multi-modal sensing unit; performing multi-head attention weighting on the modal perception information to obtain a comprehensive perception index, and generating a response strategy according to the comprehensive perception index; adapting a network based on a dynamic scene to dynamically optimize the response strategy according to the continuous multiple modal perception information and the corresponding response strategy; and in the sensing response interaction process, carrying out edge resource dynamic scheduling on the edge computing node comprising the multi-modal sensing unit based on a federated learning scheduling algorithm. Therefore, the technical bottlenecks of traditional customer service interaction in the aspects of emotion interaction accuracy, response efficiency and edge resource scheduling are solved through technical fusion of multi-modal emotion calculation, edge collaborative architecture, dynamic scene adaptation and an autonomous optimization mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and customer service, and particularly relates to an intelligent interaction method and system based on modal perception and edge collaboration. BACKGROUND

[0002] Traditional intelligent customer service systems generally adopt a response mechanism based on a rule engine or a pre-defined template, and the technical limitations mainly lie in the following aspects: 1. Lack of emotional interaction capability: Existing systems rely on a single text modality for intent recognition, lacking multi-dimensional perception capability for voice tone, facial micro-expression and context emotional evolution, resulting in inability to achieve humanized interaction.

[0003] 2. Response efficiency bottleneck: Centralized cloud computing architecture leads to interaction delay generally exceeding 200ms. The retrieval mechanism based on a fixed knowledge base is difficult to cope with dynamic business scenarios, such as real-time detection of knowledge conflicts caused by policy and regulation updates.

[0004] 3. Insufficient multi-language and scenario adaptability: Traditional intent recognition models cannot support real-time semantic alignment of multiple languages in cross-lingual scenarios.

[0005] To address the above problems, the existing technology uses BERT model to improve text understanding accuracy and introduces edge computing to reduce delay, but still does not solve the core problems of multi-modal emotional fusion and dynamic resource scheduling. Therefore, a technical fusion of emotional computing and edge collaboration is needed to construct an intelligent customer service solution with autonomous evolution capability. SUMMARY

[0006] The present application provides an intelligent interaction method and system based on modal perception and edge collaboration, which solves the technical bottlenecks of traditional customer service interaction in emotional interaction accuracy, response efficiency and edge resource scheduling through the technical fusion of multi-modal emotional computing, edge collaboration architecture, dynamic scenario adaptation and autonomous optimization mechanism. In a first aspect, an intelligent interaction method based on modal perception and edge collaboration is provided, comprising the following steps: Obtaining interaction request information of a target user; Generating modal perception information from the interaction request information based on a multi-modal perception unit; Performing multi-head attention weighting on the modal perception information to obtain a comprehensive perception index, and generating a response strategy according to the comprehensive perception index; Based on a dynamic scenario adaptation network, dynamically optimizing the response strategy according to continuous multiple modal perception information and corresponding response strategies; In the process of perceiving the response interaction, a federated learning scheduling algorithm is used to dynamically schedule edge resources of an edge computing node including a multi-modal perception unit.

[0007] In some embodiments, the step of generating modal perception information based on the multi-modal perception unit includes: The multi-modal perception unit includes a voice emotion analysis module, a facial micro-expression recognition module, and a text emotion classification module. The ECAPA-TDNN network of the voice emotion analysis module extracts target voiceprint features in the interaction request information and generates voice perception information. The HRNet network of the facial micro-expression recognition module tracks facial key point features in the interaction request information and generates AU intensity values. The BERT-Emotion network of the text emotion classification module performs semantic encoding on keyword sentences in the interaction request information to generate text emotion classification perception information.

[0008] In some embodiments, the method of multi-head attention weighting the modal perception information to obtain a comprehensive perception index is as follows: E total =αV voice +θE micro +λP text In the formula, E total is the comprehensive perception index; V voice is the voice perception information; E micro is the AU intensity value; P text is the text emotion classification perception information; and α, θ, and λ are weight coefficients.

[0009] In some embodiments, the step of generating a response strategy based on the comprehensive perception index includes: When it is detected that the comprehensive perception index is greater than or equal to a first preset threshold, a GPU computing unit is activated to accelerate the response; When it is detected that the comprehensive perception index is greater than or equal to a second preset threshold, the voice synthesis rate is reduced, the GPU computing unit response delay time is compressed, and an artificial response is activated; When it is detected that a plurality of consecutive comprehensive perception indexes are greater than or equal to a third preset threshold, a mediation flow guide map is activated and an artificial response is activated.

[0010] In some embodiments, the step of dynamically optimizing the response strategy based on the dynamic scene adaptation network according to the modal perception information and the corresponding response strategy for a plurality of consecutive times includes: The Bi-LSTM network based on the multi-granularity memory network encodes the continuous multiple modal perception information and the corresponding response strategies into dialogue entity codes, and generates dialogue entity coding correlation scores; The GAT network based on the multi-granularity memory network integrates user historical consultation semantics and behavior topology to obtain enhanced node representations of the knowledge graph topology structure; The memory coordination unit based on the multi-granularity memory network dynamically adjusts the subsequent response strategies by performing memory vector operations on the dialogue entity coding correlation scores and the enhanced node representations; The cross-language intent alignment network constructs a shared semantic space for multiple languages, and performs cross-language transfer in the shared semantic space through a contrast learning network.

[0011] In some embodiments, the federated learning scheduling algorithm is as follows: Task alloc =argmax(W cpu ×P latency ×Battery level ^0.5); wherein W cpu =0.3×log2(CPU main frequency / 1.0GHz); P latency =1 / (1+e^{-0.1×(latency-50)}); In the formula, Task alloc is edge resource allocation; W cpu is the core node weight; P latency is the network delay penalty factor; and Battery level is the battery level.

[0012] In some embodiments, the BERT-Emotion network based on the text emotion classification module performs semantic encoding on the keyword sentences in the interaction request information, and after the text emotion classification perception information is generated, the following steps are included: When it is detected that the corrugator muscle activation and the lip tightening in the AU intensity value are both greater than the corresponding threshold values, respectively, it is determined that the target user is in an angry emotional state, and a preset appeasement phrase library is called to trigger the soft voice line slow response of the digital human figure; When it is detected that the fundamental frequency fluctuation of the voice perception information is greater than or equal to a preset fluctuation threshold, the AU intensity value is greater than or equal to a preset intensity threshold, and the keyword repetition frequency in the text emotion classification perception information is abnormal, a risk warning is triggered, and a transaction operation is frozen.

[0013] In some embodiments, after the step of dynamically scheduling edge resources of the edge computing node including the multi-modal perception unit based on the federated learning scheduling algorithm in the perception response interaction process, the method comprises: predicting a fault of the multi-modal perception unit according to the historical indicators of the edge resources based on the LSTM-Transformer network, and triggering a driving update of the multi-modal perception unit if the predicted value is greater than a preset threshold.

[0014] In a second aspect, an intelligent interaction system based on modal perception and edge collaboration is provided, comprising: An interaction request module configured to obtain interaction request information of a target user; A modal perception generation module in communication connection with the interaction request module and configured to generate modal perception information based on the multi-modal perception unit and the interaction request information; A response generation module in communication connection with the modal perception generation module and configured to perform multi-head attention weighting on the modal perception information to obtain a comprehensive perception index, and generate a response strategy according to the comprehensive perception index; A dynamic optimization module in communication connection with the modal perception generation module and the response generation module and configured to dynamically optimize the response strategy based on a dynamic scene adaptation network and according to the modal perception information and the corresponding response strategy in continuous multiple times; and An edge resource scheduling module in communication connection with the dynamic optimization module and configured to dynamically schedule edge resources of an edge computing node including the multi-modal perception unit based on a federated learning scheduling algorithm in a perception response interaction process.

[0015] Compared with the prior art, the present application has the following advantages: through the technical fusion of multi-modal emotion calculation, edge collaboration architecture, dynamic scene adaptation and autonomous optimization mechanism, the technical bottlenecks of traditional customer service interaction in terms of emotional interaction accuracy, response efficiency and edge resource scheduling are solved. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 is a flowchart of an intelligent interaction method based on modal perception and edge collaboration of the present application; Figure 2 is a structural schematic diagram of an intelligent interaction system based on modal perception and edge collaboration of the present application. DETAILED DESCRIPTION

[0017] Reference will now be made in detail to the present embodiments of the application, examples of which are illustrated in the accompanying drawings. While the application will be described in conjunction with the specific embodiments, it will be understood that the application is not intended to be limited to the described embodiments. On the contrary, the application is intended to cover alternatives, modifications, and equivalents, which can be included within the spirit and scope of the application as defined by the appended claims. It should be noted that the steps of the methods described herein can be implemented by any of the functional blocks or functional arrangements, and any of the functional blocks or functional arrangements can be implemented as physical entities or logical entities, or a combination of both.

[0018] To enable persons skilled in the art to better understand the present application, the present application will be further described in detail below in conjunction with the drawings and specific embodiments.

[0019] Note: The examples to be introduced next are only specific examples, and are not intended to limit the embodiments of the present application to the specific steps, values, conditions, data, sequences, etc. Those skilled in the art can use the concept of the present application to construct more embodiments not mentioned in the present specification by reading the present specification.

[0020] Reference will now be made in detail to the present embodiments of the application, examples of which are illustrated in the accompanying drawings. While the application will be described in conjunction with the specific embodiments, it will be understood that the application is not intended to be limited to the described embodiments. On the contrary, the application is intended to cover alternatives, modifications, and equivalents, which can be included within the spirit and scope of the application as defined by the appended claims. It should be noted that the steps of the methods described herein can be implemented by any of the functional blocks or functional arrangements, and any of the functional blocks or functional arrangements can be implemented as physical entities or logical entities, or a combination of both. Figure 1

[0021] Step S100, obtaining interaction request information of a target user; Step S200, generating modal perception information based on a multi-modal perception unit.

[0022] When the user initiates an interaction request, a multi-modal perception unit is started. The multi-modal perception unit includes a voice emotion analysis module, a facial micro-expression recognition module, and a text emotion classification module.

[0023] After the voice input is processed by the WebRTC NS algorithm for noise reduction, 192-dimensional voiceprint features are extracted through the ECAPA-TDNN network, and similarity matching is performed with the voiceprint hash fingerprints in the historical interaction records, and voice perception information is generated.

[0024] The HRNet network of the facial micro-expression recognition module tracks the facial key point features in the interaction request information, and calculates the AU (Action Unit) intensity value in combination with the OpenFace algorithm.

[0025] The BERT-Emotion network of the text emotion classification module performs semantic encoding on the key word sentences in the interaction request information, such as key words including "complaint" and "urgent", and generates text emotion classification perception information.

[0026] ​Therefore, the emotion consistency fusion with KL divergence ≤0.1 is realized by cross-modal contrast learning constraint. When collecting the facial micro-expression data of the user, the HRNet network is used to dynamically track 52 key points, and at the same time, the 256-bit voiceprint hash fingerprint is generated by the local sensitive hash (LSH) bucketing technology in the voice aspect, so as to avoid the storage risk of the original biological feature data and ensure the desensitization processing in compliance with the GDPR.

[0027] At the same time, a double authorization mechanism can be implemented in the data collection stage: the primary authorization obtains the basic biological features (voiceprint, facial contour), and the secondary authorization unlocks the emotion state analysis function. The data storage adopts a hierarchical encryption strategy, the acoustic feature data is encrypted and stored in the local edge node using the AES-256-GCM algorithm, and the emotion analysis result is uploaded to the cloud after being processed by the SHA-3 hash. The user has the right to view the data usage graph in real time through the visual control panel, and can trigger the "data forgetting" instruction at one key, and the system will complete the irreversible erasure of the related features within 15 minutes.

[0028] In step S300, the multi-head attention weighting is performed on the modality perception information to obtain a comprehensive perception index, and a response strategy is generated according to the comprehensive perception index; In the voice emotion analysis module, the gradient reversal layer (GRL) is used for adversarial training to make the feature extraction V voice error reduction; and a multi-scale residual connection is designed to improve the time sequence feature aggregation capability of the ECAPA-TDNN model; in the facial micro-expression recognition module, the block captures the instantaneous contraction features of the glabella muscle and the zygomatic major muscle with a time resolution of ≤50 ms through the optical flow analysis method to generate a dynamic intensity value E micro ; after the text emotion classification module outputs the BERT encoder, the multi-head attention weighting is finally performed on all modality perception information to obtain a comprehensive perception index, which is shown in the following formula: E total =αV voice +θE micro +λP text In the formula, E total is the comprehensive perception index; V voice is the voice perception information; E micro is the AU intensity value; P text is the text emotion classification perception information; and α, θ, and λ are weight coefficients, respectively.

[0029] The response strategy generated according to the comprehensive perception index is as follows: When E total ≥ 0.7, the GPU calculation unit is activated to accelerate the response, and when E totalWhen ≥0.8, a three-level degradation strategy is executed, including reducing the speech synthesis rate to 0.8 times, compressing the GPU computing unit response delay to ≤400 ms, and transferring the human operator response.

[0030] When E total When ≥0.75, it is determined to be an emergency state, and the conflict mediation flow guide chart is automatically pushed and the artificial intervention process is started.

[0031] When the frown muscle activation (frown muscle activation > 0.6) and the lip tightening (lip tightening > 0.5) in the AU intensity value are both greater than the corresponding threshold, it is determined that the target user is in an angry emotional state, and a preset soothing phrase library (containing 2000 psychological verification phrases) is called, and the soft voice line slow response of the digital human is triggered; When the fundamental frequency fluctuation of the speech perception information is greater than or equal to the preset fluctuation threshold (fundamental frequency fluctuation ≥20%), the AU intensity value is greater than or equal to the preset intensity threshold (E micro ≥0.6), and the keyword repetition frequency in the text emotion classification perception information is abnormal (the keywords "transfer" and "password" appear ≥3 times / minute), a risk warning is triggered, and the transaction operation is frozen.

[0032] In step S400, the dynamic scene adaptation network dynamically optimizes the response strategy based on the continuous multiple modal perception information and the corresponding response strategy. The dynamic scene adaptation network realizes context association and multilingual support through a multi-granularity memory network and a cross-language intent alignment module.

[0033] The recent N consecutive dialogue contents are time-series encoded by a bidirectional long short-term memory neural network (Bi-LSTM), where N is a preset positive integer ≥5, preferably 5, and the named entity set in the dialogue flow is dynamically extracted; the semantic correlation degree between the named entities is calculated based on a self-attention mechanism to generate an entity encoding relevance score matrix, where the score value reflects the co-occurrence significance of the entity in the current dialogue context.

[0034] A dynamic knowledge graph is constructed with user historical consultation records as nodes, and the node feature vectors are extracted by a semantic encoder; a graph attention network (GAT) is used to learn the topological structure of the knowledge graph, and the semantic information of K-hop neighborhood nodes (K is an integer ≥1) is aggregated through a multi-head attention mechanism to generate an enhanced node representation with global dependency relationships; The short-term entity relevance score matrix and the long-term enhanced node representation are aligned across granularity; a gated fusion algorithm is used to dynamically weight the contribution weights of short-term memory and long-term memory to generate a fused memory vector; finally, the fused memory vector is output to the downstream dialogue decision module to generate context-aware response content.

[0035] A cross-language intent alignment network constructs a shared semantic space for 12 languages, and performs cross-language transfer, such as Japanese to Chinese intent matching, through a contrastive learning model in the shared semantic space.

[0036] In step S500, during the perception-response interaction process, edge resources of edge computing nodes, including multimodal perception units, are dynamically scheduled based on the federated learning scheduling algorithm.

[0037] Edge computing node deployment must meet regional coverage density requirements, with at least three edge servers equipped with Qualcomm QCS8550 chips per square kilometer. These servers incorporate a lightweight ECAPA-TDNN speech analysis model, a BERT-Emotion text sentiment classifier, and an HRNet facial micro-expression recognition network. Low-latency communication links are established between nodes via the Time-Sensitive Networking (TSN) protocol. The dynamic resource scheduling (also known as federated learning scheduling algorithm) formula is as follows: Task alloc =argmax(W cpu ×P latency ×Battery level ^0.5); Among them, W cpu =0.3×log2(CPU clock speed / 1.0GHz); P latency =1 / (1+e^{-0.1×(latency-50)}); In the formula, Task alloc Allocate resources to the edge; W cpu For core node weights; P latency Battery is a network latency penalty factor. level This represents the remaining battery capacity.

[0038] The federated learning collaborative mechanism adopts a three-stage optimization strategy: In the local model training stage, each edge node updates the sentiment recognition sub-model based on regional datasets (such as dialect speech databases and regional expression habit corpora); In the model aggregation stage, the central server performs a secure multi-party computation every 24 hours and aggregates the model gradients in a privacy-preserving manner using the Paillier homomorphic encryption algorithm; In the global distribution stage, the aggregated meta-model is given Laplacian noise satisfying (ε=0.5, δ=10^-5) through a differential privacy mechanism to ensure the untraceability of model parameters.

[0039] The dynamic load balancing of the edge node is implemented by a reinforcement learning algorithm, a state space is defined as a three-tuple of node CPU utilization, memory occupancy and network bandwidth, an action space includes three strategies of task allocation, model offloading and resource preemption, a reward function is designed as R = a x Throughput - β x Latency - γ x Energy, wherein the throughput weight a = 0.6, the delay penalty coefficient β = 0.3, and the energy consumption factor γ = 0.1. The training adopts a PPO algorithm, and converges to an optimal strategy after 10 iterations.

[0040] In another embodiment of the present application, after the step of dynamically scheduling edge resources of the edge computing node including the multi-modal perception unit based on the federated learning scheduling algorithm in the sensing response interaction process in the S500, the present application comprises: S600, based on the LSTM-Transformer network, the multi-modal perception unit is fault predicted according to the historical indicators of the edge resources, and if the predicted value is greater than the preset threshold, the driving update of the multi-modal perception unit is triggered.

[0041] Specifically, the multi-modal perception unit health degree prediction model processes 72-hour historical indicator data, extracts time sequence features through a bidirectional LSTM, models long-range dependency relationships through a Transformer encoder, and outputs a future 48-hour CPU load prediction curve and a fault probability P fault ∈ [0, 1], and the driving update process is triggered when the predicted value exceeds μ + 3σ.

[0042] Meanwhile, the quality of service optimization closed loop collects user click rate and satisfaction data in real time, and adjusts the response generation strategy of ChatGPT by using a PPO reinforcement learning algorithm.

[0043] The core of the emotional adaptation mechanism is to build a multi-dimensional psychological feature portrait. By integrating user click rate and satisfaction data, a user personality trait prediction model is established. For example, for users with an emotional stability score > 70, the system automatically enables high-frequency empathetic dialogues. For users with an openness trait score > 65, such as creativity tendency, innovative solutions are preferentially recommended. The physiological signal fusion analysis module integrates PPG photoplethysmogram and EDA skin conductance sensor data. When the heart rate variability HRV is detected to be lower than 50 ms and the skin conductance response continues to rise, it is determined that the user is in a stress state, and a deep breathing guidance program is immediately started, that is, a visual relaxation training in a 3D virtual scene is performed.

[0044] Meanwhile referring to Figure 2 As shown in the figure, the embodiment of the present application also provides an intelligent interaction system based on modal perception and edge collaboration, comprising: An interaction request module is configured to acquire interaction request information of a target user. The modal perception generation module is connected in communication with the interaction request module, and is configured to generate modal perception information based on the multi-modal perception unit and the interaction request information. The response generation module is connected in communication with the modal perception generation module, and is configured to perform multi-head attention weighting on the modal perception information to obtain a comprehensive perception index, and generate a response strategy according to the comprehensive perception index. The dynamic optimization module is connected in communication with the modal perception generation module and the response generation module, and is configured to dynamically optimize the response strategy based on a dynamic scene adaptation network according to the modal perception information and the corresponding response strategy in continuous multiple times. The edge resource scheduling module is connected in communication with the dynamic optimization module, and is configured to perform edge resource dynamic scheduling on edge computing nodes including the multi-modal perception unit based on a federated learning scheduling algorithm in the perception response interaction process.

[0045] The intelligent interaction system based on modal perception and edge collaboration provided by the application solves the technical bottlenecks of traditional customer service systems in terms of emotional interaction accuracy, response efficiency and privacy security through the technical fusion of multi-modal emotion computing, edge collaboration architecture, dynamic scene adaptation and autonomous optimization mechanism.

[0046] The improved ECAPA-TDNN model is used to optimize the speech emotion analysis module, the multi-scale residual connection is used to improve the mel spectrum feature extraction capability, and the feature extraction error is reduced under the noise environment with a signal-to-noise ratio of ≤10dB. The face micro-expression recognition module is based on the 3D convolution network of HRNet, and the dynamic changes of 52 facial key points, such as the instantaneous features of the inter-brow muscle contraction intensity, are captured with a time resolution of ≤50ms, and the spatio-temporal correlation of muscle movement is quantified by the optical flow analysis method. The text emotion classification module integrates the BERT-Emotion model and the multi-head attention mechanism, and performs weighted fusion on the context semantics, and outputs the emotion polarity. The cross-modal contrast learning constrains the emotional consistency of voice, expression and text through a triple loss function, ensures that the KL divergence of multi-modal fusion is ≤0.1, and significantly improves the comprehensive accuracy of emotion recognition.

[0047] The edge collaboration architecture dynamically schedules heterogeneous device resources using a lightweight federated learning protocol, and the device grading strategy is based on CPU frequency (W cpu =0.3*log2(CPU frequency / 1.0GHz)) and network delay (P latency=1 / (1+e^{-0.1×(latency-50)})) divides the nodes into core nodes and auxiliary nodes. The core nodes undertake more than 30% of the model update tasks, thereby improving the utilization of computing power. The privacy protection unit completes biometric desensitization processing in the SGX trusted execution environment. Voiceprint data is bucketed using Local Sensitive Hash (LSH) to generate irreversible 256-bit hash fingerprints. Facial features are homomorphically encrypted and compressed into 128-dimensional vectors using HE-CNN, meeting GDPR compliance requirements end-to-end.

[0048] The dynamic scene adaptation engine achieves contextual association and multilingual support through a multi-granularity memory network and a cross-language intent alignment module. It utilizes a Bi-LSTM (Bi-Long Short-Term Memory) network to encode named entities from recent multi-turn dialogues, dynamically adjusting the dialogue focus based on entity relevance scores. A semantic topology structure based on a Graph Attention Network (GAT) is constructed, integrating user history and behavioral patterns to support contextual association for more than 30 turns of dialogue. The cross-language intent alignment module constructs a shared semantic space for 12 languages, achieving cross-language transfer of intent vectors through a contrastive learning model.

[0049] The self-optimization and fault prediction closed loop achieves system self-evolution through an equipment health prediction model and a service quality feedback mechanism. The equipment health prediction model employs an LSTM-Transformer hybrid architecture, processing 72 hours of historical indicators (CPU load, memory usage) to generate a 48-hour prediction curve. A driver update is triggered when the predicted value exceeds μ+3σ. The service quality feedback closed loop collects user click-through rate and satisfaction data in real time, dynamically adjusting the ChatGPT response generation strategy through near-end strategy optimization of the PPO algorithm. The knowledge base self-verification algorithm detects version conflicts of policy and regulatory entries based on logical constraints, automatically generating a confidence level traceability report. The knowledge base update response time is reduced to within 5 minutes, ensuring the legal compliance of service content.

[0050] At the level of emotional interaction, multimodal fusion technology can improve the accuracy of emotion recognition, for example, in financial anti-fraud scenarios, it can simultaneously detect voice tremor (fundamental frequency fluctuation ≥20%) and facial muscle tension (E). micro The false alarm rate for triggering the risk warning protocol is reduced to 3% for anomalies such as ≥0.6) and keyword repetition anomalies ("transfer" and "password" appearing ≥3 times / minute). The cross-language intent alignment module supports real-time semantic transfer in 12 languages, reducing the intent misjudgment rate in Japanese customer service scenarios.

[0051] In terms of response efficiency, the edge collaboration architecture compresses end-to-end interaction latency to ≤50ms, and pre-loads a library of 500 high-frequency question-and-answer templates to achieve localized, real-time responses to common questions through semantic similarity matching. Privacy protection mechanisms ensure the security of biometric data through edge-side anonymization processing, meeting the requirements of GDPR and the Personal Information Protection Act.

[0052] In terms of system self-evolution and service continuity, the device health prediction model can provide 48-hour early warning for hardware failure, and the service interruption rate is reduced. The knowledge base conflict detection algorithm scans policy and regulation updates every 24 hours, automatically generates version change traceability reports, and ensures the real-time and compliance of the knowledge base. User satisfaction data is fed back to the ChatGPT model through reinforcement learning to optimize the response generation strategy. For example, in a certain bank case, the system synchronously analyzes voice, expression and text keywords through the sentiment computing engine, improves the accuracy of identifying abnormal transaction requests, and improves the efficiency of risk work order processing.

[0053] Specifically, the present embodiment corresponds to the above method embodiment one by one, and the functions of each module have been described in detail in the corresponding method embodiment, so they will not be repeated here.

[0054] Based on the same inventive concept, the embodiments of the present application also provide a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement all or part of the method steps of the above method.

[0055] The present application realizes all or part of the above method, and can also be completed by a computer program to instruct related hardware. The computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms. The computer readable medium can include any entity or device that can carry computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content of the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.

[0056] Based on the same inventive concept, the embodiments of the present application also provide an electronic device, comprising a memory and a processor, the memory stores a computer program running on the processor, and the processor executes the computer program to implement all or part of the method steps of the above method.

[0057] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor is a control center of the computer device, and connects all parts of the computer device through various interfaces and lines.

[0058] The memory can be used to store computer programs and / or modules, and the processor can realize various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and calling data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application program required by a function (for example, a sound playing function, an image playing function, etc.); and the data storage area can store data created according to use of the mobile phone (for example, audio data, video data, etc.). In addition, the memory can include a high-speed random access memory, and can also include a nonvolatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory devices.

[0059] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, a server or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage, etc.) containing computer-usable program code.

[0060] The present application is described in reference to the appended drawings figures and / or block diagrams of methods, apparatus (systems), servers, and computer program products according to embodiments of the application. It will be understood that each block of the flowchart and / or block diagrams, and combinations of blocks in the flowchart and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 one or more flowcharts and / or blocks

[0061] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 one or more flowcharts and / or blocks

[0062] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart and / or block diagram block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 one or more flowcharts and / or blocks

[0063] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. An intelligent interaction method based on modal perception and edge collaboration, characterized in that, include: Obtain the target user's interaction request information; The interaction request information is used to generate modal perception information based on the multimodal perception unit; The modal perception information is weighted by multi-head attention to obtain a comprehensive perception index, and a response strategy is generated based on the comprehensive perception index. The dynamic scene adaptation network dynamically optimizes the response strategy based on the modal perception information and the corresponding response strategy obtained from multiple consecutive modal perceptions. During the perception-response interaction process, edge resources are dynamically scheduled for edge computing nodes, including multimodal perception units, based on the federated learning scheduling algorithm.

2. The intelligent interaction method based on modal perception and edge collaboration as described in claim 1, characterized in that, The process of generating modal perception information from the interaction request information based on the multimodal perception unit includes: The multimodal perception unit includes a speech emotion analysis module, a facial micro-expression recognition module, and a text emotion classification module; The ECAPA-TDNN network based on the speech emotion analysis module extracts the target voiceprint features from the interaction request information and generates speech perception information. The HRNet network based on the facial micro-expression recognition module tracks the facial key point features in the interaction request information and generates AU intensity values. The BERT-Emotion network, based on the text emotion classification module, performs semantic encoding on the keywords and phrases in the interaction request information to generate text emotion classification perception information.

3. The intelligent interaction method based on modal perception and edge collaboration as described in claim 1, characterized in that, The method for performing multi-head attention weighting on the modal perception information to obtain the comprehensive perception index is shown in the following formula: E total =αV voice +θE micro +λP text In the formula, E total For comprehensive perception index; V voice For speech-sensing information; E micro AU intensity value; P text The emotional perception information is classified into texts; α, θ, and λ are the weighting coefficients, respectively.

4. The intelligent interaction method based on modal perception and edge collaboration as described in claim 1, characterized in that, The method for generating a response strategy based on the comprehensive perception index includes: When the comprehensive perception index is detected to be greater than or equal to the first preset threshold, the GPU computing unit is activated to accelerate the response. When the comprehensive perception index is detected to be greater than or equal to the second preset threshold, the following actions are taken: reducing the speech synthesis rate, compressing the GPU computing unit response latency, and activating human response. When multiple consecutive comprehensive perception indices are detected to be greater than or equal to the third preset threshold, the mediation process guide diagram and manual response are activated.

5. The intelligent interaction method based on modal perception and edge collaboration as described in claim 1, characterized in that, The dynamic scene adaptation network dynamically optimizes the response strategy based on multiple consecutive modal perception information and the corresponding response strategy, including: The Bi-LSTM network, based on a multi-granularity memory network, encodes dialogue entities for the modality perception information and the corresponding response strategy multiple times in succession, and generates a dialogue entity encoding relevance score. The graph attention GAT network based on multi-granularity memory network integrates the semantics and behavioral topology of user historical consultations to obtain enhanced node representations of the knowledge graph topology. The memory collaboration unit based on the multi-granularity memory network remembers the relevance score of the dialogue entity encoding and the enhanced node representation to dynamically adjust the subsequent response strategy; A shared semantic space for multiple languages ​​is constructed based on a cross-language intent alignment network, and cross-language transfer is performed in the shared semantic space through a contrastive learning network.

6. The intelligent interaction method based on modal perception and edge collaboration as described in claim 1, characterized in that, The federated learning scheduling algorithm is shown in the following formula: Task alloc =argmax(W cpu ×P latency ×Battery level ^0.5); Among them, W cpu =0.3×log2(CPU clock speed / 1.0GHz); P latency =1 / (1+e^{-0.1×(latency-50)}); In the formula, Task alloc Allocate resources to the edge; W cpu For core node weights; P latency Battery is a network latency penalty factor. level This represents the remaining battery capacity.

7. The intelligent interaction method based on modal perception and edge collaboration as described in claim 2, characterized in that, The BERT-Emotion network, based on the text emotion classification module, performs semantic encoding on the keywords and phrases in the interaction request information to generate text emotion classification perception information, including: When the activation of the corrugator supercilii muscle and the lip contraction in the AU intensity value are both greater than the corresponding thresholds, it is determined that the target user is in an angry state, and the preset soothing speech library is invoked to trigger the soft voice of the digital humanoid to respond slowly. When the fundamental frequency fluctuation of the voice perception information is greater than or equal to the preset fluctuation threshold, the AU intensity value is greater than or equal to the preset intensity threshold, and the frequency of keyword repetition in the text sentiment classification perception information is abnormal, a risk warning and transaction freezing operation will be triggered.

8. The intelligent interaction method based on modal perception and edge collaboration as described in claim 1, characterized in that, During the perception-response interaction process, after dynamically scheduling edge resources for edge computing nodes including multimodal perception units based on the federated learning scheduling algorithm, the process includes: Based on the LSTM-Transformer network, fault prediction of the multimodal sensing unit is performed according to the historical indicators of edge resources. If the predicted value is greater than the preset threshold, the driving update of the multimodal sensing unit is triggered.

9. An intelligent interaction system based on modal perception and edge collaboration, characterized in that, include: The interaction request module is used to obtain interaction request information from the target user; The modal perception generation module is communicatively connected to the interaction request module and is used to generate modal perception information from the interaction request information based on the multimodal perception unit; The response generation module is communicatively connected to the modality perception generation module and is used to perform multi-head attention weighting on the modality perception information to obtain a comprehensive perception index, and generate a response strategy based on the comprehensive perception index. The dynamic optimization module is communicatively connected to the modality perception generation module and the response generation module, and is used to dynamically optimize the response strategy based on the modality perception information and the corresponding response strategy through the dynamic scene adaptation network. as well as, The edge resource scheduling module is communicatively connected to the dynamic optimization module and is used to dynamically schedule edge resources for edge computing nodes, including multimodal sensing units, based on a federated learning scheduling algorithm during the perception-response interaction process.

Citation Information

Patent Citations

  • Multi-modal digital employee reception system and method based on edge calculation

    CN119693750A