Enterprise service method and system based on artificial intelligence

By monitoring multimodal dialogue messages within enterprise groups, performing multimodal fusion and personalized processing, and utilizing an industry-level federated knowledge base to determine the response mechanism, the problem of low efficiency in traditional enterprise digital consulting services has been solved, achieving efficient and accurate information processing and secure enterprise communication.

CN120952803APending Publication Date: 2025-11-14HUNAN JINGRUI INTELLIGENT TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511445300.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional enterprise digital consulting services are inefficient and fragmented in knowledge transfer. Intelligent customer service systems struggle to cope with complex and ever-changing consulting scenarios and have significant deficiencies in terms of professionalism, multimodal information understanding, and contextual understanding, thus failing to meet the in-depth digital consulting needs of enterprise clients.

Method used

By listening to multimodal dialogue messages from enterprise groups, identity information is determined, multimodal fusion is performed, an industry-level federated knowledge base is used to determine the response mechanism, and the entire lifecycle of dialogue is monitored. Combined with cross-modal attention mechanisms and differential privacy technology, personalized processing and information storage are achieved.

Benefits of technology

It has enabled efficient and accurate information processing and interaction, improved corporate communication efficiency, ensured information security, optimized decision-making processes, and enhanced business management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952803A_ABST
    Figure CN120952803A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an enterprise service method and system based on artificial intelligence. The method comprises the following steps: monitoring dialogue messages corresponding to a plurality of enterprise groups, the dialogue messages comprising voice messages, and determining identity information corresponding to the voice messages; performing multi-modal fusion on the dialogue message to obtain a fused dialogue message; determining a target event corresponding to the fused dialogue message, and determining a response mechanism corresponding to the target event based on an industry-level federal knowledge base; and delivering the target event based on the response mechanism and the identity information, monitoring a full-life-cycle dialogue of the target event, and storing generated log information. According to the invention, the enterprise group communication efficiency can be improved, the decision process can be optimized, and the overall business management level can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of enterprise service technology, specifically to an enterprise service method and system based on artificial intelligence. Background Technology

[0002] Traditional enterprise digital consulting services primarily rely on human consultant teams, providing one-on-one services through offline meetings or instant messaging tools. As enterprise clients grow in size, this model has revealed problems such as low service efficiency and fragmented knowledge transfer.

[0003] In recent years, intelligent customer service systems have attempted to automate responses through rule engines. However, limited by preset dialogue flows and fixed knowledge base architectures, they struggle to handle complex and ever-changing consultation scenarios. Furthermore, their technical architectures are mostly built on web pages or standalone applications. In recent years, while chatbots based on natural language processing can handle simple questions and answers, they still have significant shortcomings in terms of professionalism, multimodal information understanding, and contextual understanding, failing to meet the needs of enterprise clients for in-depth digital consultation. Summary of the Invention

[0004] To address the aforementioned technical problems, embodiments of this application provide an enterprise service method and system based on artificial intelligence.

[0005] According to one aspect of the embodiments of this application, an artificial intelligence-based enterprise service method is provided, comprising: listening to dialogue messages corresponding to multiple enterprise groups, the dialogue messages including voice messages, and determining the identity information corresponding to the voice messages; performing multimodal fusion on the dialogue messages to obtain fused dialogue messages; determining a target event corresponding to the fused dialogue messages, and determining a response mechanism corresponding to the target event based on an industry-level federated knowledge base; delivering the target event based on the response mechanism and the identity information, and monitoring the entire lifecycle dialogue of the target event and storing the generated log information.

[0006] According to one aspect of the embodiments of this application, the method further includes: obtaining context information of the dialogue message and historical service records; and determining the dialogue sequence corresponding to the dialogue message based on the context information and the historical service records.

[0007] According to one aspect of the embodiments of this application, the step of performing multimodal fusion on the dialogue message to obtain a fused dialogue message includes: acquiring a multimodal data stream in the dialogue message, the multimodal data stream including multiple elements such as images, text, and speech; performing cross-modal fusion on the multimodal data stream, and determining the fused dialogue message based on the fusion result and the dialogue timing.

[0008] According to one aspect of the embodiments of this application, determining the target event corresponding to the fused dialogue message includes: determining cross-modal attention output features corresponding to text features and image features in the dialogue message based on a cross-modal attention mechanism, wherein the cross-modal attention output features represent the attention result of the text features on the image features; mapping the cross-modal attention output features and the text features to a gating space respectively to obtain corresponding gating coefficients; determining the fusion ratio of the text features and the image features in cross-modal feature fusion based on the gating coefficients, so as to determine a target fusion strategy between the text features and the image features based on the fusion ratio; and determining the target event corresponding to the fused dialogue message through the target fusion strategy.

[0009] According to one aspect of the embodiments of this application, determining the response mechanism corresponding to the target event based on an industry-level federated knowledge base includes: collecting historical consultation cases through differential privacy technology based on the underlying industry-level federated knowledge base to obtain an industry knowledge graph; determining the triggering engine corresponding to the target event, and obtaining the associated content corresponding to the target event based on the triggering engine and the industry knowledge graph; and determining the response mechanism corresponding to the target event based on the associated content.

[0010] According to one aspect of the embodiments of this application, the step of sending the target event based on the response mechanism and the identity information, and monitoring the full lifecycle dialogue of the target event and storing the generated log information includes: triggering a full lifecycle dialogue management system based on the response mechanism, the full lifecycle dialogue management system including a dialogue state tracker, an intent association analyzer, and a service strategy optimizer; determining the service strategy corresponding to the target event based on the full lifecycle dialogue management system, and adjusting the robot intervention depth and response detail level based on the job responsibilities and permission level of the target employee through the service strategy optimizer.

[0011] According to one aspect of the embodiments of this application, the method further includes: parsing the voice message to determine the voiceprint information corresponding to the voice message based on the parsing result; determining the target employee and the identity identifier corresponding to the target employee based on the voiceprint information; obtaining an enterprise organizational structure map, and determining the job responsibilities and authority level corresponding to the target employee based on the enterprise organizational structure map and the identity identifier.

[0012] According to one aspect of the embodiments of this application, an artificial intelligence-based enterprise service system is provided. The system includes: a monitoring module, configured to monitor dialogue messages corresponding to multiple enterprise groups, the dialogue messages including voice messages, and determine the identity information corresponding to the voice messages; a fusion module, configured to perform multimodal fusion of the dialogue messages to obtain fused dialogue messages; a determination module, configured to determine a target event corresponding to the fused dialogue messages, and determine a response mechanism corresponding to the target event based on an industry-level federated knowledge base; and a delivery module, configured to deliver the target event based on the response mechanism and the identity information, and monitor the entire lifecycle of the target event's dialogue and store the generated log information.

[0013] In the technical solution provided by the embodiments of this application, by listening to multiple enterprise group dialogue messages and clarifying the identity information corresponding to the voice messages, the accuracy and traceability of the information source are ensured, which helps to personalize the processing for different identities in the future. Multimodal fusion of dialogue messages can comprehensively utilize multiple information forms such as text and voice to obtain more comprehensive and accurate fused dialogue messages, avoiding the one-sidedness that may exist with single-modal information. The target event corresponding to the fused dialogue message is determined, and the response mechanism is determined with the help of an industry-level federated knowledge base. The industry-level federated knowledge base gathers extensive and professional industry knowledge, which can ensure the scientific, professional, and targeted nature of the response mechanism. Finally, the target event is delivered based on the response mechanism and identity information, and the entire lifecycle of the dialogue is monitored and stored. This not only achieves effective processing and feedback of the target event, but the stored full lifecycle dialogue data can also provide rich material for subsequent analysis and optimization, which helps to improve the efficiency of enterprise group communication, optimize decision-making processes, and enhance the overall business management level.

[0014] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings: Figure 1 This is a schematic diagram illustrating an implementation environment for an AI-based enterprise service, as shown in an exemplary embodiment of this application. Figure 2 This is a flowchart illustrating an exemplary embodiment of the present application of an artificial intelligence-based enterprise service method; Figure 3This is a flowchart illustrating an artificial intelligence-based enterprise service method, as shown in another exemplary embodiment of this application; Figure 4 This is a flowchart illustrating an artificial intelligence-based enterprise service method, as shown in another exemplary embodiment of this application; Figure 5 This is a flowchart illustrating an artificial intelligence-based enterprise service method, as shown in another exemplary embodiment of this application; Figure 6 This is a flowchart illustrating an artificial intelligence-based enterprise service method, as shown in another exemplary embodiment of this application; Figure 7 This is a flowchart illustrating an artificial intelligence-based enterprise service method, as shown in another exemplary embodiment of this application; Figure 8 This is a flowchart illustrating an artificial intelligence-based enterprise service method, as shown in another exemplary embodiment of this application; Figure 9 This is a block diagram illustrating an artificial intelligence-based enterprise service system, as shown in an exemplary embodiment of this application. Figure 10 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown. Detailed Implementation

[0016] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0017] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0018] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0019] In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0020] First, it's important to note that existing mass messaging technologies generally employ static tag matching mechanisms and timed broadcasting, distributing content in a coarse-grained manner based on enterprise category tags. This fails to consider the stage-specific characteristics of enterprise digitalization processes and the timeliness of industry policies, leading to a mismatch between pushed content and customers' real-time needs. Within the WeChat ecosystem, due to the highly interactive and fragmented nature of group chats, traditional push technologies face two major challenges: first, the relevance of pushed content to customers' immediate business needs rapidly diminishes over time; second, the lack of an effective knowledge transfer mechanism for cross-group content dissemination results in low efficiency in spreading high-quality information. Furthermore, existing systems often use independent content management modules, failing to form a data loop with consulting services and hindering continuous optimization of push strategies.

[0021] Enterprise consultation scenarios within the WeChat ecosystem involve multimodal information interaction, including text, voice, images, and emojis, which places higher demands on the semantic understanding capabilities of chatbots. Current mainstream solutions primarily employ a modular processing strategy: speech recognition, image OCR recognition, and text processing operate independently. This sequential processing mode not only leads to the loss of contextual information but also makes it difficult to resolve semantic relationships between multimodal information. Furthermore, while general-purpose large language models possess multimodal processing potential, their open-domain nature can easily cause responses to deviate from the enterprise's proprietary knowledge system, posing a risk of insufficient reliability in professional knowledge.

[0022] In the area of ​​platform ecosystem integration, although WeChat has become an important entry point for enterprise customer service, existing technical solutions have significant shortcomings in deep integration with the WeChat ecosystem. Most systems use plug-in message middleware for basic connection, failing to natively support key features of WeChat group chat, such as context state management, precise capture of specified robot commands, and isolated access to multiple enterprise knowledge bases. This leads to systemic risks such as frequent message response delays, cross-group dialogue confusion, and leakage of sensitive information during service delivery. Furthermore, traditional architectures struggle to support service stability under high concurrency scenarios. When handling consultation requests from hundreds of enterprise groups simultaneously, message loss or response timeouts are common, severely limiting scalable service capabilities.

[0023] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an implementation environment for an artificial intelligence-based enterprise service, as shown in an exemplary embodiment of this application. Figure 1As shown, server 110 can monitor dialogue messages corresponding to multiple enterprise groups. These dialogue messages include voice messages, text messages, and image messages within the enterprise groups. Server 110 can determine the corresponding identity information based on the voice messages. Furthermore, server 110 can perform multimodal fusion of the dialogue messages to obtain fused dialogue messages. Then, it can determine the target event corresponding to the fused dialogue message. Server 110 then determines the response mechanism corresponding to the target event by calling the industry federated knowledge base. Based on the response mechanism and identity information, it delivers the target event and monitors the entire lifecycle of the target event's dialogue, storing the generated log information. This enables the processing of dialogue messages from multiple enterprise groups.

[0024] in, Figure 1 The server 110 shown can be, for example, a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. No restrictions are imposed here.

[0025] Traditional enterprise digital consulting services primarily rely on human consultant teams, providing one-on-one services through offline meetings or instant messaging tools. As enterprise clients grow in size, this model has revealed problems such as low service efficiency and fragmented knowledge transfer.

[0026] In recent years, intelligent customer service systems have attempted to automate responses through rule engines. However, limited by preset dialogue flows and fixed knowledge base architectures, they struggle to handle complex and ever-changing consultation scenarios. Furthermore, their technical architectures are mostly built on web pages or standalone applications. In recent years, while chatbots based on natural language processing can handle simple questions and answers, they still have significant shortcomings in terms of professionalism, multimodal information understanding, and contextual understanding, failing to meet the needs of enterprise clients for in-depth digital consultation.

[0027] To address these issues, embodiments of this application propose an artificial intelligence-based enterprise service method, an artificial intelligence-based enterprise service device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail below.

[0028] Please see Figure 2 , Figure 2 This is a flowchart illustrating an exemplary embodiment of an artificial intelligence-based enterprise service method. This method can be applied to... Figure 1The implementation environment shown is specifically executed by server 110 within that implementation environment. It should be understood that this method can also be applied to other exemplary implementation environments and executed by devices in other implementation environments; this embodiment does not limit the implementation environment to which the method is applicable.

[0029] like Figure 2 As shown, in an exemplary embodiment, the AI-based enterprise service method includes at least steps S210 to S240, which are described in detail below: Step S210: Listen to the dialogue messages corresponding to multiple enterprise groups. The dialogue messages include voice messages, and determine the identity information corresponding to the voice messages.

[0030] For example, within the framework of WeChat's deep integration and the construction of a message hub as a unified service entry point, and with this hub bearing the important mission of enabling compliant listening and pushing of messages from multiple enterprise groups, the following process can be followed to achieve the goal of listening to conversation messages (including voice messages) from multiple enterprise groups and determining the identity information corresponding to the voice messages: First, enterprises need to complete a rigorous registration and authentication process on the WeChat Open Platform, submitting comprehensive and authentic information such as business licenses and legal representative information. After careful review and approval by WeChat officials, they will obtain legal permissions to conduct business within the WeChat ecosystem. Next, a dedicated application should be created for the enterprise, and key interface permissions such as "obtaining enterprise group messages" and "voice recognition related" should be precisely applied for based on actual business needs. At the same time, the callback address of the application should be reasonably configured. This address is used to accurately receive message notifications pushed by the WeChat server. Furthermore, a strict security strategy covering data encryption storage and transmission should be formulated to ensure that the entire process complies with national laws and regulations and WeChat ecosystem security standards. After completing the initial preparations, the system periodically sends requests to the WeChat server using the interfaces provided by the WeChat Open Platform to obtain a detailed list of all associated enterprise groups under the enterprise, including group ID, group name, number of group members, etc. This information is then properly stored in the enterprise's own database for subsequent management and retrieval. Subsequently, detailed message subscription configurations are performed in the application for each enterprise group that needs to be monitored, specifying the message types to be monitored, such as text, voice, images, and files. At the same time, the frequency and method of message push are reasonably set to ensure timely and accurate acquisition of messages in the group. Meanwhile, a highly available and stable callback service is built internally within the enterprise. This service should be able to easily handle a large number of concurrent message requests. When the WeChat server detects a new message in the enterprise group, it will push the message content to the enterprise callback service according to the preset callback address.

[0031] Step S220: Perform multimodal fusion on the dialogue messages to obtain the fused dialogue messages.

[0032] For example, after listening to multiple enterprise group dialogue messages and determining the identity information of the voice messages, the following specific implementation process can be followed to obtain the fused dialogue message through multimodal fusion: First, a multimodal fusion processing framework is constructed. This framework needs to have core functional modules such as preprocessing of different modal messages, feature extraction, feature fusion, and information integration, and it must be tightly integrated with the previously constructed message hub system to ensure seamless reception of various dialogue message data from multiple enterprise groups. For the received text messages, natural language processing techniques are used for preprocessing, including word segmentation, stop word removal, and part-of-speech tagging, to eliminate noise and redundant information in the text, making the text content more standardized and easier to analyze. Then, word embedding techniques, such as Word2Vec and GloVe, are used to convert each word in the text into a high-dimensional vector representation. These vectors can capture the semantic and syntactic features of words, thereby extracting the key feature information of the text message. For image messages, preprocessing is performed using image processing libraries such as OpenCV, including image scaling, cropping, and grayscale conversion, to meet subsequent feature extraction requirements. Then, deep learning models such as Convolutional Neural Networks (CNNs), including VGG and ResNet, are used to extract high-level features from the images. These features reflect the image's content, structure, and semantic information. For speech messages, based on previously determined identity information, further speech signal processing is performed, including pre-emphasis, framing, and windowing, to enhance the quality and stability of the speech signal. Next, feature extraction methods such as Mel-frequency cepstral coefficients (MFCC) are used to convert the speech signal into a series of feature parameters that characterize the timbre, pitch, and other features of the speech. If the speech has already been converted to text, this text information can also be used in the fusion process along with the speech features. After feature extraction, the feature fusion stage begins, employing various feature fusion methods such as simple concatenation, weighted fusion, and attention-based fusion.

[0033] Step S230: Determine the target event corresponding to the merged dialogue message, and determine the response mechanism corresponding to the target event based on the industry-level federated knowledge base.

[0034] For example, after completing the multimodal fusion of dialogue messages to obtain the fused dialogue message, to determine its corresponding target event and the response mechanism based on the industry-level federated knowledge base, the specific implementation process can be as follows: First, a target event recognition model is constructed. This model needs to combine natural language processing, machine learning, and specific industry knowledge to accurately extract key information and identify potential events from the fused dialogue message. When a new fused dialogue message is input, the target event recognition model first preprocesses it, including word segmentation, stop word removal, and part-of-speech tagging, to extract keywords and phrases in the dialogue. Then, the trained model is used to extract features and classify the preprocessed dialogue message to determine its target event type. To improve recognition accuracy, an ensemble learning method can be used to combine multiple different classification models and integrate their prediction results. At the same time, a rule engine is combined to quickly judge dialogues with obvious specific formats or keywords, assisting the event recognition process. After identifying the target event, the next step is to determine the response mechanism based on an industry-level federated knowledge base. This industry-level federated knowledge base is a knowledge-sharing platform jointly built and maintained by multiple enterprises or institutions. It contains a wealth of industry knowledge, best practices, policies, and regulations. During the knowledge base construction process, each participant organizes and uploads its accumulated industry knowledge according to unified data standards and formats to ensure the consistency and usability of the knowledge. At the same time, federated learning technology is used to achieve collaborative updates and optimization of the knowledge base while protecting the data privacy of each participant.

[0035] Furthermore, to provide personalized services to businesses, an industry-knowledge-guided semantic disambiguation module can be introduced. This module transforms industry tags in the business registration information into semantic constraints, dynamically adjusting the fusion weights of various modal features. This process simultaneously outputs an identity-bound intent vector, enabling the system to possess role-based contextual understanding capabilities.

[0036] Step S240: Based on the response mechanism and identity information, the target event is delivered, and the entire lifecycle dialogue of the target event is monitored and the generated log information is stored.

[0037] For example, event delivery can be pre-built, which needs to be tightly integrated with the previously determined response mechanism and identity information module. It should have the function of accurately locating the delivery target based on the identity information, and at the same time, it should be able to generate delivery content that conforms to business specifications and communication habits according to the response mechanism. Before delivery, the identity information will be verified and parsed to ensure accurate identification of the target recipient. For example, if the identity information is of an internal employee, the specific delivery channel will be determined based on the employee's department, position, etc., such as WeChat or email; if the identity information is of an external customer, the channel will be selected based on the customer's preferences and business scenario, such as SMS or instant messaging tools. When generating delivery content, it will combine the processing flow and standard answer information in the response mechanism to automatically generate clear, accurate and targeted messages. For some complex events, additional information such as relevant documents and links can also be provided to help the recipient better understand and handle the event. After determining the recipient and content, the event is delivered through the selected channel. During the delivery process, key information such as the delivery time, channel, and recipient is recorded, and a delivery record is generated. After delivery, feedback from the recipient is awaited. If no feedback is received within a certain period of time, a reminder or re-delivery will be issued according to preset rules to ensure that the event is handled in a timely manner. Simultaneously, a full lifecycle dialogue monitoring system is established, capable of capturing all dialogue messages related to the target event in real time, including responses after submission and further communication content. The monitoring system connects with the message hub and event submission deeply integrated with the WeChat ecosystem to obtain dialogue data streams. After the dialogue data stream enters the monitoring system, the data undergoes preprocessing, such as noise removal and identification of dialogue participants, for subsequent analysis. Natural language processing technologies, such as sentiment analysis and keyword extraction, are used to conduct in-depth analysis of the dialogue content, understanding the participants' attitudes, concerns, and the progress of event handling. During the monitoring process, the dialogue is evaluated in real time according to preset rules and indicators, such as determining whether the dialogue has deviated from the topic or whether conflicts have occurred. If an anomaly is detected, an alarm will be issued in a timely manner to notify relevant personnel to intervene and handle the situation.

[0038] In addition, in order to achieve a complete record of the entire lifecycle of the target event dialogue, all information generated during the dialogue, including text, voice, and images, will be stored in a certain format and structure. The storage must have high reliability, high scalability, and data security to ensure that the log information can be preserved for a long time and not be tampered with. When storing log information, metadata such as timestamps, event identifiers, and dialogue participants will be added to each piece of information for subsequent querying and analysis.

[0039] In some embodiments of this application, by listening to multi-enterprise group dialogue messages and accurately determining identities, multimodal fusion of dialogue messages, relying on an industry-level federated knowledge base to determine response mechanisms, delivering target events according to identities, and monitoring and storing full lifecycle dialogue logs, efficient and accurate information processing and interaction can be achieved, improving enterprise communication efficiency, ensuring information security, and assisting in decision optimization.

[0040] Furthermore, based on the above embodiments, please refer to... Figure 3 In one exemplary embodiment provided in this application, the specific implementation process of the above-mentioned artificial intelligence-based enterprise service method further includes steps S310 and S320, which are described in detail below: Step S310: Obtain the context information of the dialogue message and historical service records; Step S320: Determine the dialogue sequence corresponding to the dialogue message based on context information and historical service records.

[0041] For example, after completing the delivery of the target event and starting to monitor the entire lifecycle of the dialogue, it is necessary to obtain the context information of the dialogue messages and historical service records, and determine the dialogue sequence corresponding to the dialogue messages based on this information. The specific implementation process can be carried out as follows: First, a data acquisition module is built. This module needs to be tightly integrated with the message hub deeply coupled with the WeChat ecosystem, the enterprise's own business system, and the log information management system to obtain the required context information and historical service records. For context information, the message hub captures the context content of the current dialogue in real time, including the start time of the dialogue, participants, previous dialogue segments, etc. At the same time, natural language processing technology is used to analyze the context content and extract key information, such as topics, keywords, sentiment, etc., in order to better understand the background and intent of the dialogue. For historical service records, historical service data related to the current dialogue is obtained from the customer relationship management (CRM) module, service ticket system, etc. in the enterprise's business system. This data includes the customer's previous consultation records, problem handling progress, service evaluation, etc. The historical service records are imported into the data acquisition module through the data interface, and data cleaning and preprocessing are performed to remove duplicate, erroneous and incomplete data to ensure the accuracy and consistency of the data.

[0042] Then, after obtaining the context information and historical service records, the dialogue sequence determination stage begins. First, the dialogue messages are initially sorted using timestamp information, arranging the dialogues in chronological order to form a preliminary dialogue sequence framework. Next, semantic association analysis is performed on the dialogues by combining keywords and themes in the context information to determine the logical relationships and coherence between the dialogues. For example, if two dialogues revolve around the same theme and their content has a causal or progressive relationship, they can be grouped into the same dialogue sequence. At the same time, the dialogue sequence is further optimized by referring to the service processes and business rules in the historical service records. For example, when handling customer complaints, it is usually necessary to first understand the details of the problem, then propose and implement solutions, and finally conduct customer feedback and satisfaction surveys. According to this business rule, the relevant dialogue messages can be sorted according to this process to determine their position in the dialogue sequence.

[0043] In some embodiments of this application, the dialogue sequence is determined by using the context of the dialogue messages and historical service records, which can comprehensively grasp the dialogue context and the evolution of user needs, making dialogue processing more coherent and targeted, and improving service accuracy and user satisfaction.

[0044] Furthermore, based on the above embodiments, please refer to... Figure 4 In one exemplary embodiment provided in this application, the specific implementation process of performing multimodal fusion on the dialogue messages to obtain the fused dialogue messages may further include steps S410 and S420, which are described in detail below: Step S410: Obtain the multimodal data stream from the dialogue message. The multimodal data stream includes multiple components such as images, text, and speech. Step S420: Perform cross-modal fusion on the multimodal data stream, and determine the fused dialogue message based on the fusion result and dialogue timing.

[0045] For example, after determining the dialogue sequence, to obtain the multimodal data stream from the dialogue messages and perform cross-modal fusion, and then determine the fused dialogue message based on the fusion result and the dialogue sequence, the following specific implementation process can be followed: First, build a multimodal data stream acquisition module. This module needs to be tightly integrated with the message hub deeply coupled with the WeChat ecosystem and the previously built dialogue monitoring system. Since the WeChat ecosystem supports multiple message types, such as images, text, and voice, the message hub can capture data from these different modalities in real time and transmit it to the data stream acquisition module. For image data, the message hub will detect the image message sent by the user during the dialogue and transmit it completely to the data stream. The acquisition module performs preliminary image processing, such as format conversion and compression, to ensure efficiency and compatibility for subsequent processing. For text data, the message processing unit directly extracts the text content from the dialogue, including plain text messages and text converted through speech recognition. The data stream acquisition module performs preprocessing operations such as word segmentation and stop word removal to extract key information. For audio data, the message processing unit acquires the original audio file. The data stream acquisition module performs operations such as pre-emphasis, framing, and windowing on the audio to enhance the quality and stability of the audio signal, while extracting audio feature parameters, such as Mel-frequency cepstral coefficients (MFCC), for subsequent cross-modal fusion. After acquiring the multimodal data stream, the cross-modal fusion stage begins. The goal of cross-modal fusion is to organically combine data from different modalities to extract feature information that comprehensively reflects the content of the dialogue.

[0046] In some embodiments of this application, acquiring and fusing multimodal dialogue data streams across modalities, and combining the dialogue timing to determine the fused dialogue messages, can comprehensively and accurately capture the core content and logic of the dialogue, improve the integrity and accuracy of information processing, and optimize the dialogue interaction effect.

[0047] Furthermore, based on the above embodiments, please refer to... Figure 5 In one exemplary embodiment provided in this application, the specific implementation process of the above-mentioned artificial intelligence-based enterprise service method may further include steps S510 to S540, which are described in detail below: Step S510: Based on the cross-modal attention mechanism, determine the cross-modal attention output features corresponding to the text features and image features in the dialogue message. The cross-modal attention output features represent the attention result of the text features on the image features. Step S520: Map the cross-modal attention output features and text features to the gating space respectively to obtain the corresponding gating coefficients; Step S530: Determine the fusion ratio of text features and image features in cross-modal feature fusion based on the gating coefficient, so as to determine the target fusion strategy between text features and image features based on the fusion ratio; Step S540: Perform cross-modal fusion of multimodal data streams based on the target fusion strategy.

[0048] For example, the goal of a cross-modal attention mechanism is to enable text features to focus on important information in image features, thereby obtaining cross-modal attention output features. First, the text features and image features are linearly transformed to make their dimensions consistent, facilitating subsequent calculations. Then, the similarity matrix between the text features and image features is calculated, using methods such as dot product and cosine similarity. The similarity matrix reflects the correlation between each element in the text features and image features. Next, the similarity matrix is ​​normalized to obtain the attention weight matrix. Each element in the attention weight matrix represents the degree of attention of the text features to the corresponding element in the image features. The attention weight matrix and the image features are weighted and summed to obtain the cross-modal attention output features. This feature represents the attention result of the text features to the image features and can capture the semantic association between the text and the image.

[0049] Optionally, in some feasible embodiments, after obtaining the cross-modal attention output features and text features, they need to be mapped to a gating space to obtain the corresponding gating coefficients. The gating space is a space used to control the feature fusion ratio. Features can be mapped to the gating space through a fully connected layer. During the mapping process, activation functions such as the Sigmoid function can be used to restrict the mapped values ​​to between 0 and 1 as gating coefficients. The gating coefficients represent the importance of the corresponding features in cross-modal feature fusion. For the cross-modal attention output features and text features, the above mapping operation is performed to obtain their corresponding gating coefficients. Based on the obtained gating coefficients, the fusion ratios of text features and image features in cross-modal feature fusion are determined. The fusion ratio can be calculated using a simple weighted average method, where the gating coefficients are used as weights to sum the cross-modal attention output features and text features to obtain the fused features. For example, if the gating coefficient of the text features is 0.6 and the gating coefficient of the cross-modal attention output features is 0.4, then the fused features can be represented as 0.6 × text features + 0.4 × cross-modal attention output features. In this way, the dynamic fusion ratio determination based on the gating coefficients is achieved, enabling the fused features to better combine the information from text and images. Based on the determined fusion ratio, a target fusion strategy is formulated between text features and image features. The target fusion strategy can be adjusted according to specific business needs and fusion effects. For example, in some scenarios, more emphasis may be placed on the semantic information of the text, in which case the fusion ratio of text features can be increased; while in other scenarios, the visual information of the image may be more important, in which case the fusion ratio of cross-modal attention output features can be increased. At the same time, some constraints, such as feature sparsity and orthogonality, can also be introduced to optimize the fusion strategy and improve the fusion effect.

[0050] Then, based on the target fusion strategy, cross-modal fusion is performed on the multimodal data stream. In addition to text and image features, if the multimodal data stream also contains data from other modalities such as speech, similar methods can be used for feature extraction and fusion. For speech feature extraction, methods such as Mel-frequency cepstral coefficients (MFCC) can be used to extract speech feature parameters, which are then mapped to the gating space to obtain the corresponding gating coefficients. These are then fused with features from other modalities according to the fusion strategy. During the fusion process, attention needs to be paid to the time synchronization between features from different modalities to ensure that the fused features can accurately reflect the actual content of the dialogue. In this way, comprehensive fusion of the multimodal data stream is achieved, resulting in fused features that comprehensively reflect the dialogue messages, providing richer information for subsequent dialogue understanding and analysis.

[0051] In some embodiments of this application, cross-modal attention output features are determined based on a cross-modal attention mechanism, and these features are mapped to a gating space with text features to obtain gating coefficients. This allows for the determination of the fusion ratio and target fusion strategy, which can accurately capture the correlation and weight between text and image features, achieve efficient and accurate fusion of multimodal data streams, improve the quality and information richness of fused features, and provide a more reliable data foundation for subsequent processing.

[0052] Furthermore, based on the above embodiments, please refer to... Figure 6 In one exemplary embodiment provided in this application, the specific implementation process of determining the response mechanism corresponding to the target event based on the industry-level federated knowledge base may further include steps S610 to S630, which are described in detail below: Step S610: Based on the underlying industry-level federated knowledge base, historical consulting cases are collected using differential privacy technology to obtain an industry knowledge graph; Step S620: Determine the triggering engine corresponding to the target event, and obtain the related content corresponding to the target event based on the triggering engine and the industry knowledge graph; Step S630: Determine the response mechanism corresponding to the target event based on the related content.

[0053] For example, the first step is to build a foundational industry-level federated knowledge base. This is a vast knowledge system that brings together the knowledge resources of multiple relevant industry participants. Under the premise of protecting their own data privacy, each participant can achieve knowledge sharing and collaboration through technologies such as federated learning. Since data from different institutions is involved, differential privacy technology is used when collecting historical consulting cases to ensure data security and privacy. Differential privacy adds carefully designed noise to the data, which effectively prevents individual data records from being identified and leaked without affecting the overall statistical characteristics of the data. For example, the historical consulting case data provided by each participant is first preprocessed, including data cleaning and format standardization. Then, differential privacy algorithms, such as the Laplace mechanism, are used to perturb the data. During the collection process, the amount of noise added is strictly controlled to achieve a balance between privacy protection and data usability. The collected historical consulting case data, processed with differential privacy, is integrated into an industry-level federated knowledge base. Subsequently, knowledge graph construction technology is used to extract entities, attributes, and relationships from these cases to construct an industry knowledge graph. The knowledge graph can clearly display various concepts, entities, and their relationships within the industry, providing structured knowledge support for subsequent analysis and processing. Next, the triggering engine corresponding to the target event is determined. The triggering engine is one of the key components of the entire system. It is responsible for identifying and triggering the processing flow related to the target event according to preset rules and conditions.

[0054] Optionally, when determining the trigger engine, multiple factors such as the type, source, and urgency of the target event need to be considered comprehensively. For example, for customer consultation events, the trigger engine can match and judge based on keywords and semantic features in the consultation content. For system alarm events, the trigger engine can trigger based on parameters such as the alarm level and type. Once the target event is identified by the trigger engine, the system will initiate an interaction process with the industry knowledge graph. By querying and reasoning in the knowledge graph, it will find related content to the target event. Based on the entities and relationships of the target event, it will perform deep traversal and correlation analysis in the knowledge graph. This related content may include similar historical cases, relevant industry standards, solutions, etc., which provide rich reference for the handling of the target event. Finally, based on the obtained related content, the corresponding response mechanism for the target event is determined. The response mechanism is a set of processing procedures and strategies for the target event, which needs to comprehensively consider various information in the related content, such as the processing results of historical cases and the requirements of industry standards. At the same time, the response mechanism also needs to have a certain degree of flexibility and adjustability to adapt to the special needs of different scenarios. During implementation, the response mechanism can be optimized and improved by continuously collecting feedback information to ensure that it can handle various target events efficiently and accurately.

[0055] In some embodiments of this application, an industry knowledge graph is constructed by leveraging an underlying industry-level federated knowledge base and differential privacy technology. This allows for the secure aggregation of industry knowledge. Then, a triggering engine is used to combine the knowledge graph with relevant content to determine a response mechanism. This enables efficient and accurate provision of appropriate and secure response strategies for target events, thereby improving industry service and decision-making levels.

[0056] Furthermore, based on the above embodiments, please refer to... Figure 7 In one exemplary embodiment provided in this application, the specific implementation process of sending the target event based on the response mechanism, monitoring the entire lifecycle dialogue of the target event, and storing the generated log information may further include steps S710 and S720, which are described in detail below: Step S710: Trigger the full lifecycle dialogue management system based on the response mechanism. The full lifecycle dialogue management system includes a dialogue state tracker, an intent association analyzer, and a service strategy optimizer. Step S720: Based on the full lifecycle dialogue management system, determine the service strategy corresponding to the target event, and adjust the depth of robot intervention and the level of detail in the response based on the job responsibilities and permission level of the target employee through the service strategy optimizer.

[0057] For example, when the response mechanism is triggered, the full lifecycle dialogue management system begins to operate. The dialogue state tracker, as one of the core components of the system, is responsible for monitoring the progress and status of the dialogue in real time. Through close integration with the message hub, it obtains real-time data of the dialogue, including the start time of the dialogue, participants, and current dialogue content. The dialogue state tracker records and analyzes this data in detail to identify key nodes and turning points in the dialogue. For example, when a customer raises a new question or request, the dialogue state tracker marks it as a new dialogue stage and stores the relevant information of that stage. At the same time, the dialogue state tracker also judges the urgency and importance of the dialogue based on its semantics and context, providing basic data for subsequent intent association analysis. Next, the intent association analyzer comes into play. Utilizing natural language processing and machine learning algorithms, it performs in-depth analysis of the dialogue content to identify customer intents and needs. The intent association analyzer extracts keywords, phrases, and semantic features from the dialogue text and matches them with a pre-trained intent classification model to determine the customer's primary intent. Simultaneously, it combines information from the dialogue state tracker to analyze the correlation and evolution of customer intents. For example, if a customer first inquires about product price and then mentions after-sales service, the intent association analyzer will identify the potential correlation between these two intents, indicating that the customer may be more concerned about the overall cost-effectiveness of the product and service guarantees. Through this analysis, the intent association analyzer can more accurately understand the customer's true needs, providing strong support for developing service strategies. Based on the analysis results from the dialogue state tracker and intent association analyzer, the service strategy optimizer begins to formulate service strategies corresponding to the target events. The service strategy optimizer comprehensively considers multiple factors, including customer intent, needs, dialogue state, and historical service records, and uses technologies such as decision trees and rule engines to generate optimal service strategies. For example, if the customer's needs are urgent and important, the service strategy optimizer may suggest prioritizing the intervention of human customer service to improve the efficiency of problem-solving; if the customer's needs are relatively simple and clear, the service strategy optimizer may suggest that a chatbot provide an automated response to save labor costs. At the same time, the service strategy optimizer will also optimize and adjust the service strategies according to business rules and objectives to ensure that they conform to the overall interests of the enterprise and the needs of customers.

[0058] Optionally, during the service strategy development process, the service strategy optimizer also considers the job responsibilities and authority levels of the target employees to adjust the depth of the robot's intervention and the level of detail in its responses. Employees with different job responsibilities and authority levels possess varying levels of authority and expertise when handling customer issues. Therefore, the service strategy optimizer personalizes the robot's behavior based on these differences. For example, frontline customer service staff may primarily handle common questions and needs, and the robot can provide more detailed responses and solutions to help them resolve issues quickly. Senior managers, on the other hand, may focus more on strategic decisions and business analysis, and the robot can provide more concise and generalized information to help them quickly understand the key points of the conversation and make decisions. In this way, the service strategy optimizer can flexibly adjust the robot's depth of intervention and level of detail in its responses according to the characteristics and needs of different employees, thereby improving service efficiency and quality.

[0059] Furthermore, in some feasible embodiments, the defined service strategies are applied to actual dialogue processing. During the dialogue, the chatbot automatically or semi-automatically responds to customer questions based on the service strategies. Simultaneously, the service strategy optimizer monitors the effectiveness of the dialogue and customer feedback in real time, dynamically adjusting the service strategies according to the actual situation. For example, if the customer is dissatisfied with the chatbot's response, the service strategy optimizer may promptly adjust the chatbot's response strategy, increasing the intervention of human customer service or optimizing the chatbot's response content. Through this dynamic adjustment mechanism, the full lifecycle dialogue management system can continuously optimize service strategies, improve customer satisfaction and service quality, and create greater value for the enterprise. Throughout the implementation process, a robust monitoring and management mechanism is also required to monitor and manage the operational status of the full lifecycle dialogue management system and the execution of service strategies in real time, promptly identifying and resolving potential problems to ensure system stability and reliability. Additionally, the system needs to be regularly evaluated and optimized, continuously improving and refining its functionality and performance based on business development and changes in customer needs.

[0060] In one of the exemplary embodiments provided in this application, a full lifecycle dialogue management system covering multiple modules is triggered based on a response mechanism to determine service strategies and adjust the robot's intervention and response according to the target employee's job responsibilities and permission level. This can achieve intelligent and personalized handling of target events, improve service efficiency and quality, and at the same time ensure enterprise information security and compliance.

[0061] Furthermore, based on the above embodiments, please refer to... Figure 8In one exemplary embodiment provided in this application, the specific implementation process of the above-mentioned artificial intelligence-based enterprise service method may further include steps S810 to S830, which are described in detail below: Step S810: parse the voice message to determine the voiceprint information corresponding to the voice message based on the parsing result; Step S820: Determine the target employee and the corresponding identity identifier based on the voiceprint information; Step S830: Obtain the enterprise organizational structure diagram and determine the job responsibilities and authority levels of the target employees based on the enterprise organizational structure diagram and identity identifiers.

[0062] For example, in a smart office scenario, to achieve accurate processing and access control of employee voice commands, the following process is required: First, voice message parsing is performed to determine voiceprint information. When an employee issues a voice command through a smart device, the system preprocesses the voice message, using filtering algorithms to remove background noise, such as keyboard clicks and conversations in the office environment. Then, pre-emphasis processing is performed to enhance the energy of the high-frequency part of the voice signal, making the voice spectrum more balanced. After that, the voice is segmented into short time frames and windowed, typically using a Hamming window to reduce spectral leakage. After preprocessing, the Mel-frequency cepstral coefficient (MFCC) algorithm is used to extract voice features. MFCC can simulate the characteristics of human hearing and extract feature parameters that reflect the essence of voice. Then, voiceprint recognition technology is used to match the extracted features with a pre-built voiceprint model library. The voiceprint model library is trained by collecting a large number of voice samples from enterprise employees. By calculating the similarity between the feature vector and the model, the voiceprint information corresponding to the voice message is determined. Voiceprints are unique and stable, and can accurately identify the speaker. After confirming the voiceprint information, the target employee and their identity are further identified. The system has a built-in database mapping employee voiceprints to identity identifiers. When voiceprint information is obtained, it is queried and matched in the database. To ensure accurate identity recognition, a multiple verification mechanism is used, such as having the employee repeat the voice command or cross-verifying with facial recognition. If multiple voiceprint matching results are consistent and facial recognition passes, the target employee's identity identifier is confirmed, such as employee ID "EMP1001". Next, the company's organizational structure diagram is obtained. This diagram can be exported from the company's human resource management system and is represented by a tree structure. The root node is the company's top management, the branch nodes are the departments, and the leaf nodes are specific employees. The nodes are connected through reporting relationships, clearly showing the company's internal organizational structure and hierarchical relationships. Finally, based on the enterprise's organizational chart and identity identifier, the job responsibilities and permission levels are determined. The system locates the employee node in the enterprise's organizational chart based on the identity identifier "EMP1001". It finds that the employee belongs to the "Technology R&D Department" and the position is "Senior Software Engineer". Combined with the predefined job description of the enterprise, it is determined that the employee's main responsibilities include software system design, code writing and testing. According to the permission level settings in the organizational chart, the employee is at the middle permission level. In the system, the employee can access and operate functional modules and data related to software development, such as code repositories and testing environments, but cannot access sensitive information such as financial data and personnel decisions that can only be accessed by senior management.

[0063] In some embodiments of this application, the system achieves accurate processing and reasonable access control of employee voice commands in this way, ensuring enterprise information security and efficient business operation. At the same time, a monitoring and management mechanism needs to be established throughout the process to track the operating status of each link in real time and regularly evaluate and optimize the system to ensure accuracy and effectiveness.

[0064] Figure 9 This is a block diagram illustrating an exemplary embodiment of this application. The device can be applied to... Figure 1 The implementation environment shown is specifically configured in server 110. This device can also be applied to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment to which this device is applicable.

[0065] like Figure 9 As shown, this exemplary AI-based enterprise service system includes: a listening module 910, used to listen to dialogue messages corresponding to multiple enterprise groups, the dialogue messages including voice messages, and determine the identity information corresponding to the voice messages; a fusion module 920, used to perform multimodal fusion of the dialogue messages to obtain fused dialogue messages; a determination module 930, used to determine the target event corresponding to the fused dialogue messages, and determine the response mechanism corresponding to the target event based on an industry-level federated knowledge base; and a delivery module 940, used to deliver the target event based on the response mechanism and identity information, and monitor the entire lifecycle of the target event dialogue and store the generated log information.

[0066] According to one aspect of the embodiments of this application, the listening module 910 is further configured to: obtain context information of the dialogue message and historical service records; and determine the dialogue sequence corresponding to the dialogue message based on the context information and historical service records.

[0067] According to one aspect of the embodiments of this application, the fusion module 920 is further configured to: acquire a multimodal data stream in the dialogue message, the multimodal data stream including multiple items such as images, text and speech; perform cross-modal fusion on the multimodal data stream; and determine the fused dialogue message based on the fusion result and the dialogue timing.

[0068] According to one aspect of the embodiments of this application, the fusion module 920 is further configured to: determine cross-modal attention output features corresponding to text features and image features based on a cross-modal attention mechanism, wherein the cross-modal attention output features represent the attention result of text features on image features; map the cross-modal attention output features and text features to a gating space respectively to obtain corresponding gating coefficients; determine the fusion ratios of text features and image features in cross-modal feature fusion based on the gating coefficients, so as to determine the target fusion strategy between text features and image features based on the fusion ratios; and perform cross-modal fusion on the multimodal data stream based on the target fusion strategy.

[0069] According to one aspect of the embodiments of this application, the determining module 930 is further configured to: collect historical consulting cases based on the underlying industry-level federated knowledge base using differential privacy technology to obtain an industry knowledge graph; determine the triggering engine corresponding to the target event, and obtain the associated content corresponding to the target event based on the triggering engine and the industry knowledge graph; and determine the response mechanism corresponding to the target event based on the associated content.

[0070] According to one aspect of the embodiments of this application, the delivery module 940 is further configured to: trigger a full lifecycle dialogue management system based on a response mechanism; the full lifecycle dialogue management system includes a dialogue state tracker, an intent association analyzer, and a service strategy optimizer; determine the service strategy corresponding to the target event based on the full lifecycle dialogue management system; and adjust the robot intervention depth and response detail level based on the job responsibilities and permission level of the target employee through the service strategy optimizer.

[0071] According to one aspect of the embodiments of this application, the above-mentioned monitoring module 910 is further configured to: parse voice messages to determine the voiceprint information corresponding to the voice messages based on the parsing results; determine the target employee and the identity identifier corresponding to the target employee based on the voiceprint information; obtain the enterprise organizational structure map, and determine the job responsibilities and authority level corresponding to the target employee based on the enterprise organizational structure map and the identity identifier.

[0072] It should be noted that the AI-based enterprise service system and the AI-based enterprise service method provided in the above embodiments belong to the same concept. The specific methods by which each module and unit performs operations have been described in detail in the method embodiments and will not be repeated here. In practical applications, the AI-based enterprise service system provided in the above embodiments can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this is not a limitation here.

[0073] Embodiments of this application also provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to implement the artificial intelligence-based enterprise service method provided in the above embodiments.

[0074] Figure 10 A schematic diagram of a computer system suitable for implementing the embodiments of this application is shown. It should be noted that... Figure 10 The computer system 1000 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0075] like Figure 10As shown, the computer system 1000 includes a Central Processing Unit (CPU) 1001, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in Read-Only Memory (ROM) 1002 or programs loaded from storage portion 1008 into Random Access Memory (RAM) 1003. The RAM 1003 also stores various programs and data required for system operation. The CPU 1001, ROM 1002, and RAM 1003 are interconnected via a bus 1004. An Input / Output (I / O) interface 1005 is also connected to the bus 1004.

[0076] The following components are connected to I / O interface 1005: an input section 1006 including a keyboard, mouse, etc.; an output section 1007 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to I / O interface 1005 as needed. Removable media 1011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1010 as needed so that computer programs read from them can be installed into storage section 1008 as needed.

[0077] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1009, and / or installed from removable medium 1011. When the computer program is executed by central processing unit (CPU) 1001, it performs various functions defined in the system of this application.

[0078] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0079] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0080] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0081] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned AI-based enterprise service method. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not incorporated into the electronic device.

[0082] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the AI-based enterprise service method provided in the various embodiments described above.

[0083] The above description is merely a preferred exemplary embodiment of this application and is not intended to limit the implementation of this application. Those skilled in the art can easily make corresponding modifications or alterations based on the main concept and spirit of this application. Therefore, the scope of protection of this application should be determined by the scope of protection claimed in the claims.

Claims

1. An enterprise service method based on artificial intelligence, characterized in that, include: Monitor conversation messages corresponding to multiple enterprise groups, including voice messages, and determine the identity information corresponding to the voice messages; The dialogue messages are fused using multimodal methods to obtain the fused dialogue messages. The target event corresponding to the fused dialogue message is determined, and the response mechanism corresponding to the target event is determined based on an industry-level federated knowledge base; Based on the response mechanism and the identity information, the target event is delivered, and the entire lifecycle of the target event's dialogue is monitored and the generated log information is stored.

2. The method as described in claim 1, characterized in that, The method further includes: Obtain the context information of the dialogue message and historical service records; The dialogue sequence corresponding to the dialogue message is determined based on the context information and the historical service records.

3. The method as described in claim 2, characterized in that, The process of performing multimodal fusion on the dialogue messages to obtain fused dialogue messages includes: Obtain the multimodal data stream from the dialogue message, wherein the multimodal data stream includes multiple components such as images, text, and speech; The multimodal data stream is fused across modes, and the fused dialogue message is determined based on the fusion result and the dialogue timing.

4. The method as described in claim 3, characterized in that, The method further includes: Based on the cross-modal attention mechanism, the cross-modal attention output features corresponding to the text features and image features in the dialogue message are determined, and the cross-modal attention output features represent the attention result of the text features on the image features; The cross-modal attention output features and the text features are mapped to the gating space to obtain the corresponding gating coefficients; Based on the gating coefficient, the fusion ratios of the text features and the image features in cross-modal feature fusion are determined, and the target fusion strategy between the text features and the image features is determined based on the fusion ratios. Cross-modal fusion of the multimodal data stream is performed based on the target fusion strategy.

5. The method as described in claim 1, characterized in that, The response mechanism for determining the target event based on an industry-level federated knowledge base includes: An industry knowledge graph is obtained by collecting historical consulting cases based on an underlying industry-level federated knowledge base using differential privacy technology. Identify the triggering engine corresponding to the target event, and obtain the associated content corresponding to the target event based on the triggering engine and the industry knowledge graph; The response mechanism corresponding to the target event is determined based on the aforementioned related content.

6. The method as described in claim 5, characterized in that, The step of delivering the target event based on the response mechanism, monitoring the entire lifecycle of the target event's dialogue, and storing the generated log information includes: The response mechanism triggers the full lifecycle dialogue management system, which includes a dialogue state tracker, an intent association analyzer, and a service strategy optimizer. Based on the full lifecycle dialogue management system, the service strategy corresponding to the target event is determined, and the service strategy optimizer adjusts the depth of robot intervention and the level of detail in the response based on the job responsibilities and permission level of the target employee.

7. The method as described in claim 6, characterized in that, The method further includes: The voice message is parsed to determine the voiceprint information corresponding to the voice message based on the parsing results; The target employee and the corresponding identity identifier are determined based on the voiceprint information. Obtain the enterprise organizational structure diagram, and determine the job responsibilities and authority levels corresponding to the target employee based on the enterprise organizational structure diagram and the identity identifier.

8. An enterprise service system based on artificial intelligence, characterized in that, The system includes: The monitoring module is used to monitor conversation messages corresponding to multiple enterprise groups, including voice messages, and to determine the identity information corresponding to the voice messages; The fusion module is used to perform multimodal fusion on the dialogue messages to obtain fused dialogue messages; The determination module is used to determine the target event corresponding to the fused dialogue message, and to determine the response mechanism corresponding to the target event based on an industry-level federated knowledge base; The delivery module is used to deliver the target event based on the response mechanism and the identity information, and to monitor the entire lifecycle of the target event's dialogue and store the generated log information.

Citation Information

Patent Citations

  • Voice service pushing methods, device and systems based on artificial intelligence

    CN109618068A

  • Service provider dialogue processing method and device, storage medium and electronic equipment

    CN117171316A

  • Heat supply enterprise intelligent customer service implementation system based on artificial intelligence

    CN117829845A

  • Intelligent customer service method and system based on multi-modal dynamic fusion

    CN120471628A