Branch operation method, device, equipment, medium and program product
By collecting and analyzing users' multimodal features, combining machine learning models to output emotion intensity values and implement emotion intervention plans, the problem of uneven resource allocation in offline bank branch operations is solved, and operational efficiency and service quality are improved.
Patent Information
- Application Number
- CN202411901185.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, the operating efficiency and service level of offline bank branches are low, mainly due to the subjectivity and lag of user evaluations, which leads to uneven resource allocation and inability to promptly identify and handle abnormal user emotions, resulting in low operational efficiency.
By obtaining the user's multimodal feature authorization, collecting voice, facial image and body image features, and using machine learning models to output the emotion intensity value, the corresponding preset emotion intervention plan is executed when the emotion intensity is in the abnormal range, including monitoring, artificial intelligence soothing and manual intervention, and resource allocation is adjusted in real time.
It achieves more accurate emotional state judgment, monitors user emotional changes in real time, improves branch processing efficiency, releases human resources, avoids resource waste, and improves operational efficiency and service levels.
Smart Images

Figure CN120672176A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and financial technology, and specifically to a network operation method, device, equipment, medium and program product. Background Art
[0002] As bank branches' operational requirements increase, they often collect customer reviews to improve customer service. This feedback is used to improve service levels. However, user reviews often have limitations, including subjectivity and lags, making them increasingly difficult to meet branch operational needs.
[0003] Currently, some technologies exist that use multimodal feature fusion to capture customers' true status and objectively score them. However, existing technologies simply use objective user scores as the end / closed loop for branch operations, without considering the various operational needs of branches. This, combined with the persistent technical issue of uneven resource allocation across the entire branch operation system, ultimately leads to low branch operating efficiency and service levels. Summary of the Invention
[0004] In view of the above problems, the present disclosure provides a network operation method, device, equipment, medium and program product for improving network operation efficiency and service level.
[0005] According to a first aspect of the present disclosure, a network operation method is provided, including: obtaining a user's authorization for multimodal feature collection; collecting the user's multimodal features when the user has authorized the multimodal feature collection; outputting the user's emotion intensity value based on the multimodal features; and executing a corresponding preset emotion intervention plan when the user's emotion intensity value is within a preset abnormal emotion range.
[0006] According to an embodiment of the present disclosure, after executing the corresponding preset emotion intervention plan, the method further includes: counting the changes in emotion intensity values, the amplitude of emotion fluctuations, and the time for processing abnormal emotions; and outputting a service capability evaluation based on the changes in emotion intensity values, the amplitude of emotion fluctuations, and the time for processing abnormal emotions.
[0007] According to an embodiment of the present disclosure, the preset abnormal emotion interval includes a first abnormal emotion interval, and when the user's emotion intensity value is in the preset abnormal emotion interval, the corresponding preset emotion intervention plan is executed, including: when the user's emotion intensity value is in the first abnormal emotion interval, based on the multimodal features, by matching the preset subject content, the cause of the abnormal emotion is derived.
[0008] According to an embodiment of the present disclosure, the preset abnormal emotion interval also includes a second abnormal emotion interval, and when the user's emotion intensity value is in the preset abnormal emotion interval, executing the corresponding preset emotion intervention plan also includes: when the user's emotion intensity value is in the second abnormal emotion interval, calling an artificial intelligence service; and outputting soothing content based on the artificial intelligence service.
[0009] According to an embodiment of the present disclosure, the preset abnormal emotion interval also includes a third abnormal emotion interval, and when the user's emotion intensity value is in the preset abnormal emotion interval, executing the corresponding preset emotion intervention plan also includes: when the user's emotion intensity value is in the third abnormal emotion interval, issuing an artificial intervention signal.
[0010] According to an embodiment of the present disclosure, wherein the multimodal features include N different modal features, N is a positive integer, and outputting the user's emotion intensity value based on the multimodal features includes: inputting the N different modal features into a preset machine learning model to obtain N different emotion intensity values; and performing weighting on the N different emotion intensity values to obtain the emotion intensity value.
[0011] According to an embodiment of the present disclosure, the preset machine learning model includes N preset machine learning sub-models, and the preset machine learning sub-models correspond one-to-one to any modality in the multimodal features. The training method of the preset machine learning sub-model includes: for any of the above, obtaining an initial machine sub-learning model and a corresponding training data set, wherein the training data set includes an emotion intensity value label and a corresponding modal feature; and obtaining the preset machine learning sub-model based on the initial machine sub-learning model and the training data set.
[0012] A second aspect of the present disclosure provides an outlet operation device, comprising: an authorization module for obtaining a user's authorization for multimodal feature collection; a collection module for collecting the user's multimodal features upon obtaining the user's authorization for multimodal feature collection; an emotion intensity output module for outputting the user's emotion intensity value based on the multimodal features; and an emotion intervention module for executing a corresponding preset emotion intervention plan when the user's emotion intensity value is within a preset abnormal emotion range.
[0013] According to an embodiment of the present disclosure, the device also includes: a service evaluation module, which is used to count the changes in emotion intensity values, the amplitude of emotion fluctuations and the time for processing abnormal emotions; and output a service capability evaluation based on the changes in emotion intensity values, the amplitude of emotion fluctuations and the time for processing abnormal emotions.
[0014] According to an embodiment of the present disclosure, the preset abnormal emotion interval includes a first abnormal emotion interval, and the emotion intervention module is specifically used to derive the cause of the abnormal emotion based on the multimodal features by matching the preset subject content when the user's emotion intensity value is in the first abnormal emotion interval.
[0015] According to an embodiment of the present disclosure, the preset abnormal emotion interval also includes a second abnormal emotion interval, and the emotion intervention module is further specifically used to call an artificial intelligence service when the user's emotion intensity value is in the second abnormal emotion interval; and output soothing content based on the artificial intelligence service.
[0016] According to an embodiment of the present disclosure, the preset abnormal emotion interval further includes a third abnormal emotion interval, and the emotion intervention module is further specifically configured to send a manual intervention signal when the user's emotion intensity value is in the third abnormal emotion interval.
[0017] According to an embodiment of the present disclosure, the multimodal features include N different modal features, where N is a positive integer, and the emotion intensity output module is specifically used to input the N different modal features into a preset machine learning model to obtain N different emotion intensity values; and perform weighting on the N different emotion intensity values to obtain the emotion intensity value.
[0018] According to an embodiment of the present disclosure, the preset machine learning model includes N preset machine learning sub-models, and the preset machine learning sub-models correspond one-to-one to any modality in the multimodal features. The device also includes a training module for obtaining an initial machine sub-learning model and a corresponding training data set for any of the above, wherein the training data set includes an emotion intensity value label and a corresponding modal feature; and obtaining the preset machine learning sub-model based on the initial machine sub-learning model and the training data set.
[0019] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0020] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0021] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0022] In the embodiment of the present disclosure, in order to solve the technical problem of low operating efficiency and quality of the outlets, the embodiment of the present disclosure first collects the multimodal features of the user, then outputs an accurate emotion intensity value in combination with the multimodal features, and finally arranges a corresponding emotion intervention plan for the user according to the interval of the emotion intensity value, which can ensure the improvement of operational efficiency and quality. The beneficial effects are: 1. Outputting the emotion intensity value through multimodal features can provide additional information beyond a single modality, so as to more accurately judge the user's emotional state; 2. Real-time monitoring of the user's emotional changes and providing instant feedback based on the emotional state ensures the efficiency of the outlet processing, releases a large amount of human resources, helps the outlet staff to understand the user's emotional state more comprehensively, and focus on processing to avoid the fermentation of user emotions; 3. Through the emotional triggering objective allocation plan, various resources in the outlet system are allocated to where they are currently needed, avoiding waste of resources in the outlet operation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0024] Figure 1 A diagram schematically illustrates an application scenario of the network operation method according to an embodiment of the present disclosure;
[0025] Figure 2 The flowchart of the network operation method according to the embodiment of the present disclosure is schematically shown;
[0026] Figure 3 A schematic diagram of a structure of a network operation device according to an embodiment of the present disclosure is shown; and
[0027] Figure 4 A block diagram of an electronic device suitable for implementing a network operation method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0028] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0029] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0031] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0032] In the technical solutions disclosed herein, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0033] In scenarios where personal information is used for automated decision-making, the methods, devices, and systems provided by the embodiments of the present disclosure all provide users with corresponding operation portals for them to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge, and skills, and have reached a certain level of professionalism.
[0034] To better serve customers, offline branches conduct satisfaction assessments to obtain user feedback on the branches, monitor the service levels of branch staff, and improve branch service capabilities. However, the current assessment of the service levels of offline branches and branch staff still has the following limitations:
[0035] First, offline outlets currently rely primarily on post-service user ratings to gauge customer satisfaction. This method is subjective and single-dimensional, making it difficult to truly reflect the outlet's service level.
[0036] Secondly, most outlets currently use a post-evaluation method, that is, after the service is completed, users are asked to make a subjective evaluation of the service. The node for obtaining the evaluation is delayed, and it is impossible to identify risk events caused by abnormal user attitudes and intervene in time, which can easily cause users to leave the outlet with negative emotions, which has a negative impact on the image of the outlet and the entire enterprise.
[0037] In summary, the existing network operation solutions are inefficient and have low service levels.
[0038] Therefore, it is very necessary to have a system that can timely identify the user's emotional state, issue early warnings for abnormal conditions and provide appropriate intervention, while also being able to more objectively evaluate the service level of outlets.
[0039] In order to solve the technical problems existing in the prior art, an embodiment of the present disclosure provides a network operation method, which includes: obtaining the user's authorization for multimodal feature collection; collecting the user's multimodal features when the user's authorization for multimodal feature collection is obtained; outputting the user's emotion intensity value based on the multimodal features; and executing a corresponding preset emotion intervention plan when the user's emotion intensity value is within a preset abnormal emotion range.
[0040] In the embodiment of the present disclosure, in order to solve the technical problem of low operating efficiency and quality of the outlets, the embodiment of the present disclosure first collects the multimodal features of the user, then outputs an accurate emotion intensity value in combination with the multimodal features, and finally arranges a corresponding emotion intervention plan for the user according to the interval of the emotion intensity value, which can ensure the improvement of operational efficiency and quality. The beneficial effects are: 1. Outputting the emotion intensity value through multimodal features can provide additional information beyond a single modality, so as to more accurately judge the user's emotional state; 2. Real-time monitoring of the user's emotional changes and providing instant feedback based on the emotional state ensures the efficiency of the outlet processing, releases a large amount of human resources, helps the outlet staff to understand the user's emotional state more comprehensively, and focus on processing to avoid the fermentation of user emotions; 3. Through the emotional triggering objective allocation plan, various resources in the outlet system are allocated to where they are currently needed, avoiding waste of resources in the outlet operation system.
[0041] Figure 1 The application scenario diagram of the network operation method according to the embodiment of the present disclosure is schematically shown.
[0042] like Figure 1As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0043] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0044] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0045] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0046] It should be noted that the outlet operation method provided in the embodiments of the present disclosure can generally be executed by the server 105. Accordingly, the outlet operation apparatus provided in the embodiments of the present disclosure can generally be disposed in the server 105. The outlet operation method provided in the embodiments of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and that is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Accordingly, the outlet operation apparatus provided in the embodiments of the present disclosure can also be disposed in a server or server cluster that is different from the server 105 and that is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0047] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0048] The following will be based on Figure 1 The scene described by Figure 2 The network operation method of the disclosed embodiment is described in detail.
[0049] Figure 2 The flowchart of the network operation method according to the embodiment of the present disclosure is schematically shown.
[0050] like Figure 2 As shown, the outlet operation method of this embodiment includes operations S210 to S240 , and the outlet operation method can be executed by the server 105 .
[0051] In operation S210 , the user's authorization for multimodal feature collection is obtained.
[0052] In an embodiment of the present disclosure, before obtaining the user's information in operation S210, the user's consent or authorization may be obtained. For example, a request to obtain the user's information may be issued to the user before the operation S210. If the user agrees or authorizes the acquisition of the user's information, the operation S220 is performed.
[0053] In operation S220 , after obtaining authorization from the user for collecting multimodal features, the multimodal features of the user are collected.
[0054] Specifically, the multimodal features include: one or more of speech features, facial image features, and body image features.
[0055] In a typical scenario, the user's multimodal features can be collected with the user's knowledge and permission. For example, by installing cameras in the lobby, self-service terminals, ATM machines and counter business equipment, the user's verbal and non-verbal behavior can be recorded, and the image records can be transmitted to the background processing system for real-time data processing.
[0056] In operation S230 , an emotion intensity value of the user is output based on the multimodal features.
[0057] Specifically, the user's emotion intensity value is output after the multimodal features are fused.
[0058] In a typical scenario, voice features, facial image features, and body image features are fused to output the emotion intensity value. It can be understood that the emotion intensity value output by fusing multiple different modal features can accurately and objectively reflect the user's emotions.
[0059] According to an embodiment of the present disclosure, wherein the multimodal features include N different modal features, N is a positive integer, and outputting the user's emotion intensity value based on the multimodal features includes: inputting the N different modal features into a preset machine learning model to obtain N different emotion intensity values; and performing weighting on the N different emotion intensity values to obtain the emotion intensity value.
[0060] Specifically, multimodal features are input into respective corresponding machine learning models to output corresponding emotion intensity values, and these different emotion intensities are weighted through pre-set weight distribution to obtain a comprehensive emotion intensity value.
[0061] In a typical scenario, voice features, facial image features, and body image features are collected and their corresponding emotion intensity values are calculated. The details are as follows:
[0062] 1. Perform speech emotion recognition on voice features. Specifically: First, collect customer voice information through the microphone embedded in the self-service terminal and counter business equipment; then, convert the user's voice into text and phonetic symbols through voice recognition, and extract words or sounds that express emotions; finally, classify the text and sounds into predefined emotion categories through machine learning, thereby identifying the user's current emotional state, and use the matching degree with the emotion category as the user's current emotional state value.
[0063] 2. Perform facial expression recognition on facial image features. Specifically: First, use cameras at various locations to record the network conditions, and transmit the detected image data containing human faces to the background analysis system; then, the background system preprocesses the image, including size processing, clarity processing, angle processing, etc., to unify the image for analysis and improve analysis accuracy; finally, encode and measure the facial features in the preprocessed image, such as action units, facial features, muscle movement and strength, etc., and learn and recognize based on machine vector (SVM) or convolutional neural network (CNN), classify the current expression pattern with predefined emotion categories, thereby identifying the user's current emotional state, and use the degree of match as the user's current emotional state value.
[0064] 3. Perform limb movement recognition on limb image features. Specifically: first, use cameras at various locations to record the situation at the network site, and transmit the detected image data containing human faces to the background analysis system; then, the background system pre-processes the image, including size processing, clarity processing, angle processing, etc., to unify the image for analysis and improve analysis accuracy; finally, locate key points in the body, including but not limited to the head, shoulders, elbows, wrists, hips, knees, ankles, etc., and extract posture features, classify posture patterns into predefined emotion categories to identify the user's current emotional state, and use the matching degree as the user's current emotional state value.
[0065] After calculating the emotional outlier value, the average of the user's speech, facial expression, and body movement emotional state values in each emotional category is used as the user's emotional value for that emotion. Comprehensive calculation of emotional values has been achieved.
[0066] According to an embodiment of the present disclosure, the preset machine learning model includes N preset machine learning sub-models, and the preset machine learning sub-models correspond one-to-one to any modality in the multimodal features. The training method of the preset machine learning sub-model includes: for any of the above, obtaining an initial machine sub-learning model and a corresponding training data set, wherein the training data set includes an emotion intensity value label and a corresponding modal feature; and obtaining the preset machine learning sub-model based on the initial machine sub-learning model and the training data set.
[0067] Specifically, the pre-set machine learning model includes N pre-set machine learning sub-models, each corresponding to a modality. The training method uses a hybrid-level fusion approach, combining the characteristics of feature-level and decision-level fusion. A separate prediction model is trained for each modality, and predictions are then obtained from each modality. Ultimately, when the model is used, the emotion intensity values output by each model are integrated.
[0068] In the embodiments of the present disclosure, a corresponding operation portal can be provided for the user to choose to agree or reject the automated decision result. That is, before the branch operation processing / decision is performed on the user information, an instruction to agree or reject the processing / decision can be obtained from the user through the corresponding operation portal. If the user agrees to the processing / decision, the branch operation processing / decision is performed on the user information, that is, step S230 is executed. If the user rejects the processing / decision, the expert decision process is entered.
[0069] In operation S240 , when the emotion intensity value of the user is within a preset abnormal emotion range, a corresponding preset emotion intervention plan is executed.
[0070] Specifically, the abnormal emotion interval includes multiple levels of classification, which are divided into levels according to the intensity of the abnormal emotion. Different levels of abnormal emotion intervals correspond to different preset emotion intervention plans, among which the preset emotion intervention plans include: key emotion monitoring, artificial intelligence soothing, and manual intervention soothing, as shown below:
[0071] According to an embodiment of the present disclosure, the preset abnormal emotion interval includes a first abnormal emotion interval, and when the user's emotion intensity value is in the preset abnormal emotion interval, the corresponding preset emotion intervention plan is executed, including: when the user's emotion intensity value is in the first abnormal emotion interval, based on the multimodal features, by matching the preset subject content, the cause of the abnormal emotion is derived.
[0072] Specifically, when the emotion intensity value is within the first abnormal emotion range, the user's multimodal features are continuously monitored and matched against a preset theme. This theme represents the causes of abnormal emotions that frequently occur during branch operations. These causes can then be recorded to iterate on branch operations. This matching operation can be achieved using pre-trained multimodal models, which will not be further detailed here.
[0073] In a typical scenario, when the intensity of a user's abnormal emotion deviates by 1 standard deviation from the mean, the cause of the abnormal emotion is determined. This is done by combining the specific actions taken by the user at the time of the abnormal emotion and the subject matter contained in their speech. This is then matched against pre-defined subject matter, such as "slow service," "unable to operate," or "poor service staff attitude," to pinpoint the cause of the problem.
[0074] According to an embodiment of the present disclosure, the preset abnormal emotion interval also includes a second abnormal emotion interval. When the user's emotion intensity value is in the second abnormal emotion interval, an artificial intelligence service is called; and soothing content is output based on the artificial intelligence service.
[0075] Specifically, when the emotion intensity value is in the second abnormal emotion interval, an external artificial intelligence service is called, wherein the artificial intelligence service is trained to soothe emotions through conversation, and the artificial intelligence service can be configured on a carrier that can interact with the user.
[0076] In a typical scenario, it is determined whether to prioritize emotional intervention provided by an artificial intelligence service in combination with the intensity of the user's abnormal emotion. When the intensity of the user's abnormal emotion deviates from the average value by 2 standard deviations, the artificial intelligence service prioritizes providing appropriate emotional intervention in combination with the user's emotional reasons and customer group classification, and conducts verbal comfort + emotion processing. For example, for young users whose negative emotion reason is "slow service", the artificial intelligence service will first comfort the user's emotion and express apology, and then use content such as jokes and interesting stories to help the user relieve the emotion; for elderly users whose negative emotion reason is "slow service", the artificial intelligence service will first comfort the user's emotion and express apology, and then guide the user to participate in marketing activities such as wool薅羊毛 (a Chinese term for taking advantage of promotional offers), fill the user's waiting time, and weaken the user's perception of the waiting time.
[0077] According to an embodiment of the present disclosure, wherein the preset abnormal emotion interval further includes a third abnormal emotion interval, and when the emotion intensity value of the user is within the third abnormal emotion interval, an artificial intervention signal is issued.
[0078] Specifically, when the emotion intensity value is within the third abnormal emotion interval, an artificial intervention signal is issued for early warning to inform the staff to participate in emotional intervention.
[0079] In a typical scenario, if the intensity of the user's abnormal emotion deviates from the average value by more than 3 standard deviations, or the current problem of the user is difficult to solve by the artificial intelligence service, a risk warning is given to the network staff, requiring them to arrive on the scene as soon as possible to handle it and prevent the user's emotion from deteriorating continuously.
[0080] According to an embodiment of the present disclosure, after executing the corresponding preset emotion intervention plan, the method further includes: statistically analyzing the change in the emotion intensity value, the amplitude of emotion fluctuation, and the abnormal emotion processing time; and outputting a service ability evaluation based on the change in the emotion intensity value, the amplitude of emotion fluctuation, and the abnormal emotion processing time.
[0081] Specifically, by obtaining the change in the emotion intensity value, the amplitude of emotion fluctuation, and the abnormal emotion processing time during the execution of the emotion intervention plan, and by comprehensively considering the change in the emotion intensity value, the amplitude of emotion fluctuation, and the abnormal emotion processing time, a service ability evaluation value is output.
[0082] In a typical scenario, the user's emotional state changes at the branch, the magnitude of emotional fluctuations, and the time it takes to process abnormal emotions are analyzed. The change in the user's emotional state at the branch is measured by the average difference between the emotional value of each emotion when the user leaves the branch and the emotional value of each emotion when the user enters the branch; the magnitude of emotional fluctuation is measured by the average difference between the highest and lowest values of each emotion; and the time it takes to process abnormal emotions is measured from the time the AI service identifies the user's abnormal emotion and issues a warning to the time the user's emotion returns to normal. Each level is assigned a score of 1 to 5. Finally, the three scores are converted to a 5-point scale, and the average of the three scores is used as the branch's service level score.
[0083] It is understandable that by monitoring user emotions to evaluate services, objectivity can be ensured while the service evaluation can be used as an evaluation basis for business iteration, or the service capability evaluation can be used as an evaluation basis for iterative artificial intelligence, which can better upgrade the level of branch operations.
[0084] In the embodiment of the present disclosure, in order to solve the technical problem of low operating efficiency and quality of the outlets, the embodiment of the present disclosure first collects the multimodal features of the user, then outputs an accurate emotion intensity value in combination with the multimodal features, and finally arranges a corresponding emotion intervention plan for the user according to the interval of the emotion intensity value, which can ensure the improvement of operational efficiency and quality. The beneficial effects are: 1. Outputting the emotion intensity value through multimodal features can provide additional information beyond a single modality, so as to more accurately judge the user's emotional state; 2. Real-time monitoring of the user's emotional changes and providing instant feedback based on the emotional state ensures the efficiency of the outlet processing, releases a large amount of human resources, helps the outlet staff to understand the user's emotional state more comprehensively, and focus on processing to avoid the fermentation of user emotions; 3. Through the emotional triggering objective allocation plan, various resources in the outlet system are allocated to where they are currently needed, avoiding waste of resources in the outlet operation system.
[0085] Based on the above-mentioned network operation method, the present disclosure also provides a network operation device. Figure 3 The device is described in detail.
[0086] Figure 3 The structural block diagram of the network operation device according to an embodiment of the present disclosure is schematically shown.
[0087] like Figure 3 As shown, the network operation device 300 of this embodiment includes an authorization module 310 , a collection module 320 , an emotion intensity output module 330 and an emotion intervention module 340 .
[0088] The authorization module 310 is used to obtain the user's authorization for multimodal feature collection. In one embodiment, the authorization module 310 can be used to perform the operation S210 described above, which will not be repeated here.
[0089] The collection module 320 is used to collect the user's multimodal features when the user authorizes the collection of multimodal features. In one embodiment, the collection module 320 can be used to perform the operation S220 described above, which will not be repeated here.
[0090] The emotion intensity output module 330 is used to output the user's emotion intensity value based on the multimodal features. In one embodiment, the emotion intensity output module 330 can be used to perform the operation S230 described above, which will not be repeated here.
[0091] The emotion intervention module 340 is used to execute a corresponding preset emotion intervention plan when the user's emotion intensity value is within a preset abnormal emotion range. In one embodiment, the emotion intervention module 340 can be used to execute the operation S240 described above, which will not be repeated here.
[0092] According to an embodiment of the present disclosure, the device also includes: a service evaluation module, which is used to count the changes in emotion intensity values, the amplitude of emotion fluctuations and the time for processing abnormal emotions; and output a service capability evaluation based on the changes in emotion intensity values, the amplitude of emotion fluctuations and the time for processing abnormal emotions.
[0093] According to an embodiment of the present disclosure, the preset abnormal emotion interval includes a first abnormal emotion interval, and the emotion intervention module is specifically used to derive the cause of the abnormal emotion based on the multimodal features by matching the preset subject content when the user's emotion intensity value is in the first abnormal emotion interval.
[0094] According to an embodiment of the present disclosure, the preset abnormal emotion interval also includes a second abnormal emotion interval, and the emotion intervention module is further specifically used to call an artificial intelligence service when the user's emotion intensity value is in the second abnormal emotion interval; and output soothing content based on the artificial intelligence service.
[0095] According to an embodiment of the present disclosure, the preset abnormal emotion interval further includes a third abnormal emotion interval, and the emotion intervention module is further specifically configured to send a manual intervention signal when the user's emotion intensity value is in the third abnormal emotion interval.
[0096] According to an embodiment of the present disclosure, the multimodal features include N different modal features, where N is a positive integer, and the emotion intensity output module is specifically used to input the N different modal features into a preset machine learning model to obtain N different emotion intensity values; and perform weighting on the N different emotion intensity values to obtain the emotion intensity value.
[0097] According to an embodiment of the present disclosure, the preset machine learning model includes N preset machine learning sub-models, and the preset machine learning sub-models correspond one-to-one to any modality in the multimodal features. The device also includes a training module for obtaining an initial machine sub-learning model and a corresponding training data set for any of the above, wherein the training data set includes an emotion intensity value label and a corresponding modal feature; and obtaining the preset machine learning sub-model based on the initial machine sub-learning model and the training data set.
[0098] In the embodiment of the present disclosure, in order to solve the technical problem of low operating efficiency and quality of the outlets, the embodiment of the present disclosure first collects the multimodal features of the user, then outputs an accurate emotion intensity value in combination with the multimodal features, and finally arranges a corresponding emotion intervention plan for the user according to the interval of the emotion intensity value, which can ensure the improvement of operational efficiency and quality. The beneficial effects are: 1. Outputting the emotion intensity value through multimodal features can provide additional information beyond a single modality, so as to more accurately judge the user's emotional state; 2. Real-time monitoring of the user's emotional changes and providing instant feedback based on the emotional state ensures the efficiency of the outlet processing, releases a large amount of human resources, helps the outlet staff to understand the user's emotional state more comprehensively, and focus on processing to avoid the fermentation of user emotions; 3. Through the emotional triggering objective allocation plan, various resources in the outlet system are allocated to where they are currently needed, avoiding waste of resources in the outlet operation system.
[0099] According to an embodiment of the present disclosure, any multiple modules among the authorization module 310, the acquisition module 320, the emotion intensity output module 330, and the emotion intervention module 340 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to an embodiment of the present disclosure, at least one of the authorization module 310, the acquisition module 320, the emotion intensity output module 330, and the emotion intervention module 340 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware by any other reasonable means of integrating or packaging circuits, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, at least one of the authorization module 310 , the collection module 320 , the emotion intensity output module 330 , and the emotion intervention module 340 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0100] Figure 4 A block diagram of an electronic device suitable for implementing a network operation method according to an embodiment of the present disclosure is schematically shown.
[0101] like Figure 4 As shown, the electronic device 400 according to an embodiment of the present disclosure includes a processor 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage portion 408 into a random access memory (RAM) 403. The processor 401 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 401 may also include onboard memory for caching purposes. The processor 401 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.
[0102] Various programs and data required for the operation of the electronic device 400 are stored in the RAM 403. The processor 401, ROM 402, and RAM 403 are connected to each other via a bus 404. The processor 401 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 402 and / or RAM 403. It should be noted that the programs may also be stored in one or more memories other than the ROM 402 and RAM 403. The processor 401 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0103] According to an embodiment of the present disclosure, electronic device 400 may further include an input / output (I / O) interface 405, which is also connected to bus 404. Electronic device 400 may also include one or more of the following components connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 408 including a hard disk; and a communication section 409 including a network interface card such as a LAN card or modem. Communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. Removable media 411, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 410 as needed, so that computer programs read from the removable media can be installed into storage section 408 as needed.
[0104] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0105] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 402 and / or RAM 403 described above, and / or one or more memories other than ROM 402 and RAM 403.
[0106] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiments of the present disclosure.
[0107] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the processor 401 executes the computer program. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0108] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 409, and / or installed from a removable medium 411. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0109] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by the processor 401, the above-mentioned functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0110] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0112] Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of the present disclosure. All such combinations and / or couplings fall within the scope of the present disclosure.
[0113] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A network operation method, characterized in that: The method comprises: Obtain user authorization for multimodal feature collection; Collecting the user's multimodal features after obtaining the user's authorization for the collection of multimodal features; Based on the multimodal features, outputting the user's emotion intensity value; and When the user's emotion intensity value is within a preset abnormal emotion range, a corresponding preset emotion intervention plan is executed.
2. The method according to claim 1, characterized in that After executing the corresponding preset emotion intervention plan, the method further includes: Count changes in emotion intensity, emotion fluctuations, and the time it takes to process abnormal emotions; and Based on the change in the emotion intensity value, the emotion fluctuation amplitude, and the abnormal emotion processing time, a service capability evaluation is output.
3. The method according to claim 1 or 2, characterized in that in, The preset abnormal emotion interval includes a first abnormal emotion interval, When the user's emotion intensity value is within a preset abnormal emotion range, executing a corresponding preset emotion intervention plan includes: When the emotion intensity value of the user is in the first abnormal emotion interval, the cause of the abnormal emotion is obtained by matching preset subject content based on the multimodal features.
4. The method according to claim 3, characterized in that in, The preset abnormal emotion interval also includes a second abnormal emotion interval. When the user's emotion intensity value is within a preset abnormal emotion range, executing a corresponding preset emotion intervention plan further includes: In a case where the user's emotion intensity value is in the second abnormal emotion interval, calling an artificial intelligence service; and Output soothing content based on the artificial intelligence service.
5. The method according to claim 3, characterized in that in, The preset abnormal emotion interval also includes a third abnormal emotion interval, When the user's emotion intensity value is within a preset abnormal emotion range, executing a corresponding preset emotion intervention plan further includes: When the emotion intensity value of the user is in the third abnormal emotion interval, a manual intervention signal is issued.
6. The method according to any one of claims 1, 2, 4 and 5, characterized in that in, The multimodal features include N different modal features, where N is a positive integer. Outputting the user's emotion intensity value based on the multimodal features includes: Inputting the N different modal features into a preset machine learning model to obtain N different emotion intensity values; and Weighting is performed on the N different emotion intensity values to obtain the emotion intensity value.
7. The method according to claim 6, characterized in that The preset machine learning model includes N preset machine learning sub-models, each of which corresponds one-to-one to any modality in the multimodal features. The training method of the preset machine learning sub-model includes: For any of the above, obtaining an initial machine learning sub-model and a corresponding training dataset, wherein the training dataset includes emotion intensity value labels and corresponding modality features; and Based on the initial machine sub-learning model and the training data set, the preset machine learning sub-model is obtained.
8. A network operation device, characterized in that: The device comprises: Authorization module, used to obtain user authorization for multimodal feature collection; A collection module, configured to collect the user's multimodal features upon obtaining the user's authorization for the collection of multimodal features; an emotion intensity output module, configured to output the user's emotion intensity value based on the multimodal features; and The emotion intervention module is used to execute a corresponding preset emotion intervention plan when the emotion intensity value of the user is within a preset abnormal emotion range.
9. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.