Audio and video data processing method and device, program product and electronic device

By analyzing audio and video data of customers and tellers at financial institution branches in real time, constructing status change curves and triggering prompts, the subjective and real-time issues of customer satisfaction monitoring are resolved, thereby improving customer satisfaction and service quality.

CN119835458BActive Publication Date: 2025-12-26INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411873632.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-12-26
Estimated Expiration
2044-12-18

AI Technical Summary

Technical Problem

Existing methods for monitoring customer satisfaction at financial institution branches suffer from issues such as high subjectivity, low data collection efficiency, inability to monitor in real time, and limited information dimensions, leading to delays in service quality improvement.

Method used

By collecting audio and video data from customers and tellers, using timestamps for real-time analysis, constructing state change curves, and triggering prompts before the emotional state reaches a threshold, the system combines a multimodal emotion recognition model to monitor the emotional state of customers and tellers in real time.

Benefits of technology

It enables real-time and accurate monitoring of customers' emotional state, timely intervention to prevent emotional deterioration, and improved customer satisfaction and service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119835458B_ABST
    Figure CN119835458B_ABST
Patent Text Reader

Abstract

The application discloses a kind of audio and video data processing method and its device, program product, electronic equipment, it is related to artificial intelligence field or other related fields, wherein, the processing method includes: in the case where it is detected that target object handles financial business, and obtains target object authorization, first audio and video data of target object is collected, wherein, first audio and video data carries time stamp;Determine the target state of target object on each time point and the target state probability value corresponding to target state based on first audio and video data;Based on the target state on each time point and the target state probability value corresponding to target state, the state change curve of target object is constructed;Before monitoring that state change curve reaches first probability threshold line or second probability threshold line, state prompt information is triggered.The present application solves the technical problem that the accuracy of determining customer emotional state in related art is low.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular to a method and device for processing audio and video data, a program product and an electronic device. BACKGROUND

[0002] With the rapid development of financial technology, financial institutions are facing the problem of improving customer satisfaction as key providers of financial services. With the continuous improvement of customer expectations, the current satisfaction monitoring method gradually reveals its shortcomings. Currently, most of the counter of financial institutions relies on manual point selection to collect customer satisfaction feedback, usually through paper questionnaires or electronic devices set at the counter, allowing customers to make choices after completing business. Although this method is simple, it has the following problems in the implementation process:

[0003] (1) Strong subjectivity, lack of objectivity: The manual point selection satisfaction method mainly depends on the subjective feelings of customers, and customers may be influenced by various factors when selecting, such as personal emotions, brand loyalty to financial institutions, etc., resulting in a certain subjectivity and uncertainty of the satisfaction result.

[0004] (2) Low data collection efficiency, long feedback cycle: The manual point selection satisfaction method usually requires customers to provide feedback through paper questionnaires or electronic questionnaires after completing business, which not only has low data collection efficiency, but also has a long feedback cycle, making it difficult to reflect the real situation of the counter service of financial institutions in a timely manner.

[0005] (3) Unable to monitor in real time, lack of timeliness: The manual point selection satisfaction method cannot monitor the changes in customer satisfaction during the business process in real time, and can only provide feedback after the business is completed, which makes it difficult for financial institutions to discover and solve problems in the service process in a timely manner, affecting the continuous improvement of service quality.

[0006] (4) Single information dimension, difficult to evaluate comprehensively: The manual point selection satisfaction method usually only focuses on the overall service satisfaction evaluation of customers, lacks in-depth analysis of multiple dimensions such as customer behavior characteristics, emotional changes, and transaction results, and is difficult to comprehensively and accurately evaluate the quality of the counter service of financial institutions.

[0007] Currently, there is no effective solution to the above problems. SUMMARY

[0008] The embodiments of the present application provide a method and device for processing audio and video data, a program product and an electronic device to at least solve the technical problem of low accuracy in determining the emotional state of customers in related technologies.

[0009] According to an aspect of some embodiments of the present application, a method for processing audio and video data is provided. The method comprises: collecting first audio and video data of a target object when it is detected that the target object is conducting a financial transaction and the target object has authorized the collection, wherein the first audio and video data carries a timestamp; determining a target state of the target object at each time point and a target state probability value corresponding to the target state based on the first audio and video data; constructing a state change curve of the target object based on the target state at each time point and the target state probability value corresponding to the target state; and triggering a state prompt message before the state change curve reaches a first probability threshold line or a second probability threshold line.

[0010] Further, the step of determining the target state of the target object at each time point and the target state probability value corresponding to the target state based on the first audio and video data comprises: separating the first audio and video data to obtain video data and audio data; processing the video data to obtain target video data; extracting facial feature data and action feature data of the target object based on the target video data; processing the facial feature data, the action feature data, and the audio data respectively to obtain a first state result, a second state result, and a third state result; and determining the target state of the target object at each time point and the target state probability value corresponding to the target state based on the first state result, the second state result, and the third state result.

[0011] Further, the step of processing the video data to obtain target video data comprises: labeling each frame of image in the video data to obtain a rectangular frame on each frame of image; and cropping each frame of image based on the rectangular frame on each frame of image to obtain the target video data.

[0012] Further, the step of processing the facial feature data, the action feature data and the audio data respectively to obtain a first state result, a second state result and a third state result comprises: processing the facial feature data by using a first model to obtain the first state result, wherein the first model is a model trained by using first historical data, and the first historical data comprises a set of historical facial feature data and a state label result of labeling each historical facial feature data in the set of historical facial feature data; processing the action feature data by using a second model to obtain the second state result, wherein the second model is a model trained by using second historical data, and the second historical data comprises a set of historical action feature data and a state label result of labeling each historical action feature data in the set of historical action feature data; processing the audio data by using a third model to obtain the third state result, wherein the third model is a model trained by using third historical data, and the second historical data comprises a set of historical audio data and a state label result of labeling each historical audio data in the set of historical audio data.

[0013] Further, before processing the audio data by using the third model to obtain the third state result, the method further comprises: filtering the audio data to obtain target audio data; and extracting features from the target audio data to obtain spectral features, pitch features and volume features.

[0014] Further, the first state result, the second state result and the third state result each at least comprise a positive state and a positive probability value corresponding to the positive state, a negative state and a negative probability value corresponding to the negative state. Based on the first state result, the second state result and the third state result, the step of determining the target state of the target object at each time point and the target state probability value corresponding to the target state comprises: comparing the positive probability value with the negative probability value in a state result, and determining the state corresponding to the state result as the positive state in a case where the positive probability value is greater than or equal to the negative probability value, or determining the state corresponding to the state result as the negative state in a case where the positive probability value is less than the negative probability value, wherein the state result is the first state result, the second state result or the third state result; determining the target state of the target object at a current time point based on the states respectively corresponding to all the state results; and determining the target state probability value corresponding to the target state based on the probability values corresponding to the states in all the state results consistent with the target state.

[0015] Further, the processing method further comprises: collecting second audio-video data of a target teller, wherein the target teller is a teller who handles the financial business for the target object, and the second audio-video data carries a timestamp; processing the second audio-video data to obtain a teller state of the target teller at each time point and a teller state probability value corresponding to the teller state; constructing a teller state change curve of the target teller based on the teller state at each time point and the teller state probability value corresponding to the teller state; and displaying time points at which the teller state is a negative state on the teller state change curve in a preset form.

[0016] Further, after collecting the second audio-video data of the target teller, the method further comprises: integrating audio data in the first audio-video data and audio data of the second audio-video data according to the timestamps to obtain a dialogue text; determining dialogue content corresponding to all time points at which the target teller has the negative state in the dialogue text based on the teller state change curve; and inputting the dialogue text into a preset language model to output summary content, wherein the preset language model is a model trained using a historical dialogue text set and annotated summary content annotated for each historical dialogue text in the historical dialogue text set.

[0017] Further, before the state prompt information reaches the first probability threshold line or the second probability threshold line, the step of triggering the state prompt information comprises: before the state change curve reaches the first probability threshold line or the second probability threshold line, triggering the state prompt information based on the dialogue content before the current time point and the summary content.

[0018] According to another aspect of the embodiments of the present application, an audio-video data processing apparatus is also provided, comprising: a collection unit configured to collect first audio-video data of a target object when it is detected that the target object handles a financial business and it is obtained that the target object is authorized, wherein the first audio-video data carries a timestamp; a determination unit configured to determine a target state of the target object at each time point and a target state probability value corresponding to the target state based on the first audio-video data; a construction unit configured to construct a state change curve of the target object based on the target state at each time point and the target state probability value corresponding to the target state; and a triggering unit configured to trigger a state prompt information before the state change curve reaches a first probability threshold line or a second probability threshold line.

[0019] Further, the determining unit comprises: a first separation module, configured to separate the first audio-video data to obtain video data and audio data; a first processing module, configured to process the video data to obtain target video data; a first extraction module, configured to extract facial feature data and action feature data of the target object based on the target video data; a second processing module, configured to process the facial feature data, the action feature data and the audio data respectively to obtain a first state result, a second state result and a third state result; and a first determining module, configured to determine the target state of the target object at each time point and the target state probability value corresponding to the target state based on the first state result, the second state result and the third state result.

[0020] Further, the first processing module comprises: a first labeling submodule, configured to label each frame of image in the video data to obtain a rectangular frame on each frame of image; and a first cropping submodule, configured to crop each frame of image based on the rectangular frame on each frame of image to obtain the target video data.

[0021] Further, the second processing module comprises: a first processing submodule, configured to process the facial feature data by using a first model to obtain the first state result, wherein the first model is a model trained by using first historical data, and the first historical data comprises a set of historical facial feature data and a state labeling result of labeling each historical facial feature data in the set of historical facial feature data; a second processing submodule, configured to process the action feature data by using a second model to obtain the second state result, wherein the second model is a model trained by using second historical data, and the second historical data comprises a set of historical action feature data and a state labeling result of labeling each historical action feature data in the set of historical action feature data; and a third processing submodule, configured to process the audio data by using a third model to obtain the third state result, wherein the third model is a model trained by using third historical data, and the second historical data comprises a set of historical audio data and a state labeling result of labeling each historical audio data in the set of historical audio data.

[0022] Further, the processing device further comprises: a first filtering module, configured to filter the audio data to obtain target audio data before processing the audio data by using the third model to obtain the third state result; and a second extraction module, configured to extract features of the target audio data to obtain spectral features, pitch features and volume features.

[0023] Further, the first state result, the second state result and the third state result each at least include: a positive state and a positive probability value corresponding to the positive state, a negative state and a negative probability value corresponding to the negative state, the first determining module comprises: a first comparison submodule, configured to compare the positive probability value with the negative probability value in a state result, and determine a state corresponding to the state result as the positive state in a case that the positive probability value is greater than or equal to the negative probability value, or determine the state corresponding to the state result as the negative state in a case that the positive probability value is less than the negative probability value, wherein the state result is the first state result, the second state result or the third state result; a first determining submodule, configured to determine the target state of the target object at a current time point based on the states corresponding to all the state results; and a second determining submodule, configured to determine a target state probability value corresponding to the target state based on probability values corresponding to the states in all the state results consistent with the target state.

[0024] Further, the processing apparatus further comprises: a first collecting module, configured to collect second audio-video data of a target clerk, wherein the target clerk is a clerk who handles the financial business for the target object, and the second audio-video data carries a time stamp; a third processing module, configured to process the second audio-video data to obtain a clerk state of the target clerk at each time point and a clerk state probability value corresponding to the clerk state; a first constructing module, configured to construct a clerk state change curve of the target clerk based on the clerk state at each time point and the clerk state probability value corresponding to the clerk state; and a first displaying module, configured to display a time point at which the clerk state is a negative state on the clerk state change curve in a preset form.

[0025] Further, the processing apparatus further comprises: a first integrating module, configured to integrate audio data in the first audio-video data and audio data of the second audio-video data to obtain a dialogue text according to the time stamp after collecting the second audio-video data of the target clerk; a second determining module, configured to determine dialogue content corresponding to all time points at which the target clerk has the negative state in the dialogue text based on the clerk state change curve; and a first inputting module, configured to input the dialogue text into a preset language model to output summary content, wherein the preset language model is a model trained by using a historical dialogue text set and annotated summary content annotated for each historical dialogue text in the historical dialogue text set.

[0026] Further, the triggering unit comprises a first triggering module, configured to trigger the state prompt information based on the dialogue content before the current time point and the summary content before the state change curve reaches the first probability threshold line or the second probability threshold line.

[0027] According to another aspect of the embodiments of the present application, a computer program product is also provided, comprising a non-volatile computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement any of the above-mentioned audio and video data processing methods.

[0028] According to another aspect of the embodiments of the present application, an electronic device is also provided, comprising one or more processors and a memory, the memory is configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-mentioned audio and video data processing methods.

[0029] In the present application, when the target object handles the financial business and the target object's authorization is obtained, the first audio and video data of the target object is collected, wherein the first audio and video data carries a time stamp; based on the first audio and video data, the target state of the target object at each time point and the target state probability value corresponding to the target state are determined; based on the target state at each time point and the target state probability value corresponding to the target state, a state change curve of the target object is constructed; before the state change curve reaches the first probability threshold line or the second probability threshold line is monitored, the state prompt information is triggered, thereby solving the technical problem that the accuracy of determining the emotional state of the customer in the related art is low.

[0030] In the present application, by analyzing the collected audio and video data of the target object in real time, the target state of the target object at each time point and the target state probability value corresponding to the target state are obtained, and then a state change curve changing with time is constructed, so that the state prompt information is triggered in time before the state change curve reaches the first probability threshold line or the second probability threshold line, thus not only the emotional state of the customer can be determined in real time and accurately, but also the intervention can be made in time before the emotional state of the customer deteriorates further, which is beneficial to improve the customer satisfaction, and the technical effects of monitoring the emotional state of the customer in real time and improving the customer satisfaction are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0031] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:

[0032] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the processing method of audio and video data is shown;

[0033] Figure 2 is a flow chart of the processing method of audio and video data according to Embodiment 1 of the present application;

[0034] Figure 3 is a coordinate system diagram of a state change curve according to Embodiment 1 of the present application;

[0035] Figure 4 is a schematic diagram of the audio and video processing flow of a user according to Embodiment 1 of the present application;

[0036] Figure 5 is a schematic diagram of an optional processing device of audio and video data according to Embodiment 1 of the present application;

[0037] Figure 6 is a structure block diagram of an electronic device according to Embodiment 1 of the present application. DETAILED DESCRIPTION

[0038] In order to make the personnel in the art better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0040] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) collected and related to the present application are all information and data authorized by the user or authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data comply with relevant laws, regulations and standards in relevant regions, necessary security measures are taken, do not violate public order and good customs, and provide corresponding operation portal for user to choose authorization or refusal. For example, the system and related users or institutions are provided with an interface, and before obtaining the relevant information, the interface needs to send a request to the aforementioned user or institution, and after receiving the consent information feedback from the aforementioned user or institution, the relevant information is obtained.

[0041] In the present application, a financial institution network point counter user emotional state early warning method is proposed, which overcomes the problem that the current user's satisfaction is only manually selected through the interactive device after the business is completed, and cannot monitor the user's state change in real time during the business process. By monitoring the emotional state change of the user in the counter business process in real time and giving timely warning prompts, the user's emotional state can be intervened in time before it further deteriorates, the user's satisfaction is improved, and the emotional state change of the user is analyzed by monitoring the state change of the related counter staff and the dialogue content between them, and the subsequent service quality is improved in a targeted manner, and the user's business experience is improved.

[0042] The present application will be described in detail below in conjunction with various embodiments.

[0043] Embodiment 1

[0044] According to the embodiments of the present application, an embodiment of a method for processing audio and video data is also provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0045] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for processing audio and video data is shown. As shown in Figure 1 The computer terminal 10 (or mobile device) can include one or more (CPU) central processing units 1001, memories 1002, storage devices 1003, communication interfaces 1004, input devices 1005, output devices 1006 and the like. Figure 1The processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions may also be included. In addition, it may include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera, wherein the network interface can be connected to wired and / or wireless networks. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0046] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0047] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the audio and video data processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned audio and video data processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0048] The transmission device 106 is configured to receive or send data via a network. The network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module that is configured to communicate with the Internet wirelessly.

[0049] The display can be a liquid crystal display (LCD) that is touch screen type, for example, which can enable a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0050] In the above operating environment, the present application provides a method for processing audio and video data as shown in Figure 2 Figure 2 is a flowchart of the method for processing audio and video data according to an embodiment of the present application, which includes the following steps: Figure 2

[0051] In step S201, first audio and video data of a target object is collected when the target object is detected to handle a financial service and the target object is authorized, wherein the first audio and video data carries a time stamp.

[0052] Optionally, the method for processing audio and video data can be performed by an emotion warning system.

[0053] In the embodiment of the present application, the emotion warning system includes a device outside a counter, a device inside the counter, and a processing device.

[0054] The device outside the counter is arranged outside a counter surface and can collect audio and video of a user currently handling a service in real time. In order to protect the privacy of the user, an authorization button can be arranged on the device outside the counter to collect audio and video of the user when the user is authorized.

[0055] The device inside the counter is arranged inside the counter surface and is bound to the corresponding device outside the counter one by one. The device inside the counter can collect audio and video of a teller currently handling a service in real time.

[0056] The processing device is connected to the device outside the counter and the device inside the counter. When the device outside the counter collects audio and video of a user in front of the corresponding counter, the processing device starts to analyze the audio and video of the user and the teller to monitor the emotional state of the user in real time.

[0057] ​​In the embodiment of the present application, after the emotion warning system detects that the target object (i.e., the user who needs to handle financial business at the counter) handles financial business and obtains authorization of the target object, the first audio and video data of the target object can be collected by the out-of-counter collection device, and the collected first audio and video data carries a timestamp, which means that the emotion warning system can accurately record the time of the user's emotional change and the timestamp of each word exchanged with the teller. The existence of the timestamp ensures the real-time and traceability of the data, and provides accurate timing information for subsequent emotion analysis and dialogue text integration. This helps to identify the moment of emotional fluctuation in real time, so as to respond and intervene more quickly.

[0058] In step S202, the target state of the target object at each time point and the target state probability value corresponding to the target state are determined based on the first audio and video data.

[0059] In the embodiment of the present application, the processing device in the emotion warning system can be used to process the first audio and video data to determine the target state (positive emotion or negative emotion) of the target object at each time point and the target probability value corresponding to the target emotional state.

[0060] In the embodiment of the present application, positive emotion and negative emotion are two basic classifications used to describe the emotional state of individuals, and are usually used to reflect the inner feelings and reactions of a person in a specific situation. In the emotion warning system, these two concepts are used to evaluate the emotional state of the user in real time, so that the financial institution network point can respond quickly and optimize the user experience.

[0061] Among them, positive emotion refers to a positive, pleasant or satisfactory emotional experience. In the financial institution network service scene, positive emotion may include user satisfaction with service, gratitude for the friendly attitude of the staff, and a relaxed and pleasant feeling of the business handling process, etc. When the system identifies that the user expresses positive facial expressions, body language or voice features such as smiling, nodding, and relaxed tone, it will be classified as positive emotion.

[0062] Negative emotion is the opposite of positive emotion, which refers to a negative, unpleasant or unsatisfactory emotional experience. In the financial institution network service scene, negative emotion may be caused by user dissatisfaction with service, frustration with long waiting time, anger at the attitude of the staff, or disappointment with the result of some transactions, etc. The system identifies negative non-verbal and verbal signals such as frowning, shaking head, speeding up speech, raising volume or using negative words in speech to determine whether the user is in a negative emotional state.

[0063] In this embodiment, by monitoring the changes of positive emotions and negative emotions in real time, instant feedback on the counter service can be obtained, service strategies can be adjusted in time, customer experience can be optimized, customer dissatisfaction can be prevented and reduced, and customer trust and satisfaction with the financial institution outlets can be further enhanced.

[0064] In step S203, a state change curve of the target object is constructed based on the target state at each time point and the target state probability value corresponding to the target state.

[0065] In the embodiment of the present application, a memory is correspondingly configured in the processing device. After the processing device extracts the emotional state (i.e., the target state, including positive emotions or negative emotions) of the user at different time points and the corresponding state probability value from the collected first audio and video data, these data can be recorded and saved in the memory in time sequence to provide a basis for constructing the state change curve.

[0066] In the embodiment of the present application, the state change curve (i.e., the emotion change curve) of the target object can be generated with time as the horizontal coordinate and the emotional state and the corresponding probability value as the vertical coordinate, and the probability value reflects the confidence of the user's emotional state. By connecting these data points to form a curve, the fluctuation of the user's emotions over time can be intuitively seen.

[0067] In step S204, a state prompt information is triggered before the state change curve reaches the first probability threshold line or the second probability threshold line.

[0068] In the embodiment of the present application, two probability thresholds, i.e., the first probability threshold and the second probability threshold, can be set in advance. The first probability threshold is a threshold for judging whether the transition from positive emotions to negative emotions is about to occur, and the second probability threshold is a threshold for monitoring whether the negative emotions have reached a severity that requires immediate intervention.

[0069] In the embodiment of the present application, the first probability threshold line (i.e., the value of the horizontal coordinate changes with time, and the value of the vertical coordinate is the fixed first probability threshold) can be drawn on the coordinate axis of the emotion change curve according to the first probability threshold, and the second probability threshold line (i.e., the value of the horizontal coordinate changes with time, and the value of the vertical coordinate is the fixed second probability threshold) can be drawn on the coordinate axis of the emotion change curve according to the second probability threshold.

[0070] In the embodiment of the present application, the mood change curve can be continuously monitored, and when it is monitored that the curve is about to reach the first probability threshold line (i.e., the probability of positive mood decreases to a certain level, indicating that the mood may change from positive to negative) or the second probability threshold line (i.e., the probability of negative mood rises to a certain level, indicating that the mood has reached a level that requires urgent intervention), the mood warning system triggers the state prompt information in advance. In this way, by triggering the warning before the mood change curve reaches the threshold line, enough time window can be provided for the financial institution outlet to take measures to avoid further deterioration of the user's mood problem, which enables the financial institution outlet to more actively manage the customer's emotional experience rather than passively responding.

[0071] Figure 3 is a coordinate system diagram of the state change curve according to Embodiment 1 of the present application, as Figure 3 shown, the coordinate system of the mood change curve is constructed with time as the horizontal coordinate and the probability value of the mood as the vertical coordinate, for example, the probability value of positive mood can be represented above the time axis, and the probability value increases from bottom to top; the probability value of negative mood can be represented below the time axis, and the probability value decreases from top to bottom. And draw a dashed line representing the first probability threshold on the upper side of the time axis, and draw a dashed line representing the second probability threshold on the lower side of the time axis.

[0072] In this embodiment, by triggering the state prompt information in advance, the financial institution outlet can take action before the user's mood problem occurs or worsens, not only improving customer satisfaction, but also promoting the continuous optimization of service processes.

[0073] In summary, by analyzing the collected audio and video data of the target object in real time, the target state of the target object at each time point and the target state probability value corresponding to the target state can be obtained, and then a state change curve changing over time is constructed to trigger the state prompt information in time before the state change curve reaches the first probability threshold line or the second probability threshold line. In this way, not only can the emotional state of the customer be determined in real time and accurately, but also intervention can be made in time before the customer's emotional state further deteriorates, which is beneficial to improve customer satisfaction, achieves the technical effects of real-time monitoring of the emotional state of the customer and improving customer satisfaction, and further solves the technical problem of low accuracy in determining the emotional state of the customer in the related art.

[0074] In order to improve the accuracy of determining the target state of the target object at each time point and the target state probability value corresponding to the target state, in the method for processing audio and video data provided in Embodiment 1 of the present application, the first audio and video data is separated to obtain video data and audio data; the video data is processed to obtain target video data; the facial feature data and the action feature data of the target object are extracted based on the target video data; the facial feature data, the action feature data and the audio data are processed respectively to obtain a first state result, a second state result and a third state result; and the target state of the target object at each time point and the target state probability value corresponding to the target state are determined based on the first state result, the second state result and the third state result.

[0075] In the embodiments of the present application, the first audio and video data collected originally can be separated into two independent data streams of video and audio. This step is based on the need for multi-modal emotion recognition, because the video data can provide information of facial expressions and body language, and the audio data contains voice features. Then the separated video data is processed to obtain target video data. After that, the facial feature data is extracted from the target video data, including but not limited to the morphological changes of key parts such as eyes, mouth and eyebrows, which can reflect the emotional state of the user. At the same time, the action feature data is extracted from the target video data, including the feature data of the user's body movements and postures, to further supplement the emotional information, because some emotions may be conveyed through body language rather than facial expressions.

[0076] In the embodiments of the present application, the facial feature data, the action feature data and the audio data are processed respectively by using corresponding models to obtain the first state result, the second state result and the third state result based on different information sources, and the three results correspond to the emotional analysis of the user in the video and audio modalities respectively.

[0077] In the embodiments of the present application, considering that the user may have emotional suppression in public, such as controlling facial expressions, or the emotions expressed by the face and the action are inconsistent, or the emotional expressions of different users are inconsistent, therefore, the first emotion, the second emotion and the third emotion can be used to comprehensively consider the current emotion of the user to determine the target emotional state (positive or negative) of the target object at each time point and the confidence (i.e. the target state probability value) corresponding to the state.

[0078] In order to improve the accuracy of determining the target video data, in the method for processing audio and video data provided in Embodiment 1 of the present application, each frame of image in the video data is labeled to obtain a rectangular frame on each frame of image; and each frame of image is cropped based on the rectangular frame on each frame of image to obtain the target video data.

[0079] In the embodiments of the present application, each frame of image in the video stream can be processed, and the upper body of the target object is located and labeled using computer vision technology such as a target detection algorithm. For example, a pre-trained model such as YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), etc. can be used, which can identify and frame the human body part in the image, especially the upper body, because facial expressions and body language are important basis for emotion recognition. In this way, when extracting features (i.e. facial features and action features) later, the user's upper body area can be focused on, reducing the influence of background interference and other irrelevant information, thereby improving the accuracy of subsequent facial feature and action feature extraction.

[0080] And if the rectangular frame on each frame of image is determined, each frame of image can be cropped based on the frame to remove irrelevant areas and only keep image information of the user's upper body. The image data processed in this way is named "target video data".

[0081] In the embodiments, the cropped target video data not only reduces the data volume and speeds up the subsequent processing, but also ensures that the focus of the emotion recognition algorithm is on the key emotional expression area of the user, i.e. the face and the action of the upper body. This step improves the efficiency and accuracy of emotion recognition, because the algorithm does not have to find the key features in a large amount of background information, but can directly process image data directly related to emotional expression.

[0082] In order to improve the accuracy of determining multiple state results, in the method for processing audio and video data provided in Embodiment 1 of the present application, a first model is used to process the facial feature data to obtain a first state result, wherein the first model is a model trained using first historical data, and the first historical data includes a set of historical facial feature data and a state label result for labeling each historical facial feature data in the set of historical facial feature data; a second model is used to process the action feature data to obtain a second state result, wherein the second model is a model trained using second historical data, and the second historical data includes a set of historical action feature data and a state label result for labeling each historical action feature data in the set of historical action feature data; a third model is used to process the audio data to obtain a third state result, wherein the third model is a model trained using third historical data, and the second historical data includes a set of historical audio data and a state label result for labeling each historical audio data in the set of historical audio data.

[0083] In the embodiments of the present application, a first model can be constructed for processing facial feature data to identify emotional states. The model is trained through a large amount of first historical data, which includes a set of historical facial feature data and corresponding emotional annotation results. The emotional annotation results are based on expert or supervised learning algorithm marking the emotional states (such as happy, sad, angry, etc.) of the historical facial feature data, which provides a reference standard for the model to learn. Through deep learning techniques such as convolutional neural networks, the model can learn the association between facial features and emotional states, so that in actual application, it can output emotional recognition results based on newly collected facial feature data, i.e. first state results.

[0084] And a second model can be constructed for processing action feature data. The second model is trained based on a large amount of second historical data, which contains historical action feature data and its corresponding emotional annotation results. Through recurrent neural networks or long short-term memory networks, the second model can capture the dynamic changes of body movements over time and the relationship between these changes and emotional states, thereby outputting second state results.

[0085] A third model can also be constructed for processing audio data. The third model identifies emotions by analyzing features such as pitch, volume, speech rate, etc. The training of the third model relies on third historical data, which contains historical audio data and its emotional annotation results. The model can use deep learning techniques such as recurrent neural networks or gated recurrent units to learn the complex relationship between sound features and emotional states, thereby outputting third state results.

[0086] To facilitate subsequent processing of audio data, in the method for processing audio and video data provided in Embodiment 1 of the present application, before the third model is used to process the audio data to obtain the third state results, the audio data is filtered to obtain target audio data; the target audio data is feature extracted to obtain spectral features, pitch features and volume features.

[0087] In the embodiments of the present application, the user's audio data is extracted, and the audio of others is filtered out, leaving only the user's audio, thereby obtaining the target audio data.

[0088] In this embodiment, the filtering of audio data aims to remove irrelevant or interfering background sounds, ensuring that the third model can focus on the user's sound signal, thereby improving the accuracy and efficiency of subsequent emotional recognition. Filtering usually uses noise suppression algorithms or speech activity detection techniques, which can identify and separate the user's voice while reducing or completely removing interfering signals such as other customer conversations, environmental noise at the counter, etc. For example, detecting whether there is speech activity in the audio to distinguish between user speech and background noise.

[0089] In the embodiments of the present application, if the target audio data (i.e., the user's clear speech signal) is obtained, the audio data can be then subjected to feature extraction to obtain audio features (i.e., spectral features, pitch features, and volume features) that can represent the emotional state.

[0090] In the embodiments, spectral features in the target audio data can be extracted. The spectral features reflect the frequency composition of the audio signal. In emotion recognition, an important spectral feature is the Mel-frequency cepstral coefficient, which can capture the frequency information in the speech signal that is closely related to human perception. The Mel-frequency cepstral coefficient is usually calculated through fast Fourier transform and a Mel-frequency filter bank, and can reflect the timbre characteristics of speech, which is a key to recognizing different emotional states. In addition, pitch features in the target audio data can be extracted. The pitch feature refers to the fundamental frequency or main frequency of the sound, which is related to the intonation of speech and can reflect the speaker's tone changes. In emotion recognition, the rise or fall of the pitch can imply the speaker's emotional state, such as different changes in pitch when happy or angry. Furthermore, volume features in the target audio data can be extracted. The volume feature is the size of the sound. The emotional state is related to the volume, for example, when angry or excited, the volume of the sound will increase; while when sad or tired, the volume may decrease. The extraction of the volume feature helps the model to identify the intensity of the user's emotion.

[0091] Optionally, the first state result, the second state result, and the third state result each at least includes: a positive state and a positive probability value corresponding to the positive state, a negative state and a negative probability value corresponding to the negative state. In order to further improve the accuracy of determining the target state of the target object at each time point and the target state probability value corresponding to the target state, in the method for processing audio and video data provided in Embodiment 1 of the present application, the positive probability value and the negative probability value in the state result are compared, in the case that the positive probability value is greater than or equal to the negative probability value, it is determined that the state corresponding to the state result is the positive state, or, in the case that the positive probability value is less than the negative probability value, it is determined that the state corresponding to the state result is the negative state, wherein the state result is the first state result, the second state result, or the third state result; based on the states respectively corresponding to all the state results, the target state of the target object at the current time point is determined; and based on the probability values corresponding to the states in all the state results consistent with the target state, the target state probability value corresponding to the target state is determined.

[0092] In the embodiments of the present application, the first state result, the second state result and the third state result each at least includes: a positive emotion state and a positive probability value corresponding to the positive emotion state, a negative emotion state and a negative probability value corresponding to the negative emotion state, and the sum of the positive probability value and the negative probability value in the same state result is 1, for example, the probability of outputting a positive emotion is 0.7, and the probability of outputting a negative emotion is 0.3.

[0093] In the embodiments of the present application, for each emotion recognition result (the first state result, the second state result or the third state result), the positive probability value (positive probability value) and the negative probability value (negative probability value) can be compared. If the positive probability value is greater than or equal to the negative probability value, the emotion state corresponding to the state result is determined to be a positive emotion state; otherwise, if the positive probability value is less than the negative probability value, the emotion state is determined to be a negative emotion state. This process ensures that the system can make a clear emotion classification, i.e., positive or negative, based on the probability value.

[0094] In the embodiments of the present application, after obtaining the emotion states of all emotion recognition results, the system will perform comprehensive analysis to determine the target emotion state of the target object at the current time point. If the emotion states of most emotion recognition results are positive, the target emotion state is determined to be positive; otherwise, if the emotion states of most emotion recognition results are negative, the target emotion state is determined to be negative. For example, the results of the first emotion and the second emotion recognition are positive emotions, and the result of the third emotion recognition is a negative emotion, and it is considered that the final emotion recognition result is a positive emotion. This majority voting or weighted average decision mechanism fully utilizes the complementarity of multi-modal information, avoids errors caused by single method emotion recognition, and improves the robustness and accuracy of emotion recognition.

[0095] In the embodiments of the present application, a target state probability value corresponding to the target emotion state can be calculated, which is determined based on the state probability values in all emotion recognition results consistent with the target emotion state. If the target emotion state is positive, the system will calculate the final positive state probability value based on the positive probability values of all positive state results; if the target emotion state is negative, the final negative state probability value will be calculated based on the negative probability values of all negative state results. This process usually involves weighted average of probability values to reflect the contribution degree of different modal information to emotion recognition. For example, the first emotion is a positive emotion and the first probability value is 0.7, the second emotion is a positive emotion and the first probability value is 0.6. The third emotion is a negative emotion and the second probability value is 0.6, and the final emotion recognition result is a positive emotion, and the corresponding probability value is (0.7+0.6) / 2=0.65.

[0096] In order to improve the accuracy of constructing the curve of the teller state change, in the method for processing audio and video data provided in Embodiment 1 of the present application, the second audio and video data of the target teller is collected, wherein the target teller is the teller who handles the financial business for the target object, and the second audio and video data carries a time stamp; the second audio and video data is processed to obtain the teller state of the target teller at each time point and the teller state probability value corresponding to the teller state; based on the teller state at each time point and the teller state probability value corresponding to the teller state, the curve of the teller state change of the target teller is constructed; and the time point at which the teller state is the negative state is displayed on the curve of the teller state change in a preset form.

[0097] In the embodiment of the present application, the in-counter collection device in the emotion warning system can collect the audio and video data (i.e., the second audio and video data) of the target teller (i.e., the teller who handles the financial business for the target object) in real time, and mark these data with a time stamp. The addition of the time stamp makes the audio and video data of the teller consistent with the audio and video data of the user in time, facilitating subsequent emotion synchronization analysis.

[0098] In the embodiment of the present application, similar to the recognition process of the user emotion, the system separates the audio and video data of the target teller to obtain two parts of video and audio. The video data is used to extract facial features and action features, and the audio data is used to extract sound features. These features are then input into the corresponding model to identify the emotional state (positive or negative) of the teller at each time point and its probability value. This step fully utilizes the multi-modal analysis capability of the system for emotion recognition, ensuring the accuracy and real-time performance of the teller emotion state recognition.

[0099] Then, based on the teller emotional state at each time point and its corresponding teller state probability value, the emotional change curve (i.e., the teller emotional change curve) of the target teller is constructed. This curve takes time as the horizontal coordinate and the emotional state and its probability value as the vertical coordinate, directly showing the fluctuation of the teller emotion over time. In addition, the system can also set a warning threshold for negative emotions, and when the teller emotional state reaches or exceeds this threshold, the system will trigger the warning mechanism.

[0100] In the embodiment of the present application, on the teller emotional change curve, the system highlights the time points at which the teller emotional state is the negative emotional state in a preset form (such as different colors, bold, flashing, etc.), which facilitates the management personnel to quickly identify the abnormal period of the teller emotional fluctuation, so that appropriate intervention measures can be taken in time.

[0101] In this embodiment, the emotional warning system not only monitors the user's emotional changes, but also pays attention to the emotional state of the teller, thereby comprehensively improving the emotional management level of the counter service, providing a comprehensive and real-time emotional monitoring and warning platform for financial institution outlets. This not only helps financial institutions to timely handle emotional problems, but also promotes the optimization of service processes and improves the overall service quality.

[0102] In order to facilitate subsequent determination of warning prompt information, in the audio and video data processing method provided in Embodiment 1 of the present application, after collecting the second audio and video data of the target teller, the audio data in the first audio and video data and the audio data of the second audio and video data are integrated according to the time stamp to obtain the dialogue text; based on the teller state change curve, the dialogue content corresponding to all time points with negative state of the target teller is determined in the dialogue text; the dialogue text is input into a preset language model, and the summary content is output, wherein the preset language model is a model trained by using a historical dialogue text set and the annotated summary content annotated for each historical dialogue text in the historical dialogue text set.

[0103] In the embodiment of the present application, the audio of the user (first audio and video data) and the teller (second audio and video data) can be aligned and integrated according to the time stamp on each audio data to obtain the dialogue text. This integration process ensures that the context information of the user and the teller dialogue remains consistent, and the generated dialogue text can accurately reflect the communication content of both parties at different time points. Through voice recognition technology, the audio is converted into text, so that the dialogue content can be further analyzed.

[0104] Then, based on the teller emotional change curve, all time points of the target teller in the negative emotional state are screened out. For these time points, the system extracts the corresponding dialogue content from the integrated dialogue text, i.e. the dialogue record of the teller with the user in the negative emotional state. This analysis process helps to understand the specific reasons for emotional changes and provides more targeted data support for subsequent emotional management.

[0105] In the embodiment of the present application, since the dialogue text may be long and contain some routine dialogue content, viewing the complete dialogue text may consume too much time and result in failure to intervene in time. Therefore, in order to more efficiently understand the key information in the dialogue text, the dialogue text can be input into a preset language model, which is trained by using a historical dialogue text set and corresponding annotated summary content. The model can automatically identify the theme, emotional trigger points and important details in the dialogue, and output concise dialogue summary content. This summary content not only contains the main information of the dialogue, but also highlights the specific dialogue of the teller in the negative emotional state, facilitating quick understanding of the details of emotional problems.

[0106] To improve the accuracy of the state prompt information, in the method for processing audio and video data provided in Embodiment 1 of the present application, the state prompt information is triggered based on the conversation content before the current time point and the summary content before the current time point when it is monitored that the state change curve reaches the first probability threshold line or the second probability threshold line.

[0107] In the embodiments of the present application, the emotional change curves of the user and the teller can be continuously monitored, which reflect the fluctuations of the emotional state over time and are attached with the probability values of the emotional state. The first probability threshold line and the second probability threshold line on the emotional change curve represent the warning lines of the transition from positive emotion to negative emotion and the increase in the intensity of negative emotion, respectively. When the system monitors that the emotional change curve is about to touch these warning lines, it does not wait until the threshold is actually crossed to take action, but triggers the state prompt information in advance based on the conversation content before the current time point and the summary content generated by the preset language model. The design of this early warning mechanism aims to capture signs of emotional change in advance and provide enough time for the branch staff to take preventive measures to avoid the deterioration of emotional problems.

[0108] In the embodiments of the present application, when the early warning is triggered, the system provides all the relevant conversation content before the current time point and the summary content, which contains the background and specific context of triggering the early warning, helping the staff to understand the root cause of the emotional problem, so as to be able to intervene more accurately. For example, the staff can see which topics or service links in the conversation may have triggered the negative emotion of the user, and which emotional states the teller has shown in this process.

[0109] In the embodiments of the present application, if the first intervention opportunity (i.e., before the emotional change curve reaches the first probability threshold line) is missed, or it is subjectively judged that there is no need for intervention, a bottom-up intervention can also be made at the second intervention opportunity (i.e., before the emotional change curve reaches the second probability threshold line). And since the early warning prompt occurs before the first probability threshold or the second probability threshold is actually reached, it also gives enough preparation time to the relevant personnel.

[0110] In the present embodiment, by triggering the state prompt information before the emotional change reaches the warning line, the emotional early warning system not only can monitor the emotional fluctuations in real time, but also can actively warn and provide decision-making basis for the staff, so as to realize the prevention and management of emotional problems, thereby effectively improving the quality of the counter service and the satisfaction of the user.

[0111] Figure 4 is a schematic diagram of the user's audio and video processing flow according to Embodiment 1 of the present application, like Figure 4As shown, when the user conducts business at the counter, the audio and video data of the user can be acquired in real time by the collection device outside the counter. Then, the audio and video data of the user are separated to obtain user video and user audio. The user video is processed by a labeled rectangular frame to generate a cropped video. The video focuses on the upper body of the user, especially the face, to improve the accuracy of emotion recognition. Subsequently, based on the cropped video, facial features and action features are extracted, and the facial features are input into a first emotion recognition model (i.e., a first model) to obtain a first emotion, and the action features are input into a second emotion recognition model (i.e., a second model) to obtain a second emotion. Meanwhile, the user audio is filtered to remove background noise and ensure that only the user's voice is retained to obtain filtered audio. The filtered audio data is input into a third emotion recognition model (i.e., a third model) to obtain a third emotion. Moreover, the filtered audio is converted to obtain timestamped text, which is saved in the form of text + timestamp.

[0112] The method for processing audio and video data provided in the embodiments of the present application can obtain the target emotional state of the target object at each time point and the target state probability value corresponding to the target emotional state by performing real-time analysis on the collected audio and video data of the target object, and then construct an emotional change curve that changes over time, so as to timely trigger a state prompt information before the emotional change curve reaches the first probability threshold line or the second probability threshold line. In this way, not only can the emotional state of the customer be determined in real time and accurately, but also the intervention can be performed in time before the emotional state of the customer deteriorates further, which is beneficial to improving the customer satisfaction, achieves the technical effects of monitoring the customer emotion in real time and improving the customer satisfaction, and further solves the technical problem of low accuracy in determining the emotional state of the customer in the related art.

[0113] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0114] Embodiment 2

[0115] The embodiments of the present application also provide a processing device for audio and video data. It should be noted that the processing device for audio and video data of the embodiments of the present application can be used to execute the method for processing audio and video data provided by the embodiments of the present application. The processing device for audio and video data provided by the embodiments of the present application is introduced as follows.

[0116] According to the embodiments of the present application, a device for implementing the above-mentioned method for processing audio and video data is also provided. Figure 5 is a schematic diagram of an optional processing device for audio and video data according to an embodiment of the present application, likeFigure 5 The processing device can include: a collection unit 50, a determination unit 51, a construction unit 52, and a triggering unit 53.

[0117] The collection unit 50 is configured to collect first audio-video data of the target object under the condition that the target object is detected to handle a financial service and the target object is authorized, wherein the first audio-video data carries a timestamp.

[0118] The determination unit 51 is configured to determine a target state of the target object at each time point and a target state probability value corresponding to the target state based on the first audio-video data.

[0119] The construction unit 52 is configured to construct a state change curve of the target object based on the target state at each time point and the target state probability value corresponding to the target state.

[0120] The triggering unit 53 is configured to trigger a state prompt message before the state change curve reaches a first probability threshold line or a second probability threshold line.

[0121] The processing device for audio-video data provided by the embodiments of the present application can obtain the target state of the target object at each time point and the target state probability value corresponding to the target state by performing real-time analysis on the collected audio-video data of the target object, and then construct a state change curve that changes with time, so as to timely trigger a state prompt message before the state change curve reaches a first probability threshold line or a second probability threshold line. In this way, the emotional state of the customer can be determined in real time and accurately, and intervention can be performed in time before the emotional state of the customer deteriorates further, which is beneficial to improving customer satisfaction, achieves the technical effects of real-time monitoring of the emotional state of the customer and improvement of customer satisfaction, and further solves the technical problem of low accuracy in determining the emotional state of the customer in the related art.

[0122] Optionally, the determination unit includes: a first separation module configured to separate the first audio-video data to obtain video data and audio data; a first processing module configured to process the video data to obtain target video data; a first extraction module configured to extract facial feature data and action feature data of the target object based on the target video data; a second processing module configured to process the facial feature data, the action feature data, and the audio data respectively to obtain a first state result, a second state result, and a third state result; and a first determination module configured to determine the target state of the target object at each time point and the target state probability value corresponding to the target state based on the first state result, the second state result, and the third state result.

[0123] Optionally, the first processing module comprises: a first labeling submodule configured to label each frame of image in the video data to obtain a rectangular frame on each frame of image; and a first cropping submodule configured to crop each frame of image based on the rectangular frame on each frame of image to obtain the target video data.

[0124] Optionally, the second processing module comprises: a first processing submodule configured to process the facial feature data by using a first model to obtain a first state result, wherein the first model is a model trained by using first historical data, and the first historical data comprises a set of historical facial feature data and a state label result of labeling each historical facial feature data in the set of historical facial feature data; a second processing submodule configured to process the action feature data by using a second model to obtain a second state result, wherein the second model is a model trained by using second historical data, and the second historical data comprises a set of historical action feature data and a state label result of labeling each historical action feature data in the set of historical action feature data; and a third processing submodule configured to process the audio data by using a third model to obtain a third state result, wherein the third model is a model trained by using third historical data, and the second historical data comprises a set of historical audio data and a state label result of labeling each historical audio data in the set of historical audio data.

[0125] Optionally, the processing device further comprises: a first filtering module configured to filter the audio data to obtain target audio data before processing the audio data by using the third model to obtain the third state result; and a second extraction module configured to extract features from the target audio data to obtain a spectrum feature, a pitch feature and a volume feature.

[0126] Optionally, the first state result, the second state result and the third state result each at least comprises: a positive state and a positive probability value corresponding to the positive state, a negative state and a negative probability value corresponding to the negative state, the first determining module comprises: a first comparison submodule configured to compare the positive probability value and the negative probability value in the state result, and determine that the state corresponding to the state result is the positive state in a case that the positive probability value is greater than or equal to the negative probability value, or determine that the state corresponding to the state result is the negative state in a case that the positive probability value is less than the negative probability value, wherein the state result is the first state result, the second state result or the third state result; a first determination submodule configured to determine a target state of the target object at a current time point based on states respectively corresponding to all state results; and a second determination submodule configured to determine a target state probability value corresponding to the target state based on probability values corresponding to states in all state results consistent with the target state.

[0127] Optionally, the processing apparatus further comprises: a first acquisition module, configured to acquire second audio-video data of a target teller, wherein the target teller is a teller who handles a financial service for the target object, and the second audio-video data carries a timestamp; a third processing module, configured to process the second audio-video data to obtain a teller state of the target teller at each time point and a teller state probability value corresponding to the teller state; a first construction module, configured to construct a teller state change curve of the target teller based on the teller state at each time point and the teller state probability value corresponding to the teller state; and a first display module, configured to display a time point at which the teller state is a negative state on the teller state change curve in a preset form.

[0128] Optionally, the processing apparatus further comprises: a first integration module, configured to integrate audio data in the first audio-video data and audio data of the second audio-video data according to the timestamp after the second audio-video data of the target teller is acquired, to obtain conversation text; a second determination module, configured to determine conversation content corresponding to all time points at which the target teller has a negative state in the conversation text based on the teller state change curve; and a first input module, configured to input the conversation text into a preset language model to output summary content, wherein the preset language model is a model trained by using a historical conversation text set and a labeled summary content labeled for each historical conversation text in the historical conversation text set.

[0129] Optionally, the triggering unit comprises: a first triggering module, configured to trigger the state prompt information based on the conversation content before the current time point and the summary content before the state change curve reaches the first probability threshold line or the second probability threshold line.

[0130] The processing apparatus described above can further comprise a processor and a memory, and the above-mentioned acquisition unit 50, determination unit 51, construction unit 52, and triggering unit 53 are all stored in the memory as program units, and the corresponding functions are realized by the processor executing the above-mentioned program units stored in the memory.

[0131] The above-mentioned processor comprises a core, and the core retrieves the corresponding program units from the memory. The core can be set to one or more, and the core parameters are adjusted to trigger the state prompt information before the state change curve reaches the first probability threshold line or the second probability threshold line.

[0132] The above-mentioned memory can comprise a non-permanent memory in a computer readable medium, a random access memory (RAM) and / or a non-volatile memory such as a read-only memory (ROM) or a flash memory (flash RAM), and the memory comprises at least one memory chip.

[0133] It should be noted that the aforementioned acquisition unit 50, determination unit 51, construction unit 52, and triggering unit 53 correspond to steps S201 to S204 in Embodiment 1. The instances and application scenarios implemented by the aforementioned units and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should be noted that the aforementioned units may be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The aforementioned units may also be part of a device and can run in the computer terminal 10 provided in Embodiment 1.

[0134] Example 3

[0135] Embodiments of this application may provide an electronic device. Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 (Only one is shown) processor 602, memory 604, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0136] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the audio and video data processing method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned audio and video data processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0137] The processor can access information and applications stored in the memory via a transmission device to execute the following steps: upon detecting that a target object is conducting financial business and has obtained authorization from the target object, the processor collects first audio and video data of the target object, wherein the first audio and video data carries a timestamp; based on the first audio and video data, the processor determines the target object's target state at each time point and the target state probability value corresponding to the target state; based on the target state at each time point and the target state probability value corresponding to the target state, the processor constructs a state change curve of the target object; before the state change curve is detected to reach a first probability threshold line or a second probability threshold line, the processor triggers a state prompt message.

[0138] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: separating the first audio and video data to obtain video data and audio data; processing the video data to obtain target video data; extracting facial feature data and action feature data of the target object based on the target video data; processing the facial feature data, the action feature data, and the audio data respectively to obtain a first state result, a second state result, and a third state result; and determining a target state of the target object at each time point and a target state probability value corresponding to the target state based on the first state result, the second state result, and the third state result.

[0139] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: labeling each frame of image in the video data to obtain a rectangular frame on each frame of image; and cropping each frame of image based on the rectangular frame on each frame of image to obtain the target video data.

[0140] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: processing the facial feature data by using a first model to obtain a first state result, wherein the first model is a model trained by using first historical data, and the first historical data includes a set of historical facial feature data and a state label result of labeling each historical facial feature data in the set of historical facial feature data; processing the action feature data by using a second model to obtain a second state result, wherein the second model is a model trained by using second historical data, and the second historical data includes a set of historical action feature data and a state label result of labeling each historical action feature data in the set of historical action feature data; and processing the audio data by using a third model to obtain a third state result, wherein the third model is a model trained by using third historical data, and the second historical data includes a set of historical audio data and a state label result of labeling each historical audio data in the set of historical audio data.

[0141] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: filtering the audio data to obtain target audio data; and extracting a frequency spectrum feature, a pitch feature, and a volume feature from the target audio data.

[0142] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: comparing the positive probability value and the negative probability value in the state result, determining that the state corresponding to the state result is a positive state when the positive probability value is greater than or equal to the negative probability value, or determining that the state corresponding to the state result is a negative state when the positive probability value is less than the negative probability value, wherein the state result is the first state result, the second state result or the third state result; determining the target state of the target object at the current time point based on the states corresponding to all state results; and determining the target state probability value corresponding to the target state based on the probability values corresponding to the states in all state results consistent with the target state.

[0143] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: collecting second audio and video data of the target teller, wherein the target teller is a teller who handles financial business for the target object, and the second audio and video data carries a time stamp; processing the second audio and video data to obtain the teller state of the target teller at each time point and the teller state probability value corresponding to the teller state; constructing a teller state change curve of the target teller based on the teller state at each time point and the teller state probability value corresponding to the teller state; and displaying the time point at which the teller state is a negative state on the teller state change curve in a preset form.

[0144] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: integrating the audio data in the first audio and video data and the audio data of the second audio and video data according to the time stamp to obtain a dialogue text; determining the dialogue content corresponding to all time points at which the target teller has a negative state in the dialogue text based on the teller state change curve; and inputting the dialogue text into a preset language model to output a summary content, wherein the preset language model is a model trained using a historical dialogue text set and a labeled summary content labeled for each historical dialogue text in the historical dialogue text set.

[0145] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: based on the dialogue content before the current time point and the summary content, triggering a state prompt information before the state change curve reaches the first probability threshold line or the second probability threshold line is monitored.

[0146] By adopting the embodiment of the application, a scheme for processing audio and video data is provided. By performing real-time analysis on the collected audio and video data of the target object, the target state of the target object at each time point and the target state probability value corresponding to the target state can be obtained, and then a state change curve changing over time is constructed, so as to timely trigger the state prompt information before the state change curve reaches the first probability threshold line or the second probability threshold line. In this way, not only the emotional state of the customer can be determined in real time and accurately, but also intervention can be performed in time before the emotional state of the customer deteriorates further, which is beneficial to improving the customer satisfaction, and the technical effects of monitoring the emotional state of the customer in real time and improving the customer satisfaction are achieved, and the technical problem of low accuracy in determining the emotional state of the customer in the related art is solved.

[0147] Those skilled in the art can understand that Figure 6 The structure shown is only schematic, and the electronic device can also be a terminal device such as a smart phone, a tablet computer, a palm computer, a mobile Internet device (MID), and the like. Figure 6 It does not limit the structure of the electronic device. For example, the electronic device can further include more or fewer components (such as a network interface, a display device, and the like) than those shown in the embodiment, or have a different configuration from that shown in the embodiment. Figure 6 It does not limit the structure of the electronic device. For example, the electronic device can further include more or fewer components (such as a network interface, a display device, and the like) than those shown in the embodiment, or have a different configuration from that shown in the embodiment. Figure 6 It does not limit the structure of the electronic device. For example, the electronic device can further include more or fewer components (such as a network interface, a display device, and the like) than those shown in the embodiment, or have a different configuration from that shown in the embodiment.

[0148] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the related hardware of the terminal device, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like.

[0149] Embodiment 4

[0150] The embodiment of the application further provides a storage medium. Optionally, in the embodiment, the storage medium can be used to save the program code executed by the processing method of audio and video data provided in the embodiment 1.

[0151] Optionally, in the embodiment, the storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0152] The application further provides a computer program product adapted to execute the steps of the processing method of audio and video data when executed on a data processing device.

[0153] The above-mentioned sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0154] In the above-mentioned embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0155] In the several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the unit embodiment described above is only illustrative, and for example, the division of units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.

[0156] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0157] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of software functional unit.

[0158] If the integrated unit is realized in the form of software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the embodiments of the present application. The above-mentioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and various program codes that can be stored in the medium.

[0159] The above merely preferred embodiments of the present application, it should be noted that for those of ordinary skill in the art, without departing from the principles of the present application, can make several improvements and refinements, these improvements and refinements should also be considered as the scope of protection of the present application.

Claims

1. A method for processing audiovisual data, characterized in that, The method comprises the following steps: In the case of detecting that a target object handles a financial business and obtaining authorization of the target object, first audio and video data of the target object is collected, wherein the first audio and video data carries a timestamp; Based on the first audio and video data, a target state of the target object at each time point and a target state probability value corresponding to the target state are determined; Based on the target state at each time point and the target state probability value corresponding to the target state, a state change curve of the target object is constructed; Before monitoring that the state change curve reaches a first probability threshold line or a second probability threshold line, a state prompt information is triggered; The processing method further comprises the following steps: Second audio and video data of a target clerk is collected, wherein the target clerk is a clerk handling the financial business for the target object, and the second audio and video data carries a timestamp; The second audio and video data is processed to obtain a clerk state of the target clerk at each time point and a clerk state probability value corresponding to the clerk state; Based on the clerk state at each time point and the clerk state probability value corresponding to the clerk state, a clerk state change curve of the target clerk is constructed; The time point at which the clerk state is a negative state is displayed on the clerk state change curve in a preset form; After collecting the second audio and video data of the target clerk, the following steps are further included: According to the timestamp, audio data in the first audio and video data and audio data of the second audio and video data are integrated to obtain a dialogue text; Based on the clerk state change curve, dialogue content corresponding to all time points at which the target clerk has the negative state is determined in the dialogue text; The dialogue text is input into a preset language model to output a summary content, wherein the preset language model is a model trained by using a historical dialogue text set and a labeled summary content labeled for each historical dialogue text in the historical dialogue text set.

2. The treatment method according to claim 1, characterized in that, The step of determining the target state of the target object at each time point and the target state probability value corresponding to the target state based on the first audio and video data comprises the following steps: The first audio and video data is separated to obtain video data and audio data; The video data is processed to obtain target video data; Based on the target video data, facial feature data and action feature data of the target object are extracted; The facial feature data, the action feature data and the audio data are processed respectively to obtain a first state result, a second state result and a third state result; Based on the first state result, the second state result and the third state result, the target state of the target object at each time point and the target state probability value corresponding to the target state are determined.

3. The treatment method according to claim 2, characterized in that, The step of processing the video data to obtain target video data comprises the following steps: Each frame of image in the video data is labeled to obtain a rectangular frame on each frame of image; Based on the rectangular frame on each frame of image, each frame of image is cropped to obtain the target video data.

4. The treatment method of claim 2, wherein The steps of processing the face feature data, the action feature data and the audio data respectively to obtain the first state result, the second state result and the third state result, comprising: The first model is used to process the face feature data to obtain the first state result, wherein the first model is a model trained by first historical data, and the first historical data includes a set of historical face feature data and a state label result of labeling each historical face feature data in the set of historical face feature data; The second model is used to process the action feature data to obtain the second state result, wherein the second model is a model trained by second historical data, and the second historical data includes a set of historical action feature data and a state label result of labeling each historical action feature data in the set of historical action feature data; The third model is used to process the audio data to obtain the third state result, wherein the third model is a model trained by third historical data, and the second historical data includes a set of historical audio data and a state label result of labeling each historical audio data in the set of historical audio data.

5. The treatment method according to claim 4, characterized in that, Before the third model is used to process the audio data to obtain the third state result, it further comprises: Filtering the audio data to obtain target audio data; Extracting features from the target audio data to obtain spectral features, pitch features and volume features.

6. The treatment method of claim 2, wherein, The first state result, the second state result and the third state result all include at least a positive state and a positive probability value corresponding to the positive state, a negative state and a negative probability value corresponding to the negative state, based on the first state result, the second state result and the third state result, the steps of determining the target state of the target object at each time point and the target state probability value corresponding to the target state, comprising: Comparing the positive probability value and the negative probability value in the state result, in the case that the positive probability value is greater than or equal to the negative probability value, determining that the state corresponding to the state result is the positive state, or in the case that the positive probability value is less than the negative probability value, determining that the state corresponding to the state result is the negative state, wherein the state result is the first state result, the second state result or the third state result; Based on the states corresponding to all the state results respectively, determining the target state of the target object at the current time point; Based on the probability values corresponding to the states in all the state results consistent with the target state, determining the target state probability value corresponding to the target state.

7. The treatment method of claim 1, wherein Before monitoring the state change curve to reach the first probability threshold line or the second probability threshold line, the step of triggering the state prompt information, comprising: Before monitoring that the state change curve reaches the first probability threshold line or the second probability threshold line, triggering the state prompt information based on the dialogue content before the current time point and the summary content.

8. An apparatus for processing audiovisual data, characterized in that The method comprises the steps of: The acquisition unit is configured to acquire first audio and video data of a target object under the condition that the target object is detected to handle a financial service and authorization of the target object is obtained, wherein the first audio and video data carries a timestamp. The determination unit is configured to determine a target state of the target object at each time point and a target state probability value corresponding to the target state based on the first audio and video data. The construction unit is configured to construct a state change curve of the target object based on the target state at each time point and the target state probability value corresponding to the target state. The triggering unit is configured to trigger a state prompt information before monitoring that the state change curve reaches the first probability threshold line or the second probability threshold line. The device further comprises a first acquisition module configured to acquire second audio and video data of a target teller, wherein the target teller is a teller handling a financial service for the target object, and the second audio and video data carries a timestamp; a third processing module configured to process the second audio and video data to obtain a teller state of the target teller at each time point and a teller state probability value corresponding to the teller state; a first construction module configured to construct a teller state change curve of the target teller based on the teller state at each time point and the teller state probability value corresponding to the teller state; and a first display module configured to display a time point at which the teller state is a negative state on the teller state change curve in a preset form. The device further comprises a first integration module configured to integrate audio data in the first audio and video data and audio data of the second audio and video data according to the timestamps to obtain dialogue text after acquiring the second audio and video data of the target teller; a second determination module configured to determine dialogue content corresponding to all time points at which the target teller has a negative state in the dialogue text based on the teller state change curve; and a first input module configured to input the dialogue text into a preset language model to output summary content, wherein the preset language model is a model trained by using a historical dialogue text set and annotated summary content annotated for each historical dialogue text in the historical dialogue text set.

9. A computer program product, characterised in that, The non-transitory computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the audio and video data processing method in any one of claims 1 to 7.

10. An electronic device, comprising: The one or more processors and the memory are configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the audio and video data processing method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-mode intelligent emotion sensing system

    CN107220591A

  • Multi-modal supervision service system and method

    CN112700255A

  • Transaction risk judgment method and device, equipment and medium

    CN115689776A