Customer Satisfaction Identification Method, Device, Equipment and Medium

By acquiring customer images and generating emotion evaluation vectors using pose and expression prediction results, the problem of difficulty in accurately obtaining customer emotions in the prior art is solved, and more accurate customer satisfaction recognition is achieved.

CN114973419BActive Publication Date: 2025-06-17INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210701055.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2025-06-17
Estimated Expiration
2042-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately obtain customer emotional changes in a specific time period, resulting in the inability to form a comprehensive evaluation image of customer satisfaction and it is difficult to improve product or service levels.

Method used

By obtaining customer images of M moments in a predetermined time period, using the gesture and expression prediction results to generate an emotion evaluation vector, assemble it into a sequence of time-sequential emotion vectors, and input it into a pre-trained satisfaction recognition model to obtain customer satisfaction recognition results.

Benefits of technology

Customer satisfaction recognition is achieved by combining posture and expression, considering customer mood changes, avoiding the problem of inaccurate single prediction results, and obtaining more accurate customer satisfaction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973419B_ABST
    Figure CN114973419B_ABST
Patent Text Reader

Abstract

The present disclosure provides a customer satisfaction recognition method, which relates to the field of artificial intelligence. The method includes: obtaining customer images of a customer at M moments within a predetermined time period; obtaining a posture prediction result and an expression prediction result of the customer at each of the M moments according to the customer images at each of the M moments; obtaining an emotion evaluation vector at each of the moments according to the posture prediction result and the expression prediction result at each of the moments; assembling the emotion evaluation vectors at each of the moments in the time sequence of the M moments to obtain a time-sequence emotion vector sequence; and inputting the time-sequence emotion vector sequence into a pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result. The present disclosure also provides a customer satisfaction recognition device, equipment, storage medium and program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more particularly, to a method, apparatus, device, medium, and program product for identifying customer satisfaction. Background Art

[0002] Customer satisfaction refers to the degree of customer satisfaction, which can represent the different degrees of satisfaction obtained by customers after purchasing and consuming corresponding products or services. Enterprises or individuals providing products or services mainly collect customer satisfaction through offline, telephone, or Internet means. For example, a bank can obtain feedback through direct means such as business halls, counter work, and suggestion boxes, and indirect means such as complaint hotlines. However, it may not cover every customer, and the proportion of negative satisfaction in these feedback messages may be greater than that of positive satisfaction, resulting in the inability to form a comprehensive evaluation portrait of customer satisfaction within a specific range (such as business halls, business personnel, or products), and it is also difficult to obtain positive feedback on whether a certain convenience measure should be continued, thereby improving the product or service level. Therefore, how to obtain accurate customer satisfaction results is an urgent problem to be solved currently. Summary of the Invention

[0003] In view of the above problems, the present disclosure provides a method, apparatus, device, medium, and program product for identifying customer satisfaction by obtaining the emotional changes of customers to determine customer satisfaction.

[0004] An aspect of an embodiment of the present disclosure provides a method for identifying customer satisfaction, including: obtaining customer images of a customer at M moments within a predetermined time period, where M is an integer greater than or equal to 2; obtaining a posture prediction result and an expression prediction result of the customer at each of the M moments according to the customer images at each of the M moments; obtaining an emotion evaluation vector at each of the moments according to the posture prediction result and the expression prediction result at each of the moments; assembling the emotion evaluation vectors at each of the moments in the time order corresponding to the moments in the M moments to obtain a time-series emotion vector sequence, where the time-series emotion vector sequence includes M emotion evaluation vectors; inputting the time-series emotion vector sequence into a pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result, where the customer satisfaction recognition result is used to represent the emotional change trend of the customer within the predetermined time period.

[0005] According to an embodiment of the present disclosure, obtaining the emotion evaluation vector for each moment based on the posture prediction result and the expression prediction result for each moment includes: inputting the posture prediction result and the expression prediction result for each moment into an emotion evaluation knowledge model to obtain an output vector of the emotion evaluation knowledge model, where the emotion evaluation knowledge model is constructed based on a long short-term memory network; obtaining index data according to the output vector; and retrieving an emotion evaluation vector matching the index data from an emotion knowledge base, where the emotion knowledge base includes at least one emotion evaluation vector to be matched.

[0006] According to an embodiment of the present disclosure, before retrieving the emotion evaluation vector matching the index data from the emotion knowledge base, obtaining the emotion knowledge base is further included, specifically including: inputting an emotion text data set into a differentiable neural computer model for processing to obtain an external storage matrix of the differentiable neural computer model, where the emotion text data set includes at least one text for training the differentiable neural computer model; and obtaining the emotion knowledge base according to the external storage matrix.

[0007] According to an embodiment of the present disclosure, the satisfaction recognition model includes a first neural network model, a second neural network model, and a classification model, and the classification model is a machine learning model. Inputting the time-series emotion vector sequence into a pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result includes: enabling the first neural network model to process the time-series emotion vector sequence to obtain an emotion change heat vector; enabling the second neural network model to process the time-series emotion vector sequence to obtain an emotion evaluation intermediate vector; and inputting the emotion change heat vector and the emotion evaluation intermediate vector into the classification model to obtain the customer satisfaction recognition result.

[0008] According to an embodiment of the present disclosure, obtaining customer images at M moments within a predetermined time period includes: obtaining video data within the predetermined time period, where at least two image frames in the video data include the customer; performing customer re-identification on N image frames in the video data, N being an integer greater than or equal to 2 and N being greater than or equal to M; and obtaining the customer images at the M moments from the N image frames based on the result of the customer re-identification.

[0009] According to an embodiment of the present disclosure, obtaining the posture prediction result of the customer at each moment according to the customer image at each moment among the M moments includes: inputting the customer image at each moment into a posture recognition model to obtain the posture prediction result at each moment, where the posture prediction result at each moment is used to characterize the first emotion of the customer at that moment, and the posture recognition model is constructed by a densely connected method.

[0010] According to an embodiment of the present disclosure, obtaining the expression prediction result of the customer at each of the M moments based on the customer images at each of the M moments includes: inputting the customer images at each of the M moments into an expression recognition model to obtain the expression prediction result at each of the M moments, where the expression prediction result at each of the M moments is used to characterize the second emotion of the customer at that moment, and the first emotion and the second emotion are the same or different, and the expression recognition model is constructed by a hierarchical feature aggregation method.

[0011] According to an embodiment of the present disclosure, inputting the customer images at each of the M moments into the expression recognition model includes: performing super-resolution reconstruction on the customer images at each of the M moments to obtain super-resolution customer images at each of the M moments; inputting the super-resolution customer images at each of the M moments into the expression recognition model.

[0012] Another aspect of the embodiments of the present disclosure provides a customer satisfaction recognition device, including: an image acquisition module, configured to acquire customer images of a customer at M moments within a predetermined time period, where M is an integer greater than or equal to 2; a posture and expression prediction module, configured to obtain a posture prediction result and an expression prediction result of the customer at each of the M moments according to the customer images at each of the M moments; a first vector module, configured to obtain an emotion evaluation vector at each of the M moments according to the posture prediction result and the expression prediction result at each of the M moments; a second vector module, configured to assemble the emotion evaluation vectors at each of the M moments in the time sequence of the corresponding moments in the M moments to obtain a time-sequence emotion vector sequence, where the time-sequence emotion vector sequence includes M emotion evaluation vectors; a satisfaction recognition module, configured to input the time-sequence emotion vector sequence into a pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result, where the customer satisfaction recognition result is used to characterize the emotion change trend of the customer within the predetermined time period.

[0013] Another aspect of the embodiments of the present disclosure provides an electronic device, including: one or more processors; a storage device, configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method as described above.

[0014] Another aspect of the embodiments of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method as described above.

[0015] Another aspect of the embodiments of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.

[0016] One or more of the above embodiments have the following beneficial effects: obtaining the pose prediction results and expression prediction results at each of the M moments, and obtaining the emotion evaluation vector at each moment. Then, the M emotion evaluation vectors are assembled in the time sequence of the M moments according to the corresponding moments to obtain a time-sequence emotion vector sequence. The time-sequence emotion vector sequence is input into a pre-trained satisfaction recognition model to obtain the customer satisfaction recognition result. Therefore, it is possible to combine the pose and the expression, avoiding inaccurate prediction results obtained by only using one of them, and taking into account the emotional changes of the customer within a specific time period. By using the time-sequence emotion vector sequence obtained in the time series dimension for recognition, an accurate customer satisfaction result can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, the above content and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:

[0018] Figure 1 Schematically shows an application scenario diagram of the customer satisfaction recognition method according to an embodiment of the present disclosure;

[0019] Figure 2 Schematically shows a flowchart of the customer satisfaction recognition method according to an embodiment of the present disclosure;

[0020] Figure 3 Schematically shows a flowchart of obtaining a customer portrait according to an embodiment of the present disclosure;

[0021] Figure 4 Schematically shows an architecture diagram of a pose recognition model according to an embodiment of the present disclosure;

[0022] Figure 5 Schematically shows an architecture diagram of an expression recognition model according to an embodiment of the present disclosure;

[0023] Figure 6 Schematically shows a flowchart of reconstructing a customer image according to an embodiment of the present disclosure;

[0024] Figure 7 Schematically shows an architecture diagram of an emotion evaluation knowledge model according to an embodiment of the present disclosure;

[0025] Figure 8 Schematically shows a flowchart of obtaining an emotion evaluation vector according to an embodiment of the present disclosure;

[0026] Figure 9 Schematically shows a flowchart of obtaining an emotion knowledge base according to an embodiment of the present disclosure;

[0027] Figure 10 Schematically shows an architecture diagram of a satisfaction recognition model according to an embodiment of the present disclosure;

[0028] Figure 11 Schematically shows a flowchart of obtaining a customer satisfaction recognition result according to an embodiment of the present disclosure;

[0029] Figure 12 Schematically shows a flowchart of a customer satisfaction recognition method according to another embodiment of the present disclosure;

[0030] Figure 13 Schematically shows a structural block diagram of a customer satisfaction recognition device according to an embodiment of the present disclosure;

[0031] Figure 14 Schematically shows a block diagram of an electronic device suitable for implementing the customer satisfaction recognition method according to an embodiment of the present disclosure. Detailed implementation manners

[0032] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.

[0033] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0034] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0035] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0036] In the technical solutions of the present disclosure, the authorization or consent of the user is obtained before acquiring or collecting the user's personal information. The processing of the collected, stored, used, processed, transmitted, provided, disclosed, and applied user personal information complies with the provisions of relevant laws and regulations, takes necessary confidentiality measures, and does not violate public order and good customs.

[0037] Figure 1 FIG. schematically shows an application scenario diagram of a customer satisfaction recognition method according to an embodiment of the present disclosure.

[0038] As Figure 1 shown, the application scenario 100 according to this embodiment may include cameras 111, 112, a server 120, customer A 131, customer B 132, and a network 140. The network 140 is used to provide a medium for a communication link between the cameras 111, 112 and the server 120. The network 140 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0039] A terminal device (not shown) may be used to receive videos or images captured by the cameras 111, 112, and the terminal device may also be used to interact with the server 120, such as sending a customer satisfaction recognition instruction or receiving a customer satisfaction recognition result, etc. Various communication client applications may be installed on the terminal device, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0040] The terminal device may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. In some embodiments, the cameras 111, 112 may be implemented by a terminal device having a camera function.

[0041] The server 120 may be a server providing various services, such as a background management server that supports the websites browsed by users using the terminal device or runs an artificial intelligence model to process data to identify customer satisfaction (only as an example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0042] It should be noted that the customer satisfaction recognition method provided by the embodiments of the present disclosure can generally be executed by the server 120. Correspondingly, the customer satisfaction recognition device provided by the embodiments of the present disclosure can generally be set in the server 120. The customer satisfaction recognition method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 120 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 120. Correspondingly, the customer satisfaction recognition device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 120 and capable of communicating with the terminal devices 101, 102, 103, and / or the server 120.

[0043] It should be understood that the application scenario 100 can be an indoor environment or an outdoor environment. Figure 1 The positions and numbers of the cameras in are only illustrative. According to the implementation requirements, there can be any number of cameras or cameras can be set at any position. Similarly, the numbers of customers, networks, and servers are only illustrative. According to the implementation requirements, there can be any number of customers, networks, and servers.

[0044] The following will be based on Figure 1 the described scenario, and will Figures 2 to 12 describe in detail the customer satisfaction recognition method of the embodiments of the present disclosure through

[0045] Figure 2 FIG. schematically shows a flowchart of the customer satisfaction recognition method according to an embodiment of the present disclosure.

[0046] As Figure 2 shown, the customer satisfaction recognition method of this embodiment includes operations S210 to S250.

[0047] In operation S210, customer images at M moments within a predetermined time period are obtained, where M is an integer greater than or equal to 2.

[0048] Referring to Figure 1 , the customer images of customer A 131 or customer B 132 can be obtained by using cameras 111 and 112. For example Figure 1 in the scenario of a bank business hall, the predetermined time period can be the time period from when the customer enters the business hall to when the customer leaves the business hall, or the time between the start and end of receiving a certain service (such as loan, deposit, or card handling service).

[0049] Exemplarily, one or more images can be obtained at each moment, and the purpose is to obtain the body image (such as head, torso, and limbs) and face image of the customer at the current moment (the face image can be obtained from the body image or separately). The body image is used for gesture recognition, and the face image is used for expression recognition.

[0050] In operation S220, according to the customer images at each of the M moments, obtain the pose prediction result and the expression prediction result of the customer at each moment.

[0051] Exemplarily, the pose prediction result and the expression prediction result can be vectors or emotion categories. The above vectors can be feature maps obtained by inputting the images into a neural network model. The emotion categories can be the results obtained by classifying the feature maps, such as categories like angry, disgusted, afraid, happy, neutral, sad, or surprised. Among them, the emotion that can be reflected by the body language (e.g., posture, movement, etc.) obtained from the customer's pose, and the emotion that can be applied by micro - expressions obtained from the customer's expression changes.

[0052] In operation S230, according to the pose prediction result and the expression prediction result at each moment, obtain the emotion evaluation vector at each moment.

[0053] Exemplarily, in this operation, the satisfaction category at each moment is not directly obtained, but an intermediate vector can be obtained to prepare for obtaining data in the time - series dimension, which can save computing resources. Obtaining the time - series emotion vector sequence according to the emotion evaluation vector at each moment, rather than according to the satisfaction category at each moment, can, to a certain extent, avoid the adverse impact on the final result caused by the inaccuracy of the satisfaction category at each moment.

[0054] In some embodiments, the recording of the customer can also be obtained, and the customer's emotion is predicted through the customer's tone change, mood change, or recording content, etc. And combine the recording prediction result, the pose prediction result, and the expression prediction result (such as using the three as the input of operation S230 to obtain the emotion evaluation vector) to determine the final emotion of the customer.

[0055] In operation S240, assemble the emotion evaluation vectors at each moment in the time order of the corresponding moment among the M moments to obtain a time - series emotion vector sequence, where the time - series emotion vector sequence includes M emotion evaluation vectors.

[0056] Exemplarily, the time - series emotion vector sequence can be considered as a sequence containing M emotion evaluation vectors corresponding one - to - one to the M moments, and the order of the M emotion evaluation vectors in the sequence is consistent with the time order of the corresponding moments.

[0057] In operation S250, input the time - series emotion vector sequence into a pre - trained satisfaction recognition model to obtain a customer satisfaction recognition result, where the customer satisfaction recognition result is used to characterize the trend of the customer's emotion change within a predetermined time period.

[0058] For example, customer A131 enters the bank business hall to handle a loan. When entering the hall, the customer is in a sad mood. As the loan business is processed, the mood gradually changes from sad to neutral and the customer leaves the business hall. The mood change trend of customer A131 during the business process is a positive change, so it can be considered that the customer satisfaction is positive. Therefore, the customer satisfaction recognition result takes into account the dynamic mood changes of the customer within a predetermined time period, rather than determining that the customer satisfaction is low based on the customer's final neutral mood. Thus, the customer satisfaction score can be obtained by assigning values according to the mood change trend within the predetermined time period.

[0059] According to an embodiment of the present disclosure, the pose prediction result and the expression prediction result at each of the M moments are obtained, and the emotion evaluation vector at each moment is obtained. Then, the M emotion evaluation vectors are assembled in the time sequence of the corresponding moments in the M moments to obtain a time-sequence emotion vector sequence. The time-sequence emotion vector sequence is input into a pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result. Thus, by combining the pose and the expression, it is possible to avoid inaccurate prediction results obtained by only using one of them, and take into account the customer's mood change situation within a specific time period, and perform recognition through the time-sequence emotion vector sequence obtained in the time series dimension to obtain an accurate customer satisfaction result.

[0060] Figure 3 Schematically shows a flowchart of obtaining a customer portrait according to an embodiment of the present disclosure.

[0061] As Figure 3 shown, obtaining the customer images at M moments of the customer within a predetermined time period in operation S210 includes operations S310 to S330.

[0062] In operation S310, video data within a predetermined time period is obtained, where at least two image frames in the video data include the customer.

[0063] Referring Figure 1 , the video data can be a video file captured by one camera or a video file captured by multiple cameras. It is possible that a certain customer is not captured in some frames by each camera, then the frames captured by other cameras can be used. The purpose is to be able to determine the identity image and face image of the same customer at each moment for pose recognition and expression recognition.

[0064] In operation S320, customer re-identification is performed on N image frames in the video data, where N is an integer greater than or equal to 2 and N is greater than or equal to M.

[0065] In some embodiments, customer re-identification is implemented using person re-identification technology. Person Re-identification, also known as pedestrian re-identification, is a technology that uses computer vision technology to determine whether a specific pedestrian exists in an image or video sequence. In other words, pedestrian re-identification refers to identifying a target pedestrian in a video sequence from possible sources with non-overlapping camera fields of view.

[0066] In operation S330, based on the result of customer re-identification, customer images at M moments are obtained from N image frames.

[0067] Exemplarily, for example, if N is 100 and M is 50, in operation S330, 50 image frames capturing the customer can be determined and used as customer images, or customer images can be cropped from these image frames.

[0068] Figure 4 The architecture diagram of the pose recognition model according to an embodiment of the present disclosure is schematically shown.

[0069] According to an embodiment of the present disclosure, obtaining the pose prediction result of the customer at each moment based on the customer images at each moment of the M moments includes: inputting the customer image at each moment into the pose recognition model to obtain the pose prediction result at each moment, where the pose prediction result at each moment is used to characterize the first emotion of the customer at that moment, and the pose recognition model is constructed by a densely connected method.

[0070] Exemplarily, referring to Figure 4 , in the densely connected method, all layers in the pose recognition model are connected pairwise, so that each layer in the model receives the features of all the previous layers as input.

[0071] The pose recognition model receives the body image of the customer and gives the pose prediction result of the current frame. Considering that this task is a classification task and the target is relatively large, it is necessary to be able to take into account high-level semantics (such as the body contour of the customer) and low-level visual information (such as details of the customer's clothing style, clothing color, and accessories), so a densely connected classification network is constructed as the pose recognition model. Among them, the first emotion can be any one of categories such as angry, disgusted, afraid, happy, neutral, sad, or surprised.

[0072] In some cases, relying solely on body movements may not be able to fully and accurately evaluate the customer's current emotion, because human postures vary from person to person and are greatly affected by customs, habits, etc. When the user base, customer group, and the social strata and industries that the users span are large, the differences in the expression of emotions by body movements will be very large. The judgment result obtained solely based on pose indicators may be greatly distorted from the actual situation. Therefore, customer expression recognition indicators are introduced to comprehensively evaluate customer satisfaction through expressions and postures.

[0073] Figure 5 Schematically shows an architecture diagram of an expression recognition model according to an embodiment of the present disclosure.

[0074] According to an embodiment of the present disclosure, obtaining an expression prediction result of a customer at each of M moments based on the customer image at each moment includes: inputting the customer image at each moment into an expression recognition model to obtain an expression prediction result at each moment, where the expression prediction result at each moment is used to characterize the second emotion of the customer at that moment, and the first emotion and the second emotion are the same or different, and the expression recognition model is constructed by a hierarchical feature aggregation method.

[0075] Wherein, the second emotion can be any one of categories such as angry, disgusted, afraid, happy, neutral, sad or surprised. Since the first emotion is obtained based on gesture recognition and the second emotion is obtained based on expression recognition, there may be a situation where the two are different.

[0076] Refer to Figure 5 , a hierarchical feature aggregation network / Feature Pyramid Network can be constructed by using the hierarchical feature aggregation method as an expression recognition model, which has a top-down network structure and lateral connections, and can fuse low-level features with high resolution and high-level features with rich semantic information. Considering that the human face is local pixels in the customer image and more refined local modeling is required, a hierarchical network structure with decreasing feature mapping in Figure 5 is adopted to obtain features at different levels and then aggregate them to give an expression prediction result. In some embodiments, face localization can be performed during the process of training the expression recognition model. Adding face localization is to impose an additional mandatory constraint during the process of the model learning human face features and improve the reliability of the prediction result during learning.

[0077] Figure 6 Schematically shows a flowchart for reconstructing a customer image according to an embodiment of the present disclosure.

[0078] As Figure 6 shown, inputting the customer image at each moment into the expression recognition model in this embodiment includes inputting the customer image into the model after super-resolution reconstruction, specifically including operations S610 to S620.

[0079] In operation S610, perform super-resolution reconstruction on the customer image at each moment to obtain a super-resolution customer image at each moment.

[0080] The super-resolution reconstruction technology of an image refers to restoring a given low-resolution image into a corresponding high-resolution image through a specific algorithm. Super-resolution reconstruction of the customer image at each moment is the process of reconstructing a high-resolution face image from the given low-resolution face image.

[0081] In operation S620, the super-resolution customer image at each moment is input into the facial expression recognition model.

[0082] In an indoor environment, considering that the camera is generally located at the top of the room and is relatively far from the customer, there is a problem of imaging distortion, especially for local areas such as the human face that occupy a smaller proportion of pixels. In an outdoor environment, it is relatively open, and the camera may also be far from the customer. To obtain clearer facial expression information, first, a super-resolution network is used to perform super-resolution reconstruction on the input customer image (the overall body image or the face image, which can be obtained through the ROI (region of interest) area in the re-identification process) to obtain a higher-precision customer image, and then the facial expression recognition model is used to model the facial expression information. This overcomes or compensates for problems such as blurred imaging images, low quality, and insignificant regions of interest caused by the limitations of the camera or the acquisition environment itself.

[0083] Figure 7 Schematically shows the architecture diagram of the emotion assessment knowledge model according to an embodiment of the present disclosure. As Figure 7 shown, the emotion assessment model 700 receives the facial expression and pose prediction results as inputs, outputs the index of the emotion vector in the emotion knowledge base after constructing a transformation function by a long short-term memory network (LSTM, Long Short-Term Memory), and finally takes out the emotion vector from the emotion knowledge base according to the index as the emotion assessment result output.

[0084] Figure 8 Schematically shows the flowchart of obtaining the emotion assessment vector according to an embodiment of the present disclosure.

[0085] As Figure 8 shown, in operation S230, obtaining the emotion assessment vector at each moment according to the pose prediction result and the facial expression prediction result at each moment includes operations S810 to S830.

[0086] In operation S810, the pose prediction result and the facial expression prediction result at each moment are input into the emotion assessment knowledge model to obtain the output vector of the emotion assessment knowledge model, where the emotion assessment knowledge model is constructed according to the long short-term memory network.

[0087] In operation S820, index data is obtained according to the output vector.

[0088] Referring toFigure 7 , the emotion evaluation model 700 includes a variant model 710 and an emotion knowledge base 720. Among them, the controller 711 is constructed by a variant of LSTM based on the controller structure of the Differentiable neural computer. Among them, h is the hidden vector, c is the memory vector, v is the output obtained by transforming the current input, o is the output vector, x is the emotion evaluation vector, the subscript t represents the time step, and i represents the index of the emotion vector in the emotion knowledge base. In some embodiments, o t can be directly used as i, or o t can be further subjected to feature extraction to obtain i.

[0089] In operation S830, an emotion evaluation vector matching the index data is retrieved from the emotion knowledge base, where the emotion knowledge base includes at least one emotion evaluation vector to be matched.

[0090] Exemplarily, o t is transformed through a connection layer and a softmax layer (not shown) to obtain i, and then the emotion evaluation vector is obtained from the emotion knowledge base. For example, the emotion knowledge base includes the mapping relationship between multiple different i and multiple emotion evaluation vectors to be matched.

[0091] According to an embodiment of the present disclosure, in the execution stage of the emotion evaluation knowledge model, the direct result of the emotion evaluation is not directly given at each moment, but an emotion evaluation vector is output to describe a finer-grained emotion state. The emotion evaluation vectors to be matched in the emotion knowledge base can be predetermined prior knowledge, and retrieving the matching data from them can improve the reliability.

[0092] Figure 9 Schematically shows a flowchart of obtaining an emotion knowledge base according to an embodiment of the present disclosure.

[0093] As Figure 9 shown, before retrieving the emotion evaluation vector matching the index data from the emotion knowledge base, obtaining the emotion knowledge base is further included, and the obtaining of the emotion knowledge base in this embodiment includes operation S910 to operation S920.

[0094] In operation S910, the emotion text data set is input into the differentiable neural computer model for processing to obtain the external storage matrix of the differentiable neural computer model, where the emotion text data set includes at least one text for training the differentiable neural computer model.

[0095] Exemplarily, the text in the emotion text data set can include words that can express emotions, or can also be the predicted content obtained according to gesture recognition or facial expression recognition.

[0096] The differentiable neural computer model has one or more cells, and any cell is mainly composed of a controller and a memory. Among them, the controller can be an artificial neural network or other machine learning models. The memory can be understood as consisting of a read / write head, memory cells, and some cells for storing storage states (i.e., the external storage matrix).

[0097] In operation S920, an emotion knowledge base is obtained according to the external storage matrix.

[0098] The external storage matrix after the differentiable neural computer model converges on the emotion text dataset can be taken as the emotion knowledge base. For example, first, a part of the emotion text dataset is used to train the differentiable neural computer model. After the loss function of the differentiable neural computer model tends to converge, the remaining part of the emotion text dataset is input into the differentiable neural computer model, and the corresponding external storage matrix of this part is taken as the emotion knowledge base.

[0099] In some embodiments, the emotion knowledge base does not participate in the training of the emotion evaluation model, but only provides emotion evaluation results as an external knowledge base. It itself is obtained by the differentiable neural computer model converging on the emotion evaluation dataset, and artificial-coded emotion evaluation knowledge vectors can also be mixed in to enhance the accuracy of the evaluation results. Its role is that the emotion knowledge base can be maintained separately, making it more convenient to integrate expert knowledge and avoiding the unreliability caused by pure machine learning.

[0100] Figure 10 The architecture diagram of the satisfaction recognition model according to an embodiment of the present disclosure is schematically shown. Figure 11 The flowchart of obtaining the customer satisfaction recognition result according to an embodiment of the present disclosure is schematically shown.

[0101] As Figure 11 shown, in operation S250, the time-series emotion vector sequence is input into the pre-trained satisfaction recognition model, and obtaining the customer satisfaction recognition result includes operations S1110 to S1130.

[0102] The satisfaction recognition model includes a first neural network model, a second neural network model, and a classification model, and the classification model is a machine learning model. Referring to Figure 10 , x represents the emotion evaluation vector, and the subscript t represents the time step. GRU (Gated Recurrent Unit) is a gated recurrent unit. The first neural network model is a recurrent neural network composed of two GRUs, the second neural network model is an MLP (Multilayer Perceptron) model, and the classification model is an SVM (Support Vector Machine) model.

[0103] It should be noted thatFigure 10 The medium satisfaction recognition model is only exemplary. According to actual needs, other network structures or types can be adopted to implement the first neural network model, the second neural network model, and the classification model.

[0104] In operation S1110, the first neural network model processes the time-series emotion vector sequence to obtain an emotion change heat vector.

[0105] In operation S1120, the second neural network model processes the time-series emotion vector sequence to obtain an emotion evaluation intermediate vector.

[0106] In operation S1130, the emotion change heat vector and the emotion evaluation intermediate vector are input into the classification model to obtain a customer satisfaction recognition result.

[0107] Among them, the emotion change heat vectors on the M time series can represent the customer's emotion change trend, and the emotion evaluation intermediate vector is obtained by further performing feature fusion on multiple emotion evaluation vectors.

[0108] Exemplarily, as described in operations S1110 to S1130, first, the emotion evaluation vectors within a predetermined time period for a single customer (target) are taken out to form a continuous emotion evaluation sequence as the time-series emotion vector sequence. Then, the time-series emotion vector sequence passes through a recurrent neural network composed of two GRU units to obtain an emotion change heat vector. At the same time, the time-series emotion vector sequence passes through an MLP to obtain a summary intermediate result (for example, the MLP extracts features from the M emotion evaluation vectors to obtain an intermediate vector, which is the emotion evaluation intermediate vector). Finally, the emotion change heat vector and the summary intermediate result are connected and put into an SVM to obtain the final evaluation result.

[0109] According to an embodiment of the present disclosure, the state of the customer's emotion mutation within a predetermined time period, as well as the turning point or specific moment of the intermediate change, can be obtained from the emotion change heat vector, and the emotion evaluation intermediate vector is used as an overall summary of the customer's emotion change within the predetermined time period, so that a more accurate customer satisfaction recognition result can be obtained, and the emotion change trend can be better reflected.

[0110] Figure 12 The flowchart of a customer satisfaction recognition method according to another embodiment of the present disclosure is schematically shown.

[0111] As Figure 12 shown, the customer satisfaction recognition of this embodiment includes operations S1201 to S1216. Among them, the detector is implemented by using a complete detection network. Feature extraction, pose prediction, and expression prediction are constructed based on residual modules. On this basis, pose prediction is also constructed based on a dense connection method, and expression prediction is also constructed based on a hierarchical feature aggregation method.

[0112] In operation S1201 to operation S1209, the customer is tracked and re-identified / pedestrians are re-identified. Tracking and re-identification are used to detect and calibrate customers entering the venue, provide basic feature information for subsequent customer posture and modality recognition, and monitor and mark customers with fixed tags during the service process (the tags remain unchanged from the customer's appearance to the departure to lock on specific targets), so as to obtain dynamic changes in user satisfaction in the complete service process (such as characterization by changes in emotions), more accurately and comprehensively evaluate service quality, find out omissions and make up for deficiencies, and improve the quality and image of services and products. The details are as follows.

[0113] In operation S1201, customer images at M moments are obtained from the video data captured by the camera. For example, the images are collected from the camera of a bank business hall, discretized after periodic sampling at intervals of S, and then transmitted to the server for detection and evaluation of user satisfaction.

[0114] In operation S1202, a customer image is input to a detector.

[0115] In operation S1203, the input image is passed through a pedestrian detector to monitor and identify the customer, and one or more detection frames are obtained.

[0116] In operation S1204, one or more detection frames are used to capture a region of interest (RoI) where the customer is located on the original image. For example, the region of interest may be multiple local regions of a human body structure, such as three large regions (head, upper body, lower body) and four small regions of limbs.

[0117] In operation S1205 , one or more RoI(s) are fed into a feature extractor to extract features.

[0118] In operation S1206, the cosine similarity is calculated between the features extracted in operation S1205 and the cached historical features to obtain a re-identification metric (cosine metric) of the same client in the upper and lower frames.

[0119] In operation S1207, the re-identification metric is fed into a linear Kalman filter.

[0120] In operation S1208, one or more detection boxes are sent to a linear Kalman filter.

[0121] In operation S1209, based on the re-identification metric and one or more detection bounding boxes, linear Kalman filtering finally obtains the overall prediction bounding box and the re-identification identifier (i.e., the fixed label) of the customer in the current frame. For example, local features corresponding to each ROI region are processed, and then multiple local features are concatenated. Finally, a pedestrian re-identification feature that combines the global feature and multiple-scale local features is obtained, so as to determine the overall prediction bounding box and the re-identification identifier. (x’, y, r, h, i’) represents the output of re-identification tracking prediction, (x’, y) represents the center coordinates of the detected target, r represents the aspect ratio of the overall prediction bounding box, h represents the height of the overall prediction bounding box, and i’ represents the re-identification identifier.

[0122] In operations S1210 to S1214, since emotion assessment is a relatively complex task, it is difficult to effectively obtain accurate emotion assessment results by simply retrieving combinations of elements. Therefore, in the emotion assessment stage, the expression prediction result and the pose prediction result of the image tracking network are received, and instead of directly giving the direct result of emotion assessment, an emotion assessment vector is output to describe a finer-grained emotion state. Specifically as follows.

[0123] In operation S1210, the pose recognition model constructed using the Figure 4 shown structure is used to perform pose recognition on the ROI region screenshot to obtain the pose recognition result.

[0124] In operation S1211, the super-resolution network is used to perform super-resolution reconstruction on the ROI region screenshot.

[0125] In operation S1212, the expression recognition model constructed using the Figure 5 shown structure is used to perform expression recognition on the ROI region screenshot after super-resolution reconstruction to obtain the expression recognition result.

[0126] In operation S1213, the pose recognition result and the expression recognition result are input into the emotion assessment knowledge system to output the emotion assessment vector at each moment. The emotion assessment knowledge system is used to run the Figure 7 shown emotion assessment knowledge model.

[0127] In operations S1214 to S1216, the satisfaction degree of this customer is obtained through the time-series emotion vector sequence, and the final product or service satisfaction portrait is obtained through summarization and can be visually displayed on the front end. Specifically as follows.

[0128] In operation S1214, the emotion assessment vectors at M moments are cumulatively modeled. The process of cumulative modeling is to run the Figure 10 shown satisfaction recognition model to execute operations S1110 to S1130.

[0129] In operation S1215, obtain the satisfaction recognition result of the customer, and obtain the final product or service satisfaction evaluation result.

[0130] In operation S1216, perform front-end visualization on the satisfaction of each customer and the satisfaction of the product or service.

[0131] According to an embodiment of the present disclosure, visual application technologies such as re-identification and action capture are introduced. In combination with a monitoring system, the emotional changes of a user within a certain range (such as a business hall, a product or service, etc.) are captured and recognized from the actions and expressions of the crowd, and the satisfaction score of the user is obtained, and the final user satisfaction portrait is obtained by summarization.

[0132] Based on the above customer satisfaction recognition method, the present disclosure also provides a customer satisfaction recognition device. The following will be combined with Figure 13 Describe this device in detail.

[0133] Figure 13 Schematically shows a structural block diagram of a customer satisfaction recognition device according to an embodiment of the present disclosure.

[0134] As Figure 13 shown, the customer satisfaction recognition device 1300 of this embodiment includes an image acquisition module 1310, a posture and expression prediction module 1320, a first vector module 1330, a second vector module 1340, and a satisfaction recognition module 1350.

[0135] The image acquisition module 1310 can perform operation S210 to obtain customer images of the customer at M moments within a predetermined time period, where M is an integer greater than or equal to 2.

[0136] According to an embodiment of the present disclosure, the image acquisition module 1310 can also perform operations S310 to S330, which will not be elaborated here.

[0137] The posture and expression prediction module 1320 can perform operation S220 to obtain the posture prediction result and the expression prediction result of the customer at each moment according to the customer image at each of the M moments.

[0138] According to an embodiment of the present disclosure, the customer satisfaction recognition device 1300 may further include a super-resolution module, and this module can perform operations S610 to S620, which will not be elaborated here.

[0139] The first vector module 1330 can perform operation S230 to obtain an emotional evaluation vector at each moment according to the posture prediction result and the expression prediction result at each moment.

[0140] According to an embodiment of the present disclosure, the first vector module 1330 may also perform operations S810 to S830, and operations S910 to S920 are not described herein.

[0141] The second vector module 1340 may perform operation S240 to assemble the emotion assessment vectors at each moment in the time sequence corresponding to the moment among the M moments to obtain a time-sequential emotion vector sequence, where the time-sequential emotion vector sequence includes M emotion assessment vectors.

[0142] The satisfaction recognition module 1350 may perform operation S250 to input the time-sequential emotion vector sequence into a pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result, where the customer satisfaction recognition result is used to characterize the emotional change trend of the customer within a predetermined time period.

[0143] According to an embodiment of the present disclosure, the satisfaction recognition module 1350 may also perform operations S1110 to S1130, which are not described herein.

[0144] It should be noted that the implementation manners, the technical problems solved, the functions achieved, and the technical effects achieved by each module / unit / sub-unit, etc. in the device part embodiments are the same as or similar to those of the corresponding steps in the method part embodiments, and will not be described herein again.

[0145] According to an embodiment of the present disclosure, any multiple of the image acquisition module 1310, the pose and expression prediction module 1320, the first vector module 1330, the second vector module 1340, and the satisfaction recognition module 1350 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module.

[0146] According to an embodiment of the present disclosure, at least one of the image acquisition module 1310, the pose and expression prediction module 1320, the first vector module 1330, the second vector module 1340, and the satisfaction recognition module 1350 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any suitable combination of several of them. Alternatively, at least one of the image acquisition module 1310, the pose and expression prediction module 1320, the first vector module 1330, the second vector module 1340, and the satisfaction recognition module 1350 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.

[0147] Figure 14 Schematically shows a block diagram of an electronic device suitable for implementing the customer satisfaction recognition method according to an embodiment of the present disclosure.

[0148] As Figure 14 shown, the electronic device 1400 according to an embodiment of the present disclosure includes a processor 1401, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 1402 or a program loaded from a storage section 1408 into a random access memory (RAM) 1403. The processor 1401 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 1401 can also include on-board memory for caching purposes. The processor 1401 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0149] In the RAM 1403, various programs and data required for the operation of the electronic device 1400 are stored. The processor 1401, the ROM 1402, and the RAM 1403 are connected to each other through a bus 1404. The processor 1401 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 1402 and / or the RAM 1403. It should be noted that the program can also be stored in one or more memories other than the ROM 1402 and the RAM 1403. The processor 1401 can also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in one or more memories.

[0150] According to an embodiment of the present disclosure, the electronic device 1400 may further include an input / output (I / O) interface 1405, and the input / output (I / O) interface 1405 is also connected to the bus 1404. The electronic device 1400 may further include one or more of the following components connected to the I / O interface 1405: an input portion 1406 including a keyboard, a mouse, etc.; an output portion 1407 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1408 including a hard disk, etc.; and a communication portion 1409 including a network interface card such as a LAN card, a modem, etc. The communication portion 1409 performs communication processing via a network such as the Internet. The drive 1410 is also connected to the I / O interface 1405 as needed. A removable medium 1411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1410 as needed so that a computer program read therefrom can be installed into the storage portion 1408 as needed.

[0151] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the methods according to the embodiments of the present disclosure are implemented.

[0152] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, device, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 1402 and / or RAM 1403 and / or one or more memories other than ROM 1402 and RAM 1403.

[0153] Embodiments of the present disclosure also include a computer program product, which includes a computer program, and the computer program includes program codes for executing the methods shown in the flowcharts. When the computer program product runs in a computer system, the program codes are used to cause the computer system to implement the methods provided by the embodiments of the present disclosure.

[0154] When the computer program is executed by the processor 1401, the above functions defined in the system / apparatus of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, apparatus, module, unit, etc. can be implemented by computer program modules.

[0155] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program can also be transmitted and distributed in the form of signals on a network medium, and be downloaded and installed through the communication part 1409, and / or be installed from the removable medium 1411. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0156] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1409, and / or be installed from the removable medium 1411. When the computer program is executed by the processor 1401, the above functions defined in the system of the embodiments of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.

[0157] According to an embodiment of the present disclosure, the program code for executing the computer program provided in the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0158] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0159] Those skilled in the art will appreciate that the features recited in the various embodiments and / or claims of the present disclosure may be combined in various ways and / or combinations, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure may be combined in various ways and / or combinations. All such combinations and / or combinations fall within the scope of the present disclosure.

[0160] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.

Claims

1. A method for identifying customer satisfaction, comprising: Obtain customer images at M moments within a predetermined time period, where M is an integer greater than or equal to 2; Based on the customer images at each of the M moments, obtain the pose prediction result and expression prediction result of the customer at each of the moments; Based on the pose prediction result and expression prediction result at each of the moments, obtain the emotion evaluation vector at each of the moments, specifically including: inputting the pose prediction result and expression prediction result at each of the moments into an emotion evaluation knowledge model to obtain the output vector of the emotion evaluation knowledge model, where the emotion evaluation knowledge model is constructed based on a long short-term memory network; obtaining index data according to the output vector; retrieving the emotion evaluation vector matching the index data from an emotion knowledge base, where the emotion knowledge base includes at least one emotion evaluation vector to be matched; Assemble the emotion evaluation vectors at each of the moments in the time order of the corresponding moments in the M moments to obtain a time-series emotion vector sequence, where the time-series emotion vector sequence includes M emotion evaluation vectors; Input the time-series emotion vector sequence into a pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result, where the customer satisfaction recognition result is used to characterize the emotion change trend of the customer within the predetermined time period; Wherein, the satisfaction recognition model includes a first neural network model, a second neural network model, and a classification model, and the classification model is a machine learning model. Inputting the time-series emotion vector sequence into the pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result includes: Making the first neural network model process the time-series emotion vector sequence to obtain an emotion change heat vector; Making the second neural network model process the time-series emotion vector sequence to obtain an emotion evaluation intermediate vector; Inputting the emotion change heat vector and the emotion evaluation intermediate vector into the classification model to obtain the customer satisfaction recognition result.

2. The method according to claim 1, wherein Before retrieving the emotion evaluation vector matching the index data from the emotion knowledge base, it further includes obtaining the emotion knowledge base, specifically including: Inputting an emotion text data set into a differentiable neural computer model for processing to obtain the external storage matrix of the differentiable neural computer model, where the emotion text data set includes at least one text for training the differentiable neural computer model; Obtaining the emotion knowledge base according to the external storage matrix.

3. The method according to claim 1, wherein The obtaining of the customer images at M moments within a predetermined time period includes: Obtaining video data within the predetermined time period, where at least two image frames in the video data include the customer; Performing customer re-identification on N image frames in the video data, where N is an integer greater than or equal to 2 and N is greater than or equal to M; Based on the result of the customer re-identification, obtaining the customer images at the M moments from the N image frames.

4. The method according to claim 3, wherein The obtaining of the pose prediction result of the customer at each of the moments based on the customer images at each of the M moments includes: Input the customer image at each moment into a pose recognition model to obtain the pose prediction result at each moment, where the pose prediction result at each moment is used to characterize the first emotion of the customer at that moment, and the pose recognition model is constructed by a densely connected method.

5. The method according to claim 4, wherein Obtaining the expression prediction result of the customer at each of the M moments according to the customer images at each of the M moments includes: Input the customer image at each moment into an expression recognition model to obtain the expression prediction result at each moment, where the expression prediction result at each moment is used to characterize the second emotion of the customer at that moment, and the first emotion and the second emotion are the same or different, and the expression recognition model is constructed by a hierarchical feature aggregation method.

6. The method according to claim 5, wherein Inputting the customer image at each moment into the expression recognition model includes: Perform super-resolution reconstruction on the customer image at each moment to obtain a super-resolution customer image at each moment; Input the super-resolution customer image at each moment into the expression recognition model.

7. A device for identifying customer satisfaction, comprising: An image acquisition module, configured to acquire customer images at M moments within a predetermined time period, where M is an integer greater than or equal to 2; A pose and expression prediction module, configured to obtain the pose prediction result and the expression prediction result of the customer at each of the M moments according to the customer images at each of the M moments; A first vector module, configured to obtain an emotion evaluation vector at each moment according to the pose prediction result and the expression prediction result at each moment, specifically including: inputting the pose prediction result and the expression prediction result at each moment into an emotion evaluation knowledge model to obtain an output vector of the emotion evaluation knowledge model, where the emotion evaluation knowledge model is constructed according to a long short-term memory network; obtaining index data according to the output vector; extracting an emotion evaluation vector matching the index data from an emotion knowledge base, where the emotion knowledge base includes at least one emotion evaluation vector to be matched; A second vector module, configured to assemble the emotion evaluation vectors at each moment in the time order of the corresponding moments in the M moments to obtain a time-series emotion vector sequence, where the time-series emotion vector sequence includes M emotion evaluation vectors; A satisfaction recognition module, configured to input the time-series emotion vector sequence into a pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result, where the customer satisfaction recognition result is used to characterize the emotional change trend of the customer within the predetermined time period; Wherein, the satisfaction recognition model includes a first neural network model, a second neural network model, and a classification model, and the classification model is a machine learning model. Inputting the time-series emotion vector sequence into the pre-trained satisfaction recognition model to obtain a customer satisfaction recognition result includes: Making the first neural network model process the time-series emotion vector sequence to obtain an emotion change heat vector; Making the second neural network model process the time-series emotion vector sequence to obtain an emotion evaluation intermediate vector; Input the emotional change heat vector and the emotional assessment intermediate vector into the classification model to obtain the customer satisfaction recognition result.

8. An electronic device, comprising: One or more processors; A storage device for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to perform the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, which when executed by a processor implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Customer satisfaction recognition method and device based on micro-expression, terminal and medium

    CN110705349A

  • Business data pushing method and device, and server

    CN113221821A