Score determination method and device, storage medium and electronic device
By using a behavioral simulation deep learning model and multimodal feature fusion technology, standard videos are generated and actual videos are analyzed, which solves the problems of high consumption and bias caused by manual evaluation and achieves more accurate business compliance evaluation.
Patent Information
- Application Number
- CN202511840544.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-02-24
AI Technical Summary
Existing technologies that rely on manual assessment of staff's professional conduct suffer from high labor costs and subjective judgment bias.
A behavioral simulation deep learning model is used to generate standard videos. The actual videos are analyzed by combining convolutional neural networks and speech recognition models. Visual and linguistic features are fused through an attention mechanism to generate target scores.
It reduces the need for manual evaluation, lowers subjective judgment bias, and provides a more accurate and standardized assessment of business compliance.
Smart Images

Figure CN121564622A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital human interaction, and more specifically, to a method and apparatus for determining a score, a storage medium, and an electronic device. Background Technology
[0002] The assessment of teller business compliance in related technologies is generally conducted by arranging relevant personnel to conduct practical exercises or video retrospective analysis. These methods have problems such as high labor costs, subjective judgment bias, and lack of ability to generate standardized test scenarios.
[0003] There is currently no effective solution to the problem that relying on manual assessment to evaluate staff's professional conduct in related technologies leads to high labor costs and subjective judgment bias. Summary of the Invention
[0004] This application provides a method and apparatus for determining scores, a storage medium, and an electronic device to at least solve the problem in related technologies where the evaluation of staff's professional conduct is conducted through manual assessment, resulting in high labor costs and subjective judgment bias.
[0005] According to one embodiment of this application, a method for determining a score is provided, comprising: determining a standard video of a business operation performed in a preset business scenario based on a behavioral simulation deep learning model, and determining standard features of the business operation performed in the preset business scenario based on the standard video, wherein the standard features include at least one of the following: standard visual features and standard language features; obtaining an actual video of a target object performing the business operation in the preset business scenario, and determining actual features of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual features include at least one of the following: actual visual features and actual language features; and determining a target score for the target object performing the business operation in the preset business scenario based on the standard features and the actual features.
[0006] In an exemplary embodiment, determining a standard video for performing business operations under a preset business scenario based on a behavior simulation deep learning model includes: obtaining scene text corresponding to the preset business scenario; extracting a first feature vector corresponding to the scene text based on a semantic parsing model; and inputting the first feature vector into the behavior simulation deep learning model so that the behavior simulation deep learning model outputs the standard video.
[0007] In an exemplary embodiment, determining the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video includes: inputting the actual video into a convolutional neural network to extract content from the actual video, obtaining the actual behavior of the target object performing the business operation in the preset business scenario, and determining the actual behavior characteristics corresponding to the actual behavior; inputting the actual video into a speech recognition model to perform semantic transcription on the actual video, obtaining the actual speech of the target object performing the business operation in the preset business scenario and the target text corresponding to the actual speech, and determining the actual language characteristics corresponding to the actual speech; inputting the target text into an emotion computing model to output the actual emotion of the target object performing the business operation in the preset business scenario, and determining the actual emotion characteristics corresponding to the actual emotion, wherein the actual visual characteristics include: the actual behavior characteristics and the actual emotion characteristics.
[0008] In an exemplary embodiment, determining a target score for the target object performing the business operation in the preset business scenario based on the standard features and the actual features includes: comparing the standard visual features and the actual visual features to score the target object performing the business operation in the preset business scenario, obtaining a first score; comparing the standard language features and the actual language features to score the target object performing the business operation in the preset business scenario, obtaining a second score; and detecting the target score of the target object performing the business operation in the preset business scenario based on the first score and the second score.
[0009] In an exemplary embodiment, detecting the target score of the target object performing the business operation in the preset business scenario based on the first score and the second score includes: fusing the standard visual features and the standard language features according to an attention mechanism to generate a standard fused feature; fusing the actual visual features and the actual language features according to an attention mechanism to generate an actual fused feature; comparing the standard fused feature and the actual fused feature to score the target object performing the business operation in the preset business scenario, obtaining a third score; and constructing a scoring matrix corresponding to the target object performing the business operation in the preset business scenario based on the first score, the second score, and the third score, wherein the scoring matrix is used to indicate the target score of the target object performing the business operation in the preset business scenario.
[0010] In an exemplary embodiment, feature fusion of the actual visual features and the actual language features according to an attention mechanism to generate actual fused features includes: aligning the actual visual features and the actual language features in a time series using a time warping algorithm; and inputting the aligned actual visual features and actual language features into the attention mechanism so that the attention mechanism outputs the actual fused features.
[0011] In an exemplary embodiment, after determining the target score for the target object performing the business operation in the preset business scenario based on the standard features and the actual features, the method further includes: performing business violation detection on the target object performing the business operation based on the scoring matrix corresponding to the target object performing the business operation in the preset business scenario, to determine whether the target object has a business violation at each time point, wherein the scoring matrix is used to indicate the target score of the target object performing the business operation in the preset business scenario; if it is determined that the target object has a business violation at any time point, generating a violation report corresponding to the target object based on the business violation and the timestamp corresponding to the business violation.
[0012] According to another embodiment of this application, a scoring determination apparatus is also provided, comprising: a first determining module, configured to determine a standard video of a business operation performed in a preset business scenario based on a behavioral simulation deep learning model, and to determine standard features of the business operation performed in the preset business scenario based on the standard video, wherein the standard features include at least one of the following: standard visual features and standard language features; an acquisition module, configured to acquire an actual video of a target object performing the business operation in the preset business scenario, and to determine actual features of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual features include at least one of the following: actual visual features and actual language features; and a second determining module, configured to determine a target score for the target object performing the business operation in the preset business scenario based on the standard features and the actual features.
[0013] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the above-described scoring determination method when it is run.
[0014] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described scoring determination method through the computer program.
[0015] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program wherein the computer program is executed by a processor to perform the above-described scoring determination method.
[0016] In this embodiment, a standard video of a business operation performed under a preset business scenario is determined based on a behavioral simulation deep learning model. Based on the standard video, standard features, including at least standard visual features and standard language features, are determined for the business operation performed under the preset business scenario. An actual video of a target object performing a business operation under the preset business scenario is obtained, and the actual features of the target object performing the business operation under the preset business scenario are determined based on the actual video. The actual features include at least one of the following: actual visual features and actual language features. A target score for the target object performing the business operation under the preset business scenario is determined based on the standard features and the actual features. According to this application, the problem of high labor costs and subjective judgment bias in the related technology of evaluating staff's business compliance through manual assessment can be solved, thereby reducing the labor costs and subjective judgment bias in the evaluation of staff's business compliance. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a hardware structure block diagram of a computer terminal for a scoring determination method according to an embodiment of this application.
[0020] Figure 2 This is a flowchart of a scoring method according to an embodiment of this application;
[0021] Figure 3 This is an architecture diagram of a digital human simulation interactive teller business compliance assessment system based on multimodal behavior recognition, according to an optional embodiment of this application.
[0022] Figure 4 This is a structural block diagram of a scoring determination device according to an embodiment of this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0025] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing systems. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal for a scoring determination method according to an embodiment of this application. For example... Figure 1 As shown, a computer terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0026] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the processing strategy determination method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage systems, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0027] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0028] This embodiment provides a method for determining a score, applied to the aforementioned computer terminal. Figure 2 This is a flowchart of a scoring method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0029] Step S202: Determine the standard video for performing business operations under a preset business scenario based on the behavior simulation deep learning model, and determine the standard features for performing the business operations under the preset business scenario based on the standard video, wherein the standard features include at least one of the following: standard visual features and standard language features;
[0030] Standard visual features can include: standard behavior, standard posture, etc.
[0031] Step S204: Obtain the actual video of the target object performing the business operation in the preset business scenario, and determine the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual characteristics include at least one of the following: actual visual characteristics and actual language characteristics;
[0032] Step S206: Determine the target score for the target object to perform the business operation in the preset business scenario based on the standard features and the actual features.
[0033] Through the above steps, a standard video of performing business operations under a preset business scenario is determined based on the behavioral simulation deep learning model. Based on the standard video, standard features, including at least standard visual features and standard linguistic features, are determined for performing business operations under the preset business scenario. The actual video of the target object performing business operations under the preset business scenario is obtained, and the actual features of the target object performing business operations under the preset business scenario are determined based on the actual video. The actual features include at least one of the following: actual visual features and actual linguistic features. Based on the standard features and actual features, a target score for the target object performing business operations under the preset business scenario is determined. According to this application, the problem of high labor costs and subjective judgment bias caused by manually evaluating staff's business compliance in related technologies can be solved, thereby reducing the labor costs and subjective judgment bias in evaluating staff's business compliance.
[0034] Optionally, step S202 above, which determines the standard video for performing business operations under a preset business scenario based on the behavior simulation deep learning model, includes: obtaining the scene text corresponding to the preset business scenario; extracting the first feature vector corresponding to the scene text based on the semantic parsing model; and inputting the first feature vector into the behavior simulation deep learning model so that the behavior simulation deep learning model outputs the standard video.
[0035] Understandably, the technical solution for generating standard videos that conform to preset business scenarios includes: selecting or receiving a specific scenario description from a preset list of business scenarios, such as the service process of "a customer coming to apply for a credit card." This scenario description includes the service context, expected role behavior, and possible dialogue scripts, i.e., scenario text.
[0036] The scene text is input into the semantic parsing model. The semantic parsing model encodes representations from transformers (BERT) models for pre-trained transformers. BERT models understand the deep semantics of natural language, converting scene text into a series of numerical codes, known as first feature vectors. These first feature vectors capture key information from the scene text, such as character interactions, specific business process steps, and expected behavioral patterns. For example, in the scenario of "a customer applying for a credit card," the first feature vector might encode semantic information about specific actions and dialogues, such as "the customer presents their ID card" and "the teller verifies the customer's identity."
[0037] The first feature vector is input into the behavioral simulation deep learning model. This model integrates advanced video generation technologies, such as Sora (a deep learning model for generating simulations of human behavior) or other similar text-based video models. Based on the received first feature vector, the model generates a standard video demonstrating how a teller should act and communicate appropriately in a scenario where a customer applies for a credit card. For example, the standard video shows the teller greeting the customer with a smile, handing over documents with both hands, clearly explaining the application steps, and politely inquiring about the customer's information—a series of actions conforming to banking service standards and best practices.
[0038] Optionally, step S204 above, which determines the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video, includes: inputting the actual video into a convolutional neural network to extract content from the actual video, obtaining the actual behavior of the target object performing the business operation in the preset business scenario, and determining the actual behavior characteristics corresponding to the actual behavior; inputting the actual video into a speech recognition model to perform semantic transcription on the actual video, obtaining the actual speech of the target object performing the business operation in the preset business scenario and the target text corresponding to the actual speech, and determining the actual language characteristics corresponding to the actual speech; inputting the target text into an emotion computing model to output the actual emotion of the target object performing the business operation in the preset business scenario, and determining the actual emotion characteristics corresponding to the actual emotion, wherein the actual visual characteristics include: the actual behavior characteristics and the actual emotion characteristics.
[0039] Understandably, multiple deep learning models can be used to perform multimodal analysis on actual videos, including behavior recognition, speech recognition, and sentiment analysis, to gain a comprehensive understanding of the teller's (i.e., the target audience) actual performance in specific business scenarios (e.g., preset business scenarios). Specifically:
[0040] Behavior recognition: Actual video is input into a convolutional neural network (a deep learning model used for image and video content recognition). The convolutional neural network extracts the action features of the target object by performing layer-by-layer convolution and pooling operations on video frames, such as arm movements, head rotation, and eye gaze direction. For example, in the scenario of a customer applying for a credit card, the convolutional neural network might recognize actions such as the teller smiling at the customer, handing over documents with both hands, and nodding to confirm information. Once these actions are recognized, they are converted into digital feature vectors, i.e., actual behavioral features.
[0041] Speech recognition and semantic transcription: The audio stream from the actual video is input into a speech recognition model, such as Whisper-large-v3 (an automatic speech recognition model). Whisper-large-v3 converts the audio signal into text, thereby capturing the script spoken by the teller during the transaction. For example, the transcription result might include standard customer service phrases such as "Welcome, what service do you need?" and "Please show your ID." The obtained text is further input into a semantic parsing model, such as the BERT model, to understand the content and context of the script, determine the actual language features, including the compliance, professionalism, and fluency of the dialogue.
[0042] Affective computing: The target text (i.e., the speech-to-text result of the teller) is fed into an affective computing model, which can analyze the emotional tendency behind the text, such as positive, negative, or neutral. For example, even if the teller uses standard phrases, if the speech-to-text shows that their tone is impatient or indifferent, the affective computing model can capture this and reflect the teller's actual emotional characteristics.
[0043] Optionally, step S206 above, determining the target score for the target object performing the business operation in the preset business scenario based on the standard features and the actual features, includes: comparing the standard visual features and the actual visual features to score the target object performing the business operation in the preset business scenario, obtaining a first score; comparing the standard language features and the actual language features to score the target object performing the business operation in the preset business scenario, obtaining a second score; performing feature fusion on the standard visual features and the standard language features according to an attention mechanism to generate a standard fused feature; performing feature fusion on the actual visual features and the actual language features according to an attention mechanism to generate an actual fused feature; comparing the standard fused feature and the actual fused feature to score the target object performing the business operation in the preset business scenario, obtaining a third score; and constructing a scoring matrix corresponding to the target object performing the business operation in the preset business scenario based on the first score, the second score, and the third score, wherein the scoring matrix is used to indicate the target score for the target object performing the business operation in the preset business scenario.
[0044] Understandably, the technical solution for scoring the business operations performed by the target object under a preset business scenario to obtain the target score is as follows:
[0045] Visual Feature Comparison and Scoring (First Score): This compares standard visual features (features representing normative behavior extracted from digital human videos (i.e., standard videos) with actual visual features (behavioral features extracted from videos recorded of tellers in actual work (i.e., actual videos)). For example, in a "credit card application" scenario, standard visual features might include actions such as the teller greeting customers with a smile and handing over documents with both hands, while actual visual features reflect how the teller behaves in a real-world setting. If the teller fails to meet the standard of greeting customers with a smile in reality, this action will receive a lower score, resulting in a lower first score.
[0046] Language Feature Comparison and Scoring (Second Score): This compares standard language features (representing the transcribed text features of standard language and emotion) with actual language features. For example, standard language features require tellers to use polite and clear language when processing credit card applications, while actual language features are based on the conversation between the teller and the customer. If the teller uses inappropriate words or speaks impatiently, the second score will be affected, reflecting deficiencies in the actual language.
[0047] Feature Fusion and Comprehensive Scoring (Third Score): Standard visual and linguistic features are fused using an attention mechanism to generate a standard fused feature, comprehensively capturing behavioral and verbal patterns in standardized services. Similarly, actual visual and linguistic features are also fused into actual fused features. This fusion process considers the interaction between visual and linguistic information, such as eye contact during speech. Subsequently, the standard fused feature and the actual fused feature are compared to generate a third score.
[0048] Constructing a scoring matrix: Integrate the three scoring methods (first score, second score, and third score) to form a scoring matrix. The scoring matrix includes not only single-modal scores (visual and verbal) but also a comprehensive score after feature fusion. For example, in a "credit card application" scenario, the scoring matrix might show a teller's behavior score of 80, their sales pitch score of 85, and their overall behavior and sales pitch fusion score of 82. This matrix provides tellers with detailed feedback, indicating what they did well and where improvement is needed.
[0049] The process of fusing the actual visual features and the actual language features according to an attention mechanism to generate actual fused features includes: aligning the actual visual features and the actual language features in a time series using a time warping algorithm; and inputting the aligned actual visual features and actual language features into the attention mechanism so that the attention mechanism outputs the actual fused features.
[0050] Understandably, in order to accurately match actions in actual videos with concurrent audio information in terms of time and context, and thus more accurately assess the teller's overall performance when performing specific business operations, actual fusion features can be determined, specifically:
[0051] Time warping algorithms (e.g., Dynamic Time Warping (DTW)) align features: In multimodal behavior recognition, actions in the video stream and speech information in the audio stream do not always occur synchronously; there may be delays or discrepancies. DTW is used to adjust and align these two sets of time-series data, ensuring that comparisons of visual and linguistic features at the same point in time are reasonable. For example, in a "credit card application" scenario, a teller might first perform the action of handing over the document, and then say, "Please sign this document," a few seconds later. Without alignment, these two actions, which should be related, might be mistakenly treated as separate events, leading to inaccurate scoring.
[0052] Actual fused features are generated through an attention mechanism: aligned visual and linguistic features are input into the attention mechanism model. The model analyzes visual features (such as the teller's eye contact and nodding) and linguistic features (such as words used and tone of voice) at each moment, determining which information is most critical in the current context. For example, when a customer inquires about details, if the teller focuses their gaze on the customer and uses clear and professional vocabulary, the attention mechanism will give this moment higher weight, as this behavior demonstrates good communication skills and service awareness. Conversely, if the teller looks around or speaks vaguely while answering questions, these moments will be given lower weight, reflecting inadequate service attitude and professionalism.
[0053] Optionally, after determining the target score of the target object performing the business operation in the preset business scenario based on the standard features and the actual features in step S206 above, the method further includes: performing business violation detection on the target object performing the business operation based on the scoring matrix corresponding to the target object performing the business operation in the preset business scenario, so as to determine whether the target object has a business violation at each time point, wherein the scoring matrix is used to indicate the target score of the target object performing the business operation in the preset business scenario; if it is determined that the target object has a business violation at any time point, generating a violation report corresponding to the target object based on the business violation and the timestamp corresponding to the business violation.
[0054] Understandably, after determining the target score, it's possible to identify the corresponding business violations of the target object and generate a violation report. Specifically:
[0055] Business violation detection: Traverse the scoring matrix and analyze the scores at each time point to identify any possible business violations. For example, in the "credit card application" business scenario, if at a certain time point the teller's behavior score is low (e.g., failure to perform the standard two-handed document delivery action), and the sales pitch score is also low (using inappropriate or non-standard language), and the combined score is far below the average or preset passing score, then this time point can be marked as a business violation.
[0056] Generating Violation Reports: Once a violation is detected, a violation report can be generated based on the nature of the violation and its corresponding timestamp. The violation report details the time of the violation, its specific manifestation (e.g., failure to deliver documents according to specifications, use of non-standard language, etc.), the severity of the violation (based on a scoring matrix score), and possible improvement suggestions. For example, if a teller did not use the standard "two-handed delivery" action at 10:15:00, this can be clearly stated in the violation report, and it can be recommended to strengthen training and practice of this action.
[0057] To better understand the process of determining the above-mentioned score, the implementation flow of the above-mentioned score determination method will be described below in conjunction with optional embodiments, but this is not intended to limit the technical solution of the embodiments of this application.
[0058] This application's optional embodiments define a digital human simulation interactive teller business compliance evaluation system based on multimodal behavior recognition. Figure 3 This is an architecture diagram of a digital human simulation interactive teller business compliance assessment system based on multimodal behavior recognition, according to an optional embodiment of this application. Figure 3 As shown:
[0059] The digital human simulation interactive teller business compliance assessment system based on multimodal behavior recognition includes: a digital human generation module, an artificial intelligence (AI) processing module, and a multimodal fusion module. Among them:
[0060] 1) Input layer:
[0061] The input layer is divided into a text parsing module (which can be part of the digital human generation module) and a multimodal acquisition module (which can be part of the AI processing module).
[0062] The text parsing module receives standardized business scenario descriptions input by business personnel through a visual interface. After semantic parsing using the BERT model, it outputs JSON format (a lightweight data exchange format) feature vectors and stores them in the Kafka (an open-source stream processing platform) message queue (i.e., message queue 1).
[0063] The multimodal acquisition module collects data on tellers' facial expressions, gestures, postures, and service scripts through cameras and microphones. This data is transmitted in real time to a streaming media server and then transcoded and encapsulated into MP4 format (a standard digital multimedia container format) using FFmpeg (Fast Universal Multimedia Framework) for raw data storage (which can be stored on a cloud server). The metadata is also stored in a Kafka message queue (i.e., message queue 2, where metadata is the actual video).
[0064] 2) AI processing layer:
[0065] The AI processing layer is divided into a digital human generation module and a multimodal behavior recognition module (the multimodal behavior recognition module can be a part of the AI processing module).
[0066] The digital human generation module obtains (standard) business scenario description text from Kafka and generates digital human videos (i.e. standard videos) that conform to the business scenario through integrated Sora (a deep learning model for generating simulations of human behavior) or Korlinga models.
[0067] Digital human videos can be used to train tellers so that they can establish standard behaviors and standard language.
[0068] The video recognition module in the multimodal behavior recognition module inputs standard video and actual video into the I3D model (a three-dimensional convolutional neural network for video understanding) for content extraction and analysis, and identifies whether they conform to standard actions such as handing over with both hands and document verification gestures (i.e., obtaining the first score); the audio recognition module uses the Whisper-large-v3 model (an automatic speech recognition model) in streaming processing mode for transcription, and the transcribed text is semantically analyzed by the sentiment computing model and the BERT model to identify whether the employee's emotions and business language are appropriate (i.e., obtaining the second score).
[0069] The multimodal fusion module uses a time-series alignment module based on the DTW algorithm to align visual and speech features, and finally uses Cross-Attention (an attention mechanism) to fuse the two features.
[0070] 3) Output layer.
[0071] The results obtained from the multimodal behavior recognition module are used to construct a scoring matrix, and the output includes a timestamp analysis of violation points.
[0072] In summary, the optional embodiments of this application introduce a multimodal feature fusion mechanism to combine visual and audio modalities, achieving a more comprehensive understanding of information and better recognition results. The optional embodiments of this application also identify employee behavior based on a dynamic and static hierarchical approach, ensuring accuracy while reducing computational consumption. Specifically, the optional embodiments of this application combine visual and audio modalities through a multimodal feature fusion mechanism, achieving visual-speech feature alignment through DTW, and cross-modal feature fusion through Cross-Attention, resulting in higher recognition accuracy. Furthermore, the optional embodiments of this application rely on the Sora large model to construct digital human videos for business scenarios. Through multi-dimensional collaboration of I3D action recognition, Whisper streaming transcription, and BERT semantic analysis, the system achieves comprehensive intelligent evaluation from standard actions and compliant language to service emotions, significantly saving labor costs.
[0073] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0074] Figure 4 This is a structural block diagram of a scoring determination device according to an embodiment of this application; as shown... Figure 4 As shown, the device includes:
[0075] The first determining module 42 is used to determine the standard video of performing business operations in a preset business scenario based on the behavior simulation deep learning model, and to determine the standard features of performing the business operations in the preset business scenario based on the standard video, wherein the standard features include at least one of the following: standard visual features and standard language features.
[0076] The acquisition module 44 is used to acquire the actual video of the target object performing the business operation in the preset business scenario, and determine the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual characteristics include at least one of the following: actual visual characteristics and actual language characteristics;
[0077] The second determining module 46 is used to determine the target score of the target object performing the business operation in the preset business scenario based on the standard features and the actual features.
[0078] Using the aforementioned device, a standard video of a business operation performed under a preset business scenario is determined based on a behavioral simulation deep learning model. Based on the standard video, standard features, including at least standard visual features and standard linguistic features, are determined for the business operation performed under the preset business scenario. An actual video of a target object performing a business operation under the preset business scenario is obtained, and based on the actual video, the actual features of the target object performing the business operation under the preset business scenario are determined. These actual features include at least one of the following: actual visual features and actual linguistic features. A target score for the target object performing the business operation under the preset business scenario is determined based on the standard features and the actual features. According to this application, the problem of high labor costs and subjective judgment bias in the related technology of evaluating staff's business compliance through manual assessment can be solved, thereby reducing the labor costs and subjective judgment bias in the evaluation of staff's business compliance.
[0079] In an exemplary embodiment, the first determining module 42 is further configured to obtain scene text corresponding to the preset business scenario; extract a first feature vector corresponding to the scene text according to a semantic parsing model; and input the first feature vector into the behavior simulation deep learning model so that the behavior simulation deep learning model outputs the standard video.
[0080] In an exemplary embodiment, the acquisition module 44 is further configured to input the actual video into a convolutional neural network, so that the convolutional neural network extracts content from the actual video to obtain the actual behavior of the target object performing the business operation in the preset business scenario, and determines the actual behavior features corresponding to the actual behavior; input the actual video into a speech recognition model, so that the speech recognition model performs semantic transcription on the actual video to obtain the actual speech of the target object performing the business operation in the preset business scenario and the target text corresponding to the actual speech, and determines the actual language features corresponding to the actual speech; input the target text into an emotion computing model, so that the emotion computing model outputs the actual emotion of the target object performing the business operation in the preset business scenario, and determines the actual emotion features corresponding to the actual emotion, wherein the actual visual features include: the actual behavior features and the actual emotion features.
[0081] In an exemplary embodiment, the second determining module 46 is further configured to compare the standard visual features and the actual visual features to score the target object performing the business operation in the preset business scenario, and obtain a first score; compare the standard language features and the actual language features to score the target object performing the business operation in the preset business scenario, and obtain a second score; and detect the target score of the target object performing the business operation in the preset business scenario based on the first score and the second score.
[0082] In an exemplary embodiment, the second determining module 46 is further configured to perform feature fusion on the standard visual features and the standard language features according to an attention mechanism to generate standard fused features; perform feature fusion on the actual visual features and the actual language features according to an attention mechanism to generate actual fused features; compare the standard fused features and the actual fused features to score the target object performing the business operation in the preset business scenario to obtain a third score; and construct a scoring matrix corresponding to the target object performing the business operation in the preset business scenario based on the first score, the second score, and the third score, wherein the scoring matrix is used to indicate the target score of the target object performing the business operation in the preset business scenario.
[0083] In an exemplary embodiment, the second determining module 46 is further configured to perform time-series feature alignment of the actual visual features and the actual language features using a time warping algorithm; and input the aligned actual visual features and actual language features into the attention mechanism so that the attention mechanism outputs the actual fused features.
[0084] In an exemplary embodiment, the second determining module 46 is further configured to perform business violation detection on the target object's execution of the business operation according to the scoring matrix corresponding to the target object's execution of the business operation in the preset business scenario, so as to determine whether the target object has a business violation at each time point, wherein the scoring matrix is used to indicate the target score of the target object's execution of the business operation in the preset business scenario; if it is determined that the target object has a business violation at any time point, a violation report corresponding to the target object is generated according to the business violation and the timestamp corresponding to the business violation.
[0085] Embodiments of this application also provide a storage medium including a stored program, wherein the program executes any of the methods described above when it is run.
[0086] Optionally, in this embodiment, the storage medium may be configured to store program code for performing the following steps:
[0087] S1, determine the standard video for performing business operations in a preset business scenario based on the behavior simulation deep learning model, and determine the standard features for performing the business operations in the preset business scenario based on the standard video, wherein the standard features include at least one of the following: standard visual features and standard language features;
[0088] S2, acquire the actual video of the target object performing the business operation in the preset business scenario, and determine the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual characteristics include at least one of the following: actual visual characteristics, actual language characteristics;
[0089] S3, determine the target score for the target object to perform the business operation in the preset business scenario based on the standard features and the actual features.
[0090] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0091] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0092] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0093] S1, determine the standard video for performing business operations in a preset business scenario based on the behavior simulation deep learning model, and determine the standard features for performing the business operations in the preset business scenario based on the standard video, wherein the standard features include at least one of the following: standard visual features and standard language features;
[0094] S2, acquire the actual video of the target object performing the business operation in the preset business scenario, and determine the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual characteristics include at least one of the following: actual visual characteristics, actual language characteristics;
[0095] S3, determine the target score for the target object to perform the business operation in the preset business scenario based on the standard features and the actual features.
[0096] Embodiments of this application also provide a computer program product, including a computer program, which is executed by a processor using the steps described above:
[0097] S1, determine the standard video for performing business operations in a preset business scenario based on the behavior simulation deep learning model, and determine the standard features for performing the business operations in the preset business scenario based on the standard video, wherein the standard features include at least one of the following: standard visual features and standard language features;
[0098] S2, acquire the actual video of the target object performing the business operation in the preset business scenario, and determine the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual characteristics include at least one of the following: actual visual characteristics, actual language characteristics;
[0099] S3, determine the target score for the target object to perform the business operation in the preset business scenario based on the standard features and the actual features.
[0100] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0101] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0102] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0103] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for determining a score, characterized in that, include: Based on the behavior simulation deep learning model, standard videos for performing business operations in a preset business scenario are determined, and standard features for performing the business operations in the preset business scenario are determined based on the standard videos. The standard features include at least one of the following: standard visual features and standard language features. Obtain actual video of the target object performing the business operation in the preset business scenario, and determine the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual characteristics include at least one of the following: actual visual characteristics and actual language characteristics; The target score for the target object to perform the business operation in the preset business scenario is determined based on the standard features and the actual features.
2. The method according to claim 1, characterized in that, Based on a behavioral simulation deep learning model, standard videos for performing business operations in preset business scenarios are determined, including: Obtain the scene text corresponding to the preset business scenario; The first feature vector corresponding to the scene text is extracted based on the semantic parsing model; The first feature vector is input into the behavior simulation deep learning model so that the behavior simulation deep learning model outputs the standard video.
3. The method according to claim 1, characterized in that, Based on the actual video, the actual characteristics of the target object performing the business operation in the preset business scenario are determined, including: The actual video is input into a convolutional neural network so that the convolutional neural network can extract content from the actual video to obtain the actual behavior of the target object performing the business operation in the preset business scenario, and determine the actual behavior features corresponding to the actual behavior. The actual video is input into the speech recognition model so that the speech recognition model performs semantic transcription on the actual video to obtain the actual speech of the target object performing the business operation in the preset business scenario and the target text corresponding to the actual speech, and to determine the actual language features corresponding to the actual speech; The target text is input into the sentiment computing model so that the sentiment computing model outputs the actual emotion of the target object when performing the business operation in the preset business scenario, and determines the actual emotion feature corresponding to the actual emotion, wherein the actual visual feature includes: the actual behavioral feature and the actual emotion feature.
4. The method according to claim 1, characterized in that, Determine the target score for the target object to perform the business operation in the preset business scenario based on the standard features and the actual features, including: The standard visual features and the actual visual features are compared to score the target object performing the business operation in the preset business scenario, and a first score is obtained. The standard language features and the actual language features are compared to score the target object's execution of the business operation in the preset business scenario, and a second score is obtained. The target score is determined based on the first score and the second score when the target object performs the business operation in the preset business scenario.
5. The method according to claim 4, characterized in that, The target score is determined based on the first score and the second score, indicating that the target object performs the business operation in the preset business scenario. This includes: The standard visual features and the standard language features are fused according to the attention mechanism to generate standard fused features; The actual visual features and the actual language features are fused according to the attention mechanism to generate actual fused features; The standard fusion features and the actual fusion features are compared to score the target object performing the business operation in the preset business scenario, and a third score is obtained. A scoring matrix is constructed based on the first score, the second score, and the third score to indicate the target score of the target object performing the business operation in the preset business scenario.
6. The method according to claim 5, characterized in that, The actual visual features and the actual linguistic features are fused according to an attention mechanism to generate actual fused features, including: The actual visual features and the actual language features are aligned in time series using a time warping algorithm. The aligned actual visual features and actual linguistic features are input into the attention mechanism so that the attention mechanism outputs the actual fused features.
7. The method according to claim 1, characterized in that, After determining the target score for the target object to perform the business operation in the preset business scenario based on the standard features and the actual features, the method further includes: Based on the scoring matrix corresponding to the target object performing the business operation in the preset business scenario, the business violation point detection is performed on the target object performing the business operation to determine whether the target object has any business violation behavior at each time point. The scoring matrix is used to indicate the target score of the target object performing the business operation in the preset business scenario. If it is determined that the target object has engaged in the business violation at any point in time, a violation report corresponding to the target object is generated based on the business violation and the timestamp corresponding to the business violation.
8. A scoring determination device, characterized in that, include: The first determining module is used to determine the standard video of performing business operations in a preset business scenario based on the behavior simulation deep learning model, and to determine the standard features of performing the business operations in the preset business scenario based on the standard video, wherein the standard features include at least one of the following: standard visual features and standard language features. The acquisition module is used to acquire actual video of the target object performing the business operation in the preset business scenario, and determine the actual characteristics of the target object performing the business operation in the preset business scenario based on the actual video, wherein the actual characteristics include at least one of the following: actual visual characteristics and actual language characteristics; The second determining module is used to determine the target score of the target object performing the business operation in the preset business scenario based on the standard features and the actual features.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method described in any one of claims 1 to 7.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 7 through the computer program.
Citation Information
Cited By
Practical operation data scoring method and system
CN121859262A
A method and system for scoring practical data
CN121859262B