User intention analysis method, device, computer-readable storage medium, and processor

By integrating robots into the video interview to intelligently analyze multimedia streaming data, the problem of insufficient consideration of multiple factors due to the short interview time is solved, and efficient application approval and improved customer experience are achieved.

CN114638235BActive Publication Date: 2025-09-26DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210228040.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-08
Publication Date
2025-09-26
Estimated Expiration
2042-03-08

AI Technical Summary

Technical Problem

In the service industry, due to the short time of video interviews, customer service cannot fully consider multiple factors, resulting in misjudgment of customer applications. Existing technology cannot achieve real-time analysis, resulting in low processing efficiency.

Method used

By connecting to the robot to intelligently analyze multimedia streaming data during the face-to-face interview process, it can assist customer service in judging the customer's intentions in real time, generate analysis results and send them to the customer service terminal to assist in qualification review.

Benefits of technology

It improves the efficiency of application approval, reduces the error rate, and enhances customer experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114638235B_ABST
    Figure CN114638235B_ABST
Patent Text Reader

Abstract

The present invention discloses a user intention analysis method, device, computer-readable storage medium, and processor. The method comprises: obtaining multimedia stream data, wherein the multimedia stream data is used to represent multimedia data generated during a first user's qualification review of a second user; performing an intention analysis on the second user based on the multimedia stream data to generate a target analysis result, wherein the target analysis result is used to indicate whether the second user's intention is abnormal; and sending the target analysis result to a target terminal corresponding to the first user, wherein the target terminal is used by the first user to perform a qualification review of the second user based on the target analysis result. The present invention solves the technical problem that due to the short face-to-face signing time, the customer service staff cannot take multiple factors into consideration, thereby leading to misjudgment of the customer's application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of real-time audio and video, and in particular to a user intention analysis method, device, computer-readable storage medium and processor. Background Art

[0002] In the service industry, when a customer submits an application and is judged to be at risk, such as for fraud or phone scams, they are required to undergo a video interview. This involves a customer service representative conducting a Q&A session with the customer via video to further assess their risk. However, due to the limited number of customer service representatives within a service organization and the relatively short call times, the customer service representative may not be able to fully assess all factors during the video interview, leading to misjudgments of the customer's application. Furthermore, in the case of important applications, subsequent quality inspections require a significant amount of time for comparison and review, resulting in low processing efficiency.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] Embodiments of the present invention provide a user intent analysis method, device, computer-readable storage medium, and processor to at least solve the technical problem that the customer service staff is unable to consider multiple factors due to the short face-to-face interview time, thereby leading to misjudgment of the customer's application.

[0005] According to one aspect of an embodiment of the present invention, a user intention analysis method is provided, including: obtaining multimedia stream data, wherein the multimedia stream data is used to represent multimedia data generated during a qualification review of a second user by a first user; performing intention analysis on the second user based on the multimedia stream data to generate a target analysis result, wherein the target analysis result is used to indicate whether there is any abnormality in the intention of the second user; and sending the target analysis result to a target terminal corresponding to the first user, wherein the target terminal is used by the first user to perform qualification review of the second user based on the target analysis result.

[0006] Optionally, before obtaining the multimedia stream data, the method also includes: creating a virtual room in response to the application instruction of the second user, wherein the virtual room is a virtual environment for conducting qualification review of the second user; sending a first notification instruction to the first user in response to the second user entering the virtual room, wherein the first notification instruction is used to notify the first user to enter the virtual room; and obtaining multimedia stream data of the second user and the first user in response to the first user entering the virtual room.

[0007] Optionally, performing an intention analysis on the second user based on the multimedia streaming data to generate a target analysis result includes: sending a second notification instruction to the target robot in response to the first user entering the virtual room, wherein the second notification instruction is used to notify the target robot to enter the virtual room; controlling the target robot to perform an intention analysis on the second user based on the multimedia streaming data to generate a target analysis result.

[0008] Optionally, the multimedia stream data includes: video stream data, the target analysis results include: facial analysis results and / or motion analysis results, and the intention analysis of the second user is performed based on the multimedia stream data to generate the target analysis results, including: performing facial analysis on the second user based on the video stream data to generate facial analysis results; and / or performing motion analysis on the second user based on the video stream data to generate motion analysis results.

[0009] Optionally, the multimedia stream data also includes: audio stream data, and the target analysis result also includes: voice analysis results. The intention analysis of the second user is performed based on the multimedia stream data to generate the target analysis result, including: performing voice analysis on the second user based on the audio stream data to generate the voice analysis result.

[0010] Optionally, sending the target analysis result to a target terminal corresponding to the first user includes: controlling the target terminal to display the target analysis result based on a preset display mode.

[0011] Optionally, the method also includes: in response to the first user's exit instruction, sending a third notification instruction to the second user, and sending a fourth notification instruction to the target robot, wherein the third notification instruction is used to notify the second user that the qualification review is completed, and the fourth notification instruction is used to notify the target robot to exit the virtual room.

[0012] According to another aspect of an embodiment of the present invention, a user intention analysis device is also provided, including: an acquisition module for acquiring multimedia stream data of a second user and a first user, wherein the multimedia stream data is used to represent multimedia data generated during the qualification review of the second user by the first user; an analysis module for performing intention analysis on the second user based on the multimedia stream data, and generating a target analysis result, wherein the target analysis result is used to indicate whether there is any abnormality in the intention of the second user; a sending module for sending the target analysis result to a target terminal corresponding to the first user, wherein the target terminal is used for the first user to perform qualification review of the second user based on the target analysis result.

[0013] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is also provided, characterized in that the computer-readable storage medium includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned user intention analysis method.

[0014] According to another aspect of an embodiment of the present invention, a processor is further provided, characterized in that the processor is used to run a program, wherein the above-mentioned user intention analysis method is executed when the program is running.

[0015] In an embodiment of the present invention, an online video face-to-face interview method is adopted, and the multimedia stream data generated during the face-to-face interview process is intelligently analyzed by a robot to assist customer service in real time in determining whether to approve the user's application. This achieves efficient application approval, improves customer experience, and solves the technical problem that the customer service cannot take multiple factors into consideration due to the short face-to-face interview time, which leads to misjudgment of the customer's application. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0017] Figure 1 is a flow chart of a method for analyzing user intent according to an embodiment of the present invention;

[0018] Figure 2 is a flow chart of a method for application preparation according to an embodiment of the present invention;

[0019] Figure 3 This is a flow chart of a qualification review method according to an embodiment of the present invention;

[0020] Figure 4 This is a schematic diagram of a client interface according to an embodiment of the present invention;

[0021] Figure 5 2 is a schematic diagram of a video face-to-face signature process according to an embodiment of the present invention;

[0022] Figure 6 This is a structural block diagram of a device for analyzing user intent according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0025] In some industries, such as finance, existing video interview systems can only conduct brief interviews individually, or require time-consuming quality control and analysis after a client submits an application. Some institutions deploy robots to assist customer service during video interviews, but these robots must save and analyze the information from the video interviews, preventing real-time analysis and limiting scalability.

[0026] In order to improve the work efficiency of customer service, according to an embodiment of the present invention, a method embodiment of user intent analysis is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0027] Figure 1 is a flow chart of a method for analyzing user intent according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:

[0028] Step S102: Acquire multimedia stream data, wherein the multimedia stream data is used to represent multimedia data generated during the qualification review of the second user by the first user.

[0029] In the above steps, the first user generally refers to customer service, and the second user generally refers to the customer. When a customer submits an application to a service agency, the service agency generally conducts a risk assessment on the customer to decide whether to approve the customer's application. If the customer is determined to be not risky, the service agency can directly approve the application; if the customer is determined to be risky, the customer service will need to use the user intention analysis system to analyze whether the customer's true intention is the same as the application intention. Generally, the customer service will decide whether to approve the customer's application by conducting a video interview with the customer. The multimedia data in the above steps can generally be the information generated between the customer service and the customer during the video interview, such as the customer's call information, moving images or expression images during the video interview.

[0030] Step S104: performing intention analysis on the second user based on the multimedia stream data to generate a target analysis result, wherein the target analysis result is used to indicate whether there is any abnormality in the intention of the second user.

[0031] Intention analysis in the above steps involves using a target robot to analyze the customer's true intentions based on the multimedia streaming data generated during the video interview process and generate analysis results. Typically, this analysis involves analyzing the customer's body language, tone of voice when answering questions, facial expressions, and emotions. The robot's analysis can be used to determine whether the customer is concealing their true intentions. For example, if a customer suddenly straightens their shoulders while answering a question, or speaks incoherently or incoherently, this indicates that the customer may be lying and concealing their true intentions.

[0032] Step S106: Send the target analysis result to the target terminal corresponding to the first user, wherein the target terminal is used by the first user to conduct a qualification review of the second user based on the target analysis result.

[0033] After using the target robot to analyze multimedia streaming data, the system sends the results to the customer service video interface. Customer service can then use this analysis to further verify the customer's qualifications. For example, if the analysis indicates the customer is emotionally normal and their application intent is normal, the customer service representative can continue to ask other relevant questions. However, if the analysis indicates the customer is overly nervous and suspected of lying, the customer service representative can conduct further investigation based on previously asked questions to determine whether subjective factors may have caused the system's inaccurate judgment.

[0034] By adopting the online video interview method, the robot is connected to intelligently analyze the multimedia streaming data generated during the interview process to assist customer service in real time to determine whether the customer's application is approved. This achieves efficient application approval, improves customer experience, and solves the technical problem that the customer service cannot take multiple factors into consideration due to the short interview time, which leads to misjudgment of the customer's application.

[0035] The process from a customer submitting an application to a service agency until it is approved can generally be divided into three parts. The first part is application preparation, where the customer prepares application-related materials after submitting the application; the second part is qualification review, where the customer service representative conducts a video interview with the customer to analyze the customer's intentions; and the third part is the conclusion of the review, where the customer service representative decides whether to approve the customer's application. The following will describe the above three parts in detail:

[0036] 1. Application Preparation

[0037] Optionally, before obtaining multimedia stream data, the user intention analysis method proposed in the present application also includes: creating a virtual room in response to the application instruction of the second user, wherein the virtual room is a virtual environment for conducting qualification review of the second user; in response to the second user entering the virtual room, sending a first notification instruction to the first user, wherein the first notification instruction is used to notify the first user to enter the virtual room; in response to the first user entering the virtual room, obtaining multimedia stream data of the second user and the first user.

[0038] Figure 2 is a flow chart of a method for application preparation according to an embodiment of the present invention. Figure 2 The specific interview process includes:

[0039] Step S202: creating a virtual room in response to the application instruction of the second user, wherein the virtual room is a virtual environment for conducting qualification review on the second user.

[0040] After submitting an application through a service provider's official website or app, if a client discovers their account is at risk, they can follow the instructions on the website or app to undergo a video interview to mitigate the risk and obtain approval. Upon receiving the client's video interview request, the system creates a virtual room for the interview, verifying the client's qualifications while protecting their privacy.

[0041] In an optional embodiment, customers can select a suitable time for a video interview. If their application is urgently needed and meets the priority processing conditions, they can also request priority processing, and the system will immediately arrange for a video interview with customer service. The application instructions may include, but are not limited to: a video interview link, identity verification information, and application materials. The priority processing conditions may include, but are not limited to, urgent requests such as large-scale transfers and loans to save lives.

[0042] Step S204: In response to the second user entering the virtual room, a first notification instruction is sent to the first user, wherein the first notification instruction is used to notify the first user of entering the virtual room.

[0043] After the customer enters the virtual room, the system can immediately send a first notification instruction to notify an idle customer service representative to conduct an interview in the virtual room to ensure the efficiency of the entire application process. If there is currently no idle customer service representative available for a video interview, the customer can be prompted to wait for a moment on the client interface of the video interview.

[0044] Step S206 : In response to the first user entering the virtual room, multimedia stream data of the second user and the first user are obtained.

[0045] After both the customer and the customer service representative enter the virtual room for the video interview, they can start to obtain multimedia streaming data of the video interview process in real time.

[0046] Through the above-mentioned application preparation method, it is possible to ensure that customers' requests are processed efficiently while protecting their privacy and improving their experience.

[0047] 2. Qualification Review

[0048] Optionally, while executing the aforementioned step S204, the system also includes: in response to the first user entering the virtual room, sending a second notification instruction to the target robot, wherein the second notification instruction is used to notify the target robot to enter the virtual room; controlling the target robot to perform intention analysis on the second user on the multimedia stream data, and generating a target analysis result.

[0049] After the target robot enters the virtual room, it can obtain multimedia streaming data, analyze the second user's intention based on the multimedia streaming data, generate target analysis results, and assist customer service in conducting customer qualification review. Figure 3 is a flow chart of a qualification review method according to an embodiment of the present invention. Figure 3 The specific steps for qualification review are as follows:

[0050] Step S302: In response to the first user entering the virtual room, sending a second notification instruction to the target robot.

[0051] After the customer service representative and the customer enter the virtual room, the target robot receives a second notification instruction sent by the system and enters the virtual room.

[0052] The target robot is an intelligent analysis tool equipped with multiple recognition neural network models. It analyzes multimedia streaming data to determine a customer's level of nervousness and, therefore, whether their intentions are abnormal. These recognition neural network models can include expression analysis, motion analysis, or speech analysis, though these are not specific. Engineers can train and test these neural network models based on common nervous emotions, such as stiff expressions, sudden sitting up, or incoherent speech. For detailed training and testing methods, please refer to relevant literature.

[0053] Step S304: Control the target robot to perform intention analysis on the second user based on the multimedia stream data to generate a target analysis result.

[0054] Because customer qualifications are verified through a video interview, the available multimedia stream data can be divided into two categories: video stream data and audio stream data. After entering the virtual room, the target robot captures the multimedia stream data generated during the video interview process, analyzes the customer's intent based on this multimedia stream data, and generates analysis results.

[0055] Step S306: Send the target analysis result to the target terminal corresponding to the first user.

[0056] The target terminal mentioned above refers to the client terminal of the video face-to-face signing. The target terminal is controlled to display the target analysis results based on the preset display mode, such as Figure 4 As shown, Figure 4 It is a schematic diagram of a customer service interface according to an embodiment of the present invention. Figure 4 It mainly includes three parts. The first part is the video interview image interface, where the customer displays application-related information on the main screen ①, and the customer service staff conducts a qualification review of the customer on the secondary screen ②; the second part is the video interview status interface, which includes the status of external devices such as the camera, microphone, and speaker, and the duration of the current video interview; the third part is the interview result analysis interface, which includes the customer's nervousness and the results of the customer's intention analysis, to assist the customer service staff in determining whether to approve the customer's application.

[0057] As shown in the aforementioned step S304, the multimedia stream data includes video stream data and audio stream data. The video stream data and the audio stream data are introduced separately below:

[0058] 1. In an optional embodiment, the multimedia stream data includes video stream data, and the target analysis result includes a facial analysis result and / or a motion analysis result. Performing an intent analysis on the second user based on the multimedia stream data to generate the target analysis result includes: performing a facial analysis on the second user based on the video stream data to generate a facial analysis result; and / or performing a motion analysis on the second user based on the video stream data to generate a motion analysis result.

[0059] The video stream data is mainly composed of image data. After receiving the video stream data, the target robot can identify the customer's facial expression and body movement in the image, and use the pre-configured expression analysis model and movement analysis model to analyze the two parts to obtain the customer's current tension. Specifically, the tension can be determined by the number and amplitude of the preset actions triggered by the customer in the above analysis model in the video stream data.

[0060] For example, when a customer is lying, they may consciously open their eyes wide and look at the customer service representative, or unconsciously straighten their shoulders. Based on these two points, engineers can set the tension level to 100% if the customer intentionally makes eye contact with the customer service representative more than 10 times within a certain period, 80% if the customer makes eye contact 8 times, and so on. Another option is to set the tension level to 20% if the customer's facial expression is relaxed and smiling, and 60% if the customer frowns and looks distressed. Another option is to set the tension level to 20% if the customer's arms move slightly up and down, and 60% if the arms move frequently, and so on. The above examples are for illustrative purposes only; engineers can customize the specific tension level based on multiple factors and perspectives.

[0061] The target robot analyzes whether the customer's application intention is abnormal based on the customer's current tension. For example, if the target robot determines that the customer's current tension is 0% to 20%, it means that the customer is not nervous, and the true intention and application intention are different. Figure 1 If the target robot determines the customer's current level of nervousness is between 20% and 70%, it indicates that the customer is quite nervous, possibly due to overemphasizing the application or poor communication skills. In this case, customer service will need to conduct further inquiries with the customer to eliminate subjective factors causing nervousness and avoid misjudging the customer's application intentions. If the target robot's current level of nervousness is between 70% and 100%, it indicates that the customer is overly nervous, possibly because their true intentions are inconsistent with their application intentions. In this case, customer service will need to conduct further inquiries after eliminating subjective factors or directly reject the customer's application. The above examples are for illustrative purposes only. The specific approach can be determined by engineers and customer service personnel based on actual circumstances.

[0062] 2. In an optional embodiment, the multimedia stream data also includes: audio stream data, and the target analysis result also includes: voice analysis results. The intention analysis of the second user is performed based on the multimedia stream data to generate the target analysis result, including: voice analysis of the second user based on the audio stream data to generate the voice analysis result.

[0063] The audio stream data mainly consists of voice sentences. After receiving the audio stream data, the target robot can identify the speed, intonation, semantics and fluency of the customer's voice, and use the speech analysis model to analyze it to obtain the customer's tension. Specifically, the tension can be determined by the frequency and amplitude of the customer triggering the preset actions in the above analysis model in the audio stream data.

[0064] For example, a customer might deliberately raise their voice when lying, or speak quickly but with pauses and disjointed sentences, or have inconsistent semantics. Based on these three factors, engineers can set a stress level of 100% if the customer raises their voice more than 10 times within a certain period, 80% if it's 8 times, and so on. Alternatively, engineers can set a stress level of 100% if the customer pauses for more than 5 seconds between sentences, 80% if it's 4 seconds, and so on. Alternatively, engineers can set a stress level of 100% if the customer answers irrelevant questions or if there's a significant semantic discrepancy between sentences, and 60% if there's a slight semantic discrepancy between sentences. These examples are for illustrative purposes only; engineers can customize the specific stress level based on multiple factors and perspectives.

[0065] The way in which the target robot analyzes the application intention based on the customer's tension is shown in the video stream data and will not be repeated here.

[0066] By adding the target robot to the video interview process, considering multiple factors, directly analyzing the video stream data and audio stream data to assist customer service work, it is possible to more effectively determine whether the customer's application intention is abnormal, thereby reducing the customer service's subjective judgment error, and at the same time avoiding the problem of customer service misjudging the customer's application intention due to subjective factors on the customer's side.

[0067] 3. End of review

[0068] Optionally, after the face-to-face qualification review is completed, the system can send a third notification instruction to the second user in response to the first user's exit instruction, and send a fourth notification instruction to the target robot, wherein the third notification instruction is used to notify the second user that the qualification review is completed, and the fourth notification instruction is used to notify the target robot to exit the virtual room.

[0069] That is to say, after the customer service determines the customer's intention based on the analysis results of the target robot, he can first exit the video face-to-face signing interface. After receiving the exit instruction from the customer service, the system can send a third notification instruction and a fourth notification instruction. Among them, the third notification instruction is sent to the client to notify the customer that the qualification review is over and that the customer can exit the virtual room and check or wait for further review results; the fourth notification instruction is sent to the target robot to control the robot to exit the virtual room and enter standby mode so that it can enter the next video face-to-face signing in time to assist the customer service in the qualification review.

[0070] In order to more clearly show the entire video interview process, such as Figure 5 As shown, Figure 5 FIG. 1 is a schematic diagram of a video face-to-face signing process according to an embodiment of the present invention. Figure 5There are four objects in the process, namely, the customer, the video face-to-face signing service, the customer service and the target robot. The actions between them are as shown in the above embodiment and will not be repeated here.

[0071] According to another aspect of an embodiment of the present invention, corresponding to the embodiment of the aforementioned method for analyzing user intention, this specification also provides a device for analyzing user intention. The specific implementation method and application scenario are the same as those of the aforementioned embodiment and will not be repeated here.

[0072] Please refer to Figure 6 , Figure 6 1 is a structural block diagram of a device for analyzing user intent according to an embodiment of the present invention, the device comprising:

[0073] An acquisition module 602 is configured to acquire multimedia stream data, wherein the multimedia stream data is used to represent multimedia data generated during the qualification review of the second user by the first user;

[0074] An analysis module 604 is configured to analyze the intention of the second user based on the multimedia stream data and generate a target analysis result, wherein the target analysis result is used to indicate whether the intention of the second user is abnormal;

[0075] The sending module 606 is used to send the target analysis result to the target terminal corresponding to the first user, wherein the target terminal is used by the first user to conduct qualification review of the second user based on the target analysis result.

[0076] Optionally, before obtaining multimedia stream data, the device also includes: a virtual room creation module, used to create a virtual room in response to the application instruction of the second user, wherein the virtual room is a virtual environment for conducting qualification review of the second user; a first notification instruction sending module, used to send a first notification instruction to the first user in response to the second user entering the virtual room, wherein the first notification instruction is used to notify the first user to enter the virtual room.

[0077] Optionally, when the analysis module 604 performs intent analysis and produces target analysis results, it also includes: a second notification instruction sending module, which is used to send a second notification instruction to the target robot in response to the first user entering the virtual room, wherein the second notification instruction is used to notify the target robot to enter the virtual room.

[0078] Optionally, the multimedia stream data includes: video stream data, the target analysis results include: facial analysis results and / or motion analysis results, and the aforementioned analysis module 604 includes: a facial analysis unit, used to perform facial analysis on the second user based on the video stream data to generate a facial analysis result; and / or, a motion analysis unit, used to perform motion analysis on the second user based on the video stream data to generate a motion analysis result.

[0079] Optionally, the multimedia stream data also includes: audio stream data, the target analysis result also includes: voice analysis result, the intention analysis of the second user is performed based on the multimedia stream data to generate the target analysis result, and the aforementioned analysis module 604 also includes: a voice analysis unit, which is used to perform voice analysis on the second user based on the audio stream data to generate a voice analysis result.

[0080] Optionally, the sending module 606 includes: a target terminal display unit, configured to control the target terminal to display the target analysis result based on a preset display mode.

[0081] Optionally, after sending the analysis results, the device also includes an exit module for sending a third notification instruction to the second user in response to the exit instruction of the first user, and sending a fourth notification instruction to the target robot, wherein the third notification instruction is used to notify the second user that the qualification review is completed, and the fourth notification instruction is used to notify the target robot to exit the virtual room.

[0082] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is also provided, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to perform the user intention analysis method of the above method embodiment.

[0083] According to another aspect of an embodiment of the present invention, a processor is further provided, and the processor is used to run a program, wherein the method of analyzing user intentions in the above method embodiment is executed when the program is running.

[0084] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0085] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0086] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0087] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.

[0088] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0089] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.

[0090] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for analyzing user intention, characterized in that: include: Acquiring multimedia stream data, wherein the multimedia stream data is used to represent multimedia data generated during a qualification review of a second user by a first user; performing an intention analysis on the second user based on the multimedia stream data to generate a target analysis result, wherein the target analysis result is used to indicate whether the second user's intention is abnormal, and the target analysis result is generated by a target robot analyzing the second user's true intention based on the multimedia data. The target robot is an intelligent analysis tool configured with multiple recognition neural network models, and the target robot is used to determine the second user's level of nervousness; Sending the target analysis result to a target terminal corresponding to the first user, wherein the target terminal is used by the first user to conduct a qualification review of the second user based on the target analysis result; Wherein, the multimedia stream data includes: video stream data, the target analysis results include: facial analysis results and motion analysis results, and the intention analysis of the second user is performed based on the multimedia stream data to generate the target analysis results, including: performing facial analysis on the second user based on the video stream data to generate facial analysis results, and performing motion analysis on the second user based on the video stream data to generate motion analysis results, wherein the facial analysis results and the motion analysis results are used to obtain the tension of the second user.

2. The method according to claim 1, characterized in that The method further comprises: In response to the application instruction of the second user, creating a virtual room, wherein the virtual room is a virtual environment for conducting a qualification review of the second user; In response to the second user entering the virtual room, sending a first notification instruction to the first user, wherein the first notification instruction is used to notify the first user of entering the virtual room; In response to the first user entering the virtual room, the multimedia streaming data of the second user and the first user are obtained.

3. The method according to claim 2, characterized in that Performing an intention analysis on the second user based on the multimedia stream data to generate a target analysis result includes: In response to the first user entering the virtual room, sending a second notification instruction to the target robot, wherein the second notification instruction is used to notify the target robot of entering the virtual room; The target robot is controlled to perform intention analysis on the second user based on the multimedia stream data to generate the target analysis result.

4. The method according to claim 1, wherein The multimedia stream data further includes audio stream data, and the target analysis result further includes a voice analysis result. The intention analysis of the second user is performed based on the multimedia stream data to generate the target analysis result, including: Perform a speech analysis on the second user based on the audio stream data to generate a speech analysis result.

5. The method according to claim 1, wherein Sending the target analysis result to a target terminal corresponding to the first user includes: The target terminal is controlled to display the target analysis result based on a preset display mode.

6. The method according to claim 3, characterized in that The method further comprises: In response to the first user's exit instruction, a third notification instruction is sent to the second user, and a fourth notification instruction is sent to the target robot, wherein the third notification instruction is used to notify the second user that the qualification review is completed, and the fourth notification instruction is used to notify the target robot to exit the virtual room.

7. A user intention analysis device, characterized in that: include: an acquisition module, configured to acquire multimedia stream data of the second user and the first user, wherein the multimedia stream data is used to represent multimedia data generated during the qualification review of the second user by the first user; an analysis module, configured to perform an intent analysis on the second user based on the multimedia stream data and generate a target analysis result, wherein the target analysis result indicates whether the second user's intent is abnormal, and the target analysis result is generated by a target robot analyzing the second user's true intent based on the multimedia data. The target robot is an intelligent analysis tool configured with multiple recognition neural network models, and is configured to determine the second user's level of stress; a sending module, configured to send the target analysis result to a target terminal corresponding to the first user, wherein the target terminal is used by the first user to conduct a qualification review of the second user based on the target analysis result; The multimedia stream data includes video stream data, and the target analysis results include facial analysis results and motion analysis results. The device is also used to perform facial analysis on the second user based on the video stream data to generate facial analysis results, and perform motion analysis on the second user based on the video stream data to generate motion analysis results. The facial analysis results and the motion analysis results are used to obtain the tension of the second user.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the user intention analysis method according to any one of claims 1 to 6.

9. A processor, characterized in that: The processor is used to run a program, wherein the program executes the user intention analysis method described in any one of claims 1 to 6 when running.

Citation Information

Patent Citations

  • Audio and video data processing method and device, computer equipment and storage medium

    CN112398931A