Double-recording quality inspection method, system and device and electronic equipment
By combining audio and video separation, transcoding, and frame extraction processing with local cloud quality inspection services, the problem of low efficiency in dual-recording quality inspection is solved, and real-time and efficient dual-recording quality inspection is achieved, reducing hardware costs and improving quality inspection accuracy and customer experience.
Patent Information
- Application Number
- CN202510675724.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-12
AI Technical Summary
The existing dual-recording quality inspection solution has the disadvantages of low efficiency, poor real-time performance, high hardware cost, and complex operation, which affects customer experience and business smoothness.
By acquiring dual-recording information streams, performing audio and video separation, transcoding, and frame extraction, combined with local and cloud-based quality inspection services, preliminary and in-depth inspections are achieved, and dual-recording quality inspection reports are generated, including visual navigation and in-depth analysis.
It achieves the real-time and accuracy of dual-recording quality inspection, reduces hardware costs, improves quality inspection efficiency and customer experience, and ensures compliance and risk control in the insurance sales process.
Smart Images

Figure CN120634576A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of dual-recording quality inspection, and specifically, to a dual-recording quality inspection method, a dual-recording quality inspection system, a dual-recording quality inspection device, a computer-readable storage medium, and an electronic device. Background Art
[0002] The "Interim Measures for Traceable Management of Insurance Sales Behavior" requires that during the insurance sales process, after obtaining the insured's consent, the key and necessary links of the sales process shall be recorded by traceable means of simultaneous audio and video recording in the financial management area, referred to as "double recording", to control risks and improve the compliance of business handling.
[0003] Currently, dual-recording quality inspection models typically upload dual-recording videos from various branches to a national storage center for sample manual quality inspection or AI-powered automated quality inspection. However, this approach not only places high demands on the storage performance and bandwidth of the national center, but also lacks real-time quality inspection. When non-compliant videos require re-recording, policyholders must be contacted to re-record on-site, impacting the customer experience. To improve the real-time nature of dual-recording quality inspection, another current dual-recording quality inspection model involves deploying edge devices such as edge computing boxes and GPUs at branch locations. However, this approach is costly for institutions with a large number of branches. Furthermore, dual-recording quality inspection is coupled with the dual-recording system, resulting in complex dual-recording system functionality and high hardware requirements. Furthermore, sales representatives must manually handle numerous quality inspection-related operations at each stage, impacting the smoothness of dual-recording, burdening sales representatives, and negatively impacting the image quality of customers. Summary of the Invention
[0004] The main purpose of this application is to provide a dual recording quality inspection method, a dual recording quality inspection system, a dual recording quality inspection device, a computer-readable storage medium and an electronic device, so as to at least solve the problem of low dual recording quality inspection efficiency in the dual recording quality inspection solution of the prior art.
[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a dual-recording quality inspection method is provided, comprising: obtaining a dual-recording information stream, wherein the dual-recording information stream represents the audio and video obtained by real-time recording of the conversation between the salesperson and the customer; performing a preprocessing operation on the dual-recording information stream to obtain an audio segment and a picture frame, and performing a visual navigation classification process on the picture frame to obtain a classified picture, and then performing a preliminary detection process on the audio segment and the picture frame through a local quality inspection service to obtain a first detection result, wherein the preprocessing operation includes the separation operation of the audio and video, the transcoding operation of the audio and video, and the extraction of the audio and video. Frame processing, the preliminary detection processing includes black screen detection and / or absence detection of the picture frame, and audio detection of the audio segment; deep detection processing is performed on the audio segment and the classified picture through the cloud quality inspection service to obtain a second detection result, and the first detection result and the second detection result are merged and processed to obtain a dual-recording quality inspection report corresponding to the dual-recording information flow, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each quality inspection item.
[0006] Optionally, deep detection processing is performed on the audio segment and the classified image through a cloud-based quality inspection service, including: performing ASR processing on the audio segment through a cloud-based quality inspection service to obtain an audio text corresponding to the audio segment, and performing the speech detection, sensitive word detection and emotion detection on the audio text in turn; performing OCR detection on the classified image, wherein the OCR detection includes at least one of identity card information verification, work badge information verification, risk assessment form verification, document verification and valet operation analysis.
[0007] Optionally, a preprocessing operation is performed on the dual-recording information stream, including: using OpenCV technology to perform the audio and video separation operation on the dual-recording information stream to obtain a video stream and an audio stream; performing the transcoding operation on the audio stream to obtain multiple audio segments, and performing frame extraction processing on the video stream according to a preset time period to obtain multiple picture frames.
[0008] Optionally, in the process of performing preliminary detection processing on the audio segment and the picture frame through the local quality inspection service, the method also includes: extracting the picture frame at a preset time step to perform the black screen detection, determining that the picture frame identified as the black screen picture has failed the black screen detection, and ending the black screen detection of the video stream when the picture frame extracted for a preset number of consecutive times is the black screen picture, wherein the video stream is composed of all the picture frames; performing the absence detection on the picture frame that passes the black screen detection, and when the number of people in the picture frame is lower than the preset number, determining that the picture frame has failed the absence detection, and when people leave continuously within a preset time period, determining that the video stream has failed the absence detection; and determining that the audio segment has failed the audio detection when it is detected that the average decibel of the audio of the audio segment is lower than the decibel threshold.
[0009] Optionally, performing the OCR detection on the classified image includes: identifying the risk level information of the risk assessment table display diagram based on OCR, and determining whether the risk level information matches the product risk level of the insurance product, wherein the classified image includes the risk assessment table display diagram; if the risk level information matches the product risk level, determining that the risk assessment table display diagram passes the OCR detection; if the risk level information does not match the product risk level, determining that the risk assessment table display diagram fails the OCR detection.
[0010] Optionally, the classification pictures include: the salesperson's work badge confirmation picture, the customer's ID card confirmation picture, the risk assessment form display picture, the insurance application form reading picture, the insurance terms introduction picture, the insurance product manual reading picture, the disclaimer reading picture and at least one of the signature and display pictures.
[0011] According to another aspect of the present application, a dual-recording quality inspection system is provided, comprising: a control terminal, the control terminal being used to execute any one of the dual-recording quality inspection methods; a dual-recording system, the dual-recording system being used to collect dual-recording information streams, the dual-recording system being deployed on the control terminal; a local-end quality inspection service, the local-end quality inspection service being used to perform preliminary detection processing on the dual-recording information streams, the local-end quality inspection service being deployed on the control terminal, and the local-end quality inspection service being communicatively connected to the dual-recording system based on WebSocket technology; a cloud-end quality inspection service, the cloud-end quality inspection service being used to perform deep detection processing on classified audio and classified images, and the cloud-end quality inspection service being communicatively connected to the local-end quality inspection service.
[0012] According to another aspect of the present application, a dual-recording quality inspection device is provided, comprising: an acquisition unit for acquiring a dual-recording information stream, wherein the dual-recording information stream represents audio and video obtained by real-time recording of a conversation between a salesperson and a customer; a preliminary detection processing unit for performing a preprocessing operation on the dual-recording information stream to obtain audio segments and picture frames, and performing visual navigation classification processing on the picture frames to obtain classified pictures, and then performing preliminary detection processing on the audio segments and the picture frames through a local quality inspection service to obtain a first detection result, wherein the preprocessing operation includes an audio and video separation operation, a audio and video transcoding operation, and a audio and video frame extraction operation. The preliminary detection processing includes black screen detection and / or absence detection of the picture frame and audio detection of the audio segment; a deep detection processing unit is used to perform deep detection processing on the audio segment and the classified picture through the cloud quality inspection service to obtain a second detection result, and merge the first detection result and the second detection result to obtain a dual recording quality inspection report corresponding to the dual recording information flow, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item and the time anchor point corresponding to each quality inspection item.
[0013] According to another aspect of the present application, a computer-readable storage medium is provided, which includes a stored program, wherein when the program is run, the device where the computer-readable storage medium is located is controlled to execute any one of the dual-recording quality inspection methods.
[0014] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for executing any one of the dual-recording quality inspection methods.
[0015] By applying the technical solution of the present application, a dual-recording information stream is obtained, wherein the dual-recording information stream represents the audio and video obtained by real-time recording of the conversation between the salesperson and the customer; a preprocessing operation is performed on the dual-recording information stream to obtain audio segments and picture frames, and visual navigation classification processing is performed on the picture frames to obtain classified pictures, and then preliminary detection processing is performed on the audio segments and picture frames through the local quality inspection service to obtain a first detection result, wherein the preprocessing operation includes an audio and video separation operation, an audio and video transcoding operation, and an audio and video frame extraction processing, and the preliminary detection processing includes black screen detection and / or absence detection of the picture frames, and audio detection of the audio segments; deep detection processing is performed on the audio segments and the classified pictures through the cloud-based quality inspection service to obtain a second detection result, and the first detection result and the second detection result are merged to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing, and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each quality inspection item. By separating, transcoding, and extracting frames from audio and video data, effective data preprocessing is achieved, providing high-quality input data for subsequent in-depth quality inspections. At the same time, preliminary quality inspections on the local side can quickly identify basic problems in audio and video, effectively avoiding inaccurate quality inspections due to data quality issues, and solving the technical problems of low efficiency and slow feedback in traditional quality inspection methods. Furthermore, through the collaborative quality inspection model between local and cloud-based systems, not only is real-time quality inspection of dual recordings achieved, but the hardware investment cost on the local side is also greatly reduced, and the accuracy and timeliness of quality inspections are improved, thereby improving business compliance and enhancing customer experience. At the same time, through the introduction of visual navigation technology, dual-recording quality inspections can more accurately locate key links, significantly improving the accuracy and speed of quality inspections, ensuring risk control capabilities in the insurance sales process, and solving the problem of low efficiency of dual-recording quality inspections in existing technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings that constitute part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application. In the drawings:
[0017] Figure 1 The following is a hardware structure diagram of a mobile terminal for performing a dual recording quality inspection method provided in an embodiment of the present application;
[0018] Figure 2 A schematic diagram of a process of a dual recording quality inspection method provided according to an embodiment of the present application is shown;
[0019] Figure 3 The structure diagram of the dual recording quality inspection system provided according to the embodiment of the present application is shown;
[0020] Figure 4 A schematic diagram of a process flow of a specific dual recording quality inspection method provided according to an embodiment of the present application is shown;
[0021] Figure 5 A schematic diagram of the structure of edge-end quality inspection and cloud-end quality inspection provided according to an embodiment of the present application is shown;
[0022] Figure 6 A schematic diagram of splitting a dual-recording video structure according to an embodiment of the present application is shown;
[0023] Figure 7 The figure shows a structural block diagram of a dual recording quality inspection device provided according to an embodiment of the present application.
[0024] The above drawings include the following reference numerals:
[0025] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device; 71. Acquisition unit; 72. Preliminary detection processing unit; 73. Depth detection processing unit. DETAILED DESCRIPTION
[0026] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] As introduced in the background technology, the dual recording quality inspection scheme in the prior art has the problem of low dual recording quality inspection efficiency. In order to solve the problem of low dual recording quality inspection efficiency in the dual recording quality inspection scheme in the prior art, the embodiments of the present application provide a dual recording quality inspection method, a dual recording quality inspection system, a dual recording quality inspection device, a computer-readable storage medium and an electronic device.
[0030] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.
[0031] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure diagram of a mobile terminal for a dual recording quality inspection method according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0032] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the dual-recording quality inspection method in the embodiment of the present invention. The processor 102 executes the computer programs stored in the memory 104 to execute various functional applications and data processing, thereby implementing the above-mentioned method. The memory 104 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or transmit data via a network. Specific examples of such networks may include a wireless network provided by the mobile terminal's telecommunications provider. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0033] In this embodiment, a dual recording quality inspection method running on a mobile terminal, a computer terminal or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0034] Figure 2 This is a flow chart of the dual recording quality inspection method according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:
[0035] Step S201: obtaining a dual-recording information stream, wherein the dual-recording information stream represents the audio and video obtained by real-time recording of the conversation between the salesperson and the customer;
[0036] Step S202: Preprocessing the dual-recording information stream to obtain audio segments and image frames, performing visual navigation classification on the image frames to obtain classified images, and then performing preliminary detection on the audio segments and image frames through a local quality inspection service to obtain a first detection result. The preprocessing includes separating the audio and video, transcoding the audio and video, and extracting frames from the audio and video. The preliminary detection includes black screen detection and / or seat absence detection on the image frames and audio detection on the audio segments.
[0037] Specifically, preprocessing is the foundation of the entire quality inspection process. It separates the raw audio and video data into audio and video streams, facilitating subsequent independent processing. Transcoding ensures that the audio and video data can be effectively recognized and processed by the quality inspection system. For example, audio is converted to WAV or MP3 format, and video is converted to JPG or PNG image frames. Frame extraction extracts image frames from the video stream at a consistent frequency to reduce processing effort and improve quality inspection efficiency.
[0038] The local quality inspection service performs preliminary testing and processing, quickly identifying basic audio and video issues such as black screens, attendees leaving the room, and poor audio quality. These issues are essential for ensuring the quality of dual-recording data and a prerequisite for sales compliance. This local preliminary testing and processing not only significantly improves quality inspection efficiency but also provides real-time feedback, providing sales staff and customers with immediate corrective suggestions and ensuring compliance during the insurance product sales process.
[0039] Step S203: Perform deep detection processing on the audio segment and the classified images through the cloud-based quality inspection service to obtain a second detection result, and merge the first detection result and the second detection result to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each of the above quality inspection items.
[0040] Specifically, visual navigation uses target detection algorithms to identify all key targets appearing in the video, such as identity cards, work badges, job qualification certificates, professional registration certificates, risk assessment forms, various documents, etc. Based on the identified key targets, the video is segmented in time sequence, and different images are sent to the cloud quality inspection service for in-depth analysis using different AI algorithms, which is conducive to saving computing resources and improving the speed and accuracy of detection.
[0041] The local quality inspection service establishes a connection with the dual-recording system, capturing both recorded videos in real time and performing an initial inspection of the dual-recording information stream. The cloud performs more sophisticated AI analysis on the uploaded images and audio after visual navigation screening. This collaborative local and cloud-based quality inspection model saves bandwidth, improves site resource utilization, and supports more complex AI analysis, helping to improve the efficiency of dual-recording quality inspection.
[0042] Through this embodiment, by applying the above-mentioned steps S201, S202, and S203, by separating, transcoding, and extracting frames from audio and video data, effective data preprocessing is achieved, providing high-quality input data for subsequent in-depth quality inspection. At the same time, the preliminary quality inspection on the local side can quickly identify basic problems in the audio and video, effectively avoiding inaccurate quality inspection due to data quality issues, and solving the technical problems of low efficiency and slow feedback of traditional quality inspection methods; and through the collaborative quality inspection mode of local and cloud, not only is real-time quality inspection of dual recordings achieved, but the hardware investment cost of the local side is also greatly reduced, and the accuracy and timeliness of quality inspection are improved, thereby improving business compliance and enhancing customer experience. At the same time, through the introduction of visual navigation technology, dual-recording quality inspection can more accurately locate key links, significantly improve the accuracy and speed of quality inspection, and ensure the risk control capability in the insurance sales process; it solves the problem of low efficiency of dual-recording quality inspection in the dual-recording quality inspection scheme of the existing technology.
[0043] During the specific implementation process, deep detection processing is performed on the above-mentioned audio segment and the above-mentioned classified pictures through the cloud-based quality inspection service, including: performing ASR processing on the above-mentioned audio segment through the cloud-based quality inspection service to obtain the audio text corresponding to the above-mentioned audio segment, and performing the above-mentioned speech detection, sensitive word detection and emotion detection on the above-mentioned audio text in turn; performing the above-mentioned OCR detection on the above-mentioned classified pictures, wherein the above-mentioned OCR detection includes at least one of identity card information verification, work badge information verification, risk assessment form verification, document verification and customer service operation analysis.
[0044] Among them, the speech detection means that every link that the account manager needs to confirm must be carefully and clearly confirmed by the customer, such as: clear, understand, good, understood, no problem, etc.; sensitive word detection: based on the maintained sensitive word library and full-word matching algorithm; emotion detection, based on the recognized text to perform conversation emotion recognition, in the sales process dialogue scenario, identify the user emotions behind the text of the two parties in the dialogue, and emotions are divided into three types: positive, neutral, and negative.
[0045] The cloud-based quality inspection service in this method is the core of the entire quality inspection process. Through deep detection and processing, it conducts a more detailed analysis of sales personnel's behavior, language, and customer identity information. ASR processing, or automatic speech recognition processing, can convert audio segments into text to facilitate subsequent language detection, sensitive word detection, and emotion detection. Language detection is used to check whether sales personnel are following the company's prescribed language for sales. Sensitive word detection is used to identify whether any violations or inappropriate language occur during the sales process. Emotion detection is used to analyze the emotional state of sales personnel and customers to ensure a friendly and professional sales process. OCR detection, or optical character recognition detection, is used to identify and classify text information in images, such as identity card information verification, work badge information verification, risk assessment form verification, document verification, etc., to ensure the accuracy and compliance of documents and information during the sales process. Through deep detection and processing of the cloud-based quality inspection service, this technical solution can conduct a detailed analysis of sales personnel's behavior, language, and customer identity information, ensuring the transparency and standardization of the sales process. The combination of ASR processing and OCR detection not only improves the accuracy and efficiency of quality inspection, but also provides real-time feedback on quality inspection results, providing sales staff and customers with immediate correction suggestions, effectively avoiding sales disputes and compliance risks caused by inaccurate quality inspection.
[0046] Specifically, a preprocessing operation is performed on the dual-recording information stream, including: using OpenCV technology to perform the audio and video separation operation on the dual-recording information stream to obtain a video stream and an audio stream; performing the transcoding operation on the audio stream to obtain the multiple audio segments, and performing frame extraction processing on the video stream according to a preset time period to obtain the multiple picture frames.
[0047] The preset time period is set according to specific frame extraction requirements, for example, it can be set to 1s.
[0048] This method utilizes OpenCV technology, an open-source computer vision library that provides a wealth of image and video processing capabilities, including image analysis, video processing, and feature detection. In this technical solution, OpenCV technology is used to perform audio and video separation on the dual-recording information streams, efficiently decomposing the original audio and video data into video and audio streams, facilitating subsequent independent processing. Transcoding ensures that the audio and video data can be effectively recognized and processed by the quality inspection system. For example, audio is converted to WAV or MP3 format, and video is converted to JPG or PNG format image frames. Frame extraction extracts image frames from the video stream at a certain frequency to reduce processing and improve quality inspection efficiency. By utilizing OpenCV technology for audio and video separation, as well as transcoding and frame extraction of audio and video data, this technical solution achieves effective data preprocessing, providing high-quality input data for subsequent in-depth quality inspection. The efficiency and flexibility of OpenCV technology, along with optimized transcoding and frame extraction, not only improve quality inspection efficiency but also ensure data quality, resolving the technical issues of low efficiency and complex data processing associated with traditional quality inspection methods.
[0049] More specifically, in the process of performing preliminary detection and processing on the above-mentioned audio segment and the above-mentioned picture frame through the local quality inspection service, the above-mentioned method also includes: extracting the above-mentioned picture frame at a preset time step to perform the above-mentioned black screen detection, determining that the above-mentioned picture frame identified as the black screen picture has not passed the above-mentioned black screen detection, and ending the above-mentioned black screen detection of the video stream when the above-mentioned picture frame extracted for a preset number of consecutive times is the above-mentioned black screen picture, wherein the above-mentioned video stream is composed of all the above-mentioned picture frames; performing the above-mentioned absence detection on the above-mentioned picture frame that has passed the above-mentioned black screen detection, and when the number of people in the above-mentioned picture frame is lower than the preset number, determining that the above-mentioned picture frame has not passed the above-mentioned absence detection, and when people leave continuously within a preset time period, determining that the above-mentioned video stream has not passed the above-mentioned absence detection; when it is detected that the average decibel of the audio of the above-mentioned audio segment is lower than the decibel threshold, determining that the above-mentioned audio segment has not passed the above-mentioned audio detection.
[0050] Among them, the preset time step is set according to the black screen detection requirements and can be set to 1s, and the continuous preset number of times can be set to 5 times, 8 times, 10 times, etc.
[0051] Among them, the leave detection rule is to count the number of faces and roles in the picture. If the policyholder and the insured are the same person, two faces must appear in the picture, and they are the financial manager and customer roles respectively. If the policyholder and the insured are not the same person, three faces must appear in the picture, and they are the financial manager and customer roles respectively. If the rules are met, the leave detection is passed. Otherwise, the image detection is directly terminated. If people leave their seats continuously for a period of time, the entire video detection is terminated and the video compliance detection fails.
[0052] In this technical solution, the preliminary detection and processing of the local quality inspection service includes black screen detection, absence detection and audio detection of audio segments and picture frames. Black screen detection is to extract picture frames at a preset time step and check whether the picture frames are black screens to ensure the effectiveness of video surveillance. Absence detection is to check the number of people in the picture frame to ensure the presence of sales staff and customers during the sales process. Audio detection is to check the average decibel of the audio segment to ensure that the audio quality meets the quality inspection requirements. Through the preliminary detection and processing of the local quality inspection service, this technical solution can quickly identify basic problems in audio and video, such as black screen, absence and poor audio quality. These problems are the basis for ensuring the quality of dual-recording data and are also a prerequisite for ensuring sales compliance. By setting the preset time step, preset number of times and decibel threshold, not only can the efficiency of quality inspection be significantly improved, but also the quality of data can be ensured, solving the technical problems of low efficiency and slow feedback of traditional quality inspection methods.
[0053] Furthermore, the above-mentioned OCR detection is performed on the above-mentioned classification picture, including: based on OCR, identifying the risk level information of the risk assessment table display diagram, and determining whether the above-mentioned risk level information matches the product risk level of the insurance product, wherein the above-mentioned classification picture includes the above-mentioned risk assessment table display diagram; when the above-mentioned risk level information matches the above-mentioned product risk level, determining that the above-mentioned risk assessment table display diagram has passed the above-mentioned OCR detection; when the above-mentioned risk level information does not match the above-mentioned product risk level, determining that the above-mentioned risk assessment table display diagram has not passed the above-mentioned OCR detection.
[0054] This method of OCR detection, namely optical character recognition detection, is to check the accuracy and compliance of documents and information in the sales process by identifying text information in classified images. In this technical solution, OCR detection is used to identify the risk level information in the risk assessment form display diagram, and match it with the product risk level of the insurance product, so as to ensure the consistency of the risk assessment in the sales process and the product risk level. The risk assessment form display diagram is an assessment form that the salesperson shows to the customer during the sales process. It is used to assess the customer's risk tolerance and ensure that the product sold matches the customer's risk tolerance. Through OCR detection, this technical solution can identify the risk level information in the risk assessment form display diagram, and match it with the product risk level of the insurance product, so as to ensure the consistency of the risk assessment in the sales process and the product risk level. The accuracy and efficiency of OCR detection not only improve the efficiency of quality inspection, but also ensure the quality of data, solving the technical problems of low efficiency and complex data processing of traditional quality inspection methods.
[0055] Furthermore, the above-mentioned classification pictures include: the above-mentioned salesperson's work badge confirmation picture, the above-mentioned customer's ID card confirmation picture, the risk assessment form display picture, the insurance application form reading picture, the insurance terms introduction picture, the insurance product manual reading picture, the disclaimer reading picture and at least one of the signature and display pictures.
[0056] The classified images in this method refer to images with specific information obtained through visual navigation classification processing in the real-time audio and video data of the conversation between the salesperson and the customer recorded by video surveillance equipment during the sales process. In this technical solution, the classified images include but are not limited to the salesperson’s work badge confirmation image, the customer’s ID card confirmation image, the risk assessment form display image, the insurance application form reading image, the insurance terms introduction image, the insurance product manual reading image, the disclaimer reading image, and the signature and display image. These images contain key information in the sales process, such as the salesperson’s identity, the customer’s personal information, the product information sold, the operations during the sales process, etc. By extracting key information from the video data into classified images, this technical solution not only improves the efficiency of quality inspection, but also ensures the quality of the data.
[0057] This embodiment also includes a dual-recording quality inspection system, including: a control terminal, the control terminal is used to execute any one of the above-mentioned dual-recording quality inspection methods; a dual-recording system, the dual-recording system is used to collect dual-recording information streams, and the dual-recording system is deployed on the control terminal; a local-end quality inspection service, the local-end quality inspection service is used to perform preliminary detection and processing on the dual-recording information streams, the local-end quality inspection service is deployed on the control terminal, and the local-end quality inspection service is connected to the dual-recording system based on WebSocket technology; a cloud-end quality inspection service, the cloud-end quality inspection service is used to perform deep detection and processing on classified audio and classified images, and the cloud-end quality inspection service is connected to the local-end quality inspection service.
[0058] The dual-recording quality inspection system is a system used to perform quality inspection on audio and video data during the sales process. It includes a control terminal, a dual-recording system, a local quality inspection service, and a cloud-based quality inspection service. The control terminal is the control center of the entire system, used to execute the dual-recording quality inspection method and control the operation of the entire quality inspection process. The dual-recording system is a system used to collect dual-recording information streams. It is deployed on the control terminal and can record the audio and video data of conversations between sales personnel and customers in real time. The local quality inspection service and the cloud-based quality inspection service are services used to perform preliminary inspection and in-depth inspection on the dual-recording information streams. They are deployed on the control terminal and the cloud, respectively, and achieve communication connections via HTTP / TCP technology, capable of real-time data transmission and feedback. This technical solution achieves efficient quality inspection of audio and video data during the sales process through the combination of the control terminal, dual-recording system, local quality inspection service, and cloud-based quality inspection service. The coordination of the control function of the control terminal, the collection function of the dual recording system, the preliminary detection and processing function of the local quality inspection service, and the in-depth detection and processing function of the cloud-based quality inspection service not only improves the efficiency and accuracy of quality inspection, but also provides real-time feedback on results, providing sales staff and customers with immediate correction suggestions to ensure the compliance of the sales process.
[0059] Among them, such as Figure 3 As shown, the control terminal can be a PC recording and recording audio, and the local quality inspection service is a real-time quality inspection service at the edge. The edge includes real-time quality inspection services such as a conversation module, preprocessing module, quality inspection module, and reporting module. The quality inspection module includes visual navigation and quality inspection capabilities. The real-time quality inspection service is deployed on a PC recording and recording audio that has been deployed with a dual-recording system at the branch. The conversation module establishes a connection with the dual-recording system and receives the audio and video stream files pushed by the dual-recording system. The preprocessing module performs preprocessing such as transcoding and frame extraction to separate audio segments and image frames, which are then fed into the quality inspection module. The quality inspection module first performs black screen detection and seat absence detection on the image frames. It then performs AI-based analysis such as abnormal behavior analysis and visual navigation on videos that meet the requirements. The images selected by the visual navigation module are then uploaded to the cloud for more complex AI analysis. The quality inspection module performs channel detection, silence detection, and noise detection on the audio at the edge. The audio clips are then uploaded to the cloud for speech-to-text conversion, further performing speech detection, sensitive word detection, and sentiment analysis. Finally, the combined detection results are fed back to the dual-recording system for real-time correction, and a quality inspection report for the entire video is generated on the edge and fed back to the cloud.
[0060] The cloud includes services such as quality inspection capabilities and reporting modules. The cloud-based quality inspection capabilities perform more complex OCR analysis, intelligent image understanding, intelligent audio, natural language processing, and behavioral analysis, feeding the inspection results back to the edge for consolidation. Finally, the cloud also generates a quality inspection report for the entire video.
[0061] This embodiment also includes a dynamic threshold adjustment mechanism. To further enhance the accuracy and adaptability of visual navigation, a dynamic threshold adjustment mechanism is introduced. This mechanism automatically adjusts the recognition threshold in the target detection model based on different network environments and time periods to accommodate environmental factors such as lighting changes and crowd density. In a specific embodiment, the system dynamically calculates the recognition threshold in each network's edge-side quality inspection service based on historical quality inspection data and current environmental conditions (such as light intensity and background noise), ensuring accurate identification of key targets in a variety of complex environments. For example, in low-light environments, the system automatically lowers the target detection threshold to avoid misjudgments; during busy periods, the system increases the face recognition threshold for leave detection to reduce false positives. The dynamic threshold adjustment mechanism significantly improves the accuracy and stability of target detection, enabling dual-recording quality inspection to remain efficient in various environments. Through this mechanism, the method of the present invention can automatically adapt to environmental changes in different network locations and time periods, reducing false positives, improving the real-time nature of quality inspection, and enhancing customer experience.
[0062] This embodiment also includes an emotion recognition and feedback system. This system, integrated into the dual-recording quality inspection process, monitors the emotions of conversations during the sales process in real time, ensuring a compliant and user-friendly sales process. In a specific embodiment, the system leverages deep learning technology to analyze a customer's emotional state based on their voice and facial expressions. This analysis is then combined with quality inspection reports and provided real-time feedback to sales managers and a backend monitoring center. For example, if the system detects a customer expressing confusion or anxiety while explaining insurance terms, the quality inspection service immediately notifies the sales manager to provide emotional comfort and further explanation, ensuring the customer fully understands and agrees to all terms, thereby avoiding potential complaints and risks. The introduction of this emotion recognition and feedback system allows the quality inspection process to focus not only on compliance but also on the customer's emotional experience, thereby improving customer satisfaction and business compliance throughout the entire sales process. By providing real-time emotional monitoring and feedback, sales managers can promptly adjust their communication strategies to ensure a smooth sales process and customer comfort, effectively improving conversion rates and word-of-mouth.
[0063] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the dual recording quality inspection method of the present application will be described in detail below with reference to specific embodiments.
[0064] This embodiment relates to a specific dual recording quality inspection method, such as Figure 4 As shown, including the following:
[0065] 1) If Figure 3As shown in the figure, both the dual-recording system and the real-time quality inspection service are deployed on the business recording and video recording PC. They use WebSocket technology to establish and manage connections. A handshake is required between the dual-recording system and the real-time quality inspection service. After the handshake is completed, a persistent connection is established using a full-duplex communication protocol to achieve real-time data transmission. Once the WebSocket connection is established, the dual-recording system and the real-time quality inspection service can exchange WebRTC information, such as ICE candidates and SDP, over WebSocket, and use this information to establish the WebRTC connection.
[0066] 2) Once the WebRTC connection is successfully established, the dual recording system sends the dual recording information to the local real-time quality inspection service in a streaming manner. The real-time quality inspection service responds to the start operation of the dual recording system deployed on the PC.
[0067] 3) After receiving the information stream, the real-time quality inspection service performs audio and video separation, transcoding, and frame extraction, extracting the audio segments and picture frames, and sending the audio segments and picture frames to the quality inspection module for basic environment detection, visual navigation, and global detection.
[0068] 4) The edge further uploads the images and audio classified by visual navigation to the cloud for more complex AI analysis such as OCR detection, behavior analysis, ASR, and speech detection.
[0069] 5) The cloud will feed back the analysis results to the edge in real time, and the edge will feed back the combined detection results to the dual recording system via WebSocket for real-time correction.
[0070] 6) After the dual recording is completed, a visual quality inspection report of the entire video is generated on the edge, including whether each quality inspection item has passed and the corresponding time anchor point. For quality inspection items that have failed, the reasons for failure should be displayed, such as: *** ID not presented, *** ID information is inconsistent with the transaction information, etc., and the quality inspection report is fed back to the cloud.
[0071] Figure 5 This diagram shows the architecture of edge-side and cloud-side quality inspection. The edge preprocessing module uses OpenCV to decapsulate audio and video files, extracting the video and audio streams from the files and outputting them to separate output files. The audio and video files are transcoded separately to generate audio clips and image collections.
[0072] An image frame is extracted every n milliseconds to form a test image set. The image set first undergoes black screen detection, extracting one image every second. If the image is black, the test ends immediately. If the proportion of black screen images exceeds a threshold for a continuous period of time, the entire video detection ends. After the image passes the black screen detection, the absence detection is performed. The absence detection rule counts the number of faces and roles in the image. If the policyholder and the insured are the same person, the image must contain two faces, one for the financial manager and one for the customer. If the policyholder and the insured are different people, the image must contain three faces, one for the financial manager and one for the customer. If these rules are met, the absence detection passes; otherwise, the image detection ends immediately. If a person repeatedly leaves the table for a continuous period of time, the entire video detection ends and the video compliance test fails.
[0073] After basic detection, the image enters global visual detection, which is mainly divided into visual navigation and global detection. Global detection does not require navigation, but each extracted image frame must be inspected to analyze abnormal behaviors such as falls, sleeping, and fighting. Visual navigation is achieved through object detection, which has very fast inference speed and can also be deployed on CPU devices. An object detection model is trained to include all key objects in the video. One image frame is extracted every second for object detection. All key objects that appear are identified and classified, such as identity cards, work badges, job qualification certificates, professional registration certificates, risk assessment forms, and various documents. Visual navigation helps to decompose the entire video structure, saving computing resources and improving detection speed and accuracy. Visual navigation clearly identifies the start and end times of key segments in the dual-recording video.
[0074] Among them, the schematic diagram of the dual-recording video structure split is as follows Figure 6 As shown, visual navigation uses a target detection algorithm to classify images and sends them to different detection algorithms for in-depth analysis. For example, if an ID card appears in an image, the time period in which the ID card appears is recorded, and one image is extracted every m milliseconds to form a collection. This collection containing the ID card images is sent to the cloud-based ID card information verification module for intelligent OCR recognition of ID card information, such as ID number, name, and expiration date. This information is then compared with the transaction data using a minimum edit distance algorithm to verify the ID card number, name, and other identity information. If a risk assessment form appears in an image, the time period in which the risk assessment form appears is recorded, and one image is extracted every m milliseconds to form a collection. This collection containing the risk assessment form images is then sent to the cloud-based risk assessment form verification module for OCR recognition of the risk level information on the risk assessment form and comparison with the risk level of the insurance policy to ensure that the risk level the customer can bear matches the risk level of the product.
[0075] like Figure 5As shown, the audio segment is first detected for audio presence. If an audio file exists, further audio analysis is performed; otherwise, audio detection ends immediately. The audio file is further detected for silence to determine whether the average decibel level is below a threshold. If so, the audio is considered too low and is considered silent.
[0076] Upload audio files to the cloud-based speech-to-text service, which converts them into text for natural language processing. Sensitive word detection is performed based on a maintained sensitive word library and full-word matching algorithm. The conversational language detection module ensures that each step the account manager needs to confirm is carefully and clearly confirmed by the customer, such as: "clear," "understood," "good," "understood," "no problem," etc. The sentiment analysis module uses the recognized text to identify conversation emotions. In sales conversations, it identifies the user emotions behind the text of both parties. Emotions are categorized as positive, neutral, and negative.
[0077] This application specifically achieves the following technical effects:
[0078] 1) Use target detection algorithms to identify all key targets appearing in the video, such as ID cards, work badges, job qualification certificates, professional registration certificates, risk assessment forms, and various documents. Based on the identified key targets, the video is segmented in time sequence, and different images are sent to different AI algorithms for in-depth analysis, which helps save computing resources and improve detection speed and accuracy.
[0079] 2) Through the cloud-edge collaborative quality inspection mode, the edge real-time quality inspection service and the dual-recording system establish a connection, capture dual-recording videos in real time, and complete the initial video inspection. The cloud performs more complex AI analysis on the images and audio uploaded after the edge visual navigation screening. The cloud-edge collaborative quality inspection mode saves bandwidth, improves the utilization of network resources and supports more complex AI analysis.
[0080] 3) Real-time quality inspection: The edge-end dual-recording quality inspection service receives the dual-recording stream information of the dual-recording system in real time, performs audio and video quality inspection analysis, and feeds back the results to the dual-recording system in real time to correct non-compliant dual-recording links.
[0081] 4) Low investment cost. The edge dual-recording quality inspection service is deployed on the PC for recording and video recording at the outlets. There is no need to add additional edge computing boxes, GPUs and other hardware equipment, which facilitates rapid promotion.
[0082] 6) Flexible and efficient, the quality inspection method and the dual recording system are decoupled, eliminating the need for business managers to perform additional quality inspection operations. The dual recording process is smoother, which is conducive to improving dual recording efficiency and customer experience.
[0083] In summary, when this embodiment conducts quality inspection on the key necessary links of audio and video recording in the insurance sales process, real-time quality inspection can be achieved without adding additional equipment at the outlets. It supports more complex and multi-dimensional dual-recording quality inspection items through cloud-edge collaboration, and improves quality inspection efficiency and effectiveness based on visual navigation, thereby improving risk management capabilities and customer dual-recording experience, and better safeguarding consumer rights.
[0084] The embodiments of the present application also provide a dual recording quality inspection device. It should be noted that the dual recording quality inspection device of the embodiments of the present application can be used to execute the dual recording quality inspection method provided by the embodiments of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and the details that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and conceivable.
[0085] The following introduces the dual recording quality inspection device provided in the embodiment of the present application.
[0086] Figure 7 Schematic diagram of a dual recording quality inspection device according to an embodiment of the present application. Figure 7 As shown, the device includes:
[0087] An acquisition unit 71 is configured to acquire a dual-recording information stream, wherein the dual-recording information stream represents the audio and video acquired by real-time recording of a conversation between a salesperson and a customer;
[0088] A preliminary detection processing unit 72 is configured to perform preprocessing operations on the dual-recording information stream to obtain audio segments and image frames, perform visual navigation classification processing on the image frames to obtain classified images, and then perform preliminary detection processing on the audio segments and image frames through a local quality inspection service to obtain a first detection result. The preprocessing operations include audio and video separation, audio and video transcoding, and audio and video frame extraction. The preliminary detection processing includes black screen detection and / or seat absence detection for the image frames and audio detection for the audio segments.
[0089] The deep detection processing unit 73 is used to perform deep detection processing on the above-mentioned audio segment and the above-mentioned classified pictures through the cloud-based quality inspection service to obtain a second detection result, and merge the above-mentioned first detection result and the above-mentioned second detection result to obtain a dual-recording quality inspection report corresponding to the above-mentioned dual-recording information stream, wherein the above-mentioned deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the above-mentioned dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each of the above-mentioned quality inspection items.
[0090] In this embodiment, an acquisition unit is configured to acquire a dual-recording information stream, wherein the dual-recording information stream represents audio and video obtained by real-time recording of a conversation between a salesperson and a customer. A preliminary detection processing unit is configured to perform preprocessing operations on the dual-recording information stream to obtain audio segments and image frames, perform visual navigation classification processing on the image frames to obtain classified images, and then perform preliminary detection processing on the audio segments and image frames through a local quality inspection service to obtain a first detection result, wherein the preprocessing operations include audio and video separation operations, audio and video transcoding operations, and audio and video frame extraction processing, and the preliminary detection processing includes black screen detection and / or absence detection of the image frames and audio detection of the audio segments. A deep detection processing unit is configured to perform deep detection processing on the audio segments and the classified images through a cloud-based quality inspection service to obtain a second detection result, and merge the first detection result and the second detection result to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing, and speech detection. The dual-recording quality inspection report includes whether each quality inspection item passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each quality inspection item. By separating, transcoding, and extracting frames from audio and video data, effective data preprocessing is achieved, providing high-quality input data for subsequent in-depth quality inspections. At the same time, preliminary quality inspections on the local side can quickly identify basic problems in audio and video, effectively avoiding inaccurate quality inspections due to data quality issues, and solving the technical problems of low efficiency and slow feedback in traditional quality inspection methods. Furthermore, through the collaborative quality inspection model between local and cloud-based systems, not only is real-time quality inspection of dual recordings achieved, but the hardware investment cost on the local side is also greatly reduced, and the accuracy and timeliness of quality inspections are improved, thereby improving business compliance and enhancing customer experience. At the same time, through the introduction of visual navigation technology, dual-recording quality inspections can more accurately locate key links, significantly improving the accuracy and speed of quality inspections, ensuring risk control capabilities in the insurance sales process, and solving the problem of low efficiency of dual-recording quality inspections in existing technologies.
[0091] As an optional solution, the deep detection processing unit includes a first execution module and a second execution module. The first execution module is used to perform ASR processing on the above-mentioned audio segment through the cloud quality inspection service to obtain the audio text corresponding to the above-mentioned audio segment, and perform the above-mentioned speech detection, sensitive word detection and emotion detection on the above-mentioned audio text in turn; the second execution module is used to perform the above-mentioned OCR detection on the above-mentioned classified pictures, wherein the above-mentioned OCR detection includes at least one of identity card information verification, work badge information verification, risk assessment form verification, document verification and customer operation analysis.
[0092] An optional solution, the preliminary detection processing unit includes a separation operation module and a frame extraction processing module. The separation operation module is used to use OpenCV technology to perform the above-mentioned audio and video separation operation on the above-mentioned dual-recording information stream to obtain a video stream and an audio stream; the frame extraction processing module is used to perform the above-mentioned transcoding operation on the above-mentioned audio stream to obtain multiple audio segments, and perform frame extraction processing on the above-mentioned video stream according to a preset time period to obtain multiple picture frames.
[0093] An optional solution, the device also includes a first determination unit, a second determination unit and a third determination unit, the first determination unit is used to extract the above-mentioned picture frames at a preset time step to perform the above-mentioned black screen detection during the process of performing preliminary detection processing on the above-mentioned audio segment and the above-mentioned picture frames through the local quality inspection service, determine that the above-mentioned picture frames identified as black screen pictures have not passed the above-mentioned black screen detection, and end the above-mentioned black screen detection of the video stream when the above-mentioned picture frames extracted for a preset number of consecutive times are the above-mentioned black screen pictures, wherein the above-mentioned video stream is composed of all the above-mentioned picture frames; the second determination unit is used to perform the above-mentioned absence detection on the above-mentioned picture frames that have passed the above-mentioned black screen detection, and when the number of people in the above-mentioned picture frames is lower than the preset number, determine that the above-mentioned picture frames have not passed the above-mentioned absence detection, and when people leave continuously within a preset time period, determine that the above-mentioned video stream has not passed the above-mentioned absence detection; the third determination unit is used to determine that the above-mentioned audio segment has not passed the above-mentioned audio detection when it is detected that the average decibel of the audio of the above-mentioned audio segment is lower than the decibel threshold.
[0094] An optional solution, the second execution module includes a first determination submodule, a second determination submodule and a third determination submodule; the first determination submodule is used to identify the risk level information of the risk assessment table display diagram based on OCR, and determine whether the above risk level information matches the product risk level of the insurance product, wherein the above classification picture includes the above risk assessment table display diagram; the second determination submodule is used to determine that the above risk assessment table display diagram has passed the above OCR detection when the above risk level information matches the above product risk level; the third determination submodule is used to determine that the above risk assessment table display diagram has not passed the above OCR detection when the above risk level information does not match the above product risk level.
[0095] An optional solution, the above-mentioned classification pictures include: the above-mentioned salesperson’s work badge confirmation picture, the above-mentioned customer’s ID card confirmation picture, the risk assessment form display picture, the insurance application form reading picture, the insurance terms introduction picture, the insurance product manual reading picture, the disclaimer reading picture and at least one of the signature and display pictures.
[0096] The dual-recording quality inspection device includes a processor and memory. The acquisition unit, preliminary detection processing unit, and depth detection processing unit are all stored as program units in the memory. The processor executes these program units stored in the memory to implement the corresponding functions. All of the above modules are located in the same processor; alternatively, the above modules can be located in different processors in any combination.
[0097] The processor includes a core, which retrieves the corresponding program unit from the memory. One or more cores can be set, and the problem of low efficiency of dual recording quality inspection solutions in the existing technology can be solved by adjusting the core parameters.
[0098] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0099] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is run, the device where the computer-readable storage medium is located is controlled to execute the dual-recording quality inspection method.
[0100] Specifically, the dual-recording quality inspection method includes:
[0101] Step S201, obtaining a dual-recording information stream, wherein the dual-recording information stream represents the audio and video obtained by real-time recording of the conversation between the salesperson and the customer;
[0102] Step S202: Preprocessing the dual-recording information stream to obtain audio segments and image frames, performing visual navigation classification on the image frames to obtain classified images, and then performing preliminary detection on the audio segments and image frames through a local quality inspection service to obtain a first detection result. The preprocessing includes separating the audio and video, transcoding the audio and video, and extracting frames from the audio and video. The preliminary detection includes black screen detection and / or seat absence detection on the image frames and audio detection on the audio segments.
[0103] Step S203: Perform deep detection processing on the audio segment and the classified images through the cloud-based quality inspection service to obtain a second detection result, and merge the first detection result and the second detection result to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each of the above quality inspection items.
[0104] An embodiment of the present invention provides a processor, which is used to run a program, wherein the dual recording quality inspection method is executed when the program is run.
[0105] Specifically, the dual-recording quality inspection method includes:
[0106] Step S201, obtaining a dual-recording information stream, wherein the dual-recording information stream represents the audio and video obtained by real-time recording of the conversation between the salesperson and the customer;
[0107] Step S202: Preprocessing the dual-recording information stream to obtain audio segments and image frames, performing visual navigation classification on the image frames to obtain classified images, and then performing preliminary detection on the audio segments and image frames through a local quality inspection service to obtain a first detection result. The preprocessing includes separating the audio and video, transcoding the audio and video, and extracting frames from the audio and video. The preliminary detection includes black screen detection and / or seat absence detection on the image frames and audio detection on the audio segments.
[0108] Step S203: Perform deep detection processing on the audio segment and the classified images through the cloud-based quality inspection service to obtain a second detection result, and merge the first detection result and the second detection result to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each of the above quality inspection items.
[0109] An embodiment of the present invention provides an electronic device, comprising a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, at least the following steps are performed:
[0110] Step S201, obtaining a dual-recording information stream, wherein the dual-recording information stream represents the audio and video obtained by real-time recording of the conversation between the salesperson and the customer;
[0111] Step S202: Preprocessing the dual-recording information stream to obtain audio segments and image frames, performing visual navigation classification on the image frames to obtain classified images, and then performing preliminary detection on the audio segments and image frames through a local quality inspection service to obtain a first detection result. The preprocessing includes separating the audio and video, transcoding the audio and video, and extracting frames from the audio and video. The preliminary detection includes black screen detection and / or seat absence detection on the image frames and audio detection on the audio segments.
[0112] Step S203: Perform deep detection processing on the audio segment and the classified images through the cloud-based quality inspection service to obtain a second detection result, and merge the first detection result and the second detection result to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each of the above quality inspection items.
[0113] The devices in this article can be servers, PCs, PADs, mobile phones, etc.
[0114] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program for initializing at least the following method steps:
[0115] Step S201, obtaining a dual-recording information stream, wherein the dual-recording information stream represents the audio and video obtained by real-time recording of the conversation between the salesperson and the customer;
[0116] Step S202: Preprocessing the dual-recording information stream to obtain audio segments and image frames, performing visual navigation classification on the image frames to obtain classified images, and then performing preliminary detection on the audio segments and image frames through a local quality inspection service to obtain a first detection result. The preprocessing includes separating the audio and video, transcoding the audio and video, and extracting frames from the audio and video. The preliminary detection includes black screen detection and / or seat absence detection on the image frames and audio detection on the audio segments.
[0117] Step S203: Perform deep detection processing on the audio segment and the classified images through the cloud-based quality inspection service to obtain a second detection result, and merge the first detection result and the second detection result to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each of the above quality inspection items.
[0118] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0119] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0120] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0121] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0122] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0123] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0124] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0125] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0126] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0127] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A dual recording quality inspection method, characterized in that: include: Acquire a dual-recording information stream, wherein the dual-recording information stream represents audio and video obtained by real-time recording of a conversation between a salesperson and a customer; Performing a preprocessing operation on the dual-recording information stream to obtain audio segments and picture frames, performing visual navigation classification processing on the picture frames to obtain classified pictures, and then performing preliminary detection processing on the audio segments and the picture frames through a local quality inspection service to obtain a first detection result, wherein the preprocessing operation includes separating the audio and video, transcoding the audio and video, and extracting frames of the audio and video, and the preliminary detection processing includes black screen detection and / or seat absence detection of the picture frames and audio detection of the audio segments; The audio segment and the classified image are subjected to deep detection processing through the cloud-based quality inspection service to obtain a second detection result, and the first detection result and the second detection result are merged and processed to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each quality inspection item.
2. The method according to claim 1, characterized in that Performing deep detection processing on the audio segment and the classified image through a cloud-based quality inspection service, including: Perform ASR processing on the audio segment through the cloud quality inspection service to obtain the audio text corresponding to the audio segment, and perform the speech detection, sensitive word detection and emotion detection on the audio text in sequence; The OCR detection is performed on the classified image, wherein the OCR detection includes at least one of identity card information verification, work badge information verification, risk assessment form verification, document verification and valet operation analysis.
3. The method according to claim 1, characterized in that Performing preprocessing operations on the dual recording information stream, including: Using OpenCV technology to perform the audio and video separation operation on the dual-recording information stream to obtain a video stream and an audio stream; The transcoding operation is performed on the audio stream to obtain the plurality of audio segments, and frame extraction is performed on the video stream according to a preset time period to obtain the plurality of picture frames.
4. The method according to claim 1, wherein In the process of performing preliminary detection processing on the audio segment and the picture frame by the local end quality inspection service, the method further includes: Extracting the picture frames at a preset time step to perform the black screen detection, determining that the picture frame identified as a black screen picture fails the black screen detection, and ending the black screen detection of the video stream if the picture frames extracted consecutively for a preset number of times are the black screen pictures, wherein the video stream is composed of all the picture frames; Performing the absence detection on the picture frame that passes the black screen detection; if the number of people in the picture frame is less than a preset number, determining that the picture frame fails the absence detection; and if people continuously leave the table within a preset time period, determining that the video stream fails the absence detection; When it is detected that the average decibel of the audio of the audio segment is lower than the decibel threshold, it is determined that the audio segment fails the audio detection.
5. The method according to claim 2, characterized in that Performing the OCR detection on the classified image includes: identifying risk level information of a risk assessment table display image based on OCR, and determining whether the risk level information matches the product risk level of the insurance product, wherein the classification image includes the risk assessment table display image; If the risk level information matches the product risk level, determining that the risk assessment table display diagram passes the OCR test; In a case where the risk level information does not match the product risk level, it is determined that the risk assessment table display diagram fails the OCR test.
6. The method according to claim 1, characterized in that The classification pictures include: the salesperson's work badge confirmation picture, the customer's ID card confirmation picture, the risk assessment form display picture, the insurance application form reading picture, the insurance terms introduction picture, the insurance product manual reading picture, the disclaimer reading picture and at least one of the signature and display pictures.
7. A dual recording quality inspection system, characterized in that: include: A control terminal, the control terminal being used to execute the dual recording quality inspection method according to any one of claims 1 to 6; A dual recording system, the dual recording system is used to collect dual recording information streams, and the dual recording system is deployed on the control terminal; A local-end quality inspection service, the local-end quality inspection service being used to perform preliminary inspection processing on the dual-recording information flow, the local-end quality inspection service being deployed on the control terminal, and the local-end quality inspection service being connected to the dual-recording system in a communication manner based on WebSocket technology; A cloud-based quality inspection service is used to perform deep detection processing on classified audio and classified images, and the cloud-based quality inspection service is communicatively connected to the local quality inspection service.
8. A dual recording quality inspection device, characterized in that: include: an acquisition unit, configured to acquire a dual-recording information stream, wherein the dual-recording information stream represents audio and video acquired by real-time recording of a conversation between a salesperson and a customer; a preliminary detection processing unit, configured to perform a preprocessing operation on the dual-recording information stream to obtain audio segments and picture frames, perform visual navigation classification processing on the picture frames to obtain classified pictures, and then perform preliminary detection processing on the audio segments and the picture frames through a local quality inspection service to obtain a first detection result, wherein the preprocessing operation includes audio and video separation, audio and video transcoding, and audio and video frame extraction, and the preliminary detection processing includes black screen detection and / or seat absence detection of the picture frames and audio detection of the audio segments; A deep detection processing unit is used to perform deep detection processing on the audio segment and the classified image through a cloud-based quality inspection service to obtain a second detection result, and merge the first detection result and the second detection result to obtain a dual-recording quality inspection report corresponding to the dual-recording information stream, wherein the deep detection processing includes at least one of OCR detection, behavior analysis detection, ASR processing and speech detection, and the dual-recording quality inspection report includes whether each quality inspection item has passed, the reason for failing the quality inspection item, and the time anchor point corresponding to each quality inspection item.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the dual-recording quality inspection method according to any one of claims 1 to 6.
10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for executing the double-recording quality inspection method described in any one of claims 1 to 6.
Citation Information
Cited By
Double-recording video similar partition intelligent detection method and system
CN121661569A