Quality inspection method, quality inspection device and processor for dual-recording video data
By performing audio separation and video clustering analysis on dual-recorded video data, voice navigation and video navigation lists are generated. By integrating the quality inspection results, the problem of low accuracy in the quality inspection of dual-recorded video data in existing technologies is solved, and efficient and accurate quality inspection results are achieved.
Patent Information
- Application Number
- CN202211595536.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-13
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2042-12-13
AI Technical Summary
The accuracy of quality inspection of dual-recorded video data in existing technologies is low, which makes it impossible to guarantee the validity of compliance verification results.
By separating the audio from the dual-recorded video data, target audio data and video data are obtained. The audio data is quality inspected using preset dialogue content to generate a voice navigation list, and the video data is clustered to generate a video navigation list. Finally, the results of the two are integrated to improve the accuracy of quality inspection.
This approach improves the accuracy and efficiency of dual-recorded video data quality inspection during a highly efficient quality inspection process, ensuring the accuracy and completeness of the inspection results.
Smart Images

Figure CN116012750B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a double recording video data quality inspection method and device, computer readable storage medium and processor. BACKGROUND
[0002] With the increasingly stringent regulatory requirements, the workload of compliance detection on the recording video data of financial products in the sales process is also increasingly heavy. The traditional manual auditing method requires the staff to analyze the quality check points in various businesses, then completely browse the double recording video and judge whether each quality check point meets the standard requirements. This method not only needs to consume a lot of manpower and material resources, but also the video auditing standards are different for different people, so the effectiveness of the compliance verification result cannot be guaranteed.
[0003] With the increasing maturity of artificial intelligence technology and the rapid development of computer hardware, the application of artificial intelligence technology in the financial industry is accelerating, making it possible to use voice recognition, intelligent character recognition and natural language processing technologies to automatically detect the compliance of double recording data. In the current intelligent quality inspection technical solution of double recording data, although there are quality inspection solutions using artificial intelligence technology, for example, using voice recognition technology to obtain text information from double recording audio data, and performing dialogue matching detection to judge whether it meets the requirements, and using computer vision technology to identify personnel off-site, signature behavior and other quality inspection points.
[0004] The above intelligent quality inspection solutions all have the problems of low quality inspection efficiency and inaccurate quality inspection results. SUMMARY
[0005] The main purpose of the present application is to provide a double recording video data quality inspection method, quality inspection device, computer readable storage medium and processor, to solve the problem of low accuracy of double recording video data quality inspection in the prior art.
[0006] According to an aspect of the embodiments of the present application, a quality inspection method of dual-recording video data is provided, comprising: performing audio separation on acquired dual-recording video data to obtain target audio data and target video data; performing quality inspection on a plurality of quality inspection points by at least using the target audio data and preset speech content to obtain a first quality inspection result and a voice navigation list, the voice navigation list being obtained by marking first timestamp information of part of the quality inspection points on a time axis of the target audio data; performing cluster analysis on the target video data based on part of the quality inspection points to obtain a video navigation list, and performing quality inspection on part of the quality inspection points based on the voice navigation list and the video navigation list to obtain a second quality inspection result, the video navigation list being obtained by marking second timestamp information of part of the quality inspection points on a time axis of the target video data, the quality inspection points including a certificate quality inspection point, a behavior quality inspection point, a speech quality inspection point, a sensitive word quality inspection point, and a personnel off-site quality inspection point; integrating the first quality inspection result and the second quality inspection result to obtain a target quality inspection result, and sending the target quality inspection result to a display screen of a terminal device to enable the display screen to display the target quality inspection result.
[0007] Optionally, performing quality inspection on a plurality of quality inspection points by at least using the target audio data and preset speech content to obtain a voice navigation list comprises: performing preprocessing on the target audio data to obtain a plurality of target text paragraphs, the preprocessing being used to convert the target audio data into the plurality of target text paragraphs; performing sentence-by-sentence analysis on each of the target text paragraphs based on preset certificate quality inspection points and preset behavior quality inspection points in the preset speech content to obtain the first timestamp information of the certificate quality inspection points and the behavior quality inspection points in the target audio data, the first timestamp information being start appearance time of the certificate quality inspection points and the behavior quality inspection points in the target audio data; and marking the time axis of the target audio data based on each of the first timestamp information to obtain the voice navigation list.
[0008] Optionally, performing quality inspection on a plurality of quality inspection points by at least using the target audio data and preset speech content to obtain a first quality inspection result comprises: performing preprocessing on the target audio data to obtain a plurality of target text paragraphs, the preprocessing being used to convert the target audio data into the plurality of target text paragraphs; obtaining a speech quality inspection result based on preset speech quality inspection points in the preset speech content and the plurality of target text paragraphs, and obtaining a sensitive word quality inspection result based on a preset sensitive word library and the plurality of target text paragraphs; and constituting the first quality inspection result by the speech quality inspection result and the sensitive word quality inspection result.
[0009] Optionally, the target audio data is preprocessed to obtain a plurality of target text paragraphs, including: performing speech recognition processing on the target audio data to obtain target text information corresponding to the target audio data; and performing semantic processing on the target text information to obtain a plurality of target text paragraphs.
[0010] Optionally, based on the preset script content and the plurality of target text paragraphs, a script quality inspection result is obtained, including: based on the preset script quality inspection point, performing sentence-by-sentence analysis on the plurality of target text paragraphs to obtain a plurality of alternative hit sentences; calculating the similarity scores of the preset script quality inspection point and each of the alternative hit sentences to obtain a plurality of preset similarity scores; and calculating the sum of the plurality of preset similarity scores to obtain a target similarity score of the double recording video data, wherein the target similarity score constitutes the script quality inspection result.
[0011] Optionally, calculating the similarity scores of the preset script quality inspection point and each of the alternative hit sentences to obtain a plurality of preset similarity scores includes: performing word segmentation processing on the plurality of alternative hit sentences to obtain a plurality of target keywords, wherein one of the alternative hit sentences corresponds to at least one of the target keywords; using each of the target keywords to traverse the sentence corresponding to the preset script quality inspection point to obtain a plurality of keyword matching degrees, and using each of the alternative hit sentences to traverse the sentence corresponding to the preset script quality inspection point to obtain a plurality of text distance matching degrees, wherein one target keyword corresponds to a plurality of keyword matching degrees, and one alternative hit sentence corresponds to a plurality of text distance matching degrees; and determining the corresponding preset similarity score based on the keyword matching degrees and the text distance matching degrees corresponding to each of the alternative hit sentences.
[0012] Optionally, based on a preset sensitive word library and the plurality of target text paragraphs, a sensitive word quality inspection result is obtained, including: comparing each preset sensitive word in the preset sensitive word library with the plurality of target text paragraphs; and in a case where a target text identical to the preset sensitive word exists in the target text paragraph, determining the target text paragraph to which the target text belongs as a sensitive word paragraph, and constituting the sensitive word quality inspection result by the sensitive word paragraph.
[0013] Optionally, the target video data is analyzed based on the partial quality inspection points to obtain a video navigation list, including: performing frame extraction processing on a plurality of video frames in the target video data to obtain a plurality of first target video frames; performing clustering analysis on the plurality of first target video frames to obtain the second timestamp information of the certificate quality inspection point and the behavior quality inspection point, the second timestamp information being the start time of the certificate quality inspection point and the behavior quality inspection point in the target video data; and marking each second timestamp information in the target video data to obtain the video navigation list.
[0014] Optionally, the target video data is analyzed based on the partial quality inspection points to obtain a video navigation list, including: performing frame extraction processing on a plurality of video frames in the target video data to obtain a plurality of first target video frames; performing clustering analysis on the plurality of first target video frames to obtain the second timestamp information of the certificate quality inspection point and the behavior quality inspection point, the second timestamp information being the start time of the certificate quality inspection point and the behavior quality inspection point in the target video data; and marking each second timestamp information in the target video data to obtain the video navigation list.
[0015] Optionally, the quality inspection method further includes: performing frame extraction processing on the target video data with target timestamp information to determine a plurality of third target video frames corresponding to a personnel off-site quality inspection point; performing quality inspection on the personnel off-site quality inspection point based on a first target network model and the plurality of third target video frames to obtain the target hit image corresponding to the personnel off-site quality inspection point and the timestamp information corresponding to the target hit image, the first target network model being obtained by constructing and training a neural network model; and constructing the second quality inspection result based on the target hit image corresponding to the personnel off-site quality inspection point and the corresponding timestamp information.
[0016] Optionally, based on the plurality of second target video frames corresponding to the certificate inspection point, the certificate inspection point is inspected to obtain a target hit image corresponding to the certificate inspection point and time stamp information corresponding to the target hit image, comprising: adopting a target combination algorithm, preprocessing the plurality of second target video frames corresponding to the certificate inspection point to obtain a plurality of preprocessed second target video frames, the target combination algorithm being an algorithm obtained by combining Faster-RCNN, CNN and a traditional image processing algorithm; adopting a CTPN network, performing text detection on the plurality of preprocessed second target video frames to obtain a plurality of preset text sequences, and adopting a connection time sequence classification model, processing each of the preset text sequences to obtain a plurality of credible text sequences; determining an image corresponding to the second target video frame with the most character information in each of the credible text sequences as the target hit image, and determining time stamp information of the image corresponding to the second target video frame with the most character information in each of the credible text sequences as the time stamp information corresponding to the target hit image.
[0017] Optionally, based on the plurality of second target video frames corresponding to the behavior inspection point, the behavior inspection point is inspected to obtain the target hit image corresponding to the behavior inspection point and the time stamp information corresponding to the target hit image, comprising: based on a second target network model and the plurality of second target video frames, the behavior inspection point is inspected to obtain the target hit image corresponding to the behavior inspection point and the time stamp information corresponding to the target hit image, the second target network model being obtained by constructing and training a neural network model.
[0018] According to another aspect of the embodiments of the present application, there is also provided a quality inspection device for dual-recording video data, comprising: an audio-video separation component configured to separate audio from acquired dual-recording video data to obtain target audio data and target video data; a speech recognition component configured to perform quality inspection on a plurality of quality inspection points by using at least the target audio data and preset dialogue content to obtain a first quality inspection result and a speech navigation list, wherein the speech navigation list is obtained by marking first timestamp information of part of the quality inspection points on a time axis of the target audio data; a video frame extraction component configured to perform cluster analysis on the target video data based on part of the quality inspection points to obtain a video navigation list, and perform quality inspection on part of the quality inspection points based on the speech navigation list and the video navigation list to obtain a second quality inspection result, wherein the video navigation list is obtained by marking second timestamp information of part of the quality inspection points on a time axis of the target video data, and the quality inspection points include a certificate quality inspection point, a behavior quality inspection point, a dialogue quality inspection point, a sensitive word quality inspection point and a personnel off-site quality inspection point; and an integration component configured to integrate the first quality inspection result and the second quality inspection result to obtain a target quality inspection result, and send the target quality inspection result to a display screen of a terminal device to enable the display screen to display the target quality inspection result.
[0019] According to still another aspect of the embodiments of the present application, there is also provided a computer readable storage medium comprising a stored program, wherein the program performs any of the quality inspection methods for dual-recording video data.
[0020] According to yet another aspect of the embodiments of the present application, there is also provided a processor configured to run a program, wherein the program performs any of the quality inspection methods for dual-recording video data when running.
[0021] In this embodiment of the invention, the quality inspection method for dual-recorded video data firstly uses preset dialogue content to inspect each quality inspection point in the target audio data obtained from the separation of dual-recorded video data, obtaining a first quality inspection result and a voice navigation list; then, based on some quality inspection points, cluster analysis is performed on the target video data obtained from the separation of dual-recorded video data to obtain a video navigation list, and based on the voice navigation list and video navigation category, some quality inspection points are inspected, resulting in a second quality inspection result; finally, the first and second quality inspection results are integrated to obtain a target quality inspection result, which is then sent to the display screen of the terminal device for display. Compared with the prior art, which uses manual quality inspection of dual-recorded video data or relies on target audio data to inspect each quality inspection point to obtain a target quality inspection result, this solution inspects each quality inspection point in the target audio data of the dual-recorded video data to obtain a first quality inspection result and a voice navigation list. Based on some quality inspection points, cluster analysis is performed on the target video data to obtain a video navigation list. Based on the voice navigation list and the video navigation list, some quality inspection points are inspected to obtain a second quality inspection result. Based on the first and second quality inspection results, the target quality inspection result is obtained. This ensures that while the quality inspection of dual-recorded video data is performed efficiently, the quality inspection of each quality inspection point is also performed accurately, ensuring that the target quality inspection result is relatively accurate. This solves the problem of low accuracy in the quality inspection of dual-recorded video data in the prior art, and thus ensures high efficiency in the quality inspection of dual-recorded video data. Attached Figure Description
[0022] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0023] Figure 1 A flowchart illustrating a quality inspection method for dual-recorded video data according to an embodiment of this application is shown;
[0024] Figure 2 This application illustrates a flowchart of a document quality inspection point process according to one embodiment of the present application;
[0025] Figure 3 This invention provides a schematic diagram of the structure of a dual-recording video data quality inspection device according to an embodiment of the present application.
[0026] Figure 4 A schematic diagram of the structure of a dual-recording video data quality inspection device according to another embodiment of this application is shown;
[0027] Figure 5 A schematic diagram of the architecture of a workflow engine according to an embodiment of this application is shown;
[0028] Figure 6 A schematic diagram of containerized deployment of an embodiment of the present application is shown;
[0029] Figure 7 A flowchart of a quality inspection method of dual-recording video data of a specific embodiment of the present application is shown. DETAILED DESCRIPTION
[0030] It should be noted that the embodiments and features of the embodiments in the present application can be combined with each other without conflict. The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0031] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0032] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0033] For the convenience of description, the following describes some nouns or terms related to the embodiments of the present application:
[0034] Dual-recording video data: when a financial business institution sells self-owned financial products, personal structured deposit products, agent-sold products and agented precious metals and company opening account services, it records the sales process of each product in the sales area for management;
[0035] Dual-recording intelligent quality inspection: by using computer audio and video processing technology, combined with artificial intelligence technologies such as speech recognition, intelligent character recognition and natural language processing, the dual-recording video data is automatically checked for compliance;
[0036] Automatic Speech Recognition (ASR): Automatic speech recognition or Speech to Text (STT) is a technology that automatically converts human speech content into corresponding text by computer;
[0037] Intelligent Character Recognition (ICR): Based on Optical Character Recognition (OCR), the artificial intelligence technology of computer deep learning is introduced. Semantic reasoning and semantic analysis are adopted, and the character information of the un-recognized part can be completed according to the context sentence information, which is an improved technology of optical character recognition;
[0038] Natural Language Processing (NLP): It is a field of computer science, artificial intelligence and linguistics focusing on the interaction between computer and human (natural) language. It studies various theories and methods that can realize effective communication between people and computers using natural language, including natural language understanding and natural language generation.
[0039] Face recognition: A biometric technology based on facial feature information for identity recognition. It uses a camera or camera to collect images or video streams containing human faces, automatically detects and tracks faces in images, and then performs a series of related technologies for face recognition on the detected faces.
[0040] Behavior recognition: It refers to the use of computer image visual analysis technology combined with artificial intelligence deep learning technology to analyze the behavior being performed from the video, and to identify the pre-defined behavior from the image or video.
[0041] As mentioned in the background, the accuracy of quality inspection of double recording video data in the prior art is low. In order to solve the above problem, in a typical embodiment of the present application, a quality inspection method, quality inspection device, computer readable storage medium and processor for double recording video data are provided.
[0042] According to the embodiments of the present application, a quality inspection method for double recording video data is provided.
[0043] Figure 1 is a flowchart of the quality inspection method for double recording video data according to the embodiments of the present application. As shown in Figure 1 the quality inspection method includes the following steps:
[0044] Step S101, audio separation is performed on the obtained dual-recording video data to obtain target audio data and target video data;
[0045] Step S102, at least using the target audio data and a preset script content, quality inspection is performed on a plurality of quality inspection points to obtain a first quality inspection result and a voice navigation list, the voice navigation list being obtained by marking first timestamp information of part of the quality inspection points on a time axis of the target audio data;
[0046] Step S103, based on part of the quality inspection points, cluster analysis is performed on the target video data to obtain a video navigation list, and based on the voice navigation list and the video navigation list, quality inspection is performed on part of the quality inspection points to obtain a second quality inspection result, the video navigation list being obtained by marking second timestamp information of part of the quality inspection points on a time axis of the target video data, and the quality inspection points including a certificate quality inspection point, a behavior quality inspection point, a script quality inspection point, a sensitive word quality inspection point and a personnel off-seat quality inspection point;
[0047] Step S104, the first quality inspection result and the second quality inspection result are integrated to obtain a target quality inspection result, and the target quality inspection result is sent to a display screen of a terminal device to enable the display screen to display the target quality inspection result.
[0048] In the above double-recording video data quality inspection method, first, the preset speech content is used to inspect each quality inspection point in the target audio data separated from the double-recording video data, to obtain a first quality inspection result and a voice navigation list; then, based on part of the quality inspection points, the target video data separated from the double-recording video data is subjected to cluster analysis to obtain a video navigation list, and part of the quality inspection points are inspected based on the voice navigation list and the video navigation list to obtain a second quality inspection result; finally, the first quality inspection result and the second quality inspection result are integrated to obtain a target quality inspection result, and the target quality inspection result is sent to a display screen of a terminal device to enable the display screen to display the target quality inspection result. Compared with the prior art in which the double-recording video data is inspected by artificial inspection or the quality inspection points are inspected by relying on the target audio data, in the present scheme, the target audio data in the double-recording video data is inspected to obtain the first quality inspection result and the voice navigation list. Then, based on part of the quality inspection points, the target video data is subjected to cluster analysis to obtain the video navigation list, and part of the quality inspection points are inspected based on the voice navigation list and the video navigation list to obtain the second quality inspection result. Based on the first quality inspection result and the second quality inspection result, the target quality inspection result is obtained, which ensures that the quality inspection points are accurately inspected on the basis of efficient quality inspection of the double-recording video data, and the target quality inspection result is accurate, thereby solving the problem of low accuracy of quality inspection of the double-recording video data in the prior art, and further ensuring high efficiency of quality inspection of the double-recording video data.
[0049] Specifically, in the process of product sales by a salesperson, the entire sales process is recorded by audio and video. In actual application, the audio or video recording equipment may fail. In the case of failure of the audio or video recording equipment, complete double-recording video data is difficult to collect. Compared with the prior art in which the quality inspection points are inspected by relying on the audio data, in the quality inspection method of the present application, the preset speech content is used to inspect each quality inspection point in the target audio data separated from the double-recording video data, to obtain a first quality inspection result and a voice navigation list. Based on part of the quality inspection points, the target video data separated from the double-recording video data is subjected to cluster analysis to obtain a video navigation list, and part of the quality inspection points are inspected based on the voice navigation list and the video navigation list to obtain a second quality inspection result. Finally, based on the first quality inspection result and the second quality inspection result, a target quality inspection result is obtained. That is, the present scheme realizes that part of the quality inspection points can be inspected even if the target audio data or the target video data is missing. At the same time, based on the voice navigation list and the video navigation list, the second quality inspection result can be accurately obtained, thereby ensuring that the target quality inspection result of the present application is accurate.
[0050] In a specific embodiment of the present application, the aforementioned certificate quality inspection points can include: ID card recognition, post qualification certificate recognition, name tag recognition, risk assessment form recognition, signature form recognition, and the like. The aforementioned behavior quality inspection points can include: signature behavior recognition. The dialogue quality inspection point can be the similarity between the dialogue of the salesperson and the preset dialogue content. The sensitive word quality inspection point can be whether the dialogue of the salesperson involves relatively sensitive words. The personnel leaving quality inspection point can be used to detect whether the salesperson or the customer leaves during the sales process.
[0051] Specifically, the aforementioned preset dialogue content can be obtained by matching in the preset dialogue library according to the nature or type of the product being sold.
[0052] In actual application, any feasible clustering method in the prior art can be used to perform clustering analysis on the target video data based on the partial quality inspection points. The specific clustering method is not limited in the present application, and can be flexibly adjusted according to actual use.
[0053] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0054] In order to subsequently obtain the second quality inspection result based on the voice navigation list and the video navigation list, more accurately, and further ensure that the obtained target quality inspection result is more accurate, in an embodiment of the present application, at least the target audio data and the preset dialogue content are used to perform quality inspection on a plurality of quality inspection points to obtain a voice navigation list, including: preprocessing the aforementioned target audio data to obtain a plurality of target text paragraphs, the aforementioned preprocessing being used to convert the aforementioned target audio data into a plurality of aforementioned target text paragraphs; based on the preset certificate quality inspection points and the preset behavior quality inspection points in the aforementioned preset dialogue content, performing sentence-by-sentence analysis on each of the aforementioned target text paragraphs to obtain the aforementioned first timestamp information of the aforementioned certificate quality inspection points and the aforementioned behavior quality inspection points in the aforementioned target audio data, the aforementioned first timestamp information being the start time of the aforementioned certificate quality inspection points and the aforementioned behavior quality inspection points in the aforementioned target audio data; based on each of the aforementioned first timestamp information, marking the time axis of the aforementioned target audio data to obtain the aforementioned voice navigation list.
[0055] In actual application, the timestamps of the appearance of each inspection point, i.e., the first timestamp information, can be obtained based on the preset inspection points of the certificate and the plurality of target text fields obtained by performing speech recognition on the target audio data. That is, the voice navigation list is composed of the timestamps of the appearance of each inspection point (i.e., the first timestamp information) and the corresponding inspection points. Of course, in order to ensure the integrity of the appearance of each inspection point, in the case where the timestamp of the first appearance of each inspection point is obtained, the timestamp of the first appearance of each inspection point can be taken as the timestamp of the appearance of each inspection point (the first timestamp information) within a certain period of time (for example, 2 seconds or 3 seconds before the timestamp of the first appearance of each inspection point, etc.) according to actual conditions.
[0056] Specifically, the inspection points of the certificate in the voice navigation list include identification of an ID card, identification of a post qualification card, identification of a work card, identification of a risk assessment form, and identification of a signature form; and the behavior inspection points include identification of a signature behavior. The specific judgment process is as follows: when "this is my work card" is identified in the plurality of target text fields after speech recognition, a work card identification inspection point is generated; when "this is my post qualification card" is identified in the plurality of target text fields, a post qualification card inspection point is generated; when "show your ID card" is identified in the plurality of target text fields, an ID card inspection point is generated; when "show your risk assessment form" is identified in the plurality of target text fields, a risk assessment form inspection point is generated; when "please sign" is identified in the plurality of target text fields, a signature behavior inspection point is generated; and when "show your signature form" is identified in the plurality of target text fields, a signature form inspection point is generated.
[0057] In another embodiment of the present application, at least the target audio data and the preset script content are used to inspect the plurality of inspection points to obtain a first inspection result, including: preprocessing the target audio data to obtain a plurality of target text fields, wherein the preprocessing is used to convert the target audio data into the plurality of target text fields; obtaining a script inspection result based on the preset script inspection points in the preset script content and the plurality of target text fields, and obtaining a sensitive word inspection result based on a preset sensitive word library and the plurality of target text fields; and the first inspection result is composed of the script inspection result and the sensitive word inspection result. In this embodiment, the script inspection result is obtained based on the preset script inspection points in the preset script content and the plurality of target text fields, and the sensitive word inspection result is obtained based on the preset sensitive word library and the plurality of target text fields, so that the obtained first inspection result is more accurate, and the behavior of the salesperson is further standardized.
[0058] In order to obtain the target text paragraphs more accurately, in another embodiment of the present application, the target audio data is preprocessed to obtain the target text paragraphs, including: performing speech recognition processing on the target audio data to obtain target text information corresponding to the target audio data; and performing semantic processing on the target text information to obtain the target text paragraphs, so as to ensure that the obtained target text paragraphs are reasonable and the speech recognition effect is good.
[0059] In another embodiment of the present application, based on the preset dialogue quality inspection points in the preset dialogue content and the target text paragraphs, a dialogue quality inspection result is obtained, including: based on the preset dialogue quality inspection points, performing sentence-by-sentence analysis on the target text paragraphs to obtain a plurality of candidate hit sentences; calculating the similarity scores of the preset dialogue quality inspection points and each of the candidate hit sentences to obtain a plurality of preset similarity scores; calculating the sum of the preset similarity scores to obtain a target similarity score of the double-recording video data, and since the target similarity score constitutes the dialogue quality inspection result. Specifically, in this embodiment, based on the preset dialogue quality inspection points, the target text paragraphs are analyzed sentence by sentence to obtain a plurality of candidate hit sentences, the similarity scores of the preset dialogue quality inspection points and the candidate hit sentences are calculated to obtain a plurality of preset similarity scores, and since the sum of the plurality of preset similarity scores constitutes the target similarity score, it is ensured that the target similarity score of the double-recording video data can be obtained more accurately and simply with less calculation, the overall audio data of the salesperson is scored more accurately, and further, the double-recording video data is quality inspected more efficiently.
[0060] In order to obtain the preset similarity scores more accurately and further obtain the target similarity score more accurately, in an embodiment of the present application, the similarity scores of the preset dialogue quality inspection points and each of the candidate hit sentences are calculated to obtain a plurality of preset similarity scores, including: performing word segmentation processing on the plurality of candidate hit sentences to obtain a plurality of target keywords, wherein one of the candidate hit sentences corresponds to at least one of the target keywords; using each of the target keywords to traverse the sentence corresponding to the preset dialogue quality inspection point to obtain a plurality of keyword matching degrees, and using each of the candidate hit sentences to traverse the sentence corresponding to the preset dialogue quality inspection point to obtain a plurality of text distance matching degrees, one target keyword corresponding to a plurality of keyword matching degrees, and one candidate hit sentence corresponding to a plurality of text distance matching degrees; based on the keyword matching degrees and the text distance matching degrees corresponding to each of the candidate hit sentences, determining the corresponding preset similarity scores.
[0061] Specifically, in the above embodiment, the process of calculating the text distance matching degree can be: determining the longest common subsequence (LCS) of the candidate hit sentence and the sentence corresponding to the preset script quality inspection point. Then, based on the longest common subsequence, the text distance matching degree is obtained.
[0062] Specifically, the keyword matching degree and the text distance matching degree can be weighted to obtain a preset similarity score. In a specific embodiment of the present application, the preset similarity score = a*keyword matching degree + b*text distance matching degree. Wherein, a and b can be debugged according to the double recording video data of different regions, wherein a<=1, b<=1, a+b=1.
[0063] In another embodiment of the present application, based on the preset sensitive word library and the plurality of target text paragraphs, a sensitive word quality inspection result is obtained, including: comparing each preset sensitive word in the above-mentioned preset sensitive word library with the plurality of above-mentioned target text paragraphs respectively; in the case that there is a target text in the above-mentioned target text paragraph which is the same as the above-mentioned preset sensitive word, the above-mentioned target text paragraph to which the above-mentioned target text belongs is determined as a sensitive word paragraph, and the above-mentioned sensitive word paragraph constitutes the above-mentioned sensitive word quality inspection result, and subsequently, the sensitive word quality inspection result and the script quality inspection result constitute the first quality inspection result, which further ensures that the first quality inspection result is more accurate.
[0064] In actual application process, in order to further reduce the calculation amount of quality inspection of double recording video data, in another embodiment of the present application, based on part of the above-mentioned quality inspection points, the above-mentioned target video data is clustered and analyzed to obtain a video navigation list, including: frame extraction processing is performed on a plurality of video frames in the above-mentioned target video data to obtain a plurality of first target video frames; clustering analysis is performed on a plurality of the above-mentioned first target video frames to obtain the above-mentioned second timestamp information of the above-mentioned certificate quality inspection point and the above-mentioned behavior quality inspection point, the above-mentioned second timestamp information being the start time of the above-mentioned certificate quality inspection point and the above-mentioned behavior quality inspection point in the above-mentioned target video data; each of the above-mentioned second timestamp information is marked in the above-mentioned target video data to obtain the above-mentioned video navigation list.
[0065] Specifically, in the above-mentioned embodiment, the entire target video data can be frame extraction processed according to an average of 4 frames per second.
[0066] Specifically, in the above embodiment, the certificate inspection points in the video navigation list include ID card recognition, post qualification card recognition, work card recognition, risk assessment form recognition, and signature form recognition; the behavior inspection points include signature behavior recognition. The specific judgment process is: according to the mode of extracting 4 frames per second, frame extraction processing is performed on a plurality of video frame images in the target video image to obtain a plurality of first target video frames. And based on the plurality of first target video frames, clustering processing is performed, and according to the identity card recognition, post qualification card recognition, work card recognition, risk assessment form recognition, signature form recognition and signature behavior recognition in the clustering result, the time stamp information (i.e. second time stamp information) of the corresponding inspection point is generated. In the actual application process, in order to ensure the integrity of the appearance of each inspection point, in the case of obtaining the time stamp of the first appearance of each inspection point, the time stamp of the first appearance of each inspection point can also be taken as the time stamp (second time stamp information) of the appearance of each inspection point according to the actual situation (for example, 2 seconds or 3 seconds before the time stamp of the first appearance of each inspection point, etc.).
[0067] In another embodiment of the present application, based on the above-mentioned voice navigation list and the above-mentioned video navigation list, the above-mentioned inspection points of some points are inspected to obtain a second inspection result, including: merging the above-mentioned voice navigation list and the above-mentioned video navigation list to obtain the above-mentioned target video data with target time stamp information, the above-mentioned target time stamp information being the earliest one of the above-mentioned first time stamp information and the above-mentioned second time stamp information corresponding to the above-mentioned certificate inspection points and the above-mentioned behavior inspection points; based on the above-mentioned target time stamp information, frame extraction processing is performed on the above-mentioned target video data to determine a plurality of second target video frames corresponding to the above-mentioned certificate inspection points and a plurality of the above-mentioned second target video frames corresponding to the above-mentioned behavior inspection points; based on the plurality of the above-mentioned second target video frames corresponding to the above-mentioned certificate inspection points, the above-mentioned certificate inspection points are inspected to obtain target hit images corresponding to the above-mentioned certificate inspection points and time stamp information corresponding to the above-mentioned target hit images, based on the plurality of the above-mentioned second target video frames corresponding to the above-mentioned behavior inspection points, the above-mentioned behavior inspection points are inspected to obtain the above-mentioned target hit images corresponding to the above-mentioned behavior inspection points and the time stamp information corresponding to the above-mentioned target hit images; the above-mentioned second inspection result is constituted by the above-mentioned target hit images corresponding to the above-mentioned certificate inspection points and the corresponding time stamp information, and the above-mentioned target hit images corresponding to the above-mentioned behavior inspection points and the corresponding time stamp information. In this embodiment, the voice navigation list and the video navigation list are merged and processed, which ensures that the obtained target time stamp information is relatively accurate and ensures that the obtained video of each inspection point is relatively complete. Based on the target time stamp information, frame extraction processing is performed on the target video data, which ensures that the obtained plurality of second target video frames is relatively accurate, and ensures that the obtained second inspection result is relatively accurate.
[0068] Specifically, in the above-mentioned embodiment, in the process of merging the voice navigation list and the video navigation list, for the same quality inspection point, the timestamp information earlier (which can be the first timestamp information or the second timestamp information) in the voice navigation list and the video navigation list is taken, so as to ensure that the obtained starting time of each quality inspection point is more accurate and the video of each quality inspection point is more complete. Subsequently, the behavior quality inspection point and the certificate quality inspection point are inspected based on multiple second target videos, so as to further ensure that the obtained second quality inspection result is more accurate.
[0069] In order to more accurately inspect the personnel off-site quality inspection point, in an embodiment of the present application, the above-mentioned quality inspection method further comprises: performing frame extraction processing on the above-mentioned target video data with target timestamp information to determine multiple third target video frames corresponding to the personnel off-site quality inspection point; performing quality inspection on the above-mentioned personnel off-site quality inspection point based on a first target network model and the multiple third target video frames to obtain the target hit image corresponding to the above-mentioned personnel off-site quality inspection point and the timestamp information corresponding to the target hit image, wherein the first target network model is obtained based on a neural network model and is trained; and based on the target hit image corresponding to the above-mentioned personnel off-site quality inspection point and the corresponding timestamp information, the above-mentioned second quality inspection result is constituted.
[0070] Specifically, in the above-mentioned embodiment, frame extraction processing is performed on the target video data with target timestamp information to obtain multiple third target video frames. Then, quality inspection is performed on the personnel off-site quality inspection point based on the first target network model and the multiple third target video frames. That is, the quality inspection method of the present application does not detect all video frames from the timestamp when each quality inspection point appears, but detects multiple third target video frames, so as to ensure that the calculation amount of the first target network model is small and the obtained second quality inspection result is more accurate.
[0071] Specifically, in the above-mentioned embodiment, the target hit image in the above-mentioned embodiment can be an image in which the personnel off-site quality inspection point is clearest.
[0072] In yet another embodiment of the present application, based on the plurality of second target video frames corresponding to the above-mentioned certificate inspection point, the certificate inspection point is inspected to obtain the target hit image corresponding to the certificate inspection point and the timestamp information corresponding to the target hit image, comprising: using a target combination algorithm to preprocess the plurality of second target video frames corresponding to the certificate inspection point to obtain a plurality of preprocessed second target video frames, the target combination algorithm being an algorithm obtained by combining Faster-RCNN, CNN and traditional image processing algorithms; using a CTPN network to detect text from the plurality of preprocessed second target video frames to obtain a plurality of preset text sequences, and using a connection time sequence classification model to process each of the preset text sequences to obtain a plurality of credible text sequences; determining the image corresponding to the second target video frame with the most character information in each of the credible text sequences as the target hit image, and determining the timestamp information of the image corresponding to the second target video frame with the most character information in each of the credible text sequences as the timestamp information corresponding to the target hit image, thereby ensuring that the recognition result of the certificate inspection point is relatively accurate, and further ensuring that the target inspection result in the present scheme is relatively accurate.
[0073] Specifically, in the present scheme, as Figure 2As shown, first, a target combination algorithm based on Faster-RCNN, Convolutional Neural Networks (CNN) and traditional image processing algorithm is used to pre-process multiple second target video frames. Second, a text detection algorithm based on an optimized Connection Text Proposal Network (CTPN) is used to perform edge detection and positioning on the card certificate to be recognized in the second target video frame, and multiple preset text sequences are obtained. That is, the optimized CTPN network is optimized for scene text recognition in dual recording video data, which is different from the image scanned by a traditional scanner or captured by a high-speed scanner. The text recognition algorithm of the present scheme needs to locate the text area in the video frame, and there is a big difference between the text in the dual recording video data and the printed text recognition in the scanned or high-speed scanner image. Compared with other neural network algorithms, the CTPN algorithm is more suitable for this scene text recognition, because the CTPN has the following characteristics: 1) using a vertical anchor that is more suitable for natural scene text detection characteristics (this is because the size of the text is small compared to the object, and by using a vertical anchor that is more suitable for the size of the text, small size text candidate boxes can be detected); 2) introducing RNN to process the sequence characteristics existing in scene text detection, which can connect small size text to obtain text lines; 3) introducing Side-refinement to improve the accuracy of text box boundary detection. For character recognition in video frames, the present scheme uses a method combining convolutional neural network and recurrent neural network to extract character feature information, normalizes a numerical vector into a probability distribution vector through softmax, and uses an exponential function to calculate the probability distribution of the feature. Finally, the Connectionist Temporal Classification (CTC) model is used to process each preset text sequence to obtain multiple reliable text sequences. In the present scheme, the Connectionist Temporal Classification model based on recurrent neural network can not only obtain the most reliable feature at a specific position, but also can correct the feature result by combining context information (such as Riyue Tan and Mingtan), and finally output a reliable text sequence. The image corresponding to the second target video frame with the most text information in the reliable text sequence is determined as the target hit image. The target hit image and the time stamp information corresponding to the target hit image constitute the second quality inspection result.
[0074] Specifically, in the above-mentioned embodiments, character recognition is mainly performed on the worker card image, post qualification certificate image, identity card image, risk assessment table image and signature form image to extract text information in the images.
[0075] In order to obtain the target hit image and the timestamp information corresponding to the target hit image corresponding to the behavior quality inspection point more accurately and simply, in another embodiment of this application, the behavior quality inspection point is inspected based on multiple second target video frames corresponding to the behavior quality inspection point to obtain the target hit image and the timestamp information corresponding to the target hit image. This includes: inspecting the behavior quality inspection point based on a second target network model and multiple second target video frames to obtain the target hit image and the timestamp information corresponding to the target hit image. The second target network model is constructed and trained based on a neural network model.
[0076] Specifically, the second quality inspection result in this application consists of the target hit image and corresponding timestamp information of the behavior quality inspection point, the target hit image and corresponding timestamp information of the document quality inspection point, and the target hit image and corresponding timestamp information of the personnel leaving the seat quality inspection point.
[0077] Specifically, in this application, to improve the accuracy of face recognition technology, an optimized RetinaNet is used as the backbone network, and A-Softmax is used as the loss function to train the facial feature extractor. This not only improves the recognition speed and achieves complete end-to-end recognition, but also significantly improves the recognition accuracy by using a face recognition model trained on approximately seven million face data.
[0078] This application also provides a quality inspection device for dual-recording video data. It should be noted that this device can be used to execute the quality inspection method for dual-recording video data provided in this application. The following describes the quality inspection device for dual-recording video data provided in this application.
[0079] Figure 3 This is a schematic diagram of the structure of a dual-recording video data quality inspection device according to an embodiment of this application. Figure 3 As shown, the quality inspection device includes:
[0080] The audio-video separation component 10 is used to separate the audio from the acquired dual-recorded video data to obtain target audio data and target video data.
[0081] The speech recognition component 20 is used to perform quality inspection on multiple quality inspection points using at least target audio data and preset speech content, and to obtain a first quality inspection result and a voice navigation list. The voice navigation list is obtained by marking the first timestamp information of some of the above quality inspection points on the timeline of the above target audio data.
[0082] The video frame extraction component 30 is configured to perform clustering analysis on the target video data based on part of the aforementioned quality inspection points to obtain a video navigation list, and perform quality inspection on part of the aforementioned quality inspection points based on the voice navigation list and the video navigation list to obtain a second quality inspection result. The video navigation list is obtained by marking the second timestamp information of part of the aforementioned quality inspection points on the time axis of the target video data. The aforementioned quality inspection points include the certificate quality inspection point, the behavior quality inspection point, the dialogue quality inspection point, the sensitive word quality inspection point, and the personnel off-site quality inspection point.
[0083] The integration component 40 is configured to integrate the first quality inspection result and the second quality inspection result to obtain a target quality inspection result, and send the target quality inspection result to the display screen of the terminal device to enable the display screen to display the target quality inspection result.
[0084] In the aforementioned quality inspection device for double-recording video data, the voice recognition component is configured to perform quality inspection on each quality inspection point in the target audio data separated from the double-recording video data by using the preset dialogue content to obtain a first quality inspection result and a voice navigation list. The video frame extraction component is configured to perform clustering analysis on the target video data separated from the double-recording video data based on part of the quality inspection points to obtain a video navigation list, and perform quality inspection on part of the quality inspection points based on the voice navigation list and the video navigation list to obtain a second quality inspection result. The integration component is configured to integrate the first quality inspection result and the second quality inspection result to obtain a target quality inspection result, and send the target quality inspection result to the display screen of the terminal device to enable the display screen to display the target quality inspection result. Compared with the prior art in which manual quality inspection is performed on the double-recording video data or quality inspection is performed on each quality inspection point by relying on the target audio data to obtain a target quality inspection result, in the present solution, the target audio data in the double-recording video data is used to perform quality inspection on each quality inspection point to obtain a first quality inspection result and a voice navigation list. Then, clustering analysis is performed on the target video data based on part of the quality inspection points to obtain a video navigation list, and quality inspection is performed on part of the quality inspection points based on the voice navigation list and the video navigation list to obtain a second quality inspection result. Based on the first quality inspection result and the second quality inspection result, a target quality inspection result is obtained, which ensures that each quality inspection point is accurately inspected on the basis of efficient quality inspection of the double-recording video data, and the target quality inspection result obtained is relatively accurate, thereby solving the problem of low accuracy of quality inspection of the double-recording video data in the prior art, and further ensuring high efficiency of quality inspection of the double-recording video data.
[0085] Specifically, in the process of the sales personnel selling the product, the entire sales process is recorded by audio and video. In actual application process, the audio or video recording equipment may fail. In the case of failure of the audio or video recording equipment, it is difficult to collect complete double recording video data. Compared with the quality inspection of the prior art which relies on audio data to perform quality inspection at each quality inspection point, in the quality inspection method of the present application, the preset speech content is used to perform quality inspection on each quality inspection point in the target audio data separated from the double recording video data, to obtain a first quality inspection result and a voice navigation list. Based on part of the quality inspection points, the target video data separated from the double recording video data is subjected to cluster analysis to obtain a video navigation list, and part of the quality inspection points are subjected to quality inspection based on the voice navigation list and the video navigation list, to obtain a second quality inspection result. Finally, based on the first quality inspection result and the second quality inspection result, a target quality inspection result is obtained. That is, the present scheme realizes that even in the case of missing target audio data or target video data, quality inspection can be performed on part of the quality inspection points. At the same time, based on the voice navigation list and the video navigation list, the second quality inspection result can be obtained more accurately, thereby ensuring that the target quality inspection result of the present application is more accurate.
[0086] In a specific embodiment of the present application, the above-mentioned certificate quality inspection points can include: identification of an ID card, identification of a post qualification certificate, identification of a work card, identification of a risk assessment form, identification of a signature form, etc. The above-mentioned behavior quality inspection points can include: signature behavior identification. The speech quality inspection point can be the similarity between the speech of the sales personnel and the preset speech content. The sensitive word quality inspection point can be whether the speech of the sales personnel involves relatively sensitive words. The personnel leaving quality inspection point can be used to detect whether the sales personnel or the customer leaves during the process of selling the product.
[0087] Specifically, the above-mentioned preset speech content can be obtained by matching in a preset speech library according to the nature or type of the product being sold.
[0088] In actual application process, any feasible clustering method in the prior art can be used to perform cluster analysis on the target video data based on part of the quality inspection points. In the present application, the specific method of clustering is not limited, and it can be flexibly adjusted according to actual use.
[0089] In order to subsequently obtain the second quality inspection result more accurately based on the voice navigation list and the video navigation list, and further ensure that the target quality inspection result obtained is more accurate, in an embodiment of the present application, as shown in Figure 4As shown, the voice recognition component includes a voice navigation generation component, configured to preprocess the target audio data to obtain a plurality of target text paragraphs, the preprocessing being configured to convert the target audio data into the plurality of target text paragraphs; perform sentence-by-sentence analysis on each of the target text paragraphs based on the preset certificate inspection points and the preset behavior inspection points in the preset script content, to obtain the first timestamp information of the certificate inspection points and the behavior inspection points in the target audio data, the first timestamp information being the start time of the certificate inspection points and the behavior inspection points in the target audio data; and mark the time axis of the target audio data based on each of the first timestamp information, to obtain the voice navigation list.
[0090] In actual application, the first timestamp information of the appearance of each inspection point can be obtained based on the preset certificate inspection points and the plurality of target text paragraphs obtained by voice recognition of the target audio data. That is, the voice navigation list is composed of the timestamp (i.e., the first timestamp information) of the appearance of each inspection point and the corresponding inspection point. Of course, in order to ensure the integrity of the appearance of each inspection point, when the first timestamp of the appearance of each inspection point is obtained, the time (e.g., 2 seconds or 3 seconds before the first timestamp of the appearance of each inspection point) before the first timestamp of the appearance of each inspection point can be taken as the timestamp (the first timestamp information) of the appearance of each inspection point according to actual conditions.
[0091] Specifically, the certificate inspection points of the voice navigation list include identification card recognition, post qualification certificate recognition, work card recognition, risk assessment form recognition, and signature form recognition; and the behavior inspection points include signature behavior recognition. The specific judgment process is as follows: when "this is my work card" is recognized in the plurality of target text paragraphs after voice recognition, a work card recognition inspection point is generated; when "this is my post qualification certificate" is recognized in the plurality of target text paragraphs, a post qualification certificate inspection point is generated; when "show your identification card" is recognized in the plurality of target text paragraphs, an identification card inspection point is generated; when "show your risk assessment form" is recognized in the plurality of target text paragraphs, a risk assessment form inspection point is generated; when "please sign" is recognized in the plurality of target text paragraphs, a signature behavior inspection point is generated; and when "show your signature form" is recognized in the plurality of target text paragraphs, a signature form inspection point is generated.
[0092] In another embodiment of the present application, as shown in FIG. 6, the voice navigation generation component includes a voice navigation generation component, configured to preprocess the target audio data to obtain a plurality of target text paragraphs, the preprocessing being configured to convert the target audio data into the plurality of target text paragraphs; perform sentence-by-sentence analysis on each of the target text paragraphs based on the preset certificate inspection points and the preset behavior inspection points in the preset script content, to obtain the first timestamp information of the certificate inspection points and the behavior inspection points in the target audio data, the first timestamp information being the start time of the certificate inspection points and the behavior inspection points in the target audio data; and mark the time axis of the target audio data based on each of the first timestamp information, to obtain the voice navigation list. Figure 4As shown in the figure, the voice recognition component includes a text processing component library, which is configured to preprocess the target audio data to obtain a plurality of target text paragraphs, the preprocessing being configured to convert the target audio data into the plurality of target text paragraphs; obtain a script quality inspection result based on the preset script quality inspection point in the preset script content and the plurality of target text paragraphs, and obtain a sensitive word quality inspection result based on a preset sensitive word library and the plurality of target text paragraphs; and the first quality inspection result is composed of the script quality inspection result and the sensitive word quality inspection result. In this embodiment, the script quality inspection result is obtained based on the preset script quality inspection point in the preset script content and the plurality of target text paragraphs, and the sensitive word quality inspection result is obtained based on the preset sensitive word library and the plurality of target text paragraphs, so that the obtained first quality inspection result is more accurate, and the behavior of the salesperson is further standardized.
[0093] In order to more accurately obtain the plurality of target text paragraphs, in another embodiment of the present application, as shown in the figure, Figure 4 The voice navigation generation component or the text processing component library includes a text processing component, which is configured to perform voice recognition processing on the target audio data to obtain target text information corresponding to the target audio data; and perform semantic processing on the target text information to obtain the plurality of target text paragraphs, so that the obtained target text paragraphs are more reasonable and the voice recognition effect is better.
[0094] In another embodiment of the present application, as shown in the figure, Figure 4 The text processing component library includes a script matching component, which is configured to perform sentence-by-sentence analysis on the plurality of target text paragraphs based on the preset script quality inspection point to obtain a plurality of alternative hit sentences; calculate a similarity score of the preset script quality inspection point and each of the alternative hit sentences to obtain a plurality of preset similarity scores; calculate the sum of the plurality of preset similarity scores to obtain a target similarity score of the double-recording video data, and the target similarity score constitutes the script quality inspection result. Specifically, in this embodiment, the plurality of target text paragraphs are analyzed sentence by sentence based on the preset script quality inspection point to obtain a plurality of alternative hit sentences, the similarity score of the preset script quality inspection point and the alternative hit sentences is calculated to obtain a plurality of preset similarity scores, and the sum of the plurality of preset similarity scores constitutes the target similarity score, so that the target similarity score of the double-recording video data can be more accurately and simply obtained with less calculation, the overall audio data of the salesperson is more accurately scored, and the double-recording video data is further more efficiently quality inspected.
[0095] In order to more accurately obtain the preset similarity score and further more accurately obtain the target similarity score, in an embodiment of the present application, as shown in the figure, Figure 4As shown, the above-mentioned dialogue matching component includes a text similarity calculation component for performing word segmentation processing on a plurality of the above-mentioned candidate hit sentences to obtain a plurality of target keywords, wherein one of the above-mentioned candidate hit sentences corresponds to at least one of the above-mentioned target keywords; using each of the above-mentioned target keywords, respectively traversing the sentence corresponding to the above-mentioned preset dialogue quality inspection point to obtain a plurality of keyword matching degrees, and using each of the above-mentioned candidate hit sentences, respectively traversing the sentence corresponding to the above-mentioned preset dialogue quality inspection point to obtain a plurality of text distance matching degrees, one target keyword corresponds to a plurality of the above-mentioned keyword matching degrees, and one of the above-mentioned candidate hit sentences corresponds to a plurality of the above-mentioned text distance matching degrees; based on the above-mentioned keyword matching degrees and the above-mentioned text distance matching degrees corresponding to each of the above-mentioned candidate hit sentences, the corresponding above-mentioned preset similarity score is determined.
[0096] Specifically, in the above-mentioned embodiment, the process of calculating the text distance matching degree can be: determining the longest common subsequence (LCS for short) of the candidate hit sentence and the sentence corresponding to the preset dialogue quality inspection point. Then, based on the longest common subsequence, the text distance matching degree is obtained.
[0097] Specifically, the keyword matching degree and the text distance matching degree can be weighted to obtain the preset similarity score. In a specific embodiment of the present application, the preset similarity score = a*keyword matching degree + b*text distance matching degree. Wherein, a and b can be debugged according to the double recording video data of different regions, wherein a<=1, b<=1, a+b=1.
[0098] In another embodiment of the present application, as Figure 4 shown, the above-mentioned text processing component library includes a sensitive word matching component for using each of the above-mentioned preset sensitive words in the above-mentioned preset sensitive word library to compare with a plurality of the above-mentioned target text paragraphs; in the case that there is a target text in the above-mentioned target text paragraph which is the same as the above-mentioned preset sensitive word, the target text paragraph to which the above-mentioned target text belongs is determined as a sensitive word paragraph, and the above-mentioned sensitive word quality inspection result is formed by the above-mentioned sensitive word paragraph, and the subsequent first quality inspection result is formed based on the sensitive word quality inspection result and the dialogue quality inspection result, further ensuring that the first quality inspection result is more accurate.
[0099] In actual application process, in order to further reduce the calculation amount of double video data quality inspection, in another embodiment of the application, the video frame extraction component includes a video navigation generation component, which is used for frame extraction processing on a plurality of video frames in the target video data to obtain a plurality of first target video frames, clustering analysis on a plurality of the first target video frames to obtain the second timestamp information of the certificate quality inspection point and the behavior quality inspection point, the second timestamp information being the starting appearance time of the certificate quality inspection point and the behavior quality inspection point in the target video data, and marking each second timestamp information in the target video data to obtain the video navigation list.
[0100] Specifically, in the above embodiment, the entire target video data can be frame extraction processed at an average of 4 frames per second.
[0101] Specifically, in the above embodiment, the certificate quality inspection point in the video navigation list includes ID card recognition, post qualification certificate recognition, work card recognition, risk assessment table recognition and signature form recognition, and the behavior quality inspection point includes signature behavior recognition. The specific judgment process is: a plurality of video frame images in the target video image are frame extraction processed at an average of 4 frames per second to obtain a plurality of first target video frames, and clustering processing is performed based on a plurality of first target video frames, and the timestamp information (i.e. second timestamp information) of the corresponding quality inspection point is generated according to the ID card recognition, post qualification certificate recognition, work card recognition, risk assessment table recognition, signature form recognition and signature behavior recognition in the clustering result. In actual application process, in order to ensure the integrity of the appearance of each quality inspection point, when the timestamp of the first appearance of each quality inspection point is obtained, the timestamp of the first appearance of each quality inspection point can be taken as the timestamp (second timestamp information) of the appearance of each quality inspection point according to the actual situation.
[0102] In still another embodiment of the present application, the video frame extraction component includes an intelligent inspection scheduling component, which merges the voice navigation list and the video navigation list to obtain the target video data with target timestamp information, the target timestamp information being the one with the earliest start time among the first timestamp information and the second timestamp information corresponding to the certificate inspection point and the behavior inspection point; based on the target timestamp information, the target video data is frame-extracted to determine the second target video frames corresponding to the certificate inspection point and the second target video frames corresponding to the behavior inspection point; based on the second target video frames corresponding to the certificate inspection point, the certificate inspection point is inspected to obtain the target hit image corresponding to the certificate inspection point and the timestamp information corresponding to the target hit image, and based on the second target video frames corresponding to the behavior inspection point, the behavior inspection point is inspected to obtain the target hit image corresponding to the behavior inspection point and the timestamp information corresponding to the target hit image; the second inspection result is formed based on the target hit image corresponding to the certificate inspection point and the corresponding timestamp information, and the target hit image corresponding to the behavior inspection point and the corresponding timestamp information. In this embodiment, the voice navigation list and the video navigation list are merged to ensure that the target timestamp information is relatively accurate and the video of each inspection point is relatively complete. Then, the target video data is frame-extracted based on the target timestamp information, so as to ensure that the second target video frames are relatively accurate and the second inspection result is relatively accurate.
[0103] Specifically, in the above embodiment, during the merging of the voice navigation list and the video navigation list, for the same inspection point, the timestamp information in the voice navigation list and the video navigation list is taken as the one with the earlier time (which can be the first timestamp information or the second timestamp information), so as to ensure that the start time of each inspection point is relatively accurate and the video of each inspection point is relatively complete, and then the behavior inspection point and the certificate inspection point are inspected based on the second target video, so as to further ensure that the second inspection result is relatively accurate.
[0104] In order to more accurately perform quality inspection on the personnel-off position, in an embodiment of the present application, the quality inspection device further includes a personnel-off detection component, configured to perform frame extraction processing on the target video data having the target timestamp information, to determine a plurality of third target video frames corresponding to the personnel-off position; perform quality inspection on the personnel-off position based on a first target network model and the plurality of third target video frames, to obtain the target hit image corresponding to the personnel-off position and the timestamp information corresponding to the target hit image, wherein the first target network model is obtained based on a neural network model and is trained; and form the second quality inspection result based on the target hit image corresponding to the personnel-off position and the corresponding timestamp information.
[0105] Specifically, in the above embodiment, the target video data having the target timestamp information is subjected to frame extraction processing to obtain a plurality of third target video frames. Then, the personnel-off position is subjected to quality inspection based on the first target network model and the plurality of third target video frames. That is, the quality inspection method of the present application does not start from the timestamp of each quality inspection position to detect all video frames, but detects the plurality of third target video frames, thereby ensuring that the calculation amount of the first target network model is small and the obtained second quality inspection result is more accurate.
[0106] Specifically, in the above embodiment, the target hit image can be an image in which the personnel-off position is clearest.
[0107] In another embodiment of the present application, the intelligent quality inspection scheduling component includes a certificate quality inspection position component library, configured to perform preprocessing on a plurality of second target video frames corresponding to the certificate quality inspection position by using a target combination algorithm, to obtain a plurality of preprocessed second target video frames, wherein the target combination algorithm is an algorithm obtained by combining Faster-RCNN, CNN and a traditional image processing algorithm; perform text detection on the plurality of preprocessed second target video frames by using a CTPN network, to obtain a plurality of preset text sequences, and process each preset text sequence by using a connection time sequence classification model, to obtain a plurality of credible text sequences; determine an image corresponding to the second target video frame having the most character information in each credible text sequence as a target hit image, and determine the timestamp information of the image corresponding to the second target video frame having the most character information in each credible text sequence as the timestamp information corresponding to the target hit image, thereby ensuring that the recognition result of the certificate quality inspection position is more accurate, and further ensuring that the target quality inspection result in the present solution is more accurate.
[0108] Specifically, in the present solution, as Figure 2As shown, first, a target combination algorithm based on Faster-RCNN, Convolutional Neural Networks (CNN) and traditional image processing algorithm is used to pre-process multiple second target video frames. Second, a text detection algorithm based on an optimized Connection Text Proposal Network (CTPN) is used to detect and locate the edges of the card certificate to be recognized in the second target video frame, and multiple preset text sequences are obtained. That is, the optimized CTPN network is optimized for scene text recognition in dual-recording video data, which is different from the image scanned by a traditional scanner or captured by a high-speed scanner. The text recognition algorithm of the present scheme needs to locate the text area in the video frame, and there is a big difference between the text in the dual-recording video data and the printed text recognition in the scanned or high-speed scanner image. Compared with other neural network algorithms, the CTPN algorithm is more suitable for this scene text recognition, because the CTPN has the following characteristics: 1) using a vertical anchor that is more suitable for natural scene text detection characteristics (this is because the size of the text is small compared to the object, and by using a vertical anchor that is more suitable for the size of the text, small size text candidate boxes can be detected); 2) introducing RNN to process the sequence characteristics existing in scene text detection, which can connect small size text and obtain text lines; 3) introducing Side-refinement to improve the accuracy of text box boundary detection. For character recognition in video frames, the present scheme uses a method combining convolutional neural network and recurrent neural network to extract character feature information, normalizes a numerical vector to a probability distribution vector through softmax, and uses an exponential function to calculate the probability distribution of the feature. Finally, the Connectionist Temporal Classification (CTC) model is used to process each preset text sequence to obtain multiple reliable text sequences. In the present scheme, the Connectionist Temporal Classification model based on recurrent neural network is used, which can not only obtain the most reliable feature at a specific position, but also can correct the feature result combined with the context information (such as Sun Moon Lake and Mingtan), and finally output the reliable text sequence. The image corresponding to the second target video frame with the most text information in the reliable text sequence is determined as the target hit image. And the target hit image and the time stamp information corresponding to the target hit image constitute the second quality inspection result.
[0109] Specifically, in the above-mentioned embodiments, the character recognition is mainly performed on the worker card image, the post qualification certificate image, the ID card image, the risk assessment form image, and the signature form image to extract the text information in the images. In addition, the above-mentioned certificate quality inspection point component library further includes a worker card recognition component, a post qualification certificate recognition component, an ID card recognition component, a risk assessment form recognition component, and a signature form recognition component. Therefore, the worker card recognition component can be used to recognize the worker card image, the post qualification certificate recognition component can be used to recognize the post qualification certificate image, the ID card recognition component can be used to recognize the ID card image, the risk assessment form recognition component can be used to recognize the risk assessment form image, and the signature form recognition component can be used to recognize the signature form image.
[0110] In order to more accurately and simply obtain the target hit image corresponding to the behavior quality inspection point and the timestamp information corresponding to the target hit image, in another embodiment of the present application, the above-mentioned intelligent quality inspection scheduling component includes a signature behavior recognition component, which is used to perform quality inspection on the above-mentioned behavior quality inspection point based on a second target network model and a plurality of the above-mentioned second target video frames to obtain the above-mentioned target hit image corresponding to the above-mentioned behavior quality inspection point and the timestamp information corresponding to the above-mentioned target hit image, and the above-mentioned second target network model is obtained based on a neural network model and is trained.
[0111] Specifically, the second quality inspection result in the present application is composed of the target hit image of the behavior quality inspection point and the corresponding timestamp information, the target hit image of the certificate quality inspection point and the corresponding timestamp information, and the target hit image of the personnel off-site quality inspection point and the corresponding timestamp information.
[0112] Specifically, in the present application, in order to improve the accuracy of face recognition technology, an optimized RetinaNet is used as the backbone network, and A-Softmax is used as the loss function to train the face feature extractor. Not only can the recognition speed be improved to realize completely end-to-end recognition, but also the face recognition model trained based on about seven million face data can greatly improve the recognition accuracy.
[0113] The intelligent quality inspection technical solution in the prior art cannot respond to the changes in the quality inspection points in time, so the intelligent quality inspection solution in the prior art has the disadvantages of fixed business processes and fixed parameter configurations, which causes the quality inspection points to change and the overall system process to be adjusted, greatly affecting the usability and flexibility of intelligent quality inspection. In order to solve the technical problems of poor flexibility and low usability of the intelligent quality inspection solution in the prior art, in a specific embodiment of the present application, as Figure 5As shown, the workflow scheduling engine provided by the present application can provide an intelligent quality inspection service for the double recording data in the process of selling self-financial products and agent products, and can realize the quality inspection of double recording data of insurance, finance, fund, asset management plan, structured deposit, precious metal and other businesses. Specifically, the workflow engine is composed of three parts of component management, workflow management and task management. Among them, the component management mainly includes component configuration management, component creation and debugging and component library. The component library includes basic components and business components. Among them, the basic components include intelligent character recognition components, speech recognition components, natural language processing components and face recognition components, which can provide artificial intelligence basic processing capability for target audio data and target video data. Secondly, on the basis of the basic components, by adding business logic rules, business components for double recording video data are formed, such as audio video separation component, video frame extraction component, integration component and personnel off-site component. Through the separation of basic components and business components, the business processing function can be quickly realized based on the basic artificial intelligence processing capability; the workflow management provides the ability to create a workflow based on the component library, which includes workflow parameter setting, workflow running strategy, workflow publishing and workflow creation and debugging; the task management provides the workflow task running capability, which includes task creation, query and deletion, task intervention, task running analysis, task plan, task intervention and the like.
[0114] In a specific embodiment of the present application, as shown, Figure 6 The double recording video data quality inspection device of the present application adopts a micro-service containerization deployment method, encapsulates the basic running environment (such as operating system kernel, C++ running environment, Java running environment, etc.), Spring framework, Java local interface and model inference service into a container image package, which can be directly run on the container cloud platform built by Kubernetes and Docke, so that the intelligent quality inspection method of the present application can dynamically adjust the hardware resource occupation according to the size of the business volume.
[0115] The double recording video data quality inspection device includes a processor and a memory, and the audio video separation component, the speech recognition component, the video frame extraction component and the integration component are stored in the memory as program units. The processor executes the above-mentioned program units stored in the memory to realize the corresponding functions.
[0116] The processor includes a core, which retrieves the corresponding program unit from the memory. The core can be set to one or more, and the accuracy of the quality inspection of the double recording video data in the prior art can be improved by adjusting the core parameters.
[0117] The memory can include non-persistent memory in a computer readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory, and the memory includes at least one memory chip.
[0118] The embodiment of the present application provides a computer readable storage medium, which stores a program, and the program is executed by a processor to realize the quality inspection method of the double-recording video data.
[0119] The embodiment of the present application provides a processor, which is used for running a program, and the program is executed to perform the quality inspection method of the double-recording video data.
[0120] The embodiment of the present application provides a device, which comprises a processor, a memory and a program stored in the memory and capable of running on the processor, and the processor is executed to realize at least the following steps:
[0121] Step S101, audio separation is performed on the obtained double-recording video data to obtain target audio data and target video data;
[0122] Step S102, quality inspection is performed on a plurality of quality inspection points by at least using the target audio data and preset dialogue content, to obtain a first quality inspection result and a voice navigation list, and the voice navigation list is obtained by marking first timestamp information of part of the quality inspection points in a time axis of the target audio data;
[0123] Step S103, clustering analysis is performed on the target video data based on part of the quality inspection points to obtain a video navigation list, and quality inspection is performed on part of the quality inspection points based on the voice navigation list and the video navigation list to obtain a second quality inspection result, the video navigation list is obtained by marking second timestamp information of part of the quality inspection points in a time axis of the target video data, and the quality inspection points include a certificate quality inspection point, a behavior quality inspection point, a dialogue quality inspection point, a sensitive word quality inspection point and a personnel off quality inspection point;
[0124] Step S104, the first quality inspection result and the second quality inspection result are integrated to obtain a target quality inspection result, and the target quality inspection result is sent to a display screen of a terminal device, so that the display screen displays the target quality inspection result.
[0125] The device in the present application can be a server, a PC, a PAD, a mobile phone and the like.
[0126] The present application also provides a computer program product, which is suitable for executing a program with at least the following method steps when executed on a data processing device:
[0127] Step S101, audio separation is performed on the obtained dual-recording video data to obtain target audio data and target video data;
[0128] Step S102, at least using the target audio data and the preset script content, quality inspection is performed on a plurality of quality inspection points to obtain a first quality inspection result and a voice navigation list, wherein the voice navigation list is obtained by marking part of the first timestamp information of the quality inspection points in the time axis of the target audio data;
[0129] Step S103, based on part of the quality inspection points, clustering analysis is performed on the target video data to obtain a video navigation list, and based on the voice navigation list and the video navigation list, quality inspection is performed on part of the quality inspection points to obtain a second quality inspection result, wherein the video navigation list is obtained by marking part of the second timestamp information of the quality inspection points in the time axis of the target video data, and the quality inspection points include a certificate quality inspection point, a behavior quality inspection point, a script quality inspection point, a sensitive word quality inspection point, and a personnel off-site quality inspection point;
[0130] Step S104, the first quality inspection result and the second quality inspection result are integrated to obtain a target quality inspection result, and the target quality inspection result is sent to the display screen of the terminal device to enable the display screen to display the target quality inspection result.
[0131] In order for those skilled in the art to more clearly understand the technical solutions of the present application, the technical solutions and technical effects of the present application will be described in conjunction with specific embodiments below.
[0132] Embodiments
[0133] In a specific embodiment of the present application, as shown in Figure 7 a dual-recording video data quality inspection scheme is provided. A piece of obtained dual-recording video data is separated to obtain target audio data and target video data.
[0134] The target audio data is subjected to speech recognition processing to obtain target text information corresponding to the target audio data, and the target text information is subjected to semantic processing to obtain a plurality of target text paragraphs. Based on preset certificate quality inspection points and preset behavior quality inspection points in the preset script content, each target text paragraph is subjected to sentence-by-sentence analysis to obtain first timestamp information of the certificate quality inspection points and the behavior quality inspection points in the target audio data, so as to generate a voice navigation list. Based on preset script quality inspection points in the preset script content and the plurality of target text paragraphs, a script quality inspection result is obtained. Based on a preset sensitive word library and the plurality of target text paragraphs, a sensitive word quality inspection result is obtained. The first quality inspection result is composed of the script quality inspection result and the sensitive word quality inspection result.
[0135] Frame extraction is performed on a plurality of video frames in the target video data to obtain a plurality of first target video frames. Cluster analysis is performed on the target video data based on the partial quality inspection points, and cluster analysis is performed on the plurality of first target video frames to obtain second timestamp information of the certificate quality inspection points and the behavior quality inspection points. Each second timestamp information is marked in the target video data to obtain a video navigation list.
[0136] The voice navigation list and the video navigation list are merged to obtain target video data with target timestamp information. Based on the plurality of second target video frames, target hit images corresponding to the certificate quality inspection points and timestamp information corresponding to the target hit images, and target hit images corresponding to the behavior quality inspection points and timestamp information corresponding to the target hit images are obtained. Based on the target timestamp information, frame extraction is performed on the target video data to determine a plurality of second target video frames corresponding to the certificate quality inspection points and a plurality of second target video frames corresponding to the behavior quality inspection points. Based on the plurality of second target video frames, target hit images corresponding to the behavior quality inspection points and timestamp information corresponding to the target hit images are obtained. Based on the first target network model and the plurality of third target video frames, the personnel off-site quality inspection point is inspected to obtain target hit images corresponding to the personnel off-site quality inspection point and timestamp information corresponding to the target hit images. The target hit images corresponding to each behavior quality inspection point, certificate quality inspection point and personnel off-site quality inspection point and the timestamp information thereof constitute a second quality inspection result.
[0137] The first quality inspection result and the second quality inspection result are integrated to obtain a target quality inspection result. The target quality inspection result is sent to the display screen of the terminal device to enable the display screen to display the target quality inspection result. In this way, the quality inspection personnel can clearly and intuitively observe the quality inspection result.
[0138] In the above embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0139] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented by other ways. Among them, the device embodiments described above are only schematic, for example, the division of the above units can be a logical function division, and actual implementation can have another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between units or modules, which can be electrical or other forms.
[0140] The units described as separate components above can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0141] In addition, each functional unit in each embodiment of the application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0142] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the application, essentially or the part that contributes to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the above-mentioned method of each embodiment of the application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various program code storage media.
[0143] From the above description, it can be seen that the above-mentioned embodiments of the application achieve the following technical effects:
[0144] 1) In the quality inspection method of the dual-recording video data of the application, first, the preset speech content is used to perform quality inspection on each quality inspection point in the target audio data obtained by separating the dual-recording video data, to obtain a first quality inspection result and a voice navigation list; then, based on part of the quality inspection points, the target video data obtained by separating the dual-recording video data is subjected to cluster analysis to obtain a video navigation list, and based on the voice navigation list and the video navigation list, the quality inspection is performed on part of the quality inspection points to obtain a second quality inspection result; finally, the first quality inspection result and the second quality inspection result are integrated to obtain a target quality inspection result, and the target quality inspection result is sent to the display screen of the terminal device to enable the display screen to display the target quality inspection result. Compared with the prior art in which the dual-recording video data is manually inspected or the quality inspection is performed on each quality inspection point depending on the target audio data to obtain a target quality inspection result, in the present scheme, the quality inspection is performed on each quality inspection point in the target audio data in the dual-recording video data to obtain a first quality inspection result and a voice navigation list. Then, based on part of the quality inspection points, the target video data is subjected to cluster analysis to obtain a video navigation list, and based on the voice navigation list and the video navigation list, the quality inspection is performed on part of the quality inspection points to obtain a second quality inspection result, and based on the first quality inspection result and the second quality inspection result, a target quality inspection result is obtained, which ensures that the dual-recording video data is efficiently inspected and each quality inspection point is accurately inspected, and the target quality inspection result is accurate, thereby solving the problem of low accuracy in the quality inspection of the dual-recording video data in the prior art, and further ensuring high efficiency in the quality inspection of the dual-recording video data.
[0145] 2)、the double video data quality inspection device of the application, the voice recognition component is used to adopt the preset speech content to carry out quality inspection on each quality inspection point in the target audio data separated from the double video data, obtain the first quality inspection result and the voice navigation list; the video frame extraction component is used to carry out clustering analysis on the target video data separated from the double video data based on part of the quality inspection points, obtain the video navigation list, and carry out quality inspection on part of the quality inspection points based on the voice navigation list and the video navigation list, the second quality inspection result; the integration component is used to integrate the first quality inspection result and the second quality inspection result, obtain the target quality inspection result, and send the target quality inspection result to the display screen of the terminal device, so that the display screen displays the target quality inspection result. Compared with the prior art, the target audio data in the double video data is used to carry out quality inspection on each quality inspection point, obtain the first quality inspection result and the voice navigation list. Then, based on part of the quality inspection points, the target video data is subjected to clustering analysis to obtain the video navigation list, and part of the quality inspection points is subjected to quality inspection based on the voice navigation list and the video navigation list to obtain the second quality inspection result. Based on the first quality inspection result and the second quality inspection result, the target quality inspection result is obtained, which ensures that each quality inspection point is accurately inspected on the basis of efficiently inspecting the double video data, and ensures that the obtained target quality inspection result is relatively accurate, thereby solving the problem of low accuracy of quality inspection of the double video data in the prior art, and further ensuring the high efficiency of quality inspection of the double video data.
[0146] The above only describes the preferred embodiments of the application and is not used to limit the application. Those skilled in the art can make various modifications and changes to the application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall be included in the protection scope of the application.
Claims
1. A method for quality inspection of dual recording video data, characterized in that, The method comprises the following steps: audio separation is performed on the obtained dual-recording video data to obtain target audio data and target video data; at least the target audio data and preset dialogue content are used to perform quality inspection on multiple quality inspection points to obtain a first quality inspection result and a voice navigation list, the voice navigation list being obtained by marking first timestamp information of part of the quality inspection points on a time axis of the target audio data; based on part of the quality inspection points, cluster analysis is performed on the target video data to obtain a video navigation list, and based on the voice navigation list and the video navigation list, quality inspection is performed on part of the quality inspection points to obtain a second quality inspection result, the video navigation list being obtained by marking second timestamp information of part of the quality inspection points on a time axis of the target video data, the quality inspection points including a certificate quality inspection point, a behavior quality inspection point, a dialogue quality inspection point, a sensitive word quality inspection point, and a personnel off-site quality inspection point; the first quality inspection result and the second quality inspection result are integrated to obtain a target quality inspection result, and the target quality inspection result is sent to a display screen of a terminal device to enable the display screen to display the target quality inspection result, based on the voice navigation list and the video navigation list, quality inspection is performed on part of the quality inspection points to obtain a second quality inspection result, including: the voice navigation list and the video navigation list are merged to obtain the target video data with target timestamp information, the target timestamp information being a timestamp information with the earliest start time among the first timestamp information and the second timestamp information corresponding to the certificate quality inspection point and the behavior quality inspection point; based on the target timestamp information, frame extraction is performed on the target video data to determine a plurality of second target video frames corresponding to the certificate quality inspection point and a plurality of second target video frames corresponding to the behavior quality inspection point; based on the plurality of second target video frames corresponding to the certificate quality inspection point, quality inspection is performed on the certificate quality inspection point to obtain a target hit image corresponding to the certificate quality inspection point and timestamp information corresponding to the target hit image; based on the plurality of second target video frames corresponding to the behavior quality inspection point, quality inspection is performed on the behavior quality inspection point to obtain the target hit image corresponding to the behavior quality inspection point and timestamp information corresponding to the target hit image; the second quality inspection result is constituted by the target hit image corresponding to the certificate quality inspection point and the corresponding timestamp information, and the target hit image corresponding to the behavior quality inspection point and the corresponding timestamp information.
2. The quality inspection method of claim 1, wherein, at least the target audio data and preset dialogue content are used to perform quality inspection on multiple quality inspection points to obtain a voice navigation list, including: the target audio data is preprocessed to obtain a plurality of target text paragraphs, the preprocessing being used to convert the target audio data into a plurality of target text paragraphs; based on the preset certificate quality inspection points and the preset behavior quality inspection points in the preset dialogue content, performing sentence-by-sentence analysis on each of the target text paragraphs to obtain the first timestamp information of the certificate quality inspection points and the behavior quality inspection points in the target audio data, the first timestamp information being start time of the certificate quality inspection points and the behavior quality inspection points in the target audio data; based on each of the first timestamp information, marking a time axis of the target audio data to obtain the voice navigation list.
3. The method of claim 1, wherein, At least by using target audio data and preset dialogue content, multiple quality inspection points are inspected to obtain a first quality inspection result, including: preprocessing the target audio data to obtain multiple target text paragraphs, the preprocessing being used to convert the target audio data into multiple target text paragraphs; based on preset dialogue quality inspection points in the preset dialogue content and multiple target text paragraphs, obtaining a dialogue quality inspection result, and based on a preset sensitive word library and multiple target text paragraphs, obtaining a sensitive word quality inspection result; the dialogue quality inspection result and the sensitive word quality inspection result constitute the first quality inspection result.
4. The quality inspection method according to claim 2 or 3, characterized in that, The preprocessing of the target audio data to obtain multiple target text paragraphs includes: performing speech recognition processing on the target audio data to obtain target text information corresponding to the target audio data; performing semantic processing on the target text information to obtain multiple target text paragraphs.
5. The quality inspection method according to claim 3, characterized in that, Based on the preset dialogue quality inspection points in the preset dialogue content and multiple target text paragraphs, obtaining a dialogue quality inspection result, including: based on the preset dialogue quality inspection points, performing sentence-by-sentence analysis on multiple target text paragraphs to obtain multiple candidate hit sentences; calculating the similarity scores of the preset dialogue quality inspection points and each of the candidate hit sentences to obtain multiple preset similarity scores; calculating the sum of multiple preset similarity scores to obtain a target similarity score of the double recording video data, and since the target similarity score constitutes the dialogue quality inspection result.
6. The method of claim 5, wherein, Calculating the similarity scores of the preset dialogue quality inspection points and each of the candidate hit sentences to obtain multiple preset similarity scores includes: performing word segmentation processing on multiple candidate hit sentences to obtain multiple target keywords, wherein one candidate hit sentence corresponds to at least one target keyword; using each of the target keywords, traversing the sentence corresponding to the preset dialogue quality inspection point to obtain multiple keyword matching degrees, and using each of the candidate hit sentences, traversing the sentence corresponding to the preset dialogue quality inspection point to obtain multiple text distance matching degrees, one target keyword corresponding to multiple keyword matching degrees, and one candidate hit sentence corresponding to multiple text distance matching degrees; based on the keyword matching degrees and the text distance matching degrees corresponding to each of the candidate hit sentences, determining the corresponding preset similarity scores.
7. The quality inspection method according to claim 3, characterized in that, Based on a preset sensitive word library and multiple target text paragraphs, a sensitive word quality inspection result is obtained, including: using each of the preset sensitive words in the preset sensitive word library to compare with multiple target text paragraphs respectively; In a case where the target text in the target text paragraph is identical to the preset sensitive word, the target text paragraph to which the target text belongs is determined as a sensitive word paragraph, and the sensitive word paragraph is used to form the sensitive word quality inspection result.
8. The quality inspection method of claim 1, wherein, Based on the quality inspection points, the target video data is subjected to clustering analysis to obtain a video navigation list, including: Frame extraction processing is performed on a plurality of video frames in the target video data to obtain a plurality of first target video frames; Clustering analysis is performed on the plurality of first target video frames to obtain second timestamp information of the certificate quality inspection point and the behavior quality inspection point, the second timestamp information being start appearance time of the certificate quality inspection point and the behavior quality inspection point in the target video data; Each of the second timestamp information is marked in the target video data to obtain the video navigation list.
9. The method of claim 1, wherein, The quality inspection method further includes: Frame extraction processing is performed on the target video data with target timestamp information to determine a plurality of third target video frames corresponding to the personnel off-site quality inspection point; Based on a first target network model and the plurality of third target video frames, the personnel off-site quality inspection point is subjected to quality inspection to obtain a target hit image corresponding to the personnel off-site quality inspection point and timestamp information corresponding to the target hit image, the first target network model being obtained based on a neural network model and being trained; Based on the target hit image corresponding to the personnel off-site quality inspection point and the corresponding timestamp information, the second quality inspection result is formed.
10. The method of claim 1, wherein, Based on the plurality of second target video frames corresponding to the certificate quality inspection point, the certificate quality inspection point is subjected to quality inspection to obtain a target hit image corresponding to the certificate quality inspection point and timestamp information corresponding to the target hit image, including: A target combination algorithm is used to pre-process the plurality of second target video frames corresponding to the certificate quality inspection point to obtain a plurality of pre-processed second target video frames, the target combination algorithm being an algorithm obtained by combining Faster-RCNN, CNN and a traditional image processing algorithm; A CTPN network is used to detect text in the plurality of pre-processed second target video frames to obtain a plurality of preset text sequences, and a connection time sequence classification model is used to process each of the preset text sequences to obtain a plurality of credible text sequences; An image corresponding to the second target video frame with the most character information in each of the credible text sequences is determined as a target hit image, and timestamp information of the image corresponding to the second target video frame with the most character information in each of the credible text sequences is determined as timestamp information corresponding to the target hit image.
11. The method of claim 1, wherein, Based on the plurality of second target video frames corresponding to the behavior quality inspection point, the behavior quality inspection point is subjected to quality inspection to obtain the target hit image corresponding to the behavior quality inspection point and the timestamp information corresponding to the target hit image, including: The behavior inspection point is inspected based on the second target network model and the plurality of second target video frames, target hit images corresponding to the behavior inspection point and timestamp information corresponding to the target hit images are obtained, and the second target network model is obtained based on a neural network model and training.
12. A quality inspection apparatus for dual recording video data, characterized by comprising: Comprise: An audio-video separation component configured to separate audio from acquired dual-recording video data to obtain target audio data and target video data; A speech recognition component configured to inspect a plurality of inspection points using at least the target audio data and preset dialogue content to obtain a first inspection result and a voice navigation list, wherein the voice navigation list is obtained by marking first timestamp information of part of the inspection points on a time axis of the target audio data; A video frame extraction component configured to perform cluster analysis on the target video data based on part of the inspection points to obtain a video navigation list, and perform inspection on part of the inspection points based on the voice navigation list and the video navigation list to obtain a second inspection result, wherein the video navigation list is obtained by marking second timestamp information of part of the inspection points on a time axis of the target video data, and the inspection points comprise a certificate inspection point, a behavior inspection point, a dialogue inspection point, a sensitive word inspection point, and a personnel absence inspection point; An integration component configured to integrate the first inspection result and the second inspection result to obtain a target inspection result, and send the target inspection result to a display screen of a terminal device to enable the display screen to display the target inspection result, The video frame extraction component comprises an intelligent inspection scheduling component configured to perform merging processing on the voice navigation list and the video navigation list to obtain the target video data with target timestamp information, the target timestamp information being a timestamp information with the earliest start time among the first timestamp information corresponding to the certificate inspection point and the second timestamp information corresponding to the behavior inspection point; perform frame extraction processing on the target video data based on the target timestamp information to determine a plurality of second target video frames corresponding to the certificate inspection point and a plurality of second target video frames corresponding to the behavior inspection point; perform inspection on the certificate inspection point based on the plurality of second target video frames corresponding to the certificate inspection point to obtain target hit images corresponding to the certificate inspection point and timestamp information corresponding to the target hit images; and perform inspection on the behavior inspection point based on the plurality of second target video frames corresponding to the behavior inspection point to obtain the target hit images corresponding to the behavior inspection point and the timestamp information corresponding to the target hit images; and the second inspection result is formed based on the target hit images corresponding to the certificate inspection point and the corresponding timestamp information, and the target hit images corresponding to the behavior inspection point and the corresponding timestamp information.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium comprises a stored program, wherein the program performs the inspection method of the dual-recording video data according to any one of claims 1 to 11.
14. A processor, comprising: The processor is configured to run a program, wherein the program performs the quality inspection method of the dual-recording video data according to any one of claims 1 to 11 when running.
Citation Information
Patent Citations
Double-recording quality inspection method and device based on artificial intelligence, computer equipment and medium
CN112101311A
Method and apparatus for quality detection of double-recorded video, and computer device and storage medium
WO2020140665A1