A smart dual-recording quality inspection method and device
By using an intelligent dual-recording quality inspection method that combines facial recognition, video action recognition, and voice recognition, the problem of low efficiency in manual quality inspection during the document collection process has been solved, enabling real-time supervision of document collection and ensuring business compliance.
Patent Information
- Application Number
- CN202310307977.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-03-27
AI Technical Summary
In the existing technology, the document collection process lacks effective supervision measures, which makes manual quality inspection time-consuming, labor-intensive and prone to errors, and difficult to provide strong support for business processes.
The system employs an intelligent dual-recording quality inspection method, using facial recognition, video motion recognition, and voice recognition to monitor the document collection process in real time, ensuring that the document is collected by the person in question.
It has improved the efficiency and accuracy of the document collection process, standardized business procedures, reduced the risk of document loss and miscollection, and provided dual protection during and after the process.
Smart Images

Figure CN116453015B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to an intelligent dual-recording quality inspection method and device. Background Technology
[0002] In recent years, government service halls and bank branches have seen a surge in document issuance, leading to a significant increase in the use of self-service document issuance terminals and portable instant document issuance devices. In addition, there are numerous door-to-door document processing scenarios, resulting in issues such as lost or incorrectly issued documents. Currently, there are no effective measures to supervise the document collection process, which is generally conducted manually. However, manual quality inspection is time-consuming, labor-intensive, inefficient, and prone to errors. Furthermore, it is a post-event audit and cannot provide strong support for the document collection process. Summary of the Invention
[0003] This application provides an intelligent dual-recording quality inspection method to solve the problems of time-consuming, labor-intensive, and inefficient post-event manual quality inspection.
[0004] In a first aspect, embodiments of this application provide an intelligent dual-recording quality inspection method, including:
[0005] In response to the user's instruction to confirm the card production information, after confirming that the card production was successful, the process of the user receiving the card is recorded to obtain video and audio data;
[0006] The video data is subjected to facial recognition to obtain the recognition result of the video data;
[0007] When the recognition result is successfully matched with the facial data corresponding to the card information, action recognition is performed on the video data;
[0008] Once it is determined that the recognition action corresponding to the video data is a normal card-collecting action, and the audio data is identified as compliant, the process of the user collecting the card is determined to be the user collecting the card themselves.
[0009] Based on the above solution, intelligent quality inspection is carried out through facial recognition, video action recognition, and voice recognition to solve the problems of slow and error-prone manual review, standardize and safeguard the entire card issuance process, and provide strong support for solving issues such as lost or misissued documents and business compliance.
[0010] In one possible implementation, the action recognition of the video data includes:
[0011] The video data is input into a first neural network for feature extraction to determine the feature matrix of the human torso corresponding to the multiple images included in the video data.
[0012] The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to determine the recognition action corresponding to the video data.
[0013] In one possible implementation, the identification of the audio data includes:
[0014] The audio data is identified using a third neural network to obtain the audio content corresponding to the audio data;
[0015] Determining the compliance of the audio data includes:
[0016] The audio content is compared with preset first audio data and preset second audio data respectively. The first audio data is the audio data that needs to be included in the card retrieval process, and the second audio data includes audio data that is prohibited from appearing in the card retrieval process.
[0017] When it is determined that the audio content contains audio data from the first audio data and the audio content does not contain audio data from the second audio data, the audio content is determined to be compliant.
[0018] In one possible implementation, the method further includes:
[0019] When the recognition result fails to match the facial data, it is determined that the user's card-collecting process was not performed by the user themselves, and a first alarm message is generated; or,
[0020] When it is determined that the recognition action corresponding to the video data is an abnormal card-collecting action, the user's card-collecting process is determined to be abnormal, and a second alarm prompt is generated; or,
[0021] When the audio data is non-compliant, it is determined that the language used during the user's card collection process is non-compliant, and a third alarm prompt is generated.
[0022] In one possible implementation, the first neural network and the second neural network constitute a total neural network model, which is obtained as follows:
[0023] Obtain a training sample set, which includes multiple training samples labeled as normal card-collecting actions and multiple training samples labeled as abnormal card-collecting actions.
[0024] For each training sample, each training sample is input into the first neural network to be trained for feature extraction, thereby obtaining the feature matrix of the human torso corresponding to the multiple images included in each training sample.
[0025] The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to be trained to determine the predicted recognition action corresponding to each training sample.
[0026] The loss value is determined based on the predicted recognition action and label corresponding to each training sample;
[0027] The network parameters of the first neural network to be trained and the second neural network to be trained are adjusted using the loss value to obtain the overall neural network model.
[0028] Secondly, embodiments of this application provide an intelligent dual-recording quality inspection device, comprising:
[0029] The acquisition module is used to respond to the user's instruction to confirm the card production information. After confirming that the card production is successful, it records the process of the user receiving the card to obtain video and audio data.
[0030] The recognition module is used to perform facial recognition on the video data to obtain the recognition result of the video data;
[0031] When the recognition result is successfully matched with the facial data corresponding to the card information, action recognition is performed on the video data;
[0032] The determination module is used to determine that the user's card collection process is a personal collection process after determining that the recognition action corresponding to the video data is a normal card collection action and after recognizing the audio data to determine that the audio data is compliant.
[0033] In one possible implementation, the recognition module, when performing action recognition on the video data, is specifically used for:
[0034] The video data is input into a first neural network for feature extraction to determine the feature matrix of the human torso corresponding to the multiple images included in the video data.
[0035] The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to determine the recognition action corresponding to the video data.
[0036] In one possible implementation, the recognition module, when recognizing the audio data, is specifically used for:
[0037] The audio data is identified using a third neural network to obtain the audio content corresponding to the audio data;
[0038] The determining module, when determining that the audio data is compliant, is specifically used for:
[0039] The audio content is compared with preset first audio data and preset second audio data respectively. The first audio data is the audio data that needs to be included in the card retrieval process, and the second audio data includes audio data that is prohibited from appearing in the card retrieval process.
[0040] When it is determined that the audio content contains audio data from the first audio data and the audio content does not contain audio data from the second audio data, the audio content is determined to be compliant.
[0041] In one possible implementation, the determining module is further configured to:
[0042] When the recognition result fails to match the facial data, it is determined that the user's card-collecting process was not performed by the user themselves, and a first alarm message is generated; or,
[0043] When it is determined that the recognition action corresponding to the video data is an abnormal card-collecting action, the user's card-collecting process is determined to be abnormal, and a second alarm prompt is generated; or,
[0044] When the audio data is non-compliant, it is determined that the language used during the user's card collection process is non-compliant, and a third alarm prompt is generated.
[0045] In one possible implementation, the first neural network and the second neural network constitute a total neural network model, which is obtained as follows:
[0046] Obtain a training sample set, which includes multiple training samples labeled as normal card-collecting actions and multiple training samples labeled as abnormal card-collecting actions.
[0047] For each training sample, each training sample is input into the first neural network to be trained for feature extraction, thereby obtaining the feature matrix of the human torso corresponding to the multiple images included in each training sample.
[0048] The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to be trained to determine the predicted recognition action corresponding to each training sample.
[0049] The loss value is determined based on the predicted recognition action and label corresponding to each training sample;
[0050] The network parameters of the first neural network to be trained and the second neural network to be trained are adjusted using the loss value to obtain the overall neural network model.
[0051] Thirdly, embodiments of this application provide an execution device, including:
[0052] Memory, used to store program instructions;
[0053] A processor is configured to retrieve program instructions from the memory and execute the method described in the first aspect and different implementations of the first aspect according to the retrieved program instructions.
[0054] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect and different implementations of the first aspect.
[0055] The technical effects of any of the implementation methods in the second to fourth aspects can be found in the first aspect and the technical effects of different implementation methods of the first aspect, which will not be repeated here. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0057] Figure 1 A flowchart illustrating an intelligent dual-recording quality inspection method provided in an embodiment of this application;
[0058] Figure 2 A schematic diagram of the training process of a total neural network model provided in an embodiment of this application;
[0059] Figure 3 This is a schematic diagram of the structure of an intelligent dual-recording quality inspection device provided in an embodiment of this application;
[0060] Figure 4 This is a schematic diagram of the structure of an execution device provided in an embodiment of this application. Detailed Implementation
[0061] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can be arranged and designed in various different configurations.
[0062] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0063] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0064] In recent years, government service halls and bank branches have seen a surge in document issuance, leading to a significant increase in the use of self-service document issuance terminals and portable instant document issuance devices. In addition, there are numerous door-to-door document processing scenarios, resulting in issues such as lost or incorrectly issued documents. Currently, there are no effective measures to supervise the document collection process, which is generally conducted manually. However, manual quality inspection is time-consuming, labor-intensive, inefficient, and prone to errors. Furthermore, it is a post-event audit and cannot provide strong support for the document collection process.
[0065] To address the aforementioned issues, this embodiment, in response to a user's instruction to confirm card issuance information, records the user's card collection process after successful card issuance, obtaining video and audio data. Facial recognition is then performed on the video data to obtain the recognition result. When the recognition result successfully matches the facial data corresponding to the card issuance information, action recognition is performed on the video data. Furthermore, when it is determined that the recognized action corresponding to the video data is a normal card collection action, and the audio data is confirmed to be compliant, the user's card collection process is confirmed to be their own. This embodiment utilizes facial recognition, video action recognition, and voice recognition for intelligent quality inspection, which can solve the problems of slow and error-prone manual review, standardize and ensure the entire card collection process, and provide strong support for resolving issues such as lost or mis-collected documents and business compliance.
[0066] In this embodiment, video and audio data can be sent to an intelligent dual-recording quality inspection and auditing platform, which can be deployed on an execution device. In some scenarios, the execution device can be a self-service kiosk, which also includes a data acquisition device that can acquire video data in real time and send it to the intelligent dual-recording quality inspection and auditing platform for face recognition and video action recognition. In other scenarios, the execution device can also be a server, and the intelligent dual-recording quality inspection and auditing platform can be deployed on the server. The self-service kiosk can send the real-time acquired video data to the server, and the intelligent dual-recording quality inspection and auditing platform deployed on the server can then perform face recognition and video action recognition.
[0067] This application provides an intelligent dual-recording quality inspection method. Figure 1 The flowchart of the intelligent dual-recording quality inspection method is illustrated executively. This process can be executed by an execution device, which can be a self-service kiosk machine or a server. The specific process is as follows:
[0068] 101. In response to the user's instruction to confirm the card production information, after confirming that the card production was successful, the process of the user receiving the card is recorded to obtain video and audio data.
[0069] In some embodiments, after the user confirms the card issuance information, they can click the card issuance confirmation control on the self-service kiosks. Then, in response to the user's confirmation instruction and after confirming successful card issuance based on the user's information, the process of the user collecting the card can be recorded to obtain video and audio data. Specifically, the process of the user collecting the card can be recorded using audio and video capture devices or recording tools, including capturing images of the card dispensing area.
[0070] In some embodiments, audio and video data to be quality inspected can be generated according to the requirements of format, resolution, and recording duration.
[0071] 102. Perform face recognition on the video data to obtain the recognition results of the video data.
[0072] In some embodiments, video data can be sent to an intelligent dual-recording quality inspection and auditing platform, which then performs facial recognition on the video data. Specifically, the video data can be input into a neural network to perform facial recognition on the video data, thereby obtaining the facial recognition results.
[0073] 103. When the recognition result successfully matches the facial data corresponding to the card information, perform action recognition on the video data.
[0074] In some embodiments, the facial data corresponding to the card issuance information can be sent to an intelligent dual-recording quality inspection and auditing platform, and the recognition result can be matched with the facial data corresponding to the card issuance information. In some scenarios, when the recognition result successfully matches the facial data corresponding to the card issuance information, it can be determined that the user is the one who received the card, and then action recognition can be performed on the video data.
[0075] In some embodiments, action recognition of video data can be achieved in the following ways:
[0076] The video data is input into a first neural network for feature extraction, determining the feature matrices of the human torso corresponding to each of the multiple images in the video data. Further, the feature matrices of the human torso corresponding to each of the multiple images can be input into a second neural network to determine the recognition action corresponding to the video data.
[0077] In some embodiments, the first neural network and the second neural network constitute the overall neural network model, see [link to relevant documentation]. Figure 2 As shown, the overall neural network model can be obtained in the following way:
[0078] 201, Obtain the training sample set.
[0079] The training sample set includes multiple training samples labeled as normal card-collecting actions and multiple training samples labeled as abnormal card-collecting actions.
[0080] 202. For each training sample, input each training sample into the first neural network to be trained for feature extraction, and obtain the feature matrix of the human torso corresponding to the multiple images included in each training sample.
[0081] 203. Input the feature matrices of the human torso corresponding to multiple images into the second neural network to be trained, and determine the predicted recognition action corresponding to each training sample.
[0082] 204. Determine the loss value based on the predicted recognition action and label corresponding to each training sample.
[0083] 205. The network parameters of the first neural network to be trained and the second neural network to be trained are adjusted by using the loss value to obtain the overall neural network model.
[0084] In some embodiments, face recognition and action recognition can be performed on video data using a total neural network model to determine the face recognition results and action recognition results.
[0085] 104. Once it is determined that the recognition action corresponding to the video data is a normal card-collecting action, and the audio data is recognized as compliant, the process of the user collecting the card is determined to be the user collecting the card themselves.
[0086] In some embodiments, after performing action recognition on the video data using a general neural network model and determining that the action corresponding to the video data is a normal card-collecting action, the video data can be identified. As an example, audio data identification can be achieved by using a third neural network to identify the audio data and obtain the corresponding audio content.
[0087] Furthermore, the compliance of audio data can be determined based on its content. Specifically, audio data compliance can be determined by comparing the audio content with preset first audio data and preset second audio data. The first audio data includes the audio data required during the card retrieval process, while the second audio data includes audio data prohibited during the card retrieval process. For example, the first audio data may include the user's own promises and specific introductory remarks from the salesperson, while the second audio data may include misleading or coercive language used during the card retrieval process. Further, if the audio content contains audio data from the first audio data but does not contain audio data from the second audio data, the audio content is deemed compliant.
[0088] In some embodiments, once the audio data is determined to be compliant, the process of the user collecting the card is determined to be the user collecting the card themselves.
[0089] Based on the above solution, the process of users receiving cards is recorded to obtain video and audio data. The audio and video data, along with the facial information corresponding to the card production information, are uploaded to the intelligent dual-recording quality inspection and audit platform. Intelligent quality inspection is carried out through facial recognition, motion recognition, and audio recognition to ensure business compliance and that the certificate is received by the person in question. This provides strong support for standardizing and safeguarding the entire certificate production and supervision process and resolving issues such as lost or mis-received certificates and business compliance.
[0090] In some embodiments, the acquisition device can send the acquired video data to the intelligent dual-recording quality inspection and auditing platform in real time. Upon receiving the video data, the platform can perform facial and motion recognition on the video data in real time. In some scenarios, when the recognition result fails to match the facial data, it is determined that the user's card-collecting process was not performed by the account holder, and a first alarm is generated. At this time, the card-dispensing action of the self-service kiosks can be paused to prevent the prepared cards from being collected by someone other than the user.
[0091] In other scenarios, when the recognized action corresponding to the video data is determined to be an abnormal card-collecting action, the user's card-collecting process is deemed abnormal, and a second alarm is generated. This alarm can be displayed as an alarm control on the screen, or it can be triggered via voice prompts or ringing.
[0092] In other scenarios, when the audio data is non-compliant, it is determined that the language used by the user during the card collection process is non-compliant, and a third alarm prompt is generated.
[0093] Based on the above solution, it can be ensured that the applicant receives the certificate at the certificate issuance point, and the post-event quality inspection can be moved to real-time quality inspection during the process, ensuring that the business is compliant and the certificate is received by the person in charge. This guarantees that every business operation meets the requirements, solves the problem of remediation after the fact, and improves service quality.
[0094] Based on the same technical concept, this application provides an intelligent dual-recording quality inspection device 300, see [link to relevant documentation]. Figure 3 As shown. The device 300 can perform any step in the above-described intelligent dual-recording quality inspection method. To avoid repetition, it will not be described again here. The device 300 includes a data acquisition module 301, an identification module 302, and a determination module 303.
[0095] The acquisition module 301 is used to respond to the user's instruction to confirm the card production information. After confirming that the card production is successful, it records the process of the user receiving the card to obtain video data and audio data.
[0096] The recognition module 302 is used to perform face recognition on the video data to obtain the recognition result of the video data;
[0097] When the recognition result is successfully matched with the facial data corresponding to the card information, action recognition is performed on the video data;
[0098] The determination module 303 is used to determine that the user's card collection process is a personal collection process after determining that the recognition action corresponding to the video data is a normal card collection action and after recognizing the audio data to determine that the audio data is compliant.
[0099] In some embodiments, the recognition module 302, when performing action recognition on the video data, is specifically used for:
[0100] The video data is input into a first neural network for feature extraction to determine the feature matrix of the human torso corresponding to the multiple images included in the video data.
[0101] The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to determine the recognition action corresponding to the video data.
[0102] In some embodiments, the recognition module 302, when recognizing the audio data, is specifically used for:
[0103] The audio data is identified using a third neural network to obtain the audio content corresponding to the audio data;
[0104] The determining module 303, when determining that the audio data is compliant, is specifically used for:
[0105] The audio content is compared with preset first audio data and preset second audio data respectively. The first audio data is the audio data that needs to be included in the card retrieval process, and the second audio data includes audio data that is prohibited from appearing in the card retrieval process.
[0106] When it is determined that the audio content contains audio data from the first audio data and the audio content does not contain audio data from the second audio data, the audio content is determined to be compliant.
[0107] In some embodiments, the determining module 303 is further configured to:
[0108] When the recognition result fails to match the facial data, it is determined that the user's card-collecting process was not performed by the user themselves, and a first alarm message is generated; or,
[0109] When it is determined that the recognition action corresponding to the video data is an abnormal card-collecting action, the user's card-collecting process is determined to be abnormal, and a second alarm prompt is generated; or,
[0110] When the audio data is non-compliant, it is determined that the language used during the user's card collection process is non-compliant, and a third alarm prompt is generated.
[0111] In some embodiments, the first neural network and the second neural network constitute a total neural network model, which is obtained in the following manner:
[0112] Obtain a training sample set, which includes multiple training samples labeled as normal card-collecting actions and multiple training samples labeled as abnormal card-collecting actions.
[0113] For each training sample, each training sample is input into the first neural network to be trained for feature extraction, thereby obtaining the feature matrix of the human torso corresponding to the multiple images included in each training sample.
[0114] The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to be trained to determine the predicted recognition action corresponding to each training sample.
[0115] The loss value is determined based on the predicted recognition action and label corresponding to each training sample;
[0116] The network parameters of the first neural network to be trained and the second neural network to be trained are adjusted using the loss value to obtain the overall neural network model.
[0117] Based on the same technical concept, this application provides an execution device 400, see [link to relevant documentation]. Figure 4 As shown. The device 400 can perform any step in the above-described intelligent dual-recording quality inspection method. To avoid repetition, it will not be described again here. The device 400 includes a memory 401 and a processor 402.
[0118] Memory 401 is used to store program instructions;
[0119] The processor 402 is used to call the program instructions stored in the memory and execute any step included in the above-mentioned intelligent dual-recording quality inspection method according to the obtained program instructions.
[0120] In the embodiments of this application, the processor 402 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, capable of implementing or executing the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.
[0121] Memory 401, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 401 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 401 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 401 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0122] Based on the same technical concept, embodiments of this application provide a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by a computer, cause the computer to perform the intelligent dual-recording quality inspection method described above. Since the principle by which the above-described computer-readable storage medium solves the problem is similar to that of the intelligent dual-recording quality inspection method, the implementation of the above-described computer-readable storage medium can be referred to the implementation of the method; repeated details will not be elaborated further.
[0123] Based on the same technical concept, this application provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to execute the intelligent dual-recording quality inspection method described above. Since the principle by which the above computer program product solves the problem is similar to that of the intelligent dual-recording quality inspection method, the implementation of the above computer program product can refer to the implementation of the method, and repeated details will not be repeated.
[0124] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0125] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0128] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. An intelligent dual-recording quality inspection method, characterized in that, This method is applied to self-service kiosks, and the method includes: In response to the user's instruction to confirm the card production information, after confirming that the card production was successful, the process of the user receiving the card is recorded in real time to obtain video and audio data; The video data is subjected to facial recognition to obtain the recognition result of the video data; When the recognition result does not match the facial data corresponding to the card issuance information, a first alarm is generated and the card dispensing action on the self-service kiosks is paused. When the recognition result is successfully matched with the facial data corresponding to the card information, action recognition is performed on the video data; Once it is determined that the recognition action corresponding to the video data is a normal card-collecting action, and the audio data is identified as compliant, the process of the user collecting the card is determined to be the user collecting the card themselves.
2. The method as described in claim 1, characterized in that, The action recognition process for the video data includes: The video data is input into a first neural network for feature extraction to determine the feature matrix of the human torso corresponding to the multiple images included in the video data. The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to determine the recognition action corresponding to the video data.
3. The method as described in claim 1, characterized in that, The process of identifying the audio data includes: The audio data is identified using a third neural network to obtain the audio content corresponding to the audio data; Determining the compliance of the audio data includes: The audio content is compared with preset first audio data and preset second audio data respectively. The first audio data is the audio data that needs to be included in the card retrieval process, and the second audio data includes audio data that is prohibited from appearing in the card retrieval process. When it is determined that the audio content contains audio data from the first audio data and the audio content does not contain audio data from the second audio data, the audio content is determined to be compliant.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: When the recognition result fails to match the facial data, it is determined that the user's card-collecting process was not performed by the user themselves, and a first alarm message is generated; or, When it is determined that the recognition action corresponding to the video data is an abnormal card-collecting action, the user's card-collecting process is determined to be abnormal, and a second alarm prompt is generated; or, When the audio data is non-compliant, it is determined that the language used during the user's card collection process is non-compliant, and a third alarm prompt is generated.
5. The method as described in claim 2, characterized in that, The first neural network and the second neural network constitute a total neural network model, which is obtained in the following manner: Obtain a training sample set, which includes multiple training samples labeled as normal card-collecting actions and multiple training samples labeled as abnormal card-collecting actions. For each training sample, each training sample is input into the first neural network to be trained for feature extraction, thereby obtaining the feature matrix of the human torso corresponding to the multiple images included in each training sample. The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to be trained to determine the predicted recognition action corresponding to each training sample. The loss value is determined based on the predicted recognition action and label corresponding to each training sample; The network parameters of the first neural network to be trained and the second neural network to be trained are adjusted using the loss value to obtain the total neural network model.
6. An intelligent dual-recording quality inspection device, characterized in that, The device is deployed on a self-service kiosks, and the device includes: The acquisition module is used to respond to the user's instruction to confirm the card production information. After confirming that the card production is successful, it records the user's card collection process in real time to obtain video and audio data. The recognition module is used to perform facial recognition on the video data to obtain the recognition result of the video data; When the recognition result does not match the facial data corresponding to the card issuance information, a first alarm is generated and the card dispensing action on the self-service kiosks is paused. When the recognition result is successfully matched with the facial data corresponding to the card information, action recognition is performed on the video data; The determination module is used to determine that the user's card collection process is a personal collection process after determining that the recognition action corresponding to the video data is a normal card collection action and after recognizing the audio data to determine that the audio data is compliant.
7. The apparatus as claimed in claim 6, characterized in that, The recognition module, when performing action recognition on the video data, is specifically used for: The video data is input into a first neural network for feature extraction to determine the feature matrix of the human torso corresponding to the multiple images included in the video data. The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to determine the recognition action corresponding to the video data.
8. The apparatus as claimed in claim 6, characterized in that, The recognition module, when recognizing the audio data, is specifically used for: The audio data is identified using a third neural network to obtain the audio content corresponding to the audio data; The determining module, when determining that the audio data is compliant, is specifically used for: The audio content is compared with preset first audio data and preset second audio data respectively. The first audio data is the audio data that needs to be included in the card retrieval process, and the second audio data includes audio data that is prohibited from appearing in the card retrieval process. When it is determined that the audio content contains audio data from the first audio data and the audio content does not contain audio data from the second audio data, the audio content is determined to be compliant.
9. The apparatus according to any one of claims 6-8, characterized in that, The determining module is also used for: When the recognition result fails to match the facial data, it is determined that the user's card-collecting process was not performed by the user themselves, and a first alarm message is generated; or, When it is determined that the recognition action corresponding to the video data is an abnormal card-collecting action, the user's card-collecting process is determined to be abnormal, and a second alarm prompt is generated; or, When the audio data is non-compliant, it is determined that the language used during the user's card collection process is non-compliant, and a third alarm prompt is generated.
10. The apparatus as claimed in claim 7, characterized in that, The first neural network and the second neural network constitute a total neural network model, which is obtained in the following manner: Obtain a training sample set, which includes multiple training samples labeled as normal card-collecting actions and multiple training samples labeled as abnormal card-collecting actions. For each training sample, each training sample is input into the first neural network to be trained for feature extraction, thereby obtaining the feature matrix of the human torso corresponding to the multiple images included in each training sample. The feature matrices of the human torso corresponding to the multiple images are input into the second neural network to be trained to determine the predicted recognition action corresponding to each training sample. The loss value is determined based on the predicted recognition action and label corresponding to each training sample; The network parameters of the first neural network to be trained and the second neural network to be trained are adjusted using the loss value to obtain the total neural network model.
11. An execution device, characterized in that, include: Memory, used to store program instructions; A processor is configured to retrieve program instructions from the memory and execute the method of any one of claims 1-5 according to the retrieved program instructions.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Double-record quality inspection method and device, computer device and storage medium
CN109767335A
Human body behavior identification method based on residual-recurrent neural network
CN111814661A
Audio and video quality inspection processing method and system
CN115631448A