Single-direction video-based witness account opening method and device and storage medium

By performing facial movement, voice liveness detection, and document image detection during the one-way video witnessing account opening process, recording complete audio and video information and sending it to the server for analysis, the problem of low authenticity and reliability of user intentions is solved, and the smoothness of operation and review efficiency are improved.

CN115410106BActive Publication Date: 2026-07-21北京中关村科金技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
北京中关村科金技术有限公司
Filing Date
2021-05-11
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In existing technologies, the one-way video witnessing account opening method cannot ensure the authenticity and reliability of the user's intentions, the operation process is inconvenient, the workload of the review personnel is heavy, and there is a risk of fraud.

Method used

The client records the audio and video information of the entire account opening process of the target object, performs face action detection, voice liveness detection, and document image quality detection to ensure the integrity of the recording, and sends the results to the server for comprehensive analysis after the detection results meet the requirements.

Benefits of technology

It ensures the authenticity and reliability of user intentions, reduces the risk of fraud, improves operational fluency and review efficiency, and safeguards the legitimate rights and interests of securities companies and users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115410106B_ABST
    Figure CN115410106B_ABST
Patent Text Reader

Abstract

The application discloses a one-way video-based witnessing account opening method and device and a storage medium. The method comprises the following steps: a client starts recording audio and video information, wherein the audio and video information is used for recording the whole account opening process of a target object; in the process of recording the audio and video information, the client acquires face image information and voice information of the target object, performs face action detection on the face image information, performs voice living body detection on the voice information, judges whether the target object is a living body according to the face action detection result, and judges whether the target object is a living body according to the voice living body detection result; the client acquires certificate image information corresponding to a certificate displayed by the target object, performs quality detection on the certificate image information, and ends the recording of the audio and video information in the case that the obtained quality detection result meets the quality requirement; and the client sends the recorded audio and video information to a server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular to a witness account opening method, apparatus and storage medium based on one-way video. Background Technology

[0002] On February 9, 2021, China Securities Depository and Clearing Corporation Limited (CSDC) revised and released the "Implementation Rules for Non-Face-to-Face Securities Account Opening," clarifying that securities firms can use one-way video account opening. This means that the pilot program has been officially and fully opened, allowing each securities firm to choose to activate this service, making one-way video account opening a focus of attention. One-way video account opening, because it does not rely on online employee witnessing, allows for 24 / 7 account opening with good service continuity; it reduces the need for more verification staff when dealing with large volumes of accounts, resulting in lower costs; it is easy to drive traffic to accounts, as it can be embedded purely in H5; and the system architecture is relatively simple, with low operation and maintenance costs. This facilitates quick and easy verification of customer identity and the authenticity of their account opening intentions.

[0003] However, the existing method of opening accounts through video verification has several limitations. On the one hand, it selects one or a few frames of images and sends them to the network device for facial recognition. One or a few frames of images cannot witness the entire account opening process of the user, cannot determine the authenticity of the user's intentions, and cannot determine whether there are any abnormalities in the account opening process when the auditor reviews it. Therefore, it cannot serve as effective evidence to prevent users from denying their rights, which creates management risks and cannot ensure the legitimate rights and interests of securities companies and users.

[0004] On the other hand, during the entire witnessing process, the uploading of documents such as ID cards and bank cards and the recording of user audio and video are not continuous and there are interruptions. The entire account opening process cannot be recorded, and the user's true intentions cannot be expressed. The reliability is low and there is a risk of fraud.

[0005] Furthermore, if the reviewer discovers an error during the review process after the witness video is sent, the user must start from the beginning, redo the video review, and then send a new video to the reviewer. This creates an inconvenient workflow and may lead to customer loss. Additionally, the lack of processing of uploaded audio and video significantly increases the workload of the reviewers.

[0006] There are currently no effective solutions to the technical problems existing in the above-mentioned technologies, such as the inability to verify the authenticity of users' intentions, low reliability, inconvenient operation process, and heavy workload of auditing personnel when using video witnessing for account opening. Summary of the Invention

[0007] The embodiments of this application provide a method, apparatus, and storage medium for account opening based on one-way video witnessing, so as to at least solve the technical problems existing in the prior art, such as the inability to determine the authenticity of the user's intentions, low reliability, inconvenient operation process, and heavy workload of the review personnel when using video witnessing for account opening.

[0008] According to one aspect of the embodiments of this application, a witness account opening method based on one-way video is provided, applied to a witness account opening system. The witness account opening system includes a client and a server, comprising: the client responding to a start request for recording audio and video information sent by a target object, starting to record audio and video information, wherein the audio and video information is used to record the entire account opening process of the target object; during the recording of audio and video information, the client acquires the facial image information of the target object, performs facial motion detection on the facial image information, and determines whether the target object is a live object based on the obtained facial motion detection result; during the recording of audio and video information, the client acquires the voice information of the target object, performs voice liveness detection on the voice information, and determines whether the target object is a live object based on the obtained voice liveness detection result; during the recording of audio and video information, the client acquires the document image information corresponding to the document displayed by the target object, performs quality detection on the document image information, and ends the recording of audio and video information if the obtained quality detection result meets the quality requirements; and the client sends the recorded audio and video information to the server.

[0009] According to another aspect of the embodiments of this application, a storage medium is also provided, the storage medium including a stored program, wherein, when the program is running, the processor executes any of the methods described above.

[0010] According to another aspect of the embodiments of this application, a witness account opening device based on one-way video is also provided, applied to a witness account opening system. The witness account opening system includes a client and a server, including: a start request response module, used by the client to respond to a start request sent by the target object to record audio and video information, and start recording audio and video information, wherein the audio and video information is used to record the entire account opening process of the target object; a face action detection module, used by the client to acquire the face image information of the target object during the recording of audio and video information, perform face action detection on the face image information, and determine whether the target object is a live object based on the obtained face action detection result; a voice liveness detection module, used by the client to acquire the voice information of the target object during the recording of audio and video information, perform voice liveness detection on the voice information, and determine whether the target object is a live object based on the obtained voice liveness detection result; a document image information acquisition module, used by the client to acquire the document image information corresponding to the document displayed by the target object during the recording of audio and video information, perform quality detection on the document image information, and end the recording of audio and video information if the obtained quality detection result meets the quality requirements; and a sending module, used by the client to send the recorded audio and video information to the server.

[0011] According to another aspect of the embodiments of this application, a witness account opening device based on one-way video is also provided, applied to a witness account opening system. The witness account opening system includes a client and a server, including: a processor; and a memory connected to the processor, used to provide the processor with instructions to process the following steps: the client responds to a start request for recording audio and video information sent by the target object, and starts recording audio and video information, wherein the audio and video information is used to record the entire account opening process of the target object; during the recording of audio and video information, the client obtains the facial image information of the target object, performs facial motion detection on the facial image information, and determines whether the target object is a live object based on the obtained facial motion detection result; during the recording of audio and video information, the client obtains the voice information of the target object, performs voice liveness detection on the voice information, and determines whether the target object is a live object based on the obtained voice liveness detection result; during the recording of audio and video information, the client obtains the document image information corresponding to the document displayed by the target object, performs quality detection on the document image information, and ends the recording of audio and video information if the obtained quality detection result meets the quality requirements; and the client sends the recorded audio and video information to the server.

[0012] In this embodiment, the client records audio and video information to document the entire account opening process for the target object. During the recording process, facial motion detection and voice liveness detection are performed on the target object. Based on the facial motion detection results and voice liveness detection results, the client determines whether the target object is alive. Additionally, the client performs quality checks on the document image information during recording. Recording ends when the quality check results meet the requirements. The client then sends the recorded audio and video information to the server, which performs comprehensive analysis and determines whether to provide account opening services to the target object based on the analysis results.

[0013] Therefore, this application provides a one-way video-based account opening witnessing method. The client records the entire account opening process of the target (including document display and audio / video input), thus obtaining complete audio / video information. This audio / video information can witness the entire account opening process of the target, ensuring the integrity of the audio / video recording and verifying the authenticity of the target's intention to open an account. Furthermore, during the recording process, multiple detection methods, such as facial motion detection and voice liveness detection, are used to verify the target's liveness, ensuring the authenticity of the user's intention and reducing the risk of fraud. In addition, the client monitors the user's actions at each stage in real time during recording to ensure compliance. When errors are detected, the user is promptly reminded to correct them until a satisfactory audio / video recording is completed. Compared to the method where reviewers discover errors during the witnessing process and require the user to re-witness the video, this application can monitor the user's actions at each stage in real time during the recording process and promptly remind the user to correct errors, effectively reducing the user's operational steps and providing a smoother user experience.

[0014] Furthermore, after recording the entire audio and video information, the client no longer sends only one or a few frames to the server, but instead sends the complete audio and video information. This allows reviewers to determine if there are any anomalies in the account opening process based on this complete audio and video information during the review. The complete audio and video information also serves as valid evidence to prevent users from denying wrongdoing, thus ensuring the legitimate rights and interests of both the securities company and the user. Moreover, after receiving the complete audio and video information from the client, the server performs comprehensive analysis on it. The results of this analysis assist the reviewers in their review process. Compared to directly sending unanalyzed audio and video information to the reviewers, the server-side audio and video analysis method effectively assists the reviewers in their review work, greatly improving the efficiency and accuracy of the review process.

[0015] In summary, this technical solution records the uploading of ID cards and bank cards, along with uninterrupted audio and video recording of the user during the witnessing process, as a complete audio-visual message. This ensures the integrity of the audio-visual recording, improves the reliability of expressing the user's true intentions, and reduces the risk of forgery. After collecting the complete audio-visual information used to record the entire account opening process, the client sends this information to the server to prevent user denial and protect the legitimate rights and interests of the securities company. The client and server combine multiple liveness detection algorithms, working collaboratively to prevent unauthorized operation, ensuring security and controllability, guaranteeing the account opening is based on the user's true intentions, and protecting the legitimate rights and interests of the securities company. Furthermore, the server-side audio-visual analysis method effectively assists reviewers in their verification work, greatly improving the efficiency and accuracy of the review process. This solves the technical problems of existing technologies that, when using video witnessing for account opening, cannot determine the authenticity of the user's intentions, have low reliability, inconvenient operation procedures, and impose a heavy workload on reviewers. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0017] Figure 1 This is a hardware structure block diagram of a computing device used to implement the method described in Embodiment 1 of this application;

[0018] Figure 2 This is a schematic diagram of a witness account opening system based on one-way video, as described in Embodiment 1 of this application.

[0019] Figure 3 This is a flowchart illustrating the witness account opening method based on one-way video according to the first aspect of Embodiment 1 of this application;

[0020] Figure 4 This is another schematic flowchart of the witness account opening method based on one-way video as described in the first aspect of Embodiment 1 of this application;

[0021] Figure 5 This is a schematic diagram of the process by which the server performs comprehensive analysis of audio and video information according to the first aspect of Embodiment 1 of this application;

[0022] Figure 6 This is a schematic diagram of a witness account opening device based on one-way video according to Embodiment 2 of this application; and

[0023] Figure 7This is a schematic diagram of a witness account opening device based on one-way video as described in Embodiment 3 of this application. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0026] Example 1

[0027] According to this embodiment, a witness account opening method based on one-way video is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0028] The method embodiments provided in this example can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Figure 1 A hardware block diagram of a computing device for implementing a one-way video-based witness account opening method is shown. Figure 1 As shown, a computing device may include one or more processors (processors may include, but are not limited to, microprocessors such as MCUs or programmable logic devices such as FPGAs), memory for storing data, and transmission devices for communication functions. In addition, it may also include: a display, input / output interfaces (I / O interfaces), a universal serial bus (USB) port (which may be included as one of the ports in the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, a computing device may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0029] It should be noted that the aforementioned one or more processors and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits can be implemented wholly or partially as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits can be a single, independent processing module, or wholly or partially integrated into any other element in the computing device. As involved in the embodiments of this application, the data processing circuit serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0030] The memory can be used to store software programs and modules of application software, such as the program instruction / data storage device corresponding to the one-way video-based witness account opening method in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the one-way video-based witness account opening method of the aforementioned application. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the computing device via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0031] The transmission device is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the computing device's communications provider. In one example, the transmission device includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0032] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows users to interact with the user interface of the computing device.

[0033] It should be noted here that, in some optional embodiments, the above... Figure 1The computing device shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computing devices.

[0034] Figure 2 This is a schematic diagram of the witness account opening system based on one-way video as described in this embodiment. (Refer to...) Figure 2 As shown, the system includes a client 100, a server 200, and a data storage device 300. The client 100 records audio and video information to document the entire account opening process of the target object, and then sends this audio and video information to the server 200. The server 200 then performs comprehensive analysis on the audio and video information and writes the analysis results into the data storage device 300.

[0035] It should be noted that the hardware structure described above can be used for the client 100, server 200, and data storage device 300 in the system.

[0036] In the above operating environment, according to the first aspect of this embodiment, a witness account opening method based on one-way video is provided, applied to a witness account opening system, which includes a client and a server. The method is... Figure 2 The client 100 and server 200 shown are implemented. Figure 3 A flowchart illustrating the method is shown below. (Refer to...) Figure 3 As shown, the method includes:

[0037] S302: The client responds to the start request for recording audio and video information sent by the target object and begins recording audio and video information, which is used to record the entire account opening process of the target object;

[0038] S304: During the recording of audio and video information, the client obtains the facial image information of the target object, performs facial motion detection on the facial image information, and determines whether the target object is a live object based on the obtained facial motion detection results;

[0039] S306: During the recording of audio and video information, the client obtains the voice information of the target object, performs voice liveness detection on the voice information, and determines whether the target object is alive based on the obtained voice liveness detection result;

[0040] S308: During the recording of audio and video information, the client obtains the image information of the ID card displayed by the target object, performs quality checks on the ID card image information, and terminates the recording of audio and video information if the obtained quality check result meets the quality requirements; and

[0041] S310: The client sends the recorded audio and video information to the server.

[0042] Specifically, client 100 responds to the start request for recording audio and video information sent by target object 110 and begins recording audio and video information. This audio and video information is used to record the entire account opening process of target object 110. Upon receiving the start request for recording audio and video information sent by target object 110, client 100 responds to the request and begins recording audio and video information. During the recording process, target object 110 will perform corresponding actions based on prompts from client 100 (S302).

[0043] Furthermore, during the recording of audio and video information, the client 100 acquires the facial image information of the target object 110, performs facial motion detection on the facial image information, and determines whether the target object 110 is a live person based on the obtained facial motion detection results. Specifically, during the recording of audio and video information, the client 100 detects that the face of the target object 110 is in a screen prompt box and provides a prompt to the target object 110, such as "blink" or "nod". The target object 110 performs the specified action according to the prompt, the client 100 acquires the facial image information of the target object 110, performs facial motion detection on the facial image information, and then determines whether the facial motion of the target object 110 matches the prompt information based on the facial motion detection results. If the obtained facial motion detection results show that the facial motion of the target object 110 matches the prompt information, the client 110 determines that the target object 110 is a live person. If the face action detection result shows that the face action of the target object 110 does not match the prompt information, the client 110 determines that the target object 110 is not a live object (S304).

[0044] Furthermore, during the recording of audio and video information, the client 100 acquires the voice information of the target object 110, performs voice liveness detection on the voice information, and determines whether the target object 110 is alive based on the obtained voice liveness detection result. Specifically, during the recording of audio and video information, the target object 110 generates voice information. After the client 100 determines that the target object 110 is alive through facial movements, it acquires the voice information of the target object 110 and performs voice liveness detection on the voice information using methods such as voiceprint recognition to obtain a voice liveness detection result. Then, the client 100 determines whether the target object 110 is alive based on the voice liveness detection result. If the voice liveness detection result matches human voice, the client 100 determines that the target object 110 is alive, that is, the target object 110 is a human. Otherwise, the client 100 determines that the target object 110 is not alive, that is, the target object 110 is not a human (S306).

[0045] Furthermore, during the recording of audio and video information, the client 100 acquires the image information of the document corresponding to the document displayed by the target object 110, performs quality detection on the document image information, and ends the recording of audio and video information if the obtained quality detection result meets the quality requirements. Specifically, after determining that the target object 110 is a living person based on the voice information of the target object 110, the client 100 prompts the target object 110 to enter the document photo display stage. During the recording of audio and video information, the client 100 prompts the target object 110 to hold the document (such as an ID card and bank card). Then, the client 100 acquires the image information of the document corresponding to the document displayed by the target object 110, uses a document photo quality detection algorithm to perform quality detection on the document image information in real time, and determines whether the document image information meets the quality requirements based on the quality detection result. When the document image information does not meet the quality requirements, the client 100 immediately requests the target object 110 to re-display the document until the quality detection result meets the quality requirements, after which the client 100 ends the recording of audio and video information (S308). In addition, the client 100 will perform face detection in real time throughout the recording process and provide face-in-frame prompts to ensure that the face of the target object 110 is recorded in the video throughout the entire process.

[0046] Furthermore, after recording the audio and video information, the client 100 sends the recorded audio and video information to the server 200 (S310).

[0047] As described in the background section, the existing method of opening an account through video witnessing has several limitations. On the one hand, it selects one or a few frames of images and sends them to a network device for facial recognition. One or a few frames of images cannot witness the entire account opening process of the user, cannot determine the authenticity of the user's intentions, and cannot determine whether there are any abnormalities in the account opening process when the auditor reviews it. Therefore, it cannot serve as effective evidence to prevent users from denying their rights, which creates management risks and cannot ensure the legitimate rights and interests of securities companies and users.

[0048] On the other hand, during the entire witnessing process, the uploading of documents such as ID cards and bank cards and the recording of user audio and video are not continuous and there are interruptions. The entire account opening process cannot be recorded, and the user's true intentions cannot be expressed. The reliability is low and there is a risk of fraud.

[0049] Furthermore, if the reviewer discovers an error during the witnessing process after the video is sent, the user must start from the beginning, redo the video witnessing, and then send a new video to the reviewer. This creates an inconvenient workflow and may lead to customer churn.

[0050] To address the aforementioned technical problems, the technical solution of this application embodiment involves a client recording audio and video information to document the entire account opening process of a target object. During the recording process, facial motion detection and voice liveness detection are performed on the target object. Based on the results of these two detections, the client determines whether the target object is alive. Furthermore, the client also performs quality checks on the document image information during recording. Recording ends when the quality check results meet the requirements. The client then sends the recorded audio and video information to the server.

[0051] Therefore, this application provides a one-way video-based account opening witnessing method. The client records the entire account opening process of the target (including document display and audio / video input), thus obtaining complete audio / video information. This audio / video information can witness the entire account opening process of the target, ensuring the integrity of the audio / video recording and verifying the authenticity of the target's intention to open an account. Furthermore, during the recording process, multiple detection methods, such as facial motion detection and voice liveness detection, are used to verify the target's liveness, ensuring the authenticity of the user's intention and reducing the risk of fraud. In addition, the client monitors the user's actions at each stage in real time during recording to ensure compliance. When errors are detected, the user is promptly reminded to correct them until a satisfactory audio / video recording is completed. Compared to the method where reviewers discover errors during the witnessing process and require the user to re-witness the video, this application can monitor the user's actions at each stage in real time during the recording process and promptly remind the user to correct errors, effectively reducing the user's operational steps and providing a smoother user experience.

[0052] Furthermore, after the client completes the recording of the entire audio and video information, it no longer selects one or a few frames of video images to send to the server, but sends a complete audio and video message to the server. This allows the reviewers to determine whether there are any abnormalities in the entire account opening process based on this complete audio and video message during the review. At the same time, the complete audio and video message can also serve as effective evidence to prevent users from denying their rights, thereby ensuring the legitimate rights and interests of securities companies and users.

[0053] In summary, this technical solution records the uploading of ID cards and bank cards, along with uninterrupted audio and video recording of the user during the witnessing process, as a single, complete audio-visual message. This ensures the integrity of the audio-visual recording, improves the reliability of expressing the user's true intentions, and reduces the risk of forgery. After collecting the complete audio-visual information used to record the entire account opening process, the client sends this information to the server to prevent user denial and protect the legitimate rights and interests of the securities company. This solves the technical problem of low reliability in existing technologies that cannot verify the authenticity of the user's intentions when using video witnessing for account opening.

[0054] Optionally, the method further includes: during the recording of audio and video information, the client uses a speech recognition algorithm to perform speech recognition detection on the speech information; and based on the obtained speech recognition results, it determines whether the target object reads out the prompt content for account opening as required.

[0055] Specifically, refer to Figure 4As shown, during the recording of audio and video information, client 100 displays a text message, which is a prompt for account opening. Client 100 requests target 110 to read this prompt. Then, while target 110 reads the prompt, client 100 acquires the target 110's voice information. Client 100 then uses a speech recognition algorithm to perform speech recognition detection on this voice information. Specifically, the acquired voice information is converted into corresponding text information, and then compared with the prompt content to obtain the speech recognition result. When the similarity between the text information and the prompt content reaches a preset similarity value, such as greater than 80%, client 100 determines that target 110 has read the prompt for account opening as required. Otherwise, client 100 determines that target 110 has not read the prompt for account opening as required. Thus, during the recording of audio and video information, client 100 determines whether target 110 has read the prompt for account opening as required by performing speech recognition detection on the voice information. Therefore, when the target 110 reads the account opening prompts, the client 100 acquires the voice information in real time and performs voice recognition detection. If the client 100 determines, based on the voice recognition results, that the target 110 has not read the prompts as required, it will require the target 110 to immediately reread the prompts until the target 110's voice information meets the requirements. This ensures that the target 110 reads the prompts correctly. Compared to the method where reviewers discover errors in the user's voice information during the verification process and then require the user to re-verify via video, this application can detect the user's voice information in real time during the recording of audio and video information and promptly remind the user to correct any errors, thereby effectively reducing the user's operational steps and providing a smooth user experience.

[0056] Optionally, the method further includes: during the recording of audio and video information, the client performs voice emotion analysis on the voice information based on a voice emotion analysis algorithm; and based on the obtained voice emotion analysis results, determines whether the target object voluntarily opens an account.

[0057] Specifically, during the recording of audio and video information, client 100 acquires the voice information of target 110 and performs voice sentiment analysis on the voice information based on a voice sentiment analysis algorithm to obtain a voice sentiment analysis result regarding the emotional state of target 110. The voice sentiment analysis result is categorized into three types: positive, negative, and normal. When the obtained voice sentiment analysis result is positive or normal, client 100 determines that target 110 voluntarily opened the account. When the obtained voice sentiment analysis result is negative, it indicates that target 110 recorded the audio and video information with negative emotions such as fear, and client 100 determines that target 110 did not voluntarily open the account. Therefore, during the recording of audio and video information, client 100 determines whether target 110 voluntarily opened the account by performing voice sentiment analysis on the voice information. In this way, client 100 can detect the target subject 110's emotions in real time during the recording of audio and video information. When client 100 determines that the target subject 110's emotions are negative based on the obtained voice sentiment analysis results, it will require the target subject 110 to immediately adjust their emotions until client 100 detects that the voice sentiment analysis results of the target subject 110's voice information meet the requirements. This ensures that the target subject 110 is recording in an emotionally stable state. Compared to the method where reviewers discover abnormal emotions in the user during the witnessing process and then require the user to re-witness the video, this application can detect the user's emotions in real time during the recording of audio and video information and promptly remind the user to adjust their emotions when abnormal emotions are detected, thereby effectively reducing the user's operation process and providing a smooth user experience.

[0058] Optionally, the method further includes: the server performing comprehensive analysis on the audio and video information received from the client to obtain comprehensive analysis results; and the server determining whether to provide account opening services to the target object based on the comprehensive analysis results.

[0059] Specifically, refer to Figure 5 As shown, after the audio and video information is recorded, the client 100 sends the audio and video information to the server 200. The server 200 then receives the audio and video information and performs a comprehensive analysis to obtain the analysis result. Based on this result, the server 200 determines whether to provide account opening services to the target 110. If the analysis result meets the account opening requirements, the server 200 provides account opening services to the target 110. Otherwise, the server 200 does not provide account opening services to the target 110. Thus, through the collaborative work of the client and server, the workload of the review personnel is reduced, and work efficiency is improved.

[0060] Optionally, the server performs a comprehensive analysis of the audio and video information received from the client to obtain the comprehensive analysis results. This includes: the server performs face motion detection on the audio and video information to determine the time period of the face motion detection of the target object by the client and the best video frame of the target object's face; and the server writes the face motion time period and the best video frame of the face into a preset data storage device.

[0061] Specifically, server 200 performs facial motion detection on audio and video information, determining the start and end times of the facial motions detected by client 100 on target 110, and then determines the facial motion time period based on these times. Server 200 then acquires the best video frame of target 110's face within this facial motion detection time period. Server 200 then writes the facial motion time period and the best video frame to a pre-defined data storage device 300 (e.g., a database list). By determining the facial motion time period and the best video frame and writing them to the data storage device 300, this information can be provided to reviewers, facilitating the review process by allowing them to examine the video content within the facial motion time period based on the facial motion detection time points, thus improving review efficiency.

[0062] Optionally, the operation of the server performing comprehensive analysis on the audio and video information received from the client to obtain the comprehensive analysis results also includes: the server performing speech recognition on the audio and video information to determine the speech reading time period of the prompt content for account opening read by the target object and the speech recognition result corresponding to the read speech information; and the server writing the speech reading time period and the speech recognition result into a preset data storage device.

[0063] Specifically, server 200 performs speech recognition on the audio and video information to determine the start and end times of the speech input for the account opening prompts read to target object 110, and determines the speech input time period based on the start and end times. Server 200 also determines the speech recognition result corresponding to the read speech information. Then, server 200 writes the speech input time period and speech recognition result to a preset data storage device 300 (e.g., a database list). Thus, by determining the speech input time period in the audio and video information and writing it to the data storage device 300, server 200 can provide this information to reviewers, facilitating the review process by allowing them to view the video content within the speech input time period based on the speech input time points, thereby improving review efficiency.

[0064] Optionally, the operation of the server performing comprehensive analysis on the audio and video information received from the client to obtain the comprehensive analysis results also includes: the server performing ID photo detection on the audio and video information to determine the ID display time period for the target object and the best video frame of the ID; and the server writing the ID display time period and the best video frame of the ID to a preset data storage device.

[0065] Specifically, server 200 performs ID photo detection on the audio and video information to determine the start and end times of the ID display for target object 110. Based on the start and end times, it determines the ID display period and identifies the optimal video frame for the ID. Server 200 then writes the ID display period and optimal video frame into a pre-defined data storage device 300 (e.g., a database list). By determining the ID display period and writing it to data storage device 300, server 200 can provide this information to reviewers, facilitating the review process by allowing them to easily view the video content within the ID display period based on the displayed time points, thus improving review efficiency.

[0066] Optionally, the method further includes: the server performing OCR recognition on the document in the best video frame of the document to obtain the user information of the target object; the server performing face recognition on the face image on the document in the best video frame of the document, and determining whether the document is the document of the target object based on the face recognition result.

[0067] Specifically, server 200 performs OCR recognition on the document in the best video frame to obtain the user information of target object 110. For example, server 200 performs OCR recognition on the ID card image in the best video frame to identify the name and ID number on the ID card of target object 110. Then, server 200 performs face recognition on the face image on the document (e.g., ID card) in the best video frame to obtain the corresponding face recognition result. Server 200 compares the face recognition result with the best video frame of the face of target object 110 obtained during the face action time period to obtain a face comparison score. When the face comparison score reaches the required score threshold, server 200 determines that the document belongs to target object 110. Otherwise, server 200 determines that the document does not belong to target object 110. Server 200 writes the verification result and face comparison score into a preset data storage device 300 (e.g., a database list). This allows the system to determine whether the document belongs to the target 110 based on the facial recognition results, thus ensuring the authenticity of the user information.

[0068] Optionally, the method further includes: the server performing single-person frame detection on the audio and video information to obtain single-person frame detection results; the server performing multi-person frame detection on the audio and video information to obtain multi-person frame detection results; the server determining whether the target object meets the account opening operation requirements based on the single-person frame detection results and the multi-person frame detection results; and the server sending alarm information to the reviewer's terminal device if it determines that the target object does not meet the account opening operation requirements.

[0069] Specifically, server 200 performs single-person frame detection on the audio and video information, obtaining a single-person frame detection result. Server 200 also performs multi-person frame detection on the audio and video information, obtaining a multi-person frame detection result. Then, server 200 determines whether the target object meets the account opening operation requirements based on the single-person frame detection result and the multi-person frame detection result. When the proportion of the multi-person frame time to the total duration of the audio and video information does not exceed a preset range, server 200 determines that target object 110 meets the account opening operation requirements. Otherwise, server 200 determines that target object does not meet the account opening operation requirements. Furthermore, server 200 also writes the single-person frame time proportion and the multi-person frame time proportion to the data storage device 300 (e.g., a database list). Then, if server 200 determines that target object 110 does not meet the account opening operation requirements, it sends an alarm message to the reviewer's terminal device. In addition, the manual review time needs to be extended. Therefore, in this technical solution, server 200 determines whether target object 110 meets the account opening operation requirements based on the single-person frame detection result and the multi-person frame detection result. This method ensures that the account opening process is recorded as a single-person operation in the audio and video information.

[0070] In addition, refer to Figure 1 As shown, according to a second aspect of this embodiment, a storage medium is provided. The storage medium includes a stored program, wherein, when the program is executed, a processor performs any of the methods described above.

[0071] According to this embodiment, the client records audio and video information to document the entire account opening process for the target object. During the recording process, facial motion detection and voice liveness detection are performed on the target object. Based on the results of facial motion detection and voice liveness detection, the client determines whether the target object is alive. Furthermore, the client also performs quality checks on the document image information during recording. Recording ends when the quality check results meet the requirements. The client then sends the recorded audio and video information to the server, which performs comprehensive analysis and determines whether to provide account opening services to the target object based on the analysis results.

[0072] Therefore, this application provides a one-way video-based account opening witnessing method. The client records the entire account opening process of the target (including document display and audio / video input), thus obtaining complete audio / video information. This audio / video information can witness the entire account opening process of the target, ensuring the integrity of the audio / video recording and verifying the authenticity of the target's intention to open an account. Furthermore, during the recording process, multiple detection methods, such as facial motion detection and voice liveness detection, are used to verify the target's liveness, ensuring the authenticity of the user's intention and reducing the risk of fraud. In addition, the client monitors the user's actions at each stage in real time during recording to ensure compliance. When errors are detected, the user is promptly reminded to correct them until a satisfactory audio / video recording is completed. Compared to the method where reviewers discover errors during the witnessing process and require the user to re-witness the video, this application can monitor the user's actions at each stage in real time during the recording process and promptly remind the user to correct errors, effectively reducing the user's operational steps and providing a smoother user experience.

[0073] Furthermore, after recording the entire audio and video information, the client no longer sends only one or a few frames to the server, but instead sends the complete audio and video information. This allows reviewers to determine if there are any anomalies in the account opening process based on this complete audio and video information during the review. The complete audio and video information also serves as valid evidence to prevent users from denying wrongdoing, thus ensuring the legitimate rights and interests of both the securities company and the user. Moreover, after receiving the complete audio and video information from the client, the server performs comprehensive analysis on it. The results of this analysis assist the reviewers in their review process. Compared to directly sending unanalyzed audio and video information to the reviewers, the server-side audio and video analysis method effectively assists the reviewers in their review work, greatly improving the efficiency and accuracy of the review process.

[0074] In summary, this technical solution records the uploading of ID cards and bank cards, along with uninterrupted audio and video recording of the user during the witnessing process, as a complete audio-visual message. This ensures the integrity of the audio-visual recording, improves the reliability of expressing the user's true intentions, and reduces the risk of forgery. After collecting the complete audio-visual information used to record the entire account opening process, the client sends this information to the server to prevent user denial and protect the legitimate rights and interests of the securities company. The client and server combine multiple liveness detection algorithms, working collaboratively to prevent unauthorized operation, ensuring security and controllability, guaranteeing the account opening is based on the user's true intentions, and protecting the legitimate rights and interests of the securities company. Furthermore, the server-side audio-visual analysis method effectively assists reviewers in their verification work, greatly improving the efficiency and accuracy of the review process. This solves the technical problems of existing technologies that, when using video witnessing for account opening, cannot determine the authenticity of the user's intentions, have low reliability, inconvenient operation procedures, and impose a heavy workload on reviewers.

[0075] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0076] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0077] Example 2

[0078] Figure 6 A witness account opening device 600 based on one-way video is shown according to this embodiment, which corresponds to the method described according to the first aspect of Embodiment 1. Reference Figure 6As shown, the device 600 includes: a start request response module 610, used by the client to respond to a start request for recording audio and video information sent by the target object, and to start recording audio and video information, wherein the audio and video information is used to record the entire account opening process of the target object; a face action detection module 620, used by the client to acquire the face image information of the target object during the recording of audio and video information, perform face action detection on the face image information, and determine whether the target object is a live object based on the obtained face action detection result; a voice liveness detection module 630, used by the client to acquire the voice information of the target object during the recording of audio and video information, perform voice liveness detection on the voice information, and determine whether the target object is a live object based on the obtained voice liveness detection result; a document image information acquisition module 640, used by the client to acquire the document image information corresponding to the document displayed by the target object during the recording of audio and video information, perform quality detection on the document image information, and end the recording of audio and video information if the obtained quality detection result meets the quality requirements; and a sending module 650, used by the client to send the recorded audio and video information to the server.

[0079] Optionally, the device 600 further includes: a speech recognition detection module, used by the client to perform speech recognition detection on the speech information using a speech recognition algorithm during the recording of audio and video information; and a first judgment module, used to judge whether the target object reads out the prompt content for account opening as required based on the obtained speech recognition result.

[0080] Optionally, the device 600 further includes: a voice emotion analysis module, used by the client to perform voice emotion analysis on the voice information based on a voice emotion analysis algorithm during the recording of audio and video information; and a second judgment module, used to determine whether the target object voluntarily opens an account based on the obtained voice emotion analysis results.

[0081] Optionally, the device 600 further includes: a comprehensive analysis module, used by the server to perform comprehensive analysis on the audio and video information received from the client to obtain comprehensive analysis results; and a third judgment module, used by the server to determine whether to provide account opening services to the target object based on the comprehensive analysis results.

[0082] Optionally, the comprehensive analysis module includes: a first determination submodule, used by the server to perform face action detection on the audio and video information, and determine the face action time period and the best video frame of the target object by the client; and a first writing submodule, used by the server to write the face action time period and the best video frame of the target object into a preset data storage device.

[0083] Optionally, the comprehensive analysis module further includes: a second determining submodule, used by the server to perform speech recognition on the audio and video information, determine the speech reading time period of the prompt content for account opening read by the target object and the speech recognition result corresponding to the read speech information; and a second writing submodule, used by the server to write the speech reading time period and the speech recognition result into a preset data storage device.

[0084] Optionally, the comprehensive analysis module further includes: a third determination submodule, used by the server to perform ID photo detection on the audio and video information, determine the ID display time period for the target object and the best video frame of the ID; and a third writing submodule, used by the server to write the ID display time period and the best video frame of the ID into a preset data storage device.

[0085] Optionally, the comprehensive analysis module also includes: an identification submodule, used by the server to perform OCR recognition on the document in the best video frame of the document to obtain the user information of the target object; and a face recognition submodule, used by the server to perform face recognition on the face image on the document in the best video frame of the document, and to determine whether the document is the document of the target object based on the face recognition result.

[0086] Optionally, the comprehensive analysis module also includes: a single-person frame detection submodule, used by the server to perform single-person frame detection on the audio and video information and obtain single-person frame detection results; a multi-person frame detection submodule, used by the server to perform multi-person frame detection on the audio and video information and obtain multi-person frame detection results; a judgment submodule, used by the server to determine whether the target object meets the account opening operation requirements based on the single-person frame detection results and the multi-person frame detection results; and an information sending submodule, used by the server to send alarm information to the terminal device of the reviewer if the target object does not meet the account opening operation requirements.

[0087] Therefore, according to this embodiment, the uploading of ID cards and bank cards during the witnessing process, along with the uninterrupted recording of the user's audio and video, is captured as a complete audio-visual message. This records the entire account opening process, ensuring the integrity of the audio-visual information recording, improving the reliability of expressing the user's true intentions, and reducing the risk of forgery. After the client collects the complete audio-visual information used to record the entire account opening process, it sends this information to the server to prevent user denial and protect the legitimate rights and interests of the securities company. The client and server combine multiple liveness detection algorithms, working collaboratively to prevent unauthorized operation, ensuring security and controllability, guaranteeing the account opening is for the user's true intentions, and protecting the legitimate rights and interests of the securities company. Furthermore, the server-side audio-visual analysis method effectively assists reviewers in their verification work, greatly improving the efficiency and accuracy of the verification process. This solves the technical problems existing in the prior art where video witnessing for account opening cannot determine the authenticity of the user's intentions, resulting in low reliability, inconvenient operation procedures, and a heavy workload for reviewers.

[0088] Example 3

[0089] Figure 7 A witness account opening device 700 based on one-way video is shown according to this embodiment, which corresponds to the method described according to the first aspect of Embodiment 1. Reference Figure 7 As shown, the device 700 includes: a processor 710; and a memory 720 connected to the processor 710, used to provide the processor 710 with instructions to process the following steps: In response to a start request for recording audio and video information sent by a target object, the client begins recording audio and video information, wherein the audio and video information is used to record the entire account opening process of the target object; during the recording of audio and video information, the client acquires the facial image information of the target object, performs facial motion detection on the facial image information, and determines whether the target object is a live object based on the obtained facial motion detection result; during the recording of audio and video information, the client acquires the voice information of the target object, performs voice liveness detection on the voice information, and determines whether the target object is a live object based on the obtained voice liveness detection result; during the recording of audio and video information, the client acquires the document image information corresponding to the document displayed by the target object, performs quality detection on the document image information, and ends the recording of audio and video information if the obtained quality detection result meets the quality requirements; and the client sends the recorded audio and video information to the server.

[0090] Optionally, the memory 720 is also used to provide the processor 710 with instructions for processing the following steps: during the recording of audio and video information, the client uses a speech recognition algorithm to perform speech recognition detection on the speech information; and based on the obtained speech recognition results, it determines whether the target object reads out the prompts for account opening as required.

[0091] Optionally, the memory 720 is also used to provide the processor 710 with instructions to process the following steps: during the recording of audio and video information, the client performs voice emotion analysis on the voice information based on a voice emotion analysis algorithm; and based on the obtained voice emotion analysis results, determines whether the target object voluntarily opens an account.

[0092] Optionally, the memory 720 is also used to provide the processor 710 with instructions to process the following steps: the server performs comprehensive analysis on the audio and video information received from the client to obtain a comprehensive analysis result; and the server determines whether to provide account opening service to the target object based on the comprehensive analysis result.

[0093] Optionally, the server performs a comprehensive analysis of the audio and video information received from the client to obtain the comprehensive analysis results. This includes: the server performs face motion detection on the audio and video information to determine the time period of the face motion detection of the target object by the client and the best video frame of the target object's face; and the server writes the face motion time period and the best video frame of the face into a preset data storage device.

[0094] Optionally, the operation of the server performing comprehensive analysis on the audio and video information received from the client to obtain the comprehensive analysis results also includes: the server performing speech recognition on the audio and video information to determine the speech reading time period of the prompt content for account opening read by the target object and the speech recognition result corresponding to the read speech information; and the server writing the speech reading time period and the speech recognition result into a preset data storage device.

[0095] Optionally, the operation of the server performing comprehensive analysis on the audio and video information received from the client to obtain the comprehensive analysis results also includes: the server performing ID photo detection on the audio and video information to determine the ID display time period for the target object and the best video frame of the ID; and the server writing the ID display time period and the best video frame of the ID to a preset data storage device.

[0096] Optionally, the memory 720 is also used to provide the processor 710 with instructions to process the following steps: the server performs OCR recognition on the document in the best video frame of the document to obtain the user information of the target object; and the server performs face recognition on the face image on the document in the best video frame of the document, and determines whether the document is the document of the target object based on the face recognition result.

[0097] Optionally, the memory 720 is also used to provide the processor 710 with instructions to process the following steps: the server performs single-person frame detection on the audio and video information and obtains the single-person frame detection result; the server performs multi-person frame detection on the audio and video information and obtains the multi-person frame detection result; the server determines whether the target object meets the account opening operation requirements based on the single-person frame detection result and the multi-person frame detection result; and the server sends an alarm message to the terminal device of the reviewer if it determines that the target object does not meet the account opening operation requirements.

[0098] Therefore, according to this embodiment, the uploading of ID cards and bank cards during the witnessing process, along with the uninterrupted recording of the user's audio and video, is captured as a complete audio-visual message. This records the entire account opening process, ensuring the integrity of the audio-visual information recording, improving the reliability of expressing the user's true intentions, and reducing the risk of forgery. After the client collects the complete audio-visual information used to record the entire account opening process, it sends this information to the server to prevent user denial and protect the legitimate rights and interests of the securities company. The client and server combine multiple liveness detection algorithms, working collaboratively to prevent unauthorized operation, ensuring security and controllability, guaranteeing the account opening is for the user's true intentions, and protecting the legitimate rights and interests of the securities company. Furthermore, the server-side audio-visual analysis method effectively assists reviewers in their verification work, greatly improving the efficiency and accuracy of the verification process. This solves the technical problems existing in the prior art where video witnessing for account opening cannot determine the authenticity of the user's intentions, resulting in low reliability, inconvenient operation procedures, and a heavy workload for reviewers.

[0099] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0100] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0101] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0102] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0103] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0104] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0105] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A witness account opening method based on one-way video, applied to a witness account opening system, the witness account opening system comprising a client and a server, characterized in that, include: The client responds to the recording audio and video information start request sent by the target object and begins recording the audio and video information, wherein the audio and video information is used to record the entire account opening process of the target object; During the recording of the audio and video information, the client obtains the facial image information of the target object, performs facial motion detection on the facial image information, and determines whether the target object is a live object based on the obtained facial motion detection results; During the recording of the audio and video information, the client obtains the voice information of the target object, performs voice liveness detection on the voice information, and determines whether the target object is alive based on the obtained voice liveness detection result; During the recording of the audio and video information, the client obtains the document image information corresponding to the document displayed by the target object, performs quality detection on the document image information, and ends the recording of the audio and video information if the obtained quality detection result meets the quality requirements. as well as The client sends the recorded audio and video information to the server; it also includes: During the recording of the audio and video information, the client utilizes a speech recognition algorithm to perform speech recognition and detection on the speech information; and Based on the obtained speech recognition results, it is determined whether the target object reads out the prompts for account opening as required; it also includes: During the recording of the audio and video information, the client performs voice emotion analysis on the voice information based on a voice emotion analysis algorithm; and Based on the obtained voice sentiment analysis results, it is determined whether the target subject voluntarily opened an account; it also includes: The server performs comprehensive analysis on the audio and video information received from the client to obtain a comprehensive analysis result; and The server determines whether to provide account opening services to the target object based on the comprehensive analysis results; the operation of the server performing comprehensive analysis on the audio and video information received from the client to obtain the comprehensive analysis results includes: The server performs facial motion detection on the audio and video information to determine the time period of facial motion detection performed by the client on the target object and the best video frame of the target object's face. The server writes the facial action time period and the best video frame of the face into a preset data storage device.

2. The method according to claim 1, characterized in that, The operation of the server performing comprehensive analysis on the audio and video information received from the client to obtain the comprehensive analysis result also includes: The server performs speech recognition on the audio and video information to determine the time period during which the target object reads the prompt content for account opening and the speech recognition result corresponding to the read speech information; and The server writes the speech reading time period and the speech recognition result into a preset data storage device.

3. The method according to claim 2, characterized in that, The operation of the server performing comprehensive analysis on the audio and video information received from the client to obtain the comprehensive analysis result also includes: The server performs ID photo detection on the audio and video information to determine the time period during which the target object displays the ID and the optimal video frame of the ID; and The server writes the document display time period and the best video frame of the document into a preset data storage device.

4. The method according to claim 3, characterized in that, Also includes: The server performs OCR recognition on the document in the best video frame of the document to obtain the user information of the target object; The server performs facial recognition on the face image on the document in the best video frame of the document, and determines whether the document is the document of the target object based on the result of the facial recognition.

5. The method according to claim 4, characterized in that, Also includes: The server performs single-person frame detection on the audio and video information to obtain the single-person frame detection result. The server performs multi-person co-frame detection on the audio and video information to obtain multi-person co-frame detection results; The server determines whether the target object meets the account opening operation requirements based on the single-person frame detection result and the multiple-person frame detection result. as well as If the server determines that the target object does not meet the requirements for account opening, it sends an alarm message to the terminal device of the reviewer.

6. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, the method described in any one of claims 1 to 5 is performed by a processor.