Online signing method and system for informed consent of clinical trials
Through the informed consent process of online clinical trials, multiple risk verification steps and biometric detection are adopted to solve the problem of identity forgery, and informed consent signing with high security and low complexity is achieved.
Patent Information
- Application Number
- CN202510465586.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-15
AI Technical Summary
During the informed consent process of online clinical trials, there are problems of identity forgery, which affects the standardization and scientificity of the trial.
Multiple progressive risk verification steps are adopted, including video clip acquisition, three-dimensional facial network hypothesis generation, key point and texture feature extraction, and geometric consistency detection, combined with biometric verification, to ensure the authenticity of subject identity.
It significantly improves the safety and accuracy of the signing process, reduces operational complexity, reduces manual intervention, and improves the standardization of signing.
Smart Images

Figure CN119993351B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, system, and device for online signing of informed consent for clinical trials. Background Art
[0002] Informed consent is an essential step in clinical trials. Ensuring that subjects fully understand the trial's objectives, methods, potential risks, and other relevant information before signing the informed consent form is a fundamental requirement for safeguarding their rights and interests. To improve the standardization of the informed consent process and save manpower, online systems are being used for subject informed consent. Subjects can interact with the online system as needed using smartphone apps and cameras.
[0003] However, the use of online systems also brings potential risks. Since subjects interact online through cameras, identity forgery may occur. For example, some people may forge their identities through images or videos. This behavior greatly undermines the standardization of clinical trials and has a huge impact on the scientific nature of the trial results.
[0004] Therefore, there is an urgent need to propose a new online signing method for informed consent in clinical trials. Summary of the Invention
[0005] The present application provides an online signing method, system and device for informed consent for clinical trials, which solves the technical problem of identity forgery using online systems in related technologies and achieves the technical effect of more accurate subject identity authentication in online systems.
[0006] In order to achieve the above objectives, the main technical solutions adopted in this application include:
[0007] In a first aspect, embodiments of the present application provide a method for online signing of informed consent for clinical trials. The online signing process is divided into multiple, sequentially progressive risk verification steps, the method comprising:
[0008] After completing the interactive process of the current risk verification step, obtain a video clip of the subject signing the clinical trial informed consent form online;
[0009] Inputting the video frames in the video clip into a three-dimensional mesh generation network to generate a hypothesis to obtain a three-dimensional facial network hypothesis of the subject, and projecting the three-dimensional facial network hypothesis onto a preset two-dimensional plane to obtain projected key points and projected texture features of the subject's face;
[0010] Performing key point extraction and texture feature extraction on the video frame to obtain biometric key points and local microtexture features of the subject's face;
[0011] Performing geometric consistency detection based on the projected key points, the projected texture features, the biometric key points, and the local microtexture features to obtain key point alignment errors and texture consistency between the three-dimensional facial network hypothesis and the video frame;
[0012] If both the key point alignment error and the texture consistency meet the living body detection conditions, proceed to the next risk verification step of the current risk verification step.
[0013] Optionally, the plurality of risk verification steps include informed information display and preliminary verification steps, and the method further comprises:
[0014] In the informed information display and preliminary verification step, if the subject passes the face verification, a preliminary verification interactive interface is displayed; wherein the preliminary verification interactive interface contains basic information of the clinical trial;
[0015] If it is detected that the subject has completed the review of the basic information, it is confirmed that the informed information display and preliminary verification steps are completed.
[0016] Optionally, the plurality of risk verification steps include key clause explanation and dynamic behavior verification steps, and the method further includes:
[0017] In the key clause explanation and dynamic behavior verification step, voice data of the key clause explanation is played one by one, and a first action reminder message is issued; wherein, the first action reminder message is used to require the subject to confirm understanding of the key clause through a first specified action behavior.
[0018] Optionally, the plurality of risk verification steps include question-and-answer confirmation and in-depth verification steps, and the method further comprises:
[0019] In the question-answer confirmation and in-depth verification step, obtaining the subject's question voice data regarding the clinical trial;
[0020] In response to the question voice data, reply voice data for the question voice data is generated, and a second action reminder message is issued; wherein the second action reminder message requires the subject to perform a second specified action behavior to ensure dynamic participation of the living body.
[0021] Optionally, the plurality of risk verification steps include a signing step, and the method further comprises:
[0022] In the signing step, in response to the electronic signature entry operation, tamper-proof signature record data is generated;
[0023] The process data of each of the plurality of risk verification steps, the biometrics of the subject and the signed record data are bound and stored in an encrypted manner.
[0024] Optionally, the method further includes:
[0025] If either the key point alignment error or the texture consistency condition does not meet the living body detection condition, an abnormality alarm is triggered and a re-verification reminder message is sent.
[0026] Optionally, the performing geometric consistency detection based on the projected key points, the projected texture features, the biometric key points, and the local microtexture features to obtain key point alignment errors and texture consistency between the three-dimensional facial network hypothesis and the video frame includes:
[0027] Matching the projected key points with the biometric key points to obtain an average Euclidean distance, and determining the key point alignment error based on the average Euclidean distance;
[0028] The projected texture features and the local micro-texture features are compared using a structural similarity index and cosine similarity to obtain the texture consistency.
[0029] Optionally, inputting the video frames in the video clip into a three-dimensional mesh generation network to generate a hypothesis to obtain a three-dimensional facial network hypothesis of the subject includes:
[0030] Inputting video frames in the video clip into a three-dimensional mesh generation network to generate hypotheses, thereby obtaining a plurality of facial network hypotheses;
[0031] Each facial network hypothesis is scored, and the facial network hypothesis with the highest score is used as the three-dimensional facial network hypothesis.
[0032] In a second aspect, an embodiment of the present application provides an online signing system for informed consent for clinical trials, the system comprising:
[0033] A video clip acquisition module is used to obtain a video clip of the subject signing the clinical trial informed consent form online when the interactive process of the current risk verification step is completed;
[0034] a hypothesis generation and projection module, configured to input the video frames in the video clip into a three-dimensional mesh generation network to generate a hypothesis, thereby obtaining a three-dimensional facial network hypothesis of the subject, and projecting the three-dimensional facial network hypothesis onto a preset two-dimensional plane to obtain projected key points and projected texture features of the subject's face;
[0035] A video frame feature extraction module is used to extract key points and texture features from the video frame to obtain biometric key points and local microtexture features of the subject's face;
[0036] a geometric consistency detection module that performs geometric consistency detection based on the projected key points, the projected texture features, the biometric key points, and the local microtexture features to obtain key point alignment errors and texture consistency between the three-dimensional facial network hypothesis and the video frame;
[0037] The liveness detection result confirmation module is configured to proceed to the next risk verification step of the current risk verification step if both the key point alignment error and the texture consistency satisfy the liveness detection conditions.
[0038] In a third aspect, an embodiment of the present application provides a computer device, including:
[0039] A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method described in any of the above embodiments by executing the computer instructions.
[0040] In an embodiment of the present application, first, after completing the interactive process of the current risk verification step, a video clip of a subject signing an informed consent form online for a clinical trial is obtained. Second, the video frames in the video clip are input into a three-dimensional mesh generation network for hypothesis generation, resulting in a three-dimensional facial network hypothesis for the subject. The three-dimensional facial network hypothesis is then projected onto a preset two-dimensional plane to obtain projected key points and projected texture features of the subject's face. Next, key point extraction and texture feature extraction are performed on the video frame to obtain biometric key points and local microtexture features of the subject's face. Next, geometric consistency detection is performed based on the projected key points, projected texture features, biometric key points, and local microtexture features to obtain the key point alignment error and texture consistency between the three-dimensional facial network hypothesis and the video frame. Finally, if both the key point alignment error and texture consistency meet the liveness detection conditions, the next risk verification step of the current risk verification step is continued. By using the three-dimensional facial network hypothesis generation and projection technology, combined with the extraction and comparison of biometric key points and local microtexture features, it is possible to effectively prevent identity forgery by non-liveness attacks and significantly improve the accuracy of identity authentication during the signing process. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific implementation methods or the description of the prior art. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 A schematic diagram of the clinical trial informed consent system provided in the embodiments of this specification;
[0043] Figure 2 A flowchart of the online signing method for informed consent for clinical trials provided in the embodiments of this specification;
[0044] Figure 3 A flowchart of the online signing method for informed consent for clinical trials provided in the embodiments of this specification;
[0045] Figure 4 A flowchart of the online signing method for informed consent for clinical trials provided in the embodiments of this specification;
[0046] Figure 5 A flowchart of the online signing method for informed consent for clinical trials provided in the embodiments of this specification;
[0047] Figure 6 A flowchart of the online signing method for informed consent for clinical trials provided in the embodiments of this specification;
[0048] Figure 7a A flowchart of the online signing method for informed consent for clinical trials provided in the embodiments of this specification;
[0049] Figure 7b A schematic diagram of a noise estimation model provided in an embodiment of this specification;
[0050] Figure 8 Schematic diagram of the online signing device for informed consent for clinical trials provided in the embodiments of this specification;
[0051] Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0052] To make the purpose, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0053] In clinical trials, ensuring that subjects fully understand the purpose, methods, and risks of the trial is a fundamental requirement for safeguarding their rights and interests. To improve the standardization of the informed consent process and conserve human resources, adopting an online system for signing informed consent for clinical trials is a direction worth exploring. Subjects can complete the informed consent process by interacting with the system through smartphone applications and cameras. However, online systems also have technical issues. Interacting with subjects through cameras can lead to identity forgery, such as using images or videos to impersonate others. This not only affects the standardization of the trial, but may also affect the scientific nature of the results.
[0054] Based on this, the present application proposes an online signing method for informed consent of clinical trials. The online signing process is divided into multiple sequentially progressive risk verification steps. The method includes: first, after completing the interactive process of the current risk verification step, obtaining a video clip of the subject during the online signing process of the informed consent form of the clinical trial; second, inputting the video frame in the video clip into a three-dimensional mesh generation network for hypothesis generation to obtain a three-dimensional facial network hypothesis of the subject, and projecting the three-dimensional facial network hypothesis into a preset two-dimensional plane to obtain the projected key points and projected texture features of the subject's face; then, performing key point extraction and texture feature extraction on the video frame to obtain the biometric key points and local micro-texture features of the subject's face; then, performing geometric consistency detection based on the projected key points, projected texture features, biometric key points and local micro-texture features to obtain the key point alignment error and texture consistency between the three-dimensional facial network hypothesis and the video frame; finally, if the key point alignment error and texture consistency both meet the liveness detection conditions, proceed to the next risk verification step of the current risk verification step. This method reduces operational complexity by dividing the process into multiple, progressive risk verification steps. Through automated processes and intelligent processing, it reduces manual intervention and improves the standardization of informed consent signing for clinical trials. By combining 3D facial network hypothesis generation and projection technology with the extraction and comparison of key biometric points and local microtexture features, it effectively prevents identity forgery by non-live individuals and significantly enhances the security of the signing process. This provides a low-complexity, highly standardized, and highly secure method for online informed consent signing for clinical trials.
[0055] This application scenario example proposes an online signing method for informed consent for clinical trials. Figure 1, this method is applied to the platform management terminal 120 of the clinical trial informed consent system, and the system also includes a first user terminal 110 and a second user terminal 130 that are communicatively connected to the platform management terminal 120. Among them, the platform management terminal 120 is responsible for the overall management of the system, which may include registration of subjects in the system after recruitment, guiding subjects to perform informed consent and signing processes after passing preliminary verification. The first user terminal 110 is provided for subjects to use, and its functions include initiating subject registration applications, obtaining recruitment results, and subsequent clinical trial records. The second user terminal 130 is used by real doctors, and its functions include assessing whether the subjects are fully informed about the clinical trial and confirming their consent to participate in the clinical trial, as well as guiding them to complete the electronic signature entry.
[0056] After a subject initiates a registration application through the first user terminal 110, the platform management terminal 120 receives the request and immediately guides the subject through the informed consent process. First, the subject is prompted to capture a video of the process and obtain their consent. The subject's face is then photographed to confirm that the lighting is appropriate—not too bright or too dark to affect the video capture. If not, the subject is prompted via voice prompts to move to a more lit area and confirm the information. After receiving the confirmation, the platform management terminal 120 captures the subject's face again using the first user terminal 110's camera. Once the lighting check is passed, the informed consent information display and preliminary verification steps can proceed. For example, if the subject passes facial verification, a preliminary verification interface is displayed, which includes basic clinical trial information. The subject can scroll through the screen to review the lengthy basic information. Once completed, they can click the "Read" button to submit the information. After the platform management terminal 120 detects the subject's submission of the information, it analyzes and verifies the video clip for that step. If the verification passes, the subject can proceed to the next step.
[0057] The next step may be the explanation of key clauses and dynamic behavior verification. The first user terminal 110 can play a line-by-line explanation of the key clauses via voice. After the playback is complete, the subject is asked whether they understand and a first action reminder message (such as blinking or opening their mouth) is issued, indicating that if the subject understands, they need to perform the first action. After the platform management terminal 120 detects that the subject has performed the first action, it analyzes and tests the video clip of this step. If the test passes, the next step can be entered. If the subject is not detected to perform the action within a certain period of time, such as 10 seconds, the verification fails and the subject can choose to retry or exit.
[0058] The next step may be a question-and-answer confirmation and in-depth verification step. The subject can send questions to the platform management terminal 120 by voice. After the platform management terminal 120 obtains the voice data of the question, it can generate the corresponding answer through intelligent analysis and play it through voice. After the playback is completed, the subject is asked whether he has understood it, and a second action reminder message (such as blinking, opening the mouth) is issued. That is, if the subject has understood it, he needs to make the second action. After the platform management terminal 120 detects that the subject has made the second action, it analyzes and detects the video clip of this step. After the test passes, it can enter the last step. If the subject is not detected to make the action within a certain period of time, such as 10 seconds, the verification fails and the subject can choose to retry or exit.
[0059] The last step may be a signing step, where the platform management terminal 120 initiates a call request to a real doctor. After the doctor logs in to the second client 130, he communicates with the subject via video through the camera. First, the doctor asks the subject if he has any questions about the clinical trial. If he has any questions, the doctor will answer them one by one until the subject confirms that he has no questions. Then, the doctor asks the subject if he agrees to participate in the clinical trial, and the subject needs to confirm it clearly. Finally, the doctor guides the subject to complete the electronic signature entry. The subject can enter the electronic signature on the screen of the mobile phone terminal 110, and the first user terminal 110 can upload the signature to the platform management terminal 120. After the platform management terminal 120 detects that the subject has submitted the electronic signature, it analyzes and detects the video clip of this step. After the detection is passed, the process data of each of the multiple risk verification steps, the subject's biometrics and the signature record data are bound and encrypted for storage.
[0060] This scenario example reduces operational complexity by dividing the process into multiple, progressive risk verification steps. It also reduces manual intervention and improves the standardization of informed consent signing for clinical trials through automated processes and intelligent processing. It also uses dynamic behavioral verification technology to effectively prevent identity forgery by non-living individuals, significantly improving the security of the signing process.
[0061] According to an embodiment of the present application, an embodiment of a method for online signing of informed consent for clinical trials is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0062] See also Figure 2 In this embodiment, a method for online signing of informed consent for clinical trials is provided. The online signing process is divided into multiple progressive risk verification steps. The method includes:
[0063] S110. When the interactive process of the current risk verification step is completed, a video clip of the subject signing the clinical trial informed consent form online is obtained.
[0064] Risk verification steps can be multiple stages during the online signing process to ensure that the subject fully understands and agrees to the clinical trial information. Video clips can be video recordings of any risk verification step during the signing process, typically used to verify the subject's identity and confirm their active participation in the process.
[0065] In some embodiments, the online signing method for clinical trial informed consent is applied to the platform management end of the clinical trial informed consent system, which also includes a first user end in communication with the platform management end. During the entire clinical trial informed consent process, the platform management end can capture video of the subject through the camera of the first user end. The first risk verification step is the informed information display and preliminary verification step. After the subject passes the face verification and agrees to the basic information of the clinical trial informed, the platform management end obtains the video clip of the subject at this step; the second risk verification step is the key clause explanation and dynamic behavior verification step. After the subject confirms that he is aware of the key clauses, the platform management end obtains the video clip of the subject at this step; the third risk verification step is the question and answer confirmation and in-depth verification step. After the subject confirms that he is aware of the questions asked, the platform management end obtains the video clip of the subject at this step; the fourth risk verification step is the signing step. After the subject submits the electronic signature, the platform management end obtains the video clip of the subject at this step.
[0066] S120: Input the video frames in the video clip into a three-dimensional mesh generation network for hypothesis generation to obtain a three-dimensional facial network hypothesis of the subject, and project the three-dimensional facial network hypothesis onto a preset two-dimensional plane to obtain projected key points and projected texture features of the subject's face.
[0067] The 3D mesh generation network is used to generate a 3D facial network hypothesis of the subject from video frames in a video clip and can be a deep learning model. The 3D facial network hypothesis can be the 3D geometric structure data of the subject's face. Projected key points and projected texture features refer to the key points and texture information obtained by projecting the 3D facial network hypothesis onto a preset 2D plane. For example, projected key points can include the coordinates of key points of facial features such as the eyes, nose, and mouth, while projected texture features can include information such as facial skin texture.
[0068] In some implementations, 3D face modeling techniques, such as the 3D Face Alignment Network (3D-FAN), can be used to generate hypotheses about the subject's face. 3D-FAN first extracts the subject's facial key points from video frames; then, using a deep neural network, it maps the facial information in the 2D image into 3D space, generating a 3D network hypothesis of the subject's face, including shape, texture, and so on.
[0069] In some embodiments, after obtaining the three-dimensional facial network hypothesis of the subject, the three-dimensional facial network hypothesis of the subject is projected onto a two-dimensional plane using perspective projection technology to obtain a two-dimensional facial projection. In order to extract facial key points from the two-dimensional facial projection, Dlib can be used. It is an open source machine learning library that is widely used in facial recognition and processing tasks. Dlib provides an efficient facial key point detection algorithm that can accurately identify and locate the 68 feature points of the face. Based on this, Dlib can be used to extract the coordinates of the facial key points from the two-dimensional facial projection image and use them as the projected key points of the subject's face. Finally, a convolutional neural network can be used to extract texture features from the two-dimensional facial projection to obtain projected texture features.
[0070] S130 , performing key point extraction and texture feature extraction on the video frame to obtain biometric key points and local micro-texture features of the subject's face.
[0071] Key point extraction involves identifying and extracting the coordinates of specific facial feature points of a subject, such as eyes, nose, and mouth, from video frames. Texture feature extraction involves analyzing facial details to obtain local microtexture features, such as skin texture and pore distribution.
[0072] In some embodiments, Dlib can be used to extract key points of the subject's face from the video frame, such as 68 points of facial features. After extraction, the coordinates of each key point are obtained as the biometric key point.
[0073] In some embodiments, a convolutional neural network may be used to extract texture features from the subject's face in a video frame to obtain local microtexture features of the subject.
[0074] S140, performing geometric consistency detection based on the projected key points, projected texture features, biometric key points, and local micro-texture features to obtain key point alignment errors and texture consistency between the three-dimensional facial network hypothesis and the video frame.
[0075] Geometric consistency testing can be performed by comparing the differences between projected keypoints and biometric keypoints, as well as comparing the differences between projected texture features and local microtexture features, to determine the degree of match between the 3D facial network hypothesis and the subject's features in the video frame. Keypoint alignment error can be the error between projected keypoints and biometric keypoints. Texture consistency can be the difference between projected texture features and local microtexture features.
[0076] In some implementations, keypoint alignment is performed using a landmark-based alignment method. For example, the eye is used as a reference to geometrically align projected keypoints with biometric keypoints. After alignment, the keypoint alignment error is calculated. Specifically, the Euclidean distance between each pair of keypoints can be calculated as an error metric. Alternatively, the mean squared error (MSE) can be calculated.
[0077] In some embodiments, the texture consistency between the 3D facial network hypothesis and the video frame is obtained by comparing the similarity between the projected texture features and the local micro-texture features. Specifically, the Euclidean distance or cosine similarity between the projected texture features and the local micro-texture features can be calculated.
[0078] It's important to note that real faces have a natural three-dimensional structure, and their projections from different viewpoints adhere to strict perspective geometry. In contrast, forgeries (such as image or video replays) contain only planar information or a fixed two-dimensional perspective and therefore fail to present a three-dimensional structure that conforms to projection rules. Specifically, the 3D facial network of a real face assumes that the face exhibits natural facial contours and depth variations when projected, while a 2D forged image may appear flattened or distorted. Based on this principle, when the 3D keypoints of a real face are projected onto a 2D image, the error with the keypoints of the facial image in the video frame is minimal. In contrast, due to the lack of a reasonable 3D structure, the projection error of a forged attack is significantly increased. Furthermore, the texture of real skin typically closely matches the image in the video frame, while forged attacks often produce noticeable visual differences due to material reflectivity, insufficient texture resolution, or synthesis issues. Therefore, keypoint alignment error and texture consistency can be used to determine whether a subject in a video frame is forged.
[0079] S150: If both the key point alignment error and the texture consistency meet the living body detection conditions, proceed to the next risk verification step of the current risk verification step.
[0080] Liveness detection conditions can be pre-set thresholds or standards used to determine whether the subject's facial key point alignment error and texture consistency meet liveness detection requirements. These conditions can be used to assess the authenticity of facial features, ensuring that the subject participating in the signing process is a living person with biometric characteristics, rather than a fabricated image or video.
[0081] In some implementations, if the keypoint alignment error is less than a keypoint alignment error threshold and the texture consistency is greater than a texture consistency threshold, it is determined that both the keypoint alignment error and texture consistency meet the liveness detection criteria, and the next risk verification step can be performed. For example, if both the keypoint alignment error and texture consistency meet the liveness detection criteria during the informed information display and preliminary verification steps, the key clause explanation and dynamic behavior verification steps can be performed.
[0082] In some embodiments, a video frame is extracted for detection at a certain time step (e.g., 5 seconds), and the key point alignment errors and texture consistency obtained multiple times are averaged and compared with the liveness detection conditions to determine whether the subject in the video frame is a forged attack.
[0083] The above embodiment reduces operational complexity by dividing the risk verification process into multiple, progressive steps. Automated processes and intelligent processing reduce manual intervention, improving the standardization of informed consent signing for clinical trials. Three-dimensional facial network hypothesis generation and projection technology, combined with the extraction and comparison of key biometric points and local microtexture features, effectively prevents identity forgery by non-live attacks and significantly improves the accuracy of identity verification during the signing process. This embodiment provides a low-complexity, highly standardized, and highly accurate method for online signing of informed consent for clinical trials.
[0084] See also Figure 3 In some embodiments, the multiple risk verification steps include an informed information display and a preliminary verification step, and the method further includes:
[0085] S210. In the informed information display and preliminary verification step, if the subject passes the face verification, a preliminary verification interactive interface is displayed.
[0086] S220: If it is detected that the subject has completed the review of the basic information, confirm that the informed information display and preliminary verification steps are completed.
[0087] The informed information display and preliminary verification step is the first of multiple risk verification steps. It is used to complete preliminary verification and present basic clinical trial information to subjects. The preliminary verification interactive interface contains basic clinical trial information. Basic information can be the core content of the clinical trial, such as the study purpose and background, drug information, subject eligibility, and the study process (including adverse reactions, potential benefits, damages, etc.). Subjects must review this information in full to ensure they fully understand the upcoming trial. Facial verification can verify the subject's identity through facial recognition technology.
[0088] In some embodiments, face verification can first obtain an image of the subject through a camera; then perform feature extraction on the face to obtain the subject's facial features; then obtain the face image corresponding to the subject's identity information from the ID card information database through an interface, perform feature extraction to obtain the subject's identity features; finally, compare the subject's facial features and identity features for similarity to determine whether the identities match.
[0089] In some embodiments, when the basic information of the clinical trial is displayed in the preliminary verification interactive interface, the subject can scroll through the long basic information by sliding the screen. After completing the review, they can click the "Read" button to submit the information. After the platform management detects the subject's submission of information, it confirms that the subject has completed the informed information display and preliminary verification steps and obtains a video clip of this step. Then, according to steps S120 to S150, a liveness test is performed on the subject in the video clip. If the test passes, the next step is continued.
[0090] In the above embodiment, facial verification is used to initially confirm the accuracy of the subject's identity. After the subject has reviewed their basic information, geometric consistency testing is used to further confirm the accuracy of their identity during the informed information display and initial verification steps. This dual verification mechanism improves the reliability and accuracy of identity confirmation.
[0091] In some embodiments, multiple risk verification steps include key clause explanation and dynamic behavior verification steps. The method also includes: in the key clause explanation and dynamic behavior verification steps, playing voice data of the key clause explanation one by one, and issuing a first action reminder message.
[0092] Key terms may be content that requires special attention in clinical trials, such as subject eligibility, privacy protection, and exit mechanisms. The first action reminder message is used to require the subject to confirm understanding of the key terms through a first designated action, such as requiring the subject to nod or blink to confirm understanding of the key terms.
[0093] In some embodiments, the mobile phone plays the voice data of the explanation of the key terms one by one. After the playback is completed, the subject is asked to blink through voice to confirm that he has understood the content of the key terms. The mobile phone first obtains the eye key points through Dlib, and then calculates the aspect ratio (Eye Aspect Ratio, EAR) of the subject's eyes when open and closed, and monitors its dynamic changes to determine whether it is a live operation, effectively preventing non-live attacks. After the test is passed, the confirmation information is submitted to the platform management end. The platform management end obtains the video clip of this step, and then extracts a video frame for detection at a certain time step (for example, 5s). The key point alignment error and texture consistency obtained multiple times are averaged, and then compared with the live detection conditions to determine whether the subject in the video frame is a forged attack. If the test passes, proceed to the next step.
[0094] In the above embodiment, in the key clause explanation and dynamic behavior verification steps, a preliminary non-living attack prevention is first performed through the first specified action, and then the prevention is further strengthened through geometric consistency detection, and the double detection is used to ensure the accuracy of the subject's identity.
[0095] See also Figure 4 In some embodiments, the multiple risk verification steps include question-answer confirmation and in-depth verification steps, and the method further includes:
[0096] S410: In the question-answer confirmation and in-depth verification steps, voice data of questions raised by the subjects regarding the clinical trial are obtained.
[0097] S420: In response to the question voice data, generate reply voice data for the question voice data, and send a second action reminder message.
[0098] The question voice data can be a recording of any question asked by the subject during the Q&A confirmation and in-depth verification steps. The response voice data refers to the voice response generated by the platform management terminal based on the subject's question. The second action reminder message requires the subject to perform a second specified action to ensure dynamic participation.
[0099] In some embodiments, the mobile phone prompts the subject to voice-enter their questions about the clinical trial, and the voice information is collected in real time through the phone's microphone. After collection is complete, the question voice data is obtained and submitted to the platform management end. The platform management end first pre-processes the question voice data using natural language processing (NLP) technology, including speech recognition (ASR) to convert speech into text, and performs word segmentation, noise removal, and semantic parsing on the text to accurately extract the core content of the subject's question voice data. Next, the platform management end generates a targeted reply text based on the extracted question content based on the trained clinical trial question-answering model. Finally, the platform management end uses speech synthesis (TTS) technology to convert the reply text into natural and fluent reply voice data, and plays it back to the subject through the mobile phone end to clearly and accurately answer the question and ensure the subject's understanding of the clinical trial content.
[0100] It's important to note that clinical trial question-answering models can be fine-tuned based on a pre-trained Transformer model. For example, using GPT, the model can be fine-tuned using clinical trial-related question-answer pairs as training data, including information on trial objectives, processes, risks, and privacy protections. Through fine-tuning, the model can learn the specific semantics and contextual relationships within the clinical trial domain, generating more accurate and professional responses.
[0101] In some embodiments, after the mobile phone obtains the subject's confirmation that there is no doubt, it asks the subject to open his mouth through voice. The mobile phone first obtains the key points of the mouth through the Dlib tool, then calculates the aspect ratio (Mouth Aspect Ratio, MAR) of the subject's mouth when it is open and closed, and monitors its dynamic changes to determine whether it is a live operation, effectively preventing non-live attacks. After the test is passed, the confirmation information is submitted to the platform management end. The platform management end obtains the video clip of this step, and then extracts a video frame for detection at a certain time step (for example, 5 seconds). The key point alignment error and texture consistency obtained multiple times are averaged, and then compared with the live detection conditions to determine whether the subject in the video frame is a forged attack. If the test passes, proceed to the next step.
[0102] In the above embodiment, in the question-answer confirmation and depth verification steps, preliminary non-living attack prevention is first performed through the second specified action, and then the prevention is further strengthened through geometric consistency detection, and the double detection is used to ensure the accuracy of the subject's identity.
[0103] See also Figure 5 In some embodiments, the plurality of risk verification steps include a signing step, the method further comprising:
[0104] S510. In the signing step, in response to the electronic signature entry operation, tamper-proof signature record data is generated.
[0105] S520: Bind and encrypt the process data of each of the multiple risk verification steps, the subject's biometrics, and the signature record data for storage.
[0106] Electronic signature entry can involve the subject electronically signing an informed consent form on their mobile phone. Process data can include interactive data generated during each risk verification step, such as video clips, voice data, and action recordings. Encrypted storage can involve storing data in a secure database using an encryption algorithm to prevent data leakage or tampering. The subject's biometric features can include key biometric points and local microtexture features of the subject's face.
[0107] In some implementations, before the electronic signature is entered, the platform administrator can automatically call the live physician responsible for informed consent for the clinical trial. Once online via their mobile device, the live physician can communicate with the subject via video link. First, the physician will ask the subject if they have any questions about the clinical trial. If so, the physician will answer them one by one until the subject confirms they have no further questions. Then, the physician will inquire about the subject's consent to participate in the clinical trial and obtain explicit confirmation from the subject. Finally, the physician will instruct the subject to complete the electronic signature entry.
[0108] In some embodiments, the subject can enter an electronic signature on the mobile phone screen. After the first user terminal detects that the entry is complete, it can upload the electronic signature to the platform management terminal. The platform management terminal can use blockchain technology to generate tamper-proof signature record data. First, the electronic signature and related signature information (such as the identity of the signatory, the signing timestamp, the signing content, etc.) are hashed to generate a unique hash value. The hash value is then recorded in the block of the blockchain. Each block forms a chain structure by including the hash value of the previous block, ensuring that the data cannot be tampered with once it is on the chain. Any modification to the signature record will cause the hash value to change, which will be detected by the blockchain network. At the same time, the distributed ledger characteristics of the blockchain enable the signature record to be stored on multiple nodes, further enhancing the security and tamper-proofness of the data, and ultimately achieving the tamper-proof and high credibility of the electronic signature record.
[0109] In some implementations, after the platform management detects a subject's electronic signature, it captures a video clip of that step. Then, at regular time intervals (e.g., 5 seconds), it extracts a video frame for testing. The keypoint alignment error and texture consistency obtained multiple times are averaged and compared with liveness detection criteria to determine whether the subject in the video frame is a forgery attack. If the test passes, the process data from each of the multiple risk verification steps, the subject's biometrics, and the signature record data are combined and encrypted using an encryption algorithm (e.g., Advanced Encryption Standard (AES)). The data is then stored in a highly secure database, ensuring traceability throughout the informed consent process for clinical trials.
[0110] In the above embodiment, tamper-proof signed record data is generated through electronic signature entry, and liveness detection technology is used to authenticate the subject's identity, effectively preventing the risk of identity forgery through non-liveness attacks. Furthermore, the signed record data is securely stored along with process data from the risk verification step and the subject's biometrics, enabling traceability throughout the entire clinical trial informed consent process and further ensuring the security of online informed consent signing for clinical trials.
[0111] In some embodiments, the method further includes: if any one of the key point alignment error and the texture consistency condition does not meet the living body detection condition, triggering an abnormality alarm and sending a re-verification reminder message.
[0112] Among them, abnormal alarms are used to remind platform managers to conduct further intervention and processing, effectively preventing the risk of non-living attacks.
[0113] In some implementations, if the keypoint alignment error is less than a keypoint alignment error threshold, or if the texture consistency is less than a texture consistency threshold, an anomaly alert is sent to the platform administrator. The administrator can then further analyze the subject's video clip and test results to determine whether the anomaly is caused by a non-live attack or a platform defect. If the anomaly is due to a platform defect, the administrator can implement improvements and optimizations accordingly, thereby continuously improving the platform's accuracy and reliability and effectively mitigating the risk of non-live attacks.
[0114] In some embodiments, the online clinical trial informed consent signing system allows subjects to attempt verification up to three times at each step, i.e., two re-verification attempts. If the subject fails the previous two verification attempts, the system will send a re-verification reminder message, and the subject can proceed with the current step again according to the reminder.
[0115] The above embodiment sets thresholds for key point alignment error and texture consistency, triggers an abnormal alarm when any condition is not met, and allows the subject to re-verify, which can effectively prevent the risk of non-living attacks and improve the accuracy and reliability of identity authentication.
[0116] See also Figure 6 In some embodiments, geometric consistency detection is performed based on projected key points, projected texture features, biometric key points, and local micro-texture features to obtain key point alignment errors and texture consistency between the 3D facial network hypothesis and the video frame, including:
[0117] S710 , matching the projected key points with the biometric key points to obtain an average Euclidean distance, and determining a key point alignment error based on the average Euclidean distance.
[0118] S720 , using the structural similarity index and cosine similarity to compare the projected texture features with the local micro-texture features to obtain texture consistency.
[0119] The average Euclidean distance can be the average geometric distance between corresponding points of the projected keypoints and the biometric keypoints. The structural similarity index can be used to simulate the subjective perception of image quality by the human visual system. It combines the comparison of brightness, contrast, and structural information. Values closer to 1 indicate higher similarity, while values closer to 0 indicate lower similarity.
[0120] In some implementations, the projected key points and the biometric key points are first obtained using the Dlib tool, and then the geometric distances between the corresponding points are calculated and averaged to obtain the average Euclidean distance, which is used as the key point alignment error.
[0121] In some embodiments, statistics representing brightness, contrast, and structural information can be obtained from image features, and then a structural similarity index between two images can be calculated. For example, by projecting texture features and local microtexture features, the brightness similarity, contrast similarity, and structural similarity can be obtained, respectively, using the following formulas:
[0122]
[0123]
[0124]
[0125] The structural similarity index is the product of brightness similarity, contrast similarity, and structural similarity:
[0126]
[0127] Among them, μx 、μ y are the mean values of the two image pixels (reflecting brightness), σ x 2 , σ y 2 are the variance of the pixel values of the two images (reflecting the contrast), σ xy is the covariance of the pixel values of the two images (reflecting structural similarity), C1, C2, and C3 are constants used to avoid the denominator being zero.
[0128] In some embodiments, the cosine similarity between the projected texture feature and the local microtexture feature is calculated using the projected texture feature vector and the local microtexture feature vector, and the formula is:
[0129]
[0130] Among them, A is the projected texture feature vector, B is the local micro-texture feature vector, and C is the cosine similarity.
[0131] In some embodiments, the structural similarity index and the cosine similarity are comprehensively considered. For example, when both the structural similarity index and the cosine similarity exceed their respective thresholds, the texture consistency is determined to be consistent; otherwise, it is determined to be inconsistent.
[0132] In the above embodiment, the projected keypoints are matched with the biometric keypoints to obtain the average Euclidean distance, and the keypoint alignment error is determined based on the average Euclidean distance. Furthermore, the projected texture features and local microtexture features are compared using the structural similarity index and cosine similarity to determine texture consistency. This provides strong data support for liveness detection.
[0133] See also Figure 7a In some embodiments, inputting video frames from a video clip into a 3D mesh generation network to generate a hypothesis to obtain a 3D facial network hypothesis of the subject includes:
[0134] S810: Inputting video frames in the video clip into a three-dimensional mesh generation network to generate hypotheses, thereby obtaining a plurality of facial network hypotheses.
[0135] S820: Score each facial network hypothesis, and use the facial network hypothesis with the highest score as the three-dimensional facial network hypothesis.
[0136] In some embodiments, the 3D mesh generation network includes a first face network hypothesis generation module and a second face network hypothesis generation module. The first face network hypothesis generation module may be a 3D Face Alignment Network (3D-FAN), which is configured to generate an initial 3D face network hypothesis based on an image frame. The second face network hypothesis generation module is configured to generate multiple face network hypotheses using a reverse diffusion process based on a denoising diffusion probabilistic model (DDPM).
[0137] The forward diffusion process of DDPM can be to gradually add noise to the initial three-dimensional facial network hypothesis x0, and after T time steps, it gradually becomes a standard Gaussian distributed noise x T , the forward diffusion process can be expressed as:
[0138]
[0139] in, .
[0140] By gradually adding noise, It can be expressed as:
[0141]
[0142] in, , .
[0143] The reverse diffusion process of DDPM can be obtained from Gaussian distribution Stepwise denoising and generating facial network hypotheses At each time step T i ,from Can get After multiple iterations, the noise is gradually reduced.
[0144]
[0145] in, is the noise variance, To estimate the noise, ~N(0,1).
[0146] To get , the noise at each step needs to be known , the noise can be obtained through a noise estimation model, which can output a noise estimate at each step and provide guidance for the reverse sampling process. Figure 7bThe model consists of an encoding module 811, a hypothesis generation module 812, and a decoding module 806. The encoding module 811 includes an encoder 801 and an encoder 802. The encoder 801 can be a three-dimensional convolutional neural network (3D CNN). Encoding is performed to obtain a noisy facial feature vector. Encoder 802, which can be a two-dimensional convolutional neural network (2D CNN), extracts features from the image frame to obtain local facial features (such as the eyes, nose, forehead, chin, etc.) and global facial features. These features can serve as diffusion conditions to guide the diffusion process. Assume that generation module 812 includes a multi-head self-attention unit 803, a cross-attention unit 804, and a feedforward network 805. The local facial features and the noisy facial feature vector are input into the multi-head self-attention unit 803. This helps the model better understand the relationship between local regions, thereby more accurately recognizing faces. In the cross-attention unit 804, the global facial features serve as key and value features, and the output of the multi-head self-attention unit 803 serves as the query feature. This allows for a better fusion of global and local information, enabling the diffusion process to adjust to the content of the video frame and more accurately estimate noise. Decoding module 806, which can be a multi-layer perceptron, decodes the output noise estimate.
[0147] It should be noted that a training dataset can be pre-constructed. The data samples in the training dataset use a noisy 3D facial network and corresponding 2D face images. The initial noise estimation model is trained using the training dataset until the training stop condition is met. The noise estimation model is obtained. The loss function can use the mean squared error (MSE) loss function, which continuously optimizes the model parameters by comparing the noise estimated by the model with the actual noise added.
[0148] In some embodiments, a sample is taken from the standard Gaussian distributed noise obtained in the forward diffusion process, and the sample is used as the initial value. The reverse diffusion process is iterated, wherein the noise estimation model is used to obtain the noise estimation of each step. ,After multiple steps of iteration, the facial network hypothesis is finally obtained.
[0149] In some embodiments, multiple sampling is performed from the standard Gaussian distributed noise obtained in the forward diffusion process, and the initial value obtained each time is different. It also randomly samples different values, and finally obtains multiple different facial network hypotheses.
[0150] It's important to note that in situations where poor lighting results in unclear video frames, especially when non-liveness attacks intentionally exploit low-light conditions, generating facial network hypotheses can effectively increase the probability of capturing correct facial features. Furthermore, when the subject's face is partially occluded, generating facial network hypotheses can improve the system's robustness to occluded areas through diversified predictions. While unobstructed portions can be directly captured, generating hypotheses can help the system better infer features from occluded areas, thereby improving the integrity and accuracy of overall facial features. Especially in situations where non-liveness attacks intentionally implement partial occlusion, generating facial network hypotheses can effectively reduce the attack's success rate and enhance system security. Therefore, the 3D facial network hypotheses generated by the second facial network hypothesis generation module are more accurate than the initial 3D facial network hypotheses generated by the first facial network hypothesis generation module. After the initial 3D facial network hypothesis is generated by the first facial network hypothesis generation module, multiple 3D facial network hypotheses are generated by the second facial network hypothesis generation module.
[0151] In some embodiments, multiple facial network hypotheses are scored for macroscopic features, and the facial network hypothesis with the highest score is selected as the 3D facial network hypothesis. Specifically, the macroscopic feature scoring may include a shape consistency score and a facial symmetry score. For example, a standard facial template (e.g., an average facial shape) is predefined, which contains typical facial contours and facial feature proportions. For each 3D facial model, the geometric difference between it and the standard facial template is calculated. Key points (e.g., the locations of the eyes, nose, and mouth) of the 3D facial model and the standard facial template are first extracted, and then the Euclidean distance between these key points is calculated. Based on this distance, shape consistency is assessed; smaller distance values indicate higher shape consistency scores. For example, bilaterally symmetrical key points of the 3D facial model (e.g., left and right eyes, left and right eyebrows, and left and right corners of the mouth) are first extracted, and then the Euclidean distance between these symmetrical key points is calculated. Symmetry is assessed based on the distance difference between the key points; smaller differences indicate higher facial symmetry scores. For example, the shape consistency score and facial symmetry score of each facial network hypothesis are added together as its macro-feature score, and the facial network hypothesis with the highest score is taken as the three-dimensional facial network hypothesis.
[0152] In the above embodiment, by inputting video frames from a video clip into a 3D mesh generation network, multiple facial network hypotheses are generated, thereby increasing the probability of capturing correct facial features. Each facial network hypothesis is scored, and the highest-scoring one is selected as the 3D facial network hypothesis. This provides a strong data foundation for projecting it onto a pre-defined 2D plane for detailed liveness detection.
[0153] See also Figure 8The embodiment of the present application further provides an online signing system 900 for informed consent for clinical trials. The online signing system 900 for informed consent for clinical trials includes:
[0154] The video clip acquisition module 910 is used to acquire a video clip of the subject signing the clinical trial informed consent form online when the interactive process of the current risk verification step is completed;
[0155] Hypothesis generation and projection module 920 is used to input the video frames in the video clip into the three-dimensional mesh generation network to generate a hypothesis, obtain a three-dimensional facial network hypothesis of the subject, and project the three-dimensional facial network hypothesis onto a preset two-dimensional plane to obtain projected key points and projected texture features of the subject's face;
[0156] The video frame feature extraction module 930 is used to extract key points and texture features from the video frame to obtain biometric key points and local micro-texture features of the subject's face;
[0157] A geometric consistency detection module 940 performs geometric consistency detection based on projected key points, projected texture features, biometric key points, and local microtexture features to obtain key point alignment errors and texture consistency between the 3D facial network hypothesis and the video frame;
[0158] The liveness detection result confirmation module 950 is configured to proceed to the next risk verification step of the current risk verification step if both the key point alignment error and the texture consistency meet the liveness detection conditions.
[0159] In some embodiments, the online signing system 900 for clinical trial informed consent further includes:
[0160] An interactive interface display module is used to display a preliminary verification interactive interface when the subject passes facial verification during the informed information display and preliminary verification steps; wherein the preliminary verification interactive interface contains basic information about the clinical trial;
[0161] The confirmation module is used to confirm the completion of the informed information display and preliminary verification steps if it is detected that the subject has completed the review of the basic information.
[0162] In some embodiments, the online signing system 900 for clinical trial informed consent further includes:
[0163] The voice playback module is used to play the voice data of the key terms one by one during the key terms explanation and dynamic behavior verification steps, and issue a first action reminder message; wherein the first action reminder message is used to require the subject to confirm understanding of the key terms through a first specified action behavior.
[0164] In some embodiments, the online signing system 900 for clinical trial informed consent further includes:
[0165] The subject voice acquisition module is used to obtain the subject's voice data regarding clinical trial questions during the question-answer confirmation and in-depth verification steps;
[0166] The reply voice generation module is used to generate reply voice data for the question voice data in response to the question voice data, and send a second action reminder message; wherein the second action reminder message requires the subject to perform a second specified action behavior to ensure dynamic participation of the living body.
[0167] In some embodiments, the online signing system 900 for clinical trial informed consent further includes:
[0168] A signing record generating module, configured to generate tamper-proof signing record data in response to an electronic signature input operation during the signing step;
[0169] The data binding and encryption module is used to bind and encrypt the process data of multiple risk verification steps, the subject's biometrics and signature record data.
[0170] In some embodiments, the online signing system 900 for clinical trial informed consent further includes:
[0171] The alarm module is used to trigger an abnormal alarm and send a re-verification reminder message if either the key point alignment error or the texture consistency does not meet the liveness detection conditions.
[0172] In some implementations, the geometric consistency detection module 940 further includes:
[0173] an alignment error determination unit, configured to match the projected key points with the biometric key points to obtain an average Euclidean distance, and determine the key point alignment error based on the average Euclidean distance;
[0174] The texture condition determination unit is used to compare the projected texture feature with the local micro-texture feature using a structural similarity index and a cosine similarity to obtain texture consistency.
[0175] In some implementations, it is assumed that the projection generation module 920 further includes:
[0176] a video frame input unit, configured to input video frames in a video clip into a three-dimensional mesh generation network for hypothesis generation, thereby obtaining a plurality of facial network hypotheses;
[0177] The scoring unit is used to score each facial network hypothesis and take the facial network hypothesis with the highest score as the three-dimensional facial network hypothesis.
[0178] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0179] The online signing system for informed consent for clinical trials in this embodiment is presented in the form of functional units, where the units refer to ASIC (Application Specific Integrated Circuit) circuits, processors and memories that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0180] See also Figure 9 , Figure 9 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 9 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of a GUI on an external input / output device (such as, a display device coupled to an interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 9 A processor 10 is taken as an example.
[0181] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0182] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0183] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0184] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0185] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0186] The embodiments of the present application also provide a computer-readable storage medium. The above-mentioned method according to the embodiment of the present application can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0187] An embodiment of the present application provides a computer program product, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a method according to any embodiment of the present application.
[0188] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.
[0189] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0190] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0191] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0192] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0193] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0194] It is understandable that in the specific implementation of this application, related data such as user information, location information, navigation data, etc. are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0195] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0196] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0197] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0198] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0199] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0200] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0201] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0202] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0203] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
[0204] Although the embodiments of the present application have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations shall fall within the scope defined by the appended claims.
Claims
1. A method for online signing of informed consent for clinical trials, characterized in that: The online signing process is divided into multiple progressive risk verification steps, including: After completing the interactive process of the current risk verification step, obtain a video clip of the subject signing the clinical trial informed consent form online; Inputting the video frames in the video clip into a three-dimensional mesh generation network to generate a hypothesis to obtain a three-dimensional facial network hypothesis of the subject, and projecting the three-dimensional facial network hypothesis onto a preset two-dimensional plane to obtain projected key points and projected texture features of the subject's face; Performing key point extraction and texture feature extraction on the video frame to obtain biometric key points and local microtexture features of the subject's face; A geometric consistency test is performed based on the projected key points, the projected texture features, the biometric key points, and the local micro-texture features to obtain a key point alignment error and texture consistency between the three-dimensional facial network hypothesis and the video frame. The geometric consistency test determines the degree of match between the three-dimensional facial network hypothesis and the subject's features in the video frame by comparing the differences between the projected key points and the biometric key points, and comparing the differences between the projected texture features and the local micro-texture features. Specifically, the projected texture features and the local micro-texture features are compared using a structural similarity index and a cosine similarity to obtain the texture consistency. The texture consistency is obtained by comprehensively considering the structural similarity index and the cosine similarity. When both the structural similarity index and the cosine similarity exceed their respective thresholds, the texture consistency is determined to be consistent. If both the key point alignment error and the texture consistency meet the living body detection conditions, proceed to the next risk verification step of the current risk verification step.
2. The method according to claim 1, characterized in that The plurality of risk verification steps include an informed information display and a preliminary verification step, and the method further comprises: In the informed information display and preliminary verification step, if the subject passes the face verification, a preliminary verification interactive interface is displayed; wherein the preliminary verification interactive interface contains basic information of the clinical trial; If it is detected that the subject has completed the review of the basic information, it is confirmed that the informed information display and preliminary verification steps are completed.
3. The method according to claim 1, characterized in that The plurality of risk verification steps include key clause explanation and dynamic behavior verification steps, and the method further includes: In the key clause explanation and dynamic behavior verification step, voice data of the key clause explanation is played one by one, and a first action reminder message is issued; wherein, the first action reminder message is used to require the subject to confirm understanding of the key clause through a first specified action behavior.
4. The method according to claim 1, wherein The plurality of risk verification steps include question-answer confirmation and in-depth verification steps, and the method further comprises: In the question-answer confirmation and in-depth verification step, obtaining the subject's question voice data regarding the clinical trial; In response to the question voice data, reply voice data for the question voice data is generated, and a second action reminder message is issued; wherein the second action reminder message requires the subject to perform a second specified action behavior to ensure dynamic participation of the living body.
5. The method according to claim 1, wherein The plurality of risk verification steps include a signing step, and the method further comprises: In the signing step, in response to the electronic signature entry operation, tamper-proof signature record data is generated; The process data of each of the plurality of risk verification steps, the biometrics of the subject and the signed record data are bound and stored in an encrypted manner.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: If either the key point alignment error or the texture consistency condition does not meet the living body detection condition, an abnormality alarm is triggered and a re-verification reminder message is sent.
7. The method according to any one of claims 1 to 5, characterized in that The performing of geometric consistency detection based on the projected key points, the projected texture features, the biometric key points, and the local micro-texture features to obtain key point alignment errors and texture consistency between the three-dimensional facial network hypothesis and the video frame further includes: The projected key points are matched with the biometric key points to obtain an average Euclidean distance, and the key point alignment error is determined based on the average Euclidean distance.
8. The method according to any one of claims 1 to 5, characterized in that The step of inputting the video frames in the video clip into a three-dimensional mesh generation network to generate a hypothesis to obtain a three-dimensional facial network hypothesis of the subject includes: Inputting video frames in the video clip into a three-dimensional mesh generation network to generate hypotheses, thereby obtaining a plurality of facial network hypotheses; Each facial network hypothesis is scored, and the facial network hypothesis with the highest score is used as the three-dimensional facial network hypothesis.
9. An online signing system for informed consent for clinical trials, characterized by: The system comprises: A video clip acquisition module is used to obtain a video clip of the subject signing the clinical trial informed consent form online when the interactive process of the current risk verification step is completed; a hypothesis generation and projection module, configured to input the video frames in the video clip into a three-dimensional mesh generation network to generate a hypothesis, thereby obtaining a three-dimensional facial network hypothesis of the subject, and projecting the three-dimensional facial network hypothesis onto a preset two-dimensional plane to obtain projected key points and projected texture features of the subject's face; A video frame feature extraction module is used to extract key points and texture features from the video frame to obtain biometric key points and local microtexture features of the subject's face; A geometric consistency detection module performs geometric consistency detection based on the projected key points, the projected texture features, the biometric key points, and the local micro-texture features to obtain key point alignment errors and texture consistency between the three-dimensional facial network hypothesis and the video frame. The geometric consistency detection determines the degree of match between the three-dimensional facial network hypothesis and the subject features in the video frame by comparing the differences between the projected key points and the biometric key points, and comparing the differences between the projected texture features and the local micro-texture features. Specifically, the projected texture features and the local micro-texture features are compared using a structural similarity index and a cosine similarity to obtain texture consistency. The texture consistency is obtained by comprehensively considering the structural similarity index and the cosine similarity. When both the structural similarity index and the cosine similarity exceed their respective thresholds, the texture consistency is determined to be consistent. The liveness detection result confirmation module is configured to proceed to the next risk verification step of the current risk verification step if both the key point alignment error and the texture consistency satisfy the liveness detection conditions.
10. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 8 by executing the computer instructions.
Citation Information
Patent Citations
Drug clinical trial monitoring method and system based on block chain, device and medium
CN109065101A
Clinical test informed agreement signing method, signing system and electronic equipment
CN118964543A
Method and apparatus for facial recognition
US20160070952A1
Perspective distortion characteristic based facial image authentication method and storage and processing device thereof
US20200026941A1