Method and verification method for generating identity confirmation of events of interest from a video stream

CN122290029BActive Publication Date: 2026-08-11SUZHOU DEEPSIGHT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本申请提供了一套全新的、集成了媒体锁定、信任链溯源和独立可验证凭证机制的综合解决方案,以尝试解决现有技术在数据真实性、可审计性和可信赖性方面的多重不足

Benefits of technology

[0015]This application discloses a media-linked identity verification method and authentication method. Its significant technical effect lies in constructing a highly three-dimensional identity verification process with multi-level verification and a non-repudiable evidence chain. This solution overcomes the pain points of existing media analysis systems, such as the susceptibility of identity information to interference, lack of process credibility, and lack of data traceability. It elevates identity recognition from simple target matching to event-based composite evidence generation. Its core technical effect lies in constructing a three-dimensional information anchoring mechanism. First, a dedicated event detection unit performs high-precision temporal analysis on continuous raw video streams, accurately focusing attention on video segments of interest with specific behavioral patterns and precisely determining their time range. Subsequently, the system does not rely solely on target recognition results but uses a first summarization unit to calculate a highly abstract scene fingerprint for the event segment. This fingerprint forcibly associates identity information with specific spatiotemporal background patterns, fundamentally improving the contextual depth of identity recognition. In the identity verification stage, this system not only obtains the original role identifier of the person, but also transforms it into a unique and secure reference role identifier through an irreversible first encryption algorithm, greatly enhancing the data's resistance to cracking and its security during transmission and storage. Regarding data integration, this application designs a second digest unit, whose key value lies in performing highly complex semantic abstraction on three sets of core information: scene fingerprint, encrypted role identifier, and precise start and end times, generating an indispensable "original digest." This digest perfectly integrates identity information with the event's context and timing, forming the first preliminary identity credential that the system can automatically verify. Furthermore, this system introduces a human-machine collaborative identity verification mechanism, allowing expert-level human judgment to serve as an independent, high-weight evidence stream, which is then superimposed and merged with the machine-generated original digest to generate a verified identity verification with the highest level of credibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290029B_ABST
    Figure CN122290029B_ABST
Patent Text Reader

Abstract

This application provides a method for identity verification and authentication of generating events of interest from video streams. The identity verification method includes: extracting and preprocessing frames from the input video stream to remove noise interference and enhance the image quality of keyframes; identifying objects of interest in the video stream using a target detection algorithm and extracting their biometric information; matching the extracted features with a preset identity feature database to determine the object's preliminary identity; verifying the preliminary identity by combining it with scene information of the event of interest, and outputting the identity verification result. The authentication method includes receiving the identity verification result and verifying the accuracy of the identity verification through multimodal feature cross-validation, time series consistency checks, etc., to ensure the reliability of identity recognition in events of interest. The method of this application can effectively improve the reliability of identity recognition for events of interest in video streams and is applicable to various scenarios such as sports videos and intelligent transportation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the intersection of information security technology, computer vision technology, and digital evidence storage technology, and in particular to a method for identity recognition and tamper-proof evidence storage based on a multimodal data source of video streams. Background Technology

[0002] Video recording systems have become essential tools for documenting human activities and behaviors, and their applications have expanded dramatically, encompassing key areas such as sports, medical diagnosis, and business process management. In these fields, video recordings not only serve as a record of facts but also as the basis for critical decisions, such as athlete recruitment, security incident investigations, and medical diagnoses. Therefore, there is a widespread industry need to reliably link specific events and individuals within video recordings; that is, to permanently and verifiably bind identity information to video evidence in a highly credible, auditable, and tamper-proof manner.

[0003] Existing technologies for associating identity information with media evidence suffer from serious systemic and fundamental flaws, resulting in a lack of reliability, auditability, and legal validity. First, most mainstream identity labeling systems currently employ post-hoc metadata recording, lacking a mechanism for encrypted commitment to the original media content before identity claims are generated. This means the relationship between identity records and video content is merely external, lacking an inseparable encrypted causal link. Second, existing systems generally cannot distinguish the source and evolution history of evidence, lacking a mechanism for differentiating trust levels. Once data is modified manually or by the system, the original confidence level and derivation history information are lost, making complete traceability and reliability assessment of evidence impossible. Furthermore, these records are typically stored in traditional databases, lacking tamper-proof mechanisms; any authorized access party can modify historical records without leaving a trace. Most importantly, all existing verification processes heavily rely on the centralized platform database that issues the claim, lacking a mechanism that allows third-party users to independently verify its authenticity using only the original evidence and portable credentials. In summary, existing technologies fail to simultaneously address the four core issues of media evidence locking, identity history tracing, data tamper-proofing, and independent external verification within a unified framework.

[0004] This application provides a novel, comprehensive solution that integrates media locking, trust chain tracing, and independent verifiable credential mechanisms to address the multiple shortcomings of existing technologies in terms of data authenticity, auditability, and trustworthiness. Summary of the Invention

[0005] The first aspect of this application provides a method for identity verification by generating events of interest from a video stream, wherein an event detection unit detects events of interest in the video stream to obtain one or more video segments of the events of interest and the start and end times of the video segments, characterized in that, for any video segment, the following steps are performed: S1: Calculate the fingerprint of the video segment using the first digest algorithm: S2: Based on the target recognition unit, identify the characters in the video clip and obtain the reference character identifier and confidence level; S3: Calculate the original digest based on the fingerprint, the reference role identifier, and the start and end times using the second digest algorithm; S4: Use the fingerprint, the reference role identifier, the confidence level, the real-time timestamp, and the original digest as the original identity verification.

[0006] Preferably, the steps further include: S5: The person or role identified by the target recognition unit is manually verified, and the fingerprint, the reference role identifier, the real-time timestamp, the verifyer information, and the original summary are used as the verification identity confirmation.

[0007] In one embodiment, when there are multiple video segments of the events of interest, the original identity verification of subsequent video segments also includes the original summary of the previous video segment.

[0008] In one embodiment, the fingerprint, the start and end times of the video segment, the reference role identifier, and the original digest are packaged together and a trusted chain signature is added as a verifiable identity verification credential.

[0009] In one embodiment, the event detection unit is an AI computer vision model trained to detect specific motion actions.

[0010] Preferably, the reference character identifier is calculated from the character identification information using an irreversible first encryption algorithm.

[0011] The second aspect of this application provides a verification method for identity confirmation, wherein the data to be verified includes an identity verification credential, a person's role, and an original video stream, the identity verification credential being obtained based on the method described in the first aspect of this application; the verification method includes: Based on the start and end times of the video segments in the identity verification credential, the corresponding video segments in the original video stream are read, the video segments are calculated using a first digest algorithm to obtain a second fingerprint, and the second fingerprint is matched with the fingerprint in the identity verification credential to confirm the consistency of the video segments. The character identification information is calculated using a first encryption algorithm to obtain first encrypted information, and the first encrypted information is matched with the reference character marker in the identity verification certificate to confirm the consistency of the character. The authenticity of the identity verification credential is confirmed through a trusted chain signature.

[0012] A third aspect of this application provides an identity management system for generating events of interest from a video stream, comprising: An event detection unit is adapted to detect events of interest in a video stream and obtain one or more video segments of the events of interest and their start and end times; The first digest unit is adapted to calculate the fingerprint of the video segment using a first digest algorithm; The target recognition unit is adapted to identify the characters and their confidence levels in the video clip; The first encryption unit is adapted to calculate character identification information using an irreversible first encryption algorithm to obtain a reference character identifier. The second digest unit is adapted to calculate the original digest based on the fingerprint, the reference role identifier, and the start and end times using a second digest algorithm. The original identity generation unit is adapted to generate an original identity verification based on the fingerprint, the reference role identifier, the confidence level, the real-time timestamp, and the original digest; The identity verification generation unit, in response to the human verifier's verification operation on the character role, generates a verification identity confirmation based on the fingerprint, the reference character identifier, the verifier information, the original summary, and the real-time timestamp.

[0013] A fourth aspect of this application provides a computing device, characterized in that the computing device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the method of the first or second aspect of this application.

[0014] The fifth aspect of this application provides a computer-readable storage medium storing a computer program that, when run on a processor, executes the methods of the first or second aspect of this application.

[0015] This application discloses a media-linked identity verification method and authentication method. Its significant technical effect lies in constructing a highly three-dimensional identity verification process with multi-level verification and a non-repudiable evidence chain. This solution overcomes the pain points of existing media analysis systems, such as the susceptibility of identity information to interference, lack of process credibility, and lack of data traceability. It elevates identity recognition from simple target matching to event-based composite evidence generation. Its core technical effect lies in constructing a three-dimensional information anchoring mechanism. First, a dedicated event detection unit performs high-precision temporal analysis on continuous raw video streams, accurately focusing attention on video segments of interest with specific behavioral patterns and precisely determining their time range. Subsequently, the system does not rely solely on target recognition results but uses a first summarization unit to calculate a highly abstract scene fingerprint for the event segment. This fingerprint forcibly associates identity information with specific spatiotemporal background patterns, fundamentally improving the contextual depth of identity recognition. In the identity verification stage, this system not only obtains the original role identifier of the person, but also transforms it into a unique and secure reference role identifier through an irreversible first encryption algorithm, greatly enhancing the data's resistance to cracking and its security during transmission and storage. Regarding data integration, this application designs a second digest unit, whose key value lies in performing highly complex semantic abstraction on three sets of core information: scene fingerprint, encrypted role identifier, and precise start and end times, generating an indispensable "original digest." This digest perfectly integrates identity information with the event's context and timing, forming the first preliminary identity credential that the system can automatically verify. Furthermore, this system introduces a human-machine collaborative identity verification mechanism, allowing expert-level human judgment to serve as an independent, high-weight evidence stream, which is then superimposed and merged with the machine-generated original digest to generate a verified identity verification with the highest level of credibility.

[0016] The technical effectiveness of this application lies in constructing a closed-loop lifecycle from "data capture" to "evidence generation." This process not only enhances continuity by adding a summary of the previous event as prior knowledge for subsequent identification, but also packages key elements such as fingerprints, time, and role identifiers, and attaches a trusted chain signature, making the generated identity verification not merely a data result, but a complete, verifiable, auditable, and timestamped legal-grade evidentiary credential. Furthermore, the system provides independent methodological support, allowing for multiple cross-verifications by re-matching the original video stream, fingerprints, and encrypted information, ensuring the authenticity and consistency of data at each stage, thereby fundamentally solving the long-standing problem of insufficient evidence credibility in the field of media tracing. Attached Figure Description

[0017] Figure 1This is a flowchart of a method for generating an identity verification method for an event of interest from a video stream, as described in this application.

[0018] Figure 2 This is a schematic diagram of the structure of the terminal device or server in the embodiments of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0021] The identity verification method provided in this application can be run directly in a smart shooting device, analyzing events of interest in real time during shooting and identifying targets in event video clips to generate identity verification; it can also be run in a computing device, transmitting real-time video streams or complete videos to the computing device to execute this method.

[0022] One embodiment of this application discloses an identity verification method for generating events of interest (ROIs) from a video stream. The video stream can be a real-time video stream in progress or a complete, pre-recorded video stream. ROIs refer to events that meet preset rules and are video segments extracted based on actual business scenario requirements. The specific configuration of ROIs varies in different business scenarios. For example, in a sports competition scenario, ROIs can be configured as athletes completing technical movements of preset difficulty in a designated event, such as gymnastic moves, scoring points in a ball game, or spectacular goals, thereby generating a highlight reel of the designated athlete. In professional skills training, the system can be configured with ROIs to determine the completion rate of key skill steps. It can record whether trainees have completed each precise movement according to the preset optimal operating sequence; every detail is a condition for triggering the event, thus generating a more comprehensive and objective skills assessment report than traditional manual assessment.

[0023] Identity verification is a type of data whose core value lies in using encryption technology and digest algorithms to irreversibly bind the content features, character information, timestamps, and trust chain elements of video clips, forming a traceable and tamper-proof identity association record. This data not only includes the initial results of AI-automated recognition but also integrates authoritative verification information from human review, providing a complete chain of evidence for subsequent independent verification. In practical applications, identity verification data can serve as publicly recognized electronic evidence or be used for internal system access control, event tracing, and other scenarios, effectively solving the problems of insufficient credibility, high risk of tampering, and reliance on centralized platforms for verification inherent in traditional video identity association. By embedding the original digest with the correlation information of preceding and following events, identity verification data can also construct a continuous evolution trajectory of events, achieving cross-segment identity consistency tracking, further enhancing the integrity and application value of the data.

[0024] During execution, the method first requires an event detection unit to detect events of interest in the video stream, obtaining one or more video segments of the events of interest and their start and end times. Specifically, the event detection unit can employ a deep learning-based temporal event detection model or an AI computer vision model trained to detect specific motion actions. The model consists of a feature extraction subunit and an event classification subunit: the feature extraction subunit extracts spatial and temporal features from the input video frame sequence, extracts single-frame spatial features through a convolutional neural network, and captures inter-frame temporal dependencies using a long short-term memory network, outputting a multi-dimensional event feature vector; the event classification subunit calculates the cosine similarity between this feature vector and a predefined feature library of events of interest. When the similarity is greater than a preset confidence threshold (e.g., 85%), an event of interest is detected, and the start and end timestamps of the corresponding video segment are recorded simultaneously. For consecutive events of the same type of interest, if the time interval between adjacent events is less than a preset merging threshold (e.g., 2 seconds), the event detection unit will merge them into a single complete event to avoid duplicate recording. If the interval exceeds the merging threshold, the events will be split into multiple independent events, ensuring clear boundaries for each event. During the training phase, the event detection unit's model can employ supervised learning. The training dataset contains a large number of video samples labeled with the types of events of interest, start and end times, and key features. By optimizing the cross-entropy loss function to minimize detection errors, the model's generalization ability in complex scenarios is improved. For example, in a sports competition scenario, when an athlete completes a jump of preset difficulty, the event detection unit will automatically extract a video segment from the start of the action (e.g., both feet leaving the ground) to the end of the action (e.g., landing stably) and record the corresponding time interval, providing basic material for subsequent identity verification. The start and end times of the video segments are typically stored as timestamps.

[0025] like Figure 1As shown, for any video segment of interest detected by the event detection unit, the following steps are performed: S1: Calculate the fingerprint of the video segment using a first digest algorithm. Specifically, the first digest algorithm employs a frame sequence-based hash fusion strategy. In one possible implementation, for each frame in the video segment, a pre-trained lightweight convolutional neural network is used to extract the spatial feature vector of the frame, with the vector dimension fixed at 512. Then, the feature vectors of all frames are fused using a temporal weighting method, with the weights dynamically allocated based on the importance of the frame in the event—for example, the weight of the event start frame and key action frames is set to 1.2 times that of ordinary frames. Subsequently, the fused feature vector is input into the SHA-256 hash function to generate a fixed-length 256-bit hash value, which is the fingerprint of the video segment. This fingerprint possesses strong tamper resistance: even if the video segment undergoes minor adjustments such as single-frame pixel tampering or frame order reversal, the recalculated fingerprint will show a significant difference from the original fingerprint, effectively locking in the original state of the video content. For example, in a medical surgery scenario, if a key operation frame of a surgical video segment is tampered with, comparing the original fingerprint with the fingerprint of the modified segment can quickly identify content anomalies, ensuring the authenticity of medical records.

[0026] S2: Based on the target recognition unit, identify the characters in the video segment to obtain a reference character identifier and confidence level. Specifically, the target recognition unit can adopt a combined architecture of the YOLO series target detection model with spatial attention mechanism and the character feature extraction model based on the Transformer architecture. First, character detection is performed on keyframes in the video segment that are uniformly sampled at time intervals, the target detection box of each character is located and the character region is cropped; then, a 1024-dimensional feature vector of the character region is extracted through a pre-trained feature extraction model. This vector contains key identity information such as facial features, body features, and clothing features of the character; subsequently, the extracted feature vector is compared with a pre-built character feature library to calculate the cosine similarity and obtain the character matching confidence level. Preferably, to protect personal privacy, the matched character identification information can be input into a first encryption algorithm. The first encryption algorithm can use an asymmetric encryption algorithm such as SHA256 to generate an irreversible reference character identifier. If the confidence level is lower than a preset threshold, for example, it can be configured to 90%, then the record can be marked as "pending review" to trigger the subsequent manual review process. For example, in a sports competition scenario, when a video clip of a basketball player performing a dunk is detected, the target recognition unit will extract the athlete's facial and body features and compare them with the registered athlete feature database. If the similarity reaches 95%, a reference role marker for the athlete will be generated and a confidence level of 95% will be output. If the similarity is only 80%, the process will proceed to the manual review stage, where coaches and other professionals will confirm the athlete's identity.

[0027] Alternatively, the target recognition unit can be deployed as an end-to-end integrated detection and recognition model based on DETR (Detection Transformer). This model utilizes a Transformer encoder-decoder architecture to directly output the bounding boxes and corresponding feature embedding vectors of all person instances in the image, eliminating the need for complex anchor box design or non-maximum suppression post-processing. The extracted feature vectors are then matched against a character feature library, following the same process. This architecture simplifies the processing pipeline and is more conducive to integration on computationally limited edge devices.

[0028] S3: Based on the fingerprint, the reference role identifier, and the start and end times, calculate the original digest using a second digest algorithm. Specifically, the second digest algorithm can employ a hash generation strategy that involves ordered concatenation of multi-source data. First, the fingerprint, reference role identifier, and the start and end times of the video clip are concatenated into a byte stream in a fixed order of "fingerprint → reference role identifier → start and end time." Then, the concatenated byte stream is input into a hash function to generate the original digest. This original digest irreversibly binds the content features of the video clip, the identity identifier, and the time boundary, ensuring that the correlation between the three cannot be tampered with. For example, in a skills training scenario, the original digest generated by concatenating the fingerprint of the video clip of a trainee completing a key operation, their reference role identifier, and the operation time can serve as a unique and reliable identifier of the trainee's operation behavior. Any subsequent alteration to the video content, identity information, or time will result in a change to the original digest, thus enabling accurate identification during the verification process.

[0029] S4: The fingerprint, the reference role identifier, the confidence score, the real-time timestamp, and the original digest are used as the original identity verification. Specifically, the original identity verification is encapsulated in a structured data format, with each core field organized into a key-value pair structure in a preset fixed order: the "fingerprint" field stores the hash value generated in step S1; the reference role identifier field records the character information obtained by the target detection algorithm in step S2; the "confidence score" field stores the matching confidence score of target recognition in the form of a floating-point number between 0 and 1; the "real-time timestamp" field uses the UTC standard time format, accurate to the millisecond level, marking the specific time when the original identity verification was generated; and the "original digest" field stores the multi-source binding hash value calculated in step S3. After all fields are assembled, a transmittable and storable original identity verification data unit is generated through serialization in accordance with the JSON or Protobuf industry standard. This data unit is also appended with a globally unique identity identifier (ID), which is generated by double hashing the original digest and the real-time timestamp using the SHA-256 algorithm, ensuring uniqueness and traceability in cross-system scenarios. The generated raw identity verification data units can be directly used to trigger subsequent review processes, or stored as basic data in the identity management system's database.

[0030] Considering that the target recognition unit is an automated system, there is inevitably a certain probability of misidentification. Therefore, in a preferred embodiment, S5 is also included: after the original identity verification is generated, manual review of the identity verification is supported to obtain a reviewed identity verification. To ensure the identity verification is identical to the original identity verification and to prevent data tampering, the reviewed identity verification includes a video clip fingerprint and an original summary; the video clip fingerprint ensures that the video clip on which the reviewed identity verification is based is consistent with the original identity verification; the original summary is the target recognition summary in the original identity verification, ensuring that the conclusion reviewed by the reviewed identity verification is the same conclusion as in the original identity verification. In addition, the reviewed identity verification also includes a reference role identifier, reviewer information, and a real-time timestamp at the time of review.

[0031] In one possible implementation, there may be more than one event of interest (HOI) from the same video stream. For multiple HIOIs in the same video stream, the identity management system generates a unique identity verification data unit for each event. Based on the temporal correlation of the events, the original identity verification of the subsequent video segment includes the original summary of the preceding video segment, forming a complete identity verification chain. Specifically, for the Nth HIOI (N≥2) appearing in chronological order in the same video stream, when generating its original identity verification, in addition to including the event's fingerprint, reference role identifier, confidence level, real-time timestamp, and original summary, an additional field for the summary of the preceding event is added. The value of this field is the original summary of the (N-1)th HIOI. Through this ordered embedding method, each subsequent event forms an irreversible association with the preceding event, constructing a continuous identity verification chain. This chain not only clearly presents the temporal evolution of events but also allows for rapid identification of anomalies by verifying the consistency of the summaries before and after any tampering occurs at any stage. For example, in security scenarios, when a series of abnormal events captured by the same camera, such as a suspicious person repeatedly appearing in a restricted area, the identity verification chain binds the identity information of each event to the previous event. If the original digest of a certain event is tampered with, the digests of the preceding events in subsequent events will not match the original digest of the tampered event, thus triggering a tampering alarm in the system and ensuring the integrity and credibility of the entire event sequence. Furthermore, the identity verification chain also supports cross-event identity consistency verification. By comparing the reference role markers and feature associations of all events in the chain, the behavioral trajectory of the same target in different events can be effectively tracked, further enhancing the traceability and application value of the identity management system. When it is necessary to perform overall verification of the entire video stream's event sequence, the system can start from the beginning of the chain and sequentially verify the matching relationship between the original digest of each event and the digest of the preceding event. If all steps pass verification, it proves that the entire event sequence has not been tampered with and that the identity association logic is coherent. This chain structure not only strengthens the anti-tampering characteristics of the data but also provides structured evidence support for event analysis in complex scenarios.

[0032] Based on the above steps, this method can obtain complete identity verification, identity verification review, and identity verification chain, thus realizing the generation of identity verification. Preferably, this identity verification can also be provided to external personnel for verification. The externally provided identity verification credential includes the fingerprint, the start and end times of the video clip, the reference role identifier, and the original digest packaged together, with the platform's trusted chain signature added. Specifically, the platform uses a private key of an asymmetric encryption algorithm to digitally sign the packaged core data (fingerprint, video clip start and end times, reference role identifier, and original digest), generating a fixed-length signature value; simultaneously, the platform's trusted certificate number is encapsulated together with the core data and signature value into an identity verification credential conforming to the X.509 standard or the industry-standard JWT (JSON Web Token) format. This credential supports cross-platform transmission and parsing. External verifiers only need to obtain the platform's public key to complete the authenticity verification of the credential, including consistency verification of the video clip, consistency verification of the role, and authenticity verification of the credential.

[0033] Based on the same inventive concept, this application also discloses a verification method for identity confirmation. The method involves reading corresponding video segments from the original video stream based on the start and end times of the video segments in the identity confirmation credential; calculating a second fingerprint from the video segments using a first digest algorithm; matching the second fingerprint with the fingerprint in the identity confirmation credential to confirm the consistency of the video segments; calculating the person / role identification information using a first encryption algorithm to obtain first encrypted information; matching the first encrypted information with the reference role marker in the identity confirmation credential to confirm the consistency of the person / role; and confirming the authenticity of the identity confirmation credential through a trusted chain signature. Specifically, an external verifier decrypts the signature value in the credential using the platform's public key. If the decrypted result matches the hash value of the core data, the signature is considered authentic and valid. Simultaneously, the verifier can query the validity of the platform's trusted certificate number through an authoritative Certificate Authority (CA) to confirm that the certificate is not expired, not revoked, and correctly owned. For example, in financial risk control scenarios, when an external auditing firm verifies a customer's identity verification credentials for remote account opening, it first extracts the start and end times from the credentials to obtain the account opening operation segment from the original video stream, calculates the second fingerprint and matches it with the credential fingerprint, then extracts the customer's facial features through target recognition to generate encrypted information that matches the reference role identifier, and finally verifies the signature using the platform's public key. This confirms that the video content of the account opening operation is authentic, the identity association is correct, and the credentials have not been tampered with. After completing the above three verifications, the external verification party can determine that the events and identity information recorded on the identity verification credentials are authentic and reliable, and can be used as valid evidence for compliance review or dispute resolution.

[0034] The following example, using a football match recording, illustrates the specific implementation of this solution. In the high-intensity broadcasting and video analysis of football matches, a media-bound identity tracing mechanism with encrypted sequence constraints at its core has been developed. Its core technological highlight lies in constructing a time-series-based, multi-source fusion identity verification system, completely resolving the shortcomings of existing evidence records that only remain at the "association declaration" stage and cannot prove the integrity of the evidence subject.

[0035] When processing each game, the system does not identify individuals before recording, but strictly follows a causal sequence of first solidifying the evidence subject, then embedding the identity subject, and finally forming a commitment summary. When the system's event detection unit identifies a meaningful event such as "goal" or "foul," the first unit is triggered. Before any identification occurs, this unit immediately extracts the original video segment corresponding to the event at the byte level and performs a high-intensity one-way summary operation to generate a unique and unchanging video segment fingerprint. This fingerprint serves as the digital fingerprint anchor point for the original video content, and its generation itself locks in the physical evidence basis of the segment, ensuring that all subsequent operations are based on an unalterable video transcript.

[0036] Subsequently, after the system obtains this solidified video fingerprint, the target recognition unit begins to operate. It utilizes multimodal algorithms, such as number OCR recognition, body shape tracking, and posture analysis, to identify participants in the segment and, based on the first encryption unit, encrypts the participants' personal information, transforming their identities into stable reference role identifiers. At this point, the system possesses two independent evidentiary components with strictly controlled causal order: the video fingerprint and the reference role identifier.

[0037] The final solidification of the entire evidence chain is accomplished by the original identity generation unit. The key calculation performed by this generator is based on the second digest unit, which concatenates the two components—the video fingerprint and the person's identity marker—with a timestamp accurate to milliseconds, and then performs a final hash encapsulation using the SHA-256 algorithm to generate the final identity verification. Identity verification is not a simple concatenation of video and identity, but a single, highly complex cryptographic object. Any minor modification to the original video data will alter the video fingerprint, thus rendering the original digest completely invalid; conversely, even if the video remains unchanged, if the identity reference marker is altered, the original digest will also undergo an irreversible change. This causal dependency makes identity verification itself a self-verifying and self-constraining form of evidence.

[0038] To address the complexity of the on-site environment, the system employs a sophisticated evidence trust grading system. Initially, the results automatically generated by the target identification unit are marked as Trust Level 1 original identity verification. If the seriousness of the evidence or the adversarial requirements increase, the system can initiate a human-machine collaborative review mechanism. At this point, after manual review by an authorized administrator (such as a coach or manager), the initial Trust Level 1 record is not modified. Instead, a new Trust Level 2 record is generated based on the original video fingerprint and identity reference markers. The new Trust Level 2 record explicitly points to the upstream Trust Level 1 record in the data flow, forming a permanently auditable overlay record chain. This ensures that the entire evidence generation and verification process forms an unbreakable and tamper-proof complete audit trail. Ultimately, all key evidence generated throughout the process is batch-signed using trusted chain technology and encapsulated into independent, cross-platform credentials. This allows external third-party verifiers to independently verify the integrity of the evidence, the accuracy of the identity, and the temporal logic of the entire evidence chain using only these credentials, without relying on the platform's database, by recalculating the video hash value.

[0039] Based on the same inventive concept, one embodiment of this application discloses an identity management system for generating events of interest from a video stream, including: The event detection unit is suitable for detecting events of interest in a video stream, obtaining one or more video segments of the events of interest and their start and end times. The event detection unit can be implemented as an embedded intelligent analysis module based on an edge computing architecture, its core consisting of a preprocessing subunit, an event feature extraction subunit, and a decision subunit. The preprocessing subunit is responsible for real-time frame decoding and dimensionality reduction of the input video stream. For example, it uses an H.265 hardware decoder to quickly parse video frames and adjusts the frame resolution to a fixed size through bilinear interpolation, while applying Gaussian filtering to the frame sequence to reduce noise interference. The event feature extraction subunit uses an event detection model based on a large artificial intelligence model or an embedded spatiotemporal convolutional network. This model integrates the temporal dynamic features captured by 3D convolutional layers with the spatial semantic features extracted by 2D convolutional layers, effectively identifying the feature patterns of preset events of interest. The decision subunit, based on a pre-trained event classifier, classifies and judges the extracted spatiotemporal features. When the confidence of an event exceeds a preset threshold, it triggers a video segment extraction operation and records its start and end timestamps. Furthermore, the event detection unit supports dynamically updating the feature template library for events of interest, and continuously optimizes detection accuracy through incremental learning by receiving externally labeled data. For example, in a smart factory scenario, the event detection unit can monitor worker violations in real time. When such an event is detected, it automatically extracts the corresponding video clip and marks the time interval, providing accurate material for subsequent identity verification. This module can be deployed on edge cameras or local servers, featuring low latency and high throughput, meeting the needs of real-time video stream analysis.

[0040] A first digest unit is adapted to calculate the fingerprint of the video segment using a first digest algorithm, and a second digest unit is adapted to calculate the original digest based on the fingerprint, the reference role identifier, and the start and end times using a second digest algorithm. The first and second digest units can be implemented as a lightweight hash calculation module integrated into an edge computing node or cloud server. The core of this module includes a data preprocessing subunit, an optimized hash execution subunit, and a result caching subunit: the data preprocessing subunit is responsible for converting the input video segment byte stream (used by the first digest unit) or multi-source structured data (fingerprint, reference role identifier, start and end times, used by the second digest unit) into a unified binary format and performing data integrity verification; the optimized hash execution subunit uses hardware-accelerated SHA256 or SM3 algorithms, improving computational efficiency through pipelined parallel processing to ensure digest generation is completed within milliseconds; the result caching subunit caches recently generated hash results, supporting fast querying and reuse, and reducing redundant computation overhead. This module also supports dynamic algorithm switching, allowing users to select different hash algorithms based on application scenario requirements. It also provides standard API interfaces for data interaction with other units in the identity management system, ensuring efficient collaboration throughout the entire process. For example, in real-time scenarios, this module can quickly process fingerprint generation from continuous video clips and calculate raw digests from multi-source data, providing reliable foundational data support for the subsequent construction of the identity verification chain.

[0041] The target recognition unit, suitable for identifying characters and their confidence levels in the video clip, can be implemented as an intelligent recognition module integrating multimodal features. Its core consists of a character detection subunit, a feature extraction subunit, an identity matching subunit, and a confidence evaluation subunit. The character detection subunit employs a lightweight YOLO model or a DETR (DetectionTransformer) end-to-end architecture to quickly locate character instances in the video clip and output accurate bounding boxes. The feature extraction subunit integrates a Transformer-based pre-trained model to extract facial features, key pose features, and clothing features, generating high-dimensional feature embedding vectors. The identity matching subunit calculates the cosine similarity between the extracted feature vectors and templates in the character feature library to obtain preliminary matching results. The confidence evaluation subunit combines the matching similarity, the output confidence of the detection model, and the template quality score from the feature library to generate a final identity matching confidence level in the 0-1 range. This module supports dynamic updates to the feature library, incrementally learning from correctly identified identity data verified by manual review to continuously optimize matching accuracy. Simultaneously, it employs model quantization and pruning techniques to compress the model size to less than 30% of its original size, allowing deployment on edge cameras or embedded devices for low-latency real-time recognition. In one possible embodiment, within a smart campus scenario, the target recognition unit can quickly identify teachers and students in classroom interaction video clips. When the confidence level is above 90%, identity verification is automatically completed; when it is below 80%, the data is pushed to the administrator for manual review, effectively improving the efficiency and accuracy of campus behavior analysis.

[0042] The first encryption unit is adapted to calculate character identification information using an irreversible first encryption algorithm to obtain a reference character identifier. The original identity generation unit is suitable for generating original identity verification based on the fingerprint, the reference role identifier, the confidence level, the real-time timestamp, and the original digest. The first encryption unit can be implemented as a lightweight encryption computing module, integrated into the backend of the target recognition unit or deployed independently on an edge computing node. Its core consists of a feature preprocessing subunit, an encryption algorithm execution subunit, and a result standardization subunit: the feature preprocessing subunit receives the high-dimensional character feature vector output by the target recognition unit, performs dimensionality compression and feature normalization through principal component analysis, removes redundant information, and maps feature values ​​to a fixed range; the encryption algorithm execution subunit uses the SM3 algorithm conforming to national cryptographic standards or the internationally accepted SHA-256 irreversible hash algorithm to perform encryption calculations on the preprocessed feature vector, ensuring that the original feature information cannot be reverse-engineered; the result standardization subunit converts the encrypted hash value into a hexadecimal string format, performs length verification, and finally outputs a reference role identifier of fixed length. This module supports dynamic configuration of algorithm parameters, allowing adjustment of encryption strength according to the security level requirements of different scenarios, and provides a real-time monitoring interface to record key logs of the encryption process for audit tracing. For example, in a remote financial account opening scenario, after the target recognition unit extracts the customer's facial feature vector, the first encryption unit preprocesses and encrypts the vector with SM3 to generate a unique and irreversible reference role identifier, which protects the customer's biometric privacy and provides a reliable role identifier for subsequent identity verification.

[0043] The identity verification generation unit, in response to a human verifier's verification operation on a user's role, generates a verification identity confirmation based on the fingerprint, the reference role identifier, the verifier information, the original summary, and the implementation timestamp. This unit can be implemented as an intelligent verification module that integrates human interaction and automated processing. Its core consists of a verification request receiving subunit, a verification information collection subunit, a verification data encapsulation subunit, and a verification result storage subunit. The review request receiving subunit is responsible for listening to low-confidence identity verification requests or manually initiated review instructions from the target recognition unit, parsing key information such as video clip ID and original identity verification data in the request, and generating a unique review task identifier. The review information collection subunit displays the video clip to be reviewed, the original identity data, and confidence details through a visual interactive interface, supporting reviewers to perform secondary confirmation, correction, or rejection of the role, while automatically collecting the reviewer's employee number, operation notes, and real-time implementation timestamp. The review data encapsulation subunit associates and integrates the collected review information with the fingerprint, reference role identifier, and original summary in the original identity verification, generating a review identity verification data packet containing the review conclusion according to a preset structured format. The review result storage subunit synchronizes the review identity verification data to a distributed trusted database and establishes a bidirectional association index with the original identity verification, supporting subsequent fast queries and full-link traceability. Furthermore, this module supports multi-level review permission management, with reviewers at different levels having different operational permissions (e.g., junior reviewers can perform routine confirmations, while senior reviewers can correct identity information and add approval comments). It also integrates an operation log auditing function, recording the complete trajectory of each review operation to ensure the traceability and compliance of the review process. For example, in a smart transportation scenario, when the target recognition unit's confidence level in identifying the driver of a violating vehicle is below 85%, the review identity generation unit automatically pushes video footage and original identity data of the violation to the traffic management personnel's review terminal. After confirming the driver's identity by viewing the high-definition video, the management personnel input the review result and remarks, and the system automatically generates a review identity confirmation and synchronizes it to the traffic violation processing system, providing authoritative identity evidence for subsequent penalty decisions. This module can be deployed in the cloud or on a local server, possessing low-latency interactive response capabilities to meet the needs of real-time review scenarios.

[0044] The methods suitable for execution by each unit in the aforementioned system are described in detail in the rest of the specification and will not be repeated here.

[0045] like Figure 2As shown, based on the same inventive concept described above, this application also provides a computing device, which can be implemented as a terminal device or a server. The computing device includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the terminal device or server. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0046] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 510 as needed so that computer programs read from it can be installed into storage section 508 as needed.

[0047] Specifically, according to embodiments of this application, the above method flow steps can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the functions defined in the system of this application.

[0048] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, or any suitable combination thereof.

[0049] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0050] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be located in a processor. The names of these units or modules do not, in certain circumstances, constitute a limitation on the unit or module itself.

[0051] In another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable storage medium stores one or more programs that are used by one or more processors to execute the methods described in this application.

[0052] In another aspect, embodiments of this application also provide a computer program product that, when executed by a processor, implements the methods of any of the above embodiments.

[0053] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions claimed in this application.

Claims

1. A method for identifying events of interest from a video stream, wherein, The event detection unit detects events of interest in the video stream to obtain one or more video segments of the events of interest and the start and end times of the video segments. The method is characterized in that, for any video segment, the following steps are performed: S1: Calculate the fingerprint of the video segment using the first digest algorithm: S2: Based on the target recognition unit, identify the characters in the video segment and obtain a reference character identifier and confidence level. The reference character identifier is calculated by the character's identification information through an irreversible first encryption algorithm. S3: Calculate the original digest based on the fingerprint, the reference role identifier, and the start and end times using the second digest algorithm; S4: Use the fingerprint, the reference role identifier, the confidence level, the real-time timestamp, and the original digest as the original identity verification.

2. The method as described in claim 1, characterized in that, Also includes: S5: The person or role identified by the target recognition unit is manually verified, and the fingerprint, the reference role identifier, the real-time timestamp, the verifyer information, and the original summary are used as the verification identity confirmation.

3. The method as described in claim 1, characterized in that, When there are multiple video clips of the events of interest, the original identity verification of subsequent video clips also includes the original summary of the previous video clip.

4. The method according to any one of claims 1-3, characterized in that, The fingerprint, the start and end times of the video clip, the reference role identifier, and the original digest are packaged together and a trusted chain signature is added to serve as a verifiable identity verification credential.

5. The method according to any one of claims 1-3, characterized in that, The event detection unit is an AI computer vision model that has been trained to detect specific motion actions.

6. The verification method for identity confirmation, wherein, The data to be verified includes identity verification credentials, character roles, and the original video stream, wherein the identity verification credentials are obtained based on the method described in claim 4; the verification method includes: Based on the start and end times of the video segments in the identity verification credential, the corresponding video segments in the original video stream are read, the video segments are calculated using a first digest algorithm to obtain a second fingerprint, and the second fingerprint is matched with the fingerprint in the identity verification credential to confirm the consistency of the video segments. The character identification information is calculated using a first encryption algorithm to obtain first encrypted information, and the first encrypted information is matched with the reference character marker in the identity verification certificate to confirm the consistency of the character. The authenticity of the identity verification credential is confirmed through a trusted chain signature.

7. An identity management system for generating events of interest from video streams, including: An event detection unit is adapted to detect events of interest in a video stream and obtain one or more video segments of the events of interest and their start and end times; The first digest unit is adapted to calculate the fingerprint of the video segment using a first digest algorithm; The target recognition unit is adapted to identify the characters and their confidence levels in the video clip; The first encryption unit is adapted to calculate character identification information using an irreversible first encryption algorithm to obtain a reference character identifier. The second digest unit is adapted to calculate the original digest based on the fingerprint, the reference role identifier, and the start and end times using a second digest algorithm. The original identity generation unit is adapted to generate an original identity verification based on the fingerprint, the reference role identifier, the confidence level, the real-time timestamp, and the original digest; The identity verification generation unit, in response to the human verifier's verification operation on the character role, generates a verification identity confirmation based on the fingerprint, the reference character identifier, the verifier information, the original summary, and the real-time timestamp.

8. A computing device, characterized in that, The computing device includes a processor and a memory, the memory storing a computer program, and the processor executing the computer program to implement the method of any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a processor, performs the method of any one of claims 1-6.

Citation Information

Patent Citations

  • Video processing method and related equipment

    CN116980715A

  • Video file management system and method based on block chain network

    CN120196787A