Operation and maintenance operation tracing method, program product, electronic equipment and storage medium

By using spatiotemporal alignment and blockchain evidence storage technology, the problem of low security in the operation and maintenance process is solved, and the entire operation and maintenance process is traceable and the behavior is associative, which enhances the security of the operation and maintenance process and the immutability of audit records.

CN121815033APending Publication Date: 2026-04-07BEIJING TOPSEC NETWORK SECURITY TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies have low security during operation and maintenance, with risks of identity theft, disconnect between operation and identity, and lack of real-time monitoring, making it impossible to generate complete security audit videos.

Method used

By using spatiotemporal alignment technology to accurately identify the identity of maintenance personnel, embedding video anchor watermarks and storing them on the blockchain, the immutability of video content and operational behavior is ensured, and tamper-proof verification information is generated.

Benefits of technology

It enables full traceability and correlation of operations and maintenance, enhances the non-repudiation of the evidence chain, and meets the requirements of judicial-level non-repudiation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815033A_ABST
    Figure CN121815033A_ABST
Patent Text Reader

Abstract

The invention provides an operation and maintenance operation tracing method, a program product, electronic equipment and a storage medium, and is applied to the technical field of network security, and the operation and maintenance operation tracing method comprises the following steps: obtaining initial video data and an operation behavior log in an operation and maintenance operation process; performing space-time alignment on the initial video data and the operation behavior log to obtain association information between the initial video data and the operation behavior log; embedding a video anchoring watermark in a video frame in the initial video data based on the associated information to obtain target video data; and generating tamper-proof verification information based on the target video data, and uploading the tamper-proof verification information to the block chain network. By constructing a complete and verifiable evidence chain from a physical operation to a digital instruction, the problem of video and log unhooking in a traditional operation and maintenance process can be solved, whole-course traceability and behavior association of the operation and maintenance operation are realized, and the final non-tampering property of an audit record is ensured through a block chain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and more specifically, to an operation and maintenance traceability method, program product, electronic device, and storage medium. Background Technology

[0002] With the development of information technology, bastion hosts, as important equipment for operation and maintenance security management, are widely used for access control and operation auditing of core assets such as servers and network devices. In particular, the emergence of portable bastion hosts has met the needs of flexible scenarios such as mobile operation and maintenance and on-site debugging. As enterprise informatization and portable bastion hosts become more widespread in the industrial control industry, Remote Desktop Protocol (RDP) operation and maintenance systems have become the main means of daily operation and maintenance. In order to ensure the security of the operation and maintenance process and the traceability after the operation and maintenance is completed, the industrial control industry needs to ensure security control during the operation and maintenance process and generate complete security audit videos after the operation and maintenance is completed.

[0003] Current technologies typically ensure security during operation and maintenance by authenticating users using information such as usernames, UKeys, and facial recognition. Once authenticated, users can perform operation and maintenance on the target assets, generating corresponding audit videos. However, these existing technologies offer relatively low security during the operation and maintenance process. Summary of the Invention

[0004] The purpose of this application is to provide a method, program product, electronic device and storage medium for tracing the source of operation and maintenance, so as to solve the technical problem of low security in the operation and maintenance process of the prior art.

[0005] In a first aspect, embodiments of this application provide a method for tracing the source of operation and maintenance, comprising: acquiring initial video data and operation behavior logs during the operation and maintenance process; performing spatiotemporal alignment on the initial video data and the operation behavior logs to obtain association information between the initial video data and the operation behavior logs; embedding video anchor watermarks into video frames in the initial video data based on the association information to obtain target video data; generating anti-tampering verification information based on the target video data, and uploading the anti-tampering verification information to a blockchain network.

[0006] In the above solution, a complete and verifiable chain of evidence, from physical operation to digital instruction, is constructed by spatiotemporally aligning initial video data with operation behavior logs, embedding video anchor watermarks based on correlation information, and storing the evidence on the blockchain. Therefore, it solves the problem of decoupling initial video data from operation behavior logs in traditional operations and maintenance processes, achieving full traceability and correlation of operations and maintenance activities, and ensuring the ultimate immutability of audit records through blockchain.

[0007] In an optional implementation, the spatiotemporal alignment of the initial video data and the operation behavior log to obtain the association information between the initial video data and the operation behavior log includes: synchronizing and correcting the timestamps of the initial video data and the operation behavior log; identifying the operation object, operator, and operation behavior in the initial video data, and mapping and associating the identification results with the semantic information in the operation behavior log to obtain the association information. In the above scheme, by explicitly including two sub-steps—high-precision time synchronization and spatial semantic association—the consistency between the initial video data and the operation behavior log in both time and semantic dimensions is ensured through proactive, precise calculation and verification, rather than a simple timestamp comparison.

[0008] In an optional implementation, the synchronization correction of the timestamp of the initial video data and the timestamp of the operation behavior log includes: determining the timestamp of the video frame based on the time signal of a multi-source high-precision clock; and periodically inserting a time reference signal frame into the initial video data based on the timestamp of the video frame. In the above scheme, the use of multi-source time synchronization fusion and the insertion of time reference signal frames improves the time accuracy within the video stream and effectively reduces time jitter introduced by transmission, encoding, and other processes, thereby giving the timestamp of the video frame high reliability and stability.

[0009] In an optional implementation, identifying the operation object and operation behavior in the initial video data and mapping the identification result to the semantic information in the operation behavior log includes: identifying the location information of the operation object and the real-time biometrics of the operator using a target detection model; determining the behavior type of the operation behavior based on continuous inter-frame changes using a motion analysis algorithm; and matching the location information and the behavior type with the semantic information using a semantic mapping engine. In the above scheme, through target detection, motion analysis, and semantic mapping, intelligent parsing and conversion from video pixels to operation semantics is achieved, thereby enabling the association of the initial video data with the operation behavior log.

[0010] In an optional implementation, the operation and maintenance traceability method further includes: calculating the spatiotemporal matching degree between the current frame of the initial video data and the current entry of the operation behavior log; and triggering an anomaly handling mechanism when the spatiotemporal matching degree is less than a matching degree threshold. In the above scheme, the spatiotemporal matching degree is introduced as a quantitative indicator, and a threshold is set to trigger anomaly handling, enabling the system to detect alignment deviations in real time and automatically. Compared with existing technologies that only allow for manual verification after the fact, this allows for dynamic quality monitoring and immediate intervention during the operation process.

[0011] In an optional implementation, calculating the spatiotemporal matching degree between the current frame of the initial video data and the current entry of the operation behavior log includes: calculating the spatiotemporal matching degree using the following formula: ; in, For the spatiotemporal matching degree, The timestamp of the operation behavior log. The timestamp of the initial video data. This refers to the operation location in the operation behavior log. This refers to the operation location within the initial video data. and These are the weighting coefficients. The intersection-union-comparison function is used. In the above scheme, by providing a specific mathematical formula to calculate the spatiotemporal matching degree, an objective, repeatable, and quantifiable scientific metric is provided for operational consistency, thus meeting the stringent requirements of objectivity and accuracy for judicial evidence.

[0012] In an optional implementation, the anomaly handling mechanism includes at least one of the following: if the time offset between the timestamp of the operation behavior log and the timestamp of the initial video data is greater than a time tolerance value, a sliding window resynchronization mechanism is triggered to calculate the optimal time alignment parameters within a local time window; if the real-time biometrics do not match the pre-stored biometrics, an alarm message is generated. In the above scheme, by distinguishing between the two core anomaly causes—time deviation and identity discrepancy—and taking targeted measures, the system can not only detect problems but also intelligently diagnose the root cause and perform precise repairs or evidence collection, thereby improving the system's self-healing capabilities and the effectiveness of its security response.

[0013] In an optional implementation, embedding a video anchor watermark in the video frames of the initial video data based on the associated information includes: embedding a visually visible video anchor watermark in the spatial domain of the video frame, wherein the video anchor watermark includes one or more of the operator's identification information, the timestamp of the video frame, and the associated information; and / or performing a frequency domain transformation on the video frame and embedding a machine-readable video anchor watermark in the transform domain coefficients, wherein the video anchor watermark includes the operation behavior corresponding to the video frame. In the above scheme, by combining a dual embedding strategy of visible watermarks in the spatial domain and hidden watermarks in the frequency domain, the convenience requirement for rapid manual identification and the deep security requirement for machine anti-counterfeiting verification are simultaneously met. Furthermore, since visible watermarks can provide intuitive evidence, while hidden watermarks can provide tamper-resistant deep binding, the combination of the two enhances the effectiveness of the video as evidence.

[0014] In an optional implementation, generating tamper-proof verification information based on the target video data includes: segmenting the target video data and calculating the hash value of each video segment; constructing an integrity verification tree based on the hash values ​​of the video segments, wherein the tamper-proof verification information includes the root hash value of the integrity verification tree. In the above scheme, by segmenting the target video data, calculating hashes, and constructing an integrity verification tree, the integrity verification of massive amounts of video data is transformed into the verification of a single root hash value, thereby improving the efficiency of integrity verification. Furthermore, by combining it with blockchain storage, the judicial-grade evidence preservation requirements of tamper-proof and easily verifiable video evidence can be achieved.

[0015] Secondly, embodiments of this application provide an operation and maintenance traceability device, comprising: an acquisition module for acquiring initial video data and operation behavior logs during the operation and maintenance process; an alignment module for performing spatiotemporal alignment on the initial video data and the operation behavior logs to obtain association information between the initial video data and the operation behavior logs; an embedding module for embedding video anchor watermarks into video frames in the initial video data based on the association information to obtain target video data; and a generation module for generating anti-tampering verification information based on the target video data and uploading the anti-tampering verification information to a blockchain network.

[0016] In the above solution, a complete and verifiable chain of evidence, from physical operation to digital instruction, is constructed by spatiotemporally aligning initial video data with operation behavior logs, embedding video anchor watermarks based on correlation information, and storing the evidence on the blockchain. Therefore, it solves the problem of decoupling initial video data from operation behavior logs in traditional operations and maintenance processes, achieving full traceability and correlation of operations and maintenance activities, and ensuring the ultimate immutability of audit records through blockchain.

[0017] In an optional implementation, the alignment module is specifically used to: synchronize and correct the timestamps of the initial video data and the operation behavior log; identify the operation object, operator, and operation behavior in the initial video data, and map and associate the identification results with the semantic information in the operation behavior log to obtain the associated information. In the above scheme, by explicitly defining spatiotemporal alignment as including two sub-steps—high-precision time synchronization and spatial semantic association—it ensures that the consistency between the initial video data and the operation behavior log in both time and semantic dimensions is actively, accurately calculated, and verified, rather than a simple timestamp comparison.

[0018] In an optional implementation, the alignment module is further configured to: determine the timestamp of the video frame based on the time signal of a multi-source high-precision clock; and periodically insert time reference signal frames into the initial video data based on the timestamp of the video frame. In the above scheme, the use of multi-source time synchronization fusion and the insertion of time reference signal frames improves the time accuracy within the video stream and effectively reduces time jitter introduced by transmission, encoding, and other processes, thereby giving the timestamp of the video frame high reliability and stability.

[0019] In an optional implementation, the alignment module is further configured to: identify the location information of the object being operated on and the real-time biometric features of the operator using a target detection model; determine the behavior type of the operation based on continuous inter-frame changes using a motion analysis algorithm; and match the location information and the behavior type with the semantic information using a semantic mapping engine. In the above scheme, through target detection, motion analysis, and semantic mapping, intelligent parsing and conversion from video pixels to operational semantics is achieved, thereby enabling the association of initial video data with operation behavior logs.

[0020] In an optional implementation, the operation and maintenance traceability device further includes: a calculation module for calculating the spatiotemporal matching degree between the current frame of the initial video data and the current entry of the operation behavior log; and a triggering module for triggering an anomaly handling mechanism when the spatiotemporal matching degree is less than a matching degree threshold. In the above scheme, the spatiotemporal matching degree is introduced as a quantitative indicator, and a threshold is set to trigger anomaly handling, enabling the system to detect alignment deviations in real time and automatically. Compared with existing technologies that only allow for manual verification after the fact, this allows for dynamic quality monitoring and immediate intervention during the operation process.

[0021] In an optional implementation, the calculation module is specifically used to: calculate the spatiotemporal matching degree using the following formula: ; in, For the spatiotemporal matching degree, The timestamp of the operation behavior log. The timestamp of the initial video data. This refers to the operation location in the operation behavior log. This refers to the operation location within the initial video data. and These are the weighting coefficients. The intersection-union-comparison function is used. In the above scheme, by providing a specific mathematical formula to calculate the spatiotemporal matching degree, an objective, repeatable, and quantifiable scientific metric is provided for operational consistency, thus meeting the stringent requirements of objectivity and accuracy for judicial evidence.

[0022] In an optional implementation, the anomaly handling mechanism includes at least one of the following: if the time offset between the timestamp of the operation behavior log and the timestamp of the initial video data is greater than a time tolerance value, a sliding window resynchronization mechanism is triggered to calculate the optimal time alignment parameters within a local time window; if the real-time biometrics do not match the pre-stored biometrics, an alarm message is generated. In the above scheme, by distinguishing between the two core anomaly causes—time deviation and identity discrepancy—and taking targeted measures, the system can not only detect problems but also intelligently diagnose the root cause and perform precise repairs or evidence collection, thereby improving the system's self-healing capabilities and the effectiveness of its security response.

[0023] In an optional implementation, the embedding module is specifically used to: embed a visually visible video anchor watermark in the spatial domain of the video frame, wherein the video anchor watermark includes one or more of the operator's identification information, the timestamp of the video frame, and the associated information; and / or, perform a frequency domain transformation on the video frame, embedding a machine-readable video anchor watermark in the transform domain coefficients, wherein the video anchor watermark includes the operation behavior corresponding to the video frame. In the above scheme, by combining a dual embedding strategy of visible watermarks in the spatial domain and hidden watermarks in the frequency domain, the convenience requirement for rapid manual identification and the deep security requirement for machine anti-counterfeiting verification are simultaneously met. Furthermore, since visible watermarks can provide intuitive evidence, while hidden watermarks can provide tamper-resistant deep binding, the combination of the two enhances the effectiveness of the video as evidence.

[0024] In an optional implementation, the generation module is specifically used to: segment the target video data and calculate the hash value of each video segment; construct an integrity verification tree based on the hash values ​​of the video segments, wherein the anti-tampering verification information includes the root hash value of the integrity verification tree. In the above scheme, by segmenting the target video data, calculating hashes, and constructing an integrity verification tree, the integrity verification of massive amounts of video data is transformed into the verification of a single root hash value, thereby improving the efficiency of integrity verification. Furthermore, by combining it with blockchain storage, the judicial-grade evidence preservation requirements of tamper-proof and easily verifiable video evidence can be achieved.

[0025] Thirdly, embodiments of this application provide a computer program product, including computer program instructions, which are read and executed by a processor to perform the operation and maintenance traceability method as described in the first aspect.

[0026] Fourthly, embodiments of this application provide an electronic device, including: a processor, a memory, and a bus; the processor and the memory communicate with each other through the bus; the memory stores computer program instructions that can be executed by the processor, and the processor can execute the operation and maintenance traceability method as described in the first aspect by calling the computer program instructions.

[0027] Fifthly, embodiments of this application provide a computer-readable storage medium that stores computer program instructions. When the computer program instructions are executed by a computer, the computer performs the operation and maintenance traceability method as described in the first aspect.

[0028] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, embodiments of this application are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 A flowchart of an operation and maintenance traceability method provided in this application embodiment; Figure 2 This application provides a structural block diagram of an operation and maintenance traceability device. Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0031] To ensure the security and traceability of the operation and maintenance process, the industrial control system needs to guarantee security control during operation and maintenance and generate complete security audit videos after completion. However, existing bastion hosts have the following main problems: First, attackers can bypass authentication by stealing accounts or forging biometrics, thus posing a risk of identity theft; second, because video recordings and operation logs are stored separately, it is difficult to trace the specific person responsible for the operation, thus posing a risk of disconnect between operation and identity; third, because the authenticity of personnel identity during the operation cannot be dynamically verified, there is a lack of real-time monitoring.

[0032] For example, existing technologies can implement operation and maintenance through the following process: First, the user enters the corresponding account and password, the facial features of the operation and maintenance personnel are captured by the camera, and a UKey is inserted to authenticate the user; second, the transmitted images in the RDP protocol are parsed and the images are used to generate a standard video stream, generating an operation and maintenance audit video file; finally, a digital watermark is embedded in the header of the video file, and the watermark content includes the operator ID and the start and end time of the operation.

[0033] The above-mentioned operation and maintenance process has the following problems: First, it relies on one-way identity authentication (such as facial recognition) during the initial login and does not continuously verify the identity of personnel during the operation. Attackers can steal the Session Token through session hijacking and switch the actual operator after authentication. Second, video recordings and operation logs are stored in different systems (such as local hard drives and cloud databases), lacking a spatiotemporal alignment mechanism. Third, it is impossible to prove that the operator in the video is actually operating the equipment on site, which poses a risk of remote forgery.

[0034] Therefore, existing technologies have shortcomings in terms of security during operation and maintenance processes and in generating complete security audit videos. In view of this, this application provides an operation and maintenance traceability method. It accurately determines the identity of operation and maintenance personnel through spatiotemporal alignment technology to prevent impersonation; it uses video anchoring technology to create a physical-logical dual binding between video content and the operator's identity and actions, enhancing the non-repudiation of the evidence chain; and it generates tamper-proof verification information for key frames of the video and log files, stores it in the blockchain, and ensures that tampering is detected in real time by the blockchain verification mechanism, meeting judicial-level non-repudiation requirements.

[0035] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0036] Please refer to Figure 1 , Figure 1 This application provides a flowchart of an operation and maintenance traceability method, which can be, but is not limited to, executed by an electronic device. Figure 3 The possible structure of this electronic device is shown below; for details, please refer to the following section. Figure 3 The above-mentioned methods for tracing the source of operational operations can specifically include: S101: Obtain initial video data and operation logs during the operation and maintenance process.

[0037] Operations and maintenance (O&M) operations refer to the technical instructions or interactive behaviors executed by O&M personnel with specific system privileges on target assets through professional O&M management tools or protocols to ensure, monitor, repair, or configure information systems, resulting in state changes or information acquisition. Target assets refer to critical computing devices, systems, services, or data resources within an organization's Information Technology (IT) or Operational Technology (OT) environment, which are remotely accessed, configured, managed, maintained, or monitored by O&M personnel via portable bastion hosts.

[0038] Understandably, initial video data and operation behavior logs will be generated during the operation and maintenance process. The initial video data refers to the raw video stream directly captured by the built-in camera of the portable bastion host without being processed by this method, including the screen of the operator performing the task, the operation interface, and interactive actions. The operation behavior log refers to the structured text record of the portable bastion host in the operation and maintenance session, arranged in chronological order, which includes the instructions, parameters, timestamps and target assets of each operation.

[0039] It should be noted that the embodiments of this application do not specifically limit the specific implementation methods for obtaining the aforementioned initial video data and operation behavior logs, and those skilled in the art can make appropriate adjustments according to the actual situation. For example, the initial video data and operation behavior logs sent by an external device can be received; or, the initial video data and operation behavior logs stored in advance can be read from local storage or the cloud; or, the aforementioned initial video data and operation behavior logs can be collected in real time, etc.

[0040] For example, when maintenance personnel log in to the target asset through a portable bastion host to perform operations, the system starts two recording channels in parallel: one channel continuously collects initial video data, including the operator, the operation interface, and interactive actions, through the device's camera at a fixed frame rate; the other channel captures and parses keyboard, mouse commands, and command line input through a protocol proxy to form a structured operation behavior log.

[0041] S102: Perform spatiotemporal alignment on the initial video data and operation behavior logs to obtain the correlation information between the initial video data and the operation behavior logs.

[0042] Spatiotemporal alignment refers to the process of establishing a precise correspondence between visual events in the initial video data and text records in the operation behavior log in terms of time and spatial location through technical means. The associated information refers to the structured data generated after spatiotemporal alignment, which describes the correspondence between video frames in the initial video data and log entries in the operation behavior log, such as: matching video frame identifiers (ID), log entry IDs, matching degree scores, etc.

[0043] The above-mentioned S102 can perform fine-grained spatiotemporal alignment processing on the acquired initial video data and operation behavior logs. For example, this processing first calibrates the time base of both through a high-precision time synchronization service to ensure alignment in the time dimension; secondly, it achieves alignment in the spatial dimension based on target detection and tracking; and finally, it achieves alignment in both the time and spatial dimensions through joint time and space verification.

[0044] For example, nanosecond-level time synchronization between video streams and operation commands is achieved through atomic clock timing and NTP protocol; combined with YOLOv8 object detection algorithm, a spatiotemporal mapping relationship between operation objects (keyboard / monitor / mouse) and log semantics is established.

[0045] S103: Based on the association information, embed video anchor watermarks into the video frames of the initial video data to obtain the target video data.

[0046] Video anchor watermarks are digital markers embedded in video frames that carry specific information. They are divided into visually visible video anchor watermarks (i.e., visible watermarks) and visually invisible but machine-readable video anchor watermarks (i.e., hidden watermarks). Correspondingly, target video data refers to the final video file generated after embedding video anchor watermarks and used for auditing.

[0047] Based on the association information obtained in S102, the above-mentioned S103 can dynamically embed a video anchor watermark in the aligned video frames. For example, the text "Operator: Zhang San | Operation: Submit Form | Time: 2023-10-27 10:30:25.123" can be displayed in semi-transparent text in the lower right corner of the video screen; at the same time, a hidden watermark containing the hash value of the log can be embedded in the video frequency domain.

[0048] For example, visible watermarking: embedding semi-transparent text containing biometric hash values ​​in the lower right corner of the video (transparency ≤30%); hidden watermarking: replacing the hash value of the embedded operation type with the least significant bit of the Discrete Wavelet Transform (DWT) domain, resisting JPEG compression (Q≥80) and frame cropping (retaining ≥70% of the content).

[0049] S104: Generate tamper-proof verification information based on the target video data and upload the tamper-proof verification information to the blockchain network.

[0050] Tamper-proof verification information refers to a data digest or signature used to verify the integrity and authenticity of target video data; for example, tamper-proof verification information is a hash value. A blockchain network is a distributed, immutable data storage and verification system maintained by multiple nodes, forming a chain of data blocks linked in chronological order.

[0051] One implementation method is to use Hyperledger Fabric to build a private blockchain, with nodes including bastion host vendors, user organizations, and forensic institutions. This blockchain supports the national cryptographic algorithms SM2 / SM3 / SM4, meeting domestic compliance requirements. Correspondingly, the evidence storage process can include: uploading video segment hash values ​​to the blockchain via a TLS 1.3 encrypted channel; and triggering on-chain hash value retrieval for each audit operation to verify video integrity.

[0052] For example, the Hyperledger Fabric consortium blockchain is used to store video frame hash values, operation instruction hash values, and biometric hash values; a Merkle tree is constructed to verify data integrity, and the national cryptographic SM3 hash algorithm and SM2 digital signature are supported.

[0053] In the above solution, a complete and verifiable chain of evidence, from physical operation to digital instruction, is constructed by spatiotemporally aligning initial video data with operation behavior logs, embedding video anchor watermarks based on correlation information, and storing the evidence on the blockchain. Therefore, it solves the problem of decoupling initial video data from operation behavior logs in traditional operations and maintenance processes, achieving full traceability and correlation of operations and maintenance activities, and ensuring the ultimate immutability of audit records through blockchain.

[0054] The specific implementation of the above-mentioned spatiotemporal alignment is described below, that is, S102 may specifically include: S201: Synchronize and correct the timestamps of the initial video data and the operation behavior logs.

[0055] A timestamp is a numerical value that identifies the moment when data is generated or processed; synchronization correction is the process of adjusting the time of two or more independent systems to keep them consistent.

[0056] For example, the system can maintain a high-precision local clock, which is timed by an atomic clock (e.g., Global Positioning System (GPS), BeiDou satellite system, or rubidium atomic clock) to ensure that the local clock accuracy is no greater than ±1μs. For each frame of initial video data and each operation behavior log, the system can timestamp them based on this high-precision clock when they are generated.

[0057] S202: Identify the operation object, operator, and operation behavior in the initial video data, and map and associate the identification results with the semantic information in the operation behavior log to obtain the associated information.

[0058] The operation object refers to the entity related to the operation and maintenance in the video frame, such as: the monitor screen area, keyboard, mouse, or specific buttons or windows on the graphical user interface (GUI). Semantic information refers to the meaning and context of the operation instructions recorded in the operation behavior log. Mapping association refers to establishing a correspondence from one data form (such as visual coordinates) to another data form (such as semantic description).

[0059] For example, the system can use a deep learning-based object detection model to perform real-time analysis of video frames. This model can identify the object being manipulated (such as detecting and outlining screen areas and hand areas) and the operator's facial features. Simultaneously, by analyzing the movement trajectories of the hand or mouse pointer between consecutive frames, it determines the operation behavior (such as moving, clicking, or dragging). On the other hand, the system parses the operation behavior logs and extracts semantic information. Finally, the results obtained from the visual analysis are matched with the semantic information of the logs.

[0060] In the above scheme, by clearly defining that spatiotemporal alignment includes two sub-steps: high-precision time synchronization and spatial semantic association, it is ensured that the consistency between the initial video data and the operation behavior log in both time and semantic dimensions is actively and accurately calculated and verified, rather than simply a timestamp comparison.

[0061] Furthermore, based on the above embodiments, S201 may specifically include: S301: Determine the timestamp of a video frame based on the time signal of a multi-source high-precision clock.

[0062] Multi-source high-precision clocks refer to two or more independent sources that can provide high-precision time signals, such as GPS satellite signals, BeiDou satellite signals, and rubidium atomic clock hardware modules. In other words, portable bastion hosts can have built-in or external multi-source high-precision clocks connected to determine the timestamps of video frames using these clock signals.

[0063] As one implementation method, at least three high-precision clocks can be deployed, and signal deviations can be eliminated through a weighted algorithm. The specific calculation formula is as follows: ; in, This indicates the time signal components of three high-precision clocks. The comprehensive local metric result obtained by weighted summation Indicates subscript The summation is performed on all items from 1 to 3, which includes... The sum of the three items, Indicates the first The weight of each component, Indicates the first The target parameter values ​​for each component.

[0064] Weighting coefficient Defined as: Where the weights are non-negative (i.e., ), and the sum of all weights is 1 (i.e. ); Indicates the first Each component variance ( Standard deviation (Variance) measures the dispersion of the component data. The larger the value, the higher the corresponding weight. The higher the value, the better the component contributes to the overall result. The greater the contribution.

[0065] S302: Time reference signal frames are periodically inserted into the initial video data based on the timestamps of the video frames.

[0066] A time reference signal frame is a special frame (such as a completely black frame or a frame with a specific encoding pattern) that is regularly inserted into a video stream. It serves as an internal time scale to detect and correct time jitter during video playback. In the embodiments of this application, a time reference signal frame (such as inserting one black synchronization frame per second) can be embedded in the initial video data to compensate for latency jitter caused by network transmission or encoding.

[0067] For example, in the video encoding stage, the system can periodically (e.g., at integer intervals of every second) force a time reference signal frame into the original video frame sequence based on a predetermined timestamp sequence of video frames. This frame's content is simple and clear, facilitating accurate detection at the decoding end. During network transmission or local playback, even if delays or jitter occur due to packet loss or decoding latency, the receiving end can detect the actual arrival intervals of these reference frames, deduce the timeline distortion, and dynamically compensate for the timestamps of other ordinary frames, thereby reconstructing a high-precision timeline at the receiving end.

[0068] In the above scheme, the technology of multi-source time synchronization fusion and insertion of time reference signal frames is adopted to improve the time accuracy within the video stream and effectively reduce the time jitter introduced by transmission, encoding and other links, thereby making the timestamp of the video frame have high reliability and stability.

[0069] Furthermore, based on the above embodiments, S202 may specifically include: S401: Through a target detection model, the location information of the object being operated on and the real-time biometric features of the operator are identified.

[0070] Object detection models are deep learning-based computer vision models used to identify specific objects in an image and output their category and bounding box location. Examples include the You Only Look Once (YOLO) model and the Single Shot MultiBox Detector (SSD) model. Real-time biometrics refer to the physiological or behavioral characteristics of an operator extracted in real-time from the initial video data being processed, such as facial feature vectors and hand geometric features.

[0071] For example, the system loads a pre-trained object detection model (e.g., YOLO v8-tiny, balancing speed and accuracy); for each frame of input video, the model performs inference, identifies key operational objects (e.g., monitor, keyboard, mouse, etc.) and operator features (e.g., face, left hand, etc.), and outputs their positional information (bounding box coordinates). Simultaneously, for the detected face region, a face recognition network is used to extract a feature vector as the real-time biometric feature for that frame.

[0072] S402: Determine the behavior type of the operation based on the changes between consecutive frames using a motion analysis algorithm.

[0073] Motion analysis algorithms are techniques used to analyze the motion information of objects in a continuous sequence of images. For example, optical flow is used to calculate the speed and direction of motion of pixels.

[0074] For example, the system caches consecutive video frames. Using motion analysis algorithms (e.g., Farneback optical flow), it calculates the pixel motion vector field within the hand or mouse pointer detection box area between the current frame and the previous frame. By analyzing the overall direction and variation pattern of this vector field, the system determines the type of action, such as: stationary (preparing to click), linear movement (moving the cursor), or instantaneous motion convergence (click action).

[0075] S403: Matches location information and behavior type with semantic information through a semantic mapping engine.

[0076] A semantic mapping engine is a software module whose function is to map low-level, perception-level data (such as coordinates and action types) into high-level, business-meaning descriptions.

[0077] For example, the semantic mapping engine receives location information from S401 and behavior type from S402; simultaneously, the semantic mapping engine obtains the semantic information of the currently matched entry in the operation behavior log. The engine executes the matching logic: first, it performs a coarse time screening (time stamps are close), then calculates whether the spatial distance between the location information and the log coordinates is within a threshold, and checks whether the behavior types are consistent; if the match is successful, it generates an association information, recording the association information between the video frame ID and the log entry ID.

[0078] For example, the detected operation object is associated with the semantic instructions in the operation log (such as "mouse click coordinates (100,200)" corresponding to "click the 'delete' button"), and the target position in the occluded scene is predicted by Kalman filtering.

[0079] In the above scheme, intelligent parsing and conversion from video pixels to operation semantics is achieved through object detection, motion analysis, semantic mapping and other processing, so that the initial video data can be associated with operation behavior logs.

[0080] Furthermore, based on the above embodiments, the operation and maintenance traceability method provided in this application may further include: S501: Calculate the spatiotemporal matching degree between the current frame of the initial video data and the current entry of the operation behavior log.

[0081] For example, the spatiotemporal matching degree can be calculated using the following formula: ; in, For spatiotemporal matching degree, Timestamps for operation behavior logs. The timestamp of the initial video data. This refers to the location of the operation in the operation behavior log. This refers to the operation location within the initial video data. and These are the weighting coefficients. This is the Intersection over Union (IoU) function.

[0082] The Intersection over Union (IoU) function is a commonly used similarity metric in computer vision and image processing, used to measure the degree of overlap between two regions. Its core idea is to evaluate spatial consistency by calculating the ratio of the intersection area to the union area of ​​two regions. In this formula, the IoU function is responsible for calculating the location parameters. and Spatial overlap. Weighting coefficients. and Optimization is typically achieved through the training set, usually by taking... .

[0083] S502: When the spatiotemporal matching degree is less than the matching degree threshold, the exception handling mechanism is triggered.

[0084] The system compares the calculated spatiotemporal matching degree with a pre-set matching degree threshold. When the spatiotemporal matching degree is less than the matching degree threshold, it indicates that the alignment attempt has failed, which may mean: first, the initial video data and the operation behavior log did not record the same event at that moment; second, there was a significant deviation in time synchronization; third, spatial detection or mapping errors; fourth, identity impersonation leading to inconsistent behavior patterns. In this case, the system can trigger an exception handling mechanism, which will attempt self-repair or issue a security alarm depending on the specific circumstances.

[0085] For example, an exception handling mechanism may include at least one of the following: The first approach is to trigger a sliding window resynchronization mechanism if the time offset between the timestamp of the operation behavior log and the timestamp of the initial video data is greater than the time tolerance value, and calculate the optimal time alignment parameters within a local time window.

[0086] During the calculation of spatiotemporal matching, if the system detects that the time offset is continuously greater than the time tolerance value, it determines that a systematic time baseline drift has occurred. At this time, the system triggers a sliding window resynchronization mechanism. This mechanism expands a certain range forward and backward from the current anomaly point to form a sliding window. Within the local time window, the locally optimal alignment parameters are recalculated. The system then uses these new parameters to compensate for the timestamps of subsequent video frames, achieving self-repair.

[0087] The second method is to generate an alarm message if the real-time biometrics do not match the pre-stored biometrics.

[0088] The system continuously compares real-time biometric features (the current operator's facial features) with pre-stored biometric features of legitimate users of the account during login (e.g., calculating the cosine similarity between feature vectors). If a mismatch is found between the real-time and pre-stored biometric features, it is determined that the operation was not performed by the user, and an alarm message is immediately generated. Simultaneously, the system can record the image in the operation video frame as direct evidence for subsequent accountability.

[0089] In the above scheme, spatiotemporal matching degree is introduced as a quantitative indicator, and a threshold is set to trigger anomaly handling, enabling the system to detect alignment deviations in real time and automatically. Compared with the existing technology that can only be manually verified after the fact, it can realize dynamic quality monitoring and real-time intervention in the operation process.

[0090] Furthermore, based on the above embodiments, S103 may specifically include: S601: Embed a visually visible video anchor watermark in the spatial domain of a video frame, wherein the video anchor watermark includes one or more of the operator's identification information, the video frame's timestamp, and associated information.

[0091] The spatial domain refers to the image representation domain directly composed of pixel color values. Watermarks embedded in this domain are implemented by modifying pixel values.

[0092] For example, the system selects a region in the spatial domain of a video frame that does not affect the main operation content (e.g., the lower right corner, a non-operational hotspot area (avoiding the keyboard and mouse operation areas)) and embeds a visually visible text or graphic watermark therein. The information contained in this watermark can be extracted from associated information, such as: the operator's identification information (e.g., name, employee ID, biometric hash value, etc.), the precise timestamp corresponding to the video frame, and the operation semantics in the associated information (e.g., SSH login, database query, etc.).

[0093] As one implementation method, the watermark can be presented in a semi-transparent manner (e.g., transparency ≤ 30%) to ensure visual traceability without affecting the readability of the video content.

[0094] S602: Perform frequency domain transformation on the video frame and embed a machine-readable video anchor watermark in the transform domain coefficients, wherein the video anchor watermark includes the operation behavior corresponding to the video frame.

[0095] Frequency domain transformation is an image processing technique that converts an image from the spatial domain to the frequency domain (e.g., through Discrete Cosine Transform (DCT) or DWT), representing the image as a combination of different frequency components. The transform domain coefficients are a set of values ​​obtained after the image undergoes frequency domain transformation, representing the intensity of different frequency components.

[0096] For example, the system first performs a frequency domain transform on the video frames, decomposing the image into multiple frequency domain coefficients. Then, in the selected transform domain coefficients (e.g., low-frequency coefficients), the least significant bit (LSB) of the operation behavior (e.g., the operation behavior itself, the hash value corresponding to the operation behavior, the behavioral feature code extracted based on the operation behavior, etc.) is embedded, and the hidden information is carried out through bit substitution. Because the modification is extremely small and located in the frequency domain, the watermark is invisible to the human eye and has strong resistance to common video compression and cropping attacks. It is mainly used for authenticity verification and deep correlation in forensic identification.

[0097] Furthermore, the watermark density and position can be dynamically adjusted based on the camera's perspective verification results (e.g., detecting that the distance between the operator and the device is no more than 1 meter); alternatively, biometric vectors can be generated by combining keyboard keystroke force sensor data (e.g., triggering watermark intensity enhancement when the keystroke depth is no less than 2 mm) and associated with the video stream to enhance identity binding strength; alternatively, a gyroscope can be introduced to monitor changes in device posture (e.g., pausing watermark embedding when the device tilt angle is no less than 15°); or, a Long Short-Term Memory (LSTM) network can be used to learn the operator's keystroke rhythm characteristics (e.g., keystroke interval time distribution) to generate dynamic behavioral signatures.

[0098] In the above scheme, by combining a dual embedding strategy of visible watermarks in the spatial domain and hidden watermarks in the frequency domain, the convenience requirement of rapid manual identification and the deep security requirement of machine anti-counterfeiting verification are simultaneously met. In addition, since visible watermarks can provide intuitive evidence and hidden watermarks can provide tamper-resistant deep binding, the combination of the two enhances the effectiveness of video as evidence.

[0099] Furthermore, based on the above embodiments, S104 may specifically include: S701: Perform segmentation on the target video data and calculate the hash value of each video segment.

[0100] Segmentation refers to cutting complete target video data into a series of smaller, continuous segments in chronological order. A hash value is a fixed-length, unique digital fingerprint calculated by a hash function (such as SHA-256), where any tiny change to the input data will result in a completely different hash value.

[0101] For example, the system segments the target video data file (e.g., generating a hash node every 10 seconds), and then uses a cryptographic hash algorithm to calculate the hash value of each video segment.

[0102] S702: Construct an integrity verification tree based on the hash values ​​of video segments, wherein the tamper-proof verification information includes the root hash value of the integrity verification tree.

[0103] Integrity verification tree refers to a tree-like data structure (e.g., Merkle tree) where the leaf nodes are the hash values ​​of data blocks, and the non-leaf nodes are the result of hashing the hash values ​​of their child nodes. The hash value of the root node can represent the integrity of all data under the entire tree. As long as the content of any segment in the video is tampered with, its hash value will change, which will then be passed up layer by layer, eventually changing the root hash value.

[0104] For example, the system uses these sharded hash values ​​as leaf nodes to construct an integrity verification tree. As one implementation, the construction process of the integrity verification tree may include: concatenating two adjacent hash values ​​and hashing them again to generate their parent node hash values; recursively performing this process until a unique root hash value is finally generated.

[0105] Furthermore, the real-time integrity verification process may include: receiving an audit request; extracting the fragment hash; performing integrity verification tree verification to obtain integrity confirmation.

[0106] In the above scheme, by calculating hashes from the target video data and constructing an integrity verification tree, the integrity verification of massive amounts of video data is transformed into the verification of a single root hash value, thereby improving the efficiency of integrity verification. In addition, by combining blockchain storage, the judicial-grade evidence preservation requirements of tamper-proof and easily verifiable video evidence can be achieved.

[0107] The following example illustrates the operation and maintenance traceability method provided in this application. First, the deployment environment is described: A provincial power grid dispatch center is equipped with a multi-screen portable bastion host, which supports the simultaneous operation of two remote machines; the camera is equipped with an infrared thermal imaging module to monitor the distance between the operator and the equipment in real time (threshold ≤ 1.5 meters).

[0108] First, the system analyzes Zhang's finger flexion and extension force using electromyography (EMG) signals (≥3N is considered a valid operation); second, the video anchoring module captures the operation area (1080P resolution @ 60fps), and embeds the instruction hash value into the DWT field; finally, the operation behavior log and video stream are stored on the Hyperledger Fabric consortium blockchain, supporting the national cryptographic standard SM4 encryption.

[0109] Therefore, by dynamically binding real-time feature acquisition with hardware behavior (tapping force / touch pressure), the cost of forgery is high (it requires copying physiological characteristics and interactive behavior simultaneously); video anchoring technology achieves a four-in-one association of time, space, operation, and identity, meeting the requirements of the "Electronic Data Forensics Rules"; blockchain evidence storage supports cross-chain verification, and the accuracy of hash value tampering detection is high; operation latency is ≤50ms, and it supports real-time processing of 4K video streams; it complies with 12 international / national standards such as IEC 62443-4-2, GDPR, and Cybersecurity Classified Protection 2.0; the bioelectric signal acquisition module adopts a dry electrode design, allowing for seamless interaction by the operator; the dynamic watermark transparency is ≤30%, which does not affect the identification of key information on the operation screen.

[0110] Please refer to Figure 2 , Figure 2 This application provides a structural block diagram of an operation and maintenance traceability device 800, which includes: an acquisition module 801 for acquiring initial video data and operation behavior logs during the operation and maintenance process; an alignment module 802 for performing spatiotemporal alignment on the initial video data and the operation behavior logs to obtain association information between the initial video data and the operation behavior logs; an embedding module 803 for embedding video anchor watermarks into video frames in the initial video data based on the association information to obtain target video data; and a generation module 804 for generating anti-tampering verification information based on the target video data and uploading the anti-tampering verification information to a blockchain network.

[0111] In the above solution, a complete and verifiable chain of evidence, from physical operation to digital instruction, is constructed by spatiotemporally aligning initial video data with operation behavior logs, embedding video anchor watermarks based on correlation information, and storing the evidence on the blockchain. Therefore, it solves the problem of decoupling initial video data from operation behavior logs in traditional operations and maintenance processes, achieving full traceability and correlation of operations and maintenance activities, and ensuring the ultimate immutability of audit records through blockchain.

[0112] Furthermore, based on the above embodiments, the alignment module 802 is specifically used to: synchronize and correct the timestamp of the initial video data with the timestamp of the operation behavior log; identify the operation object, operator and operation behavior in the initial video data, and map and associate the identification result with the semantic information in the operation behavior log to obtain the association information.

[0113] In the above scheme, by clearly defining that spatiotemporal alignment includes two sub-steps: high-precision time synchronization and spatial semantic association, it is ensured that the consistency between the initial video data and the operation behavior log in both time and semantic dimensions is actively and accurately calculated and verified, rather than simply a timestamp comparison.

[0114] Furthermore, based on the above embodiments, the alignment module 802 is also used to: determine the timestamp of the video frame based on the time signal of the multi-source high-precision clock; and periodically insert time reference signal frames into the initial video data based on the timestamp of the video frame.

[0115] In the above scheme, the technology of multi-source time synchronization fusion and insertion of time reference signal frames is adopted to improve the time accuracy within the video stream and effectively reduce the time jitter introduced by transmission, encoding and other links, thereby making the timestamp of the video frame have high reliability and stability.

[0116] Furthermore, based on the above embodiments, the alignment module 802 is also used to: identify the location information of the operation object and the real-time biometrics of the operator through a target detection model; determine the behavior type of the operation behavior based on continuous inter-frame changes through a motion analysis algorithm; and match the location information and the behavior type with the semantic information through a semantic mapping engine.

[0117] In the above scheme, intelligent parsing and conversion from video pixels to operation semantics is achieved through object detection, motion analysis, semantic mapping and other processing, so that the initial video data can be associated with operation behavior logs.

[0118] Furthermore, based on the above embodiments, the operation and maintenance traceability device 800 further includes: a calculation module, used to calculate the spatiotemporal matching degree between the current frame of the initial video data and the current entry of the operation behavior log; and a triggering module, used to trigger an exception handling mechanism when the spatiotemporal matching degree is less than the matching degree threshold.

[0119] In the above scheme, spatiotemporal matching degree is introduced as a quantitative indicator, and a threshold is set to trigger anomaly handling, enabling the system to detect alignment deviations in real time and automatically. Compared with the existing technology that can only be manually verified after the fact, it can realize dynamic quality monitoring and real-time intervention in the operation process.

[0120] Furthermore, based on the above embodiments, the calculation module is specifically used to: calculate the spatiotemporal matching degree using the following formula: ; in, For the spatiotemporal matching degree, The timestamp of the operation behavior log. The timestamp of the initial video data. This refers to the operation location in the operation behavior log. This refers to the operation location within the initial video data. and These are the weighting coefficients. It is the intersection-union ratio function.

[0121] The above scheme provides a specific mathematical formula to calculate the spatiotemporal matching degree, thus providing an objective, repeatable, and quantifiable scientific metric for operational consistency, thereby meeting the strict requirements of judicial evidence for objectivity and accuracy.

[0122] Furthermore, based on the above embodiments, the anomaly handling mechanism includes at least one of the following: if the time offset between the timestamp of the operation behavior log and the timestamp of the initial video data is greater than the time tolerance value, a sliding window resynchronization mechanism is triggered to calculate the optimal time alignment parameter within a local time window; if the real-time biometrics do not match the pre-stored biometrics, an alarm message is generated.

[0123] In the above solution, by distinguishing between the two core anomalies of time deviation and identity discrepancy and taking targeted measures, the system can not only discover problems, but also intelligently diagnose the root causes of problems and perform precise repairs or evidence collection, thereby improving the system's self-healing ability and the effectiveness of security response.

[0124] Furthermore, based on the above embodiments, the embedding module 803 is specifically used to: embed a visually visible video anchor watermark in the spatial domain of the video frame, wherein the video anchor watermark includes one or more of the operator's identification information, the timestamp of the video frame, and the associated information; and / or, perform a frequency domain transformation on the video frame and embed a machine-readable video anchor watermark in the transform domain coefficients, wherein the video anchor watermark includes the operation behavior corresponding to the video frame.

[0125] In the above scheme, by combining a dual embedding strategy of visible watermarks in the spatial domain and hidden watermarks in the frequency domain, the convenience requirement of rapid manual identification and the deep security requirement of machine anti-counterfeiting verification are simultaneously met. In addition, since visible watermarks can provide intuitive evidence and hidden watermarks can provide tamper-resistant deep binding, the combination of the two enhances the effectiveness of video as evidence.

[0126] Furthermore, based on the above embodiments, the generation module 804 is specifically used for: performing segmentation processing on the target video data and calculating the hash value of each video segment; constructing an integrity verification tree based on the hash values ​​of the video segments, wherein the anti-tampering verification information includes the root hash value of the integrity verification tree.

[0127] In the above scheme, by calculating hashes from the target video data and constructing an integrity verification tree, the integrity verification of massive amounts of video data is transformed into the verification of a single root hash value, thereby improving the efficiency of integrity verification. In addition, by combining blockchain storage, the judicial-grade evidence preservation requirements of tamper-proof and easily verifiable video evidence can be achieved.

[0128] Please refer to Figure 3 , Figure 3 This application provides a structural block diagram of an electronic device 900, which includes at least one processor 901, at least one communication interface 902, at least one memory 903, and at least one communication bus 904. The communication bus 904 enables direct communication between these components, the communication interface 902 facilitates signaling or data communication with other node devices, and the memory 903 stores machine-readable instructions executable by the processor 901. When the electronic device 900 is running, the processor 901 communicates with the memory 903 via the communication bus 904. When the machine-readable instructions are invoked by the processor 901, the aforementioned operation and maintenance traceability method is executed.

[0129] For example, the processor 901 in this embodiment of the application can read a computer program from the memory 903 via the communication bus 904 and execute the computer program to implement the following method: acquiring initial video data and operation behavior logs during the operation and maintenance process; performing spatiotemporal alignment on the initial video data and the operation behavior logs to obtain association information between the initial video data and the operation behavior logs; embedding video anchor watermarks into the video frames in the initial video data based on the association information to obtain target video data; generating anti-tampering verification information based on the target video data, and uploading the anti-tampering verification information to the blockchain network.

[0130] The processor 901 comprises one or more, and can be an integrated circuit chip with signal processing capabilities. The processor 901 can be a general-purpose processor, including a Central Processing Unit (CPU), a Microcontroller Unit (MCU), a Network Processor (NP), or other conventional processors; it can also be a special-purpose processor, including a Neural-network Processing Unit (NPU), a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Furthermore, when there are multiple processors 901, some can be general-purpose processors, and others can be special-purpose processors.

[0131] The memory 903 includes one or more, which may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0132] Understandable. Figure 3 The structure shown is for illustrative purposes only; the electronic device 900 may also include components that are more advanced than those shown. Figure 3 The more or fewer components shown, or having the same Figure 3 The different configurations shown. Figure 3The components shown can be implemented using hardware, software, or a combination thereof. In the embodiments of this application, the electronic device 900 can be, but is not limited to, physical devices such as desktop computers, laptops, smartphones, smart wearable devices, and in-vehicle devices, or virtual devices such as virtual machines. Furthermore, the electronic device 900 is not necessarily a single device; it can be a combination of multiple devices, such as a server cluster, etc.

[0133] This application also provides a computer program product, including a computer program stored on a computer-readable storage medium. The computer program includes computer program instructions. When the computer program instructions are executed by a computer, the computer can perform the steps of the operation and maintenance traceability method described in the above embodiments, such as: S101: Obtaining initial video data and operation behavior logs during the operation and maintenance process. S102: Performing spatiotemporal alignment on the initial video data and operation behavior logs to obtain association information between the initial video data and the operation behavior logs. S103: Embedding video anchor watermarks in video frames of the initial video data based on the association information to obtain target video data. S104: Generating anti-tampering verification information based on the target video data and uploading the anti-tampering verification information to the blockchain network.

[0134] This application also provides a computer-readable storage medium that stores computer program instructions. When the computer program instructions are executed by a computer, the computer performs the operation and maintenance traceability method described in the foregoing method embodiments.

[0135] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0136] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0137] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0138] It should be noted that if the function is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0139] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.

[0140] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for tracing the source of operation and maintenance, characterized in that, include: Acquire initial video data and operation logs during the operation and maintenance process; Spatiotemporal alignment is performed on the initial video data and the operation behavior log to obtain the association information between the initial video data and the operation behavior log. Based on the aforementioned association information, a video anchor watermark is embedded into the video frames of the initial video data to obtain the target video data. Based on the target video data, anti-tampering verification information is generated and uploaded to the blockchain network.

2. The operation and maintenance traceability method according to claim 1, characterized in that, The step of performing spatiotemporal alignment on the initial video data and the operation behavior log to obtain the association information between the initial video data and the operation behavior log includes: The timestamps of the initial video data and the timestamps of the operation behavior logs are synchronized and corrected. The operation object, operator, and operation behavior in the initial video data are identified, and the identification results are mapped and associated with the semantic information in the operation behavior log to obtain the associated information.

3. The operation and maintenance traceability method according to claim 2, characterized in that, The synchronization and correction of the timestamp of the initial video data with the timestamp of the operation behavior log includes: The timestamp of the video frame is determined based on the time signal from a multi-source high-precision clock. Time reference signal frames are periodically inserted into the initial video data based on the timestamps of the video frames.

4. The operation and maintenance traceability method according to claim 2, characterized in that, The process of identifying the operation objects and operations in the initial video data and mapping and associating the identification results with the semantic information in the operation behavior log includes: The target detection model identifies the location information of the object being operated on and the real-time biometric features of the operator. The behavior type of the operation is determined based on changes between consecutive frames using a motion analysis algorithm. The location information and behavior type are matched with the semantic information using a semantic mapping engine.

5. The operation and maintenance traceability method according to claim 4, characterized in that, Also includes: Calculate the spatiotemporal matching degree between the current frame of the initial video data and the current entry of the operation behavior log; When the spatiotemporal matching degree is less than the matching degree threshold, the exception handling mechanism is triggered.

6. The operation and maintenance traceability method according to claim 5, characterized in that, The calculation of the spatiotemporal matching degree between the current frame of the initial video data and the current entry of the operation behavior log includes: The spatiotemporal matching degree is calculated using the following formula: ; in, For the spatiotemporal matching degree, The timestamp of the operation behavior log. The timestamp of the initial video data. This refers to the operation location in the operation behavior log. This refers to the operation location within the initial video data. and These are the weighting coefficients. It is the intersection-union ratio function.

7. The operation and maintenance traceability method according to claim 6, characterized in that, The exception handling mechanism includes at least one of the following: If the time offset between the timestamp of the operation behavior log and the timestamp of the initial video data is greater than the time tolerance value, the sliding window resynchronization mechanism is triggered to calculate the optimal time alignment parameters within a local time window. If the real-time biometrics do not match the pre-stored biometrics, an alarm message will be generated.

8. The operation and maintenance traceability method according to any one of claims 1-7, characterized in that, The step of embedding a video anchor watermark in the video frames of the initial video data based on the association information includes: A visually visible video anchor watermark is embedded in the spatial domain of the video frame, wherein the video anchor watermark includes one or more of the operator's identification information, the timestamp of the video frame, and the associated information; And / or, The video frame is subjected to frequency domain transformation, and a machine-readable video anchor watermark is embedded in the transform domain coefficients, wherein the video anchor watermark includes the operation behavior corresponding to the video frame.

9. The operation and maintenance traceability method according to any one of claims 1-7, characterized in that, The generation of anti-tampering verification information based on the target video data includes: The target video data is segmented, and the hash value of each video segment is calculated. An integrity verification tree is constructed based on the hash values ​​of the video segments, wherein the anti-tampering verification information includes the root hash value of the integrity verification tree.

10. A computer program product, characterized in that, It includes computer program instructions, which are read and executed by a processor to perform the operation and maintenance traceability method as described in any one of claims 1-9.

11. An electronic device, characterized in that, include: Processor, memory, and bus; The processor and the memory communicate with each other via the bus; The memory stores computer program instructions that can be executed by the processor, and the processor can execute the operation and maintenance traceability method as described in any one of claims 1-9 by calling the computer program instructions.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a computer, cause the computer to perform the operation and maintenance traceability method as described in any one of claims 1-9.