A method and system for secure management of audio and video files

By designing an audio-visual file security management system, including adaptive buffering strategies, deeply integrated source identification embedding and multi-dimensional metadata structure construction, the problems of audio-visual file transmission stability, file source identification and file management efficiency are solved, and efficient management and security protection of audio-visual files are realized.

CN119691708BActive Publication Date: 2025-06-17STATE GRID ANHUI ELECTRIC POWER CO LTD TONGCHENG POWER SUPPLY CO +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510207590.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-17
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The prior art cannot realize adaptive adjustment of buffer capacity during audio and video file acquisition and storage, resulting in data overflow or insufficient, affecting transmission stability; at the same time, file source identification embedding is not deeply integrated enough, making it difficult to achieve clear labeling and reliable traceability of key attributes; when users browse and tag files, they lack effective mechanisms to extract key interval fragments and build multi-dimensional metadata indexes, which affect file management and retrieval efficiency.

Method used

A video and video file security management system is designed, including coding processing module, key marking module and security protection module. The encoding processing module adjusts the buffer capacity through adaptive buffering strategies and embeds source identification information; the key marking module extracts key interval fragments and constructs a multi-dimensional metadata structure positioning index by monitoring user operations; the security protection module generates dynamic encryption keys based on the metadata structure and performs hierarchical encryption.

Benefits of technology

The buffer capacity adaptive adjustment during audio and video file transfer is realized, avoiding data overflow or insufficient, and ensuring transmission stability; deeply integrated source identification embedding realizes clear annotation and reliable traceability of file sources; the construction of multi-dimensional metadata structures improves file management and retrieval efficiency; dynamic encryption keys and hierarchical encryption enhance the security of files and prevents illegal acquisition and tampering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119691708B_ABST
    Figure CN119691708B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for secure management of audio-visual files, belonging to the technical field of file security. After the camera captures the audio-visual files, they are stored in a buffer area. The buffer capacity is adjusted according to the changes in network bandwidth and file byte count, and the source identification information is encoded and embedded before being stored in the secure storage area. The key marking operations of the users on the audio-visual files in the secure storage area are monitored, the key interval segments are extracted, and the key marking information is obtained through screening and identification, and then a multi-dimensional metadata structure positioning index is constructed. For the audio-visual files in the secure storage area, an initial encryption key is generated based on the metadata structure positioning index, and then the encryption level is determined according to the data volume and format of the file and hierarchical encryption is performed. It solves the technical problems such as the adaptive buffer capacity in the acquisition and storage of audio-visual files, the processing of key marking information and index construction after user operations, and ensures the transmission, traceability, management and security protection of audio-visual files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of file security, and more particularly to a method and system for secure management of video and audio files. Background Art

[0002] With the rapid development of technology, cameras have been extremely widely used in fields such as news reporting, documentary filming, and self-media creation. These application scenarios have generated a vast amount of video and audio files, thus bringing severe challenges to file management, and further highlighting the importance of secure management methods and systems. In the encoding process of camera video and audio files, a reasonable encoding method can better preserve the quality and details of video and audio with limited storage resources. For example, adopting an advanced video encoding standard can effectively compress the file size, facilitate file storage and transmission, and also lay a foundation for subsequent security management operations. Some encoding algorithms can reserve specific fields for embedding carriers of security identifiers or encrypted information. The functions of key marking and retrieval are of great significance for the management of camera-shot materials. In news gathering work, a journalist may shoot a large amount of materials in a day. By marking key information, people, locations, etc., such as marking the speech segments of important people and the occurrence scenes of emergencies, valuable content can be quickly retrieved from numerous files in the later stage, greatly improving work efficiency and reducing the time cost of material sorting. And file encryption is the key defense line for ensuring the security of camera video and audio files. Encryption technology can ensure that these files are not obtained, viewed, or tampered with by unauthorized third parties during the process of being stored in the camera memory card, transmitted to a computer or other devices, and shared with specific personnel, thereby safeguarding the rights and interests of all parties and information security, and enabling camera video and audio files to be under secure and reliable control throughout their entire life cycle.

[0003] Chinese Patent with the authorization announcement number CN107977551B discloses a method, device, and electronic device for protecting files, which relates to file security technology and can ensure the integrity of MP4 file content. The method for protecting files includes: obtaining an index box included in a target file; writing digital rights management information in a user data box in the index box; updating a mapping table box and a block offset box of samples and blocks in the index box of the target file according to the written digital rights management information and the position of the index box in the target file; and performing a digital signature on the updated target file.

[0004] Although the prior art ensures the integrity of the content of MP4 files, it still fails to solve the technical problems of adaptive adjustment of the buffer capacity during the acquisition and storage of video and audio files, embedding of source identification, extraction of key marking information and construction of multi-dimensional indexes during user browsing and playing, so as to ensure the stable transmission, traceability and convenient management of video and audio files. Therefore, in order to overcome these limitations, the present invention proposes a method and system for the security management of video and audio files. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a method and system for the security management of video and audio files, which solves the following problems: First, during the transmission of video and audio files, the buffer capacity cannot be adaptively adjusted according to the real-time transmission rate and data volume of the files, which is likely to cause data overflow or deficiency, thus affecting the continuous reception of video and audio files and the overall transmission stability; Second, the embedding of source identification information in video and audio files is not deeply integrated enough, making it difficult to clearly mark and reliably trace key attributes such as the source of the files; Third, when users browse, play and make key marking operations on the stored video and audio files, there is a lack of an effective mechanism to accurately extract key interval video and audio segments, screen key frames, identify targets and classify and extract audio data, etc., so that a comprehensive and multi-dimensional metadata structure positioning index cannot be constructed, which is not conducive to the subsequent rapid and accurate search and management of relevant video and audio files. To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A video and audio file security management system includes an encoding and processing module, a key marking module and a security protection module:

[0007] The encoding and processing module is used to store the video and audio file in the buffer when the camera captures the video and audio file, adjust the buffer capacity by monitoring the changes in network bandwidth and the number of bytes of the video and audio file, and encode and embed the source identification information of the video and audio file in the buffer into the secure storage area;

[0008] The key marking module is used to extract the key interval video and audio segments by monitoring the key marking operations when the user plays or browses the video and audio files in the secure storage area, screen the key frames by using the amount of motion change, perform target recognition, extract audio data and classify them to obtain key marking information, then define a metadata structure template according to the identification information to construct a directed graph, and generate a multi-dimensional metadata structure positioning index by processing the new identification information records through a comparison function and reviewing the enumerated type records;

[0009] The security protection module is used to encrypt the video and audio files in the secure storage area, generate an initial encryption key based on the metadata structure positioning index of the video and audio files, and determine the encryption level according to the data volume and format of the video and audio files, and encrypt the video and audio files at different levels.

[0010] Specifically, the encoding processing module includes a data buffer unit and an information embedding unit;

[0011] An adaptive buffering strategy is configured in the data buffer unit. The adaptive buffering strategy is used to adjust the buffer capacity in real time according to the transmission rate and data volume of the video and audio files transmitted by the camera. By continuously monitoring the network bandwidth and the change of the number of bytes of the video and audio files during the transmission process, a buffer state evaluation model is constructed to evaluate the buffer state, so as to dynamically allocate memory space as the buffer;

[0012] A deep fusion strategy is configured in the information embedding unit. The deep fusion strategy is used to embed the source identification information into the data structure of the video and audio files, and store the processed video and audio files in the secure storage area.

[0013] Specifically, the specific steps of the adaptive buffering strategy include:

[0014] Set the initial buffer capacity and obtain the camera parameter information;

[0015] Configure a time window and count the number of bytes of the video and audio files received within the time window;

[0016] Configure a sliding window to track the long-term trend and short-term fluctuations of the data traffic. Add the number of bytes of the video and audio files counted within the time window to the sliding window data sequence, and remove the number of bytes of the video and audio files counted in the earliest time window to keep the number of data points in the sliding window fixed;

[0017] Configure an evaluation interval. Every other evaluation interval, calculate the buffer occupancy rate by the number of bytes occupied in the buffer and the buffer capacity, and calculate the remaining data duration that the buffer can continue to accommodate;

[0018] Evaluate the buffer state, construct a buffer state evaluation model based on the buffer occupancy rate, the remaining data duration that can be accommodated, and the network bandwidth, and calculate the buffer state evaluation score;

[0019] Configure an adjustment threshold. If the buffer state evaluation score is less than the adjustment threshold, trigger the buffer capacity increase operation; otherwise, do not perform any processing.

[0020] Specifically, the buffer capacity increase operation includes:

[0021] Calculate the growth rate of the number of bytes within each sliding window based on the number of bytes of the audio - video file in the sliding window data sequence;

[0022] Dynamically increase the buffer capacity according to the growth rate of the number of bytes within the sliding window and the buffer occupancy rate, that is:

[0023]

[0024] Among them, is the adjusted buffer capacity, is the buffer capacity, is the buffer occupancy rate, is the average value of the growth rate of the number of bytes within the sliding window, and are respectively the minimum threshold and the maximum threshold of the growth rate of the number of bytes within the sliding window, is the buffer occupancy rate threshold, , and are all weighting coefficients.

[0025] Specifically, the key - point marking module includes a marking acquisition unit and an index construction unit;

[0026] The marking acquisition unit is configured with a marking capture strategy. The marking capture strategy is used to monitor the key - point marking operations made by the user during the playback or browsing of the audio - video file, calculate and screen key frames through the amount of motion change, generate a content summary through target recognition of the key frames, extract audio data and classify it to identify and extract key - point marking information;

[0027] The index construction unit is configured with an association and fusion strategy. The association and fusion strategy is used to index the key - point marking information and source identification information of the audio - video file, generate a multi - dimensional metadata structure location index according to the preset metadata structure template, add metadata structure identification information records through a comparison function, and conduct information review on the enumerated type of identification information records.

[0028] Specifically, the steps of the marking capture strategy include:

[0029] When the user starts to play or browse the audio - video file in the buffer, monitor the key - point marking operation and identify the key - point marking type;

[0030] By monitoring the user's interaction behavior, record the marking time point of the audio - video file where the key - point marking is located, and generate a time stamp for each key - point marking;

[0031] According to the key - point marking type, configure the marking extraction interval based on each key - point marking time stamp, and obtain the audio - video segments of the key interval according to the marking extraction interval;

[0032] Convert the video and audio segments in the key interval into a series of image frames, mark the first image frame as the key frame, configure the vector change threshold, and calculate the motion change amount of the image frames, that is:

[0033]

[0034] Wherein, is the th image frame and the th image frame of the motion change amount, is the motion vector at the position in the image frame , is the motion vector at the position in the image frame ;

[0035] If the motion change amount is greater than the vector change threshold, then mark as the key frame, and then, starting from , continue to calculate the adjacent motion change amount ; otherwise, skip , calculate the motion change amount , until the image frames of the video and audio segment are traversed;

[0036] Perform object recognition on the key frames, obtain the key marking information of the key frames, including text information, target objects, target scenes, and convert the recognition result of the key marking information of the key frames into a text description as the content summary;

[0037] Extract the audio data in the video and audio segments of the key interval, identify the key marking information of the audio data, classify the audio data, and identify the emotional color of the audio data.

[0038] Specifically, the steps of the association and fusion strategy include:

[0039] Define the metadata structure template according to the identification information, define the metadata structure as a directed graph, and define the attribute types for each vertex of the metadata structure. The attribute types include enumeration type, string type, and array type;

[0040] Obtain the identification information records of the video and audio files. For each obtained identification information record, perform identification information type comparison with the metadata structure vertices through the comparison function, construct the comparison function of the identification information record and the metadata structure vertices, and query whether there is the same identification information type as the identification information record in the metadata structure vertices;

[0041] If there is an identification information type that is the same as the identification information record, it indicates that the comparison of the identification information types is successful, and the identification information record is filled in the position corresponding to the vertex of the metadata structure of the same identification information type. If there is no identification information type that is the same as the identification information record, the comparison of the identification information types fails, and a temporary metadata structure vertex is created to store the identification information record for which the comparison fails.

[0042] Specifically, the steps of the association and fusion strategy further include:

[0043] When a new identification information record of an audio-visual file is added to the metadata structure , select the identification information records with the enumerated type attribute for information review to query whether the identification information records of the enumerated type exist in the enumerated type set in the metadata structure;

[0044] If in the enumerated type set in the metadata structure, it is queried that the identification information type with the identification information record existing is the same as the information type of the identification information record of the newly added audio-visual file, the information review passes; otherwise, the information review fails.

[0045] For the information identification records that fail the information review, filter out the similar information identification records in the metadata structure for replacement;

[0046] Calculate the similarity metric value between the information identification records that fail the information review and the information identification records of the same information type in the enumerated type set, and select the information identification record with the largest similarity metric value in the enumerated type set in the metadata structure to replace the identification information record of the newly added audio-visual file.

[0047] Specifically, the security protection module includes a key generation unit and an encryption execution unit;

[0048] The key generation unit is configured with a feature fusion strategy, which is used to generate an encryption key with dynamic update characteristics for the audio-visual file. By extracting feature data from the location index of the metadata structure of the audio-visual file, an original feature set is formed, and after performing a hash operation on the original feature set, key material is generated, and the initial encryption key is generated using the key material as the input;

[0049] The encryption execution unit is configured with an adaptive encryption strategy, which is used to perform hierarchical encryption on the audio-visual file according to the initial encryption key. By evaluating the data volume of the audio-visual file and the audio-visual file format, the encryption level is determined, and the audio-visual file is adaptively encrypted according to the encryption level.

[0050] A method for security management of audio-visual files includes the following steps:

[0051] Step S1: When the camera captures an audio-visual file, store the audio-visual file in the buffer. By monitoring the changes in network bandwidth and the number of bytes of the audio-visual file, adjust the buffer capacity, and encode and embed the source identification information of the audio-visual file in the buffer and then store it in the secure storage area;

[0052] Step S2: By monitoring the key marking operations when the user plays or browses the audio-visual files in the secure storage area, extract the key interval audio-visual segments according to the key marking operations, use the amount of motion change to screen key frames and perform target recognition, extract audio data and classify them to obtain key marking information, then define a metadata structure template based on the identification information to construct a directed graph, and generate a multi-dimensional metadata structure positioning index by processing the newly added identification information records through a comparison function and reviewing the enumeration type records;

[0053] Step S3: Encrypt the audio-visual files in the secure storage area, generate an initial encryption key based on the metadata structure positioning index of the audio-visual files, and determine the encryption level according to the data volume and format of the audio-visual files, and encrypt the audio-visual files at different levels.

[0054] Advantages of the present invention:

[0055] 1. The adaptive buffering strategy in the encoding processing module can, according to the transmission rate and data volume of the audio-visual files transmitted by the camera, monitor the changes in network bandwidth and the number of file bytes in real time, construct an evaluation model and dynamically allocate memory space as the buffer accordingly, effectively avoiding data overflow or shortage, ensuring the stability and integrity of the audio-visual files during the transmission process, and guaranteeing that the data can be continuously received and successfully stored in the secure storage area, which is particularly crucial for scenarios that require long-term and large-scale audio-visual data collection and storage.

[0056] 2. The key marking module, through the marking capture strategy, monitors the key marking operations when the user plays or browses the audio-visual files, and then screens key frames, identifies targets, extracts audio data and classifies them, etc., to accurately obtain key marking information, which is convenient for users to quickly locate the content of interest subsequently, and also provides a basis for further exploring the value of audio-visual files based on user feedback. For example, personalized recommendations and other application scenarios can be generated based on these key markings, and the key marking information can be indexed with the source identification information to construct a multi-dimensional metadata structure positioning index, facilitating more efficient retrieval and query of files that meet specific requirements in a large number of audio-visual files, and improving the overall management efficiency.

[0057] 3. The key generation unit of the security protection module, whose feature fusion strategy can locate and extract multiple types of features from the metadata structure of the audio-visual file to generate an encryption key with the characteristic of dynamic update. Compared with the fixed key, the dynamic update can greatly reduce the risk of the key being cracked, increase the security of the audio-visual file encryption, and ensure that the data is not illegally obtained and tampered with during storage and subsequent possible transmission and other links. Description of the Drawings

[0058] Figure 1 It is a schematic structural diagram of a security management system for audio-visual files according to the present invention;

[0059] Figure 2 It is a flowchart of the specific steps of the adaptive buffering strategy according to the present invention;

[0060] Figure 3 It is a schematic diagram of the structure of the camera according to the present invention Figure 1 ;

[0061] Figure 4 It is a flowchart of the specific steps of the marker capture strategy according to the present invention;

[0062] Figure 5 It is a schematic diagram of the structure of the camera according to the present invention Figure 2 ;

[0063] Figure 6 It is a flowchart of the specific steps of the association fusion strategy according to the present invention;

[0064] Figure 7 It is a flowchart of a security management method for audio-visual files according to the present invention. Detailed Description of the Invention

[0065] Example 1

[0066] Please refer to Figure 1 , this example introduces a security management system for audio-visual files, including an encoding processing module, a key marking module, and a security protection module;

[0067] The encoding processing module is used to store the audio-visual file in the buffer when the camera captures the audio-visual file, adjust the buffer capacity by monitoring the network bandwidth and the change of the number of bytes of the audio-visual file, and encode and embed the source identification information of the audio-visual file in the buffer;

[0068] Preferably, the encoding processing module includes a data buffer unit and an information embedding unit;

[0069] An adaptive buffering strategy is configured in the data buffer unit. The adaptive buffering strategy is used to adjust the buffer capacity in real time according to the transmission rate and data volume of the video and audio files transmitted by the camera. By continuously monitoring the network bandwidth and the change in the number of bytes of the video and audio files during the transmission process, a buffer status evaluation model is constructed to evaluate the buffer status, so as to dynamically allocate memory space as the buffer, avoid data overflow or shortage, ensure the stability and integrity of data transmission, and guarantee the continuous reception of video and audio files;

[0070] A deep fusion strategy is configured in the information embedding unit. The deep fusion strategy is used to embed the identification information into the data structure of the video and audio files, combine the identification information with the video and audio files, and transmit it along with the video and audio files to be stored in the secure storage area, so as to realize the clear marking and reliable traceability of key attributes such as the source of the video and audio files.

[0071] Please refer to Figure 2 , preferably, the specific steps of the adaptive buffering strategy include:

[0072] Please refer to Figure 3 , when the WIFI indicator light is on, the video and audio files obtained by the camera are transmitted to the buffer, the initial buffer capacity is set, and the camera parameter information is obtained, including resolution, frame rate, encoding format, sampling rate and number of channels;

[0073] Configure a time window , and count the number of bytes of the video and audio files received within the time window, which is calculated through the video and audio encoding format, resolution and frame rate. By calculating the number of bytes within a fixed duration, the change of data traffic in the short term can be obtained in time, providing basic data points for subsequent sliding window data update and data traffic analysis, and being able to monitor the fluctuation of data traffic with a smaller time granularity so as to capture the change trend of data traffic more sensitively;

[0074] Configure a sliding window , which is used to track the long-term trend and short-term fluctuation of data traffic. The sliding window is larger than the time window. Add the number of bytes of the video and audio files counted within the time window to the sliding window data sequence , and remove the number of bytes of the video and audio files counted in the earliest time window, keeping the number of data points within the sliding window fixed at , where is the The number of bytes of an audio-visual file changes as the sliding window is updated; the number of bytes of the audio-visual file counted within multiple time windows is integrated into a data sequence, which can more comprehensively reflect the overall change characteristics of the data traffic over a period of time, including statistical indicators such as the mean and standard deviation of the data traffic, thereby providing strong support for accurately evaluating the stability and trend of the data traffic, enabling the adjustment decision of the buffer capacity not to be based solely on the current instantaneous traffic situation, but comprehensively considering the data traffic change pattern over the past period of time, effectively avoiding misjudgments caused by short-term burst traffic.

[0075] Configure the evaluation interval. Every other evaluation interval, calculate the buffer occupancy rate based on the number of bytes occupied in the buffer and the buffer capacity and calculate the remaining data duration that the buffer can continue to accommodate : :

[0076]

[0077] where, is the mean number of bytes of the current sliding window data sequence, that is: ; By calculating the buffer occupancy rate and the remaining data duration that can be accommodated, timely grasp the usage situation of the buffer and the sustainability of the remaining capacity, providing key quantitative indicators for judging whether the buffer needs to adjust its capacity.

[0078] Conduct a status evaluation of the buffer. Based on the buffer occupancy rate, the remaining data duration that can be accommodated, and the network bandwidth, construct a buffer status evaluation model and calculate the buffer status evaluation score, that is:

[0079]

[0080] where, is the buffer status evaluation score, is the current network bandwidth, is the set standard network bandwidth, and are non-negative weighting coefficients respectively, used to balance the dimensions of the buffer occupancy rate and the remaining data duration that can be accommodated; comprehensively quantify the overall status of the buffer considering multiple key factors, and the calculated buffer status evaluation score provides an intuitive and unified indicator to measure the quality of the buffer.

[0081] Configure the adjustment threshold. If the buffer status evaluation score is less than the adjustment threshold, trigger the buffer capacity increase operation; otherwise, do not perform any processing;

[0082] The buffer capacity increase operation includes:

[0083] Calculate the growth rate of the number of bytes within the sliding window, i.e.:

[0084]

[0085] wherein, is the growth rate of the number of bytes in the th time window within the sliding window, is the number of bytes of the th audio - video file in the sliding window data sequence, is the number of bytes of the th audio - video file in the sliding window data sequence;

[0086] Dynamically increase the buffer capacity according to the growth rate of the number of bytes within the sliding window and the buffer occupancy rate, i.e.:

[0087]

[0088] wherein, is the adjusted buffer capacity, is the average value of the growth rate of the number of bytes within the sliding window, , and are respectively the minimum threshold and the maximum threshold of the growth rate of the number of bytes within the sliding window, is the buffer occupancy rate threshold, , and are non - negative weighting coefficients obtained through testing and verification. When the buffer status evaluation score is less than the adjustment threshold, trigger the buffer capacity increase operation. By calculating the growth rate of the number of bytes within the sliding window and combining the buffer occupancy rate, adopting the strategy of dynamically increasing the buffer capacity can accurately adjust the buffer capacity according to the growth trend of the data flow and the current usage of the buffer, ensure that when the data flow increases, the buffer has enough space to accommodate more data, avoid data overflow, maintain the stability and continuity of data transmission, and at the same time avoid waste of system resources caused by excessive capacity increase, improve the utilization efficiency of resources, and optimize the performance of the entire audio - video file transmission system.

[0089] Preferably, the specific steps of the deep fusion strategy include:

[0090] Collect source identification information, including reading the device identification code of the camera, obtaining location information using the built - in positioning module of the camera, and obtaining the time information when the audio - video file is shot, and construct an information set based on the collected source identification information;

[0091] Perform binary coding conversion on the information set, convert the information set into a binary code stream sequence, and record the length of the binary code stream sequence;

[0092] Parse the audio and video files, screen the set of positions where the identification information can be embedded, and for each position, evaluate and obtain the embeddable capacity corresponding to each position; Exemplarily, a calculation method based on information entropy is used to measure the data redundancy, so as to determine the maximum number of bytes of the identification information that can be accommodated in each position, ensuring that the embedding operation will not damage the basic coding structure and playback function of the audio and video.

[0093] Perform source identification information embedding, and use the least significant bit replacement method or the embedding method based on the transform domain to replace the binary bits of the source identification information into the positions where the identification information can be embedded;

[0094] Configure the evaluation threshold, calculate the evaluation metrics of the audio and video files after embedding the identification information, including signal-to-noise ratio, structural similarity index, and perceptual audio quality evaluation. If the evaluation metrics are lower than the evaluation threshold, change the embedding position or intensity until the evaluation metrics of the audio and video files after embedding the identification information reach the evaluation threshold.

[0095] The key marking module extracts the key interval audio and video segments according to the key marking operation by monitoring the key marking operation when the user plays or browses the audio and video files in the secure storage area, filters the key frames by using the amount of motion change and performs target recognition, extracts the audio data and classifies it to obtain the key marking information, and then constructs a directed graph according to the metadata structure template defined by the identification information. Through the comparison function, the newly added identification information records are processed, and the enumeration type records are reviewed to generate a multi-dimensional metadata structure positioning index;

[0096] Preferably, the key marking module includes a marking acquisition unit and an index construction unit;

[0097] The marking acquisition unit is configured with a marking capture strategy, which is used to monitor the key marking operations made by the user during the process of playing or browsing the audio and video files, calculate and filter the key frames by using the amount of motion change, generate a content summary through target recognition of the key frames, extract the audio data and classify it to identify and extract the key marking information;

[0098] The index construction unit is configured with an association and fusion strategy, which is used to index the key marking information and the source identification information of the audio and video files, and generate a multi-dimensional metadata structure positioning index according to the preset metadata structure template. The metadata structure identification information records are newly added through the comparison function, and the information of the enumeration type identification information records is reviewed.

[0099] Please refer to Figure 4 , preferably, the specific steps of the marking capture strategy include:

[0100] Please refer to Figure 5When the user starts playing or browsing the video and audio files in the buffer, monitor the key marking operations and identify the key marking types, including favorites, likes, loves, and annotations;

[0101] By monitoring the interactive behaviors of the user, record the marking time points of the video and audio files where the key markings are located, and generate time stamps for each key marking;

[0102] According to the key marking types, based on each key marking time stamp, configure the marking extraction interval, which is used to determine the number of key frames to be extracted near the marking point. According to the marking extraction interval, obtain the key interval of the video and audio segments, where is the time stamp of the th key marking, is the marking extraction interval of the key marking type of the th key marking;

[0103] Convert the video and audio segments in the key interval into a series of image frames, mark the first image frame as the key frame, configure the vector change threshold, and calculate the motion change amount of the image frames, that is:

[0104]

[0105] where is the th image frame and the th image frame 's motion change amount, is the motion vector at the position in the image frame , representing the motion direction and amplitude at that position, is the motion vector at the position in the image frame ;

[0106] If the motion change amount is greater than the vector change threshold, then mark as the key frame, and then, starting from , continue to calculate the adjacent motion change amount ; otherwise, skip , and calculate the motion change amount ; until the image frames of the video and audio segments are traversed;

[0107] Perform object recognition on the key frames, obtain the key marking information of the key frames, including text information, target objects, and target scenes, and convert the recognition result of the key marking information of the key frames into a text description as the content summary;

[0108] Extract the audio data within the video and audio segments of the key intervals, identify the key marking information of the audio data, including the dialogue introduction, sound effects, and music types, classify the audio data, including dialogue, background music, ambient sounds, and special sound effects, and identify the emotional color of the audio data, including joy, sadness, and anger.

[0109] Please refer to Figure 6 , preferably, the specific steps of the association and fusion strategy include:

[0110] Define a metadata structure template according to the identification information. The identification information includes the source identification information and the key marking information. Among them, the source identification information includes the device identification code of the camera, the location information of the video and audio file shooting, and the time information of the video and audio file shooting; the key marking information includes the target object, the target scene, the text information, the audio data type, and the audio data color;

[0111] Define the metadata structure as a directed graph, where the vertex set is , is the th identification information, is the number of identification information types. The elements in the vertex set correspond one by one to the identification information of the video and audio file. Exemplarily, the device identification code is , the location information of the video and audio file shooting is , the time information of the video and audio file shooting is , the target object is , and so on, thus constructing a metadata structure framework with clear logic and distinct levels, laying a foundation for subsequent information filling and association operations.

[0112] Define the attribute types for each vertex of the metadata structure. The attribute types include enumeration type, string type, and array type. Exemplarily, the device identification code belongs to the string type, and the target object, the target scene, and the audio data type belong to the enumeration type;

[0113] Obtain the identification information record of the video and audio file , where is the number of obtained identification information, is the th identification information record;

[0114] For each obtained identification information record, construct a comparison function between the identification information record and the metadata structure vertex through a comparison function, and query whether there is the same identification information type as the identification information record in the metadata structure vertex;

[0115] That is:

[0116]

[0117] Among them, is the th comparison function between the identification information record and the vertex of the metadata structure . is the identification information type of the th identification information record;

[0118] If the value of the comparison function is , that is, there exists an identification information type identical to the identification information record, it indicates that the comparison of the identification information type is successful. Then, the identification information record is filled into the corresponding position of the vertex of the metadata structure with the same identification information type, that is, is filled into the corresponding position of the vertex of the metadata structure . If there does not exist an identification information type identical to the identification information record, that is, the value of the comparison function is , then the comparison of the identification information type fails. A temporary metadata structure vertex is created to store the identification information record for which the comparison fails; if the identification information comparison fails, that is, the obtained identification information cannot find a completely matching type in the existing metadata structure, a temporary metadata structure vertex is created specifically to store this identification information for which the comparison fails, and at the same time, the relevant reasons for the comparison failure and information characteristics are recorded for subsequent further analysis and processing. This flexible processing method can effectively handle various complex and changeable identification information situations and ensure the integrity and adaptability of the metadata structure.

[0119] When an identification information record of a new audio - video file is added to the metadata structure , select the identification information record whose attribute is of the enumeration type for information review. Among them, is the th identification information record whose attribute is of the enumeration type. Among them, is the number of identification information records whose attribute is of the enumeration type, is less than . If in the enumeration type set in the metadata structure, it is found that there exists an identification information type of the identification information record identical to that of the identification information record of the newly added audio - video file, then the information review passes; otherwise, the information review fails. The formula for information review is:

[0120]

[0121] Among them, is the review result of the th identification information record whose attribute is of the enumeration type, It is a set of enumeration types in the metadata structure. It is the identification information type of the th identification information record in the set of enumeration types; if the value is 1, it means that there are duplicate records and the information review is passed. If the value is 0, it means that there are no duplicate records and the information review is not passed; through this review mechanism, duplicate records or potential data inconsistency problems can be discovered in a timely manner, ensuring the uniqueness and accuracy of the enumeration type information in the metadata structure, and avoiding interference with subsequent data retrieval, association, and analysis operations due to duplicate or incorrect enumeration type records. This helps to maintain the high quality of the metadata structure and the high credibility of the data, enabling various applications based on this information to obtain reliable results.

[0122] For the information identification record with the information review not passed , among which, is the th information identification record that fails the review. is less than or equal to , filter the similar information identification records in the metadata structure for replacement, calculate the similarity measurement value between the information identification record with the information review not passed and the information identification records of the same information type in the set of enumeration types, select the information identification record with the largest similarity measurement value in the set of enumeration types in the metadata structure, and replace the identification information record of the newly added video-audio file, that is:

[0123]

[0124] Among them, is the replacement result of the th information identification record that fails the review. is the information identification record in the set of enumeration types in the metadata structure that has the largest similarity with . is the similarity measurement value between and the information identification records in the set of enumeration types. Select the information identification record with the largest similarity measurement value. Optimize and correct the potential non-standard and inaccurate enumeration type information existing in the metadata structure, ensure the rationality and consistency of the data, further improve the overall quality of the metadata structure, enable it to more accurately reflect the actual situation of the video-audio file, enhance the effectiveness and reliability of the metadata in supporting application scenarios such as precise retrieval and in-depth analysis, and ensure the continuous and efficient operation of the entire video-audio file metadata management system and the full play of the data value.

[0125] The security protection module is used to encrypt the video and audio files in the secure storage area, generate an initial encryption key based on the positioning index of the metadata structure of the video and audio files, and determine the encryption level according to the data volume and format of the video and audio files, and encrypt the video and audio files at different levels;

[0126] Preferably, the security protection module includes a key generation unit and an encryption execution unit;

[0127] The key generation unit is configured with a feature fusion strategy. The feature fusion strategy is used to generate an encryption key with dynamic update characteristics for the video and audio files. By extracting feature data from the positioning index of the metadata structure of the video and audio files, an original feature set is formed, and after hashing the original feature set and combining it with a dynamic salt value, key material is generated, and the initial encryption key is generated with the key material as the input;

[0128] The encryption execution unit is configured with an adaptive encryption strategy. The adaptive encryption strategy is used to encrypt the video and audio files at different levels according to the initial encryption key. By evaluating the data volume and format of the video and audio files, the encryption level is determined, and the video and audio files are adaptively encrypted according to the encryption level.

[0129] Preferably, the specific steps of the feature fusion strategy include:

[0130] Extract multiple features from the positioning index of the metadata structure of the video and audio files, including the specific character segment of the device identification code of the camera, the hash value of the time information when the video and audio file was taken, and the target object category code, to form an original feature set;

[0131] For the device identification code of the camera, select a representative specific character segment. Exemplarily, select a continuous segment of characters in the middle, so as to inject unique information closely related to the device into the key;

[0132] For the time when the video and audio file was taken, use the hash algorithm for processing to obtain the time hash value of the video and audio file when it was taken;

[0133] For the target object category, convert it according to the predefined coding table, and assign specific and unique digital codes to various possible target objects, so that the information of the target objects can be incorporated into the original feature set in a standardized and easy-to-process form;

[0134] Perform a hash operation on the original feature set to generate an intermediate hash value with a fixed length, and perform exclusive OR, bit expansion, and circular shift in combination with the dynamic salt value to generate key material;

[0135] Set the security strength parameter, and use the HKDF key derivation function. With the key material as the input, according to the security strength parameter, generate the initial encryption key.

[0136] Preferably, the specific steps of the adaptive encryption strategy include:

[0137] Evaluate the features of the audio-visual file, locate the index through the audio-visual file and its metadata structure, and obtain the key attributes of the audio-visual file, including the data volume of the audio-visual file and the audio-visual file format;

[0138] According to the data volume and format of the audio-visual file, construct an encryption level evaluation model. The encryption level evaluation model classifies the audio-visual file into different encryption levels according to the data volume of the audio-visual file and the audio-visual file format, including: low level, medium level, and high level; Exemplarily, an audio-visual file with a large data volume and a complex audio-visual file format is determined to be at the high encryption level;

[0139] According to the determined encryption level, select an encryption algorithm. For the high encryption level, use an encryption algorithm with a longer key length, such as: AES-256; For the low encryption level, use an encryption algorithm with a shorter key length, such as: AES-128;

[0140] Call the initial encryption keys of different lengths generated by the key generation unit, use the selected encryption algorithm, and the generated encryption key to encrypt the audio-visual file. During the encryption process, adopt different encryption modes such as block encryption and stream encryption to adapt to audio-visual files of different sizes and formats;

[0141] After encryption, verify the encrypted audio-visual file to ensure that the file content is not damaged during the encryption process, and verify the effectiveness of the encryption key and algorithm through decryption tests;

[0142] Store the encrypted file and key information, store the encrypted audio-visual file in the secure storage area, and associate the encrypted key and its related information, including the encryption level and algorithm type, with the metadata structure location index of the audio-visual file.

[0143] Embodiment 2

[0144] Please refer to Figure 7 , this embodiment introduces a method for managing the security of audio-visual files, including the following steps:

[0145] Step S1: When the camera captures an audio-visual file, store the audio-visual file in the buffer area, adjust the buffer capacity by monitoring the network bandwidth and the change in the number of bytes of the audio-visual file, and encode and embed the source identification information of the audio-visual file in the buffer area and then store it in the secure storage area;

[0146] Step S2: By monitoring the key marking operations when the user plays or browses the audio-visual files in the secure storage area, extracting the audio-visual segments of the key intervals according to the key marking operations, screening the key frames by using the amount of motion change and performing target recognition, extracting the audio data and classifying them to obtain the key marking information, then defining a directed graph by constructing a metadata structure template based on the identification information, processing the newly added identification information records through a comparison function and reviewing the enumerated type records to generate a multi-dimensional metadata structure positioning index;

[0147] Step S3: Encrypt the audio-visual files in the secure storage area, generate an initial encryption key based on the metadata structure positioning index of the audio-visual files, and determine the encryption level according to the data volume and format of the audio-visual files, and encrypt the audio-visual files at different levels.

[0148] Preferably, the specific steps for adjusting the buffer capacity include:

[0149] Set the initial buffer capacity, and obtain the camera parameter information, including resolution, frame rate, encoding format, sampling rate and number of channels;

[0150] Configure a time window , and count the number of bytes of the audio-visual files received within the time window;

[0151] Configure a sliding window , which is used to track the long-term trend and short-term fluctuations of the data traffic. Add the number of bytes of the audio-visual files counted within the time window to the sliding window data sequence, and remove the number of bytes of the audio-visual files counted in the earliest time window, so as to keep the number of data points within the sliding window fixed at ;

[0152] Configure an evaluation interval. Every time an evaluation interval passes, calculate the buffer occupancy rate through the number of bytes already occupied in the buffer and the buffer capacity , and calculate the remaining data duration that the buffer can continue to accommodate ;

[0153] Evaluate the status of the buffer. Based on the buffer occupancy rate, the remaining data duration that can be accommodated, and the network bandwidth, construct a buffer status evaluation model and calculate the buffer status evaluation score, that is:

[0154]

[0155] wherein, is the buffer status evaluation score, is the current network bandwidth, is the set standard network bandwidth, and are non-negative weighting coefficients respectively;

[0156] Configure an adjustment threshold. If the buffer status evaluation score is less than the adjustment threshold, trigger an operation to increase the buffer capacity; otherwise, do nothing.

[0157] Calculate the growth rate of the number of bytes within the sliding window, that is:

[0158]

[0159] where is the growth rate of the number of bytes in the th time window within the sliding window, is the number of bytes of the th audio-visual file in the sliding window data sequence, is the number of bytes of the th audio-visual file in the sliding window data sequence;

[0160] Dynamically increase the buffer capacity according to the growth rate of the number of bytes within the sliding window and the buffer occupancy rate, that is:

[0161]

[0162] where is the adjusted buffer capacity, is the average value of the growth rate of the number of bytes within the sliding window, and are respectively the minimum threshold and the maximum threshold of the growth rate of the number of bytes within the sliding window, is the buffer occupancy rate threshold, , and are respectively the weighting coefficients.

[0163] Preferably, the specific steps for generating the metadata structure location index include:

[0164] Define a metadata structure template according to the identification information, and define the metadata structure as a directed graph, where the vertex set is , is the th identification information, is the number of identification information types; define the attribute types for each vertex of the metadata structure, and the attribute types include enumeration type, string type, and array type;

[0165] Obtain the identification information record of the audio-visual file , where is the number of obtained identification information, is the An identification information record; for each obtained identification information record, use a comparison function to compare the identification information type with the vertex of the metadata structure, that is:

[0166]

[0167] Among them, is the th identification information record and the vertex of the metadata structure comparison function, is the identification information type of the th identification information record;

[0168] If the value of the comparison function is , it means that the identification information type comparison is successful, then is filled into the corresponding position of the vertex of the metadata structure . If the value of the comparison function is , the identification information type comparison fails, and a new temporary metadata structure vertex is created to store the identification information record with the comparison failure;

[0169] When adding an identification information record of an audio - video file to the metadata structure , select the identification information record whose attribute is of the enumeration type for information review. Among them, is the th identification information record whose attribute is of the enumeration type. Among them, is the number of identification information records whose attribute is of the enumeration type, is less than , and the formula for information review is:

[0170]

[0171] Among them, is the review result of the th identification information record whose attribute is of the enumeration type, is the set of enumeration types in the metadata structure, is the identification information type of the th identification information record in the set of enumeration types; if takes the value of 1, it means that there are duplicate records and the information review passes. If takes the value of 0, it means that there are no duplicate records and the information review fails;

[0172] For the information identification record with the information review failed , among them, is the th information identification record that fails, Less than or equal to , filter and replace the similar information identification records in the metadata structure, that is:

[0173]

[0174] Among them, is the replacement result of the th unpassed information identification record, is the information identification record with the greatest similarity to in the enumeration type set in the metadata structure, is and the similarity measurement value of the information identification record in the enumeration type set, is to select the information identification record with the largest similarity measurement value.

[0175] Working principle and its effect:

[0176] The working principle and effect of a video and audio file security management system and method are as follows:

[0177] After the camera captures a video and audio file, the encoding processing module first stores it in the buffer. During this process, the data buffer unit relies on the adaptive buffering strategy to continuously monitor the changes in network bandwidth and file byte count based on the transmission rate and data volume of the video and audio file transmitted by the camera. It calculates relevant data through setting time windows and sliding windows, combines the evaluation interval to calculate the buffer occupancy rate, the remaining data duration that can be accommodated, etc., and then constructs a buffer state evaluation model to calculate the evaluation score. It dynamically adjusts the buffer capacity according to the adjustment threshold to avoid data transmission problems. At the same time, the information embedding unit uses the deep fusion strategy to embed the source identification information into the file data structure and then stores it in the secure storage area, achieving the effects of ensuring stable transmission and facilitating source tracing.

[0178] Then the key marking module comes into play. When the user plays or browses the video and audio file in the secure storage area, the marking acquisition unit monitors the key marking operation through the marking capture strategy, records the time stamp, configures the extraction interval to obtain the key interval video and audio segments, converts them into image frames, filters the key frames according to the motion change amount, and then performs operations such as target recognition and audio data extraction and classification to obtain the key marking information. Then, the association and fusion strategy of the index construction unit will define a metadata structure template based on this information and the source identification information to construct a directed graph, and generate a multi-dimensional metadata structure positioning index during the process of processing new identification information records and reviewing enumeration type records. This helps to accurately extract the key information that users are concerned about and achieve the effective integration and association of multi-dimensional information, improving the file management and retrieval efficiency.

[0179] Finally, the last security protection module operates. The key generation unit uses a feature fusion strategy to extract feature data from the metadata structure positioning index of the audio-visual file to form an original feature set, generates key materials through a hashing operation combined with a dynamic salt value, and then generates an initial encryption key. The encryption execution unit, according to the adaptive encryption strategy, determines the encryption level with reference to the data volume and format of the audio-visual file and then performs hierarchical encryption. In this way, using the dynamically updated encryption key and the adapted hierarchical encryption method enhances the security of the audio-visual file in storage and other links, preventing data from being illegally obtained and tampered with. The overall system and method work together in multiple aspects such as data transmission, user interaction information integration, and data security guarantee, effectively manage the entire life cycle of the audio-visual file, and improve its usage, management, and security protection levels.

[0180] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A video and audio file security management system, characterized in that: Including coding processing module, key marking module and security protection module: The encoding processing module is used to store the video and audio files in the buffer when the camera captures the video and audio files, adjust the buffer capacity by monitoring the network bandwidth and the number of bytes of the video and audio files, and encode and embed the source identification information of the video and audio files in the buffer and store them in the secure storage area; The key mark module is used to monitor the key mark operation of the user when playing or browsing the audio and video files in the secure storage area, extract the audio and video segments in the key interval according to the key mark operation, use the motion change to screen the key frames and perform target recognition, extract audio data and classification to obtain the key mark information, and then define the metadata structure template according to the identification information to build a directed graph, and generate a multi-dimensional metadata structure positioning index by processing the newly added identification information records and reviewing the enumeration type records through the comparison function; The security protection module is used to encrypt the audio and video files in the secure storage area, generate an initial encryption key based on the metadata structure positioning index of the audio and video files, and determine the encryption level according to the data volume and format of the audio and video files, and encrypt the audio and video files in a hierarchical manner; The key point marking module includes a marking collection unit and an index building unit; The tag collection unit is configured with a tag capture strategy, which is used to monitor the key tag operation made by the user during the playback or browsing of the audio-visual file, select key frames through motion variation calculation, generate content summaries for key frame target recognition, extract and classify audio data, so as to identify and extract key tag information; The index building unit is configured with an associated fusion strategy, which is used to index the key mark information and source identification information of the audio and video files, generate a multi-dimensional metadata structure positioning index according to a preset metadata structure template, add metadata structure identification information records through a comparison function, and perform information review on the identification information records of the enumeration type; The steps of associating the fusion strategy include: Define a metadata structure template according to the identification information, define the metadata structure as a directed graph, and define an attribute type for each vertex of the metadata structure, the attribute type includes an enumeration type, a string type, and an array type; Obtain identification information records of the video and audio files, and for each obtained identification information record, compare the identification information type with the metadata structure vertex through a comparison function, construct a comparison function between the identification information record and the metadata structure vertex, and query whether there is an identification information type that is the same as the identification information record in the metadata structure vertex; If there is an identification information type that is the same as the identification information record, it means that the identification information type comparison is successful, and the identification information record is filled in the position corresponding to the metadata structure vertex of the same identification information type. If there is no identification information type that is the same as the identification information record, the identification information type comparison fails, and a new temporary metadata structure vertex is created to store the identification information record that failed the comparison.

2. A video and audio file security management system as claimed in claim 1, characterized in that: The encoding processing module includes a data buffer unit and an information embedding unit; The data buffer unit is configured with an adaptive buffer strategy, which is used to adjust the buffer capacity in real time according to the transmission rate and data volume of the video and audio files transmitted by the camera, and to build a buffer state evaluation model by continuously monitoring the network bandwidth and the number of bytes of the video and audio files during the transmission of the video and audio files, and to evaluate the buffer state so as to dynamically allocate memory space as a buffer; The information embedding unit is configured with a deep fusion strategy, which is used to embed source identification information into the data structure of the video and audio file, and store the processed video and audio file in a secure storage area.

3. A video and audio file security management system as claimed in claim 2, characterized in that: The specific steps of the adaptive buffering strategy include: Set the initial buffer capacity and obtain the camera parameter information; Configure the time window and count the bytes of the audio and video files received within the time window; Configure a sliding window to track the long-term trend and short-term fluctuation of data traffic. Add the number of bytes of audio and video files counted in the time window to the sliding window data sequence, and remove the number of bytes of audio and video files counted in the earliest time window to keep the number of data points in the sliding window fixed. Configure the evaluation interval. At each evaluation interval, calculate the buffer occupancy rate based on the number of occupied buffer bytes and the buffer capacity, and calculate the length of time the buffer can continue to accommodate data. Evaluate the status of the buffer, build a buffer status evaluation model based on the buffer occupancy rate, the duration of data that can be continuously accommodated, and the network bandwidth, and calculate the buffer status evaluation score; Configure the adjustment threshold. If the buffer status evaluation score is less than the adjustment threshold, the buffer capacity increase operation is triggered; otherwise, no processing is performed.

4. A video and audio file security management system as claimed in claim 3, characterized in that: The buffer capacity increasing operation includes: Based on the number of bytes of the audio and video files in each sliding window data sequence, calculate the growth rate of the number of bytes in the sliding window; According to the growth rate of the number of bytes in the sliding window and the buffer occupancy rate, the buffer capacity is dynamically increased, that is: in, is the adjusted buffer capacity, is the buffer capacity, is the buffer occupancy, is the average growth rate of the number of bytes in the sliding window, and are the minimum and maximum thresholds for the growth rate of the number of bytes in the sliding window, is the buffer occupancy threshold, , and are all weighting coefficients.

5. The video and audio file security management system according to claim 1, characterized in that: The steps of the marking capture strategy include: When the user starts playing or browsing the video and audio files in the buffer, the focus marking operation is monitored and the focus marking type is identified; By monitoring the user's interactive behavior, the marking time point of the video and audio files where the key marks are located is recorded, and a timestamp is generated for each key mark; According to the key mark type, on the basis of each key mark timestamp, a mark extraction interval is configured, and according to the mark extraction interval, the video and audio clips of the key interval are obtained; Convert the audio and video clips in the key interval into a series of image frames, mark the first image frame as the key frame, configure the vector change threshold, and calculate the motion change of the image frame, that is: in, It is Image frames and Image frames The amount of movement change, is the image frame The middle position is The motion vector of is the image frame The middle position is The motion vector of If the amount of motion changes If it is greater than the vector change threshold, Marked as a keyframe, As the starting point, continue to calculate the adjacent motion changes Otherwise, skip , calculate the change in motion , until the traversal of the image frames of the video and audio clip is completed; Perform target recognition on key frames to obtain key mark information of key frames, including text information, target objects, and target scenes, and convert the key mark information recognition results of key frames into text descriptions as content summaries; Extract audio data from audio and video clips in key intervals, identify key tag information of the audio data, classify the audio data, and identify the emotional color of the audio data.

6. The video and audio file security management system according to claim 1, characterized in that: The step of associating the fusion strategy also includes: When the identification information record of the video and audio files is added to the metadata structure , select the identification information record whose attribute is the enumeration type for information review, so as to query whether the identification information record of the enumeration type exists in the enumeration type set in the metadata structure; If the identification information type of the existing identification information record is found to be the same as the information type of the identification information record of the newly added audio-visual file in the enumerated type set in the metadata structure, the information review is passed, otherwise the information review is not passed; For information identification records that fail the information review, similar information identification records in the metadata structure are screened and replaced; Calculate the similarity measure between the information identification record that fails the information review and the information identification record of the same information type in the enumeration type set, select the information identification record with the largest similarity measure in the enumeration type set in the metadata structure, and replace the identification information record of the newly added audio and video file.

7. The video and audio file security management system according to claim 1, characterized in that: The security protection module includes a key generation unit and an encryption execution unit; The key generation unit is provided with a feature fusion strategy, which is used to generate an encryption key with dynamic update characteristics for the video and audio files, by extracting feature data from the metadata structure positioning index of the video and audio files to form an original feature set, and generating key materials after hashing the original feature set, and using the key materials as input to generate an initial encryption key; The encryption execution unit is configured with an adaptive encryption strategy, which is used to perform hierarchical encryption on audio and video files according to an initial encryption key, determine the encryption level by evaluating the data volume and format of the audio and video files, and adaptively encrypt the audio and video files according to the encryption level.

8. A method for managing video and audio files safely, which is implemented based on a video and audio file safety management system according to any one of claims 1 to 7, characterized in that: The following steps are involved: Step S1: When the camera captures an audio and video file, the audio and video file is stored in a buffer, the buffer capacity is adjusted by monitoring the network bandwidth and the number of bytes of the audio and video file, and the source identification information of the audio and video file in the buffer is encoded and embedded and then stored in a secure storage area; Step S2: by monitoring the key marking operation of the user when playing or browsing the audio and video files in the secure storage area, extracting the audio and video segments of the key interval according to the key marking operation, using the motion variation to screen the key frames and perform target recognition, extracting audio data and classification, so as to obtain the key marking information, and then defining the metadata structure template according to the identification information to construct a directed graph, and processing the newly added identification information records and reviewing the enumeration type records through the comparison function to generate a multi-dimensional metadata structure positioning index; Step S3: Encrypt the audio and video files in the secure storage area, generate an initial encryption key based on the metadata structure positioning index of the audio and video files, determine the encryption level according to the data volume and format of the audio and video files, and encrypt the audio and video files in a hierarchical manner.

Citation Information

Patent Citations

  • A method, apparatus and electronic device for protecting documents

    CN107977551B

  • Cryptograph index structure based on blocking organization and management method thereof

    CN101655858A

  • Multi-dimensional metadata management method and system based on association characteristics

    CN103218404A