Video processing method and device, intelligent equipment and storage medium
By identifying scene types and dynamically determining the algorithm embedding timing, the target algorithm data is embedded into the preset position of the video frame, solving the mismatch problem of video processing technology in different scenarios, realizing the high efficiency of video stream processing and data synchronization, and improving the reliability of video analysis.
Patent Information
- Application Number
- CN202511130633.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-28
AI Technical Summary
Existing video processing technologies lack flexibility in handling different scenarios, and the synchronization between the data after video stream processing and analysis and video frames is insufficient, resulting in reduced data validity and affecting the reliability of video analysis.
By identifying the current scene type, the target algorithm is determined and executed on each original video frame of the real-time video stream to generate target algorithm data. This data is then embedded into a preset location for storage, generating the target video stream. The embedding timing is dynamically determined based on the algorithm type and scene type to ensure a close correlation between the algorithm data and the video frames.
It improves the flexibility and effectiveness of video processing, ensures accurate correspondence between algorithm data and video frames, avoids desynchronization issues, and enhances the processing quality and data validity of real-time video streams.
Smart Images

Figure CN121037641A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent terminals, and in particular to a video processing method and device, an intelligent device, and a storage medium. BACKGROUND
[0002] With the wide application of artificial intelligence technology in intelligent driving, security monitoring and other fields, the demand for video processing is increasing. In these application scenarios, real-time video stream processing and analysis are crucial to improving the intelligent level of the system. The data after video stream processing and analysis not only provides support for system decision-making, but also is used for important links such as accident analysis, system optimization, responsibility determination, and evidence collection. For example, in the vehicle driving scenario, the intelligent driving function of the vehicle depends on the processing and analysis of real-time video streams, and the data after processing and analysis plays a key role in safety decision-making and accident tracing.
[0003] However, the existing video processing technology lacks flexibility in processing different scenarios, and the synchronization of the data after video stream processing and analysis with the video frames is insufficient, resulting in a decrease in the effectiveness of the data, thereby seriously affecting the reliability of video analysis.
[0004] Therefore, how to flexibly process video, improve the processing quality of real-time video streams, and ensure the effectiveness of video processing data is a problem to be solved at present. SUMMARY
[0005] The embodiments of the present application provide a video processing method, device, intelligent device, and storage medium, which can flexibly process video, improve the processing quality of real-time video streams, and ensure the effectiveness of video processing data.
[0006] In a first aspect, the embodiments of the present application provide a video processing method, comprising:
[0007] identifying a current scene type and determining a target algorithm corresponding to the scene type;
[0008] executing the target algorithm on each original video frame of a real-time video stream to generate target algorithm data;
[0009] embedding the target algorithm data into a preset position corresponding to the original video frame for storage to generate a target video stream containing the target algorithm data.
[0010] In a possible implementation manner of the first aspect, the step of embedding the target algorithm data into a preset position corresponding to the original video frame for storage to generate a target video stream containing the target algorithm data comprises:
[0011] determine the embedding timing of the target algorithm data according to the type of the target algorithm and / or the type of the scene, the embedding timing comprising performing embedding before encoding of the original video frames or performing embedding after encoding of the original video frames;
[0012] embed the target algorithm data into the preset storage location based on the determined embedding timing, to generate a target video stream containing the target algorithm data.
[0013] In a possible implementation of the first aspect, the step of determining the embedding timing of the target algorithm data according to the type of the target algorithm comprises:
[0014] obtaining a preset time consumption threshold corresponding to the type of the target algorithm;
[0015] determining an end-to-end processing duration from capturing the original video frames to generating the target algorithm data;
[0016] determining whether the end-to-end processing duration exceeds the preset time consumption threshold;
[0017] if the end-to-end processing duration does not exceed the preset time consumption threshold, determining to embed the target algorithm data before encoding of the original video frames;
[0018] if the end-to-end processing duration exceeds the preset time consumption threshold, determining to embed the target algorithm data after encoding of the original video frames.
[0019] In a possible implementation of the first aspect, the step of determining the embedding timing of the target algorithm data according to the type of the scene comprises:
[0020] determining the embedding timing of the target algorithm data based on a mapping relationship between scene types and embedding timings in a preset scene type timing table.
[0021] In a possible implementation of the first aspect, the step of determining the embedding timing of the target algorithm data according to the type of the target algorithm and the type of the scene comprises:
[0022] obtaining a preset time consumption threshold corresponding to the type of the target algorithm, and a priority corresponding to the type of the scene;
[0023] determining an end-to-end processing duration from capturing the original video frames to generating the target algorithm data;
[0024] determining whether the end-to-end processing duration exceeds the preset time consumption threshold;
[0025] If the end-to-end processing time does not exceed the preset time consumption threshold and the priority of the scene type is higher than the preset priority threshold, then it is determined that the target algorithm data is embedded before the encoding of the original video frame.
[0026] If the end-to-end processing time exceeds the preset time consumption threshold or the priority of the scene type is lower than or equal to the preset priority threshold, then the target algorithm data is determined to be embedded after the original video frame is encoded.
[0027] In one possible implementation of the first aspect, the preset location is a private data area reserved inside the original video frame, or the preset location is the metadata area of the encoded video frame after the original video frame is encoded.
[0028] In one possible implementation of the first aspect, the video processing method further includes:
[0029] During video playback, the target parser is invoked according to the scene type or the target algorithm, and the target parser extracts and parses the target algorithm data from the target video stream.
[0030] Secondly, embodiments of this application provide a video processing apparatus, including:
[0031] The identification and determination unit is used to identify the current scene type and determine the target algorithm corresponding to the scene type;
[0032] The data generation unit is used to execute the target algorithm on each raw video frame of the real-time video stream and generate target algorithm data.
[0033] The video processing unit is used to embed the target algorithm data into a preset location corresponding to the original video frame for storage, thereby generating a target video stream containing the target algorithm data.
[0034] Thirdly, embodiments of this application provide a smart device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video processing method as described in the first aspect above.
[0035] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the video processing method described in the first aspect above.
[0036] Fifthly, embodiments of this application provide a computer program product that, when run on a smart device, causes the smart device to execute the video processing method described in the first aspect above.
[0037] In this embodiment, the target algorithm is determined by identifying the current scene type, enabling video processing to flexibly select the appropriate algorithm based on different scenes. This avoids the mismatch problem that may occur when using a fixed algorithm to process different scenes, laying the foundation for improving the processing quality of real-time video streams. By executing the target algorithm on each original video frame of the real-time video stream, targeted target algorithm data can be generated for the actual situation of each frame, ensuring the accurate correspondence between the algorithm data and the video frame content, effectively improving the effectiveness of video processing. The target algorithm data is then embedded into the preset location corresponding to the original video frame for storage, forming a close association between the algorithm data and the original video frame, avoiding the problem of asynchronous algorithm data and video frames, and further ensuring the effectiveness of the algorithm data in subsequent use. Thus, the overall flexible processing of real-time video streams is realized, improving video processing quality and ensuring the effectiveness of algorithm data. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart illustrating the implementation of the video processing method provided in the embodiments of this application;
[0040] Figure 2 This is a flowchart illustrating a specific implementation of step S103 in the video processing method provided in this application embodiment;
[0041] Figure 3 This is a flowchart illustrating a specific implementation of the video processing method for determining the embedding timing provided in this application embodiment;
[0042] Figure 4 This is another specific implementation flowchart of determining the embedding timing in the video processing method provided in the embodiments of this application;
[0043] Figure 5 This is a structural block diagram of the video processing apparatus provided in the embodiments of this application;
[0044] Figure 6 This is a schematic diagram of the smart device provided in the embodiments of this application. Detailed Implementation
[0045] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0046] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0047] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0048] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0049] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0050] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0051] By way of example and not limitation, the video processing method provided in this application can be applied to various types of smart devices or servers that require video processing, specifically including in-vehicle terminals, mobile phones, tablets, laptops, wearable devices, ultra-mobile personal computers (UMPCs), desktop computers, etc. The application embodiments do not impose any restrictions on the specific types of smart devices and servers.
[0052] Figure 1 The implementation flow of the video processing method provided in this application embodiment is illustrated. The method flow includes steps S101 to S103. The specific implementation principle of each step is as follows:
[0053] Step S101: Identify the current scene type and determine the target algorithm corresponding to the scene type.
[0054] Scene type refers to the specific application environment in which the video is located. Scene types include, but are not limited to, urban traffic jam following scenes, highway cruising scenes, parking lot parking scenes, driver monitoring scenes, indoor environments, outdoor environments, and nighttime scenes.
[0055] Target algorithms refer to video analysis algorithms that match scene types. Target algorithms include, but are not limited to, environmental perception algorithms (such as lane line recognition, object detection, position analysis, distance detection, blind spot detection, traffic sign recognition, etc.), image enhancement algorithms (such as low-light enhancement, rain and fog removal algorithms, etc.), behavior prediction algorithms (such as trajectory prediction, collision prediction, object tracking algorithms, driver status monitoring, gaze analysis, behavior analysis, etc.), and human recognition algorithms (such as pedestrian detection, facial recognition, etc.). One or more target algorithms may be matched within the same scene type.
[0056] In one possible implementation, environmental data is collected through multi-source sensors (such as vehicle-mounted cameras, surveillance cameras, and radar). In intelligent driving, vehicle driving information is also collected. The scene type is determined based on the collected environmental data and preset classification rules, or the scene type is determined based on the collected environmental characteristics, vehicle driving characteristics, and preset classification rules.
[0057] For example, when the vehicle speed is greater than 80 km / h and the lane lines are clear, it is identified as a highway cruising scenario; when the vehicle speed is less than 30 km / h and the distance between vehicles is less than 5 meters, it is identified as an urban traffic jam following scenario; when the vehicle speed is less than 20 km / h and the reversing signal or turn signal is activated, it is identified as a parking lot scenario; when the vehicle speed is greater than 80 km / h and the continuous driving time exceeds 1 hour, it is identified as a driver monitoring scenario; and when the light intensity is less than 50 lux, it is identified as a nighttime scenario.
[0058] In one possible implementation, scene type identification can be achieved by using a pre-trained deep learning model to perform feature analysis on raw, real-time captured video frames, extracting key features, and identifying the scene type. This embodiment does not limit the type of deep learning model. For example, a pre-trained CNN model can be used to analyze the captured video frames and identify the current scene type as an urban traffic congestion scenario.
[0059] After identifying the scene type, a target algorithm is selected based on a preset mapping relationship (the matching mapping relationship between scene type and target algorithm). For example, in urban congestion following scenarios, target algorithms such as pedestrian detection, blind spot detection, object tracking, and distance detection are matched; in highway cruising scenarios, target algorithms such as object detection, lane recognition, position analysis, and trajectory prediction are matched; in parking lot scenarios, target algorithms such as collision prediction and object detection are matched; in driver monitoring scenarios, target algorithms such as facial recognition, driver status monitoring, gaze analysis, and behavior analysis are matched; in indoor security monitoring scenarios, facial recognition algorithms are matched; and in nighttime scenarios, low-light enhancement algorithms are matched.
[0060] In this embodiment of the application, by dynamically adapting the scene and the target algorithm, the mismatch problem that may occur when a fixed algorithm processes different scenes is avoided, which improves the flexibility and adaptability of video processing and lays the foundation for improving the processing quality of real-time video streams.
[0061] Step S102: Execute the target algorithm on each original video frame of the real-time video stream to generate target algorithm data.
[0062] Raw video frames are video frames output by an image sensor that have not undergone encoding or other subsequent processing. For example, raw video frames are raw frame data output by a CMOS sensor or YUV / RGB frames that have undergone ISP (Image Signal Processing) preprocessing.
[0063] The original video frame is calculated and analyzed using a target algorithm determined based on the current scene type to obtain corresponding target algorithm data. This data is generated based on the image content of the original video frame and includes the execution results of the target algorithm, specifically real-time processing information such as object detection results, lane recognition results, behavior analysis results, and face recognition results. In this embodiment, the target algorithm is executed strictly frame-by-frame, and the generated target algorithm data shares the same timestamp as the original video frame (deviation < 1ms), ensuring a strict temporal correspondence between the algorithm data and the original video frame.
[0064] One possible implementation is that the execution results of the target algorithm are organized into structured data, meaning the generated target algorithm data is structured data. Lightweight, efficiently binary-serializable formats such as JSON, Protocol Buffers, Flat Buffers, and MessagePack can be used, but this embodiment does not impose any limitations.
[0065] In one possible implementation, the data structure of the target algorithm data has a strictly defined schema.
[0066] For example, the data structure of the target algorithm data includes at least a unique incrementing frame identifier frame_id (uint64) and a high-precision timestamp (double), which is synchronized with the acquisition time of the original video frame. Depending on the type of the target algorithm and its execution result, the target algorithm data also includes a list of detected targets object_list (repeated), where each target in the object_list contains class_id (uint32) / class_name (string), normalized bounding box coordinates bbox (float[4]), detection confidence (float), target tracking ID tracking_id (uint64), lane information lane_info, event flags (such as emergency braking, collision warning) event_flags, etc.
[0067] For example, in a traffic sign recognition algorithm, the target algorithm data includes the location and category of the traffic sign; in a behavior analysis algorithm, the target algorithm data includes the behavior type and the time period in which it occurred.
[0068] By executing the target algorithm on each frame of video, targeted target algorithm data can be generated based on the actual situation of each frame. Since the target algorithm matches the scene, the generated target algorithm data can accurately reflect the key information of the video frame and ensure the accurate correspondence between the algorithm data and the video frame content, effectively improving the targeting and effectiveness of video processing data.
[0069] Step S103: Embed the target algorithm data into a preset location corresponding to the original video frame and store it to generate a target video stream containing the target algorithm data.
[0070] In this embodiment, embedded storage refers to integrating the target algorithm data into a preset location to form a target video stream containing video frames and algorithm data. The preset location is a pre-defined storage area associated with the original video frames.
[0071] By embedding the target algorithm data into the preset position corresponding to the original video frame, the algorithm data and the original video frame are closely related, avoiding the problem of the algorithm data and the video frame being out of sync. Moreover, the embedded target algorithm data can remain effective in subsequent use, supporting video playback and analysis.
[0072] As one possible implementation of this application Figure 2 A specific implementation flow of step S103 in the video processing method provided in this application embodiment is shown below:
[0073] A1: Determine the embedding timing of the target algorithm data based on the type of the target algorithm and / or the type of the scene. The embedding timing includes performing embedding before encoding the original video frames or performing embedding after encoding the original video frames.
[0074] The type of the target algorithm refers to the specific category and computational complexity of the algorithm. For example, the vehicle distance calculation algorithm in the ADAS algorithm is a lightweight computation, while the collision prediction algorithm in the AEB algorithm may be a heavyweight computation. The scenario type refers to the highway cruise scenario, parking lot scenario, etc., as described in step S101. The embedding timing is the timing node when data is written to the storage medium (before the original video frame is encoded / after the original video frame is encoded).
[0075] A2: Based on the determined embedding timing, the target algorithm data is embedded into the preset location for storage, generating a target video stream containing the target algorithm data.
[0076] In this embodiment, the embedding timing of the target algorithm data is not fixed, but dynamically determined based on the type of the target algorithm and / or the type of the scene. For example, when the target algorithm is a computationally simple, lightweight algorithm (such as a basic object detection algorithm) and the scene is a parking lot scenario with high real-time requirements, it may be determined to embed it before encoding; when the algorithm is a computationally complex, heavyweight algorithm (such as a multi-object trajectory prediction algorithm), or the scene has lower real-time requirements, it may be determined to embed it after encoding. The embedding operation is closely bound to the determined timing, ensuring the orderliness of the embedding process.
[0077] In this embodiment, by refining the logic and execution method for determining the embedding timing, the embedding process of the target algorithm data is made more in line with the algorithm characteristics and scenario requirements in actual applications. If embedded before encoding, the target algorithm data can undergo encoding processing along with the original video frames, ensuring that the two are synchronized in subsequent transmission and storage. If embedded after encoding, the interference of the target algorithm data on the encoding process can be avoided, while not affecting the correspondence between the target algorithm data and the encoded video frames. This strengthens the correlation between the target algorithm data and the video frames, ensuring the integrity and effectiveness of the target video stream, and providing a more reliable foundation for subsequent playback and analysis.
[0078] As one possible implementation of this application Figure 3 This application illustrates a specific implementation of a video processing method that determines the embedding timing of the target algorithm data based on the type of the target algorithm, as detailed below:
[0079] B1: Obtain the preset time consumption threshold corresponding to the type of the target algorithm.
[0080] The preset time consumption threshold is a pre-set upper limit for processing time for different algorithm types. This threshold is determined based on the algorithm's historical processing data, computational complexity, computational requirements, and the real-time requirements of the actual application scenario. For example, the preset time consumption threshold for lightweight algorithms can be set to 50 milliseconds, while that for heavyweight algorithms can be set to 200 milliseconds. Different target algorithms can have different preset time consumption thresholds set according to their characteristics, improving the system's flexibility. The preset time consumption threshold can be stored in a configuration file or a database. The preset time consumption threshold provides a clear time reference standard for different types of algorithms, providing a basis for subsequent judgments on the reasonableness of algorithm processing time, ensuring the objectivity and consistency of the judgment.
[0081] B2: Determine the end-to-end processing time from acquiring the original video frames to generating the target algorithm data.
[0082] End-to-end processing time refers to the total time elapsed from the moment a raw video frame is captured and entered into the system until the target algorithm completes processing that frame and generates the target algorithm data. This includes the total time spent on data transmission, algorithm computation, and result generation. For example, in an autonomous driving scenario, the time taken from when a camera captures a raw video frame to when a distance calculation algorithm calculates the distance to the vehicle in front is the end-to-end processing time. End-to-end processing time precisely quantifies the actual time it takes for an algorithm to process a single frame of video.
[0083] In one possible implementation, the video processing system records the timestamps from the acquisition of raw video frames to the generation of target algorithm data. The end-to-end processing time is determined by calculating the difference between the two timestamps. For example, if the timestamp for acquiring the raw video frames is t1 and the timestamp for generating the target algorithm data is t2, then the end-to-end processing time is t2-t1. By measuring the end-to-end processing time, the system's performance can be evaluated, ensuring the efficiency of the processing.
[0084] B3: Determine whether the end-to-end processing time exceeds the preset time consumption threshold.
[0085] For example, if the preset time threshold is 10 milliseconds, and the calculated end-to-end processing time is 8 milliseconds, then the processing time does not exceed the threshold; if the calculated end-to-end processing time is 12 milliseconds, then the processing time exceeds the threshold. With clear judgment criteria, it is possible to quickly determine whether the processing efficiency of the current target algorithm meets expectations, providing a crucial basis for the subsequent selection of embedding timing.
[0086] B4: If the end-to-end processing time does not exceed the preset time consumption threshold, then the target algorithm data is embedded before the original video frame is encoded. Embedding the target algorithm before encoding the original video frame is suitable for processing high-efficiency algorithms; when the target algorithm processes quickly (within the preset time consumption threshold), embedding before encoding can ensure that the data is processed synchronously with the video frame, reducing subsequent association costs.
[0087] B5: If the end-to-end processing time exceeds the preset time consumption threshold, then the target algorithm data is determined to be embedded after the original video frame is encoded. Embedding after the original video frame is suitable for algorithms with long processing times. When the algorithm is slow (exceeding the preset time consumption threshold), embedding after encoding can avoid blocking the encoding process due to excessive algorithm processing time, ensuring the overall efficiency of video processing.
[0088] The embodiments of this application reasonably select the embedding timing based on the type of the target algorithm, thereby improving the flexibility and adaptability of the system.
[0089] As one possible implementation of this application, the embedding time sequence of the target algorithm data is determined based on the mapping relationship between scene type and embedding time sequence in a preset scene type time sequence lookup table.
[0090] The preset scenario type timing lookup table is a pre-built structured data table that stores the correspondence between scenario types and embedding timing sequences. The table covers various specific application environments, and each scenario type corresponds to a unique embedding timing sequence. For example, highway cruising scenarios correspond to post-encoding embedding, urban congestion following scenarios correspond to post-encoding embedding, parking lot scenarios correspond to pre-encoding embedding, and driver monitoring scenarios correspond to pre-encoding embedding, etc. These correspondences are set based on the scenario's requirements for real-time performance and data density. For instance, parking lot scenarios require high-precision real-time obstacle detection, necessitating pre-encoding embedding to ensure data immediacy. Because the mapping relationship is preset based on scenario requirements, it ensures that the embedding timing sequence matches the scenario's real-time performance and data accuracy requirements. By establishing a direct association between scenario types and embedding timing sequences through the preset lookup table, the determination of the embedding timing sequence does not require complex calculations; only a query and matching are needed. The system can flexibly select the embedding timing sequence according to different scenario types, improving the system's adaptability.
[0091] As one possible implementation of this application Figure 4This paper illustrates another specific implementation of the video processing method provided in this application, which determines the embedding timing of the target algorithm data based on the type of the target algorithm and the scene type. Details are as follows:
[0092] C1: Obtain the preset time consumption threshold corresponding to the type of the target algorithm, and the priority corresponding to the scene type.
[0093] The priority of each scenario type is determined by its importance and real-time requirements. For example, parking lot scenarios involve vehicle safety and might have the highest priority (e.g., level 5), while ordinary city road scenarios might have a medium priority (e.g., level 3). Obtaining the preset time consumption threshold follows the steps outlined above and will not be repeated here.
[0094] C2: Determine the end-to-end processing time from acquiring the original video frames to generating the target algorithm data.
[0095] C3: Determine whether the end-to-end processing time exceeds the preset time consumption threshold.
[0096] The specific details of steps C2 and C3 are as described above and will not be repeated here.
[0097] C4: If the end-to-end processing time does not exceed the preset time consumption threshold and the priority of the scene type is higher than the preset priority threshold, then the target algorithm data is determined to be embedded before the encoding of the original video frame. The preset priority threshold is a critical value used to classify the importance of scenes, for example, set to level 3. Scenes with a priority higher than this value are considered high priority.
[0098] To embed target algorithm data before encoding the original video frame, two conditions must be met simultaneously: "end-to-end processing time exceeds the preset time consumption threshold" and "scene type priority is higher than the preset priority threshold." In other words, both efficient algorithm processing and high scene priority must be satisfied before pre-encoding embedding is chosen. For high-priority scenes with fast algorithm processing, pre-encoding embedding ensures data synchronization with the video frame, meeting the scene's high requirements for real-time performance and accuracy. For example, in a parking lot scenario, fast-processing collision prediction algorithm data embedded before encoding can provide timely support for parking decisions.
[0099] C5: If the end-to-end processing time exceeds the preset time consumption threshold or the priority of the scene type is lower than or equal to the preset priority threshold, then the target algorithm data is determined to be embedded after the original video frame is encoded.
[0100] Embedding target algorithm data after encoding the original video frames only requires meeting one of the following conditions: "end-to-end processing time exceeds the preset time consumption threshold" or "scene type priority is lower than or equal to the preset priority threshold". In other words, if either condition is met (algorithm processing timeout or scene priority is low), embedding after encoding is chosen. This avoids the impact of excessively long algorithm processing time on high-priority scenes, or the avoidance of using pre-encoding resources in low-priority scenes, thus balancing processing efficiency and resource allocation. For example, in urban road scenes, if algorithm processing timeout occurs, embedding after encoding can avoid delaying video encoding while meeting the general real-time requirements of this scene.
[0101] In this embodiment, reference standards of both algorithm and scenario are introduced to determine the embedding time of the target algorithm data, taking into account both technical performance and scenario requirements, making the embedding decision more comprehensive and reasonable, providing a more comprehensive basis for subsequent judgment, and avoiding decision bias caused by considering a single factor.
[0102] In one possible implementation, the target algorithm data is compressed; the compressed target algorithm data is digitally signed; and the compressed and digitally signed target algorithm data is embedded into a preset position corresponding to the original video frame and stored.
[0103] Compression processing refers to using specific compression algorithms (such as lossless compression algorithms like DEFLATE and LZ77, or dedicated compression algorithms selected based on data characteristics) to reduce the storage space and transmission bandwidth of target algorithm data without losing critical information. Target algorithm data typically contains a large amount of structured information, such as object coordinates, velocity, and angles. This data exhibits redundancy during storage and transmission; compression processing can remove this redundancy. For example, in intelligent driving scenarios, ADAS algorithms generate object tracking IDs across multiple consecutive frames with minimal ID changes between adjacent frames. Compression algorithms can reduce the data volume by recording the differences rather than the complete IDs. By compressing the target algorithm data generated in real-time for each frame of video, the data volume is reduced, ensuring that the compressed data can be efficiently stored in video files without affecting the quality of the video content.
[0104] Digital signature processing refers to using asymmetric encryption techniques (such as RSA and ECC algorithms) to encrypt compressed target algorithm data using a private key, generating a unique digital signature. This signature is bound to the compressed data, and the recipient can verify its validity using the sender's public key to confirm that the data has not been tampered with and originated from a legitimate sender. For example, in intelligent driving systems, digitally signing critical algorithm data such as collision probability and braking intensity allows for verification during subsequent playback and analysis, ensuring that the data has not been maliciously modified during storage and transmission. Digitally signing compressed target algorithm data ensures data integrity during storage and transmission, prevents the infiltration of fraudulent data, and supports data traceability. In the event of a data dispute, the digital signature serves as valid proof of data authenticity.
[0105] In one possible implementation, the signature construction formula is: Signature = Sign(Private_Key, Frame_ID || CompressedData). Here, Signature identifies the digital signature, and Sign represents the signature operation, which is the process of processing the input data using a specific encryption algorithm to generate a digital signature. Private_Key refers to the private key, a secret key belonging to the data sender, held only by the sender, used for signing and encrypting the data; its corresponding public key can be used to verify the validity of the signature. Frame_ID refers to the frame index, a unique identifier assigned sequentially to each original video frame in the real-time video stream, used to distinguish different video frames and ensure that the signature is associated with the target algorithm data of a specific video frame. || represents binary concatenation, that is, sequentially concatenating the Frame_ID and CompressedData into a single data string, which serves as the input for the signature operation. CompressedData refers to the compressed target algorithm data, i.e., the result obtained after compressing the target algorithm data; compression reduces the data size, facilitating storage and transmission.
[0106] In one possible implementation, the preset location is a private data area reserved within the original video frame. For example, the preset location is a private data area reserved in the header of the original video frame. This private data area refers to a region specifically designated within the unencoded data structure of the original video frame for storing additional information. This area does not affect the display of the original video frame and is only used to carry auxiliary data related to the video frame. For example, in an RGB format original video frame, space can be reserved in the additional information field of each pixel or the extended field of the frame header as a private data area. The private data area is closely integrated with the original video frame itself, physically belonging to the same data structure. The private data area within the original video frame enables the target algorithm data to form the most direct association with the original video frame, making data separation less likely during processing and ensuring the synchronization of data and video frames.
[0107] For example, for low-latency algorithms, the target algorithm data is stored in the private data area reserved in the header of the original video frame (such as the 0x100-0x1FF address range of the YUV frame).
[0108] In one possible implementation, the preset location is the metadata region of the encoded video frame after the original video frame has been encoded. The target algorithm data stored in this metadata region is associated with the original video frame via a timestamp. The metadata region refers to the storage area defined within the metadata of the encoded video frame after the original video frame has been encoded. Metadata is data describing information such as video content, format, and attributes, including encoding format, frame rate, and duration. Delineating a region within this metadata region to store the target algorithm data does not alter the content of the encoded video frame. For example, in an H.264 encoded video frame, metadata fields such as Supplementary Enhancement Information (SEI) can be used as the area for storing the target algorithm data. The encoded video frame corresponds to the original video frame, and the metadata region is associated with the encoded video frame, belonging to a part of the video file structure independent of the image data. The metadata region of the encoded video frame utilizes the structured data space after encoding, avoiding modification of the original video frame structure. Furthermore, the standardized structure of the metadata region facilitates subsequent parsing and extraction.
[0109] In one possible implementation, the preset location is a separate storage area bound to the original video frame by a timestamp, frame index, or pointer.
[0110] In one possible implementation, the preset location is an independent storage area bound to the original video frame via a timestamp, frame index, or pointer. This independent storage area refers to a space dedicated to storing the target algorithm data, separate from the storage area of the original video frame. It does not occupy the storage resources within the original video frame, nor does it belong to the metadata area of the encoded video frame. This independent storage area can be an independent database table, a separate file, or a specific partition in the storage medium. A timestamp records the time information of the original video frame's acquisition; a frame index is a unique number sequentially assigned to each original video frame in the real-time video stream, like a page number in a book; a pointer is an identifier of a storage address, pointing to the specific location of the original video frame in the storage medium. Timestamps, frame indexes, or pointers are all used to indicate the independent storage area bound to the original video frame. By associating the target algorithm data with the timestamp of the corresponding original video frame, it can be ensured that the target algorithm data corresponding to a certain original video frame can be quickly found based on the timestamp during subsequent processing or playback. The frame index allows direct location of the corresponding original video frame and its target algorithm data. Through pointer binding, the target algorithm data can be directly associated with the storage address of the original video frame, achieving a fast mapping between the two. The independent storage area and the original video frames are not physically contained within each other, but rather logically bound together through association information such as timestamps, frame indices, or pointers. The independent storage area does not affect the structure of the original video frames, avoiding potential video frame corruption or formatting issues that might result from embedding data within the original video frames, thus ensuring the integrity of the original video frames. The association between the target algorithm data and the original video frames is flexible and efficient. When it is necessary to modify, update, or extract the target algorithm data separately, no operation on the original video frames is required. Furthermore, the independent storage area can be flexibly expanded according to the size and quantity of the target algorithm data, without being limited by the storage capacity of the original or encoded video frames, better adapting to the changing needs of the target algorithm data volume in different scenarios.
[0111] In this embodiment, multiple preset locations are provided, which not only meet the data storage needs of different processing scenarios, but also ensure that there is a clear correlation between the target algorithm data and the corresponding original video frame (or encoded video frame). Through close physical or logical association, the synchronization deviation problem when the algorithm data and video frame are stored separately is avoided, ensuring the traceability of data during subsequent playback and analysis, further guaranteeing the effectiveness of video processing data in practical applications, and providing a reliable foundation for subsequent extraction, parsing and application.
[0112] In this embodiment of the application, the target algorithm data after compression and signature processing is treated as a whole data block and embedded in a preset location corresponding to the original video frame for storage.
[0113] In one possible implementation, the aforementioned overall data block is embedded within a region of the original video frame. The specific physical location depends on the encoding and encapsulation process and is implemented at the video encoder output stage. A custom NALU is inserted before the video encoding unit (such as the NALU for H.264 / H.265) that generates the original video frame. The NALU type is identified using an unused reserved type value (such as 31 or 24 for user private). The NALU payload contains the constructed embedded data block. For example, the embedded data block structure includes: [Start Identifier (Magic Number)]: 4 bytes, e.g., 0x56445349; [Total Data Block Length]: 4 bytes (L) (little-endian or big-endian), representing the length from the next byte to the end of the block; [Compressed Data Block Length]: 4 bytes (L) (little-endian or big-endian), representing the length from the next byte to the end of the compressed data block; [Compressed Data]: variable length; [Total Signature Data Block Length]: 4 bytes (L) (little-endian or big-endian), representing the length from the next byte to the end of the signature block; [Signature Data]; [Checksum]: e.g., [CRC32]: 4 bytes (for quick detection of data block corruption during storage / transmission). This embedded data block is physically adjacent to the corresponding video frame, and then the embedded block and the video frame are processed further (network transmission or disk storage, etc.). In one possible implementation, the signature and compressed data are stored in the corresponding frame header area of the video frame.
[0114] As one possible implementation of this application, during video playback, a target parser is invoked according to the scene type or the target algorithm, and the target parser extracts and parses the target algorithm data from the target video stream.
[0115] Video playback refers to the process of playing and viewing a stored target video stream. In this embodiment, video playback not only requires viewing the video footage but also acquiring the target algorithm data contained within it for analysis, tracing, and other operations. For example, when playing back driving videos in an intelligent driving scenario, it is necessary to extract vehicle distance data, pedestrian detection data, etc., for accident analysis.
[0116] A target parser is a pre-configured tool used to extract and parse data from specific target algorithms. It contains parsing rules corresponding to the scene type or target algorithm. Since different scene types correspond to different target algorithms, and the format and storage location of the target algorithm data also differ, a matching parser needs to be called. For example, for vehicle distance calculation algorithm data in a highway scene, a parser adapted to the algorithm's data format needs to be called; for pedestrian detection algorithm data in an urban road scene, a parser with the corresponding pedestrian data parsing rules needs to be called.
[0117] In this embodiment, the target parser locates the storage location of the target algorithm data in the target video stream according to preset rules (such as the private data area inside the original video frame or the metadata area of the encoded video frame), separates it from the preset location, and then converts the extracted data according to the format of the target algorithm data (such as data structure, encoding method, etc.) to make it information that can be directly understood and used. For example, the encoded pedestrian coordinate data is parsed into specific pixel position information. If the target algorithm data embedded in the preset location has undergone compression and signature processing, the target parser will first decode the embedded compressed data and verify the signature during video playback to ensure the integrity and validity of the data, supporting real-time playback and subsequent data analysis.
[0118] By using a target parser to extract and parse the target algorithm data, accurate and efficient acquisition of the target algorithm data can be ensured during video playback. Since the target parser is invoked based on scene type or target algorithm, matching the storage format and generation logic of the target algorithm data, data extraction failures or parsing errors caused by incompatible parsing tools are avoided.
[0119] As can be seen from the above, in this embodiment, by identifying the current scene type and determining the corresponding target algorithm, video processing can flexibly select the appropriate algorithm according to different scenes, avoiding the mismatch problem that may occur when using a fixed algorithm to process different scenes. This lays the foundation for improving the processing quality of real-time video streams. By executing the target algorithm on each original video frame of the real-time video stream, targeted target algorithm data can be generated for the actual situation of each frame, ensuring the accurate correspondence between the algorithm data and the video frame content, effectively improving the effectiveness of video processing. The target algorithm data is then embedded into the preset location corresponding to the original video frame for storage, forming a close association between the algorithm data and the original video frame, avoiding the problem of the algorithm data and video frame being out of sync. The embedded data can be extracted and verified during video playback, ensuring the validity and reliability of the data. This is of great significance for accident analysis, system optimization, and evidence collection, further ensuring the validity of the algorithm data in subsequent use. Thus, the overall flexible processing of real-time video streams is realized, improving the video processing quality and ensuring the validity of the algorithm data.
[0120] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0121] Corresponding to the video processing method described in the above embodiments, Figure 5A structural block diagram of a video processing apparatus provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0122] Reference Figure 5 The video processing device includes: an identification and determination unit 51, a data generation unit 52, and a video processing unit 53, wherein:
[0123] The identification and determination unit 51 is used to identify the current scene type and determine the target algorithm corresponding to the scene type;
[0124] Data generation unit 52 is used to execute the target algorithm on each original video frame of the real-time video stream and generate target algorithm data;
[0125] The video processing unit 53 is used to embed the target algorithm data into a preset location corresponding to the original video frame for storage, thereby generating a target video stream containing the target algorithm data.
[0126] As one possible implementation of this application, the video processing unit 53 includes:
[0127] The timing determination module is used to determine the embedding timing of the target algorithm data according to the type of the target algorithm and / or the scene type, wherein the embedding timing includes performing embedding before encoding the original video frame or performing embedding after encoding the original video frame;
[0128] An embedding generation module is used to embed the target algorithm data into the preset storage location based on a determined embedding timing sequence, thereby generating a target video stream containing the target algorithm data.
[0129] As one possible implementation of this application, the timing determination module is specifically used for:
[0130] Obtain the preset time consumption threshold corresponding to the type of the target algorithm;
[0131] Determine the end-to-end processing time from acquiring the original video frames to generating the target algorithm data;
[0132] Determine whether the end-to-end processing time exceeds the preset time consumption threshold;
[0133] If the end-to-end processing time does not exceed the preset time consumption threshold, then the target algorithm data is determined to be embedded before the original video frame encoding;
[0134] If the end-to-end processing time exceeds the preset time consumption threshold, then the target algorithm data is determined to be embedded after the original video frame is encoded.
[0135] As one possible implementation of this application, the timing determination module is further specifically used for:
[0136] Based on the mapping relationship between scene type and embedding time sequence in the preset scene type time sequence lookup table, the embedding time sequence of the target algorithm data is determined.
[0137] As one possible implementation of this application, the timing determination module is further specifically used for:
[0138] Obtain the preset time consumption threshold corresponding to the type of the target algorithm, and the priority corresponding to the scene type;
[0139] Determine the end-to-end processing time from acquiring the original video frames to generating the target algorithm data;
[0140] Determine whether the end-to-end processing time exceeds the preset time consumption threshold;
[0141] If the end-to-end processing time does not exceed the preset time consumption threshold and the priority of the scene type is higher than the preset priority threshold, then it is determined that the target algorithm data is embedded before the encoding of the original video frame.
[0142] If the end-to-end processing time exceeds the preset time consumption threshold or the priority of the scene type is lower than or equal to the preset priority threshold, then the target algorithm data is determined to be embedded after the original video frame is encoded.
[0143] As one possible implementation of this application, the preset location is a private data area reserved inside the original video frame, or the preset location is the metadata area of the encoded video frame after the original video frame is encoded.
[0144] As one possible implementation of this application, the video processing apparatus further includes:
[0145] The parsing processing unit is used to call the target parser according to the scene type or the target algorithm during video playback, and extract and parse the target algorithm data from the target video stream through the target parser.
[0146] As can be seen from the above, in this embodiment, by identifying the current scene type and determining the corresponding target algorithm, video processing can flexibly select the appropriate algorithm according to different scenes, avoiding the mismatch problem that may occur when using a fixed algorithm to process different scenes. This lays the foundation for improving the processing quality of real-time video streams. By executing the target algorithm on each original video frame of the real-time video stream, targeted target algorithm data can be generated for the actual situation of each frame, ensuring the accurate correspondence between the algorithm data and the video frame content, effectively improving the effectiveness of video processing. Furthermore, the target algorithm data is embedded into the preset location corresponding to the original video frame for storage, so that the algorithm data and the original video frame form a close association, avoiding the problem of the algorithm data and the video frame being out of sync, further ensuring the effectiveness of the algorithm data in subsequent use. Thus, the overall flexible processing of real-time video streams is realized, improving the video processing quality and ensuring the effectiveness of the algorithm data.
[0147] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0148] This application embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements... Figures 1 to 4 This represents the steps of any video processing method.
[0149] This application embodiment also provides a smart device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements... Figures 1 to 4 This represents the steps of any video processing method.
[0150] This application also provides a computer program product that, when run on a smart device, causes the smart device to execute the implementation of... Figures 1 to 4 This represents the steps of any video processing method.
[0151] Figure 6 This is a schematic diagram of a smart device provided in an embodiment of this application. Figure 6 As shown, the smart device 6 in this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, it implements the steps in the various video processing method embodiments described above, for example... Figure 1Steps S101 to S103 are shown. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 5 The functions of units 51 to 53 are shown.
[0152] For example, the computer program 62 may be divided into one or more modules / units, which are stored in the memory 61 and executed by the processor 60 to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program 62 in the smart device 6.
[0153] The intelligent device 6 may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will understand that... Figure 6 This is merely an example of smart device 6 and does not constitute a limitation on smart device 6. It may include more or fewer components than shown, or combine certain components, or different components. For example, smart device 6 may also include input / output devices, network access devices, buses, etc.
[0154] The processor 60 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0155] The memory 61 can be an internal storage unit of the smart device 6, such as a hard drive or memory of the smart device 6. The memory 61 can also be an external storage device of the smart device 6, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the smart device 6. Furthermore, the memory 61 can include both internal and external storage units of the smart device 6. The memory 61 is used to store the computer program and other programs and data required by the smart device. The memory 61 can also be used to temporarily store data that has been output or will be output.
[0156] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0157] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0158] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0159] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0160] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A video processing method, characterized in that, include: Identify the current scene type and determine the target algorithm corresponding to the scene type; The target algorithm is executed on each raw video frame of the real-time video stream to generate target algorithm data. The target algorithm data is embedded into a preset location corresponding to the original video frame and stored to generate a target video stream containing the target algorithm data.
2. The method according to claim 1, characterized in that, The step of embedding the target algorithm data into a preset location corresponding to the original video frame and storing it to generate a target video stream containing the target algorithm data includes: The embedding timing of the target algorithm data is determined based on the type of the target algorithm and / or the type of the scene, wherein the embedding timing includes performing embedding before encoding the original video frame or performing embedding after encoding the original video frame; The target algorithm data is embedded into the preset location for storage based on a determined embedding timing, thereby generating a target video stream containing the target algorithm data.
3. The method according to claim 2, characterized in that, The step of determining the embedding time sequence of the target algorithm data based on the type of the target algorithm includes: Obtain the preset time consumption threshold corresponding to the type of the target algorithm; Determine the end-to-end processing time from acquiring the original video frames to generating the target algorithm data; Determine whether the end-to-end processing time exceeds the preset time consumption threshold; If the end-to-end processing time does not exceed the preset time consumption threshold, then the target algorithm data is determined to be embedded before the original video frame encoding; If the end-to-end processing time exceeds the preset time consumption threshold, then the target algorithm data is determined to be embedded after the original video frame is encoded.
4. The method according to claim 2, characterized in that, The step of determining the embedding time sequence of the target algorithm data based on the scenario type includes: Based on the mapping relationship between scene type and embedding time sequence in the preset scene type time sequence lookup table, the embedding time sequence of the target algorithm data is determined.
5. The method according to claim 2, characterized in that, The step of determining the embedding time sequence of the target algorithm data based on the type of the target algorithm and the scene type includes: Obtain the preset time consumption threshold corresponding to the type of the target algorithm, and the priority corresponding to the scene type; Determine the end-to-end processing time from acquiring the original video frames to generating the target algorithm data; Determine whether the end-to-end processing time exceeds the preset time consumption threshold; If the end-to-end processing time does not exceed the preset time consumption threshold and the priority of the scene type is higher than the preset priority threshold, then it is determined that the target algorithm data is embedded before the encoding of the original video frame. If the end-to-end processing time exceeds the preset time consumption threshold or the priority of the scene type is lower than or equal to the preset priority threshold, then the target algorithm data is determined to be embedded after the original video frame is encoded.
6. The method according to claim 1, characterized in that, The preset location is a private data area reserved inside the original video frame, or the preset location is the metadata area of the encoded video frame after the original video frame is encoded.
7. The method according to any one of claims 1 to 6, characterized in that, The video processing method further includes: During video playback, the target parser is invoked according to the scene type or the target algorithm, and the target parser extracts and parses the target algorithm data from the target video stream.
8. A video processing apparatus, characterized in that, include: The identification and determination unit is used to identify the current scene type and determine the target algorithm corresponding to the scene type; The data generation unit is used to execute the target algorithm on each raw video frame of the real-time video stream and generate target algorithm data. The video processing unit is used to embed the target algorithm data into a preset location corresponding to the original video frame for storage, thereby generating a target video stream containing the target algorithm data.
9. A smart device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the video processing method as described in any one of claims 1 to 7.
Citation Information
Cited By
Video transmission method and device for bridge detection, equipment and storage medium
CN121418550A