Intelligent dynamic coding system based on multi-modal visual identification and control method

The intelligent dynamic coding system based on multimodal visual recognition solves the problems of complex core system structure and difficult maintenance, and achieves high efficiency, low power consumption, easy maintenance and multi-scenario adaptability, thereby improving the operational stability and efficiency of the equipment.

CN121834656APending Publication Date: 2026-04-10JIANGXI AWESOMEN NEW ENERGY TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGXI AWESOMEN NEW ENERGY TECHNOLOGY CO LTD
Filing Date
2025-12-18
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, the core system adopts an independent design of functional modules, which leads to complex structure, unstable operation, low efficiency, difficulty in adapting to the needs of multiple scenarios, and difficulty in maintenance, thus failing to meet the development needs of high efficiency, low consumption, and easy maintenance.

Method used

An intelligent dynamic coding system based on multimodal visual recognition is adopted. Through a data synchronization acquisition module, a spatiotemporal-adversarial feature fusion module, a dynamic trajectory prediction module, a coding path optimization module, and an execution feedback module, a closed-loop control system is formed to realize multimodal data fusion and dynamic coding path optimization.

Benefits of technology

Simplify system structure, improve operational stability and efficiency, reduce energy consumption, simplify maintenance process, adapt to diverse scenario requirements, and enhance equipment competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834656A_ABST
    Figure CN121834656A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent dynamic coding system based on multi-modal visual identification and a control method, and relates to the technical field of automatic coding, and the system comprises a data synchronous collection module which obtains multi-modal data containing a sensitive area and carries out calibration output; the space-time-confrontation feature fusion module extracts and fuses multi-modal features; the dynamic trajectory prediction module generates a precise motion trajectory; the coding path optimization module generates and optimizes a coding path; the execution and feedback module drives code printing equipment to execute, collects effect data and feeds back the effect data to the preorder module, and an iterative optimization closed loop is formed; an integrated framework is adopted, driving, regulating and monitoring units are deeply fused, the structure is simplified, stability and reliability are improved, multiple scenes are flexibly adapted, meanwhile, a multi-module cooperation mechanism is constructed, signal and energy interaction is optimized, efficiency is improved, energy consumption and cost are reduced, maintenance and upgrading are facilitated through modular design, and practicability and market competitiveness are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automated coding technology, specifically to an intelligent dynamic coding system and control method based on multimodal visual recognition. Background Technology

[0002] Currently, industrial automation, intelligent equipment manufacturing, and public service equipment are rapidly developing towards high integration, high efficiency, and scenario-based applications. Application scenarios are gradually extending from traditional fixed production environments to complex and ever-changing outdoor operations, flexible production, and public service scenarios. As the industry's demands for equipment operational stability, functional synergy, and cost control continue to increase, the performance level of the core system, as a key component of the equipment, directly determines the overall competitiveness of the equipment. However, in the existing technology system, most core systems are still designed based on the idea of ​​stacking single functional modules, which makes it difficult to meet the needs for structural simplification, functional synergy, and rapid adaptation in multiple scenarios. The industry urgently needs technical solutions that can break down module barriers and achieve deep integration of multiple functions to promote equipment upgrades and iterations, help related fields overcome technical bottlenecks, and improve overall development efficiency.

[0003] Traditional core systems typically employ an architecture with independently designed and distributed functional modules. These modules rely on complex external wiring and connectors for signal transmission and energy exchange, significantly increasing overall system complexity, assembly difficulty, and error risks. Furthermore, poor wiring connections and signal interference can lead to performance degradation and system failures, severely impacting operational stability. Moreover, the distributed architecture lacks efficient collaboration mechanisms, often operating independently and prone to resource idleness or overload bottlenecks. High energy consumption during operation significantly increases costs over the long term. Additionally, traditional systems require disassembly to repair or replace individual modules, extending maintenance cycles and causing substantial downtime, negatively impacting production and service continuity. This makes them ill-suited to the current industry demands for high efficiency, low energy consumption, and ease of maintenance. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an intelligent dynamic coding system and control method based on multimodal visual recognition. The system acquires multi-dimensional data such as video streams containing sensitive areas and user voice commands through a data synchronization acquisition module, and achieves multimodal feature fusion using a spatiotemporal-adversarial feature fusion module. The system accurately predicts the motion trajectory of sensitive areas through a dynamic trajectory prediction module, and generates a low-energy, high-coverage coding path by combining a coding path optimization module. Finally, the execution and feedback module drives the coding device and optimizes parameters in real time, forming a closed-loop control system that effectively improves coding efficiency and accuracy.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: On one hand, an intelligent dynamic coding system based on multimodal visual recognition, the system comprising:

[0006] Data synchronization acquisition module: synchronously acquires video streams containing sensitive areas, user voice commands, device motion parameters, and text information in video frames; performs frame synchronization calibration, noise filtering, format unification, and timestamp alignment on the acquired data, and outputs a multimodal raw dataset;

[0007] Spatiotemporal-adversarial feature fusion module: Extracts the spatiotemporal correlation and motion trends of sensitive areas in the video stream, extracts speech semantic features and text features, and calibrates the synchronization between speech commands and video actions. It also fuses multimodal features through a multimodal adversarial weighted fusion algorithm.

[0008] Dynamic trajectory prediction module: Based on fused features, it establishes a kinematic model by combining motion vectors calculated by optical flow method with inertial sensor data, generates an initial motion trajectory, performs dynamic compensation through a dual closed-loop robust trajectory prediction algorithm, and outputs the accurate motion trajectory of the sensitive area.

[0009] Coding path optimization module: Based on the precise motion trajectory, combined with the shape, size and motion characteristics of the sensitive area, an initial coding path with coverage redundancy design is generated through the energy efficiency-coverage balanced path generation algorithm; based on the dynamic energy consumption optimization algorithm, the motion speed, power distribution and path curvature of the coding device are adjusted, and the optimized coding path is output.

[0010] Execution and Feedback Module: Drives the coding device to perform coding operations along the optimized path, collects coding effect data in real time, and feeds it back to the preceding module through a closed-loop parameter adaptive update algorithm to form an iterative optimization closed loop.

[0011] Furthermore, in the data synchronization acquisition module, the process of synchronously acquiring video streams containing sensitive areas, user voice commands, device motion parameters, and text information in video frames is as follows: The image sensor, voice acquisition device, and inertial sensor are simultaneously activated via a hardware synchronization trigger signal; the image sensor continuously acquires video streams containing sensitive areas at a frame rate of 30fps, generating a frame identifier after each frame is acquired; the voice acquisition device receives user voice commands in real time under the trigger signal, generating an instruction identifier for each valid command received; the inertial sensor acquires device motion parameters at a sampling rate of 100Hz, generating a sampling identifier for each set of data sampled; after the video frame is acquired, the OCR engine is immediately invoked to extract the text information in the frame and generate a text identifier.

[0012] Furthermore, in the spatiotemporal-adversarial feature fusion module, the expression for the multimodal adversarial weighted fusion algorithm is: ,in, This represents the multimodal fusion feature vector at time t. The terms represent modal types: V specifically refers to video modality, S specifically refers to speech modality, and T specifically refers to text modality. For modal weights, To counteract the gradient of the loss function with respect to modal features, This represents the adversarial sensitivity coefficient for the corresponding mode. This represents the effectiveness coefficient of the mode corresponding to time t.

[0013] Furthermore, in the dynamic trajectory prediction module, the specific steps for establishing a kinematic model by combining the motion vector calculated using the optical flow method with inertial sensor data are as follows:

[0014] (1) For consecutive frames containing sensitive regions in the video stream, use the optical flow method to select 50-100 feature points in the sensitive region, calculate the displacement of each feature point in the x and y directions between adjacent frames, and take the average value of the displacement of all feature points as the motion vector of the sensitive region.

[0015] (2) The device motion parameters collected by the inertial sensor are time-stamped to match the timestamps of the video frames, and high-frequency noise is removed by sliding window filtering;

[0016] (3) Based on the motion trend information of sensitive areas in the fusion features, determine the current motion state and assign fusion weights to the motion vector generated by the optical flow method and the inertial sensor data;

[0017] (4) Substitute the fused motion data into the preset kinematic equations, fit the equation coefficients using the least squares method, and generate an initial kinematic model that can characterize the motion law of the sensitive area.

[0018] Furthermore, in the dynamic trajectory prediction module, the expression for the kinematic equations is: ,in, For sensitive areas in Position coordinates at that moment for The initial position at time, The initial velocity, for acceleration at any moment The fusion coefficient is... It is a time variable.

[0019] Furthermore, in the dynamic trajectory prediction module, the expression for the dual-loop robust trajectory prediction algorithm is: ,in, This represents the precise predicted coordinates of the sensitive region at frame t+k, where k represents the prediction frame offset, indicating the number of frames predicted from the current frame t. It is generated by the kinematic equations Initial predicted trajectory coordinates of the frame. for The deviation vector between the actual and predicted positions of the frame. For real-time compensation coefficient, This is a pre-trained adversarial attack trajectory offset vector used to simulate the impact of potential attacks on trajectory prediction. To determine the gradient magnitude of the adversarial loss for fused features, To counteract the compensation coefficient.

[0020] Furthermore, in the code-decoding path optimization module, the expression for the energy efficiency-coverage balanced path generation algorithm is: ,in, for The coordinates of the coded path points of the frame. To cover redundant vectors, This represents the precise predicted location coordinates of the sensitive area at frame t+k. This is the redundancy coefficient. The gradient of the path smoothing loss function. This is the smoothing coefficient.

[0021] Furthermore, in the coding path optimization module, the expression for the dynamic energy consumption optimization algorithm is: ,in, Real-time power consumption of the coding device. For the speed of equipment movement, Indicates the inherent power consumption coefficient related to device speed. This represents the device's base power consumption factor. Indicates that the redundant vector is covered at time t. The length of the mold, The redundancy power consumption factor is... For path curvature, This is the curvature power consumption coefficient.

[0022] Furthermore, in the execution and feedback module, the expression for the closed-loop parameter adaptive update algorithm is: ,in, The updated parameters to be optimized. for Parameter values ​​at time, For dynamic learning rate, For multimodal average coverage intersection-union ratio, For the average success rate of multimodal adversarial attacks, This indicates the weight of the intersection-union ratio (IU) indicator. The weighting of the metric indicating the success rate of counter-attacks. This represents the number of consecutive iterations without adjustment of the current parameters. The attenuation coefficient is... This represents the decay term based on the number of iterations.

[0023] On the other hand, the intelligent dynamic coding system and control method based on multimodal visual recognition, the specific steps of which are as follows:

[0024] S100, synchronous data acquisition and preprocessing: synchronously acquire video streams containing sensitive areas, user voice commands and device motion parameters through image sensors, voice acquisition devices and inertial sensors, call the OCR engine to extract text information in video frames, and output multimodal raw datasets after preprocessing the acquired data.

[0025] S200, Multimodal Feature Fusion: Extracts various features from the original multimodal data and calibrates synchronicity, then outputs fused features through a multimodal adversarial weighted fusion algorithm;

[0026] S300, Precise Trajectory Prediction for Sensitive Areas: Based on fused features, a kinematic model is established to generate an initial motion trajectory, and a dual-closed-loop robust trajectory prediction algorithm is used to output the precise motion trajectory of the sensitive area.

[0027] S400, coding path energy consumption optimization: Generate an initial coding path with coverage redundancy based on the accurate motion trajectory, and calculate the optimized coding path through a dynamic energy consumption optimization algorithm;

[0028] S500, coding execution feedback optimization: Drives the coding device to perform coding operations along the optimized path, collects coding effect data in real time, and feeds it back to the preceding module through a closed-loop parameter adaptive update algorithm to iteratively optimize the parameters of each algorithm formula.

[0029] Compared with existing technologies, this intelligent dynamic coding system and control method based on multimodal visual recognition has the following advantages:

[0030] I. This invention breaks through the limitations of the dispersed layout of functional modules in traditional technologies by adopting an integrated architecture that deeply integrates core units such as drive, control, and monitoring. This eliminates redundant connection structures and signal transmission gaps between modules. This design not only greatly simplifies the overall structural complexity and reduces the operational difficulty during assembly, but also reduces performance loss and potential failures caused by the interconnection of multiple modules, thereby improving the stability and reliability of system operation. At the same time, the integrated structure can flexibly adapt to different application scenarios and can be quickly deployed without additional modifications to the usage environment, effectively broadening the scope of application of the technology, meeting diverse usage needs, and providing a more convenient solution for equipment upgrades in related fields.

[0031] Second, this invention constructs a multi-module collaborative linkage mechanism, optimizes the signal interaction and energy transfer logic between various functional units, and enables the drive output, parameter control, and status monitoring to form a highly efficient collaborative working closed loop. This avoids resource waste and efficiency bottlenecks when a single module is running, significantly improving the overall system's operating efficiency. In addition, the design incorporates a low-consumption adaptation concept, reducing energy consumption during system operation and long-term operating costs by optimizing the working mode and material selection of core components. This also reduces the impact on the environment. Furthermore, the modular component design facilitates later maintenance and upgrades, allowing for the repair or replacement of local components without overall disassembly, shortening maintenance cycles, reducing downtime, ensuring long-term stable system operation, and further enhancing the practical value and market competitiveness of the technology.

[0032] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from an examination of the following, or may be learned from the practice of the invention. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0034] Figure 1 This is a flowchart illustrating the workflow of an intelligent dynamic coding system module based on multimodal visual recognition.

[0035] Figure 2 This is a flowchart of the control method steps for an intelligent dynamic coding system based on multimodal visual recognition. Detailed Implementation

[0036] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0037] Example 1:

[0038] Dynamic masking of sensitive information on live streaming platforms.

[0039] like Figure 1 As shown, this embodiment is applied to a live streaming platform. For areas such as sensitive identifiers and privacy information that may appear during the live stream, an intelligent dynamic decryption system based on multimodal visual recognition is used to accurately track and decrypt the content, ensuring the compliance of the live stream content.

[0040] Data synchronization acquisition module:

[0041] The system synchronously triggers signals via hardware to simultaneously activate the image sensor, voice acquisition device, and inertial sensor. The image sensor continuously acquires the live video stream at 30fps, generating a unique frame identifier after each frame is acquired. The voice acquisition device receives valid voice commands from the broadcaster in real time, such as enabling or adjusting the blurring range, generating a command identifier for each received command. The inertial sensor acquires the motion parameters of the live streaming equipment at a 100Hz sampling rate, generating a sampling identifier for each set of data samples. After each video frame is acquired, the system immediately calls the OCR engine to extract any sensitive text information that may be present in the frame and generates text identifiers. Subsequently, all acquired data undergoes frame synchronization calibration, noise filtering, format unification, and timestamp alignment, outputting a multimodal raw dataset containing video, voice, equipment motion, and text. This processing allows for accurate matching of various data types in the temporal dimension, providing a reliable foundation for the effective fusion of subsequent multimodal features.

[0042] Spatiotemporal-adversarial feature fusion module:

[0043] The system extracts spatiotemporal correlation information and motion trends of sensitive regions from the video stream, semantic features from voice commands, and key features from text information, while simultaneously calibrating the synchronization between voice commands and actions in sensitive regions of the video. Then, through a multimodal adversarial weighted fusion algorithm, combining the weights of each modality's adversarial loss function, gradient adversarial sensitivity coefficient, and effectiveness coefficient, the expression for the multimodal adversarial weighted fusion algorithm is: ,in, This represents the multimodal fusion feature vector at time t. The terms represent modal types: V specifically refers to video modality, S specifically refers to speech modality, and T specifically refers to text modality. For modal weights, To counteract the gradient of the loss function with respect to modal features, This represents the adversarial sensitivity coefficient for the corresponding mode. This represents the effectiveness coefficient of the modality at time t. The system fuses features from video, speech, and text modalities, outputting a fused feature that comprehensively reflects information about sensitive areas. This process integrates effective information from different modalities, reduces data redundancy and conflicts, and allows the fused features to more comprehensively characterize the attributes and dynamics of sensitive areas.

[0044] Dynamic trajectory prediction module:

[0045] Based on the motion trend of sensitive regions in the fusion features, the system first selects 50-100 feature points within the sensitive region using optical flow, calculates the x and y displacements of each feature point between adjacent frames, and takes the average displacement of all feature points as the motion vector of the sensitive region. Next, the system timestamps the device motion parameters collected by the inertial sensor to match the video frame timestamps, and removes high-frequency noise using sliding window filtering. Then, the system determines the current motion state based on the motion trend of the sensitive region, assigns fusion weights to the motion vector and the inertial sensor data, substitutes the fused motion data into a preset kinematic equation, and generates an initial kinematic model by fitting the equation coefficients using the least squares method. The expression of the kinematic equation is as follows: ,in, For sensitive areas in Position coordinates at that moment for The initial position at time, The initial velocity, for acceleration at any moment The fusion coefficient is... The time variable is used. Finally, a dual-loop robust trajectory prediction algorithm is applied, combining the real-time position deviation vector, the trajectory offset vector against adversarial attacks, and the gradient magnitude of the adversarial loss from the fused features, to dynamically compensate for the initial trajectory and output the accurate trajectory of the sensitive area. The expression for the dual-loop robust trajectory prediction algorithm is: ,in, This represents the precise predicted location coordinates of the sensitive area at frame t+k. It is generated by the kinematic equations Initial predicted trajectory coordinates of the frame. for The deviation vector between the actual and predicted positions of the frame. For real-time compensation coefficient, This is a pre-trained adversarial attack trajectory offset vector used to simulate the impact of potential attacks on trajectory prediction. To determine the gradient magnitude of the adversarial loss for fused features, To counteract the compensation coefficient, this series of operations allows the predicted trajectory to better match the actual movement in the sensitive area, maintaining high tracking accuracy even in the presence of device jitter or sudden movements.

[0046] Captcha solving path optimization module:

[0047] Based on the precise motion trajectory of the sensitive area, combined with the shape, size, and motion characteristics of the sensitive area, the system uses an energy efficiency-coverage balanced path generation algorithm, incorporating coverage redundancy design to generate an initial coding path, ensuring that the coding area completely covers the sensitive area. The expression for the energy efficiency-coverage balanced path generation algorithm is: ,in, for The coordinates of the coded path points of the frame. To cover redundant vectors, This represents the precise predicted location coordinates of the sensitive area at frame t+k. This is the redundancy coefficient. The gradient of the path smoothing loss function. is the smoothing coefficient. Then, through a dynamic energy consumption optimization algorithm, considering the motion speed, basic power consumption, coverage redundancy vector magnitude, and path curvature of the coding device, the algorithm adjusts the device's motion speed power distribution and path curvature to reduce energy consumption while ensuring coding coverage. The optimized coding path is output. The expression for the dynamic energy consumption optimization algorithm is: ,in, Real-time power consumption of the coding device. For the speed of equipment movement, Indicates the inherent power consumption coefficient related to device speed. This represents the device's base power consumption factor. Indicates that the redundant vector is covered at time t. The length of the mold, The redundancy power consumption factor is... For path curvature, The curvature power consumption coefficient is used. The optimized path can avoid coding omissions caused by movement in sensitive areas, reduce unnecessary power consumption of the equipment, and improve the economic efficiency of system operation.

[0048] Execution and Feedback Module:

[0049] The system drives the coding device to perform coding operations along the optimized path, while simultaneously collecting coding effect data in real time. Through a closed-loop parameter adaptive update algorithm, combined with dynamic learning rate coverage of intersection-union ratio weights, adversarial attack success rate weights, and iteration decay terms, the coding effect data is fed back to the preceding data synchronization acquisition feature fusion trajectory prediction and path optimization modules. This dynamically adjusts the parameters of each module, forming an iterative optimization closed loop. The expression for the closed-loop parameter adaptive update algorithm is: ,in, The updated parameters to be optimized. for Parameter values ​​at time, For dynamic learning rate, For multimodal average coverage intersection-union ratio, For the average success rate of multimodal adversarial attacks, This indicates the weight of the intersection-union ratio (IU) indicator. The weighting of the metric indicating the success rate of counter-attacks. This represents the number of consecutive iterations without adjustment of the current parameters. The attenuation coefficient is... This represents the decay term based on the number of iterations. This closed-loop design allows the system to continuously adjust according to the actual blurring effect, maintaining stable blurring accuracy even if the motion pattern of sensitive areas changes in the live streaming scenario.

[0050] In summary, this embodiment addresses the sensitive information masking requirements of live streaming platforms by constructing a complete system through five major modules. The data synchronization and acquisition module achieves time-series alignment and standardized processing of multi-source data, laying a solid foundation for subsequent operations; the spatiotemporal-adversarial feature fusion module integrates multimodal information to generate comprehensive and accurate feature vectors; the dynamic trajectory prediction module combines optical flow and a dual-closed-loop robust trajectory prediction algorithm to accurately capture the movement patterns of sensitive areas; the masking path optimization module uses energy efficiency-coverage balancing and dynamic energy consumption optimization algorithms to balance coverage integrity and operational economy; and the execution and feedback module uses a closed-loop parameter adaptive update algorithm for iterative optimization, ensuring accurate and stable masking. The entire system efficiently adapts to dynamic live streaming scenarios, comprehensively ensuring content compliance.

[0051] Example 2:

[0052] Dynamic blurring of sensitive content in short video review scenarios.

[0053] like Figure 2 As shown, this embodiment is applied to the review process of a short video platform. By using the control method of an intelligent dynamic blurring system based on multimodal visual recognition, sensitive areas in short videos are automatically and dynamically blurred, thereby improving review efficiency and content compliance.

[0054] S100, Data Synchronization Acquisition and Preprocessing:

[0055] The short video review system uses image sensor voice acquisition devices and inertial sensors to simultaneously collect voice information and motion parameters of the shooting device from the video stream of the short video to be reviewed. Simultaneously, it calls an OCR engine to extract text information from each frame of the video. Then, the collected video stream voice commands, device motion parameters, and text information undergo frame synchronization calibration, noise filtering, format unification, and timestamp alignment preprocessing to ensure consistent temporal sequence and standardized format of all data types, ultimately outputting a multimodal raw dataset. This preprocessed dataset ensures that the video stream voice information, device motion parameters, and text information are consistent in temporal sequence and standardized in format, removing data-level obstacles for subsequent feature extraction and fusion.

[0056] S200, multimodal feature fusion:

[0057] The system extracts spatiotemporal correlation features of the video modality, motion trends of sensitive areas, semantic features of the speech modality, and key features of the text modality from the original multimodal data. Simultaneously, it calibrates the synchronization between speech information and actions in sensitive areas of the video. Then, a multimodal adversarial weighted fusion algorithm is employed to fuse the three types of features by integrating relevant parameters from each modality. This eliminates redundancy and conflicts between different modalities, outputting fused features that comprehensively represent information about sensitive areas. The expression for the multimodal adversarial weighted fusion algorithm is as follows: The fused features can integrate information from multiple aspects such as video, speech, and text, enabling the system to have a more comprehensive and accurate understanding of sensitive areas and providing high-quality input for subsequent trajectory prediction.

[0058] S300, Precise Trajectory Prediction in Sensitive Areas:

[0059] First, a kinematic model is established based on the fused features: Motion vectors are calculated by selecting 50-100 feature points within the sensitive region using optical flow. Inertial sensor data is then time-stamped and noise filtered. Fusion weights are assigned, and the kinematic equations are input to fit coefficients to generate an initial model. The expression for the kinematic equations is as follows: Then, a dual-loop robust trajectory prediction algorithm is used, combining real-time position deviation adversarial attack offset and adversarial loss gradient magnitude, to dynamically compensate for the initial motion trajectory, accurately predict the position of sensitive areas in the short video in subsequent frames, and output the accurate motion trajectory of the sensitive areas. The expression of the dual-loop robust trajectory prediction algorithm is as follows: The resulting trajectory can accurately capture the movement patterns of sensitive areas, and even if there are camera cuts or object obstructions in the short video, it can accurately predict the positional changes of sensitive areas.

[0060] S400, optimized energy consumption for coding path:

[0061] Based on the precise motion trajectory of the sensitive area, combined with its shape, size, and motion characteristics, the system generates an initial coding path with coverage redundancy design using an energy efficiency-coverage balanced path generation algorithm. The expression for the energy efficiency-coverage balanced path generation algorithm is as follows: To avoid missing coding due to movement in sensitive areas, a dynamic energy consumption optimization algorithm is used to balance the relationship between the coding device's movement speed, power distribution path curvature, and energy consumption. Relevant parameters are adjusted to optimize the coding path, reducing device energy consumption while ensuring complete coding coverage. The optimized coding path is then output. The expression for the dynamic energy consumption optimization algorithm is: The optimized path ensures complete coverage of sensitive areas while reducing device energy consumption, making the coding process more efficient and economical.

[0062] S500, optimized captcha execution feedback:

[0063] The system drives the decryption device to perform decryption operations on sensitive areas in short videos along an optimized path, collecting decryption effect data for each frame in real time, including metrics such as the average coverage, intersection ratio, and average adversarial attack success rate of the decrypted area and the sensitive area. Through a closed-loop parameter adaptive update algorithm, combined with dynamic learning rate weights and iteration decay terms, the decryption effect data is fed back to previous steps, dynamically updating the coefficients of the weighted trajectory prediction and path optimization parameters of the parameter feature fusion in the data acquisition preprocessing. The expression for the closed-loop parameter adaptive update algorithm is: This forms an iterative optimization loop. This closed-loop optimization allows the parameters of each step to be dynamically adjusted according to the actual blurring effect, maintaining stable blurring quality and improving overall review efficiency even if the types of sensitive areas or movement patterns in the short video are diverse.

[0064] In summary, this embodiment focuses on the needs of short video review and decryption, achieving automated and accurate decryption through a five-step process. Data synchronous collection and preprocessing standardizes the format and timing of multi-source data, eliminating data obstacles; the multimodal feature fusion module eliminates redundancy and conflicts through a dedicated algorithm, outputting high-quality fused features; accurate trajectory prediction for sensitive areas relies on kinematic models and a dual-closed-loop robust trajectory prediction algorithm to accurately predict positional changes; decryption path optimization balances coverage and energy consumption, generating efficient paths; execution feedback optimization dynamically adjusts parameters through a closed-loop update mechanism to adapt to diverse sensitive area scenarios. This method significantly improves the efficiency of short video review, ensures stable and reliable decryption quality, and assists in the platform's content compliance management.

[0065] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. An intelligent dynamic coding system based on multimodal visual recognition, characterized in that, The system includes: Data synchronization acquisition module: synchronously acquires video streams containing sensitive areas, user voice commands, device motion parameters, and text information in video frames; performs frame synchronization calibration, noise filtering, format unification, and timestamp alignment on the acquired data, and outputs a multimodal raw dataset; Spatiotemporal-adversarial feature fusion module: Extracts the spatiotemporal correlation and motion trends of sensitive areas in the video stream, extracts speech semantic features and text features, and calibrates the synchronization between speech commands and video actions. It also fuses multimodal features through a multimodal adversarial weighted fusion algorithm. Dynamic trajectory prediction module: Based on fused features, it establishes a kinematic model by combining motion vectors calculated by optical flow method with inertial sensor data, generates an initial motion trajectory, performs dynamic compensation through a dual closed-loop robust trajectory prediction algorithm, and outputs the accurate motion trajectory of the sensitive area. Coding path optimization module: Based on the precise motion trajectory, combined with the shape, size and motion characteristics of the sensitive area, an initial coding path with coverage redundancy design is generated through the energy efficiency-coverage balanced path generation algorithm; based on the dynamic energy consumption optimization algorithm, the motion speed, power distribution and path curvature of the coding device are adjusted, and the optimized coding path is output. Execution and Feedback Module: Drives the coding device to perform coding operations along the optimized path, collects coding effect data in real time, and feeds it back to the preceding module through a closed-loop parameter adaptive update algorithm to form an iterative optimization closed loop.

2. The intelligent dynamic coding system based on multimodal visual recognition according to claim 1, characterized in that, In the data synchronization acquisition module, the process of synchronously acquiring video streams containing sensitive areas, user voice commands, device motion parameters, and text information in video frames is as follows: The image sensor, voice acquisition device, and inertial sensor are simultaneously activated via a hardware synchronization trigger signal; the image sensor continuously acquires video streams containing sensitive areas at a frame rate of 30fps, generating a frame identifier after each frame is acquired; the voice acquisition device receives user voice commands in real time under the trigger signal, generating an instruction identifier for each valid command received; the inertial sensor acquires device motion parameters at a sampling rate of 100Hz, generating a sampling identifier for each set of data sampled. After the video frame is captured, the OCR engine is immediately invoked to extract the text information in the frame and generate a text identifier.

3. The intelligent dynamic coding system based on multimodal visual recognition according to claim 1, characterized in that, In the spatiotemporal-adversarial feature fusion module, the expression for the multimodal adversarial weighted fusion algorithm is: ,in, This represents the multimodal fusion feature vector at time t. The terms represent modal types: V specifically refers to video modality, S specifically refers to speech modality, and T specifically refers to text modality. For modal weights, To counteract the gradient of the loss function with respect to modal features, This represents the adversarial sensitivity coefficient for the corresponding mode. This represents the effectiveness coefficient of the mode corresponding to time t.

4. The intelligent dynamic coding system based on multimodal visual recognition according to claim 1, characterized in that, In the dynamic trajectory prediction module, the specific steps for establishing a kinematic model by combining the motion vector calculated by the optical flow method with inertial sensor data are as follows: (1) For consecutive frames containing sensitive regions in the video stream, use the optical flow method to select 50-100 feature points in the sensitive region, calculate the displacement of each feature point in the x and y directions between adjacent frames, and take the average value of the displacement of all feature points as the motion vector of the sensitive region. (2) The device motion parameters collected by the inertial sensor are time-stamped to match the timestamps of the video frames, and high-frequency noise is removed by sliding window filtering; (3) Based on the motion trend information of sensitive areas in the fusion features, determine the current motion state and assign fusion weights to the motion vector generated by the optical flow method and the inertial sensor data; (4) Substitute the fused motion data into the preset kinematic equations, fit the equation coefficients using the least squares method, and generate an initial kinematic model that can characterize the motion law of the sensitive area.

5. The intelligent dynamic coding system based on multimodal visual recognition according to claim 4, characterized in that, In the dynamic trajectory prediction module, the kinematic equations are expressed as follows: ,in, For sensitive areas in Position coordinates at that moment for The initial position at time, The initial velocity, for acceleration at any moment The fusion coefficient is... It is a time variable.

6. The intelligent dynamic coding system based on multimodal visual recognition according to claim 1, characterized in that, In the dynamic trajectory prediction module, the expression for the dual-loop robust trajectory prediction algorithm is: ,in, This represents the precise predicted coordinates of the sensitive region at frame t+k, where k represents the prediction frame offset, indicating the number of frames predicted from the current frame t. It is generated by the kinematic equations Initial predicted trajectory coordinates of the frame. for The deviation vector between the actual and predicted positions of the frame. For real-time compensation coefficient, This is a pre-trained adversarial attack trajectory offset vector used to simulate the impact of potential attacks on trajectory prediction. To determine the gradient magnitude of the adversarial loss for fused features, To counteract the compensation coefficient.

7. The intelligent dynamic coding system based on multimodal visual recognition according to claim 1, characterized in that, In the code-solving path optimization module, the expression for the energy efficiency-coverage balanced path generation algorithm is: ,in, for The coordinates of the coded path points of the frame. To cover redundant vectors, This represents the precise predicted location coordinates of the sensitive area at frame t+k. This is the redundancy coefficient. The gradient of the path smoothing loss function. This is the smoothing coefficient.

8. The intelligent dynamic coding system based on multimodal visual recognition according to claim 1, characterized in that, In the coding path optimization module, the expression for the dynamic energy consumption optimization algorithm is: ,in, Real-time power consumption of the coding device. For the speed of equipment movement, Indicates the inherent power consumption coefficient related to device speed. This represents the device's base power consumption factor. Indicates that the redundant vector is covered at time t. The length of the mold, The redundancy power consumption factor is... For path curvature, This is the curvature power consumption coefficient.

9. The intelligent dynamic coding system based on multimodal visual recognition according to claim 1, characterized in that, In the execution and feedback module, the expression for the closed-loop parameter adaptive update algorithm is: ,in, The updated parameters to be optimized. for Parameter values ​​at time, For dynamic learning rate, For multimodal average coverage intersection-union ratio, For the average success rate of multimodal adversarial attacks, This indicates the weight of the intersection-union ratio indicator. The weighting of the indicator representing the success rate of counter-attacks. This represents the number of consecutive iterations without adjustment of the current parameters. The attenuation coefficient is... This represents the decay term based on the number of iterations.

10. A control method for an intelligent dynamic coding system based on multimodal visual recognition, the method being used to control the intelligent dynamic coding system based on multimodal visual recognition as described in any one of claims 1-9, characterized in that, The specific steps of this method are as follows: S100, synchronous data acquisition and preprocessing: synchronously acquire video streams containing sensitive areas, user voice commands and device motion parameters through image sensors, voice acquisition devices and inertial sensors, call the OCR engine to extract text information in video frames, and output multimodal raw datasets after preprocessing the acquired data. S200, Multimodal Feature Fusion: Extracts various features from the original multimodal data and calibrates synchronicity, then outputs fused features through a multimodal adversarial weighted fusion algorithm; S300, Precise Trajectory Prediction for Sensitive Areas: Based on fused features, a kinematic model is established to generate an initial motion trajectory, and a dual-closed-loop robust trajectory prediction algorithm is used to output the precise motion trajectory of the sensitive area. S400, coding path energy consumption optimization: an initial coding path with coverage redundancy is generated based on the accurate motion trajectory, and the optimized coding path is calculated through a dynamic energy consumption optimization algorithm; S500, coding execution feedback optimization: Drives the coding device to perform coding operations along the optimized path, collects coding effect data in real time, and feeds it back to the preceding module through a closed-loop parameter adaptive update algorithm to iteratively optimize the parameters of each algorithm formula.

Citation Information

Patent Citations

  • Multi-variety vegetable harvester dynamic identification and feeding control system based on image processing

    CN120359903A

  • Data marking method and device, equipment and medium

    CN120744113A

  • Audio and video object intelligent tracking optimization method and system combined with deep learning

    CN120892764A

  • AI generation content detection and review method and device, equipment and storage medium

    CN121093102A

  • Multi-modal information tagging method, apparatus and device, and storage medium and product

    WO2025148651A1