Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

729 results about "Dynamic feature" patented technology

Dynamic Features is a Video production company based in Los angeles California.

Abnormal short message behavior detection method and system based on multi-dimensional feature fusion

The invention discloses an abnormal short message behavior detection method and system based on multi-dimensional feature fusion, and relates to the related technical field of short message security detection.The method comprises the steps that spatial and temporal distribution features, semantic association maps and equipment behavior fingerprints of short message interaction are collected, a dynamic feature pool is configured, and a cross-modal feature sequence is extracted; cascade identification is carried out; the feature fusion weight matrix is dynamically adjusted, an abnormal probability score is generated, and when the abnormal probability score exceeds a dynamic abnormal probability threshold, a multi-stage verification mechanism is triggered; and matching a time sequence mode at an edge computing node, dynamically generating a verification code triggering threshold value, performing interactive risk verification, and determining an abnormal short message behavior mark. The technical problems of insufficient detection timeliness and adaptability and high false alarm and missing report rate caused by single detection dimension and difficulty in identifying novel complex abnormal short message behaviors in the prior art are solved, and the technical effects of reducing the false alarm and missing report rate of short message anomaly detection and improving the detection timeliness and adaptability are achieved.
Owner:SHENZHEN YINGJIETONG INFORMATION TECHNOLOGY CO LTD

Support structure stress state monitoring method based on artificial intelligence

The invention relates to a supporting structure stress state monitoring method based on artificial intelligence, and belongs to the technical field of artificial intelligence and data processing. The method comprises the following steps: acquiring and marking strain data of a supporting structure; after abnormal values are removed, normalizing the multi-sensor data to generate a normalized strain sequence; a state monitoring model is constructed, a deep time sequence neural network architecture is adopted, and the state monitoring model comprises an input layer, a self-adaptive wavelet attention feature mapping layer, a time domain gating convolution module, a global maximum pooling layer, a dynamic feature importance reweighting layer and a full-connection classification layer; inputting a normalized data training model; optimizing a loss function through a quantile interval adaptive learning rate and a momentum updating strategy; after real-time monitoring data is processed, inputting the data into the training model according to time window slices, outputting four types of probabilities, and taking the maximum value as a prediction state; and if a plurality of continuous windows are early-warning and dangerous, triggering the terminal to give an alarm. The accuracy of monitoring the stress state of the supporting structure can be improved.
Owner:SHANDONG JIANZHU UNIV

Audio and video identity recognition system based on multimode clue driving

The invention relates to the technical field of audio and video identity recognition, in particular to an audio and video identity recognition system based on multimode clue driving. The audio and video identity recognition system comprises an audio feature extraction module, a video feature extraction module, a multimode clue fusion module, a living body detection module and an identity recognition and verification module. The audio feature extraction module is used for capturing a voice signal of a user and extracting key voiceprint features, the video feature extraction module is used for acquiring facial features or limb features of the user and extracting related features, and the multi-mode clue fusion module is used for performing intelligent weighted fusion on the features of audio and video modes. The method comprises the following steps: firstly, using a living body detection module to comprehensively analyze the dynamic characteristics of audio and video, judging whether a user is a real individual, and finally, completing the final verification of the identity of the user through an identity recognition and verification module. The problem of feature extraction and modal fusion in the prior art is effectively solved.
Owner:NANJING LONGYUAN INFORMATION TECH CO LTD

CVOCA feature extraction method and system based on multi-modal large model

The invention discloses a CVOCA feature extraction method and system based on a multi-modal large model, and belongs to the technical field of multi-modal feature extraction. Data preprocessing: processing the input data to generate standardized input; feature unified representation: utilizing a cross-modal unified encoder to map visual, text and time series data to a unified feature space; feature optimization: dynamically adjusting modal weight based on a dynamic feature optimizer to realize efficient feature fusion; carrying out context sensing modeling, and capturing long-range time sequence dependence; carrying out adversarial optimization, and improving the robustness of the model through adversarial training; the feature storage and updating system adopts an expert network and a gating mechanism to realize dynamic updating of a feature library; according to the method, the defects of an existing network behavior analysis technology in the aspects of multi-modal data fusion, calculation efficiency, anti-robustness and dynamic adaptability are overcome, and the technical performance and the application value are improved.
Owner:EVERSEC BEIJING TECH +1

Small-sample defect identification method based on cross-modal text semantic driving

The invention discloses a few-sample defect identification method based on cross-modal text semantic driving, and belongs to the technical field of image processing. Aiming at the problem of insufficient generalization of a detection model caused by scarcity of abnormal samples and dynamic evolution of defect types in an industrial quality inspection scene, an unknown defect type can be accurately identified only by a small amount of normal data by establishing a dynamic feature recombination mechanism and an adaptive discrimination boundary; according to the method, a simulation sample similar to a real defect in form is generated on a normal sample through a matching relation between text description and image features; when a defect type which is not seen is encountered, a comparison standard of image textures can be automatically adjusted according to text semantics, subtle differences between a normal area and an abnormal area can be accurately distinguished, dependence on real defect data is not needed, and the sample defect identification precision is further improved.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Visual language multi-modal fusion method based on parameter-free cross attention

The invention discloses a visual language multi-modal fusion method based on parameter-free cross attention, and belongs to the field of computer vision. The implementation method comprises the following steps: using a fixed pre-training language model as a trunk, using a visual encoder to extract image features, and calculating a cross attention weight between language query and visual features through a parameter-free activation function, replacing a plurality of groups of learnable projection matrixes introduced by a traditional cross attention module, and significantly reducing the model parameter scale. A multi-scale visual feature generation mechanism based on pooling operation is introduced, and rich visual semantic prompt information is provided for a language model. A dynamic feature selection module is designed in combination with cross attention, visual areas corresponding to all text tokens are screened, low-correlation areas are discarded, only visual content more contributing to the current language context is reserved, accurate information matching and efficient fusion between modals are achieved, and the accuracy and efficiency of information fusion are improved. And the performance of the visual language model in tasks such as image-text question answering, image generation and multi-modal instruction understanding is improved.
Owner:BEIJING INST OF TECH

Intelligent dialogue system and method based on AI multi-mode large model

The invention relates to an intelligent dialogue system and method based on an AI multi-mode large model. The method comprises the following steps: collecting biological characteristic data of a user in real time; cloning a personalized expression mode of the user by using the generative adversarial network; driving the audio avatar to perform multi-round dialogue interaction with the user, capturing a real-time physiological signal of the user through an expression recognition module, and generating a dynamic response content suggestion in combination with a dialogue context; and analyzing the interaction process based on the reinforcement learning model. The multi-modal deep fusion of the voice rhythm features, the language structure features and the facial dynamic features is realized, so that the limitation of single-modal or simple feature splicing in the prior art is broken through, the emotional state of the user can be accurately and comprehensively captured, the intention is expressed and the fine physiological reaction is expressed, and the user experience is improved. And a solid foundation is laid for constructing high-fidelity user representation.
Owner:NANCHANG YIJING INFORMATION TECH CO LTD

Concrete structure internal defect nondestructive testing method fusing big data feature extraction and deep learning

The invention discloses a nondestructive testing method for internal defects of a concrete structure fusing big data feature extraction and deep learning. According to the method, through multi-source data collaboration and dynamic feature fusion, the accuracy and robustness of concrete structure defect detection are remarkably improved. In a data processing link, ultrasonic electromagnetic induction infrared thermal imaging data and the like acquired by a multi-source nondestructive testing technology are subjected to collaborative preprocessing, so that the influence of noise interference and environmental fluctuation is eliminated, and standardized input is provided for feature extraction. The dynamic weight distribution network further combines the relevance of each modal feature in a historical defect sample, adjusts fusion weights of different modals in real time, reinforces ultrasonic features with great contribution to cavity recognition or infrared features sensitive to cracks, effectively compresses redundant information, and improves the accuracy of cavity recognition. According to the method, features and data-driven deep features of the fused feature vectors are manually designed at the same time, so that the limitation of single-modal data is avoided, and a model can more accurately capture multi-dimensional features of defects.
Owner:JIANGSU TESTING CENT FOR QUALITY OF CONSTR ENG

Complex equipment fault diagnosis method and system based on dynamic characteristic modeling

The invention relates to the technical field of fault diagnosis, in particular to a complex equipment fault diagnosis method and system based on dynamic feature modeling. The method comprises the following steps: acquiring multi-channel sensor data; performing data preprocessing on the acquired multi-channel sensor data; extracting initial channel features from the preprocessed data; on the basis of a direction perception mechanism, performing convolution enhancement after splicing the initial channel features, and generating global context features of structure perception; performing multi-time-sequence scale feature fusion on the global context features based on a multi-scale modulation mechanism; performing health trend guidance on the fused features based on a trend guidance supervision mechanism; the diagnosis precision, the trend perception capability and the practical response efficiency of the system are remarkably improved, the system is effectively adapted to core application scenes of multi-industry equipment in predictive maintenance, online fault perception, remote diagnosis analysis and the like, and the system has good popularization prospects and practical values.
Owner:YANTAI UNIV

Intelligent element positioning method and system based on AI and dynamic feature library

The invention discloses an intelligent element positioning method and system based on AI and a dynamic feature library, and the method comprises the steps: carrying out the anomaly detection and abnormal data capture, executing an element positioning operation through an automatic script, and recording a page DOM tree structured snapshot when the positioning is judged to be failed; multi-modal feature extraction: analyzing the DOM tree, and extracting an XPath / CSS path and semantic attributes of a target element; performing dynamic feature library retrieval and matching, calculating DOM structure similarity, and taking a high-similarity rule as a correction basis if the high-similarity rule exists; an AI intelligent rule is generated, and candidate positioning rules are generated through the classification model; performing script correction and real-time feedback, replacing the failure positioning statement and updating the dynamic feature library; and generating a structured report, and recording abnormal data. According to the method, the AI is combined with the dynamic feature library, so that the intelligent correction of element positioning in the automatic test process is realized, and the stability and efficiency of the automatic test are improved.
Owner:SHANGHAI TIANHAO INFORMATION TECH CO LTD

Multi-terminal video display adaptive method based on hierarchical traffic prediction

The invention discloses a hierarchical traffic prediction-based multi-terminal video display adaptive method, which comprises the following steps of: acquiring network traffic, user behaviors and equipment state information in real time by constructing a multi-dimensional feature acquisition module, and generating a dynamic feature matrix containing traffic, interaction and performance features; processing the feature matrix by using a deep learning model, and generating and dynamically adjusting a network traffic grading prediction result; formulating an optimized video content distribution strategy based on a prediction result and equipment performance parameters, and coordinating multi-terminal resource distribution; through user behavior analysis and network state monitoring, a self-adaptive video display control strategy is generated, dynamic adaptation of video resolution, frame rate and playing logic is realized, a full-link collaborative optimization module is constructed, terminal, network and application layer resource allocation is coordinated, a video transmission and display strategy is optimized, and network flow prediction and multi-terminal interaction requirements are combined to realize multi-terminal interaction. And generating a dynamic video layout and content switching strategy.
Owner:ZHONGKE RUANQI (WUHAN) TECH CO LTD

Long-time pedestrian re-identification method based on dual-path cooperation and key frame guided reconstruction

The invention discloses a long-time pedestrian re-identification method based on dual-path cooperation and key frame guided reconstruction. The method comprises the steps of firstly collecting a pedestrian video to be recognized, and extracting a video feature sequence; space and time position coding is introduced into the video feature sequence; capturing local fine-grained dynamic features through a local dynamic feature capturing path, and modeling long-range time sequence association through a cross-frame global feature modeling path; then, dual-path feature complementation is realized through bidirectional gating interaction; further screening out key frames, and realizing feature reconstruction through a full-frame attention propagation mechanism; and finally fusing the dual-path fusion features, the key frame guide reconstruction features and the refined features to generate pedestrian identity features. And processing pedestrian identity features to obtain standardized feature vectors, performing similarity comparison on the standardized feature vectors and pedestrian features in an image library, and returning a matching list. According to the method, video time sequence information is fully utilized, and the problem of insufficient robustness caused by appearance change in long-time pedestrian re-identification is effectively solved.
Owner:SHIJIAZHUANG TIEDAO UNIV

Multi-mode interaction control method for Bluetooth knob screen

The invention discloses a multi-modal interaction control method for a Bluetooth knob screen, and relates to the technical field of multi-modal interaction control, and the method comprises the steps: synchronously collecting and forming a multi-modal feature vector set, and outputting a final operation instruction through a multi-modal fusion model; performing function mapping according to the final operation instruction, and identifying a control command; judging whether to resend the same instruction or adjust the instruction parameter based on the control command, and dynamically adjusting the tactile feedback parameter based on the function mapping of the current scene mode; performing low power consumption management based on the dynamically adjusted tactile feedback parameters, and triggering scene mode mapping update according to the wake-up instruction type; and optimizing the instruction output accuracy and the function mapping rule of the multi-modal fusion model based on the updated user interaction data. According to the method, deep fusion and intelligent cooperation of multiple interaction modes are realized, and accurate output of operation instructions, intelligent management of tactile feedback and adaptive updating of scene modes are achieved through dynamic function mapping, intelligent feedback adjustment and continuous learning optimization.
Owner:SUZHOU YUNGAN INTELLIGENT TECHNOLOGY CO LTD

Intelligent data alignment method and system based on time sequence dynamic multi-source embedded mapping

The invention provides an intelligent data alignment method and system based on time sequence dynamic multi-source embedding mapping. The method belongs to the technical field of multi-modal data fusion and spatio-temporal information processing. The method comprises the following steps: performing time sequence dynamic feature extraction on a multi-source heterogeneous data source to generate a heterogeneous data sequence containing a time dependency relationship; and constructing a time sequence dynamic multi-source embedded manifold space based on a manifold learning theory, mapping a heterogeneous data sequence to a unified evolution geometric structure representation space, and generating embedded manifold data. Through the method, heterogeneous data from various different data sources can be effectively processed, unified mapping is carried out through time sequence dynamic feature extraction and a manifold learning technology, cross-source alignment of the data is achieved, and the method is particularly suitable for a data scene needing to consider a time dependency relationship.
Owner:ZHEJIANG STARSINO INFORMATION TECH

Non-intrusive load identification method and system based on multi-modal depth feature fusion

The invention provides a non-intrusive load identification method and system based on multi-modal depth feature fusion, and the method comprises the steps: extracting the dynamic features of voltage and current time sequence data through an LSTM network, and capturing the spatial features of an equipment operation image through a pre-training ResNet-18 model; designing a cross-modal multi-head attention mechanism, dynamically aligning semantic information of voltage and current features by taking image features as a query reference, and generating enhanced characterization; a wavelet transform fusion layer is constructed, deep fusion of time-frequency domain features is realized through a multi-stage decomposition-reconstruction mechanism, and the anti-noise capability is improved in combination with a learnable convolution enhancement module; and a cross entropy and contrast learning double-loss joint optimization strategy is adopted, multi-modal feature hidden space distribution is constrained, and model generalization is enhanced. According to the method, 97.1% of average classification accuracy is realized on a plurality of types of household appliance data sets, the accuracy requirement of load identification is met, and an efficient solution is provided for load identification.
Owner:WUHAN UNIV

Dynamic feature retrieval generation and management method based on cross attention mechanism

The invention discloses a dynamic feature retrieval generation and management method based on a cross attention mechanism, which belongs to the technical field of natural language processing, and comprises the following steps: step 1, converting knowledge base document fragments into atomic knowledge units, each atomic knowledge unit comprising question and answer pairs, codes as key value pairs, and adding dynamic priority weights; self-attention query is replaced with a double-channel query structure, one channel is used for cross attention, a cross attention query vector is generated through linear transformation, and key value pairs of atomic knowledge units are used; and step 3, training a cross attention adapter, freezing the weight of the language model, optimizing parameters of the adapter, and dynamically adjusting a loss function by using pre-judgment parameters based on input sequence context complexity. By means of the method, deep correlation description of user input and knowledge fragments can be achieved, semantic ambiguity is eliminated, knowledge injection and context information are balanced, distortion is avoided, and the strict requirement for semantic precision in the professional field is met.
Owner:ZHONGSHU (XIAMEN) INFORMATION TECH CO LTD +1

Method and system for detecting nonferrous metal target of scraped car

The invention belongs to the technical field of image processing, and discloses a nonferrous metal target detection method and system for a scraped car. According to the method, through deep coupling of a GAN bimodal fusion network and a double-backbone network, a scraped car nonferrous metal target detection model is constructed, so that nonferrous metals in scraped cars are efficiently and accurately sorted. A bimodal fusion network is introduced into a model input layer, a fusion image with infrared thermal saliency and visible light texture features is generated, and the problem of cross-modal information splitting is solved; according to the dual-backbone network, standard convolution is replaced by a three-branch structure of an MGHCM module, large target contours, middle target semantics and small target details are synchronously captured, and self-adaptive shunting and efficient fusion of multi-scale features are achieved in cooperation with dynamic feature routing of a DHFBlock module; and the detection head network is optimized into a rotating frame prediction and cross-modal cross entropy classification mechanism so as to adapt to the placement scene of the metal fragments at any angle.
Owner:KUNMING UNIVERSITY

Multi-modal time sequence anomaly analysis method and device, equipment and medium

The invention relates to the technical field of data analysis, and discloses a multi-modal time sequence anomaly analysis method, device, equipment and medium, and the method comprises the steps: collecting visual data, audio data and process text data, carrying out the preprocessing and standardization processing of different types of data, constructing a multi-modal data set with aligned timestamps, and storing the multi-modal data set in a database; the method comprises the following steps: extracting a time-frequency dynamic feature and a semantic vector feature, extracting a map structure feature, a time-frequency dynamic feature and a semantic vector feature, fusing the features by using a cross-modal attention mechanism to generate a fused feature vector, finally performing analysis processing based on the fused feature vector, and outputting an analysis result. According to the method, the multi-modal data set with consistent time is constructed, the structural features of various modals are extracted, and the cross-modal attention mechanism is introduced to realize deep fusion of the feature level, so that the problems of single information utilization and weak feature relevance of the existing detection means are solved, and the comprehensiveness of defect detection and the accuracy of fault diagnosis are improved.
Owner:SUN YAT SEN UNIV

Three-dimensional model compression transmission method and system based on dynamic feature perception

The invention relates to a three-dimensional model compression transmission method and system based on dynamic feature perception, and relates to the technical field of three-dimensional live-action modeling, and the method comprises the steps: obtaining the original three-dimensional data of a three-dimensional model, carrying out the multi-modal feature extraction based on a model scene, dividing a key region, a transition region and a non-key region, and generating a three-dimensional model partition map, compressing each region according to the partition map label and a preset compression algorithm to generate a multi-resolution LOD sequence, determining a transmission data hierarchy in combination with a network state and equipment parameters, and finally rendering the model according to the transmission data hierarchy, user behavior information and an environment state to obtain a target three-dimensional model. The technical effects of effectively compressing, transmitting and rendering the three-dimensional model according to various factors such as different area characteristics, network states, equipment parameters and user behaviors of the model and improving the transmission efficiency and the rendering quality are achieved.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD

Stacked image enhancement processing method

The invention provides a stack image enhancement processing method, and relates to the technical field of image enhancement, and the method comprises the steps: collecting and preprocessing a stack image, and constructing a training sample set containing a risk level and an inclination angle label; an encoder-classifier dual-stage model is constructed, an encoder extracts features through a multi-scale feature extraction and enhancement module, a space perception attention enhancement module, a dynamic feature pyramid fusion module and a deformation gradient enhancement residual module, and a classifier outputs a risk level and an inclination angle; during training, classification and regression loss weights are dynamically adjusted, an adversarial enhancement sample is generated to enhance robustness, and a gradient is adaptively cut based on a channel attention enhancement vector. According to the method, local illumination interference can be effectively eliminated, multi-scale features are fused, the sensitivity to a sparse risk area is enhanced, and the self-adaptive enhancement processing effect of the heaped image in a complex environment is improved.
Owner:SUZHOU UNIV +1

Robust unmanned aerial vehicle detection method based on dynamic feature fusion and context attention

The invention relates to a robust unmanned aerial vehicle detection method based on dynamic feature fusion and context attention, and belongs to the technical field of image processing. Aiming at the problems of small target feature loss, semantic gap, background noise interference and the like caused by a fixed convolution kernel scale, one-way feature fusion and a static attention mechanism in an existing unmanned aerial vehicle aerial image target detection method, the method comprises the following steps: constructing a detection model comprising a backbone network, a neck network and a detection head network; a feature rearrangement and extraction module is designed in the backbone network to enhance feature learning, an enhanced double-flow feature fusion pyramid is designed in the neck network to optimize multi-scale feature fusion, and a dynamic multi-scale context attention mechanism is designed in the detection head network to suppress irrelevant background noise. The method effectively improves the accuracy and robustness of small target detection, and achieves a clearer and more stable detection effect in a complex environment.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Industrial T-MOD module working state visualization method and system

The invention provides an industrial T-MOD module working state visualization method and system, and the method comprises the steps: carrying out the time dimension preprocessing and multi-dimensional feature analysis of original operation data collected by a T-MOD module, and generating a multi-dimensional dynamic feature vector, the method comprises the following steps: acquiring a multi-dimensional dynamic feature vector of a T-MOD module, constructing a digital twin three-dimensional model corresponding to the T-MOD module according to the multi-dimensional dynamic feature vector, inputting the multi-dimensional dynamic feature vector into a pre-trained large language model, performing semantic recognition and potential anomaly analysis on the running state of the T-MOD module, and generating structured analysis data; and generating a three-dimensional dynamic visual interface in the digital twin three-dimensional model according to the structured analysis data. According to the scheme, the working state of each component of the T-MOD module can be detected in real time through the three-dimensional dynamic visual interface, the manual inspection cost is effectively reduced, and the fault detection efficiency is improved.
Owner:GUANGZHOU SPECIAL CONTROL ELECTRONIC IND CO LTD

Image recognition security processing method and system for security screen

The invention discloses an image recognition security and protection processing method and system for a security and protection screen, relates to the technical field of security and protection screens, solves the problem that original object contour feature processing is rough, and can effectively lock a dynamic object by analyzing pixel point differences between adjacent frame images acquired by a high-definition camera. When the dynamic features are confirmed, difference regions and contour regions are accurately divided through point location analysis and gradient calculation of the system, and the feature region with the largest area is selected from a plurality of feature regions to serve as the dynamic features, so that the recognition accuracy is ensured, and the misjudgment rate is reduced; in each contour partition, feature difference values of different internal connecting lines are compared, and the connecting line with the minimum difference value is selected as an average refining connecting line, so that refining processing of the dynamic feature edge contour is realized, the extracted feature contour is more accurate, and the real form of an object can be better reflected.
Owner:NANTONG RUIXIN LIGHT TECHNOLOGY CO LTD

Multi-moving-target tracking method for remote sensing video and storage medium

The invention relates to a multi-moving-target tracking method for a remote sensing video and a storage medium. The method comprises the following steps: firstly, establishing a plurality of detection subsequences based on multiple frames of images of an acquired remote sensing video, then respectively extracting a space-time fusion feature containing a static feature and a dynamic feature of each detection subsequence, and obtaining a detection frame of each frame of image based on the space-time fusion feature; then, according to the target library information of each frame, prediction states of multiple frames of images are obtained through Kalman filtering in sequence, and an optical flow matrix of two adjacent frames of images is calculated based on background feature points of the two adjacent frames of images; based on the optical flow matrix, the prediction state of each frame of image can be corrected and matched with the detection frame of each frame of image, and a first target library sequence composed of the target libraries of each frame of image is obtained. And finally, performing post-processing operation of interpolation reconstruction and static target elimination on the first target library sequence based on the optical flow matrix to obtain a second target library sequence. Therefore, the multi-moving-target tracking precision of the remote sensing video can be improved.
Owner:BEIJING INST OF TECH

Video fact and viewpoint alignment traceability method

The invention discloses a video fact and viewpoint alignment traceability method, which relates to the technical field of information retrieval and verification, and comprises the following steps of: performing frame analysis on a video by using a computer vision technology, extracting scenes, objects, dynamic characteristics and background information in the video, extracting audios from the video by using a voice recognition technology, converting the audios into text information, and storing the text information in a database; utilizing an event extraction technology to extract event elements from the video; the event elements are identified through scenes, objects, dynamic features and background information, and event facts are obtained; and based on the text information, using a natural language processing technology to calculate semantic similarity between the viewpoints and the event facts in the video, and automatically aligning the viewpoints and the event facts in the video according to the semantic similarity. According to the method, high-dimensional semantic modeling is carried out on the text, the similarity is calculated, the internal relation between viewpoints and facts can be accurately recognized, and therefore automatic matching of the viewpoints and the facts is achieved.
Owner:CHONGQING QINGZHI NET EAGLE TECHNOLOGY CO LTD

Potential diffusion model-based cinema video score generation and style control method

The invention discloses a potential diffusion model-based cinema video score generation and style control method, which comprises the following steps of: firstly, pre-training a basic video-to-music generation model by utilizing a potential diffusion model architecture; then extracting semantic features, aesthetic features and emotional features from the movie video to obtain fused visual features; and finally, adding a style control module into a basic video-to-music generation model, taking melody and dynamic characteristics of music as local control as input of the style control module in a training stage, fusing visual characteristics, and sending the fused visual characteristics into a cross attention layer for global control, so as to obtain a music model. And sending the compressed representation of the music Mel spectrum to a pre-trained basic video to a music generation model for music modeling. In this way, local control (melody and dynamic control of music) and global control (semantic features, emotional features and aesthetic features of movie video clips) are utilized simultaneously, so that flexible and style-controllable movie game can be performed.
Owner:TIANJIN UNIVERSITY OF TECHNOLOGY

Audio and video object intelligent tracking optimization method and system combined with deep learning

The invention relates to the technical field of audio and video processing, and provides an audio and video object intelligent tracking optimization method and system combined with deep learning. The method comprises the following steps: performing cross-modal feature collaborative extraction on an audio stream and a video frame sequence by acquiring a synchronous audio and video data group, and generating a multi-modal feature set containing audio time domain dynamic features and video space structure features; inputting the multi-modal feature set into a pre-trained association enhancement network to generate a cross-modal semantic aligned association feature sequence; constructing a tracking stability evaluation model based on the associated feature sequence, and outputting a stability index; tracking parameters are dynamically adjusted according to the stability index, an initial tracking result is calibrated, and an optimized tracking trajectory is output. Therefore, the precision and stability of object tracking in a complex scene are improved by deeply fusing the dual-mode characteristics of the audio and the video, mining the internal association between the modes and combining a dynamic evaluation and calibration mechanism.
Owner:SHENZHEN ZIDOO TECH CO LTD

Industrial instrument panel target detection method and system based on optical flow characteristics and YOLOv8

The invention discloses an industrial instrument panel target detection method and system based on optical flow features and YOLOv8, and relates to the technical field of image processing. The method is an industrial dashboard category detection method suitable for a dynamic video scene, the core is to combine optical flow features and a YOLOv8 target detection model, a Transform module is introduced to perform optical flow feature extraction, correlation softmax is utilized to capture inter-frame motion correlation, dynamic features of a target area are highlighted in combination with a self-attention mechanism, and the target area is subjected to target detection. And meanwhile, background interference is suppressed, so that the optical flow features are more accurate and reliable. Through the structure optimization of the YOLOv8 model, the optical flow features are deeply fused into each key module of the detection network, the detection precision, robustness and real-time performance of the instrument panel category in a dynamic video scene are significantly improved, and an efficient and reliable solution is provided for dynamic target identification in a complex industrial environment.
Owner:NORTHEASTERN UNIV CHINA

Unified visual language model backdoor attack method based on feature hijacking

The invention discloses a feature hijacking-based unified visual language model backdoor attack method. The method comprises the following steps of: obtaining a data set and a unified visual language model; initializing back door modules such as a multi-mode trigger, a trigger detector and a dynamic feature alignment module; constructing a harmful data set, randomly selecting a part of samples from the training set, injecting a multi-mode trigger into the samples, generating poisoning samples, and mixing original training data and generated poisoning sample data to generate the harmful data set; training the model by using a harmful data set, freezing original parameters of the model, and only allowing the backdoor module to participate in training; a poisoning model generated through model reasoning and training is normally expressed on a benign test sample, but when a text and an image backdoor trigger exist at the same time, the model outputs a preset answer, and backdoor attack is achieved. According to the method provided by the invention, the problem that the attack effect on the unified visual language model is insufficient due to single-mode triggering or insufficient feature disturbance is solved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Energy efficiency evaluation method of artificial intelligence data center

The invention belongs to the technical field of data centers, discloses an energy efficiency evaluation method of an artificial intelligence data center, and solves the problem that a static evaluation method is disjointed with the real-time operation state of the data center in the prior art by constructing a dynamic feature library and an AI adaptive weight model. The dynamic feature library collects and associates service scene labels and time-space coupling data in real time, and provides a scene basis for an AI adaptive weight model, so that weight adjustment is converted from passive response index mutation in the prior art to active adaptive service scene change. Nonlinear influences of dynamic factors such as server load fluctuation and environment adjustment on energy efficiency are effectively captured, so that an energy efficiency evaluation result better fits the actual operation state of the data center, energy efficiency factor relevance changes caused by service switching of the AI data center are adapted, scene-based self-adaptive adjustment of weights is achieved through a deep learning model, and energy efficiency evaluation accuracy is improved. The limitation that the fixed weight distribution ignores the difference of different service scenes in the prior art is solved.
Owner:BEIJING QINGYUAN CHUANGYAN TECHNOLOGY CO LTD