Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

4153 results about "Video based" patented technology

Power operator behavior identification early warning system and method based on video analysis

The invention discloses a video analysis-based electric power operation personnel behavior identification and early warning system and method, which realize accurate identification and real-time early warning of electric power operation personnel behaviors by combining a video analysis technology with multi-modal data fusion, and effectively improve the safety management level of an operation site. Compared with a traditional safety supervision mode, the method employs a mode of combining deep learning target detection and time sequence behavior analysis, improves the recognition accuracy of operators and safety equipment, fuses the data of equipment worn by the operators with video data, improves the detection precision, and improves the safety supervision accuracy. And misjudgment caused by illumination change, shielding or complex environment is reduced. Besides, high-risk violation behaviors such as no safety helmet wearing, no safety belt fastening, violation climbing, high-altitude object throwing and the like are accurately recognized through the violation detection module, and different levels of alarm measures are adopted according to the severity of the violation behaviors in combination with an early warning feedback mechanism, so that the pertinence and response efficiency of early warning are improved.
Owner:PENGLAI WIND POWER BRANCH OF HUANENG SHANDONG POWER GENERATION CO LTD +1

Video feature extraction and multi-dimensional matching-based movie advertisement real-time pushing system

The invention discloses a video feature extraction and multi-dimensional matching-based movie advertisement real-time pushing system, and belongs to the technical field of digital advertisements. According to the system, multi-modal analysis is carried out on visual, audio, text and semantic features of a movie through a video feature extraction module, a dynamic interest tag constructed by a user portrait analysis module is combined, and the weighted matching degree of a video scene, user preference and an advertisement tag is calculated by utilizing a multi-dimensional matching engine; and millisecond-level advertisement putting is completed through the real-time pushing decision module. According to the method, context awareness and streaming computing technologies are fused, the relevance between advertisements and content scenes is remarkably improved, cross-platform deployment is supported, the method is suitable for short videos, live broadcast and other real-time scenes, the advertisement click rate is increased by 55%-72% through tests, the response delay is lower than 200 ms, and the method has the technical advantages of being efficient, accurate and low in delay.
Owner:BEIJING QICHUANG TECH CO LTD

Underground mine operation state analysis system and method based on video monitoring data

The invention discloses an underground mine operation state analysis system and method based on video monitoring data, and the system comprises a data collection module which is used for collecting mine video and environment parameter data through a distributed sensor network, and generating a multi-dimensional data fusion set based on a space-time label technology; the edge analysis module is used for extracting feature parameters through a convolutional neural network algorithm based on the multi-dimensional data fusion set and generating a mine operation state recognition result; the fence construction module is used for constructing a three-dimensional digital model and a dynamic safety boundary based on the mine operation state recognition result to form a real-time monitoring reference framework; and the decision execution module is used for performing hierarchical risk assessment on the monitoring data in the security boundary based on the real-time monitoring reference framework, and generating a security early warning and disposal scheme with a tracing identifier. Each piece of early warning and disposal information is attached with a unique tracing identification code, so that follow-up event backtracking analysis is facilitated, and the risk management and control capability is continuously improved.
Owner:河北省水文工程地质勘查院(河北省遥感中心) +3

Action localization method, device, electronic equipment, and computer-readable storage medium

An action localization method, device, electronic equipment, and computer-readable storage medium are provided. The action localization method includes: identifying at least one target video segment containing a target object in a video; acquiring a first action recognition result of at least one image frame in the at least one target video segment and a second action recognition result of the target video segment; and acquiring an action localization result of the video based on the first action recognition result and the second action recognition result.
Owner:SAMSUNG ELECTRONICS CO LTD

Multi-modal diffusion-based long video role scene decoupling generation method and system

The invention discloses a long video role scene decoupling generation method and system based on multi-modal diffusion, and relates to the technical field of image processing, and the method comprises the steps: S1, synthesizing the advanced features of a role and a scene through a SigLIP encoder and a DINOv2 encoder; s2, performing cross-modal feature fusion on the advanced features to obtain joint features, and compressing the joint features to obtain compact vectors; s3, generating text features according to the text prompt; s4, potential codes are generated from an input video through a causal 3D convolution encoder, the potential codes pass through a linear projection matrix and then are spliced with a memory state for dimension reduction, and a segmented potential vector sequence is obtained; s5, performing decoupling perception generation on the segmented potential vector sequence through an improved 3D-UNet, and performing deconvolution up-sampling reconstruction after deterministic sampling to obtain an RGB video segmented sequence; according to the method, the key problems of rough dynamic control, limited generation length and over-high resource consumption in long video generation are solved, and the quality and efficiency of the generated video are remarkably improved.
Owner:湖南马栏山视频先进技术研究院有限公司

Video analysis-based multi-scene operator violation behavior identification method and system

The invention discloses a video analysis-based multi-scene operator violation behavior identification method and system, and belongs to the technical field of intelligent operation safety monitoring and artificial intelligence identification, and the method comprises the steps: collecting a real-time video stream of a multi-scene operation site; recognizing a continuous action time sequence in the real-time video stream by using an action recognition depth model; constructing the continuous action time sequence into an action behavior sequence; the action behavior sequence is constructed into a directed behavior graph with time, space and action labels, the directed behavior graph is compared with a directed behavior graph corresponding to the standard action behavior sequence, and illegal behaviors are recognized; and carrying out multi-mode early warning on the identified illegal behaviors. According to the method, the bottleneck that the traditional image recognition technology is weak in action sequence semantic understanding and poor in environmental adaptability is broken through, and accurate recognition and real-time early warning of illegal behaviors in multi-scene operation are achieved.
Owner:CHENGDU HANGTIAN PHOTOELECTRIC TECH

Real-time video translation and audio and picture synchronization method and system based on multi-modal large model

The invention provides a real-time video translation and audio and picture synchronization method and system based on a multi-modal large model, and relates to the technical field of video translations, and the method comprises the steps: obtaining a source video; extracting the source video based on the multi-modal large model to obtain a multi-modal feature; fusing the multi-modal features through a cross-modal attention mechanism to generate a context semantic vector; translating into a target language text in real time based on the context semantic vector, and processing the translated language text based on the multi-modal features to obtain a translated language sound source; and performing mouth shape adjustment on the source video based on the translation language sound source, and merging the translation language sound source and the mouth shape animation video to obtain a real-time translation video with synchronous sound and picture. According to the method, the limitation of traditional single-modal translation is broken through, and the semantic accuracy of translation is remarkably improved by dynamically aligning the context information through the multi-modal features in combination with a cross-modal attention mechanism.
Owner:SHANGHAI YINGZHUO INFORMATION TECH CO LTD

Systems and methods for multimodal indexing of video using machine learning

Systems, methods, and computer-readable media are disclosed for systems and methods multimodal indexing of video using machine learning. An example method may include deceiving, by a video encoder of an audio-video transformer neural network comprising one or more computer processors coupled to memory, a first frame and a second frame associated with a first segment of a video. The example method may also include receiving, by an audio encoder of the audio-video transformer neural network, an audio spectrogram comprising first audio data associated with the first segment of the video. generating, by the video encoder, a first video embedding. The example method may also include generating, by the audio encoder, a first audio embedding. The example method may also include determining a fusion of the first video embedding and the first audio embedding using a multimodal bottleneck token. The example method may also include determining an output including the first video embedding and the first audio embedding. The example method may also include determining a classification of the first portion of the video based on the output.
Owner:AMAZON TECH INC

Virtual fitting method and device, storage medium and electronic equipment

The invention discloses a virtual fitting method and device, a storage medium and electronic equipment, and is applied to the technical field of computer image processing, and the method comprises the steps: processing a target video, and obtaining the human body posture data of each video frame in the target video; constructing a 3D human body model of each video frame by using the human body posture data of each video frame; for the 3D human body model of each video frame, acquiring clothing rendering information of the 3D human body model, and rendering virtual clothing on the 3D human body model based on a dynamic fitting algorithm and the clothing rendering information to obtain a virtual rendering video frame of the video frame; and generating an AR fitting video based on each virtual rendering video frame, and displaying the AR fitting video to the user. Therefore, the selected costume can be tried on in the AR form, the fitting effect of the costume can be displayed for the user, the problem that the size is improper due to the fact that the costume purchased online cannot be tried on is solved, the probability of refunding and changing goods is reduced, and good shopping experience is provided for the user.
Owner:小芒电子商务有限责任公司

Short video network public opinion information identification method based on image processing technology

The invention discloses a short video network public opinion information identification method based on an image processing technology, and relates to the technical field of artificial intelligence and image processing, and the method comprises the following steps: S001, through obtaining image frames, audio tracks and time sequence information of a short video, constructing a multi-modal fusion model, extracting continuous image frames with suspicious identity features, and carrying out the recognition of the short video network public opinion information; generating a forgery risk area distribution map; and S002, performing semantic consistency verification according to the counterfeit risk region distribution map, and extracting space and time anomaly features existing among facial micro-expressions, pronunciation actions and background semantics in the image frame. According to the method, a multi-modal model is constructed by fusing image, audio and time information, fine abnormal features, traceability forgery starting points and propagation paths in a deep forgery video are identified, and an identification strategy and a public opinion response mechanism are dynamically adjusted, so that accurate identification, adaptive processing and closed-loop control of short video public opinion risks are realized; and the identification accuracy and the treatment efficiency are improved.
Owner:TIBET UNIV

Semantic analysis of video data for event detection and validation

Methods, systems, and computer programs are presented to perform semantic analysis of video data for event detection and validation. One method includes an operation for training a machine-learning model with text-image pairs to create a semantic model. The semantic model generates text embeddings based on text describing events and image embeddings from video frames. The system calculates similarity values between text and image embeddings to determine event occurrences when the similarity exceeds a predetermined threshold. The system enables precise event detection by analyzing object relationships and attributes within video frames, reducing false alarms and enhancing monitoring efficiency. Further, these techniques can be used to detect any event describable by human text or by some image examples.
Owner:FIRST-CITIZENS BANK & TRUST CO

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Panoramic video three-dimensional point cloud sparse reconstruction method and device integrating key frame screening

The invention provides a panoramic video three-dimensional point cloud sparse reconstruction method and device integrated with key frame screening, and the method comprises the steps: obtaining a panoramic video of a target scene, and carrying out the video quality evaluation of the panoramic video, and obtaining a video evaluation result; determining a target resolution of downsampling according to the video evaluation result; carrying out downsampling processing on the panoramic video based on the target resolution to obtain a downsampled second video; for the second video, dynamically adjusting a frame extraction interval based on a preset key frame selection mechanism, and extracting a frame index of the key frame from the second video according to the frame extraction interval to obtain a plurality of frame indexes; and obtaining a plurality of corresponding key frames in the panoramic video based on the plurality of frame indexes, and obtaining the three-dimensional scene sparse point cloud of the target scene based on the plurality of key frames, thereby solving the technical problem of low processing efficiency in the prior art.
Owner:TIANJIN FIRE SCI & TECH RES INST OF MEM

Home decoration construction site monitoring system and method based on artificial intelligence

The invention discloses a home decoration construction site monitoring system and method based on artificial intelligence, and relates to the technical field of video processing. A video data set, an environment parameter set and an equipment state set in a historical construction process are collected in advance, an image semantic segmentation model is constructed based on the video data set, and construction scene division is performed; constructing a target detection model based on the video data set, generating a target spatial-temporal trajectory based on a detection result of the target detection model by using a multi-target tracking algorithm, and constructing a construction behavior recognition model based on the video data set and the target spatial-temporal trajectory; based on the recognition result of the construction behavior recognition model, the environmental parameter set and the equipment state set, constructing a risk early warning model; the safety condition of a construction site is evaluated in real time, and dynamic risk early warning is provided.
Owner:JIANGSU ZHONGBANG JIANTONG TECH CO LTD

Non-inductive identity recognition and real-time tracking method based on personnel in video

The invention discloses a non-inductive identity recognition and real-time tracking method based on personnel in a video. The method comprises the following steps: S1, acquiring video stream data and extracting a key frame image; s2, key point coordinates are extracted, and a time sequence skeleton sequence is constructed; s3, decomposing gait, action and attitude features to generate standardized features; s4, inputting the standardized features into an improved multi-scale image convolutional neural network, extracting identity representation features, and generating a target identity code; s5, matching identity codes based on similarity measurement, calculating a consistency score, and establishing an identity tracking trajectory; s6, tracking is carried out in combination with identity codes and consistency scores, identity mapping is dynamically updated, and non-inductive recognition is achieved; and S7, updating a track in real time, adjusting a tracking state, and ensuring identity stability. According to the method, the improved graph convolutional network and the consistency measurement technology are combined, the identity recognition precision and tracking stability are improved, and cross-scene non-inductive identity recognition is achieved.
Owner:BEIJING LIYANG ZHIGUANG TECH CO LTD

Content item video generation template

Methods and systems are disclosed for generating video by applying a template to various content items. The methods and systems select, by an interaction application, a video generation template comprising instructions for combining a set of content items into a video using one or more augmented reality (AR) elements. The methods and systems identify a subset of content items from a plurality of previously captured content items and modify one or more content items of the identified subset of content items based on the AR elements of the video generation template. The methods and systems generate a video comprising a collection of content items including the identified subset of content items and the modified one or more content items based on the instructions of the video generation template.
Owner:SNAP INC

Video generation method and device based on multi-agent cooperation and agents

The invention provides a video generation method and device based on multi-agent cooperation and agents, relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, large models, agents, AIGC, element universe and the like, and is applied to application scenes such as wisdom education, video production, animation production and the like. The method comprises the following steps: executing a target sub-task in a target task by utilizing at least one target agent to obtain a target sub-execution result; and generating a target video based on the target sub-execution result, the target agent executing the following operations: obtaining sub-task demand information for the target sub-task; determining a subtask execution element for controlling the execution process of the target subtask by using a first subagent of the target agent based on the subtask demand information; and utilizing a second sub-agent of the target agent to execute the target sub-task based on the sub-task execution element to obtain a target sub-execution result.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Commodity video intelligent generation method based on multi-modal analysis and dynamic narrative architecture

The invention relates to the technical field of video generation, in particular to an intelligent commodity video generation method based on multi-modal analysis and a dynamic narrative architecture, which comprises the following steps: inputting original commodity data and a user behavior log, and outputting a dynamic selling point weight vector through a selling point value evaluation engine; s2, inputting the dynamic selling point weight vector output in S1 into a narrative flow generator, and executing the following steps: intercepting first K core selling points according to the weight vector to form a narrative trunk chain; based on the historical preference data of the user, inserting an emotion enhancement node to generate a branch enhancement narrative flow; generating a multi-version narrative flow instruction set in combination with the real-time network bandwidth data; retrieving a preset video clip library according to the instruction set to generate an initial video sequence; detecting semantic fault regions between adjacent segments to generate compensation animation parameters; and injecting compensation animation parameters and outputting a continuous narrative video stream. According to the invention, the jamming feeling and the frame skipping phenomenon caused by the change of a narrative structure or different sources of video clips are reduced, and the overall watching experience of a user is improved.
Owner:SHENZHEN YINGMENG INTELLIGENT TECHNOLOGY CO LTD

Safety production behavior monitoring method and system based on AI video analysis

The invention provides a safety production behavior monitoring method and system based on AI video analysis. The method comprises the steps of collecting a real-time video data stream of a production area; inputting each frame of video image in the real-time video data stream into a target detection model for target detection to obtain a personnel target output by the target detection model and a target position coordinate of the personnel target in each frame of video image; cutting out a local image area of the personnel target in each frame of video image based on the target position coordinate, and performing feature recognition based on the local image area to obtain personnel features; performing comparison on the basis of the personnel characteristics and the personnel standard behavior characteristics to obtain personnel behavior states, and performing track association on the basis of the personnel behavior states corresponding to the continuous multi-frame video images to obtain personnel behavior tracks; and performing safety production behavior monitoring based on the personnel behavior state and the personnel behavior track, and generating an abnormal behavior early warning signal. According to the method and the device, the real-time performance and the accuracy of safety monitoring in a production scene are improved.
Owner:SHENZHEN YINXING INTELLIGENT DATA CO LTD

Behavior analysis method based on video recognition, processor and storage medium

The invention discloses a behavior analysis method based on video recognition, a processor and a storage medium, and belongs to the technical field of data processing, and the method comprises the following steps: obtaining an image and a video clip with an operation behavior in an electronic manufacturing process; performing static behavior recognition on the image through a static behavior recognition model in the fusion model, and extracting potential static illegal behaviors in the image; extracting a skeleton point sequence from the video clip through a dynamic behavior recognition model in the fusion model so as to perform dynamic behavior recognition, and extracting potential dynamic illegal behaviors in the video clip; and outputting a behavior category code according to the static violation behavior and the dynamic violation behavior. According to the behavior analysis method based on video recognition, the processor and the storage medium, the problems that in an existing student operation behavior analysis mode, behaviors in operation are difficult to comprehensively capture in real time, and scoring objectivity is difficult to guarantee are solved.
Owner:广州全域科技有限公司

Video quality detection method based on multi-scene self-adaption

The invention provides a video quality detection method based on multi-scene self-adaption. The method comprises the following steps: selecting an area in a video picture, extracting multi-dimensional features of the area, determining the type of a scene where the video picture is located, and dynamically adjusting video detection parameters of the video picture; detecting the quality of the video picture based on the video detection parameters; intercepting an abnormal video picture, and adding the abnormal video picture into the constructed equipment fault template library M; and performing local model retraining on the video quality detection model based on the equipment fault template library M, and detecting the quality of a to-be-detected video picture based on the video quality detection model. On the basis of deep analysis of different scene features, the scene type of the video picture is accurately judged through multi-dimensional feature calculation, and then the video quality detection parameters are dynamically adjusted according to the scene type, so that the detection system can automatically adapt to the optimal detection parameters according to scene changes, misjudgment and missed judgment caused by the scene changes are avoided, and the detection efficiency is improved. And the accuracy of video quality detection in different scenes is remarkably improved.
Owner:CHINA SHIPBUILDING LINGJIU HIGH TECH (WUHAN) CO LTD +1

Virtual object animation generation method and apparatus, electronic device, computer-readable storage medium, and computer program product

A virtual object animation generation method and apparatus including obtaining a to-be-processed video, determining initial 3D pose information, two-dimensional (2D) pose information, and foot ground contact information of a target object in the to-be-processed video, determining pose change information between two adjacent video frames in the to-be-processed video based on the initial 3D pose information of the target object, determining to-be-corrected pose information of the target object based on the initial 3D pose information of the target object and the pose change information, performing correction processing on the to-be-corrected pose information of the target object based on the 2D pose information and the foot ground contact information of the target object, to obtain corrected pose information of the target object, and retargeting the corrected pose information to a virtual object by means of motion retargeting, to generate a virtual object animation video corresponding to the to-be-processed video.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Generating speaker video and audio in multiple languages for videoconferencing

Systems and methods for generating speaker video and audio in multiple languages for videoconferencing are provided. For example, a computing device can access a speaker speech audio signal that includes a speaker speech in a first language, a video of the speaker and a translated speech audio signal of the speaker speech in a second language. The computing device generates, based on the translated speech audio signal, a converted translated speech audio signal that includes a speech in the second language having voice characteristics in the speaker speech. The computing device further generates a lip-synched speaker video based on the video of the speaker and the converted translated speech audio signal. Lip movements in the lip-synched speaker video correspond to the converted translated speech audio signal. The converted translated speech audio signal and the lip-synched speaker video are transmitted to a video conference provider configured to host the video conference.
Owner:ZOOM VIDEO COMM INC

Intelligent driving behavior identification method and system based on video analysis

The invention provides a driving behavior intelligent identification method and system based on video analysis, and the method comprises the steps: obtaining a driver face video stream and a road environment video stream collected by a vehicle-mounted camera, and reading the driving information recorded by a whole vehicle communication network; recognizing an eyelid closing state, a sight line direction and a head posture in the driver face video stream based on a posture recognition model, and performing fatigue distraction analysis to obtain driver state information; performing motion trail analysis on the road environment video stream and the driving information, and performing driving risk assessment in combination with the driver state information to obtain driving assessment information; and performing early warning construction according to the driving evaluation information, generating early warning prompt information, and synchronously writing the early warning prompt information, the driving evaluation information and the driver state information into a safety data protection memory. The fatigue and distraction states of the driver can be recognized more accurately, and the accuracy of state judgment is improved.
Owner:SHENZHEN ZHIJU CLOUD SERVICE TECH CO LTD

Interactive automatic explanation method for converting traditional video into artificial intelligence digital human

The invention provides an interactive automatic explanation method for converting a traditional video into an artificial intelligence digital human, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining original video data and an audio track, and carrying out the semantic analysis of the audio track, and obtaining multi-mode deconstruction data; generating an explanation script for each time period of the video based on the explanation text, and performing timestamp labeling on the visual elements to form a time sequence synchronization data structure; in the playing process, a virtual image generator is driven to synthesize digital human dynamic expression output in real time according to the current playing time point; after a user interruption request is received, semantic matching is carried out on a query intention in the explanation script, a target explanation fragment and visual elements are positioned, and complementary explanation content is generated; and driving the virtual image generator to synthesize dynamic output synchronized with the supplementary explanation, and after interaction is completed, recovering playing or skipping to a specified time point according to a user instruction. According to the invention, the conversion from the traditional video to the interactive intelligent explanation video is realized, and the watching experience and learning efficiency of the user are improved.
Owner:BEIJING MENGKE TECH CO LTD

Virtual fitting video generation method and system based on multi-view face fixation

The invention discloses a virtual fitting video generation method and system based on multi-view face fixation, and relates to the field of generative artificial intelligence and computer vision, and the virtual fitting video generation method based on explicit geometric constraints comprises the following steps: S1, obtaining an original fitting image and an action cue word, and constructing a multi-source input data set; s2, segmenting the original fitting image to obtain multi-view modeling, and generating a multi-angle face image and a clothing texture feature parameter; s3, analyzing the action cue word to generate a target posture sequence, and generating head and tail frame virtual fitting images; s4, performing pairing analysis on the virtual fitting images of the head frame and the tail frame to generate attitude transition parameters; s5, performing dynamic texture repair on the image to generate a transition frame image sequence; and S6, generating a virtual fitting video and outputting the virtual fitting video. According to the method, accurate segmentation and multi-angle modeling of the face area and the clothing area are realized through image segmentation and a rigid geometric transformation algorithm.
Owner:QINGDAO UNIV

Al-based video content analysis method and system

The invention relates to the technical field of content recognition, in particular to an Al-based video content analysis method and system, and the method comprises the following steps: carrying out the image segmentation learning based on a video sequence through employing a U-Net convolutional neural network, analyzing scenes and elements in a video frame, recognizing and isolating key visual elements in a video through network learning, and carrying out the recognition of the key visual elements in the video. Comprising objects and figures. According to the method, image segmentation learning is carried out by adopting the U-Net convolutional neural network, key visual elements in the video can be identified and isolated more accurately, a clearer basis is provided for follow-up scene change and key event tracking, video content analysis is carried out by applying the graph neural network, the identification capability of dynamic scenes and events is enhanced, and the identification efficiency is improved. Deeper structured understanding is provided for video content indexing, the video quality is analyzed through a structural similarity index evaluation method, and the reason of visual quality reduction can be accurately recognized and improved.
Owner:CHENDA (GUANGZHOU) NETWORK TECH CO LTD +1

Flow monitoring method and system based on video image analysis and processing

The invention relates to the technical field of water flow monitoring, and particularly discloses a flow monitoring method and system based on video image analysis and processing. The method comprises the following steps: carrying out communication acquisition of image monitoring and water level monitoring at a plurality of monitoring acquisition sites; performing video frame extraction and image preprocessing; feature point detection and displacement analysis; converting the water surface flow velocity into a plurality of station water surface flow velocities, and performing fitting correction; and calculating station flow data of the plurality of monitoring acquisition stations. Water surface image data are obtained through image monitoring and water level monitoring communication collection of a plurality of monitoring collection sites in combination with video frame extraction and image preprocessing, the site pixel flow velocity is calculated through feature point detection and displacement analysis, the site pixel flow velocity is converted into corrected water surface flow velocity through a space projection model, and site flow data are calculated. Non-contact full-flow monitoring is achieved, the potential safety hazard that equipment needs to be arranged in water in a traditional method is eliminated, the maintenance difficulty is remarkably reduced, and the method is particularly suitable for flow monitoring in flood periods and dangerous water areas.
Owner:JIANGXI SHANLIU HUILIAN TECHNOLOGY CO LTD

Dense video description method based on multi-modal memory knowledge

The invention relates to the field of video description, in particular to a dense video description method based on multi-modal memory knowledge, which comprises the following steps: extracting visual features and audio features of an input video and carrying out cross-modal fusion to generate a final audio code and a final visual code; determining event visual features and event audio features of a plurality of candidate events from the input video based on the final audio code and the final visual code; for each candidate event, retrieving matched external knowledge from an external memory knowledge base based on the corresponding event visual feature and event audio feature, and generating corresponding multi-modal external memory knowledge; based on the multi-mode external memory knowledge, the event visual features and the event audio features of each candidate event, a word embedding sequence is constructed step by step through an autoregression mechanism, and description of the input video is generated. According to the method, the corresponding relation between the event and the description can be learned from more comprehensive information, and the accuracy and richness of generating the description are remarkably improved.
Owner:JIAXING UNIV

Key frame extraction method and device based on dynamic reinforcement learning, equipment and medium

The invention relates to the technical field of computer vision, can be applied to the medical field and the financial science and technology field, and discloses a key frame extraction method, device and equipment based on dynamic reinforcement learning and a medium, which are applied to electronic application in a high-frequency transaction abnormal behavior monitoring scene or can be applied to a medical operation key frame extraction scene. The method comprises the steps of obtaining an original video stream and performing preprocessing to generate a standardized video frame; performing feature extraction and feature splicing on the standardized video frame, and performing time sequence modeling on the generated pair frame-level mixed feature vector to generate video-level time sequence representation; generating an enhancement action instruction based on the video-level time sequence representation through the strategy network, and performing enhancement processing on the standardized video frame according to the enhancement action instruction to generate an enhanced video frame; performing optimization processing on the strategy network according to the enhanced video frame to generate an updated strategy network; and performing feature extraction and optimization on the enhanced video frame to generate a target key frame. According to the invention, the key frame extraction precision is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD