Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41 results about "Video content analysis" patented technology

Video content analysis (also video content analytics, VCA) is the capability of automatically analyzing video to detect and determine temporal and spatial events. This technical capability is used in a wide range of domains including entertainment, health-care, retail, automotive, transport, home automation, flame and smoke detection, safety and security. The algorithms can be implemented as software on general purpose machines, or as hardware in specialized video processing units.

Video scene content label determination method based on knowledge graph

The invention relates to the technical field of video content analysis, and discloses a video scene content label determination method based on a knowledge graph. The method comprises the following steps: inputting a multi-modal feature sequence formed by visual, audio and text features extracted from a video stream into a knowledge graph inference engine comprising an entity relationship network and a semantic association rule base; performing node mapping through an engine to generate an initial scene entity set; performing hierarchical reasoning on the set based on a semantic association rule base to obtain a scene semantic topological structure; screening core entities according to entity weight distribution in the topological structure, and generating a candidate tag set; performing time sequence consistency verification on the candidate tag set, and correcting tag time sequence offset in combination with timestamp information; performing cross-modal disambiguation on the corrected label set by using an entity relationship network to eliminate semantic conflicts; and according to a disambiguation result, constructing a scene label knowledge sub-graph containing entity attributes and relation path constraints.
Owner:GUANGZHOU JUNHE INFORMATION TECH CO LTD

Automatic movie and television video script extraction method based on multi-modal large model

The invention relates to an automatic movie and television video script extraction method based on a multi-modal large model, and belongs to the field of artificial intelligence and video content analysis. Aiming at the problem that an existing automatic speech recognition tool cannot generate a structured script, the method comprises the following steps: firstly, carrying out multi-mode decomposition on an input video, extracting frames by adopting a scene self-adaptive strategy, and establishing sound and picture timestamp alignment mapping; then extracting features through a CLIP-ViT visual feature encoder and a Whisper audio feature encoder, and associating semantics by using a cross-modal attention mechanism; role identity recognition, scene type recognition and action character description are achieved, and finally a script file conforming to the standardization specification is generated and comprises a scene title, a time code mark and a special narrative mark. Compared with a mode of manually dictating marks and an automatic voice recognition tool, the method effectively improves the accuracy and efficiency of movie and television video script extraction.
Owner:BEIJING INST OF COMP TECH & APPL

Language-driven hour-level traffic video analysis method and system

The invention discloses a language-driven hour-level traffic video analysis method and system, relates to the technical field of video content analysis and retrieval, can generate a video-specific open world detector, can accurately identify and filter fine-grained and composite semantic targets, and improves the accuracy of video analysis. And an hour-level video analysis service driven by open vocabularies and natural languages is really realized. In order to achieve the purpose, the technical scheme of the invention comprises the following steps of: 1) receiving an hour-level traffic video, and positioning a video clip related to natural language query in the video; and 2) automatically constructing an open world target detector in the video clip positioned in the step 1). And step 3) finishing cross-time-period target trajectory generation on the video clip positioned in the step 1). The invention also provides a language-driven hour-level traffic video analysis system for executing the method. The system comprises a video clip positioning module, an open world target detector construction module and a target trajectory extraction module.
Owner:BEIJING INST OF TECH

Mine monitoring video key frame extraction method

PendingCN121459254ACharacter and pattern recognitionCluster algorithmVideo content analysis
The invention relates to the technical field of video content analysis, in particular to a mine surveillance video key frame extraction method, which comprises the following steps: extracting multi-dimensional features of each frame of a candidate frame set, obtaining a fusion value, and obtaining a candidate key frame set; based on an Euclidean distance method, calculating an Euclidean distance between continuous frames of the candidate key frame set, and determining a total clustering number according to a relationship between the Euclidean distance and a preset clustering threshold value; and clustering the frames in the candidate key frame set by using a preset mixed GWO-FCM clustering algorithm and the total clustering number to obtain a key frame set. According to the method, a hybrid clustering algorithm GWO-FCM combining grey wolf optimization and fuzzy C-means is adopted, global search and local optimization capabilities are considered, and the accuracy and representativeness of key frame clustering are ensured. The high-quality extraction of the key frames is realized, the number of redundant frames is obviously reduced, and the compactness and readability of the video abstract are improved.
Owner:XJ GRP CORP

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (DI)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real-world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Intelligent equipment linkage control method and system based on video content analysis

The invention discloses an intelligent equipment linkage control method and system based on video content analysis, and relates to the field of video content analysis, and the intelligent equipment linkage control method based on video content analysis comprises the following steps: S1, obtaining video parameters, and carrying out the preprocessing; s2, pixel motion vectors are extracted, and a time sequence motion field is formed; s3, extracting a pixel dominant motion information set, calculating an average speed value, and generating a time sequence speed parameter set; s4, performing zero crossing point identification on the time sequence speed parameter set, and generating a time sequence displacement control signal; and S5, sending the time sequence displacement control signal to the external intelligent equipment, and driving the external intelligent equipment to generate physical motion with the video content by adopting the displacement signal. According to the method, by introducing preprocessing means such as multi-scale image analysis, resolution compression and graying conversion, key information is efficiently and accurately extracted from the video, and the precision and efficiency of subsequent analysis are improved.
Owner:SHENZHEN WEIAI TECHNOLOGY CO LTD

A video intelligent slicing method based on multi-modal information fusion and semantic constraint

The application discloses a kind of video intelligent slice method based on multi-modal information fusion and semantic constraint, it is related to video content analysis and intelligent editing technical field, it is characterized in that it proposes semantic priority audiovisual collaborative slice scheme, obtains shot interval by shot boundary detection, constructs voice interval and effective word information by voice recognition, adopts voice interval as semantic mask to remove visual cut point in the process of voice expression, and based on fragment length and effective word, content density determination is carried out to long fragment, and further adopts visual re-cutting, audio energy driving and uniform cutting Secondary slicing strategy of division, the method provided by the application solves the problem of semantic fragmentation and content perception loss in the prior art, can obtain video fragment with complete semantics, high content density and appropriate length, and is suitable for various application scenarios such as automatic editing, short video generation and video content retrieval.
Owner:YUELAI INTELLIGENT MEDIA (TIANJIN) TECHNOLOGY CO LTD

Video texture mapping and real-time rendering method and system based on digital twin scene

ActiveCN121458862BResource allocationCharacter and pattern recognitionGraphicsVideo content analysis
The application provides a video texture mapping and real-time rendering method and system based on a digital twin scene, relates to the technical field of digital twin rendering, and comprises the following steps: acquiring a real scene monitoring video stream and performing real-time image content analysis; dividing the texture details of different regions of the video stream according to the features and numerical values of the image analysis and generating region detail identifiers associated with the video content; identifying the dynamic features of the video picture based on the identifiers and generating scene state codes representing the comprehensive picture conditions; dynamically selecting a target engine and corresponding rendering quality parameters from preset heterogeneous rendering engines in combination with the current load state of a graphics processor; and mapping the video stream image data to the surface of a corresponding model of the digital twin scene as dynamic texture to realize efficient rendering of the picture, so that dynamic texture mapping and adaptive real-time rendering based on video content analysis and hardware load in the digital twin scene can be realized.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (Dl)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real -world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Video processing collaboration method and device, equipment and storage medium

ActiveCN116567320BVideo content analysisVideo processing
The application provides a video processing cooperation method and device, equipment and a storage medium. The method comprises the following steps: sending a video rendering capability request and a video content analysis capability request to a terminal device; receiving a video rendering capability response and a video analysis capability response of the terminal device, wherein the video rendering capability response comprises a video rendering capability of the terminal device, and the video analysis capability response comprises a video analysis capability of the terminal device; determining an optimal video rendering cooperation configuration according to the video rendering capability of the terminal device, and determining an optimal video analysis cooperation configuration according to the video analysis capability of the terminal device. Therefore, the idle computing resources of the terminal device can be fully utilized under limited cloud server computing resources, so that a user can have a better cloud game quality experience.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A video generation method, system, device and medium based on a large language model

The application discloses a video generation method, system, device and medium based on a large language model. The method obtains product information input by a user, the product information including the name, description and selling point of the product. The product information is preprocessed, and the preprocessed product information is subjected to semantic information split-screen processing through a large language model to obtain split-screen description information corresponding to the product information. An original video is obtained, and the original video is subjected to video segment splitting to generate a plurality of video segments and picture description information corresponding to the video segments. The split-screen description information and the picture description information are subjected to semantic matching processing, and the video segments with the highest similarity to each split-screen description information are matched. The video segments are spliced according to the sequence of the split-screen description information to generate a complete video. Compared with the related art, the application can automatically and efficiently generate a product video by integrating natural language processing, video content analysis and intelligent matching algorithms.
Owner:GUANGZHOU TAIDONG TECH CO LTD

A network public opinion video content analysis method based on multi-modal fusion

The application provides a network public opinion video content analysis method based on multi-modal fusion, and relates to the technical field of network information processing.The method comprises the following steps: by splitting the public opinion video into an audio stream and an image frame sequence, structurally analyzing picture text and speech content respectively, and completing semantic embedding and cross-modal alignment under a unified time axis, introducing a bidirectional cross-modal attention and a gating fusion mechanism to realize deep fusion of multi-modal information, and further combining multi-role consistency, semantic superposition effect and emotion and fact collaborative features for joint modeling at an event level scale, so as to realize overall understanding of network public opinion video content and accurate determination of public opinion risk level.The application can fuse multi-modal information as a whole and realize joint modeling to analyze network public opinion video content.
Owner:WUHAN FIBERHOME PUTIAN INFORMATION TECH CO LTD

Video content analysis method, server, product, equipment and storage medium

PendingCN121459251ACharacter and pattern recognitionVideo content analysisEvent data
The invention discloses a video content analysis method, a server, a product, equipment and a storage medium, and relates to the field of video analysis, and the method comprises the steps: calling a preset large language model which cannot be subjected to subjective reasoning, carrying out the information extraction of the video display content of a video file at each moment, and obtaining an information extraction result; and performing risk identification on the extracted structured event data based on a risk identification rule, extracting verification evidence of the risk assessment result of each event entity from the video display content, and outputting a content analysis result based on each risk assessment result and the verification evidence. According to the method and the device, on the basis of the preset large language model which cannot be subjectively reasoned, the disassembly of the complex description logic is realized, the accuracy of each event entity obtained by disassembly is improved, the risk assessment result of each event entity is verified by using the verification evidence, and the content analysis precision of the finally output video file is improved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (DI)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real-world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Video texture mapping and real-time rendering method and system based on digital twin scene

ActiveCN121458862AResource allocationCharacter and pattern recognitionGraphicsVideo content analysis
The invention provides a video texture mapping and real-time rendering method and system based on a digital twin scene, and relates to the technical field of digital twin rendering. The method comprises the following steps: dynamically dividing texture detail levels of different regions of a video stream according to features and numerical values of image analysis, generating region detail identifiers associated with video contents, identifying dynamic features of a video picture based on the identifiers, and generating a scene state code representing a comprehensive picture condition; a target engine and corresponding rendering quality parameters are dynamically selected from preset heterogeneous rendering engines in combination with the current load state of a graphics processor, and video stream image data are mapped to the surface of a model corresponding to a digital twin scene as dynamic textures, so that efficient rendering of pictures is realized. Dynamic texture mapping and self-adaptive real-time rendering based on video content analysis and hardware load in a digital twinning scene can be realized.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD

Video abstract extraction method based on industry large model

The invention provides a video abstraction extraction method based on an industry large model, and belongs to the field of artificial intelligence and machine learning. The large model is finely adjusted and optimized on specific industry data, so that the large model can deeply understand and adapt to specific video contents of the industry, the quality and the industry adaptability of the video abstraction are remarkably improved, and the video abstraction extraction efficiency is improved. And a more accurate, efficient and reliable solution is provided for video content analysis, retrieval and management of each industry.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Multimodal model driven video recognition method and apparatus

This application provides a multimodal model-driven video recognition method and apparatus, relating to the field of video content analysis technology. The method includes: acquiring a target short video and performing frame segmentation, audio separation, and text content extraction to obtain video frame sequences, audio data, and text data; extracting visual features, speech features, and text features based on the video frame sequences, audio data, and text data, respectively; inputting these multimodal features into a multimodal fusion network based on an attention mechanism, and performing weighted fusion by calculating cross-modal interaction weights between the features to generate a fused feature representation; inputting the fused feature representation into a business order recognition classifier to obtain the business order recognition result of the target short video. The technical solution of this application can improve the accuracy of short video business order recognition and is applicable to scenarios such as short video platform content review, advertising monitoring, and business data analysis.
Owner:BEIJING STAR RIVER EXCELLENCE TECH CO LTD

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (DI)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real-world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Chinese zither performance video content intelligent analysis and retrieval system

The invention discloses a zither performance video content intelligent analysis and retrieval system, and belongs to the technical field of video content analysis and retrieval, the system comprises a multi-modal feature extraction module, a manifold topology transformation module, a semantic scene analysis module and a self-adaptive retrieval module, deep fusion of visual and audio features is realized through manifold topology transformation, and the video content is analyzed and retrieved. A topological structure is established based on Riemannian metrics, and precise recognition and time sequence segmentation of scenes such as technique display, track playing and teaching explanation are achieved. The adaptive retrieval module constructs a closed-loop feedback mechanism, dynamically adjusts feature extraction weights and measurement parameters according to retrieval confidence, and realizes continuous optimization of system performance, the problems of insufficient multi-modal information fusion, inaccurate scene recognition, low retrieval efficiency and the like in the prior art are solved, the scene recognition accuracy is improved by more than 12%, and the system performance is improved by more than 12%. And the retrieval precision is improved by more than 30%.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Vehicle passenger room remnant detection system and detection method thereof

PendingCN121982681AAddress functional limitationssolve protection problemsOptical detectionCharacter and pattern recognitionAutomatic controlVideo content analysis
The invention discloses a vehicle passenger room remnant detection system, which is applied to an automatic passenger shortcut system, and the automatic passenger shortcut system comprises a vehicle automatic control system. The detection system comprises a vehicle-mounted VCA system and a ground VCA system. VCA is video content analysis; and the vehicle-mounted VCA system communicates with the ground VCA system through a vehicle-ground wireless network. The invention further discloses a detection method of the detection system. By means of the technical scheme, whether left articles and persons exist in the compartment or not can be rapidly judged, alarm is given to inform center workers while the car is deducted, the positions of the left articles and persons are located in the video, the workers can process the left articles and persons conveniently, manpower is saved, and meanwhile operation efficiency is improved.
Owner:CRRC PUZHEN BOMBARDIER TRANSPORTATION SYST CO LTD

Method and apparatus for generating heat video, storage medium and electronic device

The present disclosure provides a heat video generation method, a heat video generation device, a computer storage medium and an electronic device, and relates to the technical field of game videos. The heat video generation method comprises the following steps: obtaining an initial heat video, performing video content analysis on the initial heat video to obtain a plurality of heat point dimension information; performing fusion processing on a heat video in which the sum of the heat of at least one same heat point dimension information in the initial heat video is greater than a preset heat value to obtain an intermediate heat video; generating a target heat video corresponding to the video identifier from the corresponding game based on the video identifier of the intermediate heat video, to obtain a new heat video. The present disclosure can improve the number of heat videos and the richness of video content.
Owner:NETEASE (HANGZHOU) NETWORK CO LTD

Video content analysis system according to scene reaction of video content

A video content analysis system according to a scene reaction of video content according to one embodiment of the present invention comprises: a video processing unit which divides video content into a plurality of scenes and classifies each of the divided scenes for each preset reaction; an information collection unit for collecting biometric data of a viewer reacting while the video content is reproduced to the viewer; a reaction analysis unit for analyzing a reaction for each scene from the biometric data collected by the information collection unit; and a scene analysis unit for comparing and analyzing the reaction classified by the video processing unit and the reaction with respect to the biometric data analyzed by the reaction analysis unit for each scene.
Owner:HISTRANGER INC

Image and light combined interaction device

ActiveCN223930663UVideo gamesVideo content analysisDisplay device
The utility model discloses an image and light combined interaction device, which comprises a transmitting unit, a receiving unit, a display unit, a light unit and a control processing unit, and is characterized in that the transmitting unit, the receiving unit, the display unit and the light unit are respectively connected with the control processing unit; the transmitting unit is used for aiming at the display unit to transmit infrared signals; the receiving unit is used for receiving the infrared signal and feeding back the infrared signal to the control processing unit; and the control processing unit is used for changing the video content of the display unit, analyzing the signal frequency of the infrared signal, executing an instruction corresponding to the signal frequency, calling the light script in the memory and controlling the light unit to realize a dynamic effect. By implementing the device, the problems that the interaction mode of the display and the transmitting device is limited by environment layout and cost, and large-scale application is difficult can be solved; and the control of the interaction mode of the lamplight and the emitting device is single, only switching between switch states can be realized, and complex visual effects are lacked.
Owner:HUAQIANG FANGTE (SHENZHEN) TECH CO LTD

Video streaming system and video streaming method

Disclosed is a video streaming system with server(s) and client device(s) that are communicably coupled. The server(s) is configured to: receive HDR video content comprising HDR images; analyse dynamic range and colour characteristics; receive, from client device(s), first metadata indicative of viewing conditions of HDR video content; adjust HDR mastering parameters, based on first metadata and analysis; perform HDR mastering and compress HDR images; send to client device, compressed HDR images and second metadata indicative of adjusted HDR mastering parameters. The client device(s) is configured to receive compressed HDR images and second metadata; decompress compressed HDR images; translate pixel values in decompressed images based on second metadata; receive real-world images from camera(s); compose XR images using decompressed HDR images and real-world images; send to server, first metadata for next HDR video content; and display XR images on display(s) of client device.
Owner:VARJO TECH OY

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (DI)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real-world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

Systems and methods for interactive content viewing and discovery

Disclosed are computerized systems and methods for a decision intelligence (DI)-based framework that automatically and / or dynamically provides an interactive content viewing and content discovery experience to users. The framework includes functionality for real-time video content analysis and interactive entity identification that, inter alia, provides novel capabilities to viewing users related to the extraction, analysis and subsequent interaction with content depicted within video frames during playback of such content. The framework implements artificial intelligence / machine learning (AI / ML) approaches for video processing, user interaction and information delivery through a series of interconnected processes and subsystems. In some implementations, rendered content can be parsed and mined for real-world and / or digital content depicted therein that relate to real-world entities and / or digital resources, whereby interaction with such entities is provided in the form of a provided interface, electronic message and / or recommendations for further information discovery, or some combination thereof.
Owner:RIGH INC

A TMAF-based audio-video-text multi-granularity fusion event positioning method

PendingCN122333164AVideo content analysisEngineering
The application is suitable for the technical field of computer vision and multimedia processing, and provides an audio and video text multi-granularity fusion event positioning method based on TMAF. First, multi-granularity text descriptions are generated for input audio and video data. Second, visual-text and audio-text features are extracted and fused by using a pre-trained model. Third, event multi-granularity continuous semantics are captured by an event-level time sequence multi-scale attention mechanism. Fourth, visual guided audio feature filtering is used to suppress cross-modal noise and enhance feature alignment. Finally, event discrimination and positioning are completed by full supervision or weak supervision training under a constructed multi-task loss function. The application realizes high-precision, high-robustness and high-efficiency audio and video event positioning, and can be widely applied to the fields of intelligent monitoring, video content analysis and retrieval.
Owner:LIAONING NORMAL UNIVERSITY

Video clip summarization system and method

The application provides a video clip abstract generation system and method, the system comprises a video content analysis module, an abstract generation module, a user interaction module and a retrieval cooperation module; the method comprises the following sub-steps: S1, video content analysis is performed and video features are extracted; S2, a clip abstract reflecting the core content of the video is generated; S3, user interaction and dynamic adjustment; S4, the abstract clip is integrated with a video retrieval system; the method improves the degree of automation through feature extraction of the video content analysis module, no longer relies on manual segment-by-segment editing, significantly improves the efficiency of abstract generation, reduces labor costs, shortens processing time and is suitable for processing massive video content; by setting a user feedback entrance, allowing users to provide real-time feedback through text or check the label, the system automatically records the feedback, adjusts the feature weight according to the user's demand, continuously improves the generated abstract according to the user's demand, and improves the intelligent level and flexibility of the system.
Owner:NANJING COMPREHENSIVE SAFETY CONSULTING CO LTD

A multi-target tracking method and system based on a query mechanism

The present application relates to a kind of multi-target tracking method and system based on query mechanism, belong to computer vision image processing technical field.Firstly, the scene video containing moving target is acquired.From the video frame obtained, the detection query of interest is acquired, for detecting the target of interest in current frame.Finally, the detection query and the last frame tracking result are used to track the target of current frame.The system includes video frame acquisition module, feature extraction module, detection query acquisition module and target query module.The present application simplifies the process of multi-target tracking, reduces the complexity of data association, improves the ease of use of multi-target tracking algorithm, improves the practicality of multi-target tracking method, and provides technical support for high-level video content analysis and processing.At the same time, the present application allows multi-target tracking to track the target of interest in an automatic identification or human designated manner, further improving the practicality.
Owner:BEIJING INST OF TECH

Intelligent processing method for public security command based on multi-modal fusion deep learning

The application discloses a public safety command intelligent processing method based on multi-modal fusion deep learning, belongs to the technical field of public safety command, and aims to solve the problems of unclear field perception, delayed abnormal early warning and insufficient joint linkage of the existing public safety command. The metadata model supporting mode adaptation and relationship customization is constructed, multi-modal and billion-level data convergence fusion is realized, cross-domain and cross-layer streaming media processing is cooperated with video content analysis and mining, the smoothness of media transmission is ensured, the accuracy of abnormal event time section and space section positioning in monitoring video is improved, and the identification accuracy under complex conditions is ensured. Based on the multi-modal information fusion of expert knowledge model and deep learning model in the decision-making stage, the research and judgment accuracy is greatly improved, and the practical ability of the professional research and judgment model is supported. The application is suitable for the public safety command intelligent processing method based on multi-modal fusion deep learning.
Owner:SICHUAN JIUZHOU VIDEO TECH