Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

8619 results about "Video streaming" patented technology

Water conservancy and hydropower engineering construction safety supervision system and method based on multi-source data fusion

The invention belongs to the technical field of water conservancy and hydropower engineering, and discloses a water conservancy and hydropower engineering construction safety supervision system based on multi-source data fusion. The system comprises a multi-source sensing acquisition module, a heterogeneous data fusion processing module, a risk identification and early warning module, a safety behavior evaluation and feedback module, and a command scheduling and visualization module. According to the invention, by fusing multi-dimensional data such as image monitoring, environment sensing, personnel positioning, equipment state and the like, a space-air-ground three-dimensional sensing network is constructed, and in a high slope area, the distributed optical fiber strain sensors are linked with thermal imaging data of the unmanned aerial vehicle, so that millimeter-level deformation and temperature field abnormity can be captured in real time; a video stream is analyzed in real time by means of a YOLOv8 algorithm, illegal operation behaviors of personnel can be accurately identified, a cross-modal fusion model of a Transform architecture is combined, the system can dynamically capture potential correlation among data, and millisecond-level response to risks such as side slope landslide, equipment faults and personnel dangerous operation is achieved.
Owner:YUNNAN TUOMEI DECORATION ENGINEERING CO LTD

Intelligent port crane operation monitoring method and system

The invention discloses an intelligent port crane operation monitoring method and system, and relates to the technical field of crane systems.The system comprises a data acquisition unit, an edge computing node, a cloud processor and an early warning monitoring unit, through deep fusion and efficient treatment of multi-source data, a dynamic twinborn data pool architecture is established, and the dynamic twinborn data pool architecture is established; cross-data-type space-time alignment is realized through multi-modal coding, and the feature fusion accuracy is improved; the dynamic twin data pool is intelligently analyzed, and the crane operation path is optimized in real time, so that the track of the crane operation path is dynamically adjusted, and the working efficiency is improved; through a digital twin model and integration of a multi-modal physical engine, the operation physical state deviation of the crane is evaluated, an operation plan is corrected in real time, and the operation precision is improved; video streams are deeply analyzed, virtual-real comparison is carried out, and model prediction video deviation is monitored, so that the analogue simulation precision is improved; by establishing a closed-loop early warning mechanism, full-process digital management and control are realized, and intelligent upgrading of a port is promoted.
Owner:YANTAI PORT GRP CO LTD +1

Multi-modal image automatic labeling system and method

The invention discloses a multi-modal image automatic labeling system and method, and relates to the technical field of image data processing. According to the multi-modal image automatic labeling system and method, time sequence alignment is carried out on video streams and laser radar point cloud data through an asymmetric dynamic time warping algorithm, semantic and geometric features are extracted, and the elastic coefficient of the algorithm is dynamically adjusted. And combining a modal perception attention mechanism, dynamically allocating fusion weights of the video stream and the laser radar according to the features, generating a cross-modal joint feature vector, and outputting a preliminary labeling result. And generating an annotation robustness index by calculating the prediction entropy and the three-dimensional intersection-to-union ratio confidence of the target detection frame, and iteratively optimizing the annotation result. And mapping the cross-modal features and the labeling result into a space-time correlation map, and displaying the three-dimensional positioning, motion trail and modal contribution degree thermodynamic diagram of the target in real time. The problems of time alignment, feature fusion and labeling robustness are effectively solved, and a high-precision and interpretable automatic labeling solution is provided.
Owner:NANJING MATERNITY & CHILD HEALTH CARE HOSPITAL

Dynamic Latent Space Adaptation Based on Spatiotemporal Kernal Context for Multiscale Rendering

A system for dynamic latent space adaptation using spatiotemporal kernel context for multiscale rendering with hierarchical and Lorentzian autoencoders. The Spatiotemporal Kernel Estimator (SKE) analyzes media through motion field, temporal recurrence, frequency band, and scene semantics analyzers to generate adaptive kernel parameters encoding content-specific importance distributions. The system dynamically adapts latent manifold geometry by modifying metric tensor properties according to kernel context, enabling content-aware compression that allocates representational capacity based on visual significance. A multiscale cache implements kernel-adaptive retention policies prioritizing important regions. An adaptive renderer provides intelligent level-of-detail selection based on zoom level and kernel-estimated importance, optimizing processing allocation. The self-optimizing architecture continuously refines kernel context and geometric adaptation based on user interaction and performance feedback, achieving superior compression ratios and perceptual quality. Applications include bandwidth-efficient video streaming, virtual reality, scientific visualization, and cognitive video analytics requiring intelligent context-aware visual processing.
Owner:ATOMBEAM TECH INC

Large-scale scene multi-level-of-detail cloud rendering processing method and device based on 3DGS

The invention provides a large-scale scene multi-level-of-detail cloud rendering processing method and device based on 3DGS, and relates to the technical field of three-dimensional modeling, and the method comprises the steps: dividing a target modeling scene into a plurality of sub-blocks; performing particle redundancy reconstruction and overlapping region marking on boundary regions between adjacent sub-blocks of each sub-block to obtain processed sub-blocks; performing multi-detail level division on each processing sub-block to generate a particle level set; determining a current visual area according to the user motion data, and scheduling a target hierarchy of a particle hierarchy set in the current visual area; the rendering tasks of all the processing sub-blocks of the target hierarchy are distributed to a plurality of rendering nodes to execute real-time rendering operation; and performing video stream coding on pictures rendered by each rendering node, decoding and displaying received video stream data, and performing particle level updating and re-rendering operation according to an interaction instruction. According to the invention, high-quality detail rendering can be realized for a model of a large scene.
Owner:MOBILE BROADCASTING & INFORMATION SERVICE IND INNOVATION RES INST (WUHAN) CO LTD

Millimeter wave radar behavior identification method based on multi-task cross-modal attention

The invention belongs to the technical field of intelligent perception and mode recognition, and particularly relates to a human body behavior recognition method based on millimeter wave radar and multi-task learning. The method comprises the steps that millimeter wave radar point cloud data and RGB video streams are synchronously collected, and a five-dimensional point cloud scene is generated through three-dimensional analysis and dynamic target extraction; a hierarchical Point Transform network is constructed to extract radar space-time features, and human body key point features are generated by using visual auxiliary attitude estimation; radar and key point features are dynamically fused through a cross-modal attention mechanism, a multi-task joint optimization strategy is combined, behavior classification serves as a main task, attitude estimation serves as an auxiliary task, and a behavior recognition result is output. According to the method, through multi-task cooperation and cross-modal feature interaction, on the premise of ensuring privacy security, the accuracy and robustness of human behavior recognition in a complex scene are remarkably improved, and efficient and reliable technical support is provided for the fields of intelligent monitoring, human-computer interaction and the like.
Owner:XIDIAN UNIV +1

Methods For Generating Advertisement Videos Consistent With The Context And Storyline Of A Primary Video Stream

Embodiments include methods for generating advertisement videos for insertion into a video stream to promote a product, service, or brand in a manner that is consistent with the context and storyline of the video stream before and at the time of ad insertion. Methods may include capturing an image from the video stream and generating caption text using an image-to-text description model. A product, service, or brand that is consistent with the context and storyline of the captured image is selected and ad video sequence description text is generated that includes descriptions and a storyline blending descriptions of the selected product, service, or brand with the context and storyline of the primary video stream. The ad video sequence description text is used to prompt a text-to-video generation model that generates a new advertisement video clip, which is inserted into the primary video stream before distribution to video content rendering devices.
Owner:CHARTER COMM OPERATING LLC

Video stream dynamic fragment encryption and block chain evidence storage method

The invention discloses a video stream dynamic fragmentation encryption and block chain evidence storage method, and relates to the technical field of video content security, and the method comprises the steps: calculating a color histogram difference value and an optical flow vector change rate between adjacent frames of an input video, marking the difference value as a scene switching point when the difference value exceeds a preset threshold value, and storing the scene switching point; the method comprises the following steps: preliminarily dividing a video into a plurality of scene segments according to scene switching points, performing content complexity evaluation on the scene segments, calculating gray level co-occurrence matrix characteristics of each frame of image through texture density analysis, calculating edge complexity to extract the number and distribution of Canny edges, and performing motion vector statistics to analyze the size and direction of inter-frame object displacement. The change rate between adjacent pixels in the color space is measured according to the color change gradient; the video stream dynamic fragment encryption and block chain evidence storage method is suitable for video contents of different types and complexities, and has relatively high detection accuracy and robustness.
Owner:HANGZHOU MEICHANG IOT TECH CO LTD

Intelligent event identification method and system based on high-speed camera

The invention provides an intelligent event identification method and system based on a high-speed camera, and the method comprises the steps: setting the collection parameters of the high-speed camera, and triggering the camera to collect a target scene video stream. And hardware acceleration decoding processing is carried out on the collected original video data stream, real-time environment illumination information of the environment illumination sensor is obtained, and dynamic brightness equalization processing is executed. And performing motion adaptive denoising processing on the video sequence. Geometric distortion correction is carried out on the image sequence through camera calibration parameters, sub-pixel-level displacement vectors and dense optical flow field data of a moving object are extracted, and multi-scale morphological features are extracted. And the features are fused to generate motion feature data, the data are processed through a spatio-temporal joint event classification model, an event identification result is output, the result is bound with a high-precision timestamp, and event identification information is output to an industrial control system display device in real time. According to the invention, the accuracy and real-time performance of event identification can be improved.
Owner:广州思林杰科技股份有限公司

Human abnormal behavior monitoring method based on large-model multi-agent

The invention discloses a human abnormal behavior monitoring method based on a large-model multi-agent, which is executed by a modular multi-agent system deployed on a back-end server, obtains information through a monitoring camera, and comprises the following steps: obtaining a video stream from the monitoring camera by a sensing agent and extracting human body posture features; analyzing the key frame by a scene understanding agent by using a visual large model, and constructing a time sequence dynamic scene graph; the core reasoning agent evaluates the scene semantic conformity based on the pre-trained large model and performs abnormal preliminary judgment; performing fine-grained classification, interpretation generation and risk assessment on the abnormal behaviors; and the report and action agent generates an alarm and records event data. According to the invention, through multi-agent cooperative work and a large model technology, efficient and accurate monitoring of human abnormal behaviors is realized, and the intelligent level of the monitoring system and the abnormal behavior identification accuracy are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

High-fidelity cloud rendering cluster scheduling method and system

The invention relates to a high-fidelity cloud rendering cluster scheduling method and system, and the method comprises the steps: receiving a rendering task, generating a multi-level priority queue according to the task complexity, timeliness and resource demand classification, and automatically optimizing a scheduling strategy; monitoring the resource state of the heterogeneous computing node in real time; allocating tasks by using preemptive and round-robin scheduling strategies, and adjusting and coping with resource fluctuation in combination with a dynamic code rate; the tasks are decomposed by adopting a spatial blocking and time framing strategy, and an execution sequence is controlled according to a topological sorting algorithm; abnormal nodes are identified through heartbeat detection, and affected subtasks are migrated through incremental task updating. According to the method, efficient resource allocation, scheduling algorithm and idle key frame scheme can be realized, the GPU directly outputs the video stream, the rendering speed is remarkably improved, resource waste and operation cost are reduced through a dynamic resource allocation mechanism, a fault-tolerant mechanism is provided, the task is ensured to be normally completed under the condition of node fault, and the task efficiency is improved. And distributed rendering and synthesis of large-scale high-resolution images are supported.
Owner:SHENZHEN TRAFFIC CONSTR ENG TEST & DETECTION CENT +1

Real-time video stream behavior identification and early warning system

The invention relates to the technical field of video behavior recognition, and discloses a behavior recognition and early warning system for a real-time video stream. The system comprises a spatio-temporal feature modeling module, a behavior fragment extraction module, an anomaly propagation modeling module, a risk area positioning module and an early warning strategy generation module. The spatial-temporal feature modeling module builds a dynamic model based on historical data, captures a skeleton key point three-dimensional coordinate sequence, a motion optical flow vector field and a micro-expression intensity spectrum, and outputs a theoretical behavior mode vector; the behavior fragment extraction module generates a multi-modal difference feature tensor through cross-modal difference analysis; the exception propagation modeling module generates an exception propagation path risk probability distribution cloud picture in combination with spatial constraint and trajectory information; the risk area positioning module identifies a high-risk area and marks a boundary; and the early warning strategy generation module dynamically configures monitoring parameters, starts high-frame-rate micro-expression capture for a high-risk area, and performs a track disturbance test on an adjacent area.
Owner:GAOZI TECHNOLOGY (SHENZHEN) CO LTD

Real-time video analysis method based on deep learning

The invention relates to the technical field of computer vision, and discloses a real-time video analysis method based on deep learning. The method comprises the following steps: acquiring a real-time video stream through image acquisition equipment, and performing frame segmentation processing to generate a continuous video frame sequence; and extracting features of the video frame sequence by using a pre-trained convolutional neural network to obtain a multi-dimensional feature vector, inputting the multi-dimensional feature vector into the time sequence analysis model to calculate dynamic relevance, and outputting an inter-frame movement track and object behavior features. And constructing a scene understanding map containing a spatial position and a time evolution relationship according to the above-mentioned data, and carrying out abnormal event detection and generating event marking data based on the map. And performing semantic analysis on the event marking data, determining an abnormal event type and a confidence score, triggering a real-time alarm signal according to a result, and updating a historical event database. In the analysis process, the resource occupancy rate of the system is continuously monitored, the calculation precision is dynamically adjusted, a degradation processing mechanism is started when a preset threshold value is exceeded, and key area analysis is preferentially guaranteed.
Owner:HANGZHOU SIYUAN INFORMATION TECH CO LTD

Airport video data real-time analysis system

The invention relates to the technical field of airport safety monitoring, and discloses an airport video data real-time analysis system. The system comprises a video stream spatial-temporal feature modeling module, a behavior trajectory map construction module, an abnormal region association analysis module, a risk level semantic judgment module and a situation structure visualization module. According to the method, multi-scale spatial-temporal feature analysis is carried out on an airport monitoring video stream, a multi-dimensional behavior trajectory map is established, abnormal behavior region association is analyzed, risk level semantics are judged, and finally an airport global risk situation thermodynamic distribution map is generated. According to the system, the whole process processing from video data acquisition to risk situation visualization is realized, the abnormal behavior area can be accurately identified, the risk level and category are clear, comprehensive and visual situation information is provided for airport safety management, and the intelligent level of airport safety management is improved.
Owner:SHAANXI GUANGHUIYUAN INTELLIGENT TECH CO LTD

End-cloud cooperative detection method and system for traffic anomalies

The invention discloses an end-cloud cooperative detection method and system for traffic anomalies, and relates to the technical field of intelligent traffic control. The method comprises the following steps: collecting traffic video streams through edge equipment, and identifying abnormal behaviors and generating structured event data by using a lightweight YOLOv3-tiny model; when the confidence exceeds a dynamic threshold and the event type is a high-risk type, uploading a video clip and data to a cloud; the cloud integrates a historical road condition map, meteorological data and a real-time traffic flow state, reconstructs a three-dimensional event scene by adopting a space-time attention pyramid network, and verifies the authenticity of an event in combination with a traffic flow sudden change detection algorithm; and generating a signal lamp forced switching instruction for the risk level overrun event, and issuing the signal lamp forced switching instruction to roadside equipment within 3 seconds to execute emergency response. According to the invention, full-link closed-loop control of traffic accidents from identification to response is realized, and the false alarm probability is greatly reduced while the identification accuracy is guaranteed.
Owner:高翔

AI-based security and protection monitoring video analysis method and system

The invention provides an AI-based security and protection monitoring video analysis method and system, and the method comprises the steps: firstly obtaining an initial video data set which corresponds to a to-be-analyzed monitoring video stream and comprises a plurality of video segment units which are continuously collected and are provided with timestamp marks, and then carrying out the video feature extraction processing of the initial video data set; the time sequence behavior characteristics of the target object and the spatial structure characteristics of the scene where the target object is located are obtained, then a pre-constructed abnormal behavior detection model is called to carry out joint abnormal behavior detection on the two characteristics, and an abnormal behavior detection result is generated; and determining an abnormal event type and spatial and temporal distribution feature information in a video picture according to the abnormal behavior detection result, and finally generating a security early warning instruction containing event positioning coordinates and sending the security early warning instruction to the target security control terminal to trigger linkage response operation, thereby improving the accuracy of abnormal behavior detection. Accurate positioning and rapid early warning of the abnormal event are realized.
Owner:CHENGDU SUNJIANG INFORMATION TECHNOLOGY CO LTD

Industrial visual anomaly identification method, system and device based on multi-modal feedback and medium

The invention provides an industrial visual anomaly recognition method, system and device based on multi-modal feedback and a medium, and belongs to the technical field of industrial visual detection.The method comprises the steps that a production line video stream is collected in real time, an ROI is automatically segmented, a picture sequence is generated, features of a set dimension are extracted, and the features of the set dimension are extracted; comparing a preset dynamic distance threshold with a feature library, and preliminarily identifying an abnormal image; a cue word of a detection requirement is constructed, the cue word and abnormal related information are input into the multi-modal large model, and a determined abnormal image is output; using the determined abnormal image to construct a data set to train a target detection model, and adjusting the learning rate in the training process; and inputting pictures captured in an industrial scene into the deep learning model, the multi-modal large model and the target detection model in sequence to obtain a recognition result, feeding back the recognition result to a user for confirmation, adding the recognition result to the data set, and optimizing training of the target detection model. According to the invention, accurate identification of abnormal images in industrial production is realized, and the identification efficiency is high.
Owner:山东浪潮智能生产技术有限公司

Edge calculation signal lamp control method and system based on traffic participant behavior analysis

The invention provides a traffic participant behavior analysis-based edge calculation signal lamp control method and system, and the method comprises the steps: obtaining video stream data of a target intersection, carrying out the frame-by-frame behavior state analysis of the video stream data through a spatial-temporal feature coding network, extracting the behavior state features of each traffic participant, and carrying out the analysis of the behavior state features of each traffic participant; and inputting the behavior state characteristics into a preset behavior matching model, generating a behavior triggering identifier of the traffic participant in the current time window, generating a signal lamp control instruction of an edge computing node according to the time sequence relevance between the behavior triggering identifier and the phase state of the current signal lamp, and sending the signal lamp control instruction to the edge computing node. And finally, a signal lamp phase switching time sequence of the target intersection is adjusted based on the signal lamp control instruction, so that the traffic participants meeting the traffic rule conflict condition obtain the traffic priority under the signal lamp phase switching time sequence. According to the method, the traffic accident risk is reduced, and meanwhile, the overall traffic efficiency of the intersection is optimized by dynamically adjusting the minimum response period and the phase duration parameters.
Owner:HEBEI JOY SMART TECH CO LTD

Video stream processing method for dynamic Gaussian compression and adaptive code rate regulation

The invention discloses a video stream processing method for dynamic Gaussian compression and adaptive code rate regulation, which is suitable for scenes such as virtual reality, augmented reality and three-dimensional video, and comprises the following steps: S1, Gaussian attribute modeling and initialization; s2, constructing a binary hash grid; s3, constructing a deformation prediction network; s4, designing a mask pruning mechanism; s5, entropy modeling and arithmetic coding and decoding module design; s6, model training; and S7, video stream transmission under multiple code rates. According to the method, a unified scheme combining Gaussian volume cloud coding and adaptive video transmission is proposed for the first time, the video data storage and transmission cost is remarkably reduced, and the comprehensive performance superior to that of an existing method is obtained on multiple real and synthetic data sets.
Owner:THE CHINESE UNIV OF HONG KONG (SHENZHEN) FUTURE NETWORK OF INTELLIGENCE INST +1

Intelligent cockpit system for fire fighting

The invention relates to the technical field of fire fighting systems, in particular to an intelligent cockpit system for fire fighting, which comprises a perception analysis layer, a decision processing layer, a command execution layer and a reinforcement learning closed-loop architecture for mixed reward shaping, and integrates video streams, audio communication and firefighter physiological data containing heart rate variability through a deep multi-modal fusion module. Generating a global fire scene situation of physical constraint verification; through a risk sensitive type three-dimensional fire scene deduction module, a fire extinguishing strategy considering efficiency and safety is generated; through an immersive augmented reality visual interface, in combination with a synchronous positioning and mapping technology, precise navigation in a complex environment and superposed display of situation information containing risk levels are realized; a Q learning algorithm continuous optimization strategy including domain knowledge intermediate process rewards is adopted, closed-loop optimization of perception-analysis-decision-execution-feedback is achieved, and therefore the overall efficiency of fire rescue operation in modern complex disaster scenes is remarkably improved.
Owner:ZHEJIANG YIMIN INFORMATION TECHNOLOGY CO LTD

Sublimation hardware-based video multi-target intelligent detection method and system

The invention discloses a video multi-target intelligent detection method and system based on mercuric chloride hardware. A hardware decoding module decodes an input video stream in real time, generates video frames and stores the video frames in a shared memory queue. And inputting the video frame into the YOLOv5 model converted by the mercuric chloride OMG tool, pre-loading a plurality of model instances into a memory by using an ACL interface of the mercuric chloride NPU, calling different model instances through a polling scheduling mechanism, and outputting structured data comprising a multi-target detection frame, a category label and confidence. And carrying out non-maximum suppression processing on the reasoning result, and judging whether the target is in an alarm monitoring area or not by adopting a central point detection method. According to the invention, real-time multi-target intelligent detection of the input video stream can be realized. And the decoded video frames are stored in a shared memory queue, so that efficient data transmission and processing are realized. The multi-model parallel reasoning improves the detection precision, and is suitable for the application scene of real-time video multi-target intelligent detection.
Owner:CHENGDU SIWEI INTERACTIVE TECH CO LTD

Cloud gateway storage

A video gateway system at a worksite is coupled to multiple cameras on a network, and backs-up video streams generated by the cameras to a backend cloud backup video storage system and a frontend (cache) video storage system. The video gateway system generates an aggregated video asset from a plurality of streams of video from the multiple cameras, and generates metadata and a backup report associated with the video asset. The video asset, metadata, and backup report are stored on the backend cloud backup video storage system in an file system, and are also stored on the frontend (cache) video storage system such that for at least a period of time, the video asset and the associated the metadata and backup report are stored on both the cloud backup video storage system, facilitating quick access and retrieval to the stored video for retrieval for streaming, activity detection, and other uses.
Owner:SAMSARA INC

Short video intelligent editing method and system based on multi-modal analysis

The invention discloses a short video intelligent editing method and system based on multi-modal analysis, and relates to the technical field of video editing. The method is used for improving editing efficiency and visual experience and comprises the following steps: extracting lip motion features of a character, visual saliency features of a commodity and a voice emotion intensity value from a target short video stream to form multi-modal time sequence data; afterwards, the voice stream is recorded, a product keyword timestamp is extracted, the alignment degree is calculated through dynamic time warping in combination with a visual saliency peak value, and a preliminary editing point set is generated through weighted evaluation in combination with an emotional intensity value; constructing an editing decision optimization model based on deep reinforcement learning, taking the multi-modal features as state input, adjusting the retention probability of editing points through a joint reward function, and selecting an optimal transition mode; and the lip movement and voice synchronization error before and after the editing point and the emotional and visual continuity of the transition section are analyzed, the discontinuous region is smoothed, and the edited finished product is output, so that precise short video intelligent editing is realized.
Owner:ANHUI XINGBANG DIGITAL TECHNOLOGY GROUP CO LTD

Video stream real-time target detection and tracking system based on deep learning

The invention relates to the technical field of data processing, in particular to a video stream real-time target detection and tracking system based on deep learning, and the system comprises a data collection module which captures an original video stream and outputs a bimodal image sequence; the adversarial generation module synthesizes low-frequency samples; the feature alignment module dynamically adjusts the channel attention weight; the adaptive detection module switches a main detection path and an auxiliary detection path according to the scene complexity index, the main path generates a target detection frame by adopting lightweight convolution, and the auxiliary path activates the re-identification sub-network to generate a supplementary detection frame; the space-time diagram tracking module uses the channel attention weight to carry out cross-modal matching to correct conflict nodes; the trusted decision module generates a probabilistic behavior decision. The system continuously compresses the distribution difference between training data and an actual scene through a closed-loop optimization mechanism of coordinate feedback driving sample synthesis, channel weight adjustment distribution calibration and trajectory continuity constraint decision reliability, and improves the robustness of target detection and tracking in a dynamic environment.
Owner:RONGAN CLOUD NETWORK (BEIJING) TECH CO LTD

Video analysis-based multi-scene operator violation behavior identification method and system

The invention discloses a video analysis-based multi-scene operator violation behavior identification method and system, and belongs to the technical field of intelligent operation safety monitoring and artificial intelligence identification, and the method comprises the steps: collecting a real-time video stream of a multi-scene operation site; recognizing a continuous action time sequence in the real-time video stream by using an action recognition depth model; constructing the continuous action time sequence into an action behavior sequence; the action behavior sequence is constructed into a directed behavior graph with time, space and action labels, the directed behavior graph is compared with a directed behavior graph corresponding to the standard action behavior sequence, and illegal behaviors are recognized; and carrying out multi-mode early warning on the identified illegal behaviors. According to the method, the bottleneck that the traditional image recognition technology is weak in action sequence semantic understanding and poor in environmental adaptability is broken through, and accurate recognition and real-time early warning of illegal behaviors in multi-scene operation are achieved.
Owner:CHENGDU HANGTIAN PHOTOELECTRIC TECH

Financial deep counterfeiting detection and prevention system and method based on multi-modal large model

The invention discloses a financial deep counterfeiting real-time detection and defense method and system based on a multi-modal large model, and the method comprises the steps: obtaining multi-modal data in a financial transaction scene, and carrying out the desensitization of an edge end; performing dynamic time sequence alignment on the multi-modal data, and calculating a synchronous error of lip motion and voice by adopting a dynamic time warping algorithm; inputting the features into a dynamic risk modeling layer, and generating dynamic risk features in combination with the updated risk feature library; analyzing the features through a double-flow GAN detector, and outputting a forgery probability; and a detection result is input into a compliance verification layer, the supervision file is analyzed through a legal BERT, a structured rule is generated, and real-time transaction interception and block chain log recording are executed. The method protects user privacy and data security, combines the risk feature library updated in real time, and has high flexibility and adaptability. According to the design of the double-flow GAN detector, image and video stream information is fully utilized, and the accuracy and reliability of detection are further improved.
Owner:HUAYING (SHANGHAI) INFORMATION TECH CO LTD

Video stream processing method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a video stream processing method which comprises the steps of collecting current environment parameters, generating a mode switching instruction and determining a target processing mode; obtaining multi-dimensional context awareness data, and selecting a target detection model; key area coordinate parameters in the video frame sequence are extracted, and grading resolution parameters are determined; and based on the target processing mode, the target detection model and the grading resolution parameter, constructing a video processing strategy matrix, executing the video processing strategy matrix to perform coding processing on the video stream, and generating a target coding video stream. According to the invention, through intelligent mode switching based on the current environment parameters, dynamic adjustment of the video processing mode is realized, and the adaptive capacity of the system in a complex environment is improved; through target detection model selection in combination with multi-dimensional context awareness data, the detection precision is optimized, and the reliability of visual analysis is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Substation operation risk identification method based on multi-view video and high-precision positioning

The invention relates to a substation operation risk identification method based on a multi-view video and high-precision positioning. Acquiring video data of a working site through a plurality of cameras with fixed visual angles and mobile video acquisition equipment; a high-precision positioning system is used for obtaining three-dimensional space coordinates of operators and equipment in real time; establishing a three-dimensional digital twinborn model of the substation equipment, and performing dynamic scene reconstruction based on the multi-view video stream to generate a real-time three-dimensional scene of the operation site; fusing the positioning data and the three-dimensional scene by adopting a space-time fusion algorithm to generate a dynamic digital portrait of the operator; and carrying out real-time analysis on behaviors and positions of operators by using a risk assessment algorithm based on a preset risk rule, calculating to obtain a risk assessment value, and setting a feedback mechanism to continuously optimize positioning and scene reconstruction precision. According to the invention, efficient, accurate and real-time identification and early warning of the operation risk of the transformer substation are realized, and the safety management level of an operation site is effectively improved.
Owner:GUANGZHOU JINGKAI TECH CO LTD

Multi-modal forged video detection method based on multi-head addition cross attention mechanism

The invention discloses a multi-mode counterfeit video detection method based on a multi-head addition cross attention mechanism, and belongs to the technical field of video counterfeit detection. The method comprises the following steps: preprocessing a video stream, decomposing a single-frame positioning face, and extracting an audio to generate a Mel spectrogram slice; the 3D convolutional network extracts video spatio-temporal features and motion differences, and the filter bank extracts audio features in combination with the residual network; audio features are mapped to a video alignment space through asymmetric projection, the video features are subjected to bidirectional interaction with an audio input multi-head addition cross attention module after being subjected to time sequence coding, and audio dominant and video dominant features are generated and are cascaded and fused with original features; and constructing cross-modal similarity loss constraint feature distribution, and fusing feature dynamic weighting and time sequence compression to output four classification probabilities of audio-visual double true, audio-visual double pseudo, video pseudo-audio true and video pseudo-audio pseudo. The multi-mode counterfeiting recognition precision is improved, and texture abnormity and audio and video mismatch features are captured.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Video stream real-time coding and decoding transmission method under cluster

The invention relates to the technical field of cluster video stream processing, and discloses a video stream real-time coding and decoding transmission method under a cluster. The method comprises the following steps: firstly, acquiring video stream coding parameter text data, link state time sequence data and equipment performance index data of multiple nodes of a target cluster to form a transmission link data set; semantic analysis is carried out on the coding parameter text data to obtain a coding semantic feature vector, dynamic fluctuation features are extracted from the link state time sequence data to obtain a link fluctuation feature vector, and cross-modal fusion is carried out to generate a fusion transmission feature set; generating an abnormal association degree score set by using a pre-trained multi-layer sensing network model, and obtaining an abnormal source node and an equipment defect type by combining root cause tracing according to the abnormal association degree score set; and finally, generating a dynamic optimization strategy and feeding back to the transmission control system to trigger parameter calibration. According to the method, the abnormal root cause can be accurately traced, the transmission parameters are optimized, and the cluster video stream transmission quality is improved.
Owner:ZHEJIANG VERSATILE MEDIA