Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

152 results about "Frame based" patented technology

Frame-based terminology is a cognitive approach to terminology developed by Pamela Faber and colleagues at the University of Granada.One of its basic premises is that the conceptualization of any specialized domain is goal-oriented, and depends to a certain degree on the task to be accomplished. Since a major problem in modeling any domain is the fact that languages can reflect different ...

Semantic modeling-based unsupervised video monitoring anomaly detection method and system

The invention provides an unsupervised video monitoring anomaly detection method and system based on semantic modeling, and belongs to the technical field of computer vision, artificial intelligence and video monitoring. Comprising the following steps: carrying out key frame identification on a monitoring video by adopting an image-text joint embedding model, and carrying out target cross-frame tracking based on a depth target detection algorithm to extract behavior semantic information so as to construct a semantic behavior map; sending the key frame sequence of the graph structure information into a frame prediction model, and predicting a next frame image or a target state; and carrying out abnormal scoring on the obtained prediction result, and carrying out threshold judgment by outputting a comprehensive abnormal score value so as to determine whether the current frame is an abnormal event or not. According to the method, the intelligent level and the overall efficiency of video anomaly detection can be effectively improved on the premise of ensuring the real-time performance and the stability, and support is provided for video monitoring anomaly detection in actual scenes such as smart cities, rail transit, industrial parks and commercial security.
Owner:SHANDONG UNIV

Digital production plan scheduling method and system

The invention discloses a digital production plan scheduling method and system, and belongs to the technical field of optimal scheduling, and the method comprises the steps: constructing a distributed storage architecture based on edge computing nodes; a central coordinator is adopted to realize cross-node data synchronization through an improved Raft consensus algorithm, multi-version concurrency control is realized based on a vector clock, and a global consistent data view is established; a visual scheduling platform is built based on a Vue3 framework, and man-machine interaction is realized by adopting a Canvas and WebGL collaborative rendering framework; establishing a dynamic coordinate conversion model based on bilinear interpolation, designing a space mapping function containing distortion compensation, establishing a multi-thread coordinate service based on WebWorker, and realizing submillimeter-level bidirectional mapping of pixel coordinates and physical coordinates; constructing a three-dimensional space-time analysis model fused with the multi-dimensional features; and all the units are subjected to feature fusion through residual connection, and finally a scheduling scheme with a confidence coefficient weight is output. The method and the device have the effect of meeting various scheduling requirements.
Owner:SHANDONG PORT EQUIPMENT GROUP CO LTD

Semantic guided efficient perspective view to BEV projection and sampling

A device for processing frame data may be configured to identify one or more semantic characteristics for the frame data, wherein the frame data comprises a plurality of frames, each frame of the plurality of frames being acquired for a same scene and by a different frame source of a plurality of frame sources; extract features from each respective frame of the plurality of frames; determine a non-uniform sampling pattern of the plurality of frames based on the one or more semantic characteristics; and project, using the non-uniform sampling pattern, a portion of the extracted features into a bird's-eye-view (BEV) space having a grid structure to generate a fused set of BEV features.
Owner:QUALCOMM INC

High-altitude operation risk early warning method and system based on camera image recognition

The invention provides a high-altitude operation risk early warning method and system based on camera image recognition, and relates to the technical field of computer vision, and the method comprises the steps: firstly collecting a video image sequence of a high-altitude operation scene, and generating a fusion feature map containing environment and operation main body features through multi-level feature extraction; performing spatial dimension segmentation and regional feature comparative analysis on the fusion feature map to obtain a spatial risk distribution map containing risk region identification information, processing the spatial risk distribution map of continuous frames based on a time sequence feature fusion rule to generate a dynamic risk evolution map, and calling a risk decision model to perform mode recognition to obtain a dynamic risk evolution map; and generating a risk level classification result and a risk position coordinate set according to the risk level classification result and the risk position coordinate set, and finally generating a risk early warning signal and sending the risk early warning signal to the monitoring terminal, thereby comprehensively, accurately and dynamically monitoring the high-altitude operation risk, and improving the accuracy and timeliness of risk early warning.
Owner:STATE GRID SHANXI POWER TRANSMISSION & DISTRIBUTION PROJECT CO

Multi-mode large model video content understanding reasoning acceleration method and system

The invention discloses a multi-mode large model video content understanding reasoning acceleration method and system, and mainly relates to the technical field of artificial intelligence reasoning acceleration. Comprising the following steps: inputting video data and preprocessing the video data to generate a video frame sequence; performing adaptive video Token compression on the generated video frame sequence, and outputting a compressed visual Token set; performing visual feature coding and Key-Value generation on the compressed visual Token set to obtain visual KV data; performing video KV cache partition management on the visual KV data; executing cross-modal reasoning based on the vLLM framework to generate a video content understanding result; and outputting a video content understanding result, and carrying out post-processing and structured mapping. The method has the beneficial effects that the obvious reasoning acceleration and throughput improvement can be realized on the premise of keeping the precision of the original large model.
Owner:海看网络科技(山东)股份有限公司

Smart frame selection via activity-based ranking and optimization

Various examples, systems, and methods are disclosed relating to frame selection via activity-based ranking and optimization. A first computing system can receive a plurality of frames and metadata from a capture device capturing a video stream. The first computing system can generate, using a ranking model, a plurality of rankings for the plurality of frames based on a plurality of video parameters of the plurality of frames and the metadata, wherein the plurality of rankings correspond to a summarization of the video stream. The first computing system can determine at least one of the plurality of frames to provide to at least one buffer based on the plurality of rankings, wherein the at least one buffer stores a subset of frames of the plurality of frames. The first computing system can provide, from the at least one buffer, the subset of frames as input to a machine-learning model.
Owner:NVIDIA CORP

Training method and device of visual text pre-training model, equipment and medium

The embodiment of the invention provides a training method and device of a visual text pre-training model, equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the steps of inputting an obtained sample video and an obtained sample text into an initial multi-modal processing model, performing segmentation processing on the sample video to obtain a plurality of initial video frames, and extracting text features from the sample text; performing space-time importance evaluation on each initial pixel block in each initial video frame to obtain space-time information; determining a masked video frame based on the spatio-temporal information, and performing feature reconstruction processing on a part in a masked state in the masked video frame based on the text features to obtain a sample reconstruction result; and calculating a loss value according to a sample reconstruction result, and adjusting model parameters of the initial multi-modal processing model according to the loss value to obtain a trained target multi-modal processing model. According to the invention, the ability of the multi-modal processing model obtained by training to understand the multi-modal information can be improved.
Owner:PENG CHENG LAB

Verification excitation automatic generation method and system based on backtracking thinking tree

The invention discloses an automatic verification excitation generation method based on a backtracking thinking tree, and the method comprises the steps: constructing a backtracking thinking tree frame based on a self-feedback mechanism based on a thinking tree for a to-be-verified processor function; under the backtracking thinking tree framework, decoupling a to-be-verified function from top to bottom through a large language model, and recursively decomposing the whole to-be-verified function into function points capable of being independently verified layer by layer through a tree structure; according to each function point, determining a normal function and a boundary of each function point, and for each function point, generating verification excitation layer by layer in a tree structure; after the verification excitation is generated, self-verification and self-backtracking are carried out layer by layer through a tree structure, and finally, the generated verification excitation is executed on the processor design to verify the correctness of the processor function. According to the method and the system, the function verification efficiency and the verification coverage rate of the processor are remarkably improved.
Owner:INST OF COMPUTING TECH CHINESE ACAD OF SCI

Real-time translation method and interaction system based on streaming voice segmentation and semantic verification

PendingCN122050365ANatural language translationSpeech recognitionSpeech segmentationFrame based
The invention provides a real-time translation method and interaction system based on streaming voice segmentation and semantic verification. The method comprises the following steps: receiving a conference audio stream, identifying effective audio frames based on dual-channel voice activity detection, and pressing the effective audio frames into a dynamic buffer area; according to a preset dynamic segmentation strategy, outputting an initial voice segment from the dynamic buffer area for voice recognition, and obtaining a corresponding initial text segment; performing multi-stage reliability verification on the initial text fragment, and dynamically correcting or complementing the initial text fragment according to a verification result to obtain a reliable recognition result; and after the reliable identification result is obtained, triggering an asynchronous parallel translation task. According to the method, a semantic correction mechanism cooperating with the adaptive truncation depth is designed, so that the semantic fragmentation problem of the long-sequence audio during streaming truncation is solved, and the translation accuracy of a complex word order language is greatly improved while low delay is ensured.
Owner:WUHAN UNIV

Computer-implemented multi-scale machine learning model for the enhancement of compressed video

The present disclosure relates to methods, apparatus, systems, and non-transitory computer-readable storage media for training and using a multi-scale machine learning model for the enhancement of compressed video. According to some examples, a computer-implemented method includes receiving a video at a content delivery service; performing an encode on a frame of the video by the content delivery service that converts the frame from a pixel domain to a transform domain and back to the pixel domain to generate first pixel values and a first residual for a block of the frame at a first resolution; generating a first set of features, by a machine learning model of the content delivery service, for an input, at a first resolution, of the first pixel values and the first residual of the block; generating a second set of features, by the machine learning model of the content delivery service, for an input, at a second lower resolution, of second pixel values and a second residual of the block; upsampling the second set of features to the first resolution to generate an upsampled second set of features; generating a modified version of the frame based on the first set of features and the upsampled second set of features; and transmitting the modified version of the frame to a frame buffer or from the content delivery service to a viewer device.
Owner:AMAZON TECH INC

Interactive labeling method for 3D dynamic object based on time series data, key frames, and interpolated frames

Disclosed is interactive labeling of a 4D dynamic object based on time series data, which aims at time series-related point cloud dynamic object data. Multi-frame local point clouds in the same time series are transformed into the same global coordinate system with corresponding poses to obtain global point clouds in the same time series are obtained, which clearly shows the moving trajectory of the dynamic object. Taggers can label key 3D boxes based on the moving trajectory of the dynamic object, and automatically generate 3D prediction boxes of other frames based on these key 3D boxes, which significantly reduces the number of frames that need manual operation and solves the problem that 3D prediction boxes generated based on deep learning model are inaccurate and efficiency can hardly be improved.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Four-side surface re-topology method based on frame field prediction and electronic equipment

The invention provides a quadrilateral surface re-topology method based on frame field prediction and electronic equipment. The quadrilateral surface re-topology method based on frame field prediction comprises the following steps: inputting a triangular mesh model into a frame field neural network model, and carrying out direction regression and amplitude diffusion taking the direction as the condition to obtain a corresponding frame field; performing deformation based on the frame field to obtain a deformed grid model; performing isotropic quadrilateral re-mesh division on the deformed mesh model to obtain an initial quadrilateral mesh model; and performing inverse deformation on the initial quadrilateral mesh model to obtain a quadrilateral mesh model corresponding to the triangular mesh model. Through the steps of frame field inference, deformation, isotropic quadrilateral grid re-division and inverse deformation, the quality of the finally generated quadrilateral grid is ensured, so that the purpose of improving the quality of the quadrilateral grid is achieved.
Owner:BEIJING WAZIDA TECH CO LTD

Two-dimensional code image recognition method based on boundary completion and uncertainty modeling

The invention relates to a two-dimensional code image recognition method based on boundary completion and uncertainty modeling, and belongs to the field of computer vision. The method comprises the following steps: firstly, acquiring an input two-dimensional code image, detecting a cutting area through gray level conversion and edge gradient analysis, complementing the cutting area by utilizing a generative network, and generating enhanced features by adopting deformable convolution; and then extracting rotation irrelevant features through rotation isovariant convolution, fusing spatial information by applying a global attention mechanism, and correcting the angle of a detection frame based on a positioning mark. Secondly, extracting semantic features of candidate areas, calculating similarity with a standard template to generate semantic scores, detecting positioning mark angle distribution to generate geometric constraint scores, and adjusting candidate box confidence; and finally, establishing a confidence coefficient normal distribution model through the multi-layer perception mechanism, calculating a high confidence interval probability, constructing combined loss including classification, positioning, confidence coefficient and uncertainty loss, optimizing the model according to specified parameters, and storing the optimal model to obtain a two-dimensional code recognition result.
Owner:FUZHOU UNIV

Software full life cycle intelligent collaborative optimization method and system based on large model

The invention discloses a software full-life-cycle intelligent collaborative optimization method and system based on a large model, belongs to the technical field of artificial intelligence and DevOps crossing, and aims to solve the technical problem of how to eliminate LLM model performance degradation caused by development, test and production environment data splitting. According to the technical scheme, the method comprises the steps that a large language model is called through a gRPC interface to generate a program code, and a requirement specification with a version mark is obtained; calling a large language model to generate a boundary value test case based on an extension plug-in of a JUnit framework, and obtaining a defect report with a severity label; an anomaly detector integrated with a Prometheus monitoring tool calls a large language model to predict a container fault, and obtains a resource utilization rate time sequence report; and constructing a structured data lake: storing development, test and operation and maintenance data by using Elasticsearch, and establishing space-time association among developers, test cases and container instances through a Neo4j graph database.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Video frame accurate positioning method based on hierarchical knowledge base

The invention discloses a video frame precise positioning method based on a hierarchical knowledge base, and the method comprises the steps: carrying out the preprocessing of each teaching video of a MOOC class in an online education platform, and extracting a representative key frame and a precise timestamp of the representative key frame in an original video; on this basis, knowledge points are extracted by means of a visual language large model, and a hierarchical knowledge base structure is constructed; then, carrying out deep semantic analysis on a natural language question input by a user by utilizing a large language model; based on an analysis result, through semantic similarity calculation and a structured reasoning mechanism, positioning a target knowledge point most related to the user intention layer by layer in a knowledge base; furthermore, a teaching video timestamp corresponding to the knowledge point is determined, and interpretable path information is generated based on a reasoning path in the positioning process. The method not only can realize accurate positioning of a key frame level, but also has traceable and explainable semantic alignment capability, and can remarkably improve intelligent understanding and positioning precision of teaching video content.
Owner:PAZHOU LAB (HUANGPU) +1

Computer vision processing system based on machine learning

The invention discloses a computer vision processing system based on machine learning, and the system comprises a dynamic topology management module which constructs and maintains a dynamic knowledge graph of a camera network in real time through a graph neural network technology, and carries out the topology self-adaption; the intelligent search module is used for dynamically selecting an optimal monitoring node for query based on a reinforcement learning technology; the multi-modal detection center is used for integrating multi-modal data, carrying out flame / smoke detection by combining an OfficientDet-D7 model with an optical flow method, integrating living body detection in a face recognition model trained in an MS1M-ArcFace data set based on an ArcFace framework to resist attacks, and carrying out multi-target tracking and cross-camera re-recognition by utilizing a DeepSORT algorithm; the distributed early warning system is used for triggering graded early warning through a dynamic threshold engine and carrying out context-aware false alarm suppression and key alarm priority response; the data persistence layer is used for storing structured data by adopting MongoDB fragmentation clusters and storing original video streams by adopting MinIO object storage; and the edge calculation node is used for locally executing lightweight calculation.
Owner:QIONGTAI TEACHERS COLLEGE

Residual frame video processing convolution method based on dynamic gating

The invention relates to the technical field of computer vision and video processing, and provides a residual frame video processing convolution method based on dynamic gating, which comprises the following steps of: carrying out gray conversion and Gaussian filtering processing on a current frame image and a previous frame image; calculating an inter-frame difference image and carrying out binarization processing to obtain a motion area mask; performing morphological expansion operation on the mask image; extracting the contour of the motion area and calculating a minimum bounding rectangle to obtain a bounding box; carrying out merging processing on the overlapped bounding boxes; updating the synthesized background frame based on the combined bounding box, and copying the pixels of the motion area of the current frame to the corresponding position of the background frame; and outputting the optimized bounding box set and the updated synthesized background frame. According to the method, the calculation efficiency and the resource utilization rate of video processing are improved, and meanwhile, the timing sequence continuity and the space consistency of a processing result are ensured through a strategy of keeping the static region unchanged and only updating the motion region.
Owner:CHANGSHA CHENGZHUO MICROELECTRONICS CO LTD

Hybrid noise removal method and system based on wavelet framework

The invention relates to the technical field of image processing, in particular to a hybrid noise removal method and system based on a wavelet framework, and the method comprises the steps: obtaining an original image, preprocessing the original image, converting the original image into a matrix, and decomposing a denoising problem into a plurality of sub-problems based on a denoising model; wherein the denoising model restrains the edge and texture structure of an image through a regular term, captures an abnormal point through a data fidelity term of impulse noise, and suppresses long-tail noise through a data fidelity term of Cauchy noise; each sub-problem is solved, the solving result of each sub-problem is applied to the next sub-problem for iteration, and when the error between the iteration results of the kth step and the (k + 1) th step is smaller than a set value, the iteration result of the last step is the restored image after noise is removed. The image is recovered by taking a summing item of the Cauchy noise and the impulse noise as a data fidelity item of the model and taking a wavelet frame as a regular item.
Owner:QINGDAO UNIV OF TECH

Information processing apparatus, information processing method, and computer-readable storage medium

The application discloses an information processing device and method based on zero-order learning and a computer readable storage medium. The information processing device comprises: an activation heat map generation unit configured to generate an activation heat map of each frame in a video based on a predetermined request by using a pre-trained multi-modal model or a pre-trained attention network; a spatial region of interest determination unit configured to determine a region of interest of each frame based on the activation heat map of the frame; and a similarity time sequence obtaining unit configured to calculate a similarity between the region of interest of each frame and the predetermined request by using the pre-trained multi-modal model to obtain a similarity time sequence of the video, which can be used to identify a target frame corresponding to the predetermined request in the video.
Owner:FUJITSU LTD

Local editing large model jailbreak attack method based on activation guidance

The invention relates to the technical field of large language models, and discloses a local editing large model jailbreak attack method based on activation guidance, which comprises the steps of data acquisition and marking work, data preprocessing and feature extraction, two-stage jailbreak attack based on an AGILE framework, a generation stage and an editing stage, and generation of a final jailbreak prompt. Malicious queries are converted into hidden jail break prompts based on a two-stage framework, and efficient jail break attacks are realized by guiding editing through activation signals in a model. According to the method, through AGILE two-stage framework design, in the generation stage, multiple rounds of dialogue history H are constructed at a time by means of a generator LLM, and in the editing stage, activation and attention scores are utilized to guide subtle and local editing of a generated text, so that two-stage decoupling design is achieved, the expandability of attacks is effectively improved, and the method has the advantages of being high in practicability and easy to popularize. In the generation stage, a generator LLM only needs to be called once, and efficient optimization is achieved by means of lightweight operation in the editing stage.
Owner:HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1

Correction method for virtual scene displacement missing for mixed reality

The invention provides a virtual scene displacement missing correction method for mixed reality. The method comprises the following steps: acquiring respective environment association information of a current frame and a previous frame of the current frame; determining whether a displacement correction operation for the current frame is triggered or not before the current frame is displayed based on the environment association information; when it is determined that the displacement correction operation for the current frame is triggered, obtaining a segmentation index of the current frame, and querying a displacement proportionality coefficient and a rotation proportionality coefficient corresponding to the segmentation index in a preset parameter statistics calibration table; determining a predicted displacement of the current frame based on the displacement proportionality coefficient, and determining a predicted rotation deviation of the current frame based on the rotation proportionality coefficient; determining an updated pose of the current frame based on the current frame, the previous frame, the predicted displacement and the predicted rotation deviation; and if the updated pose meets the preset correction amplitude limiting condition, performing displacement correction processing on the current frame based on the updated pose. According to the scheme, prediction and advanced compensation of a positioning mutation problem caused by environment mutation are realized.
Owner:BEIJING ZHIHUI HUANYU TECHNOLOGY CULTURE CO LTD

VVC reconstruction frame post-processing method based on large model

The invention relates to a VVC reconstructed frame post-processing method based on a large model, and belongs to the technical field of video coding and image processing. In order to solve the problems of compression artifacts and distortion existing in a VVC coding reconstruction frame, preliminary distortion suppression is performed through a loop filtering module, and coding side meta information is obtained; the input preprocessing module extracts visual features and meta-information features; the large model processing module fuses features by using a convolution attention mechanism, and realizes local enhancement and global modeling through a lightweight window self-attention and depth separable convolution hybrid network; and the output post-processing module adopts adaptive residual adjustment to generate a high-quality reconstructed frame. According to the method, the visual quality of the reconstructed frame can be remarkably improved, accurate self-adaptive repair is realized, good balance is achieved between high efficiency and light weight, and practical application deployment is facilitated.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Traditional art 3D content conversion method based on artificial intelligence

The invention discloses a traditional art 3D content conversion method based on artificial intelligence, and relates to the technical field of artificial intelligence, and the method comprises the following steps: constructing a cross-scene semantic anchor point graph, generating a verifiable semantic anchor point signature for the character features and background materials in a traditional art image, and carrying out the verification of the semantic anchor point signature; synchronously establishing a space-time constraint baseline based on space coordinates and a time sequence; and before scene switching, loading a millisecond preheating frame based on a space-time constraint baseline, executing time sequence calibration on the semantic anchor point signature, and outputting a phase reference for controlling a subsequent mapping process. According to the method, through semantic anchor point signature, space-time constraint, preheating frame calibration, hierarchical binding, mapping snapshot, texture springback, residual error remapping and a man-machine resonance regulation and control mechanism, precise binding and dynamic stability of semantic tags in a three-dimensional space are achieved, and the reduction degree of traditional art 3D conversion and the immersion continuity of multi-scene interaction are improved.
Owner:CHAOHU UNIV

Structure annotation

A computer-implemented method of creating one or more annotated perception inputs, the method comprising, in an annotation computer system: receiving a plurality of captured frames, each frame comprising a set of 3D structure points, in which at least a portion of a common structure component is captured; computing a reference position within at least one reference frame of the plurality of frames; generating a 3D model for the common structure component by selectively extracting 3D structure points of the reference frame based on the reference position within that frame; determining an aligned model position for the 3D model within a target frame of the plurality of frames based on an automatic alignment of the 3D model with the common structure component in the target frame; and storing annotation data of the aligned model position in computer storage, in association with at least one perception input of the target frame for annotating the common structure component therein.
Owner:FIVE AI LTD

Self-supervised optimization and adaptive deployment method for optical flow network

The invention discloses a self-supervised optimization and self-adaptive deployment method for an optical flow network, and the method enables the effective region endpoint of a KITTI data set and the error of a self-built complex scene data set to be effectively reduced through the collaborative design of a self-supervised optimization method and a self-adaptive computing power frame based on the accumulation and circulation consistency constraint of a reverse optical flow, and improves the precision. A feature multiplexing mechanism enables the inference speed of embedded equipment to be improved and the memory occupation to be reduced, a dynamic pyramid iteration module supports flexible balance precision and time delay improvement efficiency according to a computing power threshold, a progressive training strategy effectively inhibits error accumulation, the cross-domain migration precision is improved, and the optical flow direction estimation accuracy in a low-resolution scene is improved. The problem of performance degradation caused by a fixed structure in the prior art is solved, and an optical flow estimation solution considering real-time performance and high precision is provided for mobile platforms such as an unmanned aerial vehicle.
Owner:EAST CHINA INST OF COMPUTING TECH

Multi-target tracking method for self-adaptive threshold and buffer association

The invention provides a multi-target tracking method based on self-adaptive threshold and buffer correlation, which relates to the technical field of computer vision and comprises the following steps: calculating a self-adaptive confidence threshold of each frame of image, dividing a detection frame of each frame of image into a high-score detection frame and a low-score detection frame, generating an initial motion track set, and tracking the initial motion track set according to the self-adaptive confidence threshold; performing two-stage matching on the detection frame of each frame of image and the tracks in the initial motion track set, and generating a motion track set of each frame of image based on a matching result, the motion track set including a plurality of matched tracks; and generating a target tracking result of each frame of image based on the movement track set of each frame of image. According to the method, an adaptive confidence threshold mechanism and a buffer intersection-union matching strategy are introduced, so that the threshold can change in real time according to the scene, and limitation caused by a fixed threshold is avoided. And a buffer intersection-parallel ratio is introduced, the boundary of the low-score detection frame is properly expanded, and the expanded buffer frame is still possibly overlapped with the trajectory prediction frame sufficiently, so that successful matching is realized.
Owner:NORTHEASTERN UNIV CHINA

Video processing method, apparatus and electronic device

Embodiments of the present application provide a video processing method, device and electronic equipment, relating to the field of artificial intelligence, the video processing method is applied to a pre-trained video processing model, comprising: extracting a spatio-temporal feature of a video frame; the video frame comprises a current frame to be processed and a reference frame adjacent to the current frame; the spatio-temporal feature comprises a dynamic foreground feature, a noise level feature and a coding mode feature; the coding mode feature is used to represent the coding mode of each macroblock of the corresponding video frame; fusing the spatio-temporal feature of the current frame and the spatio-temporal feature of the reference frame to obtain a target spatio-temporal feature of the current frame; determining a filtering mode corresponding to the current frame according to the target spatio-temporal feature, and filtering the current frame based on the filtering mode. The present application can protect and retain the key information in the video, which is beneficial to more accurately and completely remove the noise and spatio-temporal redundancy in the video.
Owner:PEKING UNIV +1

Target tracking method based on two-stage selection

The invention relates to a target tracking method based on two-stage selection, and the method comprises the following steps: obtaining an input video stream sequence, extracting image features for a current frame, and generating at least one candidate mask and a confidence score corresponding to each candidate mask based on memory bank information; predicting the motion state of the target in the current frame based on a motion prediction model, and calculating a coincidence evaluation score between each candidate mask and the predicted motion state; selecting a first candidate result from the candidate masks based on the confidence score and the anastomosis evaluation score, and judging whether the anastomosis evaluation score of the first candidate result meets a preset stability condition or not; and if yes, determining the first candidate result as a final segmentation result of the current frame, and if not, calculating a similarity evaluation score between each candidate mask and a historical segmentation result stored in a historical state library, and selecting a second candidate result from the candidate masks as the final segmentation result of the current frame based on the similarity evaluation score.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1

structural modeling

This invention relates to a computer implementation method for modeling a common structural component. The method includes: in a modeling computer system, receiving a plurality of capture frames, each frame including a set of 3D structural points, wherein at least a portion of the common structural component is captured; calculating a first reference position within at least one first frame in the plurality of frames; selectively extracting first 3D structural points of the first frame based on the first reference position calculated for the first frame; calculating a second reference position within a second frame in the plurality of frames; selectively extracting second 3D structural points of the second frame based on the second reference position calculated for the second frame; and aggregating the first 3D structural points and the second 3D structural points to generate an aggregated 3D model of the common structural component based on the first reference position and the second reference position.
Owner:FIVE AI LTD

A method and system for determining whether a beaker experiment is scored

The application provides a method and system for judging whether a beaker experiment is successful, comprising: obtaining a video stream of the beaker experiment and performing region of interest detection; a plurality of images containing the region of interest obtained by the region of interest detection constitute a picture set of a region of interest period of the beaker experiment; a full convolutional neural network (FCN) model is trained based on the picture set of the region of interest period of the beaker experiment; the picture set of the region of interest period is subjected to strong supervision target classification frame by frame based on the trained FCN model, and a classification result of each frame in the picture set of the region of interest period is obtained; and whether the experiment operation in the video stream is successful is determined based on the classification result of each frame in the picture set of the region of interest period. The data required by the application is greatly reduced, and the application is suitable for video data size in actual experiment operation process; and the target region can be quickly focused on by detecting the region of interest, and the detection precision is effectively improved.
Owner:SHANGHAI MEDIA INTELLIGENCE TECH CO LTD