Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

895 results about "Key frame" patented technology

A keyframe in animation and filmmaking is a drawing that defines the starting and ending points of any smooth transition. The drawings are called "frames" because their position in time is measured in frames on a strip of film. A sequence of keyframes defines which movement the viewer will see, whereas the position of the keyframes on the film, video, or animation defines the timing of the movement. Because only two or three keyframes over the span of a second do not create the illusion of movement, the remaining frames are filled with inbetweens.

Long-time pedestrian re-identification method based on dual-path cooperation and key frame guided reconstruction

The invention discloses a long-time pedestrian re-identification method based on dual-path cooperation and key frame guided reconstruction. The method comprises the steps of firstly collecting a pedestrian video to be recognized, and extracting a video feature sequence; space and time position coding is introduced into the video feature sequence; capturing local fine-grained dynamic features through a local dynamic feature capturing path, and modeling long-range time sequence association through a cross-frame global feature modeling path; then, dual-path feature complementation is realized through bidirectional gating interaction; further screening out key frames, and realizing feature reconstruction through a full-frame attention propagation mechanism; and finally fusing the dual-path fusion features, the key frame guide reconstruction features and the refined features to generate pedestrian identity features. And processing pedestrian identity features to obtain standardized feature vectors, performing similarity comparison on the standardized feature vectors and pedestrian features in an image library, and returning a matching list. According to the method, video time sequence information is fully utilized, and the problem of insufficient robustness caused by appearance change in long-time pedestrian re-identification is effectively solved.
Owner:SHIJIAZHUANG TIEDAO UNIV

Animation video generation method and system based on video script

The invention discloses a cartoon video generation method and system based on a video script, and relates to the technical field of video synthesis, and the method comprises the steps: based on structured script data, combining a predefined lens rule base and a reinforcement learning model, determining the type, duration and lens operation effect of a lens, and generating a lens splitting sequence through a dynamic lens splitting automatic generation mechanism, based on the split mirror sequence, generating an animation style key frame image through a diffusion model, selecting action data matched with the emotion label through a predefined action library, generating a voice waveform matched with the emotion label through a text-to-voice model, selecting a background music audio matched with the emotion label through a music library, and generating a multi-modal content stream; according to the method, the lens type, the duration and the lens operation effect are dynamically optimized through the reinforcement learning model, the defects of a traditional lens scheduling method based on static rule mapping in time continuity and narrative continuity are overcome, and the narrative fluency and the dynamic adaptability of the lens division sequence are obviously improved.
Owner:CHANGCHUN VOCATIONAL INST OF TECH

Intelligent conference video frame dynamic coding method based on multi-mode semantic understanding

The invention relates to the technical field of computer vision, in particular to an intelligent conference video frame dynamic coding method based on multi-modal semantic understanding, which comprises the following steps: acquiring a video stream sequence and a synchronous audio stream in a conference scene in real time; performing semantic analysis and decoupling on the video stream sequence, and extracting key frames and subsequent frames; extracting a sparse motion field from a subsequent frame, and segmenting a video frame into candidate visual areas including a face, a mouth shape and a background; extracting audio semantic features, executing cross-modal semantic correlation analysis, calculating semantic correlation between the sparse motion field distribution features and the audio semantic features, and positioning a pronunciation area highly related to the voice content; and calculating a quantization offset value of each candidate visual area according to the semantic relevancy, applying the quantization offset values in different areas, and packaging the quantization offset values into a variable-code-rate video code stream. According to the invention, the multi-mode semantic understanding model is constructed to carry out deep semantic analysis on the video frame content so as to realize the dynamic coding of the conference video frame.
Owner:SHENZHEN JIKEYUAN ELECTRONIC TECH CO LTD

Video text cross-modal retrieval method based on spatio-temporal feature fusion

The invention relates to the field of artificial intelligence cross-modal retrieval, and provides a video text cross-modal retrieval method and system based on spatio-temporal feature fusion. The method comprises the following steps: carrying out key frame sampling and time sequence partitioning on an input video, extracting static visual features through a spatial feature network, and extracting motion features through a time dynamic network; a self-adaptive gating fusion module is adopted to dynamically calculate spatial-temporal feature weights and perform weighted fusion; extracting text semantic features by using a pre-training language model; constructing a double-flow projection network to map video fusion features and text features to a unified measurement space, and optimizing a feature distance by adopting a contrast loss function containing difficult negative sample mining and intra-modal constraint; and outputting a retrieval result according to the cosine similarity sequence. The system comprises four units, wherein the gating fusion module is integrated with an FPGA acceleration circuit. According to the method, mAP (at) 10 is equal to 0.78 in a UCF-101 data set, the time sequence action retrieval accuracy rate is 92.8%, and the single video retrieval delay is 23 milliseconds.
Owner:ZHEJIANG UNIV

Prompt construction method and system of multi-mode large language model, computer equipment and medium

The invention relates to the technical field of multi-modal large language model training, in particular to a prompt construction method and system for a multi-modal large language model, computer equipment and a medium. The method comprises the following steps: extracting a key frame set from an input video stream; executing a motion reconstruction process on the video stream to generate motion track information; and performing visualization processing on the motion track information to generate a track visualization graph. Performing space-time correlation coding on the key frame set and the motion track information to generate an enhanced key frame; a multi-modal prompt is constructed in a mode of integrating visual input and text input, and the multi-modal prompt is input into a preset multi-modal large language model for spatial reasoning. Through the mode, the technical problem that an existing prompting method is difficult to give consideration to the spatial reasoning precision and the calculation efficiency is solved, efficient and accurate spatial reasoning of the multi-modal large language model is achieved, and the calculation efficiency, the reasoning precision and the environmental adaptability of the model are improved.
Owner:HONG KONG UNIV OF SCI & TECH (GUANGZHOU)

Multi-modal fusion key frame extraction method and device, equipment and medium

The invention relates to the technical field of computers, and discloses a multi-modal fusion key frame extraction method and device, equipment and a medium, and the method comprises the steps: obtaining multi-modal input data, carrying out the modal feature coding, and obtaining a video modal feature, an audio modal feature and a text modal feature; carrying out attention fusion on the video modal features, the audio modal features and the text modal features to obtain fused cross-modal causal features; analyzing the cross-modal causal features through a causal reinforcement learning decision module in combination with a preset time sequence causal graph to obtain a fusion feature sequence and key frame probability distribution; and carrying out key frame selection operation on the time slice of the fusion feature sequence based on the key frame probability distribution to obtain a key frame set, and generating a space-time thermodynamic diagram and causal relationship visualization result corresponding to the key frame set. The multi-modal fusion key frame extraction method and device can be applied to financial science and technology or medical care service program systems, and the accuracy and interpretability of multi-modal fusion key frame extraction can be improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Underground pipeline intelligent detection and mapping method based on image recognition

The invention discloses an underground pipeline intelligent detection and mapping method based on image recognition, and the method comprises the following steps: 1, obtaining continuous video images, associating feature matching pairs of adjacent key frames, and forming a pose parameter set; 2, inputting the key frame into an improved YOLO-World detection network, and outputting a detection result set; 3, obtaining a geometric consistency matching set according to the detection result set; 4, performing multi-view triangularization on the geometric consistent matching set to form a weight factor; 5, introducing a weight factor, and executing incremental beam adjustment optimization on the pose parameter set and the three-dimensional sparse point set to obtain a sparse semantic point cloud; and step 6, outputting an underground pipe network topology map. According to the invention, high-precision and high-robustness intelligent identification and topological mapping in a complex underground pipeline environment are realized.
Owner:WUXI YIXING POWER TECH CO LTD

Long video understanding method capable of relieving time sequence illusion in video language large model

The invention provides a long video understanding method capable of relieving time sequence illusion in a video language large model. The long video understanding method is based on a static bias adaptive frame selection mechanism and a cross-modal feature fusion strategy. According to the static bias mechanism, inter-frame similarity is evaluated through a discriminator, redundant frames are identified, key frames are selected or a complete sequence is reserved, so that calculation overhead is reduced, and spatio-temporal information integrity is kept; a video frame and a text are mapped to a shared semantic space, the single-frame semantic understanding ability is enhanced, then an embedded sequence serves as a soft prompt to be input into a large language model, and a final answer is generated in an autoregression mode. According to the method, the efficiency and accuracy of long video understanding and video question and answer tasks can be remarkably improved; the problem of low training and reasoning efficiency caused by time sequence dependence redundancy and excessive computing resource consumption is effectively relieved; and through a dynamic multi-modal task processing framework and a space-time memory bank compression mechanism, the modeling capability and generalization performance of the model on a long video sequence are further improved.
Owner:LANZHOU UNIV

Dynamic scene robust visual SLAM method based on multi-feature collaborative optimization

The invention discloses a dynamic scene robust vision SLAM (Simultaneous Localization and Mapping) method based on multi-feature collaborative optimization, which comprises the following steps of: acquiring an image sequence, carrying out dynamic target detection and segmentation through an instance segmentation network, generating a segmentation mask containing a dynamic region mark, and identifying and separating a dynamic object and a static background; removing feature points corresponding to the dynamic object based on the segmentation mask to obtain static feature points; carrying out pose estimation based on the static feature points, and for the key frame, carrying out feature matching with the previous key frame by minimizing a re-projection error, and solving to obtain the camera pose of each key frame; for non-key frames, performing camera pose tracking and data association on the previous frame by adopting an optical flow algorithm, and accumulating solving results to obtain pose tracks of all the non-key frames; the key frames and the non-key frames are subjected to differential processing by fusing feature matching and an optical flow algorithm, so that the calculation efficiency is remarkably improved while the positioning precision is ensured, and the real-time performance is improved.
Owner:INNER MONGOLIA UNIVERSITY

A method, system, medium, device and terminal for monitoring video key frame extraction

The application belongs to the technical field of multimedia information processing, and discloses a kind of monitoring video key frame extraction method, system, medium, equipment and terminal, collect original video stream data, original video stream data is decomposed into image frame set;Image frame set obtained by decomposition is sampled, and the image frame result set obtained by sampling is filtered;The image frame set after filtering is adaptively clustered, and the result after clustering is collected to form a video abstract.In order to better utilize the memory space of storage medium and let the user better quickly browse the general content of original video stream, the application provides a key frame screening algorithm for original video stream by sampling, filtering and clustering.The key frame extraction method of the application screens out similar frames, redundant frames and fuzzy frames in original video stream through key frame screening algorithm for original video data, forms a video abstract storage of a section of original video, thereby greatly reduces the occupied storage space.
Owner:QINGDAO INST OF COMPUTING TECH XIDIAN UNIV

Video question-answering system and method based on iterative multi-mode

The invention provides a video question-answering system and method based on iterative multi-mode, and the method comprises the steps: carrying out the preprocessing of an input original video file and natural language query, extracting a key frame sequence, and generating an initial subtitle sequence; performing multi-granularity retrieval based on the preprocessed natural language query and the current subtitle sequence, and determining a candidate region; carrying out fine-grained frame selection in the candidate region by using a large language model, and identifying a key frame; based on the candidate area and natural language query, determining the type of the visual information to be supplemented and generating a multi-modal cue word corresponding to the type, and extracting the visual information by the visual language model according to the multi-modal cue word to update the subtitle sequence of the candidate area; generating a prediction answer by adopting a large language model and a visual language model; and judging the confidence of the generated predicted answer, and outputting a final answer. According to the method, the processing mode of video understanding can be optimized, and an accurate cross-modal coordination solution is provided through dynamic reasoning-sensing coordination.
Owner:EVALUATION & DEMONSTRATION RES CENT OF THE CHINESE PEOPLES LIBERATION ARMY ACAD OF MILITARY SCI

Method, system and terminal for classifying echocardiography videos

The invention discloses an echocardiogram video classification method, system and terminal, and the method comprises the steps: constructing a classification model network which comprises a feature extraction module, a feature enhancement module and a feature aggregation module; obtaining an echocardiogram video, obtaining a plurality of standard section views according to the echocardiogram video, performing interpolation processing and feature extraction on the plurality of standard section views through a feature extraction module, and outputting a plurality of video features; inputting the plurality of video features into a feature enhancement module for aggregation enhancement of spatial features and time sequence features, and outputting a plurality of enhanced features; and inputting the plurality of enhanced features into a feature aggregation module for frame-level feature weighted fusion to obtain a plurality of key frame features, selecting related features from the plurality of key frame features, obtaining fusion features according to the related features, and classifying the fusion features to obtain a classification result of the echocardiogram video. According to the method, the classification accuracy of the echocardiogram videos is effectively improved.
Owner:SHENZHEN CHILDRENS HOSPITAL

Camera state judgment method based on road surface covering and Hash comparison

The invention relates to the technical field of intelligent traffic, and discloses a camera state judgment method based on road surface covering and Hash comparison, and the method sequentially comprises the steps: extracting key frame pairs from a video stream at intervals; performing road surface region extraction on each frame of image to obtain a road surface mask; carrying out covering processing on the road surface area in each frame of image based on the mask to obtain a covered image; calculating a difference hash feature of each frame of covered image to obtain a hash code; calculating the Hamming distance between the Hash codes of the key frame pair; and judging whether the camera is in a rotating or stable state according to a comparison result of the Hamming distance and a threshold value. The method only depends on the video image data, does not need an external sensor, effectively eliminates the dynamic interference of the vehicle through covering the road surface, achieves the efficient and robust judgment of the state of the camera through combining the lightweight difference hash calculation, is low in calculation cost, and is suitable for edge equipment and complex road environments.
Owner:GUANGZHOU GUOJIAO RUNWAN TRAFFIC INFORMATION CO LTD

Cross-border logistics information credible management method based on block chain and dynamic hash chain evidence storage

The invention belongs to the technical field of cross-border logistics information credibility management, and discloses a cross-border logistics information credibility management method based on block chain and dynamic hash chain evidence storage, and the method comprises the steps: deploying a multi-modal data adapter, integrating a lightweight MobileViT model and a category exclusive feature template, extracting a key frame through a video, and generating a feature package; after sensor data is subjected to Kalman filtering processing, a timestamp and a numerical sequence are output; a historical feature library is compared through an LSH algorithm, repeated features are filtered, evidence storage redundancy is avoided, unstructured data features are associated with structured data such as customs clearance numbers, accurate data correspondence during traceability is ensured, and a verification result is real and reliable; for a temporary checking scene, a temporary node generates a temporary permission containing a 24-hour validity period, and a corresponding Hash unit can only be verified and cannot be accessed after expiration, so that an illegal data access risk is completely eradicated, subsequent verification is prevented from being influenced by excessive permission recovery, and data security and credible verification requirements are balanced.
Owner:中武(福建)跨境电子商务有限责任公司

Key frame animation display method and device, terminal, storage medium and product

The embodiment of the invention discloses a key frame animation display method and device, a terminal, a storage medium and a product, and relates to the field of human-computer interaction. The method comprises the steps that at least one of a first attribute value and a second attribute value is determined based on scene information, the first attribute value is the attribute value of an animation element in a first key frame, and the second attribute value is the attribute value of the animation element in a second key frame; at least one of the first attribute value and the second attribute value supports dynamic updating based on scene information; interpolating the first attribute value and the second attribute value to generate a key frame animation between the first key frame and the second key frame; and displaying the key frame animation. By adopting the method provided by the invention, the key frame animation matched with the scene can be displayed in different scenes, and the development cost of the animation is reduced.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Knowledge graph-based long video key frame retrieval method and device

The invention relates to the technical field of multi-mode intelligent video understanding, and provides a long video key frame retrieval method and device based on a knowledge graph. Through the processes of frame-level subtitle generation, frame-level knowledge graph construction, similarity video segmentation and fragment and abstract generation, a long video knowledge graph construction assembly line is constructed, and structured modeling of long video semantic content is realized. By setting a two-stage retrieval mechanism, higher retrieval precision can be obtained while the efficiency is ensured. Vector matching and multi-hop neighbor extension are carried out on a unified knowledge graph, and a node set strongly related to the problem is positioned, so that the semantic gap between the natural language problem and a structured graph is reduced, and the accuracy of key frame selection is improved. By setting an iterative retrieval mechanism, the retrieved key frame can be used as a basis for answering a question text to the maximum extent.
Owner:NAT UNIV OF DEFENSE TECH

Three-dimensional animation generation method and related equipment

The invention discloses a three-dimensional animation generation method and related equipment, and the method comprises the steps: carrying out the segmentation processing of a target animation, and obtaining the head and tail key frames of each segment of the target animation; establishing a corresponding relationship between the head and tail key frames and each segment obtained by segmentation; analyzing the head key frame and the tail key frame through a large model to generate prompt word information; and finally, inputting the head and tail key frames and the cue word information into a video generation model to obtain a fragment animation of each fragment. And carrying out splicing processing on the fragment animations to generate a three-dimensional animation. Through division of the key frames and the animation segments, the rendering calculation amount is reduced, key frame dependence can be reduced, the generation period can be shortened, and the method can be widely applied to the technical field of computers.
Owner:广州极点三维信息科技有限公司

Video anomaly detection method based on traffic scene

The invention discloses a video anomaly detection method based on a traffic scene, and relates to the related field of computer vision, and the method comprises the steps: obtaining an input video clip which comprises a plurality of frames of images, traversing the plurality of frames of images, carrying out the selection through combining with a text prompt, determining a plurality of key frames, and enabling the text prompt to be input by a user; performing context generation based on the plurality of key frames to obtain position and time context information; the key frame, the position, the time context information and the text prompt are synchronized to a large visual language model for visual questions and answers of diversified traffic scenes, abnormal events are extracted, and the abnormal events comprise abnormal scores; and carrying out abnormal event detection analysis according to the abnormal score, and carrying out abnormal detection on the diversified traffic scenes. The technical problem that abnormal events in diversified traffic scenes are difficult to comprehensively and accurately detect in existing traffic scene-based video anomaly detection is solved, and the technical effect of improving the accuracy and generalization of traffic scene video anomaly detection is achieved.
Owner:AIPARK TECHNOLOGY CO LTD

Underwater concrete scene three-dimensional reconstruction method, system and device and storage medium

The embodiment of the invention provides an underwater concrete scene three-dimensional reconstruction method, system and device and a storage medium, and relates to the technical field of underwater three-dimensional reconstruction, and the method comprises the steps: extracting feature points in each key frame image from a plurality of key frame images of video data, and obtaining feature matching point pairs; calculating relative poses of adjacent key frame images and three-dimensional coordinates of the matched feature points according to the pixel coordinates of the same matched feature point; calculating a pixel depth value of each pixel point according to the relative pose; and performing multi-view fusion according to the relative pose, the three-dimensional coordinates of the matched feature points and the pixel depth value of each pixel point, and generating a three-dimensional grid model of the underwater concrete scene. According to the method, the feature points in the enhanced video image are obtained, the moving distance and the translation vector of the shooting camera are calculated, and three-dimensional reconstruction is carried out in combination with the pixel depth value, so that the image acquisition quality is improved, point cloud data are saved, and the precision and efficiency of three-dimensional reconstruction of a large-range underwater scene are improved.
Owner:TIANFU YONGXING LAB

Video content auditing method, device and equipment

The embodiment of the invention provides a video content auditing method, device and equipment. According to the scheme, the method comprises the steps of obtaining first multi-modal data of a to-be-audited video, wherein the first multi-modal data comprises a key frame image and a voice recognition text of the to-be-audited video; multi-modal features of the first multi-modal data are extracted; based on retrieval enhancement generation, retrieving a plurality of target auditing tags for judging whether the to-be-audited video is illegal or not from a rule knowledge base according to the multi-modal features; calculating a matching score of the multi-modal feature and each target auditing tag by using a first multi-modal large language model; and obtaining a first auditing result by using a video auditing model based on the first multi-modal data, the target auditing tags and the matching scores corresponding to the target auditing tags.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Video key frame extraction method fused with self-supervised deep learning

The invention discloses a video key frame extraction method fused with self-supervised deep learning, and the method comprises the following steps: carrying out the standardized sampling of an input video according to a fixed interval, extracting a SuperPoint local key point and a Video MAE global semantic feature, and generating a dense descriptor and a semantic vector; constructing multi-dimensional change indexes such as local matching degree and global similarity; constructing a soft distribution matrix and a matching point set based on the fusion features; dynamically judging the key frame through a self-adaptive multi-threshold rule; and outputting the key frame set. The method fuses local and global spatio-temporal information, has robust feature extraction and key frame discrimination capabilities under a weak supervision condition, and can effectively improve the efficiency and precision of video compression, abstract and event detection.
Owner:GUANGDONG POLYTECHNIC OF IND & COMMERCE

Traffic accident scene reconstruction method and device

The invention provides a traffic accident scene reconstruction method and device, and relates to the technical field of three-dimensional reconstruction, and the method comprises the steps: obtaining a video frame sequence of a traffic accident scene collected by an unmanned plane, carrying out the information gain evaluation of the video frame sequence, and constructing a key frame set; according to the key frame set, macroscopic geometric reconstruction and microscopic normal recovery based on airborne light source dynamic change are carried out on the traffic accident scene, and a macroscopic depth map and a microscopic normal field are obtained; constructing a fusion optimization model with the macroscopic depth map as low-frequency constraint and the microscopic normal field as high-frequency gradient guidance, and fusing the macroscopic depth map and the microscopic normal field by using the fusion optimization model to obtain a fused depth map; and constructing an accident scene reconstruction model according to the fused depth map. By adopting the traffic accident scene reconstruction method and device, key details of the accident scene can be captured, the sensitivity to micromorphology is increased, and the accuracy of traffic accident scene reconstruction is improved.
Owner:ZHEJIANG EXPRESSWAY CO LTD +1

Video reasoning training data generation method based on key frame causal chain extraction

The invention discloses a video reasoning training data generation method based on key frame causal chain extraction, and the method comprises the steps: obtaining a to-be-processed video data set, and decoding the to-be-processed video data set to obtain a frame sequence; performing redundancy elimination screening on the frame sequence to obtain a key frame sequence; performing semantic structured representation on the key frame sequence to generate key frame description and entity information, and generating an event candidate set based on the key frame description and the entity information; executing multi-event causal discovery based on the event candidate set to obtain an event-level causal graph, and extracting at least one key frame causal chain from the event-level causal graph; generating a video reasoning training sample based on the key frame causal chain; and performing quality verification and screening on the video reasoning training sample, and outputting a screened training data set. According to the method, through offline key frame index multiplexing and causal structure constraint, the frame-by-frame calculation and labeling cost is reduced, data illusion is inhibited, and the verifiability of reasoning training signals is improved.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Vibration-illumination double dynamic compensation fan blade cracking early warning method and system

The invention relates to the technical field of fan blade health monitoring, in particular to a vibration-illumination double dynamic compensation fan blade cracking early warning method and system. The method comprises the following steps: triggering and capturing a video key frame at a key angle through a rotation angle sensor, and carrying out image coarse alignment by utilizing SIFT feature matching; a corrected motion vector field is generated by combining physical parameters of the rotating speed and the angular acceleration, vibration time sequence compensation is carried out, and illumination time sequence compensation is completed through Gamma correction by taking a clear frame as a reference; an improved DeepSORT algorithm is adopted to associate crack tracks with crack mask contour similarity, evolution characteristics are extracted and fused with SCADA working condition data, an expansion rate model is established, and the remaining life is predicted; early warning grades are divided according to the service life and the expansion rate, four-dimensional operation and maintenance data are fused to output differential operation and maintenance strategies, and model parameters are optimized through feedback data. According to the invention, the accuracy and early warning reliability of crack detection in a complex environment are effectively improved.
Owner:BEIJING YINGHUADA POWER ELECTRONICS ENG TECH CO LTD

Visual token compression method and device, computer equipment and storage medium

The embodiment of the invention belongs to the technical field of artificial intelligence, and relates to a visual token compression method and device, computer equipment and a storage medium, and the method comprises the steps: receiving a visual token compression request which is sent by a user terminal and carries an original video frame; performing key frame screening operation on the original video frame to obtain a key video frame; inputting the key video frame into a visual language model to generate a visual token; performing visual token screening operation on the visual token according to the attention of the text instruction to the visual area and the content importance of the visual area to obtain a target visual token; and outputting the target visual token to the user terminal. The method can be used for carrying out related video processing in service systems such as medical treatment, health and pension, and can improve the processing efficiency, accurately extract key information and enhance the interactivity.
Owner:PING AN TECH (SHENZHEN) CO LTD

Screen content video quality evaluation method and device based on frequency-space complementation and semantics

The invention discloses a screen content video quality evaluation method and device based on frequency-space complementation and semanteme, and relates to the field of computer vision, and the method comprises the steps: S1, extracting a video block and a key frame of a screen content video, and inputting the key frame into a high-frequency structure texture information extraction branch to obtain high-frequency structure texture information; s2, inputting the key frame into a noise sensing module to obtain a noise sensing feature; s3, inputting the noise perception features into a self-adaptive time sequence embedding module to obtain noise and semantic information; s4, splicing the high-frequency structure texture information and the noise and semantic information, performing quality regression to obtain a quality score of a single-frame key frame, and summing and averaging to obtain a spatial domain video quality score; s5, inputting the video blocks into a Fast-VQA-based quality evaluation branch to obtain a time sequence distortion perception degradation score; and S6, dynamically fusing the spatial domain video quality score and the time sequence distortion perception degradation score to obtain a final video quality score. According to the method provided by the invention, the screen content video quality is effectively evaluated.
Owner:XIAMEN UNIV OF TECH +1

Hybrid coding processing method, system and equipment based on video engine

The invention relates to the field of hybrid coding of video engines, and provides a hybrid coding processing method, system and equipment based on a video engine, which comprises the step of intelligently switching an H.265 inter-frame predictive coding mode and an MJPEG (Multijoint Joint Photographic Experts Group) intra-frame compression coding mode by dynamically monitoring the proportion of a motion area in a video picture. When a scene is static, discrete cosine transform is adopted to compress space redundancy, motion vector compensation is started to eliminate time redundancy when motion is violent, and meanwhile, key frames are doubly screened by using pixel difference analysis and a perceptual hash algorithm, and repeated frames of visual redundancy are eliminated. And finally, space-time association is established for the optimized double-code-stream data through a timestamp index system, and efficient storage and accurate reconstruction of the mixed code stream are realized. The problems that the coding efficiency is low, no self-adaptive coding mode exists, redundant frames are not optimized, and code stream management is insufficient are solved.
Owner:SICHUAN SILICON MICROELECTRONICS TECHNOLOGY CO LTD

Coding method and system for on-site video return

The invention provides an on-site video return coding method and system, and the method comprises the steps: collecting a video stream of an on-site scene, judging whether a current video frame in the video stream responds to a scene switching video frame or not, and marking the current video frame as a key frame when the current video frame responds to the scene switching video frame; after the current video frame is identified as a key frame, visual saliency analysis is performed on the key frame, and an entropy weight matrix representing regional information importance distribution in the key frame is constructed according to a visual saliency analysis result and texture information entropies of different blocks in the key frame; adjusting code rate allocation weights of different blocks in the key frame according to the entropy weight matrix and a code rate regulation and control strategy of the key frame to obtain adaptive coding configuration adaptive to content characteristics of the key frame; and performing optimization coding on the key frame through adaptive coding configuration to obtain a target code stream in response to a field video return demand. By adopting the scheme of the invention, the dynamic differential coding of the key video frame in a complex scene can be realized.
Owner:SHENHUA RAIL & FREIGHT WAGONS TRANSPORT

Karwning detection method based on GFFY-YOLO model and key frame selection algorithm

The invention relates to the technical field of fatigue behavior detection, and provides a GFFY-YOLO-based yawning detection model: on the basis of YOLOv11, an efficient neck global feature fusion network A-GFPN is designed, and a WIoU mechanism is introduced into a loss function part, so that the error of a positioning regression frame is reduced, and the operation speed of the model is improved while high precision is ensured. The method is reasonable and feasible, the key frame reflecting the fatigue state change of the driver can be efficiently and accurately extracted, the detection precision of fatigue behaviors such as yawning and the system response speed are effectively improved, and the calculation burden is greatly reduced while the real-time requirement is met. The method simplifies the video preprocessing process, has good adaptability and expansibility, can flexibly cope with processing challenges brought by different scenes and parameter changes, is suitable for various vehicle-mounted application scenes such as an intelligent cockpit and an ADAS system, has remarkable supplementing and improving effects on an existing fatigue detection technology, and has good application prospects. Wide market prospects and application values are realized.
Owner:HENAN UNIV OF SCI & TECH

A funnel spout and beaker position detection method, system, terminal and medium

The application provides a funnel spout and beaker position detection method, system, terminal and medium, comprising: acquiring a video of experimental operation; extracting any image in the video and inputting the image into a ResNet18 model to obtain the position relationship of the funnel spout and the beaker on both sides; inputting the key frame of the video into a CRNN model to obtain the drop state of the funnel spout; and judging whether the funnel spout and the beaker are attached to the wall according to the position relationship of the funnel spout and the beaker on both sides and / or the drop state of the funnel spout. The application performs multi-directional detection and judgment, avoids misjudgment, greatly improves the test accuracy, and saves a large amount of manpower cost.
Owner:SHANGHAI MEDIA INTELLIGENCE TECH CO LTD