Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2599 results about "Video image" patented technology

Three-dimensional dynamic scene reconstruction method and apparatus, and storage medium

The present disclosure relates to the field of computer vision and discloses a three-dimensional dynamic scene reconstruction method and apparatus, and a storage medium. The three-dimensional dynamic scene reconstruction method comprises: acquiring synchronized videos of a plurality of viewpoints of a dynamic scene; computing matching points between video images of different viewpoints, and estimating intrinsic and extrinsic parameters of each camera; obtaining a Gaussian splatting point set {p0} on the basis of a sparse point cloud constructed according to the depth of each matching point; for the first image frame of each video, using {p0} to perform static training thereon, to obtain a Gaussian splatting point set {p}; for the remaining image frames, dividing {p} into a static point set {S} and a dynamic point set {D}, performing dynamic training on {D}, and constructing a dynamic Gaussian splatting point set {P} from {p}, {S}, and the final {D}; and, in view of the intrinsic and extrinsic parameters of each camera, rendering {P} using a Gaussian splatting rendering pipeline, to obtain rendered images at different moments from new viewpoints.
Owner:TSINGHUA UNIVERSITY

Insurance claim settlement-oriented multi-modal image video evidence analysis method and system

The invention discloses an insurance claim settlement-oriented multi-modal image video evidence analysis method and system. The method comprises the following steps of: acquiring video / image and multi-source data such as metadata, audio, IMU (Inertial Measurement Unit), GPS (Global Positioning System), OBD (On-Board Diagnostic) and the like; calculating content Hash of the video and the audio according to frames, connecting the content Hash with time information in series to form chained Hash, and adding a verification digital signature and a credible timestamp; realizing cross-modal time sequence alignment based on self-adaptive time anchor-attitude coupling; tampering detection is carried out in combination with PRNU fingerprints, noise field consistency, dual compression, copy-movement and the like; multi-view geometry and monocular depth are fused, IMU scale constraint and micro rendering are introduced, three-dimensional reconstruction and re-projection optimization are completed, and collision dynamics verification is carried out; and constructing an event cause and effect graph, judging responsibility in combination with traffic rules, outputting a confidence coefficient vector and a structured report, and generating a verifiable evidence packet. The scheme has the advantages of high efficiency and traceability in the aspects of space-time restoration and interpretable responsibility judgment.
Owner:国任财产保险股份有限公司

Long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation

The invention discloses a long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation. The method comprises the following steps: 1) a multi-modal feature extraction module; 2) a multi-modal synchronization and alignment mechanism; 3) constructing a structured memory pool; 4) querying a drive generation mechanism; 5) incremental updating and memory compression strategy; and 6) unifying the multi-modal representation space. The invention provides a long video multi-mode understanding method fusing a large language model and retrieval enhancement generation, and aims to break through the limitation of a traditional method in the aspects of single-mode processing and semantic fragmentation. According to the method, video image features are extracted through a visual model (such as YOLO and ViT), voice transcription and environment voice description are obtained in combination with an audio model (such as Whisper and Qwen-Audio), and unified coding of vision, voice and audio in a long video is achieved. Then, a structured memory pool is constructed through semantic consistency segmentation and timestamp alignment technologies to store time slice data of different modalities.
Owner:GUANGZHOU BINGO SOFTWARE +1

High-precision instrument assembly fault backtracking method and system

The invention discloses a high-precision instrument assembly fault backtracking method and system, belongs to the field of precision manufacturing, and aims to solve the problems that in a traditional backtracking method, assembly data are scattered and unreliable, fault root positioning is fuzzy, and new scene adaptation depends on a large amount of data. The method comprises the following steps: collecting assembly structured data, video images and environment data in a multi-source manner, filtering out low-quality images, and distributing unique identifiers for products; fusing the multi-modal data to generate a depth feature matrix; constructing an anomaly detection model to output a risk score and a label; hashing the data and then storing the data into a product exclusive private block chain; when a fault occurs, extracting data on the chain through a unique identifier, reconstructing an assembly process by using a graph neural network, and comparing a standard positioning root; and based on the fault report incremental training model, parameters are optimized in combination with meta-reinforcement learning. According to the method, the data authenticity is guaranteed, the fault backtracking precision and efficiency are improved, a new scene is quickly adapted, the production rework rate is reduced, and the stable assembly quality is maintained.
Owner:XIAMEN ZONGNENG INSTR CO LTD

Substation three-dimensional fusion patrol method and system based on digital twinborn and autonomous identification

The invention relates to the technical field of transformer substation intelligent patrol, and provides a transformer substation three-dimensional fusion patrol method and system based on digital twinborn and autonomous identification. According to the method, a fused three-dimensional model is constructed through multi-source data acquisition and a three-dimensional Gaussian splash algorithm, and in combination with deep learning-based point cloud semantic segmentation and clustering, an equipment-patrol means coverage relationship is generated. Creating a virtual inspection proxy object based on a three-dimensional virtual environment, and controlling terminals such as an unmanned aerial vehicle to collect real-time video image data; the system carries out automatic identification on pictures, automatically completes equipment level alignment and standard point location identification, generates fine control of camera zooming, horizontal rotation, pitching and the like, and realizes standardized view finding and acquisition. By combining an enhanced recognition algorithm, traditional image processing and a deep learning model are fused, model self-evolution is realized through incremental learning, flexible expansion and collaboration of various patrol terminals are supported through a unified interface, and refined, real-time and intelligent patrol operation and maintenance requirements of an intelligent substation are met.
Owner:四川电力设计咨询有限责任公司

High-altitude operation risk early warning method and system based on camera image recognition

The invention provides a high-altitude operation risk early warning method and system based on camera image recognition, and relates to the technical field of computer vision, and the method comprises the steps: firstly collecting a video image sequence of a high-altitude operation scene, and generating a fusion feature map containing environment and operation main body features through multi-level feature extraction; performing spatial dimension segmentation and regional feature comparative analysis on the fusion feature map to obtain a spatial risk distribution map containing risk region identification information, processing the spatial risk distribution map of continuous frames based on a time sequence feature fusion rule to generate a dynamic risk evolution map, and calling a risk decision model to perform mode recognition to obtain a dynamic risk evolution map; and generating a risk level classification result and a risk position coordinate set according to the risk level classification result and the risk position coordinate set, and finally generating a risk early warning signal and sending the risk early warning signal to the monitoring terminal, thereby comprehensively, accurately and dynamically monitoring the high-altitude operation risk, and improving the accuracy and timeliness of risk early warning.
Owner:STATE GRID SHANXI POWER TRANSMISSION & DISTRIBUTION PROJECT CO

Photoelectric pod target identification and tracking system based on multi-scale attention mechanism

The invention provides a photoelectric pod target identification and tracking system based on a multi-scale attention mechanism, and belongs to the technical field of intelligent vision. Through combination of a multi-scale convolution module and a multi-head self-attention mechanism, accurate detection and tracking of a target in a photoelectric pod video image are realized. The multi-scale convolution module adopts convolution kernels of different sizes, and can extract local features of different scales to adapt to the change of the size of a target; the multi-head self-attention mechanism is used for capturing global features, especially long-distance dependency relationships between targets and backgrounds and between targets. Through fusion of local features and global features, the system improves the precision and robustness of target recognition in a complex scene. Meanwhile, by optimizing the structural design and the feature aggregation method, the calculation complexity of the system is remarkably reduced, and the requirements of the photoelectric pod for real-time performance and high efficiency are met.
Owner:GUANGDONG UNIV OF TECH

Real-time video image compression method based on deep learning

The invention provides a real-time video image compression method based on deep learning, and relates to the technical field of video image compression, and the method comprises the steps: carrying out the key feature recognition through employing an attention mechanism; performing convolution training optimization on the video image sample data set by using a deep learning network structure; a self-encoder structure is designed to carry out feature map encoding compression; a video image compression adaptive network is generated through series fusion; a real-time video image frame is collected for preprocessing, and feature compression processing is performed on a standard video image frame based on a video image compression adaptive network. According to the method and the device, the technical problem that the video compression quality is reduced due to the fact that the generalization ability is insufficient in the face of various scenes and the video compression strategy is difficult to adaptively adjust according to different scenes in the prior art can be solved, the adaptive network is constructed through the combination of deep learning and the auto-encoder, and the video compression quality is improved. And the video compression strategy is dynamically adjusted according to the contents of different video images, so that the video compression quality is improved.
Owner:NANJING STAR SHIELD INFORMATION TECH CO LTD

Three-dimensional video fusion method based on camera self-calibration and projection texture mapping

The invention relates to the technical field of computer vision and virtual reality, and discloses a three-dimensional video fusion method based on camera self-calibration and projection texture mapping, and the method comprises the following steps: S1, obtaining video data, and collecting at least one frame of two-dimensional image in a to-be-fused video stream; optionally, the two-dimensional image is preprocessed; and S2, calibrating internal reference of the camera, detecting linear features in the two-dimensional image by using an image processing algorithm, estimating the position of a vanishing point by using a least square method or other optimization algorithms based on the detected linear segment, and calculating an internal reference matrix of the virtual camera according to the optimized vanishing point position. The perspective relation between the video image and the surface of the three-dimensional model is determined through the vanishing point detection technology, accurate fusion of the video image and the three-dimensional model is achieved, the sense of reality of a virtual scene is improved, and the tedious camera calibration process and the complex three-dimensional reconstruction process based on a calibration plate are avoided.
Owner:ANHUI CIVIO INFORMATION & TECH

Real-time badminton action detection system and device based on MediaPipe and Motion Bidirectional Encoder Representation Transformer

A real-time badminton action detection system, consisting of: a video recording module configured to continuously record video images at a frame rate of at least thirty frames per second; a pose estimation processing unit configured to detect and output two-dimensional skeletal landmark coordinates for a variety of body joints, including at least wrists, elbows, shoulders, hips, knees and ankles, from each video frame; a Motion Bidirectional Encoder Representation Transformer (Motion-BERT) configured to receive sequential skeleton landmark coordinates over a defined time window and encode motion trajectories using multi-head self-attention mechanisms across past and future frames; and a classification controller module operationally coupled to the Motion-BERT, wherein the classification controller module comprises a dense neural network with a softmax output layer configured to generate real-time probability distributions over a variety of badminton-specific action classes, the end-to-end system being configured to produce recognition results with a processing latency of less than 100 milliseconds; and wherein the video recording module comprises a high-speed digital camera with a wide-angle lens positioned at a point on the perimeter of the court, the camera being calibrated with intrinsic and extrinsic parameters for perspective correction, and wherein the system includes a calibration routine that aligns detected skeletal landmarks with a reference badminton court coordinate system.
Owner:NITTE MEENAKSHI INSTITUTE OF TECHNOLOGY (DEEMED TO BE UNIVERSITY) BENGALURU +3

Tunnel face construction area safety monitoring system and method

The invention discloses a tunnel face construction area safety monitoring system and method, and belongs to the technical field of construction safety monitoring. The method comprises the following steps: collecting multiple types of monitoring data of a tunnel face construction area in real time, and preprocessing the monitoring data; extracting abnormal event features in the tunnel face video / image based on a target detection model; the integrity of the tunnel face is judged based on a tunnel face integrity judgment module; fusing the collected multi-source data, and generating a comprehensive safety coefficient S in real time; and triggering different levels of alarm response measures according to the interval in which the safety factor S is located. According to the method, the risk early warning precision of the tunnel face construction area can be effectively improved by integrating multi-modal data fusion, the self-adaptive warning rule and the risk quantitative evaluation index.
Owner:CHINA MCC17 GRP CO LTD

Urban rail transit passenger monitoring data processing and analyzing system and method

The invention relates to the technical field of urban rail transit, in particular to an urban rail transit passenger monitoring data processing and analyzing system and method, and the system comprises a data collection layer which is used for collecting video streams, gate passing records, passenger positioning data and environment parameters of temperature, humidity and illumination in real time; the spatio-temporal feature fusion layer is used for converting the video image data into analyzable feature vectors and integrating card swiping records and position information to form spatio-temporal trajectory data of passengers; the three-level anomaly detection layer comprises an individual layer detection unit, a group layer detection unit and a system layer prediction unit; the dynamic decision-making layer is used for calculating an abnormal score based on a dynamic threshold value, and dynamically adjusting a behavior coefficient according to a historical disposal effect through a PPO reinforcement learning algorithm; and the execution layer comprises an edge computing node and a cloud analysis platform. Therefore, the problems of single data acquisition, lack of comprehensive anomaly detection means, fixed and lagged decision response, ineffective utilization of resources and the like in the prior art are solved.
Owner:BEIJING MAGLEV DATA TECHNOLOGY CO LTD

Vehicular driver monitoring system with driver monitoring camera and near IR light emitter at interior rearview mirror assembly

A vehicular driver monitoring system includes a vehicular interior rearview mirror assembly having a mirror head that accommodates a mirror reflective element. A video display is disposed behind the mirror reflective element and operable to display video images captured by a rearward viewing camera of the vehicle. A driver monitoring camera and a near infrared light emitter are accommodated by and move in tandem with the mirror head. The near infrared light emitter is accommodated within the mirror head so that, with the mirror head adjusted relative to the mounting base to set the rearward view of the driver of the vehicle, a beam of near infrared light emitted by the near infrared light emitter is directed toward a driver's region of the vehicle. The driver monitoring camera is disposed adjacent to the video display screen so as to not view through the video display screen.
Owner:MAGNA MIRRORS OF AMERICA INC

Multi-mode unmanned material checking method based on combination of RFID and machine vision

The embodiment of the invention provides a multi-mode unmanned material checking method based on combination of RFID and machine vision, the method is applied to a mobile module and a camera module, an RFID read-write module is arranged on a mobile trolley, and the method comprises the following steps: in response to movement of the mobile module, the RFID read-write module scans a target material to obtain RFID identification data, and filters and de-duplicates the RFID identification data to obtain a target material; an RFID identification result is determined; dividing a shooting range of the camera module based on a scanning area of the RFID read-write module, and detecting a material box and a material surface sheet label in video image data when the video image data is shot; and identifying the material surface sheet label to obtain material data information, and matching the material data information with the RFID identification result to obtain a multi-modal material matching result.
Owner:HANGZHOU EBOYLAMP ELECTRONICS CO LTD

Video anti-shake method and system based on multi-scale fusion and adaptive smoothing

The invention discloses a video anti-shake method and system based on multi-scale fusion and adaptive smoothing, relates to the technical field of video image processing, and aims to effectively solve the image quality problem caused by shake in a video shooting process. Gradient histograms and wavelet energy distribution characteristics of video frames are extracted through graying and normalization processing, and the gradient histograms and the wavelet energy distribution characteristics are input into a jitter type recognition network to recognize translation, rotation and Z-axis jitter probabilities. And further extracting motion, frequency domain and edge features, and generating multi-modal coupling features through combination of a dynamic feature interaction network and a dot product attention mechanism. And constructing a motion trajectory by using the features, optimizing the trajectory by using a texture perception double-layer smoothing strategy, introducing an adaptive penalty term into a dynamic planning cost function, and outputting a smooth motion compensation parameter. And finally, processing the boundary region through motion compensation and image extrapolation to generate an anti-shake video frame. Through multi-scale feature fusion and a self-adaptive smoothing strategy, the video anti-shake effect is effectively improved, and the method is suitable for complex scenes.
Owner:江淮前沿技术协同创新中心

Intensive care unit video image processing method based on image semantic segmentation

The invention relates to the technical field of image segmentation, in particular to an intensive care unit video image processing method based on image semantic segmentation, which comprises the following steps of: acquiring a video image through monitoring camera shooting, dividing a semantic region, calculating a pixel displacement direction of each part, identifying an abnormal track, erasing interference, recombining a boundary contour, and comparing shape change. And constructing a trend and re-drawing a structure, evaluating stability in combination with directions, speeds and amplitudes, screening and correcting inconsistent labels, and outputting an index result. According to the method, track features are constructed by introducing pixel direction coding, a deviation region is identified and interference is eliminated in combination with a direction change trend, an edge structure is reconstructed by using a boundary connection sequence and curve fitting, contour division is optimized through deformation direction consistency, and labels are corrected by synthesizing direction change and boundary speed. Dynamic tracking and accurate labeling of the patient state are achieved, the abnormal action recognition efficiency and the label updating accuracy are improved, and the time sequence integrity and the space expression ability of image data are enhanced.
Owner:THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV

Vehicular imaging system with extendable camera

A vehicular camera monitoring system includes an electronic control unit (ECU) at a vehicle and a support arm movably disposed at a side portion of the vehicle, with the support arm having a base end attached at the side portion of the vehicle and a distal end opposite the base end. A camera is disposed at the distal end of the support arm. The support arm is movable between a stowed position and an extended position. A cover element covers an aperture at the side portion at least when the support arm is in the extended position. The camera, when the support arm is in the extended position, captures image data and provides captured image data to the ECU, which processes the provided image data for (i) display of video images derived from provided image data and / or (ii) detection of an object in the field of view of the camera.
Owner:MAGNA MIRRORS OF AMERICA INC

Intelligent alarm positioning method and system based on unmanned aerial vehicle

The invention discloses an intelligent alarm positioning method and system based on an unmanned aerial vehicle, and belongs to the technical field of unmanned aerial vehicle monitoring and geographic space information processing, and the method comprises the steps: obtaining a video stream in real time based on the unmanned aerial vehicle, and recognizing a risk point location in a video image; a three-dimensional space positioning model is constructed based on telemetry data and lens angle parameters acquired by the unmanned aerial vehicle in real time. And performing three-dimensional coordinate dynamic solution on the risk point location based on a three-dimensional space positioning model to obtain a world coordinate of the risk point location. And determining and outputting an alarm position of the risk point location based on the world coordinates. An automatic detection and artificial interaction dual-channel mechanism is adopted, an alarm target is accurately identified and positioned in an unmanned aerial vehicle real-time video, a corresponding relation between video pixels and geographic coordinates is established through real-time registration of an unmanned aerial vehicle image and the digital earth, a space coordinate conversion error caused by view angle difference is corrected by using a depth map, and an alarm target is accurately identified and positioned. Accurate conversion from video pixel coordinates to world coordinates is realized, and high-precision positioning support is provided for remote monitoring of the unmanned aerial vehicle.
Owner:CHINA TOWER CO LTD

Intelligent retrieval method and system fusing text and image semantic features

The invention discloses an intelligent retrieval method and system fusing text and image semantic features. The method comprises the following steps: S1, carrying out quality detection and preprocessing on an image scanning copy of an electronic file and a case simultaneous recording video; s2, constructing a structured electronic file directory; s3, extracting a text semantic feature, a file image semantic feature and a video image semantic feature as multi-modal features of the text and the image; s4, carrying out feature fusion and alignment on the multi-modal features of the text and the image through a multi-modal large model, and generating a cross-modal unified feature vector with semantic consistency; s5, automatically constructing a case knowledge graph, and realizing structured and semantic integration of legal information; and S6, performing semantic analysis and multi-hop reasoning based on natural language query and the case knowledge graph, and generating and presenting a retrieval result in a structured or question and answer form. According to the method, the semantic features of the text and the image are fused, so that deep knowledge mining and efficient intelligent retrieval of the electronic file are realized.
Owner:TONGFANG SAIWEIXUN INFORMATION TECH CO LTD +1

Fire detection method and device based on multi-modal perception and D-S evidence theory fusion

The invention discloses a fire detection method and device based on multi-modal perception and D-S evidence theory fusion. The method comprises the following steps: collecting multi-modal perception big data at least comprising video image data and temperature sensing data in a monitoring area; performing fire visual feature analysis on the video image data to generate first basic probability distribution; performing fire temperature characteristic analysis on the temperature sensing data to generate second basic probability distribution; taking the first basic probability distribution and the second basic probability distribution as two independent evidence sources, and fusing by adopting a D-S evidence theory to obtain a fused third basic probability distribution; converting the third basic probability distribution into a fire occurrence probability for decision making; and when the fire occurrence probability exceeds a preset alarm threshold, determining that a fire occurs and triggering an alarm. According to the fire detection method, multi-modal sensing big data are synchronously collected and fused, multi-source uncertain information is processed and decided under the framework of the D-S evidence theory, and the early stage of fire detection is remarkably improved.
Owner:CHINA IPPR INT ENG CO LTD

Underground pipeline intelligent detection and mapping method based on image recognition

The invention discloses an underground pipeline intelligent detection and mapping method based on image recognition, and the method comprises the following steps: 1, obtaining continuous video images, associating feature matching pairs of adjacent key frames, and forming a pose parameter set; 2, inputting the key frame into an improved YOLO-World detection network, and outputting a detection result set; 3, obtaining a geometric consistency matching set according to the detection result set; 4, performing multi-view triangularization on the geometric consistent matching set to form a weight factor; 5, introducing a weight factor, and executing incremental beam adjustment optimization on the pose parameter set and the three-dimensional sparse point set to obtain a sparse semantic point cloud; and step 6, outputting an underground pipe network topology map. According to the invention, high-precision and high-robustness intelligent identification and topological mapping in a complex underground pipeline environment are realized.
Owner:WUXI YIXING POWER TECH CO LTD

Video image segmentation method

The invention relates to the field of image processing, and discloses a video image segmentation method which is used for improving the precision, robustness and real-time performance of video image segmentation in a complex dynamic scene. The video image segmentation method comprises the following steps: generating an entropy generation rate map through weighted combination of light flow divergence and rotation, and quantifying motion irreversibility; a double-virtual-form prime concentration field is constructed, next-frame texture prediction is realized through iterative evolution, and the dynamic background adaptability is enhanced; fusing the projection entropy generation rate graph and the enhanced texture residual graph, dynamically distributing motion and appearance weights, and generating a high-precision boundary response graph; and extracting and persistent filtering are carried out, topological consistency maintenance of segmentation masks is realized, multi-scale collaborative segmentation and dynamic computing resource scheduling are supported, and efficiency and precision are balanced. The segmentation precision of the method is obviously superior to that of a traditional method in complex scenes such as illumination variation and rapid motion, and the method is suitable for the fields with high real-time requirements such as monitoring, medical treatment and automatic driving.
Owner:JIANGSU HUIHANG DIGITAL TECHNOLOGY CO LTD

Code rate control method and system based on video image segmentation

The invention relates to the technical field of video coding and image processing, and discloses a code rate control method and system based on video image segmentation, and the method comprises the steps: carrying out the pixel-level semantic segmentation of a to-be-coded video frame sequence; extracting a foreground region of interest; calculating a corresponding segmentation uncertainty parameter; establishing semantic mutation parameters of the foreground region of interest; generating a semantic perception weight; executing region-level target code rate redistribution and quantization parameter mapping; and executing partition coding control. In the prior art, code rate control mainly depends on motion intensity or pixel complexity, and especially when a foreground target suddenly appears or disappears in a monitoring scene, a technical problem that a key target is blurred or a background code rate is wasted is easily caused. Due to the fact that the uncertainty modeling and semantic mutation sensing mechanism of the semantic segmentation result is introduced, priority coding of the foreground interest area is achieved under the frame-level code rate constraint condition, and the video coding quality and the code rate utilization efficiency are improved.
Owner:KAIXIN CHUANGDA (SHENZHEN) TECH DEV CO LTD

Auxiliary driving device and system based on video image processing

The invention relates to the technical field of data processing, in particular to an auxiliary driving device and system based on video image processing, and the device comprises a frame buffer queue with a multi-channel video collection module aligned with timestamps; a hierarchical calculation task decoupling module identifies task types in the frame buffer queue and marks priority labels, including a first priority label of a radar data analysis and distance superposition task, a second priority label of a vehicle profile generation task and a third priority label of a video fusion task, and outputs a data stream with priority labels; the calculation module generates a task processing result; the genetic algorithm optimization scheduling module generates a resource allocation strategy and controls the calculation module to execute task processing; and the output synchronous control module synthesizes a high-priority task result and performs frame rate synchronous output on a low-priority task. According to the method, the problem of real-time degradation caused by multi-path high-definition video stream processing is solved through task grading and dynamic resource scheduling, and the image output delay of the auxiliary driving scene is reduced.
Owner:BEIJING XINGJIAN CHANGKONG MEASUREMENT CONTROL TECH

Video transmission image stitching data enhancement method and system based on deep learning

The invention discloses a video transmission image stitching data enhancement method based on deep learning, and relates to the field of video image processing. The method comprises the following steps: S1, acquiring and screening images; s2, continuously screening structures and selecting key frames; s3, correcting image distortion; s4, splicing and fusing the images; s5, performing image enhancement output; firstly, a sliding time window mechanism is adopted, multi-dimensional image quality screening is combined, fuzzy, underexposure or severely-shielded inferior frames are accurately removed, and high quality of input key frames is ensured; through intelligent splicing and enhancement of a dynamic adaptive threshold strategy and semantic guidance, the image splicing precision and efficiency of the unmanned aerial vehicle and the multi-view camera in a complex environment are remarkably improved; besides, semantic segmentation guided feature extraction is combined with a multi-band fusion technology, seamless splicing is realized, the quality of an output image is improved, and the reliability of automatic analysis and decision making is remarkably improved.
Owner:GUANGZHOU WEITUXIN ELECTRONIC TECH CO LTD

HDR video reconstruction method based on standardized stream

The invention discloses an HDR video reconstruction method based on a standardized stream, and belongs to the technical field of high dynamic range image processing. The method comprises the following steps of: firstly, constructing a convolution optical flow estimation module with a self-adaptive normalized structure, wherein the convolution optical flow estimation module is used for accurately acquiring optical flow information between adjacent frames in an alternative exposure LDR video image sequence; then, carrying out multi-level feature alignment on the image sequence through an image alignment module so as to reduce alignment errors caused by illumination difference and movement; and finally, inputting the aligned and fused multi-level LDR image features into a standardized flow reconstruction network to realize high-quality HDR video image reconstruction. Aiming at the video reconstruction problem under the alternate exposure condition, the invention designs a standardized flow modeling structure considering the optical flow estimation precision and the feature alignment effect, and effectively improves the HDR video reconstruction quality in a complex dynamic scene.
Owner:BEIHANG UNIV

LED spherical screen video image coordinate generation method and system, medium, program product and terminal

The invention provides an LED spherical screen video image coordinate generation method and system, a medium, a program product and a terminal, and the method comprises the steps: obtaining a configuration file of an LED spherical screen, and carrying out the analysis, so as to obtain the spatial geometric parameters of the LED spherical screen; obtaining a to-be-displayed video image of the LED spherical screen, and analyzing and screening the to-be-displayed video image to obtain an effective pixel data set; and calculating latitude and longitude coordinates corresponding to the pixel coordinates of the to-be-displayed video image by adopting a nonlinear mapping method based on the spatial geometric parameters of the LED spherical screen and the effective pixel data set of the to-be-displayed video image. By preloading spatial geometric parameters of a sphere, image pixels are dynamically analyzed, and efficient and accurate conversion from pixel coordinates to latitude and longitude coordinates is realized by adopting a nonlinear mapping algorithm. Meanwhile, the batch writing optimization and dynamic buffer management technology is adopted, the processing speed and the system operation stability are remarkably improved, and the method can be widely applied to the fields of LED spherical screen content manufacturing, display correction and the like.
Owner:SHANGHAI SANSI ELECTRONICS ENG +4

Water supply and drainage pipeline anomaly detection method based on image recognition

The invention discloses a water supply and drainage pipeline anomaly detection method based on image recognition. Initial video image data are acquired through an image sensor; establishing a three-channel feature extraction network model, wherein the three-channel feature extraction network model at least comprises a visible light channel, a geometrical shape channel for extracting pipeline structure deformation features through 3D convolution and a dynamic optical flow channel for capturing liquid flow abnormal features based on an RAFT algorithm; inputting an initial video image into the three-channel feature extraction network model to perform feature extraction, establishing an analysis model based on a bidirectional LSTM neural network, inputting feature image data into the analysis model to establish time sequence association, performing abnormal evolution through a GNN graph neural network, performing abnormal type classification according to a pipeline state evaluation index, and obtaining a pipeline state evaluation result. And performing early warning based on a classification result. The detection efficiency is improved, the labor cost and the safety risk are reduced, and operation and maintenance personnel can master the state of the pipeline in time.
Owner:ZIBO KAIHUI WATER SUPPLY EQUIP CO LTD

Geographic mosaicking method and apparatus for video images, and computer device and storage medium

The present application relates to a geographic mosaicking method and apparatus for video images, and a computer device and a storage medium. The method comprises: acquiring real-time live streaming images of a real-world scene that are collected by a plurality of unmanned aerial vehicles, and GNSS information of the plurality of unmanned aerial vehicles; using a visual SLAM algorithm to perform input frame tracking on the real-time live streaming images, and performing pose estimation on successfully tracked input frames by combining visual trajectories and the GNSS information, so as to generate georeferenced camera poses; using the camera poses and a surface reconstruction algorithm to perform densification processing on the input frames, so as to generate a depth map of the real-world scene, and mapping the depth map into dense three-dimensional point clouds, so as to construct a three-dimensional surface model of the real-world scene; and using the camera pose and the three-dimensional surface model to perform orthorectification on the input frames, and integrating the orthorectified input frames into a global mosaicked map, so as to generate a global image of the real-world scene. By means of the embodiments of the present application, a real-time picture of a real-world scene can be quickly acquired, thereby providing robust support for a rapid emergency response in a real-world scene.
Owner:SHENZHEN INST OF ADVANCED TECH

Pulmonary nodule display method and device, electronic equipment and storage medium

The invention provides a pulmonary nodule display method and device, electronic equipment and a storage medium, and the method comprises the steps: segmenting a three-dimensional reconstruction image of the chest of a target patient to obtain a preoperative segmentation image, the three-dimensional reconstruction image being obtained based on preoperative CT data, and the preoperative segmentation image comprising the position information of a pulmonary nodule; segmenting the video image of the intraoperative lung tissue of the target patient collected by the thoracoscope in real time to obtain an intraoperative segmented image; feature matching is conducted on the preoperative segmented image and the intra-operative segmented image, the lens pose of the thoracoscope is determined, the preoperative segmented image is mapped into a two-dimensional reference image based on the lens pose, and the two-dimensional reference image comprises the position information of the pulmonary nodule; and carrying out image registration on the two-dimensional reference image and the intra-operative segmented image, and carrying out pulmonary nodule marking on the intra-operative segmented image to obtain a video image for displaying pulmonary nodules in real time. The position of the pulmonary nodule in the operation is accurately displayed in real time in a non-invasive mode, and the operation efficiency is improved.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL