Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

704 results about "Multimodal image" patented technology

Real-time virtual reality scene system based on natural language description using multimodal artificial intelligence

A real-time system for the multimodal generation of virtual reality scenes based on artificial intelligence for the creation of immersive three-dimensional environments from natural language narratives, consisting of: a speech capture module configured to continuously record a user's spoken narrative via one or more directional microphones, preprocesses the captured signal by noise reduction and temporal alignment, and outputs a digital speech stream; A speech-to-text processing unit that is operationally coupled to the speech capture module and configured for real-time speech recognition using a continuous neural transformer model. The unit is trained to transcribe natural language utterances into structured text data while maintaining contextual continuity throughout the evolving narrative. a semantic interpretation processing unit that is communicatively linked to the speech recognition unit and configured to perform natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large language model that is fine-tuned for spatial reasoning tasks; a scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with metadata describing geometry, position, orientation, texture, and linking attributes between objects; a multimodal image-language model processor coupled with the scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the scene graph; a scene assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time visualization; The system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Road crack detection method and system based on fused image

The invention relates to the technical field of road crack detection, in particular to a road crack detection method and system based on a fused image. The method comprises the following steps: acquiring road multi-source monitoring data including a visible light image, infrared thermal imaging data and laser radar point cloud data, and performing multi-modal image fusion and road three-dimensional point cloud reconstruction to generate a fused road image and road three-dimensional modeling data; performing crack curvature analysis based on the fused road image to generate crack curvature data; performing reflection crack contour recognition and positioning on the fused road image through the crack curvature data to generate reflection crack initial positioning data; obtaining road base material data; and performing reflection crack stress field reconstruction on the road area according to the reflection crack initial positioning data to obtain a reflection crack stress field. According to the invention, through multi-modal fusion, curvature identification, stress field modeling and crack channel analysis, the accuracy and strain of road reflection crack detection are improved.
Owner:BINHAI BAY BRANCH OF DONGGUAN CITY URBAN MANAGEMENT & COMPREHENSIVE LAW ENFORCEMENT BUREAU

Intelligent monitoring system and method based on multi-modal remote sensing data and deep learning

The invention relates to the technical field of unmanned aerial vehicle remote sensing and artificial intelligence crossing, in particular to an intelligent monitoring system and method based on multi-modal remote sensing data and deep learning, and the system comprises an unmanned aerial vehicle cluster networking subsystem, a mixed feature matching subsystem and a multi-modal fusion and continuous learning subsystem. The method comprises the following steps: constructing an unmanned aerial vehicle cluster carrying a multispectral sensor and a laser radar LiDAR, and carrying out wireless networking among a plurality of unmanned aerial vehicles to realize sharing of acquired images; feature point extraction is carried out on collected images of different time phases, the extracted feature points are input into the generative adversarial network, and the feature points are matched; and receiving the matched collected images, dynamically fusing data of visible light, infrared and other multi-modal images through a space-time attention mechanism, and realizing high-precision target recognition and dynamic environment self-adaption in a small sample scene. According to the method, unmanned aerial vehicle multi-source remote sensing data acquisition, feature fusion and deep reinforcement learning are combined, and the method is used for intelligently monitoring a dynamic environment.
Owner:XINJIANG NORMAL UNIVERSITY

Underwater fish school monitoring statistical system based on image fusion

The invention relates to the technical field of underwater fish school monitoring, and discloses an underwater fish school monitoring statistical system based on image fusion. According to the system, underwater video streams and sonar reflection intensity data of different spectral bands are acquired through an underwater multi-source image acquisition module, and time-space synchronous multi-modal image data streams are generated through timestamp alignment; a fish school contour reconstruction module is used for segmenting a fish school contour boundary and fusing visible light texture and sonar geometric features to generate an underwater three-dimensional fish school distribution set; the dynamic track mapping module tracks the mass center displacement, calculates the movement rate and the direction deviation angle, and correlates the water area depth to generate a dynamic track topological graph; the behavior anomaly analysis module extracts environment data based on the track mutation node, and detects aggregation density change and direction dispersion to mark an anomaly feature cluster; and the population statistics output module integrates the data, performs classified statistics on population distribution, a quantity threshold value and a migration path overlap ratio, and finally generates a fish school quantity distribution statistics thermodynamic map.
Owner:福州海洋研究院

Image-fused end-side cloud collaborative intelligent fire-fighting fire monitoring system

The invention discloses an end-side cloud collaborative intelligent fire-fighting fire monitoring system based on image fusion, and relates to the technical field of intelligent fire-fighting, the system is composed of a plurality of functional modules, and the system comprises a multi-modal image fusion module which generates a dynamic scanning priority map based on prior data, distinguishes a natural heat source from an abnormal fire by using a dual-light fusion algorithm, and sends an image fusion result to a cloud server; a scanning area is divided according to the thermal risk grade, and the thermal imaging resolution is dynamically adjusted; the distributed edge computing module is used for carrying out space-time synchronization on cross-modal data through a multi-modal feature alignment network, and carrying out dynamic allocation on a CUDA core and CPU resources through adaptive computing scheduling; an improved artificial bee colony algorithm is adopted, the bandwidth of the multi-sensor data flow is dynamically allocated through a time-sharing multiplexing protocol, and three-dimensional path planning is carried out; and the end-side cloud collaborative decision module constructs a federated learning driven model sharing network, and each edge node trains a lightweight YOLOv5s pruning model based on local data.
Owner:HANGZHOU ZIPENG TECH CO LTD

Intelligent identification method and system for asymmetric plate shape defects

The invention provides an intelligent identification method and system for an asymmetric plate shape defect, and the method comprises the steps: collecting the multi-modal image data of a to-be-detected plate shape surface, and carrying out the preprocessing of the multi-modal image data, and obtaining a standardized image; performing feature extraction on the standardized image based on an asymmetric feature enhancement algorithm to obtain an asymmetric feature vector; inputting the asymmetric feature vector into a pre-trained asymmetric defect identification model to generate a preliminary defect classification result and a defect area thermodynamic diagram; according to the thermodynamic diagram of the defect region, segmenting a defect boundary in combination with a geometric constraint optimization algorithm, and determining morphological parameters and spatial positions of asymmetric defects; and based on the morphological parameters and the spatial positions, correcting the preliminary defect classification result through a dynamic threshold adjustment algorithm to obtain an identification result, thereby alleviating the technical problem of low accuracy of asymmetric plate shape defect identification in the prior art.
Owner:GUANXIAN ZHONGGUAN NEW MATERIALS CO LTD

Visual positioning method and system for precise connector assembly

The invention relates to the technical field of visual positioning, and provides a visual positioning method and system for precise connector assembly. Performing multi-mode HDR fusion and distortion correction on the original connector image to obtain a multi-layer fusion image matrix; and performing feature region coarse positioning in combination with a composite convolutional neural network to obtain a connector coordinate set, performing geometric feature fine positioning according to the connector coordinate set and the multilayer fusion image matrix to obtain a key point coordinate set, and performing hand-eye calibration and dynamic compensation on the key point coordinate set to generate a pose instruction. And performing vision-force control hybrid assembly processing on the pose instruction according to the sensor feedback data to obtain assembly completion state data. Through multi-modal image fusion, deep learning coarse positioning, precise geometric registration and dynamic cooperation of vision and force control, the speed, precision and stability of precise connector assembly are improved.
Owner:DONGGUAN HAIHONG INTELLIGENT TECH CO LTD

Ultrasonic vein puncture system integrating image recognition and data analysis

The invention relates to the technical field of medical intelligent image recognition data processing, and discloses an image recognition and data analysis fused ultrasonic venipuncture system which comprises a multi-modal image acquisition module, an intelligent analysis processing module, a real-time navigation execution module and a complication early warning module. By arranging a multi-mode fusion sensing end, when vein puncture real-time navigation is carried out, through ultrasonic image, thermodynamic distribution and optical characteristic three-mode data collaborative registration, the consistency of deep blood vessel recognition is guaranteed, meanwhile, a three-dimensional topological model containing blood vessel elastic parameters is dynamically constructed, blood vessel position deviation can be calibrated in real time in the puncture process, and the accuracy of vein puncture is improved. The accuracy of blood vessel positioning is guaranteed, puncture positioning errors of complex cases are further reduced, whether angle or depth deviation occurs in a puncture path or not is judged in real time by arranging a dynamic navigation end, a compensation path can be planned in real time through a blood vessel elastic characteristic matrix, and it is guaranteed that the angle errors are reduced when a needle body enters a blood vessel.
Owner:WEST CHINA HOSPITAL SICHUAN UNIV

Thyroid ultrasonic robot automatic scanning method, device and equipment based on RGB image and depth information and medium

The invention relates to the technical field of computer vision, and discloses a thyroid ultrasonic robot automatic scanning method, device and equipment based on RGB images and depth information and a medium, and the method comprises the steps: obtaining image information and depth information, coding the depth information, fusing and recognizing a target scanning area, determining an initial scanning point and an initial scanning direction, controlling the scanning probe to scan and collect a real-time scanning image, analyzing the real-time scanning image to recognize a preset target and an artifact area, adjusting a scanning posture and a scanning path based on a recognition result, monitoring a continuous existence state of the preset target, and stopping scanning when the preset target is not recognized continuously. The target area is identified by fusing the multi-modal image information, the scanning posture and path are dynamically adjusted in combination with real-time image analysis, scanning termination is intelligently controlled according to the target detection result, the positioning accuracy, image quality and standardization level of ultrasonic scanning are improved, and the method is suitable for automatic ultrasonic imaging of thyroid and superficial organs.
Owner:SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD

Three-dimensional space anaphora reasoning method and device, electronic equipment and storage medium

The invention provides a three-dimensional space anaphora reasoning method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: obtaining RGB-D image data of a target scene and a natural language instruction containing spatial constraints; wherein the RGB-D image data is multi-modal image data containing color visual information and depth information; inputting the RGB-D image data and the natural language instruction into a pre-trained visual language large model, and outputting a text containing an explicit reasoning process and target point coordinates conforming to spatial constraints; wherein the visual language large model is obtained through combined training of two-stage supervised learning fine tuning of depth alignment and space understanding enhancement and reinforcement learning fine tuning based on a display reasoning process; the visual language large model comprises an independent depth encoder, and the depth encoder is used for processing depth information. Through the method provided by the invention, the comprehensive performance in the complex space anaphora task is improved.
Owner:BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

Breeding early warning method and system based on artificial intelligence image recognition

The invention relates to the technical field of intelligent breeding monitoring, in particular to a breeding early warning method and system based on artificial intelligence image recognition. The method comprises the following steps of: 1, dynamically acquiring and adaptively preprocessing a multi-modal image; step 2, constructing a multi-branch convolutional neural network based on the preprocessed image data, respectively extracting space apparent characteristics and time behavior characteristics of a target object, and dynamically fusing characteristic weights by combining real-time environment parameters to generate an abnormal probability value; and step 3, performing multi-level dynamic early warning decision and feedback optimization, dynamically generating a multi-level early warning threshold according to environmental parameters and historical abnormal probability value distribution, triggering differential response operation, and updating parameters of the multi-branch convolutional neural network and the multi-level early warning threshold based on an artificial recheck result and equipment feedback data. The method can automatically adjust the early warning mechanism according to the change of the real-time environment, thereby effectively preventing the potential risk in the breeding process, improving the breeding management efficiency, and guaranteeing the sustainable development of the breeding industry.
Owner:HUNAN JINGTIAN AGRICULTURAL CO LTD

Interventional surgery navigation method based on multi-modal image fusion

The invention relates to the field of cardiology department operations through vascular intervention, and discloses an interventional operation navigation system based on multi-modal image fusion, which comprises a preoperative three-dimensional reconstruction and planning module, which is mainly used for generating a three-dimensional model of a heart through preoperative image data, and planning the position and section of an intraoperative ICE ultrasonic probe based on the model, the intraoperative multi-modal fusion navigation module is mainly used for fusing a three-dimensional model generated before an operation with an intraoperative real-time image and providing accurate navigation support for a doctor; and the dynamic feedback optimization module is mainly used for updating a dynamic track in real time and performing safety protection. According to the invention, through the preoperative three-dimensional reconstruction and virtual path planning module, in combination with the improved 3D U-Net network and the high-precision segmentation algorithm, a three-dimensional model containing key anatomical markers is generated, 20 groups of ICE probe alternative paths are planned, seamless connection of preoperative three-dimensional anatomical information and intraoperative real-time images is realized, doctors can quickly call pre-planned paths, and the accuracy of the ICE probe anatomy is improved. Adjusting time during operation is shortened, and operation efficiency is improved.
Owner:NINGBO FIRST HOSPITAL

Prediction method and system for breast cancer immunohistochemical index and typing

The invention discloses a breast cancer immunohistochemical index and typing prediction method and system. The prediction method comprises the following steps: acquiring a breast ultrasound image and a breast magnetic resonance image of a patient; performing preprocessing and quality control on the mammary gland ultrasonic image and the mammary gland magnetic resonance image; performing breast lesion area segmentation by adopting a deep learning model to obtain segmented lesion areas; based on the segmented lesion area, extracting multi-modal radiomics characteristics of the breast ultrasonic image and the breast magnetic resonance image; fusing the multi-modal radiomics characteristics of the breast ultrasound image and the breast magnetic resonance image, constructing a machine learning model, and performing immunohistochemical index prediction to obtain an immunohistochemical index prediction result; and performing breast cancer molecular typing analysis according to the immunohistochemical index prediction result. By fusing the multi-modal image information, the tumor features can be described more comprehensively, and the accuracy of biomarker prediction is improved.
Owner:PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)

High-voltage equipment defect positioning system and method based on multi-modal image fusion

The invention relates to the technical field of electrical equipment defect monitoring, in particular to a high-voltage equipment defect positioning system and method based on multi-modal image fusion, and the system comprises an ultraviolet triggering module, a modal synchronization module, a space positioning module, a structure recognition module and a fusion diagnosis module. According to the method, automatic calibration of a suspected defect area can be realized by performing quantitative threshold judgment on a corona discharge signal in an ultraviolet image, the consistency of three-mode data in space and time dimensions is ensured, and the accuracy of the defect area is improved by comparing three types of coordinates of a lead hot spot, boundary geometry and a spot center in the image and performing registration offset correction. The image feature alignment precision is improved, the insulator chain light spot distribution path is extracted, the continuous distribution length is calculated and corresponds to the pollution level standard, automatic identification of the pollution level is achieved, a multi-modal image fusion feature block is constructed, and a risk sorting label is given. And the defect positioning accuracy and the comprehensive evaluation capability on the defect property and the influence degree are effectively improved.
Owner:SHANGHAI ZIHONG OPTOELECTRONICS TECH CO LTD

Digital information processing method for hospital radiology department

The invention provides a digital information processing method for a hospital radiology department, which comprises the following steps: S1, multi-modal image collaborative acquisition and standardization: synchronously acquiring anatomical structure images, functional images and metabolic parameter data of a patient through radiology department imaging equipment, and converting the anatomical structure images, the functional images and the metabolic parameter data into space-time aligned three-dimensional digital matrixes; s2, image quality optimization processing: performing nonlinear contrast enhancement and noise suppression on the original image to improve the signal-to-noise ratio of a target area; s3, dynamic self-adaptive registration: according to the biomechanical characteristics of the organ, fusing the rigid transformation model and the elastic deformation model, and according to the digital information processing method for the hospital radiology department, based on the dynamic registration matrix of the biomechanical model, improving the multi-modal image fusion precision; a deep learning segmentation algorithm fused with morphological constraints improves the focus boundary recognition accuracy; a texture mapping three-dimensional reconstruction technology is mixed, and an anatomical structure and metabolism information are presented at the same time; the invention discloses a structured report automatic generation system based on an attention mechanism.
Owner:SHENZHEN SECOND PEOPLES HOSPITAL (SHENZHEN INST OF TRANSLATIONAL MEDICINE)

Malicious code detection method, device and equipment based on multi-modal feature fusion

The invention discloses a malicious code detection method, device and equipment based on multi-modal feature fusion, relates to the technical field of deep learning, and aims to solve the problem that existing malicious code detection is low in accuracy and reliability. The method comprises the steps of performing multi-modal feature extraction on original code data, generating a binary texture image, a frequency domain energy distribution image, an information entropy thermodynamic image and an operation code sequence feature, performing parallel processing on a multi-modal image through a heterogeneous convolutional neural network CNN, outputting a structure mode feature of a malicious code, and obtaining an operation code sequence feature of the malicious code. And modeling operation code sequence features by adopting a long short-term memory (LSTM) network, outputting behavior intention features of malicious codes, aligning structural mode features and the behavior intention features through a cross-modal attention mechanism, and generating malicious code classification tags, the classification tags being used for representing confidence that original code data are malicious codes.
Owner:SHANXI UNIV

Multi-modal image matching method and system based on saliency graph structure enhancement

The invention discloses a multi-modal image matching method and system based on saliency graph structure enhancement, and belongs to the field of image processing. The method comprises the following steps: firstly, innovatively constructing a pixel-level saliency confidence graph for measuring the matching potential of each region, and guiding an attention mechanism to be dynamically focused on a key region in a graph structure through the graph; secondly, multi-scale structure features and semantic segmentation information are fused, and the semantic perception ability of feature expression is enhanced; and finally, constructing two heterogeneous graph structures of an in-image structure graph and an inter-image semantic guidance graph, and realizing global-local information enhancement and cross-modal semantic alignment on the graph structures by introducing a self-attention and cross-attention mechanism of saliency modulation, so that the matching precision and stability are remarkably improved, and the matching accuracy is improved. And semi-dense matching of multi-modal images is realized.
Owner:WUHAN UNIV

Multi-modal image threshold segmentation preprocessing method based on convolutional neural network

The invention relates to a multi-modal image threshold segmentation preprocessing method based on a convolutional neural network, and the method comprises the steps: unifying an image into a standard space, carrying out the pixel value mapping, carrying out the resampling, generating high and low frequency sub-bands, carrying out the soft threshold denoising of the high frequency sub-bands, enhancing the contrast of the low frequency sub-bands, and carrying out the fusion; an optimized VGGnet framework is constructed; a noise adversarial network is generated to carry out active learning loop training on a convolutional neural network model; the image is input into the model for prediction; a local entropy and a gradient magnitude are calculated based on a prediction result; optimal segmentation is realized by setting a double-layer matrix of a feature tag-segmentation method; grey matter Dice calculation is carried out on the segmented images, and preprocessing parameters of unqualified images are optimized through a dynamic parameter adjusting module based on a Gaussian process regression model. The segmentation precision and the processing efficiency of the multi-modal image are effectively improved, and the adaptability of the model to a complex image is enhanced.
Owner:川北医学院附属医院 +1

Cervical cancer close-range radiotherapy high-risk target area sketching method fused with multi-modal image

The invention discloses a cervical cancer close-range radiotherapy high-risk target area sketching method fused with a multi-modal image. The method comprises the following steps of collecting an MRI image scanned before radiotherapy and a CT image during radiotherapy of a cervical cancer close-range radiotherapy patient and annotation data of the MRI image and the CT image; the method comprises the following steps: preprocessing an MRI image scanned before radiotherapy and a CT image during radiotherapy, and converting annotation data into a three-dimensional tag image; on the basis of the preprocessed MRI image, the preprocessed CT image and the three-dimensional label image of the preprocessed MRI image, the preprocessed CT image and the three-dimensional label image of the preprocessed MRI image, registration of the MRI image and the CT image is conducted through a pre-constructed registration neural network model, feature extraction and fusion are conducted on the registered MRI image and the registered CT image through a pre-constructed high-risk target area automatic segmentation neural network model of a multi-scale cross-modal attention mechanism, and the high-risk target area automatic segmentation neural network model of the multi-scale cross-modal attention mechanism is obtained. Automatic delineation of a cervical cancer close-range radiotherapy high-risk target area is realized; according to the method, automatic segmentation of HR-CTV in close-range radiotherapy of cervical cancer is realized, high efficiency, accuracy and generalizability are realized, and intelligent support can be provided for clinical work.
Owner:XIANGYA HOSPITAL CENT SOUTH UNIV

Robot electrical equipment defect detection system based on multi-modal image processing

The invention provides a robot electrical equipment defect detection system based on multi-modal image processing. The system improves the accuracy and reliability of electrical equipment defect identification. Infrared and visible light images are jointly collected, and through a registration algorithm of multi-source features and equipment structure priori, space-time alignment of multi-modal images is achieved. Then, a dynamic weighted fusion strategy is utilized to generate fusion features with higher discriminative ability, and abnormal features are extracted through a double-branch mechanism to be verified with thermophysical consistency; and finally, constructing a neural network model fused with physical prior, and performing defect classification and positioning output on the verified feature data. According to the method, the structure and thermal information are fused, a physical constraint mechanism and a joint training strategy are introduced, the robustness and engineering interpretability of the system under complex working conditions are remarkably improved, and the method has a wide application prospect.
Owner:NANJING DONGXIN HUIKE INFORMATION TECH CO LTD

Hyperspectral and multispectral image fusion method based on wavelet feature fusion and comparative learning

The invention discloses a high-resolution hyperspectral image reconstruction method based on wavelet domain feature fusion and contrast learning, and belongs to the technical field of image fusion and super-resolution reconstruction. The method comprises the following steps: constructing a fusion network model comprising a wavelet transformation module, a cross-modal feature fusion module, a high-frequency contrast learning module and an image reconstruction module; performing end-to-end supervised training by using a training data set containing the low-resolution hyperspectral image and the high-resolution multispectral image; and after training is completed, inputting a test image pair to realize image reconstruction. According to the method, the detail retention capability is improved by combining wavelet decomposition and a directional fusion mechanism, the cross-modal high-frequency feature alignment capability is enhanced through comparative learning, a fusion image with high spatial resolution and high spectral consistency is finally generated, and the method is suitable for multi-modal image reconstruction tasks such as remote sensing, medical and natural images.
Owner:DONGHUA UNIV

Multimodal image matching method and system, terminal device, and storage medium

Provided are a multimodal image matching method and system, terminal device, and storage medium. The method includes: performing self-supervised feature extraction on an optical image and a synthetic aperture radar (SAR) image to obtain a repetitive feature point between the optical image and the SAR image; segmenting the optical image and the SAR image into a first image block sequence based on the repetitive feature point, and performing feature extraction on the first image block sequence through a dual-branch network to obtain feature description vectors of the optical image and the SAR image respectively, where the dual-branch network includes a first branch network for extracting a global feature and a second branch network for extracting a local feature; and performing feature matching on the optical image and the SAR image based on the feature description vectors to obtain a matching point pair between the optical image and the SAR image.
Owner:SUN YAT SEN UNIV

Ultrasonic image offline acquisition and dynamic synchronous processing method and system

The invention discloses an ultrasonic image off-line acquisition and dynamic synchronous processing method and system, which realizes off-line acquisition and dynamic synchronous processing of an ultrasonic image under the condition of network interruption through six functional modules, namely a network state monitoring module, an online data synchronization module, an image acquisition processing module, an off-line mode switching module, an off-line diagnosis processing module and an increment synchronization module. The method comprises the following steps: acquiring network connection state data and judging a working mode; pulling and caching tasks and template data in an online state; obtaining multi-modal image data and performing double-writing storage; detecting network interruption and switching to an offline mode; performing AI analysis on the images acquired offline and generating a report; and detecting network recovery and executing incremental data synchronization. According to the invention, the problem of service interruption caused by dependence of a traditional PACS system on a network is solved, the network adaptability and robustness of the system are improved, the operation and maintenance cost is reduced, and the system fault recovery time is shortened.
Owner:GUIZHOU PRECISION HEALTH DATA CO LTD

Gynecological tumor image processing method and system based on AI multi-modal image analysis

The invention belongs to the field of image processing, and provides a gynecological tumor image processing method and system based on AI multi-modal image analysis, and the method comprises the steps: 1, obtaining an original image of a patient, and obtaining a structure mask and an image frame sequence after period alignment and structure normalization based on the original image; step 2, obtaining a focus mask sequence after structure limitation based on the image frame sequence; step 3, respectively acquiring a modal structure semantic tensor of each image in the image frame sequence, and acquiring a fused semantic feature tensor based on the modal structure semantic tensor; 4, obtaining a final focus mask based on the fused semantic feature tensor and the structure mask; and step 5, obtaining a response visualization graph based on the focus mask. The method is clear in technical structure, coherent in task chain and independent in model interface, has real deployment and continuous evolution capabilities, and is particularly suitable for gynecological image AI auxiliary system scenes under periodic driving.
Owner:THE THIRD AFFILIATED HOSPITAL OF SOUTHERN MEDICAL UNIV (ACAD OF ORTHOPEDICS GUANGDONG PROVINCE)

Bone model manufacturing method based on 3D printing technology

The invention discloses a skeleton model manufacturing method based on a 3D printing technology, and relates to the technical field of medical treatment, and the method comprises the steps: collecting the multi-modal image data of the skeleton of a patient, carrying out the image segmentation and reconstruction, and generating a digital three-dimensional model of the skeleton of the patient; carrying out mechanical optimization and material distribution adjustment on the digital three-dimensional model by utilizing a topological optimization algorithm, designing a bionic microstructure, and outputting a three-dimensional skeleton model subjected to optimization and bionic design; the bone model which is detected to be qualified is used for generating a personalized mold, and a prosthesis which is completely matched with the bone of the patient is customized through the mold; preoperative simulation and design of a prosthesis implantation scheme are carried out based on the skeleton model, and a personalized prosthesis and an operation scheme matched with the skeleton of the patient are obtained; the structural rigidity is maximized and the material consumption is minimized by utilizing a topological optimization algorithm, so that the bone model can better meet the biomechanical requirements of a patient, and meanwhile, the bionic characteristics and mechanical properties of the model are enhanced through the bionic microstructure design.
Owner:BEIJING KEFEI JINCHENG TECH CO LTD

Visual navigation method based on tumor interventional surgical robot

The invention relates to the technical field of tumor interventional operations, and discloses a visual navigation method based on a tumor interventional operation robot. The method comprises the following steps: acquiring real-time medical image data of a tumor area containing multi-modal imaging information so as to comprehensively present anatomical details; and performing three-dimensional reconstruction on the image data to generate a tumor area three-dimensional anatomical structure model capable of visually displaying a space structure. Key anatomical feature points are extracted based on the model, space coordinates are calculated, a surgical robot intervention path is planned according to the coordinates, and an initial navigation track is generated; and continuously collecting real-time pose data of the robot in an operation, dynamically matching the real-time pose data with the initial navigation trajectory, adjusting motion parameters according to a matching result, and generating a corrected navigation instruction. The method can reflect the intraoperative anatomy condition in real time, dynamically optimize the path, solve the problems that traditional navigation depends on preoperative static images and lacks real-time adjustment, reduce operative complications and improve the treatment effect of patients.
Owner:HE BEI SHENG ZHONG YI YUAN (FIRST AFFILIATED HOSPITAL OF HEBEI UNIVERSITY OF TRADITIONAL CHINESE MEDICINE HEBEI CENTER FOR PREVENTION & CONTROL OF SCOLIOSIS IN CHILDREN & ADOLESCENTS)

Detection method and device for intelligent visual detection of parts and storage medium

The invention discloses a detection method and device for intelligent visual detection of parts and a storage medium, belongs to the technical field of industrial quality detection, and aims to solve the problems that surface and internal defects of complex parts are difficult to identify synchronously and the identification precision is low. The method comprises the following steps: acquiring multi-modal image data obtained through structured light three-dimensional imaging and laser ultrasonic scanning; performing geometric registration and scale normalization on the image to generate a fused image; carrying out image preprocessing and edge analysis, and extracting a region of interest; dividing the candidate region of interest into a plurality of image blocks, and inputting the image blocks into an anomaly detection network; calculating a reconstruction error between the original image block and the reconstructed image block, and generating an abnormal scoring graph; and finally, extracting a defect area through image post-processing, and outputting information such as a defect type, a spatial position, a geometric dimension and a severity level. According to the method, the robustness and accuracy of multi-modal defect identification are improved, and the method is suitable for an online visual inspection task of an industrial production line.
Owner:NINGBO CITY QIQIANG PRECISION STAMPINGS +2

Face multimode image feature collaborative retrieval method

The invention provides a face multimode image feature collaborative retrieval method, and aims to solve the defects of the prior art in the aspects of multimode feature processing, complex scene adaptability and the like. The method comprises the following steps: firstly, acquiring multi-modal data such as visible light, infrared and depth images, pre-processing the multi-modal data, and inputting the pre-processed multi-modal data into a retrieval model comprising a multi-modal feature decoupling layer, a collaborative perception dynamic fusion layer and a cross-modal attitude measurement learning layer; wherein the decoupling layer separates modal exclusive features from shared identity features, eliminates semantic differences while keeping modal uniqueness, and enhances feature space consistency; when shielding is detected, the dynamic fusion layer allocates modal weights according to scenes, compensates shielding features and generates fusion features; and finally, outputting a result through metric learning optimization in combination with a hierarchical strategy. According to the scheme, through feature decoupling and dynamic fusion, the bottleneck of a traditional method under feature integration and complex scenes is broken through, and the retrieval accuracy and stability are improved.
Owner:CHONGQING UNIV OF TECH

Robot production line article grabbing method and system based on visual positioning

The invention provides a robot production line article grabbing method based on visual localization, which comprises the following steps: preprocessing a multi-modal image to obtain an original image; performing grid mapping on the original image to obtain a multi-level feature descriptor; based on the multi-level feature descriptors, a dynamic mapping relation among the three coordinate systems is established through a non-rigid coordinate system alignment algorithm, and coordinate offset of movement of the conveyor belt is compensated in real time; performing hierarchical feature matching on the multi-level feature descriptors and an article template library, identifying article categories and extracting contour geometric features; based on the contour geometric features and the surface curvature distribution, the three-dimensional pose and candidate grabbing points of the object are obtained through a geometric constraint optimization algorithm; and generating a robot obstacle avoidance track according to the candidate grabbing points and the robot motion model. According to the method, accurate coordinate compensation is realized through multi-modal data fusion and dynamic adaptive grid mapping, and the article positioning accuracy is improved in combination with hierarchical feature matching.
Owner:SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY)

Visible light and infrared image depth fusion auto-encoder network model based on Haar wavelet transform

The invention relates to a visible light and infrared image deep fusion auto-encoder network model based on Haar wavelet transform, and belongs to the technical field of multi-modal image fusion. The model comprises an image feature extraction module, an image feature fusion module and an image reconstruction module. The image feature extraction module comprises shallow shared feature extraction, coarse-grained feature extraction and fine-grained feature optimization; the image feature fusion module fuses high-frequency and low-frequency features by using Haar wavelet inverse operation; and the image reconstruction module carries out image reconstruction by using a Decoder module. The self-encoder network model has good performance, particularly, the high-frequency and low-frequency features of the infrared picture and the visible light picture reserved in the fused image are rich, fusion is fast, and help is provided for improving the accuracy of downstream visual tasks such as target detection, target tracking and target segmentation.
Owner:CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI