Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

307 results about "Bi modal" patented technology

Retrieval enhancement method based on multi-modal data fusion and modal perception

The invention relates to the technical field of information retrieval and generation, in particular to a retrieval enhancement method based on multi-modal data fusion and modal perception. According to the method, firstly, a dual-channel architecture is adopted to perform feature extraction and coding on a text and an image respectively, and mutually independent embedded representation spaces are constructed, so that high-quality collaboration and matching of cross-modal representation are realized; and a pseudo-pairing generation mechanism is introduced to effectively mine and reconstruct the existing non-paired data in the knowledge base. And designing a query modal perception and dynamic weighting mechanism for accurately controlling the fusion proportion of the image-text bimodal information in the retrieval stage so as to match the modal demand difference of different query contents. And further executing aggregation retrieval and reordering of the cross-modal information by using dynamic weighted fusion retrieval to generate a candidate set of multi-modal responses. According to the method, accurate matching and dynamic weight adjustment of the image-text content are realized, and the accuracy and expression integrity of the generated content are improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Multi-mode ultrasonic fusion pressure vessel welding seam defect nondestructive testing method and multi-mode ultrasonic fusion pressure vessel welding seam defect nondestructive testing system

The invention provides a multi-mode ultrasonic fusion pressure vessel weld defect nondestructive testing method and system, and relates to the technical field of nondestructive testing. According to the method, geometric parameters of a welding seam are obtained through three-dimensional laser scanning, and an optimal scanning parameter set is generated; driving ultrasonic phased array equipment to scan for one time and synchronously acquire shear wave full-matrix capture and longitudinal wave linear scanning data; performing energy flow angular spectrum analysis and envelope analysis on the bimodal data, extracting defect feature parameters and constructing a three-dimensional feature tensor; carrying out multi-dimensional feature fusion by adopting Tucker decomposition, and enhancing a core tensor through physical modeling; generating three types of defect indication diagrams including a defect existence possibility diagram, a defect relative scale diagram and a defect space orientation diagram from the enhanced feature tensor; and the three types of indication diagrams are visually presented for comprehensive interpretation of detection personnel. Through multi-modal data fusion and physical modeling enhancement, the defect identification accuracy and detection efficiency are remarkably improved, the false alarm rate is reduced, and reliable technical support is provided for pressure vessel welding seam safety detection.
Owner:YUNNAN SPECIAL EQUIP SAFETY TESTING RES INST

Water surface target detection method based on bimodal image feature fusion

The invention discloses a water surface target detection method based on bimodal image feature fusion, and relates to the technical field of water surface target detection, and the method comprises the steps: constructing a detection network based on bimodal feature enhancement, a backbone network of an infrared image input end of the detection network and a backbone network of a visible light image input end of the detection network are the same as YOLOv5 in structure, when the backbone network of the infrared image input end and the backbone network of the visible light image input end carry out feature transmission, feature fusion of infrared light and visible light is completed through a fusion enhancement module; the fused features are respectively transmitted back to the backbone network of the infrared image input end and the backbone network of the visible light image input end for further feature extraction; and detecting a water surface target by the trained detection network. According to the method, bimodal feature extraction is completed by expanding a backbone network of YOLOv5, a fusion enhancement module is introduced, and visible light and infrared feature fusion is completed in the feature extraction process.
Owner:CHINA SHIP DEV & DESIGN CENT

Physical examination data analysis method and system based on artificial intelligence

The invention discloses a health examination data analysis method and system based on artificial intelligence, and the method comprises the steps: receiving multi-modal health examination original data of a user, and generating a multi-modal feature matrix fused with time-space correlation; inputting the multi-modal feature matrix into a lightweight dual-channel network to obtain a decoupled dual-modal feature group; constructing a hidden Markov chain based on the bimodal feature group, and generating a health state transition trajectory diagram with a timestamp; inputting the health state transition trajectory diagram into a discriminator of the generative adversarial network, and outputting a confidence score and a pathology trigger threshold of a high-risk node; and activating a rule engine according to a pathological trigger threshold, and generating a hierarchical health intervention instruction set in combination with individual living habit data. By using the embodiment of the invention, the physical examination data can be efficiently analyzed, and the accuracy of disease early warning and health intervention is optimized.
Owner:ZHEJIANG KANGLUE SOFTWARE CO LTD

Bimodal image fusion method and device based on tuple disturbance and storage medium

The invention provides a dual-mode image fusion method and device based on tuple disturbance and a storage medium, and relates to the technical field of image fusion processing. The method comprises the following steps: constructing a model comprising an encoder, a feature fusion network and a decoder by collecting a cross-modal image group; an encoder is trained based on a tuple disturbance comparison learning framework, cross-modal sharing and complementary features can be mined, learnable positive and negative sample pairs are constructed through a tuple disturbance module, dependence on the same mode is avoided, and generalization is enhanced; the feature fusion network performs complementary feature fusion and multi-scale semantic aggregation training by using cross-modal deep features to improve the quality of a fused image; multi-scale semantic aggregation solves the problem of detail loss, and spatial feature calibration strengthens details, so that a fused image retains bimodal characteristics, and semantic consistency is enhanced; and a combined loss function is constructed, a pixel intensity and gradient structure difference optimization model is integrated, the performance and robustness are improved, key information of the fused image is ensured, and the visual effect and the accuracy are better.
Owner:WUHAN INST OF TECH +2

Audio and video object intelligent tracking optimization method and system combined with deep learning

The invention relates to the technical field of audio and video processing, and provides an audio and video object intelligent tracking optimization method and system combined with deep learning. The method comprises the following steps: performing cross-modal feature collaborative extraction on an audio stream and a video frame sequence by acquiring a synchronous audio and video data group, and generating a multi-modal feature set containing audio time domain dynamic features and video space structure features; inputting the multi-modal feature set into a pre-trained association enhancement network to generate a cross-modal semantic aligned association feature sequence; constructing a tracking stability evaluation model based on the associated feature sequence, and outputting a stability index; tracking parameters are dynamically adjusted according to the stability index, an initial tracking result is calibrated, and an optimized tracking trajectory is output. Therefore, the precision and stability of object tracking in a complex scene are improved by deeply fusing the dual-mode characteristics of the audio and the video, mining the internal association between the modes and combining a dynamic evaluation and calibration mechanism.
Owner:SHENZHEN ZIDOO TECH CO LTD

Multi-modal emotion recognition method and system for service-oriented robot

The invention belongs to the technical field of artificial intelligence, and particularly relates to a service-oriented robot-oriented multi-modal emotion recognition method and system, and the method comprises the steps: collecting audio and video stream data of emotion changes of a user, and separating visual and voice data; extracting visual and voice emotion features through a pre-training model, and calculating prediction probability distribution of each mode; constructing a bimodal confidence quantitative model based on the distribution to obtain each modal confidence; and fusing the features by adopting a sectional type dynamic weight distribution strategy so as to identify the emotional state of the user. Visual and voice modes are fused, feature alignment is realized in combination with dynamic time warping, spatial optimization performance is shared and expressed through a confidence model, a dynamic weight strategy and a cross-modal time sequence cooperation module, and the method has high recognition accuracy, high robustness and real-time processing capacity in a complex environment and is suitable for various service scenes.
Owner:SUZHOU CITY UNIV

Knowledge distillation method for target detection

The invention discloses a knowledge distillation method for target detection. The knowledge distillation method comprises two major designs including a dynamic teacher-student mutual learning mechanism and interpretable feature decoupling. The RGB detector is used as a teacher guidance event detector to learn static semantic knowledge in a conventional illumination scene, and the event detector is used as a teacher guidance RGB detector to learn dynamic robust features in an unconventional illumination scene, so that dual-mode collaborative optimization is realized; and the original teacher and student feature space is decoupled into three parts of mode universality, mode specificity and mode irrelevance, and effective knowledge migration based on the mode universality feature is carried out between the two modes by constructing bidirectional distillation loss. The objective of the invention is to enable an RGB detector and an event detector to have a more accurate target recognition effect.
Owner:HUNAN UNIV

Shield muck volume calculation method and device based on machine vision

The invention discloses a shield muck volume calculation method and device based on machine vision, and relates to the field of image processing, and the method comprises the steps: carrying out the feature extraction and fusion of a target depth image and a target RGB image obtained in real time through an image segmentation network based on bimodal input, and obtaining a target depth image and a target RGB image; performing image segmentation on the target RGB image according to the obtained fusion features to obtain a residue soil mask image of the target RGB image and coordinates of a residue soil bounding box; according to the coordinates of the muck bounding box and the muck mask image, the volume of the shield muck is calculated through a volume calculation method; according to the method, the accuracy of volume calculation of the shield muck is improved on the basis of good real-time performance, and the problem that the accuracy of the volume of the shield muck calculated by an existing non-contact measurement technology of the volume of the shield muck is low under the high real-time performance requirement is solved.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Intelligent cabin man-machine cooperative control method and device, vehicle, medium and product

The invention relates to the technical field of vehicles, in particular to an intelligent cabin man-machine cooperative control method and device, a vehicle, a medium and a product. Performing fusion processing on the voice signal data, the visual signal data and the tactile signal data of the current driver by using a preset dynamic weight distribution model to obtain a multi-modal fusion control instruction; and under the condition that the multi-mode fusion control instruction meets a preset sight-voice dual-mode condition, synchronously issuing the multi-mode fusion control instruction to a plurality of target execution devices in the intelligent cabin so as to drive each target execution device to perform cooperative control. Therefore, by constructing a cooperative control mechanism of multi-mode sensing and dynamic weight distribution, the problems that the false wake-up rate is high, the voice and graphical interface cooperative efficiency is low and the interaction mode cannot adapt to a complex driving environment due to a single or separated interaction mode in the related technology are effectively solved.
Owner:CHERY AUTOMOBILE CO LTD

Method and device for filling and quantifying multi-source data of power system based on bimodal hybrid algorithm in extreme weather

The invention discloses a method and a device for filling and quantifying multi-source data of a power system based on a bimodal hybrid algorithm in extreme weather, which belong to the technical field of power distribution networks and are based on a reverse chaos whale transfer algorithm and combined with meteorological characteristics to dynamically adjust and construct a load prediction model to fill missing data and output distribution parameters and weight coefficients. Meanwhile, a meteorological condition matrix is introduced, a Gaussian mixture model is constructed, and the uncertainty of data is quantized and filled through a Monte Carlo subset deduction algorithm. The method effectively solves the problems that a traditional method is low in processing efficiency and large in error in extreme weather, experiments show that the method is high in convergence speed and high in estimation precision, the accuracy and reliability of power system data are remarkably improved, and guarantee is provided for stable operation of a power dispatching system.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

Art and craft material detection method based on multi-modal deep learning

The invention discloses an industrial art material detection method based on multi-modal deep learning, particularly relates to the field of material analysis, and is used for solving the problems that cross-modal data alignment of curved surface utensils in highlight and multi-layer coating scenes is difficult, and material boundaries are easy to drift along with shooting batches and visual directions. The method comprises the following steps of: positioning a hyperspectral line scanning fragment; constructing a local reference grid to realize initial pairing; generating pixel-level residual image quantization dislocation; estimating luminosity mapping to form consistent bimodal data pairs after eliminating a high-reflectivity region; iteratively supplementing anchor points or adjusting grid rigidity, generating a stable material graph, updating an index table and a consistency record, ensuring cross-batch reusability and traceability, and improving the material detection precision.
Owner:WUXI GONGCHUN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Infrared and visible light image fusion method based on text semantic consistency guidance

The invention provides an infrared and visible light image fusion method based on text semantic consistency guidance, which relates to the technical field of multi-modal image fusion, and comprises the following steps: respectively carrying out fine-grained text semantic generation on infrared and visible light images and mapping the infrared and visible light images to a unified embedding space; bidirectional compensation and enhancement are carried out on text semantics through a cross-modal attention mechanism, and unified text semantic priori is constructed; a structure-intensity decoupling double-branch encoder is adopted, and structure texture features of visible light and intensity significant features of infrared light are extracted respectively; under the prior guidance of text semantics, the bimodal visual features are aligned in a shared semantic space through explicit semantic consistency constraint and implicit semantic distribution consistency constraint; and finally, taking the text semantic priori as a global modulation signal, and carrying out adaptive weighted fusion and decoding on the aligned features to generate a fused image. The problems that an existing method is insufficient in semantic modeling and poor in fusion result consistency are effectively solved.
Owner:XIAMEN UNIV OF TECH

Automatic material taking mechanism for cable processing equipment

The invention relates to the technical field of cable processing, in particular to an automatic material taking mechanism for cable processing equipment, and the mechanism comprises an incoming material arraying and conveying unit which is used for flattening scattered cables and conveying the cables through a conveying belt synchronized with an encoder; the industrial vision or AI vision unit is used for performing vision detection and vision measurement on the cable in the conveying process; the material taking execution unit comprises an execution mechanism with the X-Y-theta degree of freedom and the Z degree of freedom and a bimodal flexible gripper; the posture correcting and guiding unit is used for correcting the direction of the end of the grabbed cable and standardizing the feeding posture of the grabbed cable; according to the automatic material taking mechanism for the cable processing equipment, the supplied material arraying and conveying unit is arranged to be combined with the industrial vision or AI vision unit, and the supplied material arraying and conveying unit is arranged to be combined with the industrial vision or AI vision unit; the defects that vibration stop photographing or a general taking and placing system is low in flexible cable separation efficiency and prone to missing detection are effectively overcome.
Owner:NUO XUN (JIANGSU) CABLE TECH CO LTD

Driving authority management method and device based on biological characteristic verification and medium

The invention provides a driving authority management method and device based on biological feature verification and a medium, and belongs to the technical field of vehicles. The method comprises the steps of collecting facial features, voiceprint data and driving behavior data by detecting a starting operation of a current driver; the identity of the current driver is recognized by using a multi-modal biological recognition algorithm, and the accuracy and anti-counterfeiting capability of identity verification are improved by adopting a bimodal fusion scheme of face recognition and voiceprint recognition; when it is determined that identity recognition of the current driver succeeds, a driving account is determined to judge whether the driver has the use permission or not, a binding relation of driver identity-account-function permission is established, and accurate matching of the intelligent driving permission and the driver qualification is achieved; the open state of the intelligent driving function is dynamically controlled based on a permission verification result, if the permission exists, the intelligent driving function is automatically prepared, if the permission does not exist, the function is locked, a clear prompt is given, and the safety risk that the intelligent driving function is used before being learned is avoided from the source.
Owner:CHINA FAW CO LTD

Prompt guidance and multi-modal fusion-based class incremental learning method

The invention provides a class incremental learning method based on prompt guidance and multi-modal fusion, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: firstly, performing semantic extension on a category label, and constructing semantic enhanced text representation through a text encoder; then block embedding and hierarchical feature extraction are carried out on the input image by using a pre-trained visual encoder, a cross-modal unified embedding space is constructed, a bimodal prompt gating fusion module is introduced into the unified embedding space, and adaptive weighting is carried out on text prompt and image prompt according to gating weight to generate fusion prompt; through a bimodal prompt collaborative filtering module, screening out a prompt set most relevant to the current task according to the similarity of the semantic features of the image and the text; the pre-training backbone network is frozen in the increment stage, only prompt parameters and fusion layer weights are optimized, a joint loss function is used for parameter updating, finally, image and text data are input in the reasoning stage, cross-modal similarity is calculated, and a classification prediction result is output.
Owner:NORTHEASTERN UNIV CHINA

Vending machine multi-mode emotion perception and personalized dialogue generation method and system

The invention discloses a vending machine multi-mode emotion perception and personalized dialogue generation method and system, and the method comprises the steps: collecting a face image and a voice signal of a user through a camera and a microphone, and outputting a dual-mode emotion tag through the analysis of a face expression and a voice emotion; fusing the label and the dialogue text to generate a structured Prompt containing emotion, intention and context; based on an Agent framework multi-modal large model, combining context history and a personalized strategy to generate a reply text matched with the emotional scene; according to the text and the comprehensive emotion label, TTS parameters are dynamically adjusted, and anthropomorphic voice is output; in multiple rounds of dialogues, styles are kept consistent through an emotion smoothing formula, and emotion-verbal skill-conversion data are precipitated in combination with a user feedback optimization strategy. The interaction limitation of a traditional vending machine is broken through, multi-mode emotion perception, personalized personification interaction and continuous evolution are achieved, the vending machine is promoted to be upgraded to an emotional shopping guide platform, user experience and sales transformation are improved, and a technical model is provided for intelligent retail.
Owner:SHANGHAI QUZHI NETWORK TECH CO LTD

High-precision face recognition method and system

The invention discloses a high-precision face recognition method and system which are suitable for complex illumination and unconstrained position scenes. The method is realized through cooperation of a multi-modal acquisition guide strategy driven by a state machine and a self-adaptive fusion mechanism based on a physical multiplier array: the state machine controls logic to generate a position-attitude sequence instruction, and drives a user to synchronously provide RGB-NIR image pairs in upper, middle and lower screen areas and left and right side face attitudes; the double-flow convolutional neural network extracts an aligned bimodal feature map through shared weight constraint; the attention fusion module dynamically generates a space-channel joint attention weight based on a local signal-to-noise ratio difference through a multiplier array and a weight register realized by hardware, and performs pixel-by-pixel weighted fusion on the bimodal features. And a physical mapping relation between state machine position change and modal weight distribution is established through coupling training. According to the invention, through algorithm hardening and hardware cooperation, the accuracy and real-time performance of cross-illumination and cross-position unconstrained face recognition are significantly improved.
Owner:BEIJING AUTO SMART INFORMATION TECH CO LTD

Method and system for tracking and detecting sheltered target in electric power scene based on Leiyu-vision integrated multi-mode space-time adversarial learning

The invention discloses a method and a system for tracking and detecting an occluded target in an electric power scene based on thunder-vision integrated multi-mode spatio-temporal adversarial learning, and relates to the field of intelligent inspection of electric power equipment, and the method comprises the steps: inputting a structured point cloud matrix and a visual image into a double-branch feature extraction model, respectively extracting geometric features and semantic features by using a double-branch feature extraction model; inputting the bimodal fusion features into a dynamic weight adaptive fusion mechanism based on an adversarial training strategy to obtain a fusion weight coefficient of the target in a short-time shielding state in the power scene; and a double-buffer memory pool is constructed, and a three-level progressive trajectory prediction compensation mechanism is combined to compensate a target trajectory position in a long-time shielding state in a power scene. According to the method, the shielding state quantitative model is constructed, the fusion weight distribution proportion can be effectively and dynamically adjusted in real time according to the visual sensor, and when it is detected that the target equipment enters the shielding state, the visual data weight is automatically reduced to the safety threshold value.
Owner:NANJING STEIN SMART ENERGY TECH CO LTD

Unmanned aerial vehicle small target detection method and system based on bimodal image fusion

The invention discloses an unmanned aerial vehicle small target detection method and system based on bimodal image fusion. The target detection method comprises the steps of obtaining a to-be-detected unmanned aerial vehicle image; the to-be-detected unmanned aerial vehicle image is input to a target detection model, a detection result is obtained, and the target detection model comprises a backbone network module, a neck network module, a detection module and an output module; the backbone network module is used for acquiring feature maps of different scales; the neck network module is used for performing different-level feature fusion on the feature map to obtain a fused feature map; the detection module is used for detecting the fused feature map to obtain a detection result; and the output module is used for outputting the detection result. The problems of low detection precision and high omission ratio caused by small target size and large modal information difference in small target detection of the unmanned aerial vehicle are solved.
Owner:YANGTZE UNIVERSITY

Skin lesion auxiliary detection method and system based on multi-modal information fusion

The invention provides a skin lesion auxiliary detection method and system based on multi-modal information fusion, and belongs to the technical field of computer vision and application. The method comprises the following steps: firstly, acquiring and constructing bimodal data including a dermatoscope image and a text of a patient, and preprocessing the data; secondly, respectively adopting an improved convolutional neural network and a text coding model, and extracting deep and high-dimensional features from the skin disease image and the self-described text of the patient; afterwards, an improved symmetric gating cross attention fusion module is constructed, fine-grained interaction and alignment between visual and text features are realized through a bidirectional cross attention mechanism, and a gating unit is utilized to adaptively adjust a fusion weight; and finally, inputting the deeply fused multi-modal feature vector into a classifier to obtain a confidence score of each skin lesion type so as to provide reference for dermatologists. According to the method, the diagnosis accuracy, robustness and interpretability are synergistically improved, and the method has great clinical application value.
Owner:SHENYANG SEVENTH PEOPLES HOSPITAL

Cross-modal attention collaborative jail break attack method

The invention discloses a cross-modal attention collaborative jail break attack method, and belongs to the technical field of artificial intelligence security. The method comprises the following steps of: constructing an input sequence representation containing a system prompt, an adversarial image representation, a malicious query and an adversarial text suffix according to a causal self-attention mechanism; inputting the sequence into a visual language model to execute forward propagation; based on the designed attention-oriented loss collaborative function, optimizing adversarial image representation through a joint gradient optimization algorithm and updating an adversarial text suffix to optimize an attack target; iteratively circulating until convergence, and outputting the optimized unified multi-modal knowledge; and finally, utilizing the knowledge to construct an attack sequence to realize jailbreak. According to the method, accurate control on an internal attention mechanism of the visual language model is realized for the first time, and through visual-text dual-mode cooperative attack, the attack success rate is remarkably improved while high concealment is kept, and the important driving force for promoting the progress of a safe alignment technology is achieved.
Owner:NAT UNIV OF DEFENSE TECH

Litchi phenotype identification and weight estimation method based on multi-modal learning

The invention relates to the technical field of litchi phenotype analysis, in particular to a litchi phenotype recognition and weight estimation method based on multi-modal learning, and adopts the technical scheme that an RGB (Red, Green, Blue) image and a depth image are used as bimodal input and are input into a LitchiPhenoNet model to realize detection and segmentation of fruits; an RD fusion module is adopted to significantly improve the multi-modal feature expression capability and segmentation performance, and the fusion strategy integrates texture details of an RGB image and space structure information of a depth image, so that the accuracy and robustness of detection and segmentation are improved; on the basis of geometric features and abstract intermediate features directly extracted by LitchiPhenoNet, fruit weight estimation is performed by adopting a multi-modal regression method, a regression model combines information from two aspects of vision and physics, and by combining the two types of features, the regression model can effectively improve the utilization efficiency and prediction precision of multi-modal data, so that the prediction accuracy of the fruit weight is improved. The multi-modal regression framework effectively integrates complementary information of geometric features and intermediate visual features, so that the accuracy and robustness of litchi weight estimation are improved.
Owner:TROPICAL FRUIT RES INST HAINAN ACADOF AGRI SCI

Multi-modal model construction method and system for predicting efficacy of sorafenib in hepatocellular carcinoma

The present invention provides a multi-modal model construction method and system for predicting the efficacy of sorafenib in hepatocellular carcinoma. The method comprises: step 1, collecting clinical information of a target patient, and generating a whole slide image; step 2, preprocessing clinical data, and retaining clinical features as input for a multi-modal deep learning model; step 3, preprocessing the whole slide image; step 4, constructing an image model, acquiring patch-level scores of the pathological image on the basis of the preprocessed image and by using different aggregation algorithms, and predicting the score of the whole pathological image to obtain best model features; step 5, constructing a multi-modal model, performing modal fusion on the best model features and the clinical features, and outputting an image-level or patient-level prediction result; and step 6, testing and evaluating the model. The present invention achieves bimodal input of a pathological image and clinical information, fully utilizes the complementarity of the two types of modal data, and thus improves prediction accuracy.
Owner:CENT HOSPITAL OF MINHANG DISTRICT SHANGHAI +1

Tumor identification method based on multi-modal feature fusion

The invention provides a tumor recognition method based on multi-modal feature fusion, relates to the technical field of tumor recognition, and is used for solving the problems of missing multi-modal feature calibration, insufficient semantic association, poor generalization and insufficient clinical interpretability in the prior art. The method comprises the following steps: firstly, synchronously acquiring visual structure and electrical impedance bimodal data through an event-driven sensor array, and completing preprocessing through pulse coding, STDP rule optimization and quality verification; then, respectively extracting structured feature vectors of a tumor visual structure class and characteristic feature vectors of a numerical attribute class by a heterogeneous twin coding engine, and unifying dimensions and distribution through processes such as multi-scale dynamic perception and cross-modal calibration; then based on an attention mechanism, InfoNCE contrast loss and minority class weight gain, dynamic weighted fusion and semantic association enhancement are performed on the bimodal features, and a comprehensive feature vector is generated; and finally, through clinical logic adaptation and multi-center deviation correction, outputting a tumor benign and malignant identification result through a full-connection classifier, and synchronously generating a clinical interpretable report containing key features and weights. According to the method, the complementary advantages of bimodal information are effectively integrated, the problems of heterogeneous multi-center equipment, unbalanced samples and the like are solved, the tumor recognition accuracy and the early-stage tiny tumor detection rate are improved, the computing power consumption is reduced, edge medical equipment deployment is adapted, and the requirements for low misjudgment and traceability of clinical diagnosis are met.
Owner:SINONEEDLE INTELLIGENCE TECH CO LTD

Aerial small target image detection method based on multi-mode progressive fusion

The invention relates to the technical field of computer vision, and discloses a multi-mode progressive fusion aerial small target image detection method, which comprises the following steps of: acquiring an aerial small target image; inputting the aerial small target image into the improved cross-modal target detection model, and outputting the predicted position, category and confidence of the small target in the aerial small target image; the cross-modal target detection model comprises a backbone network, a neck network and a detection head, the backbone network comprises a visible light branch, an infrared branch and four dynamic enhancement modules, and the visible light branch and the infrared branch are finally fused by a multi-scale small target cross screening mechanism module; and the neck network replaces a C2k3 module in the YOLO11 network with a multi-scale small target cross screening mechanism module. According to the method, the cross-modal target detection model is constructed under the condition that dual-modal data are input at the same time, double-flow features are extracted and fused, and same-dimension feature enhancement and cross-dimension feature fusion at different stages are achieved.
Owner:CIVIL AVIATION FLIGHT UNIV OF CHINA

Vehicle flow prediction method and system based on bimodal optimization embedded learning model

The invention provides a vehicle flow prediction method and system based on a bimodal optimization embedded learning model, and the system comprises a data preprocessing module, a bimodal coding module, an attention alignment module and a sequence prediction module, is suitable for the field of intelligent transportation, and aims at achieving the prediction of the vehicle flow through multi-source data fusion and deep model architecture innovation. The problem that a traditional method is insufficient in prediction precision in a complex traffic scene is solved, and efficient decision support is provided for urban traffic management.
Owner:FUJIAN NORMAL UNIV

Transform and double-branch feature decoupling-based multi-modal medical image fusion method, system and equipment and medium

The invention belongs to the technical field of medical image processing, and discloses a multi-modal medical image fusion method, system and device based on Transform and double-branch feature decoupling and a medium, and the method comprises the steps: obtaining a multi-modal medical image, carrying out the normalization processing of the multi-modal medical image, and obtaining a preprocessed multi-modal medical image; inputting the preprocessed multi-modal medical image into a medical image fusion model for image fusion to obtain a fused image; wherein the medical image fusion model comprises a shallow feature extraction module, a dual-mode cross attention module, a dual-branch feature decoupling module and a splicing fusion module which are connected in sequence. According to the technical scheme, by comprehensively fusing the features of the bimodal medical image, the image fusion effect is remarkably improved, detail and structure information between different modals can be better captured and combined, and therefore more accurate and comprehensive support is provided for medical image analysis.
Owner:YANSHAN UNIV

Concrete grout performance automatic identification method and system based on multi-modal fusion

The invention relates to a multi-modal fusion-based concrete slurry performance automatic identification method and system, and belongs to the technical field of concrete detection. The invention aims to solve the defects of the traditional manual evaluation and the existing single-mode automatic detection technology, constructs a bimodal depth perception model fusing an image and sensing data, comprehensively analyzes visual features and physical parameters of slurry, preprocesses the image by using an illumination self-adaptive model, overcomes the influence of field light variation, and improves the detection accuracy. Then visual and sensing features are extracted through a double-flow network and fused, and finally objective and comprehensive slurry performance evaluation and risk grades are output through a performance index algorithm and a risk scoring formula. According to the method, the accuracy, the robustness and the intelligent level of concrete workability identification in a complex construction environment are remarkably improved, and reliable support is provided for engineering quality control.
Owner:云南华电金沙江中游水电开发有限公司

Multi-modal target detection model training method and device, equipment and medium

The invention relates to the technical field of target detection, in particular to a multi-modal target detection model training method and device, equipment and a medium. The reliability of the parameters of the bimodal feature extraction units is dynamically evaluated by quantifying the validity (contribution degree to the model) of the parameters and the sensitivity (optimization degree of the parameters) of the gradients, so that the gradient conflict between the bimodal feature extraction units is accurately identified and corrected, and the model optimization process is ensured to be always guided by more reliable modal information. According to the method, the semantic conflict problem in multi-modal data is remarkably relieved, the consistency learning of cross-modal features is promoted, the detection precision is improved, and the adaptive capacity of the model in complex scenes (such as illumination variation and shielding) is enhanced.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)