Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

232 results about "Bi modal" patented technology

Retrieval enhancement method based on multi-modal data fusion and modal perception

The invention relates to the technical field of information retrieval and generation, in particular to a retrieval enhancement method based on multi-modal data fusion and modal perception. According to the method, firstly, a dual-channel architecture is adopted to perform feature extraction and coding on a text and an image respectively, and mutually independent embedded representation spaces are constructed, so that high-quality collaboration and matching of cross-modal representation are realized; and a pseudo-pairing generation mechanism is introduced to effectively mine and reconstruct the existing non-paired data in the knowledge base. And designing a query modal perception and dynamic weighting mechanism for accurately controlling the fusion proportion of the image-text bimodal information in the retrieval stage so as to match the modal demand difference of different query contents. And further executing aggregation retrieval and reordering of the cross-modal information by using dynamic weighted fusion retrieval to generate a candidate set of multi-modal responses. According to the method, accurate matching and dynamic weight adjustment of the image-text content are realized, and the accuracy and expression integrity of the generated content are improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Multi-mode ultrasonic fusion pressure vessel welding seam defect nondestructive testing method and multi-mode ultrasonic fusion pressure vessel welding seam defect nondestructive testing system

The invention provides a multi-mode ultrasonic fusion pressure vessel weld defect nondestructive testing method and system, and relates to the technical field of nondestructive testing. According to the method, geometric parameters of a welding seam are obtained through three-dimensional laser scanning, and an optimal scanning parameter set is generated; driving ultrasonic phased array equipment to scan for one time and synchronously acquire shear wave full-matrix capture and longitudinal wave linear scanning data; performing energy flow angular spectrum analysis and envelope analysis on the bimodal data, extracting defect feature parameters and constructing a three-dimensional feature tensor; carrying out multi-dimensional feature fusion by adopting Tucker decomposition, and enhancing a core tensor through physical modeling; generating three types of defect indication diagrams including a defect existence possibility diagram, a defect relative scale diagram and a defect space orientation diagram from the enhanced feature tensor; and the three types of indication diagrams are visually presented for comprehensive interpretation of detection personnel. Through multi-modal data fusion and physical modeling enhancement, the defect identification accuracy and detection efficiency are remarkably improved, the false alarm rate is reduced, and reliable technical support is provided for pressure vessel welding seam safety detection.
Owner:YUNNAN SPECIAL EQUIP SAFETY TESTING RES INST

Multi-modal emotion recognition method and system for service-oriented robot

The invention belongs to the technical field of artificial intelligence, and particularly relates to a service-oriented robot-oriented multi-modal emotion recognition method and system, and the method comprises the steps: collecting audio and video stream data of emotion changes of a user, and separating visual and voice data; extracting visual and voice emotion features through a pre-training model, and calculating prediction probability distribution of each mode; constructing a bimodal confidence quantitative model based on the distribution to obtain each modal confidence; and fusing the features by adopting a sectional type dynamic weight distribution strategy so as to identify the emotional state of the user. Visual and voice modes are fused, feature alignment is realized in combination with dynamic time warping, spatial optimization performance is shared and expressed through a confidence model, a dynamic weight strategy and a cross-modal time sequence cooperation module, and the method has high recognition accuracy, high robustness and real-time processing capacity in a complex environment and is suitable for various service scenes.
Owner:SUZHOU CITY UNIV

Intelligent cabin man-machine cooperative control method and device, vehicle, medium and product

The invention relates to the technical field of vehicles, in particular to an intelligent cabin man-machine cooperative control method and device, a vehicle, a medium and a product. Performing fusion processing on the voice signal data, the visual signal data and the tactile signal data of the current driver by using a preset dynamic weight distribution model to obtain a multi-modal fusion control instruction; and under the condition that the multi-mode fusion control instruction meets a preset sight-voice dual-mode condition, synchronously issuing the multi-mode fusion control instruction to a plurality of target execution devices in the intelligent cabin so as to drive each target execution device to perform cooperative control. Therefore, by constructing a cooperative control mechanism of multi-mode sensing and dynamic weight distribution, the problems that the false wake-up rate is high, the voice and graphical interface cooperative efficiency is low and the interaction mode cannot adapt to a complex driving environment due to a single or separated interaction mode in the related technology are effectively solved.
Owner:CHERY AUTOMOBILE CO LTD

Art and craft material detection method based on multi-modal deep learning

The invention discloses an industrial art material detection method based on multi-modal deep learning, particularly relates to the field of material analysis, and is used for solving the problems that cross-modal data alignment of curved surface utensils in highlight and multi-layer coating scenes is difficult, and material boundaries are easy to drift along with shooting batches and visual directions. The method comprises the following steps of: positioning a hyperspectral line scanning fragment; constructing a local reference grid to realize initial pairing; generating pixel-level residual image quantization dislocation; estimating luminosity mapping to form consistent bimodal data pairs after eliminating a high-reflectivity region; iteratively supplementing anchor points or adjusting grid rigidity, generating a stable material graph, updating an index table and a consistency record, ensuring cross-batch reusability and traceability, and improving the material detection precision.
Owner:WUXI GONGCHUN ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Infrared and visible light image fusion method based on text semantic consistency guidance

The invention provides an infrared and visible light image fusion method based on text semantic consistency guidance, which relates to the technical field of multi-modal image fusion, and comprises the following steps: respectively carrying out fine-grained text semantic generation on infrared and visible light images and mapping the infrared and visible light images to a unified embedding space; bidirectional compensation and enhancement are carried out on text semantics through a cross-modal attention mechanism, and unified text semantic priori is constructed; a structure-intensity decoupling double-branch encoder is adopted, and structure texture features of visible light and intensity significant features of infrared light are extracted respectively; under the prior guidance of text semantics, the bimodal visual features are aligned in a shared semantic space through explicit semantic consistency constraint and implicit semantic distribution consistency constraint; and finally, taking the text semantic priori as a global modulation signal, and carrying out adaptive weighted fusion and decoding on the aligned features to generate a fused image. The problems that an existing method is insufficient in semantic modeling and poor in fusion result consistency are effectively solved.
Owner:XIAMEN UNIV OF TECH

Automatic material taking mechanism for cable processing equipment

The invention relates to the technical field of cable processing, in particular to an automatic material taking mechanism for cable processing equipment, and the mechanism comprises an incoming material arraying and conveying unit which is used for flattening scattered cables and conveying the cables through a conveying belt synchronized with an encoder; the industrial vision or AI vision unit is used for performing vision detection and vision measurement on the cable in the conveying process; the material taking execution unit comprises an execution mechanism with the X-Y-theta degree of freedom and the Z degree of freedom and a bimodal flexible gripper; the posture correcting and guiding unit is used for correcting the direction of the end of the grabbed cable and standardizing the feeding posture of the grabbed cable; according to the automatic material taking mechanism for the cable processing equipment, the supplied material arraying and conveying unit is arranged to be combined with the industrial vision or AI vision unit, and the supplied material arraying and conveying unit is arranged to be combined with the industrial vision or AI vision unit; the defects that vibration stop photographing or a general taking and placing system is low in flexible cable separation efficiency and prone to missing detection are effectively overcome.
Owner:NUO XUN (JIANGSU) CABLE TECH CO LTD

Driving authority management method and device based on biological characteristic verification and medium

The invention provides a driving authority management method and device based on biological feature verification and a medium, and belongs to the technical field of vehicles. The method comprises the steps of collecting facial features, voiceprint data and driving behavior data by detecting a starting operation of a current driver; the identity of the current driver is recognized by using a multi-modal biological recognition algorithm, and the accuracy and anti-counterfeiting capability of identity verification are improved by adopting a bimodal fusion scheme of face recognition and voiceprint recognition; when it is determined that identity recognition of the current driver succeeds, a driving account is determined to judge whether the driver has the use permission or not, a binding relation of driver identity-account-function permission is established, and accurate matching of the intelligent driving permission and the driver qualification is achieved; the open state of the intelligent driving function is dynamically controlled based on a permission verification result, if the permission exists, the intelligent driving function is automatically prepared, if the permission does not exist, the function is locked, a clear prompt is given, and the safety risk that the intelligent driving function is used before being learned is avoided from the source.
Owner:CHINA FAW CO LTD

Prompt guidance and multi-modal fusion-based class incremental learning method

The invention provides a class incremental learning method based on prompt guidance and multi-modal fusion, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: firstly, performing semantic extension on a category label, and constructing semantic enhanced text representation through a text encoder; then block embedding and hierarchical feature extraction are carried out on the input image by using a pre-trained visual encoder, a cross-modal unified embedding space is constructed, a bimodal prompt gating fusion module is introduced into the unified embedding space, and adaptive weighting is carried out on text prompt and image prompt according to gating weight to generate fusion prompt; through a bimodal prompt collaborative filtering module, screening out a prompt set most relevant to the current task according to the similarity of the semantic features of the image and the text; the pre-training backbone network is frozen in the increment stage, only prompt parameters and fusion layer weights are optimized, a joint loss function is used for parameter updating, finally, image and text data are input in the reasoning stage, cross-modal similarity is calculated, and a classification prediction result is output.
Owner:NORTHEASTERN UNIV CHINA

Vending machine multi-mode emotion perception and personalized dialogue generation method and system

The invention discloses a vending machine multi-mode emotion perception and personalized dialogue generation method and system, and the method comprises the steps: collecting a face image and a voice signal of a user through a camera and a microphone, and outputting a dual-mode emotion tag through the analysis of a face expression and a voice emotion; fusing the label and the dialogue text to generate a structured Prompt containing emotion, intention and context; based on an Agent framework multi-modal large model, combining context history and a personalized strategy to generate a reply text matched with the emotional scene; according to the text and the comprehensive emotion label, TTS parameters are dynamically adjusted, and anthropomorphic voice is output; in multiple rounds of dialogues, styles are kept consistent through an emotion smoothing formula, and emotion-verbal skill-conversion data are precipitated in combination with a user feedback optimization strategy. The interaction limitation of a traditional vending machine is broken through, multi-mode emotion perception, personalized personification interaction and continuous evolution are achieved, the vending machine is promoted to be upgraded to an emotional shopping guide platform, user experience and sales transformation are improved, and a technical model is provided for intelligent retail.
Owner:SHANGHAI QUZHI NETWORK TECH CO LTD

High-precision face recognition method and system

The invention discloses a high-precision face recognition method and system which are suitable for complex illumination and unconstrained position scenes. The method is realized through cooperation of a multi-modal acquisition guide strategy driven by a state machine and a self-adaptive fusion mechanism based on a physical multiplier array: the state machine controls logic to generate a position-attitude sequence instruction, and drives a user to synchronously provide RGB-NIR image pairs in upper, middle and lower screen areas and left and right side face attitudes; the double-flow convolutional neural network extracts an aligned bimodal feature map through shared weight constraint; the attention fusion module dynamically generates a space-channel joint attention weight based on a local signal-to-noise ratio difference through a multiplier array and a weight register realized by hardware, and performs pixel-by-pixel weighted fusion on the bimodal features. And a physical mapping relation between state machine position change and modal weight distribution is established through coupling training. According to the invention, through algorithm hardening and hardware cooperation, the accuracy and real-time performance of cross-illumination and cross-position unconstrained face recognition are significantly improved.
Owner:BEIJING AUTO SMART INFORMATION TECH CO LTD

Method and system for tracking and detecting sheltered target in electric power scene based on Leiyu-vision integrated multi-mode space-time adversarial learning

The invention discloses a method and a system for tracking and detecting an occluded target in an electric power scene based on thunder-vision integrated multi-mode spatio-temporal adversarial learning, and relates to the field of intelligent inspection of electric power equipment, and the method comprises the steps: inputting a structured point cloud matrix and a visual image into a double-branch feature extraction model, respectively extracting geometric features and semantic features by using a double-branch feature extraction model; inputting the bimodal fusion features into a dynamic weight adaptive fusion mechanism based on an adversarial training strategy to obtain a fusion weight coefficient of the target in a short-time shielding state in the power scene; and a double-buffer memory pool is constructed, and a three-level progressive trajectory prediction compensation mechanism is combined to compensate a target trajectory position in a long-time shielding state in a power scene. According to the method, the shielding state quantitative model is constructed, the fusion weight distribution proportion can be effectively and dynamically adjusted in real time according to the visual sensor, and when it is detected that the target equipment enters the shielding state, the visual data weight is automatically reduced to the safety threshold value.
Owner:NANJING STEIN SMART ENERGY TECH CO LTD

Skin lesion auxiliary detection method and system based on multi-modal information fusion

The invention provides a skin lesion auxiliary detection method and system based on multi-modal information fusion, and belongs to the technical field of computer vision and application. The method comprises the following steps: firstly, acquiring and constructing bimodal data including a dermatoscope image and a text of a patient, and preprocessing the data; secondly, respectively adopting an improved convolutional neural network and a text coding model, and extracting deep and high-dimensional features from the skin disease image and the self-described text of the patient; afterwards, an improved symmetric gating cross attention fusion module is constructed, fine-grained interaction and alignment between visual and text features are realized through a bidirectional cross attention mechanism, and a gating unit is utilized to adaptively adjust a fusion weight; and finally, inputting the deeply fused multi-modal feature vector into a classifier to obtain a confidence score of each skin lesion type so as to provide reference for dermatologists. According to the method, the diagnosis accuracy, robustness and interpretability are synergistically improved, and the method has great clinical application value.
Owner:SHENYANG SEVENTH PEOPLES HOSPITAL

Cross-modal attention collaborative jail break attack method

The invention discloses a cross-modal attention collaborative jail break attack method, and belongs to the technical field of artificial intelligence security. The method comprises the following steps of: constructing an input sequence representation containing a system prompt, an adversarial image representation, a malicious query and an adversarial text suffix according to a causal self-attention mechanism; inputting the sequence into a visual language model to execute forward propagation; based on the designed attention-oriented loss collaborative function, optimizing adversarial image representation through a joint gradient optimization algorithm and updating an adversarial text suffix to optimize an attack target; iteratively circulating until convergence, and outputting the optimized unified multi-modal knowledge; and finally, utilizing the knowledge to construct an attack sequence to realize jailbreak. According to the method, accurate control on an internal attention mechanism of the visual language model is realized for the first time, and through visual-text dual-mode cooperative attack, the attack success rate is remarkably improved while high concealment is kept, and the important driving force for promoting the progress of a safe alignment technology is achieved.
Owner:NAT UNIV OF DEFENSE TECH

Tumor identification method based on multi-modal feature fusion

The invention provides a tumor recognition method based on multi-modal feature fusion, relates to the technical field of tumor recognition, and is used for solving the problems of missing multi-modal feature calibration, insufficient semantic association, poor generalization and insufficient clinical interpretability in the prior art. The method comprises the following steps: firstly, synchronously acquiring visual structure and electrical impedance bimodal data through an event-driven sensor array, and completing preprocessing through pulse coding, STDP rule optimization and quality verification; then, respectively extracting structured feature vectors of a tumor visual structure class and characteristic feature vectors of a numerical attribute class by a heterogeneous twin coding engine, and unifying dimensions and distribution through processes such as multi-scale dynamic perception and cross-modal calibration; then based on an attention mechanism, InfoNCE contrast loss and minority class weight gain, dynamic weighted fusion and semantic association enhancement are performed on the bimodal features, and a comprehensive feature vector is generated; and finally, through clinical logic adaptation and multi-center deviation correction, outputting a tumor benign and malignant identification result through a full-connection classifier, and synchronously generating a clinical interpretable report containing key features and weights. According to the method, the complementary advantages of bimodal information are effectively integrated, the problems of heterogeneous multi-center equipment, unbalanced samples and the like are solved, the tumor recognition accuracy and the early-stage tiny tumor detection rate are improved, the computing power consumption is reduced, edge medical equipment deployment is adapted, and the requirements for low misjudgment and traceability of clinical diagnosis are met.
Owner:SINONEEDLE INTELLIGENCE TECH CO LTD

Aerial small target image detection method based on multi-mode progressive fusion

The invention relates to the technical field of computer vision, and discloses a multi-mode progressive fusion aerial small target image detection method, which comprises the following steps of: acquiring an aerial small target image; inputting the aerial small target image into the improved cross-modal target detection model, and outputting the predicted position, category and confidence of the small target in the aerial small target image; the cross-modal target detection model comprises a backbone network, a neck network and a detection head, the backbone network comprises a visible light branch, an infrared branch and four dynamic enhancement modules, and the visible light branch and the infrared branch are finally fused by a multi-scale small target cross screening mechanism module; and the neck network replaces a C2k3 module in the YOLO11 network with a multi-scale small target cross screening mechanism module. According to the method, the cross-modal target detection model is constructed under the condition that dual-modal data are input at the same time, double-flow features are extracted and fused, and same-dimension feature enhancement and cross-dimension feature fusion at different stages are achieved.
Owner:CIVIL AVIATION FLIGHT UNIV OF CHINA

Concrete grout performance automatic identification method and system based on multi-modal fusion

The invention relates to a multi-modal fusion-based concrete slurry performance automatic identification method and system, and belongs to the technical field of concrete detection. The invention aims to solve the defects of the traditional manual evaluation and the existing single-mode automatic detection technology, constructs a bimodal depth perception model fusing an image and sensing data, comprehensively analyzes visual features and physical parameters of slurry, preprocesses the image by using an illumination self-adaptive model, overcomes the influence of field light variation, and improves the detection accuracy. Then visual and sensing features are extracted through a double-flow network and fused, and finally objective and comprehensive slurry performance evaluation and risk grades are output through a performance index algorithm and a risk scoring formula. According to the method, the accuracy, the robustness and the intelligent level of concrete workability identification in a complex construction environment are remarkably improved, and reliable support is provided for engineering quality control.
Owner:云南华电金沙江中游水电开发有限公司

Multi-modal target detection model training method and device, equipment and medium

The invention relates to the technical field of target detection, in particular to a multi-modal target detection model training method and device, equipment and a medium. The reliability of the parameters of the bimodal feature extraction units is dynamically evaluated by quantifying the validity (contribution degree to the model) of the parameters and the sensitivity (optimization degree of the parameters) of the gradients, so that the gradient conflict between the bimodal feature extraction units is accurately identified and corrected, and the model optimization process is ensured to be always guided by more reliable modal information. According to the method, the semantic conflict problem in multi-modal data is remarkably relieved, the consistency learning of cross-modal features is promoted, the detection precision is improved, and the adaptive capacity of the model in complex scenes (such as illumination variation and shielding) is enhanced.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Multi-modal network dynamic loading and management system and method based on modal virtualization

The invention relates to a multi-modal network dynamic loading and management system and method based on modal virtualization, and belongs to the technical field of network system structures and virtualization. By constructing a modal virtualization layer and a modal description language analysis mechanism, the limitation that traditional network virtualization is only oriented to a single protocol stack is broken through, unified packaging and isolation of multiple network modalities such as IP, TSN, ICN and custom protocols are achieved, the dual-modal state snapshot and traffic mirroring technology is utilized, and the network modality is improved. The problem of service interruption in the online switching process of the network modes is solved; and further in combination with a vectorized heterogeneous resource elastic scheduling model, global optimal configuration and on-demand mapping of the CPU, the GPU, the FPGA and the P4 programmable switching chip are realized.
Owner:SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN

Improved YOLOv8-based dual-mode substation equipment overheating defect detection method

The invention relates to the technical field of power equipment fault detection, in particular to a bimodal substation equipment overheating defect detection method based on improved YOLOv8, which comprises the following steps: acquiring a visible light (RGB) image and an infrared (IR) image of equipment with overheating defects in a substation, constructing a bimodal target detection framework, and detecting the overheating defects of the equipment on the basis of YOLOv8n. A visible light (RGB) image and an infrared (IR) image are processed through two independent branches respectively, decoupling and fusion of modal information are achieved, an ACEFusion module is introduced to achieve efficient guide fusion of modal features, an RIAAdd module is integrated to enhance the expression ability of multi-scale semantic features, meanwhile, the complexity of the model is controlled, an improved YOLOv8 bimodal detection model is trained, and thermal defect detection is executed. By designing a bimodal feature extraction and fusion architecture and a lightweight enhancement module, efficient collaboration of infrared and visible light information is realized, the calculation cost is relatively low, the thermal defect detection precision and real-time performance are improved, real thermal defects and environmental interference can be effectively distinguished, and the method is suitable for actual engineering deployment scenes.
Owner:TONGHUA POWER SUPPLY COMPANY STATE GRID JILIN ELECTRIC POWER +1

Multi-modal image fusion method and system, medium and program product

The invention relates to the technical field of computer vision, in particular to a multi-modal image fusion method and system, a medium and a program product, and the method comprises the steps: firstly obtaining a bimodal image data set with consistent time and space, the bimodal image data set comprising a visible light image data set and an infrared image data set; then constructing an image fusion network by using the bimodal image data set, and constraining the image fusion network by using source image consistency loss, infrared target structure loss and visible light texture loss; performing target detection on the fused image output by the image fusion network by using a target detection network; carrying out the cooperative training of the target detection network and the image fusion network through the target detection loss, and obtaining the trained image fusion network and target detection network; and finally, low-altitude weak and small target detection is carried out by using the trained image fusion network and the target detection network. According to the invention, the detection precision and efficiency of the low-altitude weak and small target are obviously improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Power data dynamic verification method based on large model

The invention relates to the technical field of data processing, in particular to an electric power data dynamic verification method based on a large model. Obtaining to-be-verified power multi-modal data, and carrying out fusion understanding on the power multi-modal data through a pre-trained multi-modal large model to obtain multi-dimensional unified semantic representation; constructing a bimodal rule knowledge base containing a prior rule mode and a historical case mode, and matching an applicable verification rule set from the bimodal rule knowledge base based on multi-dimensional unified semantic representation; executing first significance verification on the multi-dimensional unified semantic representation based on the applicable verification rule set, identifying and marking dominant data defects, and generating a first significance verification result; on the basis of the first significance verification result and a pre-constructed multi-modal association map, executing second significance verification to obtain a second significance verification result; according to the method, the power data quality and the reliability of data application can be greatly improved.
Owner:STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO

Digital mammary gland artificial intelligence auxiliary diagnosis system based on multi-modal fusion

The invention relates to the technical field of disease auxiliary diagnosis, in particular to a digital mammary gland artificial intelligence auxiliary diagnosis system based on multi-modal fusion, and the system comprises a multi-modal data collection module which is used for obtaining digital mammary gland image data and rehabilitation scheme data of a patient; the image quality correction module is used for carrying out acquisition quality evaluation and correction on the digital mammary gland image data; the feature modeling and evaluation module is used for performing feature extraction on the corrected digital mammary gland image data and rehabilitation scheme data, and comprises an image-drug action mechanism cooperation unit and an image-rehabilitation bimodal co-learning unit; and the rehabilitation effect prediction module is used for constructing a rehabilitation effect prediction model and outputting a rehabilitation prediction result under the influence of the scheme specificity of the current rehabilitation scheme of the patient. According to the method, the digital mammary gland image and the patient rehabilitation scheme are deeply fused, so that the accuracy and reliability of mammary gland tumor rehabilitation prediction are effectively improved.
Owner:MEITIAN HUAYING MEDICAL MANAGEMENT (SHANGHAI) CO LTD

Underwater salient target detection method and system based on double-flow fusion network

The invention discloses an underwater salient target detection method and system based on a double-flow fusion network, and belongs to the technical field of computer vision. The method comprises the following steps: respectively extracting multi-scale features of an RGB image and a depth image through a double-flow encoder; in the shallow layer, fusing and enhancing the edge and detail information of the bimodal features through an edge fusion module; in a deep layer, content-adaptive cross-modal semantic fusion is realized in a frequency domain through a dynamic filtering module; fusing the multi-scale features through a cross-layer aggregation decoder to generate a rough saliency map; extracting detail features from the original RGB image through a global detail purification network; and finally, fusing the rough saliency map and the detail features, and outputting an underwater saliency target prediction map. The objective of the invention is to improve the precision and boundary definition of salient target detection in an underwater complex scene.
Owner:NANKAI UNIV

YOLOv8-based rail transit construction site multi-category engineering vehicle identification device, method, equipment and medium

The invention discloses a YOLOv8-based rail transit construction site multi-category engineering vehicle recognition device, method and equipment and a medium, and the device comprises an image obtaining module which is configured to obtain a visible light image and an infrared image of a rail transit construction site; the multi-modal feature enhancement module is configured to respectively extract texture structure features of the visible light image and temperature distribution features of the infrared image, align dual-modal features by using a spatial transformation network, and realize dual-modal feature weighted fusion through a dual-channel attention mechanism; the component-level dynamic attention mechanism module is configured to perform component-level enhancement and optimization on the bimodal features after weight fusion and output an enhanced fusion feature map; and the time sequence correlation loss function module is configured to realize target classification and bounding box prediction, output a target position distribution prediction result after time sequence correlation optimization through a target motion model in combination with a continuous frame detection result, and complete engineering vehicle identification of the rail transit construction site.
Owner:CRSC COMM & INFORMATION GRP CO LTD

Motion quality evaluation method based on video and skeleton bimodal fusion

The invention discloses an action quality evaluation method based on video and skeleton bimodal fusion, and the method comprises the steps: inputting a video frame sequence containing a complete action process into a fragmentation network, and generating continuous fragments; secondly, inputting each video clip into a video modal feature extraction network and a human body skeleton feature extraction network, respectively obtaining a video modal feature and a skeleton modal feature, fusing the video modal feature and the skeleton modal feature through a cross-modal attention mechanism, and generating a fused clip-level feature representation; and finally, inputting each segment-level feature representation into a multi-layer Transform encoder, obtaining a global time sequence feature of the action through a self-attention mechanism, inputting the global time sequence feature into a multi-layer perceptron (MLP) regression network, and generating an action quality score. According to the method, a prediction result is ensured to be more stable in numerical value and more accord with human perception in semantics, and the accuracy and consistency of action quality evaluation are remarkably improved.
Owner:HANGZHOU DIANZI UNIV

Digital human interaction method and system, storage medium and program product

The invention relates to the technical field of artificial intelligence, and discloses a digital human interaction method and system, a storage medium and a program product, which can associate a user with an interaction process by adding a target session identifier to a first voice request, ensure context coherence of multiple rounds of conversations, and realize full duplex interaction. Furthermore, the voice signal is converted into the processable target text, the command task in the target text is identified, and the first command task text is generated, so that the intention of the user can be accurately identified, the task type can be distinguished, and the semantic understanding efficiency is improved. Furthermore, through parallel processing of text sentence segmentation and speech synthesis, the mechanical feeling of single word output is avoided, the waiting time of the user is shortened, and the interaction naturalness is improved. Furthermore, the target command task voice and the second command task text are sent to the client, so that bimodal synchronous output is realized, the interaction naturalness is improved, multi-scene adaptation is realized, and the user experience is improved.
Owner:SHANGHAI INVESTIGATION DESIGN & RES INST CO LTD

Audio quality analysis method and system based on algorithm

The invention belongs to the technical field of audio quality analysis, and discloses an algorithm-based audio quality analysis method and system, and the method comprises the steps: employing a lightweight CRNN model to fuse audio-video dual-mode features through the cooperation of a multi-problem joint detection module and a self-adaptive feature fusion extraction module, and carrying out the algorithm-based audio quality analysis. The five problems of abnormal volume, noise and the like are synchronously detected in combination with a multi-task learning framework, and a bottom-layer feature network is shared to avoid parameter conflicts; meanwhile, weighted binary cross entropy loss is introduced, and the detection rate of low-probability problems such as howling and sound interruption is enhanced; compared with a traditional sub-module scheme, the design eliminates the contradiction that noise suppression excessively weakens voice, the multi-problem detection accuracy is remarkably improved, the result consistency is greatly improved, and the accurate detection requirement for concurrency of multiple problems in a conference scene is met; and full-link adaptation is realized through a dynamic parameter adjustment mechanism: the difference perception feature learning expands the normal / abnormal frame feature difference, and the sensitivity of a low signal-to-noise ratio scene is improved.
Owner:NANJING SHUZHI ENTROPY TECHNOLOGY CO LTD

Bidding document checking method and system based on multi-modal large model

The invention relates to the field of data processing, in particular to a bidding document checking method and system based on a multi-modal large model, and the method comprises the steps: obtaining bidding documents, preprocessing the bidding documents to obtain a target word sequence of each paragraph, selecting any two bidding documents as a to-be-compared document pair, obtaining word similarity through the product of the overlapping rate and the position alignment rate of words, and comparing the word similarity with the target word sequence; performing weighted calculation on the similarity of the corresponding words based on the editing distance to obtain weighted similarities, performing matching to obtain a plurality of groups of word position pairs, and taking a mean value of all the weighted similarities as text similarity; extracting image data, calculating approximation degrees and matching, and taking a mean value of all the approximation degrees as image similarity; and fusing the bimodal similarity through a preset weight to obtain a risk value, and completing bidding and tendering document inspection. According to the method, through overlapping rate and alignment rate two-dimensional word similarity evaluation, editing distance wrongly written character correction and text-image multi-modal fusion calculation, the probability of missing and misjudgment is reduced.
Owner:NINGBO GUOYAN INFORMATION TECHNOLOGY CO LTD

Multi-modal large model fine tuning method and device, terminal equipment and storage medium

The invention is suitable for the technical field of artificial intelligence, and provides a multi-modal large model fine tuning method and device, terminal equipment and a storage medium, and the method comprises the steps: firstly obtaining single-modal training data; according to the single-mode training data, carrying out single-mode training fine tuning on the TIS-MOME, and obtaining a first target TIS-MOME; then bimodal training data is obtained; according to the bimodal training data, carrying out bimodal training fine tuning on the first target TIS-MOME, and obtaining a second target TIS-MOME; and finally, obtaining multi-modal training data. And performing multi-modal training fine tuning on the second target TIS-MOME according to the multi-modal training data to obtain a third target TIS-MOME and outputting the third target TIS-MOME. According to the embodiment of the invention, through training fine tuning on the multi-modal large model TIS-MOME, the adaptive capacity of the multi-modal large model TIS-MOME can be effectively improved, and the intelligent and precise development of traditional Chinese and western medicine combined diagnosis and treatment is facilitated.
Owner:CHINESE MEDICINE GUANGDONG LABORATORY