A method and apparatus for AI-powered multimodal image semantic segmentation and target detection

CN122574368APending Publication Date: 2026-08-14XIANGTAN NUOXING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-16
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]本发明的目的在于提供一种Ai智能多模态图像语义分割与目标检测方法及装置,以解决上述背景技术中提出的传统视觉感知技术大多基于单模态图像数据开展训练与推理,仅依靠可见光图像完成检测与分割任务,在复杂工况环境下,如强光逆光、弱光夜景、雨雪雾恶劣天气、遮挡阴影、低对比度场景中,单模态可见光图像存在信息缺失、特征模糊、干扰噪声大的问题,极易导致目标漏检、误检、分割边缘模糊、语义分类错误等情况的问题

Benefits of technology

1、通过设置多模态图像采集模块,多模态图像融合感知架构,融合可见光、红外、深度多维度图像信息,弥补了单模态图像在恶劣环境、复杂场景下的信息缺陷,大幅提升目标检测与语义分割的抗干扰能力,有效降低漏检、误检概率,适配全天候、全场景视觉感知需求;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574368A_ABST
    Figure CN122574368A_ABST
Patent Text Reader

Abstract

This invention discloses an AI-powered multimodal image semantic segmentation and target detection method and apparatus, including a lifting mechanism. The lifting mechanism comprises a fixed base and a lifting cylinder. The middle of the top of the fixed base is fixedly connected to the fixed end of the lifting cylinder. A lifting platform is fixedly installed on the movable end of the lifting cylinder. A turntable is provided at the top of the lifting platform. A first displacement mechanism is fixedly installed at the top of the turntable. A second displacement mechanism is fixedly installed at the top of the first displacement mechanism. A mounting frame is fixedly installed at the top of the second displacement mechanism. This invention, by setting up a multimodal image acquisition module and a multimodal image fusion perception architecture, integrates visible light, infrared, and depth multi-dimensional image information, compensating for the information defects of single-modal images in harsh environments and complex scenes. It significantly improves the anti-interference capability of target detection and semantic segmentation, effectively reduces the probability of missed detection and false detection, and adapts to all-weather, all-scene visual perception needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, specifically to an AI-powered multimodal image semantic segmentation and target detection method and apparatus. Background Technology

[0002] Computer vision is an important branch of artificial intelligence. Object detection and semantic segmentation are two core tasks. Object detection not only identifies the category of objects in an image, but also determines their location, usually represented by bounding boxes. Semantic segmentation classifies each pixel in an image to achieve pixel-level understanding. These two tasks are of great significance in real-world applications. For example, autonomous vehicles need to detect pedestrians and vehicles on the road, and medical image analysis needs to accurately segment organs and lesion areas.

[0003] However, traditional image semantic segmentation and object detection have the following drawbacks: Traditional visual perception technologies are mostly based on single-modal image data for training and inference, relying solely on visible light images to complete detection and segmentation tasks. In complex working environments, such as strong light and backlight, low light and night scenes, rain, snow and fog, occlusion and shadows, and low contrast scenes, single-modal visible light images suffer from information loss, feature blurring, and large interference noise, which can easily lead to missed detection, false detection, blurred segmentation edges, and semantic classification errors. Summary of the Invention

[0004] The purpose of this invention is to provide an AI-powered intelligent multimodal image semantic segmentation and target detection method and apparatus to address the problems mentioned in the background art. Traditional visual perception technologies are mostly based on single-modal image data for training and inference, relying solely on visible light images to complete detection and segmentation tasks. In complex working environments, such as strong light and backlight, low light night scenes, rain, snow, fog, severe weather, occlusion and shadows, and low contrast scenes, single-modal visible light images suffer from information loss, blurred features, and high interference noise, which can easily lead to target omissions, false detections, blurred segmentation edges, and semantic classification errors.

[0005] To achieve the above objectives, the present invention provides the following technical solution: an AI intelligent multimodal image semantic segmentation and target detection device, including a lifting mechanism; The lifting mechanism includes a fixed base and a lifting cylinder. The middle part of the top of the fixed base is fixedly connected to the fixed end of the lifting cylinder. A lifting platform is fixedly installed on the movable end of the lifting cylinder. A turntable is provided at the top of the lifting platform. A first displacement mechanism is fixedly installed at the top of the turntable. A second displacement mechanism is fixedly installed at the top of the first displacement mechanism. A mounting frame is fixedly installed at the top of the second displacement mechanism. A visual analysis table is fixedly installed on one side of the mounting frame. A visual analysis probe is fixedly installed on one side of the visual analysis table. The visual analysis probe is equipped with a multimodal image acquisition module, an image preprocessing module, a cross-modal feature processing module, a multi-scale feature fusion encoding module, a dual-task collaborative reasoning module, a result optimization output module, and an intelligent deployment module.

[0006] As a further technical solution of the present invention, the image preprocessing module is electrically connected to the multimodal image acquisition module, the cross-modal feature processing module is connected to the image preprocessing module, the multi-scale feature fusion encoding module is connected to the cross-modal feature processing module, the result optimization output module is connected to the dual-task collaborative reasoning module, and the multimodal image acquisition module is connected to a visible light camera, an infrared imaging module, and a depth camera respectively, for synchronously acquiring visible light images, infrared images, and depth images of the scene, realizing real-time acquisition and synchronous transmission of multi-dimensional image data, and ensuring the temporal consistency of multimodal data; The image preprocessing module has built-in image denoising unit, size normalization unit, pixel registration unit, and brightness correction unit, which are used to complete the denoising, size unification, accurate pixel registration and image quality optimization of multimodal images, and eliminate data differences and environmental interference between modalities. The cross-modal feature processing module includes a multi-branch feature extraction unit, an attention weight allocation unit, and a feature alignment enhancement unit. It is used to extract core features of each modality in a differentiated manner, adaptively allocate modality weights, complete accurate fusion and enhancement of cross-modal features, filter out invalid noise features, and strengthen key target feature information. The multi-scale feature fusion encoding module has a built-in feature pyramid network structure, which is used to perform multi-scale sampling and bidirectional fusion of fused features, aggregate shallow detailed features and deep global semantic features, and generate a high-dimensional and robust global fusion feature map. The dual-task collaborative reasoning module connects the target detection reasoning unit and the semantic segmentation reasoning unit respectively. The two units share the front-end encoded features and simultaneously complete target localization and classification and pixel-level semantic segmentation reasoning, realizing parallel and efficient operation of the two tasks, avoiding redundant calculations and improving reasoning efficiency. The result optimization output module has a built-in post-processing optimization algorithm to remove false detections and duplicate detections, optimize segmentation edge accuracy, output standardized and high-precision detection and segmentation results, and support data visualization and data storage. The intelligent deployment module is used to perform lightweight compression, quantization pruning, and operator optimization of the model. It supports the deployment of the model on cloud servers, embedded terminals, mobile devices, and industrial control equipment, adapting to the real-time inference needs of different scenarios.

[0007] As a further technical solution of the present invention, the first displacement mechanism includes a first displacement frame and a first stepper motor. A first lead screw is rotatably connected inside the first displacement frame. A first displacement block is threadedly connected to the middle of the first lead screw and slidably connected to the first displacement frame. One side of the first displacement frame is fixedly connected to one side of the first stepper motor. The output end of the first stepper motor is fixedly connected to the end of the first lead screw that is directly opposite to it. The first stepper motor drives the first lead screw to rotate. The thread on the surface of the first lead screw matches the thread on the inner wall of the first displacement block. The first displacement block is limited by the first displacement frame, which matches its shape and size. Therefore, the first displacement block slides along the first lead screw to make the first adjustment to the position of the visual analysis probe.

[0008] As a further technical solution of the present invention, the top end of the first displacement block is fixedly connected to the second displacement mechanism, the bottom end of the first displacement frame is fixedly connected to the turntable, and the first displacement mechanism is installed between the turntable and the second displacement mechanism through the first displacement block and the first displacement frame.

[0009] As a further technical solution of the present invention, the second displacement mechanism includes a second displacement frame and a second stepper motor. A second lead screw is rotatably connected inside the second displacement frame. A second displacement block is threadedly connected to the middle of the second lead screw and slidably connected to the second displacement frame. One side of the second displacement frame is fixedly connected to one side of the second stepper motor. The output end of the second stepper motor is fixedly connected to the end of the second lead screw that is directly opposite to it. The second stepper motor drives the second lead screw to rotate. The thread on the surface of the second lead screw matches the thread on the inner wall of the second displacement block. The second displacement block is limited by the second displacement frame, which matches its shape and size. Therefore, the second displacement block slides along the second lead screw to perform secondary adjustment of the position of the visual analysis probe.

[0010] As a further technical solution of the present invention, the bottom end of the second displacement frame is fixedly connected to the first displacement mechanism, the top end of the second displacement block is fixedly connected to the mounting frame, and the second displacement mechanism is installed between the first displacement mechanism and the mounting frame through the second displacement frame and the second displacement block.

[0011] As a further technical solution of the present invention, a servo motor is fixedly installed in the middle of the top of the lifting platform, and a rotating shaft is fixedly installed at the output end of the servo motor. The top end of the rotating shaft is fixedly connected to the bottom end of the turntable. The servo motor drives the rotating shaft to rotate, and the rotating shaft drives the turntable to move synchronously to adjust the detection direction of the visual analysis probe.

[0012] As a further technical solution of the present invention, lifting rods that are slidably connected to the lifting platform are fixedly installed on both sides of the top of the fixed base. A control panel is fixedly installed on one end of the lifting platform. The lifting cylinder performs telescopic movement and pushes the lifting platform from the bottom to adjust the height of the vision analysis probe. During the lifting process, the lifting rods slide relative to the lifting platform to improve the stability of the lifting platform.

[0013] A method for using an AI-powered multimodal image semantic segmentation and object detection device includes the following steps: Step 1: Multimodal Image Acquisition and Preprocessing: The multimodal image acquisition module simultaneously acquires raw multimodal image data, including visible light images, infrared images, and depth images, corresponding to the scene. The image preprocessing module performs unified preprocessing on each modality image, including image denoising, size normalization, brightness adaptive correction, and resolution alignment operations, eliminating size deviations, pixel offsets, and environmental noise interference between different modal images to obtain a standardized multimodal image dataset. At the same time, pixel-level registration is performed on the preprocessed images to ensure that the pixel information of the same spatial location in different modalities corresponds one-to-one, laying the foundation for subsequent feature fusion. Step 2: Cross-modal Feature Alignment and Enhancement: The cross-modal feature processing module constructs a multimodal feature extraction branch, employing differentiated feature extraction strategies based on the characteristics of different modalities. For visible light images, it extracts shallow features such as texture, color, and contour, as well as semantic features of details; for infrared images, it extracts temperature differences and heat source target features; and for depth images, it extracts spatial distance and 3D contour features. Simultaneously, a cross-modal attention weight allocation mechanism is introduced to adaptively calculate the confidence weights of each modal feature. The calculation formula is as follows: , Among them, w m For the corresponding modal weights , The Sigmoid activation function maps feature responses to the 0-1 range, where v is the visible light image, r is the infrared image, and d is the depth image. The multimodal fusion feature calculation formula is: , Among them, w m F is the single-mode deep feature map corresponding to the m-th mode. m Let v be the adaptive feature weight corresponding to the m-th modality, v be the visible light image, r be the infrared image, and d be the depth image. Effective features are enhanced, and invalid noise features are suppressed to achieve accurate alignment and adaptive enhancement of multimodal features, thus solving the problems of feature redundancy and loss of effective features in traditional fusion methods. Step 3: Multi-scale Feature Fusion Encoding: The multi-scale feature fusion encoding module constructs a lightweight multi-scale feature encoding network. It performs multi-scale sampling on the enhanced multi-modal fusion features, extracting shallow fine-grained small target features, mid-level target contour features, and deep global semantic features. The bidirectional fusion of multi-scale features is completed through a feature pyramid fusion structure. The multi-scale feature weighted fusion formula is as follows: , in Adaptive weights for features at each scale, satisfying F1 is the shallow, fine-grained small target feature, F2 is the medium-level target contour feature, and F3 is the deep global semantic feature. It retains the detailed features of small targets and the global semantic features of large scenes, and generates a global fusion feature map containing multi-dimensional and multi-scale information, taking into account the detection and segmentation accuracy of targets of different sizes. Step 4: Dual-Task Collaborative Inference: The dual-task collaborative inference module, based on the globally fused feature map, initiates a dual-task collaborative inference mechanism for detection and segmentation. The target detection branch generates candidate boxes, performs bounding box regression, and classifies categories, outputting the target's location coordinates, size, category, and confidence score. The semantic segmentation branch performs pixel-level semantic classification, edge refinement, and region connectivity optimization, outputting pixel-level semantic segmentation results for the entire image, dividing different scene regions and target semantic attributes. The joint loss function formula for the dual tasks is: , Where L det To detect the loss, L scg To divide the loss, , For balancing weighting coefficients; Step 5: Result Optimization Output: The result optimization output module performs post-processing optimization on the detection and segmentation results. It removes duplicate detection boxes and low-confidence false positives using a non-maximum suppression algorithm, and optimizes semantic segmentation edges for jagged edges and holes using an edge smoothing algorithm. The segmentation result optimization employs the following edge smoothing filtering formula: , in, N represents the pixel neighborhood window, and N is the number of neighboring pixels. Finally, the intelligent deployment module outputs high-precision target detection results and pixel-level semantic segmentation results, completing multimodal image intelligent perception.

[0014] Compared with the prior art, the beneficial effects of the present invention are: 1. By setting up a multimodal image acquisition module and a multimodal image fusion perception architecture, the system integrates visible light, infrared, and depth multi-dimensional image information, which makes up for the information defects of single-modal images in harsh environments and complex scenes, greatly improves the anti-interference ability of target detection and semantic segmentation, effectively reduces the probability of missed detection and false detection, and adapts to the visual perception needs of all weather and all scenes. 2. By setting up a cross-modal feature processing module, a cross-modal adaptive attention fusion mechanism is introduced. The feature weights of different modalities are adaptively allocated according to the scene environment, effectively strengthening the effective features and suppressing noise interference. This solves the problems of feature redundancy and low alignment accuracy in traditional multimodal fusion, and significantly improves the recognition and segmentation accuracy of small targets and occluded targets in complex scenes. 3. By setting up a dual-task collaborative reasoning module, and adopting a detection and segmentation dual-task collaborative reasoning structure, we can achieve front-end feature parameter sharing and dual-task parallel reasoning. This abandons the traditional step-by-step independent operation mode, greatly reduces model computational redundancy, improves reasoning efficiency, and balances algorithm accuracy and real-time performance. Attached Figure Description

[0015] Figure 1 This is a side view of the present invention; Figure 2 This is a perspective view of the lifting mechanism of the present invention; Figure 3 This is a side view of the first displacement mechanism of the present invention; Figure 4 This is a perspective view of the second displacement mechanism of the present invention; Figure 5 This is a schematic diagram of the architecture of the visual analysis probe of the present invention; Figure 6 This is a flowchart of the present invention.

[0016] In the diagram: 1. Lifting mechanism; 11. Fixed base; 12. Lifting platform; 13. Lifting rod; 14. Lifting cylinder; 2. Servo motor; 3. Rotating shaft; 4. Turntable; 5. First displacement mechanism; 51. First displacement frame; 52. First lead screw; 53. First displacement block; 54. First stepper motor; 6. Control panel; 7. Second displacement mechanism; 71. Second displacement frame; 72. Second lead screw; 73. Second displacement block; 74. Second stepper motor; 8. Mounting frame; 9. Visual analysis table; 10. Visual analysis probe; 101. Multimodal image acquisition module; 102. Image preprocessing module; 103. Cross-modal feature processing module; 104. Multi-scale feature fusion encoding module; 105. Dual-task collaborative reasoning module; 106. Result optimization output module; 107. Intelligent deployment module. Detailed Implementation

[0017] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1-6 The present invention provides an AI intelligent multimodal image semantic segmentation and target detection device, including a lifting mechanism 1; The lifting mechanism 1 includes a fixed base 11 and a lifting cylinder 14. The middle part of the top of the fixed base 11 is fixedly connected to the fixed end of the lifting cylinder 14. The movable end of the lifting cylinder 14 is fixedly mounted with a lifting platform 12. The top of the lifting platform 12 is provided with a turntable 4. The top of the turntable 4 is fixedly mounted with a first displacement mechanism 5. The top of the first displacement mechanism 5 is fixedly mounted with a second displacement mechanism 7. The top of the second displacement mechanism 7 is fixedly mounted with a mounting frame 8. A visual analysis table 9 is fixedly mounted on one side of the mounting frame 8. A visual analysis probe 10 is fixedly mounted on one side of the visual analysis table 9. The visual analysis probe 10 includes a multimodal image acquisition module 101, an image preprocessing module 102, a cross-modal feature processing module 103, a multi-scale feature fusion coding module 104, a dual-task collaborative reasoning module 105, a result optimization output module 106, and an intelligent deployment module 107.

[0019] The image preprocessing module 102 is electrically connected to the multimodal image acquisition module 101, the cross-modal feature processing module 103 is connected to the image preprocessing module 102, the multi-scale feature fusion encoding module 104 is connected to the cross-modal feature processing module 103, and the result optimization output module 106 is connected to the dual-task collaborative reasoning module 105.

[0020] In use, the multimodal image acquisition module 101 is connected to a visible light camera, an infrared imaging module, and a depth camera respectively, to simultaneously acquire visible light images, infrared images, and depth images of the scene, realize real-time acquisition and synchronous transmission of multi-dimensional image data, and ensure the temporal consistency of multimodal data; The image preprocessing module 102 has built-in image denoising unit, size normalization unit, pixel registration unit, and brightness correction unit, which are used to complete the denoising, size unification, accurate pixel registration and image quality optimization of multimodal images, and eliminate data differences and environmental interference between modalities. The cross-modal feature processing module 103 includes a multi-branch feature extraction unit, an attention weight allocation unit, and a feature alignment enhancement unit, which are used to extract core features of each modality in a differentiated manner, adaptively allocate modality weights, complete accurate fusion and enhancement of cross-modal features, filter invalid noise features, and strengthen key target feature information. The multi-scale feature fusion coding module 104 has a built-in feature pyramid network structure, which is used to perform multi-scale sampling and bidirectional fusion of fused features, aggregate shallow detailed features and deep global semantic features, and generate a high-dimensional and robust global fusion feature map. The dual-task collaborative reasoning module 105 is connected to the target detection reasoning unit and the semantic segmentation reasoning unit respectively. The two units share the front-end encoding features and simultaneously complete the target localization classification and pixel-level semantic segmentation reasoning, realizing parallel and efficient operation of the two tasks, avoiding repeated calculations and improving reasoning efficiency. The result optimization output module 106 has a built-in post-processing optimization algorithm to remove false detections and duplicate detections, optimize segmentation edge accuracy, output standardized and high-precision detection and segmentation results, and support data visualization and data storage. The intelligent deployment module 107 is used to complete the lightweight compression, quantization pruning and operator optimization of the model, and supports the deployment of the model on cloud servers, embedded terminals, mobile devices and industrial control equipment, adapting to the real-time inference needs of different scenarios.

[0021] The first displacement mechanism 5 includes a first displacement frame 51 and a first stepper motor 54. The first displacement frame 51 is rotatably connected to a first lead screw 52. The middle part of the first lead screw 52 is threadedly connected to a first displacement block 53 that is slidably connected to the first displacement frame 51. One side of the first displacement frame 51 is fixedly connected to one side of the first stepper motor 54. The output end of the first stepper motor 54 is fixedly connected to the end of the first lead screw 52 that is directly opposite to it.

[0022] In use, the first stepper motor 54 drives the first lead screw 52 to rotate. The thread on the surface of the first lead screw 52 matches the thread on the inner wall of the first displacement block 53. The first displacement block 53 is limited by the first displacement frame 51, which matches its shape and size. Therefore, the first displacement block 53 slides along the first lead screw 52 to make the first adjustment to the position of the vision analysis probe 10.

[0023] The top end of the first displacement block 53 is fixedly connected to the second displacement mechanism 7, and the bottom end of the first displacement frame 51 is fixedly connected to the turntable 4.

[0024] In use, the first displacement mechanism 5 is installed between the turntable 4 and the second displacement mechanism 7 via the first displacement block 53 and the first displacement frame 51.

[0025] The second displacement mechanism 7 includes a second displacement frame 71 and a second stepper motor 74. A second lead screw 72 is rotatably connected inside the second displacement frame 71. A second displacement block 73, which is slidably connected to the second displacement frame 71, is threadedly connected to the middle of the second lead screw 72. One side of the second displacement frame 71 is fixedly connected to one side of the second stepper motor 74. The output end of the second stepper motor 74 is fixedly connected to the end of the second lead screw 72 that is directly opposite to it.

[0026] In use, the second stepper motor 74 drives the second lead screw 72 to rotate. The thread on the surface of the second lead screw 72 matches the thread on the inner wall of the second displacement block 73. The second displacement block 73 is limited by the second displacement frame 71, which matches its shape and size. Therefore, the second displacement block 73 slides along the second lead screw 72 to make secondary adjustments to the position of the vision analysis probe 10.

[0027] The bottom end of the second displacement frame 71 is fixedly connected to the first displacement mechanism 5, and the top end of the second displacement block 73 is fixedly connected to the mounting frame 8.

[0028] In use, the second displacement mechanism 7 is installed between the first displacement mechanism 5 and the mounting frame 8 via the second displacement bracket 71 and the second displacement block 73.

[0029] A servo motor 2 is fixedly installed at the middle of the top of the lifting platform 12. A rotating shaft 3 is fixedly installed at the output end of the servo motor 2. The top of the rotating shaft 3 is fixedly connected to the bottom of the turntable 4.

[0030] In use, the servo motor 2 drives the rotating shaft 3 to rotate, and the rotating shaft 3 drives the turntable 4 to move synchronously, adjusting the detection direction of the vision analysis probe 10.

[0031] Both sides of the top of the fixed base 11 are fixedly installed with lifting rods 13 that are slidably connected to the lifting platform 12, and a control panel 6 is fixedly installed at one end of the lifting platform 12.

[0032] In use, the lifting cylinder 14 extends and retracts, pushing the lifting platform 12 from the bottom to adjust the height of the vision analysis probe 10. During the lifting process, the lifting rod 13 slides relative to the lifting platform 12, improving the stability of the lifting platform 12.

[0033] A method for using an AI-powered multimodal image semantic segmentation and object detection device includes the following steps: Step 1: Multimodal Image Acquisition and Preprocessing: The multimodal image acquisition module 101 simultaneously acquires raw multimodal image data, including visible light images, infrared images, and depth images, corresponding to the scene. The image preprocessing module 102 performs unified preprocessing on each modal image, including image denoising, size normalization, brightness adaptive correction, and resolution alignment operations, to eliminate size deviations, pixel offsets, and environmental noise interference between different modal images, resulting in a standardized multimodal image dataset. At the same time, pixel-level registration is performed on the preprocessed images to ensure that the pixel information of the same spatial location in different modalities corresponds one-to-one, laying the foundation for subsequent feature fusion. Step 2: Cross-modal feature alignment and enhancement: The cross-modal feature processing module 103 constructs a multimodal feature extraction branch, employing differentiated feature extraction strategies based on the characteristics of different modal images. For visible light images, it extracts shallow features such as texture, color, and contour, as well as semantic features of details; for infrared images, it extracts temperature differences and heat source target features; and for depth images, it extracts spatial distance and 3D contour features. Simultaneously, a cross-modal attention weight allocation mechanism is introduced to adaptively calculate the confidence weights of each modal feature. The calculation formula is as follows: , Among them, w m For the corresponding modal weights , The Sigmoid activation function maps feature responses to the 0-1 range, where v is the visible light image, r is the infrared image, and d is the depth image. The multimodal fusion feature calculation formula is: , Among them, w m F is the single-mode deep feature map corresponding to the m-th mode. m Let v be the adaptive feature weight corresponding to the m-th modality, v be the visible light image, r be the infrared image, and d be the depth image. Effective features are enhanced, and invalid noise features are suppressed to achieve accurate alignment and adaptive enhancement of multimodal features, thus solving the problems of feature redundancy and loss of effective features in traditional fusion methods. Step 3: Multi-scale Feature Fusion Encoding: The multi-scale feature fusion encoding module 104 constructs a lightweight multi-scale feature encoding network. It performs multi-scale sampling on the enhanced multi-modal fusion features, extracting shallow fine-grained small target features, mid-level target contour features, and deep global semantic features. The bidirectional fusion of multi-scale features is completed through a feature pyramid fusion structure. The multi-scale feature weighted fusion formula is as follows: , in Adaptive weights for features at each scale, satisfying F1 is the shallow, fine-grained small target feature, F2 is the medium-level target contour feature, and F3 is the deep global semantic feature. It retains the detailed features of small targets and the global semantic features of large scenes, and generates a global fusion feature map containing multi-dimensional and multi-scale information, taking into account the detection and segmentation accuracy of targets of different sizes. Step 4: Dual-Task Collaborative Inference: The dual-task collaborative inference module 105, based on the global fusion feature map, initiates a dual-task collaborative inference mechanism for detection and segmentation. The target detection task branch generates candidate boxes, performs bounding box regression, and classifies categories, outputting the target's location coordinates, size, category, and confidence score. The semantic segmentation task branch performs pixel-level semantic classification, edge refinement, and region connectivity optimization, outputting pixel-level semantic segmentation results for the entire image, dividing different scene regions and target semantic attributes. The joint loss function formula for the dual tasks is: , Where L det To detect the loss, L scg To divide the loss, , For balancing weighting coefficients; Step 5: Result Optimization Output: The result optimization output module 106 performs post-processing optimization on the detection and segmentation results. It removes duplicate detection boxes and low-confidence false detections using a non-maximum suppression algorithm, and optimizes semantic segmentation edges for jagged edges and holes using an edge smoothing algorithm. The segmentation result optimization employs the following edge smoothing filtering formula: , in, N is the number of neighboring pixels, and the intelligent deployment module 107 outputs high-precision target detection results and pixel-level semantic segmentation results to complete multimodal image intelligent perception.

[0034] In this invention, the lifting cylinder 14 performs telescopic movement, pushing the lifting platform 12 from the bottom to adjust the height of the visual analysis probe 10. During the lifting process of the lifting platform 12, the lifting rod 13 slides relative to the lifting platform 12, improving the stability of the lifting platform 12. The servo motor 2 drives the rotating shaft 3 to rotate, and after the rotating shaft 3 rotates, it drives the turntable 4 to move synchronously, adjusting the detection direction of the visual analysis probe 10. The first stepper motor 54 drives the first lead screw 52 to rotate. The thread on the surface of the first lead screw 52 matches the thread on the inner wall of the first displacement block 53. The first displacement block 53 is limited by the first displacement frame 51, which matches its shape and size. Therefore, the first displacement block 53 slides along the first lead screw 52, ​​adjusting the detection direction of the visual analysis probe 10. The first adjustment is made to the position of the head 10. The second stepper motor 74 drives the second lead screw 72 to rotate. The thread on the surface of the second lead screw 72 matches the thread on the inner wall of the second displacement block 73. The second displacement block 73 is limited by the second displacement frame 71, which matches its shape and size. Therefore, the second displacement block 73 slides along the second lead screw 72 to make a second adjustment to the position of the visual analysis probe 10. The multimodal image acquisition module 101 is connected to the visible light camera, the infrared imaging module, and the depth camera respectively, and is used to simultaneously acquire visible light images, infrared images, and depth images of the scene to realize real-time acquisition and synchronous transmission of multi-dimensional image data and ensure the temporal consistency of multimodal data. The image preprocessing module 102 has a built-in image denoising unit and a size normalization unit. The image processing module 103 includes a pixel registration unit, a brightness correction unit, and a multimodal image denoising, size unification, precise pixel registration, and image quality optimization, eliminating data differences and environmental interference between modalities. The cross-modal feature processing module 103 includes a multi-branch feature extraction unit, an attention weight allocation unit, and a feature alignment enhancement unit, used to differentially extract core features of each modality, adaptively allocate modal weights, achieve precise cross-modal feature fusion and enhancement, filter invalid noise features, and strengthen key target feature information. The multi-scale feature fusion encoding module 104 incorporates a feature pyramid network structure for multi-scale sampling and bidirectional fusion of fused features, aggregating shallow detail features and deep global semantic features to generate a high-dimensional, robust global fusion feature map. The task collaborative reasoning module 105 connects to the target detection reasoning unit and the semantic segmentation reasoning unit respectively. The two units share front-end encoding features and simultaneously complete target localization and classification and pixel-level semantic segmentation reasoning, realizing parallel and efficient operation of the two tasks, avoiding redundant calculations, and improving reasoning efficiency. The result optimization output module 106 has a built-in post-processing optimization algorithm to remove false detections and duplicate detections, optimize segmentation edge accuracy, and output standardized and high-precision detection and segmentation results, and supports data visualization and data storage. The intelligent deployment module 107 is used to complete the lightweight compression, quantization pruning and operator optimization of the model, and supports the deployment of the model on cloud servers, embedded terminals, mobile devices and industrial control equipment, adapting to the real-time reasoning needs of different scenarios.

[0035] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An AI-powered multimodal image semantic segmentation and target detection device, comprising a lifting mechanism (1). Its features are: The lifting mechanism (1) includes a fixed base (11) and a lifting cylinder (14). The middle part of the top of the fixed base (11) is fixedly connected to the fixed end of the lifting cylinder (14). The movable end of the lifting cylinder (14) is fixedly installed with a lifting platform (12). The top of the lifting platform (12) is provided with a turntable (4). The top of the turntable (4) is fixedly installed with a first displacement mechanism (5). The top of the first displacement mechanism (5) is fixedly installed with a second displacement mechanism (7). The top of the second displacement mechanism (7) is fixedly installed with a mounting frame (8). A visual analysis table (9) is fixedly installed on one side of the mounting frame (8). A visual analysis probe (10) is fixedly installed on one side of the visual analysis table (9). The visual analysis probe (10) includes a multimodal image acquisition module (101), an image preprocessing module (102), a cross-modal feature processing module (103), a multi-scale feature fusion encoding module (104), a dual-task collaborative reasoning module (105), a result optimization output module (106), and an intelligent deployment module (107).

2. The AI-powered multimodal image semantic segmentation and target detection device according to claim 1, characterized in that: The image preprocessing module (102) is electrically connected to the multimodal image acquisition module (101), the cross-modal feature processing module (103) is connected to the image preprocessing module (102), the multi-scale feature fusion encoding module (104) is connected to the cross-modal feature processing module (103), and the result optimization output module (106) is connected to the dual-task collaborative reasoning module (105).

3. The AI-powered multimodal image semantic segmentation and target detection device according to claim 1, characterized in that: The first displacement mechanism (5) includes a first displacement frame (51) and a first stepper motor (54). The first displacement frame (51) is rotatably connected to a first lead screw (52). The first lead screw (52) is threadedly connected to the middle part of the first displacement frame (51) and slidably connected to a first displacement block (53). One side of the first displacement frame (51) is fixedly connected to one side of the first stepper motor (54). The output end of the first stepper motor (54) is fixedly connected to the end of the first lead screw (52) that is directly opposite to it.

4. The AI-powered multimodal image semantic segmentation and target detection device according to claim 3, characterized in that: The top end of the first displacement block (53) is fixedly connected to the second displacement mechanism (7), and the bottom end of the first displacement frame (51) is fixedly connected to the turntable (4).

5. The AI-powered multimodal image semantic segmentation and target detection device according to claim 1, characterized in that: The second displacement mechanism (7) includes a second displacement frame (71) and a second stepper motor (74). The second displacement frame (71) is rotatably connected to a second lead screw (72). The middle part of the second lead screw (72) is threadedly connected to a second displacement block (73) that is slidably connected to the second displacement frame (71). One side of the second displacement frame (71) is fixedly connected to one side of the second stepper motor (74). The output end of the second stepper motor (74) is fixedly connected to the end of the second lead screw (72) that is directly opposite to it.

6. The AI-powered multimodal image semantic segmentation and target detection device according to claim 5, characterized in that: The bottom end of the second displacement frame (71) is fixedly connected to the first displacement mechanism (5), and the top end of the second displacement block (73) is fixedly connected to the mounting frame (8).

7. The AI-powered multimodal image semantic segmentation and target detection device according to claim 1, characterized in that: A servo motor (2) is fixedly installed at the middle of the top of the lifting platform (12), and a rotating shaft (3) is fixedly installed at the output end of the servo motor (2). The top of the rotating shaft (3) is fixedly connected to the bottom of the turntable (4).

8. The AI ​​intelligent multimodal image semantic segmentation and target detection device according to claim 1, characterized in that: Both sides of the top of the fixed base (11) are fixedly installed with lifting rods (13) that are slidably connected to the lifting platform (12), and a control panel (6) is fixedly installed at one end of the lifting platform (12).

9. A method of using the AI ​​intelligent multimodal image semantic segmentation and target detection device according to any one of claims 1-8, characterized in that, Includes the following steps: Step 1: Multimodal image acquisition and preprocessing: The multimodal image acquisition module (101) simultaneously acquires the original multimodal image data of the scene, including visible light image, infrared image, and depth image. The image preprocessing module (102) performs unified preprocessing on each modal image, including image denoising, size normalization, brightness adaptive correction, and resolution alignment operations, to eliminate size deviation, pixel offset, and environmental noise interference between different modal images, and obtain a standardized multimodal image dataset. At the same time, pixel-level registration is performed on the preprocessed images to ensure that the pixel information of the same spatial position of different modalities corresponds one-to-one, laying the foundation for subsequent feature fusion. Step 2, Cross-modal feature alignment and enhancement: The cross-modal feature processing module (103) constructs a multimodal feature extraction branch, adopts a differentiated feature extraction strategy for the characteristics of different modal images, and extracts texture, color, shallow contour features and detail semantic features from visible light images; Extract temperature differences and heat source target features from infrared images; Spatial distance and 3D contour features are extracted from depth images. A cross-modal attention weight allocation mechanism is introduced to adaptively calculate the confidence weights of each modality feature. The calculation formula is as follows: , Among them, w m For the corresponding modal weights , The Sigmoid activation function maps feature responses to the 0-1 range, where v is the visible light image, r is the infrared image, and d is the depth image. The multimodal fusion feature calculation formula is: , Among them, w m F is the single-mode deep feature map corresponding to the m-th mode. m Let v be the adaptive feature weight corresponding to the m-th modality, v be the visible light image, r be the infrared image, and d be the depth image. Effective features are enhanced, and invalid noise features are suppressed to achieve accurate alignment and adaptive enhancement of multimodal features, thus solving the problems of feature redundancy and loss of effective features in traditional fusion methods. Step 3, Multi-scale Feature Fusion Encoding: The multi-scale feature fusion encoding module (104) constructs a lightweight multi-scale feature encoding network, performs multi-scale sampling on the enhanced multi-modal fusion features, extracts shallow fine-grained small target features, mid-level target contour features, and deep global semantic features, respectively, and completes the bidirectional fusion of multi-scale features through the feature pyramid fusion structure. The multi-scale feature weighted fusion formula is as follows: , in Adaptive weights for features at each scale, satisfying F1 is the shallow, fine-grained small target feature, F2 is the medium-level target contour feature, and F3 is the deep global semantic feature. It retains the detailed features of small targets and the global semantic features of large scenes, and generates a global fusion feature map containing multi-dimensional and multi-scale information, taking into account the detection and segmentation accuracy of targets of different sizes. Step 4, Dual-Task Collaborative Reasoning: The dual-task collaborative reasoning module (105) initiates a dual-task collaborative reasoning mechanism for detection and segmentation based on the global fusion feature map. The target detection task branch generates candidate boxes, performs bounding box regression, and classifies categories to output the target's location coordinates, size, category, and confidence score. The semantic segmentation task branch performs pixel-level semantic classification, edge refinement, and region connectivity optimization to output the entire image's pixel-level semantic segmentation results, dividing different scene regions and target semantic attributes. The joint loss function formula for the dual tasks is: , Where L det To detect the loss, L scg To divide the loss, , For balancing weighting coefficients; Step 5, Result Optimization Output: The result optimization output module (106) performs post-processing optimization on the detection and segmentation results. It removes duplicate detection boxes and low-confidence false detections using a non-maximum suppression algorithm, and optimizes semantic segmentation edge jaggedness and holes using an edge smoothing algorithm. The segmentation result optimization uses the edge smoothing filter formula: , in, N is the number of neighboring pixels, and the intelligent deployment module (107) outputs high-precision target detection results and pixel-level semantic segmentation results to complete multimodal image intelligent perception.