End-to-end multi-task three-dimensional depth and semantic fusion visual perception system and method
The end-to-end multi-task stereo depth and semantic fusion visual perception system solves the problems of low stability, high latency and large space occupation in the existing technology, and achieves high real-time and stable visual perception, which is suitable for industrial production.
Patent Information
- Application Number
- CN202511429990.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-09
Smart Images

Figure CN120913137A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of industrial embodied intelligence, and in particular to an end-to-end multi-task stereo depth and semantic fusion visual perception system and method. BACKGROUND
[0002] In the industrial embodied intelligence scene, mechanical arms, AGVs, etc. need to rely on visual perception to complete environment modeling, target recognition and positioning to meet the control requirements of production line rhythm and man-machine cooperation. The industrial field space is limited, the interference is strong, and the continuous running time is long, which puts higher requirements on the real-time and stability of the perception system.
[0003] The existing mainstream technology usually adopts a separate architecture of 3D visual perception and visual AI recognition: the stereo / ToF camera generates depth or point cloud locally, the external industrial computer completes decoding, preprocessing and model inference, and the results are transmitted to the robot controller or PLC through Ethernet, and rely on independent geometric calibration and time synchronization mechanism.
[0004] The above existing technology still has the following problems: first, the equipment and connection are many, the hardware failure points increase, and the stability is insufficient under the conditions of electromagnetic interference and vibration; second, the cross-device large data transmission and multiple memory copying lead to uncontrollable end-to-end delay and jitter, and it is difficult to maintain the consistency of time base and external parameters; third, the volume and heat dissipation of the split type occupy a large space, which is not conducive to the deployment of the end or mobile platform; fourth, the installation, debugging, recalibration and version maintenance cost is high, and the concurrent expansion is limited, which is difficult to meet the application requirements of high real-time, high stability and narrow space. SUMMARY
[0005] Therefore, the embodiments of the present application provide an end-to-end multi-task stereo depth and semantic fusion visual perception system and method to solve the problems of low stability of the separate architecture, uncontrollable end-to-end delay and synchronization caused by cross-device data transmission, and large space occupation caused by volume and wiring complexity in the prior art.
[0006] The first aspect of the embodiment of the application provides a visual perception system for end-to-end multi-task stereo depth and semantic fusion, comprising: an image acquisition module configured to acquire a pair of synchronized images under a unified time base by a binocular imaging unit and output image data carrying a calibration parameter identifier; a fusion preprocessing module configured to perform geometric correction, noise suppression and stereo registration on the pair of synchronized images according to the calibration parameter, to obtain a pair of registered images; a multi-task network module comprising a shared feature extraction layer, a depth estimation branch and a semantic segmentation branch; wherein the shared feature extraction layer is configured to generate a shared feature map from the pair of registered images, the depth estimation branch is configured to construct a cost representation based on the shared feature map and output a depth map, and the semantic segmentation branch is configured to perform multi-scale semantic analysis based on the shared feature map and output a semantic mask; a multi-task coupling unit connected to the shared feature extraction layer, the depth estimation branch and the semantic segmentation branch, configured to impose consistency constraints on shared features and branch features, and transmit geometric priors and class priors between branches; and an output interface module configured to encapsulate the depth map and the semantic mask according to a preset data structure and output them to a real-time control interface of an external execution control system through an industrial real-time communication interface for calling.
[0007] The second aspect of the embodiment of the application provides a visual perception method for an end-to-end multi-task stereo depth and semantic fusion visual perception system, comprising: acquiring a pair of synchronized images under a unified time base by binocular imaging, and associating a time stamp and a calibration parameter identifier with the pair of synchronized images; performing geometric correction and epipolar alignment on the pair of synchronized images according to the calibration parameter, executing noise suppression, and completing stereo registration based on a correlation measure and sub-pixel interpolation, to obtain a pair of registered images; performing weight-shared hierarchical convolution coding on left and right views of the pair of registered images, and forming a shared feature map through cross-layer fusion; constructing a cost representation based on an adaptive disparity range on the shared feature map, performing weighted aggregation on the cost representation, generating continuous disparity through learnable sub-pixel refinement, converting the continuous disparity into a depth map, generating semantic features based on multi-scale semantic analysis and decoding and outputting a semantic mask; calculating a consistency measure of depth edges and semantic boundaries in a boundary neighborhood of the shared feature map, inputting geometric priors to semantic decoding, using class priors to perform item-by-item weighting and incompatible item shielding on the cost representation, to jointly correct boundary regions of the depth map and the semantic mask; encapsulating the depth map and the semantic mask, the corresponding time stamp, the calibration parameter identifier and a model version identifier according to a preset data structure, and outputting them to a real-time control interface of an external execution control system through an industrial real-time communication interface for calling.
[0008] The above at least one technical solution adopted by the embodiment of the application can achieve the following beneficial effects: The image acquisition module is used for acquiring a synchronous image pair under a unified time base through a binocular imaging unit and outputting image data carrying a calibration parameter identifier; the fusion preprocessing module is used for performing geometric correction, noise suppression and stereoscopic registration on the synchronous image pair according to the calibration parameter, to obtain a registered image pair; the multi-task network module includes a shared feature extraction layer, a depth estimation branch and a semantic segmentation branch; the shared feature extraction layer is used for generating a shared feature map from the registered image pair, the depth estimation branch is used for constructing a cost representation based on the shared feature map and outputting a depth map, and the semantic segmentation branch is used for performing multi-scale semantic analysis based on the shared feature map and outputting a semantic mask; the multi-task coupling unit is connected with the shared feature extraction layer, the depth estimation branch and the semantic segmentation branch, and is used for imposing consistency constraints on shared features and branch features and transmitting geometric priors and class priors between branches; and the output interface module is used for packaging the depth map and the semantic mask according to a preset data structure and outputting them to a real-time control interface of an external execution control system through an industrial real-time communication interface for calling. The application can realize end-to-end low latency and deterministic output, time synchronization and calibration consistency maintenance, compact integration and simplified wiring, and stably provide high-precision depth maps and semantic masks to a control system. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without any creative effort.
[0010] Figure 1 is the overall structural framework schematic diagram of the end-to-end multi-task stereo depth and semantic fusion visual perception system provided by the embodiments of the present application; Figure 2 is the structural composition schematic diagram of the end-to-end multi-task stereo depth and semantic fusion visual perception system provided by the embodiments of the present application; Figure 3 is the flow schematic diagram of the visual perception method of the end-to-end multi-task stereo depth and semantic fusion visual perception system provided by the embodiments of the present application. DETAILED DESCRIPTION
[0011] In the following description, specific details are set forth such as particular system structures, techniques, etc., in order to provide a thorough understanding of the embodiments of the present application. However, it should be apparent to those skilled in the art that the present application can be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted in order not to obscure the description of the present application with unnecessary details.
[0012] In the application field of industrial embodied intelligence, a mechanical arm, an AGV or other automated equipment relies on a visual perception device to clearly see the operating object and the surrounding environment, interact with people, and perform intelligent work, and ultimately realize human-machine collaboration on an industrial production line to solve the urgent problem of labor shortage in the field of intelligent manufacturing.
[0013] In order to solve this problem, the prior art scheme separates the 3D visual perception and visual AI recognition processing into two parts: the environmental depth perception calculation process is completed in an independent stereo camera, and the visual AI recognition process is completed in an independent industrial computer.
[0014] With more and more robots and other embodied intelligent agents entering various fields of human society, especially the field of industrial production, the degree of intelligence of the embodied intelligent agent such as an industrial robot is required to be higher and higher. As an important perception means of the embodied intelligent agent, the visual perception system has a core position in the entire embodied intelligent system and is the key to realizing embodied intelligence.
[0015] Although the prior art scheme can complete the functions of environmental perception, object classification and recognition, and positioning of the embodied intelligent agent, it also has the following defects: (1) the stereo visual perception and AI recognition calculation are separated, which reduces the stability of the overall system due to too many devices and too many external connections in the system, and increases the risk of interruption of industrial production capacity; (2) the separated scheme has large data transmission delay and cannot meet the demand for timely response of the embodied intelligence, so it cannot perform some high real-time process tasks; (3) the space volume occupied by the existing separated scheme is too large, so it cannot be used in many narrow space environments.
[0016] The present application aims to solve the problems of low system stability, large data transmission delay, and excessive space occupation caused by the separation of 3D visual perception and visual AI recognition processing in the prior art, and provides an end-to-end multi-task stereo depth and semantic fusion visual perception system to realize integrated processing of stereo depth perception and semantic recognition, improve system stability, reduce data transmission delay, and reduce space occupation, so as to meet the needs of high real-time performance, high stability, and narrow space operation in industrial production.
[0017] The overall structural framework of the end-to-end multi-task stereo depth and semantic fusion visual perception system provided by the present application and the functions of each part will be described below in conjunction with the drawings and embodiments. Figure 1 The overall structural framework of the end-to-end multi-task stereo depth and semantic fusion visual perception system provided by the present application is shown in FIG. 1, which specifically can include the following contents: Figure 1 The overall structural framework of the end-to-end multi-task stereo depth and semantic fusion visual perception system provided by the present application is shown in FIG. 1, which specifically can include the following contents: I. Image acquisition module It is composed of a binocular camera, a rigid base, a fine adjustment sliding groove, and an FPGA synchronous trigger unit. It is used for simultaneous exposure and output of synchronous image pairs under a unified time base; the base and the fine adjustment sliding groove are used for baseline adjustment and locking; the FPGA generates hardware triggers and timestamp identifiers, which are output together with the calibration parameter identifiers.
[0018] II. Fusion preprocessing module It is composed of a preprocessing unit and a data cache unit (such as an 8MB DDR3). The preprocessing unit performs distortion correction, noise suppression, and stereo registration on the synchronous image pair according to the calibration parameters to generate a registered image pair; the data cache unit manages the registered image pair and the timestamp and parameter version identifier in a ring buffer to provide a constant timing data source to the downstream.
[0019] III. Multi-task network module It includes a ResNet-50 feature extraction layer, a depth estimation branch, and a semantic segmentation branch. The feature extraction layer performs hierarchical encoding and cross-layer fusion of the registered image pair to generate a shared feature map; the depth estimation branch constructs a cost representation based on the shared feature map and outputs a depth map; the semantic segmentation branch performs multi-scale analysis on the shared feature map and outputs a semantic mask; the module has multi-task coupling logic to transfer geometric and class priors between the two branches.
[0020] IV. Output interface module It is composed of an EtherCAT / PROFINET communication unit and dual redundant links. It is used to package the depth map and semantic mask according to a preset data structure and send them periodically on an industrial real-time bus via zero-copy DMA; when the main link is abnormal, it switches to the standby link and continues to output data to the external execution control system.
[0021] V. Power module It is composed of a wide voltage input unit and an isolated output unit. It is used to receive 9-36V DC input and provide multiple isolated voltage outputs for the image acquisition module, fusion processing module, multi-task network module, and output module, and provide power good and reset control signals to related modules.
[0022] The data flow of the system is sequentially connected in the order of "acquisition → preprocessing → multi-task inference → industrial interface output", and the power module provides unified power supply and timing control for each functional module.
[0023] Based on the system overall architecture diagram shown in Figure 2 , the specific structure and function of the end-to-end multi-task stereo depth and semantic fusion visual perception system provided by the embodiments of the present application are described in detail in combination with the drawings and specific embodiments.Figure 2 is a structural composition diagram of an end-to-end multi-task stereo depth and semantic fusion visual perception system provided by an embodiment of the present application, as shown in the figure, the end-to-end multi-task stereo depth and semantic fusion visual perception system can specifically include the following modules: Figure 2 An image acquisition module 201 is configured to acquire a pair of synchronized images under a unified time base through a binocular imaging unit and output image data carrying a calibration parameter identifier; A fusion preprocessing module 202 is configured to perform geometric correction, noise suppression and stereo registration on the pair of synchronized images according to the calibration parameter, and obtain a pair of registered images; A multi-task network module 203 includes a shared feature extraction layer, a depth estimation branch and a semantic segmentation branch; wherein the shared feature extraction layer is configured to generate a shared feature map from the pair of registered images, the depth estimation branch is configured to construct a cost representation based on the shared feature map and output a depth map, and the semantic segmentation branch is configured to perform multi-scale semantic analysis based on the shared feature map and output a semantic mask; A multi-task coupling unit 204 is connected to the shared feature extraction layer, the depth estimation branch and the semantic segmentation branch, configured to impose consistency constraints on shared features and branch features, and transmit geometric priors and class priors between branches; An output interface module 205 is configured to encapsulate the depth map and the semantic mask according to a preset data structure, and output them to a real-time control interface of an external execution control system through an industrial real-time communication interface for calling.
[0024] In some embodiments, the image acquisition module is configured to: generate a synchronization trigger signal based on the unified time base and send it to the left imager and the right imager respectively; collect the left image frame and the right image frame at the same time under the action of the synchronization trigger signal, and write a timestamp corresponding to the unified time base for the image frames; perform parameter calibration during assembly or operation, associate the calibration parameter identifier with the image frames and output them with the frames; when detecting that a preset recalibration trigger condition is met, perform recalibration and update the calibration parameter, use the last version of the calibration parameter for output during recalibration, and output the frames with the updated calibration parameter identifier after the recalibration is completed; configure the binocular imaging unit through the baseline adjustable mechanism and the optical axis alignment mechanism to meet the geometric relationship required for stereo registration.
[0025] Specifically, the image acquisition module provided by the embodiment is configured to complete binocular synchronization acquisition, write timestamps and calibration parameter identifiers with the frames under a unified time base, perform recalibration and parameter atomic switching under a trigger condition, and complete geometric configuration through the baseline adjustable mechanism and the optical axis alignment mechanism.
[0026] The implementation of the key technical features and the concept of important terms involved in the present embodiment are explained as follows: I. Key technical features and implementation Unified time base and synchronous trigger. The internal time base of the FPGA is used as the unified time base, and a fixed pulse width synchronous trigger signal is generated by the trigger generator, which is sent to the left and right imagers through 50Ω coaxial lines. The trigger signal has a pulse width of 10μs and a channel time delay of ≤10ns. The FPGA counter accumulates with a resolution of 1μs, serving as the full-link time stamp source.
[0027] Time stamp and calibration parameter identification associated with frames. An 8-byte time stamp and calibration parameter identification are written in the frame header of each frame of image, and the identification is formed by combining the current calibration parameter version number and the check code. It is output with the frame and used as a parameter retrieval index downstream.
[0028] Parameter calibration and re-calibration. After assembly, initial calibration is performed to obtain intrinsic, extrinsic and distortion parameters and generate calibration parameter identification V0. When the detected vibration acceleration is ≥5g or the environmental temperature difference changes by >10℃ during operation, the re-calibration process is triggered, 10 groups of checkerboard images within a fixed field of view are collected, and Zhang's calibration method is used to calculate new parameters and generate identification V1. During the re-calibration calculation period, V0 is used as the frame output, and after completion, V0 is switched to V1 at the frame boundary. The total re-calibration time is ≤30s.
[0029] Geometric configuration. The baseline adjustable mechanism is composed of an aluminum alloy rigid base and a fine adjustment slot. The base flatness is ≤0.01mm, the minimum adjustment amount of the differential head is 0.01mm, the baseline adjustment range is 0~200mm, and the repeat positioning error after locking is ≤0.1mm. The optical axis alignment mechanism uses a laser collimator (precision ±1″) to adjust the parallelism of the optical axes of the two imagers during assembly, and the parallelism error is ≤0.05°. A cross-shaped level (precision 0.02mm / m) is provided at the bottom of the base to ensure the stable relationship between the camera imaging plane and the working plane.
[0030] II. Explanation of important terms and assembly points Unified time base: refers to a single clock reference provided by the FPGA, which is used for triggering, timing and time stamp writing.
[0031] Calibration parameter identification: a compact identification indicating the current intrinsic, extrinsic and distortion parameter version, which is used for parameter consistency reference and atomic switching in the processing link.
[0032] Parameter atomic switching: refers to the replacement of identification at the boundary between the end of a frame and the beginning of the next frame, to avoid parameter mixing in the same frame.
[0033] Mechanical alignment sequence: first coarse baseline adjustment, then optical axis collimation, finally locking and recording the target baseline value and attitude record code as the initial guess for subsequent re-calibration.
[0034] The implementation process of the embodiment is described in detail below in combination with examples in specific scenarios, and the specific content is as follows: First, assembly and initial calibration. Two global shutter industrial cameras (resolution 1280x1024, frame rate 30fps) are installed on a rigid base, and the baseline is set to 120mm and locked according to the target working condition. A laser collimator is used to correct the optical axis parallelism, and a level is used to adjust the vertical relationship with the normal direction of the station. A fixed checkerboard calibration board is placed in the field of view, and 10 groups of calibration images are collected to calculate the intrinsic and extrinsic parameters and (k1k3, p1p2) distortion parameters, generate a calibration parameter identifier V0 and write it into the FPGA register.
[0035] Next, synchronous acquisition and frame-by-frame writing. The system enters the running state, and the FPGA counts with 1μs resolution and periodically outputs a 10μs trigger pulse to the left and right imagers to achieve simultaneous exposure. Every time a pair of image frames is obtained, the timestamp T and the current calibration parameter identifier V0 are sequentially written in the frame header before being sent to the downstream fusion processing module. To control the cross-machine jitter, the FPGA compares the timestamp difference after receiving the feedback from the two cameras and performs fine correction within the range of ±0.2μs.
[0036] Then, re-calibration and parameter switching. After long-term operation, when the acceleration sensor detects a vibration ≥5g or the temperature sensor detects a temperature difference >10℃, the re-calibration process is triggered. The system automatically collects 10 groups of checkerboard images without stopping acquisition, calculates new parameters and generates an identifier V1; after receiving the "new parameter ready" signal, the FPGA records the switching point, switches the frame-by-frame identifier from V0 to V1 after the current frame output ends, and keeps the old and new identifiers simultaneously for one period for downstream comparison.
[0037] Then, maintenance and reconfiguration. When the station switching requires adjustment of the working distance, loosen the locking mechanism, adjust the differential head to the desired baseline value (for example, from 120mm to 160mm), and after completion, lock again and repeat the above initial calibration process to generate a new identifier V2; the time base and trigger configuration remain unchanged, only the identifier is updated to drive the processing link to reference the new parameters.
[0038] Through the above embodiment, the image acquisition module completes the binocular synchronous acquisition under the unified time base, the association of timestamps and calibration parameter identifiers with frame-by-frame, the re-calibration and parameter atomic switching based on trigger conditions, and the baseline and optical axis configuration that meets the stereo registration geometric relationship, providing consistent input data and parameter reference for subsequent fusion preprocessing and multi-task network inference.
[0039] In some embodiments, the fusion preprocessing module is used for: The pixel mapping based on the lookup table performs radial and tangential distortion compensation according to calibration parameters, and completes line alignment according to the epipolar constraint, wherein the geometric correction and the parameter version identification are associated with the frame, and atomic switching is performed when the parameters are updated; The spatial domain filtering is performed on the vector processing unit, and the filtering parameters are adaptively selected according to image statistics; The relative displacement of left and right views is estimated based on a correlation degree metric, and the synchronous image pair is aligned to a unified pixel grid through sub-pixel interpolation, and a registered image pair is output; A preprocessing pipeline with fixed processing timing is adopted, and the registered image pair is associated with the corresponding time stamp and parameter version and stored in a ring buffer structure for sequential reading by a multi-task network module.
[0040] Specifically, the fusion preprocessing module provided in the embodiment is used to complete fixed timing processing of geometric correction, spatial domain filtering and stereoscopic registration, and to provide a registered image pair and its time stamp and parameter version identification to a multi-task network module in a ring buffer manner.
[0041] The implementation of the key technical features and the concept of important terms involved in the embodiment are explained and described as follows: I. Key technical features and implementation Geometric correction and epipolar alignment. According to the calibration parameter identification carried by the image acquisition module with the frame, the current intrinsic, extrinsic and distortion parameters are retrieved in the parameter storage area, and a multi-scale lookup table LUT is generated in advance. The LUT takes a 16x16 pixel block as an index unit and records the mapping from distorted pixels to ideal pixel coordinates. During processing, the mapping is read by block and bilinear resampling is performed to complete radial and tangential distortion compensation; then, line alignment is performed on left and right views according to the epipolar constraint, and the right view is shifted by an integer pixel according to the line offset calculated based on the baseline and the extrinsic parameters. During the geometric correction process, the time stamp and parameter version identification are written into the frame header buffer. If a parameter update request is detected, only the old identification is replaced by the new identification at the frame boundary to ensure atomic switching of parameters.
[0042] Spatial domain filtering and adaptive parameter selection. The vector processing unit performs spatial domain filtering on the corrected left and right views in parallel. First, the histogram and local variance map of each frame are calculated, and then the filtering strategy is selected according to the histogram and local variance map: a 3x3 Gaussian kernel (σ=1.0) is used to smooth high-frequency noise in low-noise scenes; and in high-reflectivity or salt-and-pepper noise scenes, adaptive median filtering is switched between 3x3 and 5x5 according to the local variance threshold. The filtering is performed in a block-in-streaming manner, and the boundary pixels are expanded using mirroring.
[0043] Stereo rectification and sub-pixel interpolation. To eliminate the subtle view displacement caused by assembly tolerance and temperature drift, phase correlation method is used to estimate the relative displacement (Δx, Δy) of left and right views in frequency domain, and the initial value of displacement is provided by the row offset after epipolar alignment. The estimated displacement is decomposed into an integer part and a fractional part, the integer part is completed by pixel shift, and the fractional part is completed by bilinear interpolation for sub-pixel resampling, so that the right view is aligned to the unified pixel grid of the left view, and a registered image pair is generated.
[0044] Fixed timing pipeline and ring buffer. The module is organized in three levels of flow water according to the order of “geometric correction → spatial domain filtering → stereo rectification”, and the inter-frame delay is fixed. The output end is configured with a ring buffer, and the buffer items include {left registered frame, right registered frame, timestamp, parameter version identifier, ready flag}. The write pointer advances by frame and sets the flag after being ready, and the read pointer is sequentially read by the multi-task network module; when a parameter version change is detected, a switching identifier is recorded at the boundary of the first frame of the new and old versions, which facilitates consistent reference by the downstream in batch processing or fusion. To avoid write-read competition, the buffer adopts single-write and single-read arbitration and minimum water level control, and when the downstream is temporarily blocked, the data is kept in order in the ready queue without repeated calculation in the pipeline.
[0045] II. Important terms and implementation details LUT pixel mapping: indexed by block to reduce random access overhead, and the coordinates are recovered by pixel-level interpolation within the block; multi-scale pointer uses a layered LUT set for different focal lengths and distortion radii.
[0046] Parameter version identifier: composed of version number and check code, used as a unique key for downstream parameter retrieval; atomic switching is limited to inter-frame switching.
[0047] Phase correlation estimation: calculate the cross-power spectrum in frequency domain and locate the main peak to get the sub-pixel estimation of (Δx, Δy), which is used as the input of sub-pixel interpolation.
[0048] Ring buffer depth and pointer: the buffer depth is configured according to the frame rate and network inference time, for example, under the condition of 30fps and 8ms of downstream inference time, the buffer depth is set to 8 to ensure sequential reading and version consistency.
[0049] The implementation process of the embodiment will be described in detail below with examples in specific scenarios, and the specific content is as follows: Take the robot assembly station as an example, the input is a pair of synchronized images with resolution of 1280x1024 and frame rate of 30 fps. The module first identifies V1 loading corresponding LUT according to the calibration parameters, performs distortion compensation and epipolar alignment; then selects 3x3 Gaussian filter according to local variance to suppress the noise on the surface of metal parts; then estimates the displacement of the right view relative to the left view by phase correlation, i.e. Δx=0.25 pixels and Δy=0.00 pixels, and completes sub-pixel level bilinear interpolation to resample the right view to a unified pixel grid. The generated registration image pair is written into the ring buffer together with the timestamp Tn and parameter version identifier V1, and the ready flag is set for the multi-task network module to read. When the subsequent re-calibration obtains a new identifier V2, the module completes the switching from V1 to V2 at the frame boundary and records the switching point in the cache metadata, ensuring that the downstream shared feature extraction layer references consistent parameters within the same batch.
[0050] Through the above embodiment, the fusion preprocessing module realizes the stereo registration of geometric correction and epipolar alignment based on lookup table, vectorized adaptive spatial filtering, phase correlation and sub-pixel interpolation, and ordered data output based on fixed timing pipeline and ring buffer, providing input data with consistent registration, consistent parameters and consistent timestamp for the multi-task network module.
[0051] In some embodiments, the shared feature extraction layer is used for: performing hierarchical convolutional feature coding on the left and right views of the registration image pair respectively with weight sharing to obtain multi-scale feature representation; fusing and unifying the tensor format of the feature of different scales based on cross-layer connection to form a shared feature map; providing the shared feature map to the depth estimation branch and the semantic segmentation branch at the same time, and outputting a feature position index corresponding to the position of the shared feature map for calling by the multi-task coupling unit.
[0052] Specifically, the shared feature extraction layer provided in the embodiment is used for hierarchical convolutional feature coding, cross-layer fusion and tensor format unification on the registration image pair, generating a shared feature map and simultaneously serving the depth estimation branch and the semantic segmentation branch; and outputting a feature position index corresponding to the position of the shared feature map for calling by the multi-task coupling unit.
[0053] First, the input is the registration image pair output by the fusion preprocessing module and its timestamp and parameter version identifier. The shared feature extraction layer enters the isomorphic coding branch according to the left and right views respectively, and adopts weight sharing to ensure the consistency of the feature space of the two views. After coding and fusion, a single-path shared feature map is generated, the tensor step is unified to 16, and the number of channels is set to a fixed value during deployment to balance the two downstream branches. The output includes the shared feature map and the feature position index, which are written into the read-only buffer together with the timestamp and the parameter version identifier.
[0054] Further, the left and right views respectively pass through the same convolutional layer and down-sampling sequence, and the convolution kernel parameters are completely shared to reduce the domain difference introduced by the disparity. The encoding levels successively extract edge, texture and high semantic features from shallow to deep, and the resolution is gradually reduced and the receptive field is improved through the combination of step and pooling. In order to maintain the traceability of geometric alignment, each layer uses symmetric padding and deterministic down-sampling, so that the left and right features of the same layer can be one-to-one corresponding in spatial index. The left and right encoding outputs are aligned and spliced or bit-averaged at the same level to form a binocular fusion feature pyramid, which is used as the input of subsequent cross-layer fusion.
[0055] Further, in order to simultaneously retain detail and semantic information, the high-resolution features in the shallow layer and the high-semantic features in the deep layer of the pyramid are fused through cross-layer connection. Specifically, the deep layer features are up-sampled to the target resolution, and then fused with the corresponding shallow layer features in the channel dimension. During the fusion process, the channel dimension is standardized and the number of channels is unified to avoid downstream interface ambiguity caused by inconsistent channels. The multi-scale fusion sequence is performed from deep to shallow, and finally the intermediate features balanced between spatial resolution and semantic abstraction are obtained.
[0056] Further, after completing multi-scale fusion, a shared feature map is generated according to a preset step and channel number; the tensor adopts a fixed memory layout and alignment strategy, which facilitates parallel reading of the two downstream branches. The shared feature map is generated only once and is accessed by the depth estimation branch and the semantic segmentation branch in read-only mode in the buffer to avoid repeated calculation and additional copying.
[0057] Further, in order to support the multi-task coupling unit to perform consistent constraint in the boundary area, a feature position index is constructed to depict the mapping relationship between each position of the shared feature map and the original image coordinates. The index is calculated according to the registration mapping and the total step, records the original image coordinate center point corresponding to each feature position and the effective sampling range, and is encoded in fixed-point format and stored in association with the parameter version identifier. The feature position index and the shared feature map have the same spatial size and row-column order, and the coupling unit can read them directly by position alignment.
[0058] Further, after the shared feature map and the feature position index are written into the read-only buffer, the depth estimation branch and the semantic segmentation branch consume them in parallel according to the order number; when a parameter version identifier change is detected, the atomic switching of the new and old versions is completed at the frame boundary to ensure that the two branches refer to consistent features and position indexes within the same frame. The buffer is set with a minimum water level and overflow protection to maintain the reading order and avoid repeated encoding when the downstream is temporarily blocked.
[0059] The implementation process of the embodiment will be described in detail below in conjunction with examples in specific scenarios, and the specific content is as follows: The spatial size of the shared feature map is 80x64 when the input resolution is 1280x1024 and the step is unified as 16 at the mechanical arm assembly station. The shared feature map is used by the depth estimation branch to construct a disparity-related cost representation and by the semantic segmentation branch for multi-scale semantic parsing; the feature position index provides the spatial correspondence required for boundary alignment for the multi-task coupling unit, thereby completing unified reference without changing the interface of the two branches.
[0060] Through the above embodiments, the shared feature extraction layer completes hierarchical coding of weight sharing, cross-layer multi-scale fusion, and tensor format unification, stably provides shared feature maps and feature position indexes to the downstream, and ensures consistent reference of the same frame data by the two branches and the coupling unit under the cooperation of frame-level parameter atomic switching and read-only buffer mechanism.
[0061] In some embodiments, the depth estimation branch is used to: construct a cost representation based on the shared feature map within an adaptive disparity range, which is dynamically determined at the pixel or region level according to scene distance distribution or image statistics; weight and aggregate the class prior delivered by the shared feature map and the multi-task coupling unit with the cost representation corresponding to the attention weight, and apply a masking constraint to the unmatched region; perform a learnable sub-pixel refinement on the aggregation result, restore continuous disparity using inverse residual convolution, and convert the continuous disparity into a depth map according to the imaging model; reuse the shared feature map cache in the inference stage and execute it in parallel with the semantic segmentation branch to output a depth map that aligns with the data structure of the output interface module.
[0062] Specifically, the depth estimation branch of the present embodiment works on the basis of the shared feature map output by the shared feature extraction layer, and sequentially completes adaptive disparity range determination, cost representation construction and weighted aggregation, unmatched region masking, learnable sub-pixel refinement, and disparity-to-depth conversion, and executes in parallel with the semantic segmentation branch in the inference stage, and outputs a depth map consistent with the field format of the output interface module.
[0063] First, the branch reuses the shared feature map and the feature position index from the read-only buffer, while reading the timestamp and calibration parameter identifier carried with the frame. The module establishes an internal context with the frame number as the key, and obtains the disparity statistical histogram and texture intensity map of the previous frame as the priori. Based on the local gradient amplitude, texture entropy and the previous frame disparity distribution, the adaptive disparity range is dynamically determined at the pixel or region level, wherein the region division adopts the same method as the shared feature Figure 1The search window is set as [dmin, dmax] for each position, which is determined by the upper and lower bounds of the grid cell (e.g. 16x16) that the position belongs to. The upper bound is set larger for the near and texture-rich cells, and shrunk for the far or low-texture cells. The upper and lower bounds are updated by exponential smoothing and truncated by the camera geometry constraints, resulting in the search window [dmin, dmax] for each position.
[0064] Then, the branch aligns the disparity candidates of each position between the shared features of left and right views, and calculates the similarity measure of each disparity layer to form the cost representation. The similarity calculation uses the upper description of intra-channel correlation or cosine similarity, and the output cost representation is fixed-length coded in the channel dimension and arranged in the order of consecutive disparity layers, which facilitates subsequent sequential access and hardware cache prefetching. To ensure data alignment with the parallel semantic branch, the spatial step of the cost representation is set to be the same as the spatial step of the shared feature Figure 1 Consequently, the index relationship is provided by the feature position index.
[0065] Next, the branch receives the category prior tensor provided by the multi-task coupling unit, and jointly generates an attention weight map with the shared feature map. The attention weight performs weighted aggregation on the cost representation in the spatial and disparity dimensions, so that different disparity candidates belonging to the same spatial position are superimposed according to the weight, and the aggregated cost response is obtained. To handle the possible unmatched regions, the branch performs forward and backward checks on the left and right consistency, and simultaneously refers to the occlusion indication and texture confidence. For positions that do not satisfy the consistency or are determined to be occluded, a masking constraint is applied, the weight of the corresponding disparity item is set to zero, and an unmatched flag is recorded. In the subsequent refinement stage, the masked candidates are bypassed accordingly.
[0066] Subsequently, the branch performs a learnable sub-pixel refinement on the aggregated discrete disparity response. The refinement network adopts a cascaded structure composed of inverse residual convolution units, and the convolution step is set to 1. If necessary, an interpolation branch is used to restore the continuous disparity higher than the shared feature step. To improve the geometric consistency at the boundary, the refinement network takes the boundary guidance tensor generated by the shared feature map as an additional input, which is concatenated with the main feature under the premise of maintaining channel alignment and jointly participates in refinement. The network output is a continuous disparity map with the same size as the shared feature map, and a per-pixel confidence plane is attached as an optional output.
[0067] After completing the continuous disparity estimation, the branch converts the disparity into depth according to the imaging model. The conversion process reads the current baseline b and effective focal length f identified by the calibration parameters, and performs per-pixel calculation according to D = bf / disp. To be compatible with the output interface module field, the calculation result is quantized in the fixed-point domain and limited within the preset depth range, and the output is a 16-bit single-channel depth map. If there is a confidence plane, it is written into the load buffer in the form of an additional plane together with the depth map, and the field order is consistent with the real-time object mapping of the output interface module.
[0068] In the inference stage, the depth estimation branch and the semantic segmentation branch are executed in parallel with the same frame number, sharing the feature map and the feature position index which are accessed by both branches in read-only buffer. When the parameter version identifier is detected to change, the switching is only performed at the frame boundary, so that both branches refer to the consistent internal and external distortion parameter set within the same frame. After the depth estimation branch completes the above calculation, the depth map, timestamp, calibration parameter identifier and model version identifier are written to the reserved area of the load buffer, waiting for the output interface module to read and send from the zero-copy address according to the set period.
[0069] In an example scenario, the input image size of the mechanical arm assembly station is 1280x1024, and the shared feature map step size is 16, corresponding to a spatial size of 80x64. The system sets the adaptive disparity upper limit to 64 in the jig area according to the statistics of the previous frame, and to 32 in the distal background area. The cost representation is stacked according to the disparity layer and aggregated with a class prior weight, and the unmatched area is marked by the joint left-right consistency check and occlusion detection. The continuous disparity output by the refinement network is converted into a depth map after D=bf / disp conversion, and is written in 16-bit fixed-point format to the load buffer, together with the timestamp and the calibration parameter identifier, to provide the output interface module for industrial real-time communication transmission. The above process constitutes the complete implementation path of the depth estimation branch of the embodiment, which is consistent with the aforementioned modules in terms of timing and data interface.
[0070] In some embodiments, the semantic segmentation branch is used for: context modeling on the shared feature map based on a multi-scale dilated convolution array to form a semantic feature map; differential decoding and refinement on the semantic boundary corresponding area according to the geometric prior provided by the multi-task coupling unit; adaptive reweighting according to the class distribution in the model training or updating stage, and selecting pixels that meet the dynamic threshold to participate in parameter updating based on online hard sample mining; generating a semantic mask by upsampling and pixel alignment on the refined semantic feature map and outputting.
[0071] Specifically, the semantic segmentation branch of the embodiment works on the basis of the shared feature map output by the shared feature extraction layer, completes context modeling, differential decoding and boundary refinement based on geometric prior, adaptive reweighting and online hard sample mining in the training or updating stage, and generates a semantic mask by upsampling and pixel alignment on the output side. The input is the shared feature map and the feature position index, as well as the geometric prior tensor provided by the multi-task coupling unit, and the output is the semantic mask and optional confidence plane consistent with the data structure of the output interface module.
[0072] Firstly, the branch performs context modeling of the shared feature map with a multi-scale dilated convolution array. The dilation rates are configured as 2, 4, 8, 16, and the kernel size and the number of channels are fixed at deployment, so that the network obtains context information of different receptive fields without reducing the spatial resolution. The outputs of the branches at different scales are aligned and fused into a semantic feature map in the channel dimension. The fusion process maintains the alignment with the shared feature Figure 1 map, and the step size and memory layout are consistent, facilitating subsequent parallel reading and index consistency maintenance.
[0073] Secondly, the branch performs differential decoding and boundary refinement according to the geometric prior provided by the multi-task coupling unit. The geometric prior tensor is composed of planarity indicators, normal vectors, and depth gradients calculated by the depth estimation branch. After size matching, it is concatenated with the semantic feature map in the channel dimension into the decoding layer. The system enables the boundary decoding path within the boundary neighborhood indexed by the feature position, and the trunk decoding path in the non-boundary area. The two paths share parameter constraints and perform selective fusion at the output end, so that finer classification results are obtained at the boundary position, while stable semantic analysis is maintained in large areas. The boundary neighborhood is determined by the joint threshold of the depth gradient and the shared feature gradient. The threshold is stored as a configurable parameter after online calibration and is associated with the parameter version identifier.
[0074] Thirdly, the branch implements adaptive reweighting and online hard sample mining during model training or online updating. Adaptive reweighting calculates class weights according to the effective sample number of class distribution. The weights are refreshed according to the statistics in each training period and written into the loss function. Online hard sample mining sorts the current batch based on pixel-level loss, and selects pixels above the dynamic threshold for backpropagation. The dynamic threshold is determined by the batch loss quantile to adapt to the difficulty changes in different scenarios. The above strategies and the consistency constraint maintained by the multi-task coupling unit take effect at the same time, ensuring that the semantic branch shares the same training step and cache queue with the depth branch when updating parameters.
[0075] Subsequently, the branch generates a semantic mask from the refined semantic feature map through upsampling and pixel alignment. The upsampling process maintains the alignment strategy with the shared feature Figure 1 map, and the pixel alignment maps the feature coordinates back to the original image grid according to the feature position index to generate a class map with the same resolution as the input image. The class map uses 8-bit class encoding, and the optional output pixel-wise confidence plane is written into the load buffer along with the timestamp, calibration parameter identifier, and model version identifier, waiting for the output interface module to read and send according to the real-time object mapping. If a parameter version identifier change is detected, the switching is only performed at the frame boundary to ensure that the semantic branch and the depth branch refer to consistent calibration parameters and shared features within the same frame.
[0076] For example, in the specific application of a mechanical arm assembly station, the input image resolution is 1280x1024, and the shared feature map space size is 80x64. The semantic segmentation branch identifies 10 types of objects including bolts, nuts, gaskets, tooling plates, and conveyors. The hole rate extracts multi-scale contexts in parallel branches with a ratio of 2, 4, 8, and 16. The planarity and normal vector in the geometric prior are used to stabilize the main stem decoding path in the conveyor and tooling plate regions, and the boundary decoding path is enabled at the bolt and nut boundary to enhance the detail judgment. In the training or updating phase, the class weight is calculated according to the class distribution, and online hard sample mining is used to select about 30% of the high-loss pixels to participate in parameter updating. The final output semantic mask and confidence plane are written into the payload buffer in the field order aligned with the output interface module, and are periodically sent to the external execution control system, together with the generated depth map, for subsequent positioning and planning.
[0077] In some embodiments, the multi-task coupling unit is used for: In the boundary neighborhood of the shared feature map, a consistency measure is calculated for the edge features of the depth estimation branch and the boundary features of the semantic segmentation branch, and the consistency measure is used as a training constraint term to constrain the boundary alignment of the two branches; A geometric prior tensor containing planarity and normal vector information is generated from the depth estimation branch, and the geometric prior tensor is input into the decoding layer of the semantic segmentation branch and fused with the feature tensor of the decoding layer to form a geometric-constrained semantic feature; A class prior tensor is generated from the semantic segmentation branch, which is mapped to the weight of the pixel and disparity combination, each term of the cost representation of the depth estimation branch is weighted, and the weight of the pixel and disparity combination term incompatible with the class prior is set to zero; In the training phase, the training constraint term, the geometric prior fusion term, and the class prior weighting term are jointly optimized, and in the inference phase, the fusion and weighting are performed according to the fixed parameters obtained by training.
[0078] Specifically, the multi-task coupling unit of the present embodiment is deployed between the shared feature extraction layer, the depth estimation branch, and the semantic segmentation branch, and the input is the shared feature map and the feature position index, the geometric features from the depth estimation branch, and the class features from the semantic segmentation branch. The output is a consistency constraint term for the training phase, a geometric prior fused semantic feature, and a weighting coefficient for the cost representation. All tensors are associated with frame numbers and parameter version identifiers to ensure consistency of the three references in the same frame.
[0079] In the construction of the boundary consistency constraint, the coupling unit first locates the boundary neighborhood on the shared feature map according to the feature position index, and the boundary neighborhood is determined by the joint threshold of the depth gradient amplitude and the shared feature gradient, and the neighborhood width is 2-3 pixels in the feature map scale. The edge feature output by the depth estimation branch is denoted as Ed, and the boundary feature output by the semantic segmentation branch is denoted as Es. The coupling unit registers Ed and Es at the same spatial position and calculates a consistency measure, which adopts the combination of amplitude difference and direction similarity to form a single-channel constraint map. In the training stage, the constraint map is added to the loss function to constrain the alignment degree of the two branches in the boundary neighborhood; in the inference stage, only the consistency measure is retained as an internal quality reference, and is not output to the external interface.
[0080] In the generation and fusion of geometric priors, the coupling unit receives a geometric prior tensor containing planarity and normal vector information from the depth estimation branch. The geometric prior is composed of local plane fitting confidence, unit normal three components and depth gradient, and is aligned with the shared feature map in spatial size and channel order. After linear mapping, the geometric prior is concatenated with the feature tensor of the semantic segmentation branch decoding layer in the channel dimension, and the decoding layer obtains the geometrically constrained semantic feature accordingly. Higher geometric channel weights are retained in the boundary neighborhood to enhance the classification judgment of small component boundaries; in the large-area near-plane region, the geometric channel weight is reduced to stabilize the semantic analysis of weak texture regions. The fusion of geometric prior and semantic feature is completed once in the same frame, avoiding prior leakage across frames.
[0081] In the generation and weighting of class priors, the coupling unit receives a class probability map from the semantic segmentation branch and maps it to a class prior tensor. The class prior is aligned to the disparity search grid of the depth estimation branch through feature position indexing. The coupling unit maps the class prior to the weight table of the pixel and disparity combination according to the consistency relationship between the class and the disparity, and weights each item of the cost representation of the depth estimation branch; for the pixel and disparity combination that is judged to be incompatible with the class prior, the weight is set to zero and a mask flag is recorded to block the propagation of unreasonable matches in the subsequent refinement stage. The above mapping rules are configured according to the scene when the system is deployed and can be upgraded with model version identification.
[0082] In the joint optimization of the training stage, the coupling unit sends the boundary consistency constraint term, the geometric prior fusion term, and the category prior weighting term into the training pipeline at the same time. The boundary consistency constraint term takes the pixel set of the boundary neighborhood as the scope, the geometric prior fusion term takes the semantic features output by the decoding layer as the scope, and the category prior weighting term takes the cost representation as the scope. The three share the same batch of feature position indexes and parameter version identifiers to ensure the consistent alignment of the gradients. In the optimization process, a fixed loss weight upper configuration is adopted to support updating the weight ratio according to the scene on a periodic basis. In the inference stage, the gradient backpropagation is stopped, only the trained fusion and weighting parameters are retained, and the atomic switching of the parameters version identifiers is performed at the frame boundary.
[0083] For example, in a specific application example, the shared feature map space size of the mechanical arm assembly station is 80x64, and the step size is 16. The coupling unit constrains the coincidence relationship between the depth edge and the semantic boundary at the interface between the bolt and the tool plate through the consistency measure, and at the same time, uses the geometric prior with high planarity and stable normal to enhance the semantic decoding of the tool plate region. At the position of the bolt edge with high curvature, the category prior sets the weight of the incompatible combination term with large disparity to zero, thereby suppressing the interference of the background and the mismatched disparity. This process is completed within a single frame and maintains the timing consistency with the parallel inference of the depth estimation branch and the semantic segmentation branch.
[0084] Through the coupling mechanism of the above-mentioned embodiments, the system realizes the alignment constraint of the depth and semantic features in the boundary neighborhood, introduces the geometric prior derived from the depth in the semantic decoding, and uses the category prior derived from the semantics in the disparity search, forming bidirectional information transmission under the same time base and parameter version. The system can stably work in both the training and inference stages and can complete the prior fusion and weighting control while maintaining a unified data interface and frame-level atomic switching. The above-mentioned implementation enables the multi-task network to obtain consistent boundary alignment, controlled cost search, and stable semantic decoding performance in complex industrial scenes.
[0085] In some embodiments, the output interface module is configured to: associate the depth map and the semantic mask with the corresponding time stamp, calibration parameter identifier, and model version identifier in the field, and frame the payload buffer according to the preset data structure; map the fields in the payload buffer to real-time cyclic data objects according to the interface mapping table, configure the communication period and priority, and establish the corresponding relationship with the real-time control interface; associate the payload buffer with the industrial real-time communication interface, and send it according to the communication period under the trigger of a hardware interrupt; carry the check code and frame sequence number in the sent frame for sequence and integrity control, and carry the updated identifier with the frame when the parameter or model version is updated and perform atomic switching at the frame boundary; Upon detecting that the primary link meets the timeout threshold, the sending path is switched to the backup physical link, and the continuity of the payload buffer is maintained during the switching.
[0086] Specifically, the output interface module of the embodiment is located between the multi-task network module and the external execution control system, and the input is a depth map, a semantic mask, and an optional confidence plane, which are accompanied by a frame timestamp, a calibration parameter identifier, and a model version identifier. The output is a structured data frame periodically sent through an industrial real-time communication interface. The module works in a zero-copy DMA and interrupt-driven mode, and realizes redundant switching between the primary and backup physical links.
[0087] Firstly, the module associates the results and metadata from the multi-task network module with fields. The association content includes a timestamp T, a calibration parameter identifier VID, a model version identifier MID, a depth map pointer and length, a semantic mask pointer and length, and an optional confidence plane pointer and length. The fields form a frame header in a preset order, the frame header has a fixed length layout and records the starting offset and total length of the payload, and then the frame header and payload pointer are written into the payload buffer. The payload buffer is located in a pre-mapped physical address interval 0x40000000-0x40010000, and a double-buffer writing strategy is used to ensure continuous readability during sending.
[0088] Secondly, the module maps each field in the payload buffer to a real-time cyclic data object according to an interface mapping table, and configures the communication period and priority. For the EtherCAT scenario, the period is set to 1ms and the jitter is guaranteed by the lower layer hardware timing; for the PROFINET scenario, it works in IRT level ClassB and uses the equal period parameter. After the mapping is established, the module registers the payload buffer physical address and length to the DMA descriptor queue of the communication stack, so that the communication stack directly reads the data object according to the object index at each period without CPU moving.
[0089] Thirdly, the module associates the payload buffer with the industrial real-time communication interface and enters the interrupt-driven sending process. When the clock reaches the period boundary, a hardware interrupt is generated, and the communication stack reads the data object of the current ready buffer and frames and sends it. Each sending frame contains a frame header, a payload index, a CRC check code, and a frame sequence number. The CRC is used for frame integrity checking, and the frame sequence number is used for sequence and loss detection. If the calibration parameter or model version is updated in this period, the module writes the updated VID and MID at the period boundary and switches to the new buffer, ensuring that the atomic switching of the parameter and model only occurs between frames.
[0090] Subsequently, the module continuously monitors the timeout count and error statistics of the main link. When the timeout time of the main link exceeds a preset threshold (e.g., 500 μs) or the error count reaches a threshold, the module triggers a link switching state machine to switch the sending path to the backup physical link. The read-only state of the current ready buffer is maintained during the switching process, and the frame resources being sent are not recycled, so that the read pointer and the write pointer of the upper layer do not compete across frames. After the switching is completed, the link identifier is carried in the first frame for the upper system to record.
[0091] Again, the module establishes signal linkage with the power supply and alarm module. When receiving a power-off alarm, the module preferentially completes the framing and sending of the current frame, and attaches an abnormal identifier and the last valid frame sequence number at the end of the frame; if the power supply is restored within a single period, the sending process is automatically reset and continues to send according to the established mapping in the next period, without the need for the upper reset communication relationship.
[0092] For example, in the application example of a mechanical arm assembly station, the system sends with an EtherCAT 1 ms period, the depth map in the load buffer is in 16-bit single-channel format, the semantic mask is in 8-bit category encoding and can optionally be accompanied by a confidence plane. The module places T, VID, MID, frame sequence number and CRC in the fixed area of the frame header, and maps the depth map and semantic mask pointer to the load index of two real-time objects, which are directly DMA taken by the communication stack in the order of objects when the interrupt arrives and are sent out. When the online model update is completed, the module writes a new MID at the next frame period boundary and switches the buffer; when it is detected that the main link timeout exceeds the threshold, it is switched to the backup link within ≤1 ms, and the load buffer remains continuously readable during the switching period. The frame sequence number received by the external execution control system is continuous and uninterrupted.
[0093] Through the above embodiments, the output interface module realizes framing based on field association and fixed-length frame header, real-time object mapping based on interface mapping table, zero-copy sending based on DMA and interrupt, parameter and model atomic switching based on frame boundary, and link redundancy switching based on state machine, thereby providing a depth map and semantic mask data stream with uniform format, sequence checkable, link switchable and consistent with the upstream parameter version to the external execution control system within the specified period.
[0094] In some embodiments, the system further comprises a power supply and alarm module for: receiving a direct current input power supply and generating multiple mutually electrically isolated stabilized outputs; monitoring the voltage, current and temperature of the direct current input and each stabilized output, determining overvoltage, overcurrent, undervoltage or overtemperature according to a preset threshold, and outputting an alarm signal to the multi-task network module and the output interface module; When detecting that the voltage of the input power supply decreases to a threshold value, the driving energy buffer unit maintains short-time power supply, outputs an abnormality identifier to the output interface module, and issues a controlled shutdown signal to the image acquisition module, the fusion preprocessing module, and the multi-task network module; A power good signal and a reset control signal are provided to establish the power-on and power-off timing of the image acquisition module, the fusion preprocessing module, the multi-task network module, and the output interface module.
[0095] Specifically, the power supply and alarm module of the embodiment is arranged between the system input power supply and each functional module, is responsible for converting 9-36V DC input into multiple isolated stabilized outputs, and completes electrical parameter monitoring, abnormality determination, power-off buffering, and power-on / power-off timing control. The module output includes three isolated power supplies for a camera (12V / 2A), a GPU (5V / 5A), and an FPGA (3.3V / 1A), with a ripple of ≤50mV, supporting overvoltage, overcurrent, and undervoltage protection, and being linked with the multi-task network module and the output interface module through hardware signals and state pins.
[0096] On the input side, the module adopts EMI filtering and surge suppression combination, arranges common-mode inductance and π-type filter at the input end to reduce conducted interference, and configures a TVS tube to suppress transient overvoltage; to prevent reverse connection, a reverse connection protection path composed of an ideal diode controller and a MOSFET is used; the input undervoltage lock threshold is set to ensure that the downstream DC-DC has a stable working voltage level, and the undervoltage release threshold is relatively moved upward to form a back difference, avoiding repeated start-stop caused by edge jitter. The input current is measured by a sampling resistor and an operational amplifier, and the input temperature is monitored by a thermosensitive element or a board-mounted temperature sensor, and the measurement values of the three are used for alarm determination and event recording.
[0097] On the power conversion architecture, the module uses an isolated DC-DC as a primary stage, converts the input voltage to each intermediate bus, and then finely stabilizes it to the target level through a synchronous step-down stabilizer, thereby obtaining lower ripple and higher dynamic response under the premise of meeting the isolation requirement. Each path is configured with fast overcurrent limiting and short-circuit return, the overvoltage protection threshold is set to 38V, and the overcurrent threshold is set to 2.5A or calculated according to the safety factor of the rated current of the path. The outputs of each path are provided with remote sampling lines to compensate for the voltage drop of the connecting line, and the output end is connected in parallel with a small capacitor array and a low ESR electrolytic capacitor to suppress the voltage drop caused by the load step, and the soft start time of the FPGA and GPU power supply is coordinated according to the sequence relationship in the device data book to prevent power-on surge and incorrect power-on sequence.
[0098] In terms of monitoring and alarm logic, the module continuously samples the input voltage, current, and three-way output voltage, current, and on-board temperature. After comparing the sampled values with preset thresholds, an alarm level is generated and sent to the multi-task network module and output interface module through the hardware interrupt line and status pin. The threshold is divided into two levels: pre-warning and failure. Pre-warning is used to prompt the boundary state that protection will be triggered soon, and failure is used to trigger the controlled shutdown process. For transient peak signals, a time window filter is used, and only when the threshold exceeds the set value for a duration longer than the set value will the failure flag be set. The alarm logic has a hold and clear mechanism. Hold is used to record events, and clear is reset after the upper layer reads and confirms to avoid repeated responses.
[0099] In terms of power-down processing, the module configures an energy buffer unit to maintain short-time power supply when detecting that the input power voltage has dropped to a threshold. The power-down detection is monitored by a comparator that watches the input voltage and outputs a power-down edge when it falls below the threshold. The module immediately outputs an exception identifier to the output interface module, while latching the current frame number and timestamp, and allows the output interface module to prioritize completing data transmission for the current frame. At the same time, it sends a controlled shutdown signal to the image acquisition module, fusion preprocessing module, and multi-task network module, making them stop new operations in a predetermined order and freeze parameter version identifiers and buffer read-write pointers to avoid inconsistent data during voltage drop.
[0100] In terms of power-on and power-off timing control, the module provides power good signals and reset control signals for each output, establishing a power-on sequence in the order of "basic control layer → calculation acceleration layer → peripheral layer". Specifically, enable FPGA 3.3V and release FPGA reset after its PG stabilizes, then enable GPU 5V and release calculation module reset after its PG stabilizes, and finally enable the camera 12V and turn on the trigger logic and light control. The power-off sequence is executed in reverse order to ensure that the communication stack and inference tasks exit first. The PG signal is sent to the output interface module as one of the conditions for allowing entry into the real-time transmission period.
[0101] For example, in the application example, taking the EtherCAT 1 ms cycle working condition of the mechanical arm assembly station as an example, when the input power supply appears a rapid drop due to bus disturbance, the power-off threshold comparator of the module first detects the voltage drop and outputs a power-off edge, the energy buffer unit is put into work, and the output interface module accelerates to complete the current frame framing and sending after receiving the abnormal identifier, the multi-task network module stops a new reasoning batch after receiving the controlled shutdown signal and keeps the shared feature map buffer read-only, and the image acquisition module stops issuing a new trigger pulse. If the input power supply recovers to above the undervoltage release threshold within a short time, the module recovers each power supply in turn according to the power-on sequence and clears the retained fault identifier; if the input power supply is not recovered, the module completes the last frame sending before the buffer energy is exhausted and locks each output, waiting for the upper reset.
[0102] Through the above embodiment, the power supply and alarm module realizes unified power supply of wide voltage input and multi-path isolated voltage stabilization, real-time monitoring and hierarchical alarm of input and output key electrical parameters, rapid detection and short-time maintenance of power-off events, and consistent power-on / power-off sequence control with the rest of the system modules, so that the system can still complete consistent packaging and orderly exit of data under the working conditions of power supply disturbance, load mutation and link abnormality, and provide traceable event and state information for subsequent safe recovery.
[0103] Through the effect of the present application in practical application, compared with the traditional scheme, the technical scheme of the present application at least has the following advantages: 1. Stability and reliability: through the baseline adjustable structure, synchronous trigger optimization and automatic re-calibration, the system mean time between failures (MTBF) reaches 10,000 hours, which is 60% higher than the separated scheme; the deep error drift under vibration / temperature difference environment is ≤0.5%, which meets the industrial stability requirements.
[0104] 2. Real-time breakthrough: the total end-to-end delay (acquisition→output) is ≤20 ms (the traditional scheme is 50 ms), among which the pre-processing link is 3 ms, the network reasoning is 8 ms, and the communication is 5 ms, which can support 100 Hz control closed loop.
[0105] 3. Space and cost optimization: the integrated design volume is 0.015 m 3 (reduced by 75% compared with the separated scheme), the power consumption is 15 W (reduced by 60%), and the hardware cost is reduced by 40% (independent industrial computer is saved).
[0106] 4. Precision and robustness: the depth estimation error is ≤1% (the traditional scheme is 15%), the semantic segmentation mIoU is ≥92% (increased by 8%); the recognition rate remains above 85% under 60 dB noise and ±30% illumination change, which is suitable for complex industrial environments.
[0107] 5. Ease of maintenance: Supports online recalibration, hot model updates, and self-supervised optimization, reducing operation and maintenance costs by 50% and ensuring that the upgrade process does not interrupt production.
[0108] The above embodiments have described in detail the specific structure and function of the end-to-end multi-task stereo depth and semantic fusion visual perception system of this application. The implementation process of the end-to-end multi-task stereo depth and semantic fusion visual perception method of this application will be described in detail below with reference to specific embodiments. Figure 3 This is a flowchart illustrating the visual perception method of the end-to-end multi-task stereo depth and semantic fusion visual perception system provided in this application embodiment. Figure 3 As shown, the visual perception method of this end-to-end multi-task stereo depth and semantic fusion visual perception system may specifically include the following steps: S301, acquires synchronized image pairs by binocular imaging under a unified time base, and associates timestamps and calibration parameter identifiers with synchronized image pairs; S302, geometric correction and epipolar alignment are performed on the synchronized image pair according to the calibration parameters, noise suppression is performed, and stereo registration is completed based on correlation measurement and sub-pixel interpolation to obtain the registered image pair; S303 performs weight-shared hierarchical convolutional encoding on the left and right views of the registered image pair and forms a shared feature map through cross-layer fusion; S304 constructs a cost representation based on adaptive disparity range on the shared feature map, performs weighted aggregation on the cost representation and generates continuous disparity through learnable sub-pixel thinning, converts the continuous disparity into a depth map, generates semantic features based on multi-scale semantic parsing and decodes and outputs a semantic mask. S305: Calculate the consistency measure between depth edge and semantic boundary in the boundary neighborhood of the shared feature map, input geometric prior to semantic decoding, and use category prior to perform item-by-item weighting and incompatible item masking on the cost representation in order to jointly correct the boundary region of the depth map and semantic mask. S306 encapsulates the depth map and semantic mask along with the corresponding timestamp, calibration parameter identifier, and model version identifier according to a preset data structure, and outputs them to the real-time control interface of the external execution control system for use through the industrial real-time communication interface.
[0109] Specifically, for ease of understanding, the following will be combined with Figure 3 The flowcharts S301 to S306 illustrate the implementation process of the end-to-end multi-task stereo depth and semantic fusion visual perception method of this application step by step. Each step corresponds to each functional module of the aforementioned system and runs collaboratively under the same time base and parameter version identifier.
[0110] S301, the synchronization acquisition of unified time base and frame association system. The internal clock of FPGA is taken as the unified time base, and the trigger generator outputs the synchronization trigger signal with pulse width of 10 μs to the left and right imagers according to the set period, so as to realize the exposure at the same time. The left and right image frames collected are written with 8-byte time stamp and calibration parameter identifier in the frame header. The calibration parameter identifier is composed of the current intrinsic parameter, extrinsic parameter and distortion parameter version number and check information. If the re-calibration is triggered during running, the system only performs atomic switching between the old and new identifiers at the frame boundary to ensure the consistency of the parameters in the same frame.
[0111] S302, geometric correction, polar alignment and pre-registration based on calibration parameters. The fusion preprocessing module loads the corresponding intrinsic parameters, extrinsic parameters and distortion parameters according to the calibration parameter identifier carried by S301, and calls the multi-scale lookup table to complete the radial and tangential distortion compensation. Then, the line alignment is performed on the left and right views according to the polar constraint, and the preliminary aligned image pair is obtained. The full-frame histogram and local variance graph are calculated on the vector processing unit, and the spatial domain filtering parameters are adaptively selected to suppress noise and glare artifacts. On this basis, the phase correlation method is used to estimate the fine relative displacement of the left and right views. The integer pixel displacement is completed by shifting, and the fractional pixel displacement is completed by bilinear interpolation, so that the right view is resampled to the pixel grid consistent with the left view, and the registered image pair is obtained. The above processing is organized into a pipeline with fixed timing, and the registered image pair, time stamp and parameter version identifier are written into the ring buffer frame by frame for subsequent reading. Figure 1
[0112] S303, hierarchical convolutional coding with weight sharing and cross-layer fusion to generate shared feature maps. The shared feature extraction layer respectively encodes the left and right views of the registered image pair, and the convolutional weights are shared between the two branches to maintain the consistency of the feature space. The edge, texture and high semantic features are extracted from shallow to deep, and the spatial index is kept traceable through deterministic down-sampling and symmetric padding. In order to balance the details and semantics, the system up-samples the deep layer features and fuses them with the shallow layer features of the corresponding scale after aligning them in the channel dimension, finally forming the shared feature maps with uniform step and fixed number of channels, and generating the feature position index to record the mapping relationship from the shared feature map position to the original image coordinates. The shared feature maps and position index are written into the read-only buffer for parallel reading by the depth estimation branch and the semantic segmentation branch.
[0113] S304, parallel reasoning of depth and semantics. The depth estimation branch first determines the adaptive disparity range [dmin, dmax] dynamically at the pixel or region level based on the shared feature map and the statistics of the previous frame, aligns the left and right shared features according to the candidate disparity and calculates the per-disparity layer similarity to construct the cost representation. Combine the class prior from the multi-task coupling unit to generate attention weights, and aggregate the cost representation in spatial and disparity dimensions while applying masking constraints to left-right inconsistent or occluded positions and recording the unmatched label. The aggregation result is sent to the learnable sub-pixel refinement network, which uses inverse residual convolution and interpolation branch to restore continuous disparity, and converts it to a depth map according to D=bf / disp and the corresponding baseline b, focal length f, and quantizes it to 16-bit format. The semantic segmentation branch performs multi-scale dilated convolution context modeling on the shared feature map to generate semantic features; according to the geometric prior input by the multi-task coupling unit, it enables differential decoding paths in the boundary neighborhood and non-boundary area, refines the boundary, and keeps the large-area region stable. Finally, the semantic features are upsampled and mapped back to the original pixel grid according to the feature position index to generate an 8-bit class-encoded semantic mask. The two branches run in parallel at the same frame number, and the shared feature map and parameter identifier remain consistent within the frame.
[0114] S305, bidirectional prior coupling and boundary joint correction. The multi-task coupling unit determines the pixel set in the boundary neighborhood of the shared feature map according to the depth gradient and feature gradient, calculates the consistency measure of the edge features from the depth estimation branch and the boundary features from the semantic segmentation branch, and uses the measure as a loss term to constrain the boundary alignment of the two branches during the training phase. On the inference path, the coupling unit generates a geometric prior tensor containing planarity and normal vector from the depth estimation branch, which is input into the semantic decoding layer after size matching and fused with the feature tensor of the semantic decoding layer, thereby constraining the semantic boundary with geometric information; at the same time, it generates a class prior tensor from the semantic segmentation branch, which is mapped into a weight table for pixel and disparity combinations, and each item of the depth branch cost representation is weighted and the weight of the incompatible combination item with the class prior is set to zero. After the above bidirectional prior transmission, the system performs consistent joint correction on the boundary area of the depth map and the semantic mask, and the corrected result covers the corresponding area of the current frame.
[0115] S306, the result is packaged and output in industrial real time. The output interface module associates the depth map and semantic mask with their time stamps, calibration parameter identifiers, and model version identifiers in fields, generates a fixed-length frame header, and records the payload start offset and length, writes the frame header and payload pointer to the payload buffer. According to the interface mapping table, each field is mapped to a real-time cyclic data object, the communication period and priority are configured, and the buffer physical address is registered to the communication stack DMA descriptor queue. When the hardware interrupt arrives at the period boundary, the communication stack directly takes the array frame from the buffer for sending, and carries the CRC and frame sequence number in the sent frame for integrity and sequence control. If the calibration parameter or model version update is detected, the module carries the updated identifier in the frame boundary and performs atomic switching. If the main link timeout exceeds the threshold, switch to the standby physical link, and keep the payload buffer continuous and readable during switching. The external execution control system can seamlessly receive data.
[0116] It should be noted that the above steps are under the same time base to ensure consistent reference of parameters, models and data with ring buffer and atomic switching strategy, S301 to S306 constitute a closed-loop process from acquisition, preprocessing, feature extraction, double-branch parallel inference, prior coupling to industrial real-time output. The online update of calibration and model version can be completed without stopping upstream acquisition and downstream communication, and the consistency and timing traceability of multi-task results are maintained within the frame.
[0117] It should be understood that the size of the serial number of each step in the above method embodiment does not mean the order of execution. The execution order of each process should be determined by its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present disclosure.
[0118] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit them; although the technical solutions of the present application are described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. An end-to-end multi-task stereo depth and semantic fusion visual perception system, characterized in that, The method comprises the following steps: An image acquisition module is configured to acquire a pair of synchronized images under a unified time base through a binocular imaging unit and output image data carrying a calibration parameter identifier; A fusion preprocessing module is configured to perform geometric correction, noise suppression and stereoscopic registration on the pair of synchronized images according to the calibration parameter, and obtain a pair of registered images; A multi-task network module comprises a shared feature extraction layer, a depth estimation branch and a semantic segmentation branch; the shared feature extraction layer is configured to generate a shared feature map from the pair of registered images, the depth estimation branch is configured to construct a cost representation based on the shared feature map and output a depth map, and the semantic segmentation branch is configured to perform multi-scale semantic analysis based on the shared feature map and output a semantic mask; A multi-task coupling unit is connected to the shared feature extraction layer, the depth estimation branch and the semantic segmentation branch, configured to impose consistency constraints on shared features and branch features, and transfer geometric priors and class priors between branches; An output interface module is configured to encapsulate the depth map and the semantic mask according to a preset data structure, and output them to a real-time control interface of an external execution control system through an industrial real-time communication interface for calling.
2. The system of claim 1, wherein, The image acquisition module is configured to: Generate a synchronized trigger signal based on the unified time base and send it to a left imager and a right imager respectively; Collect a left image frame and a right image frame simultaneously under the action of the synchronized trigger signal, and write a timestamp corresponding to the unified time base for the image frames; Perform parameter calibration during assembly or operation, associate the calibration parameter identifier with the image frames and output them together with the image frames; When a preset recalibration trigger condition is met, perform recalibration and update the calibration parameter, use the last version of the calibration parameter for output during recalibration, and output the image frames together with the updated calibration parameter identifier after the recalibration is completed; Perform geometric configuration on the binocular imaging unit through a baseline adjustable mechanism and an optical axis alignment mechanism to meet the geometric relationship required for stereoscopic registration.
3. The system of claim 1, wherein, The fusion preprocessing module is configured to: Perform radial and tangential distortion compensation according to the calibration parameter based on pixel mapping of a lookup table, and complete line alignment according to the epipolar constraint, wherein the geometric correction and the parameter version identifier are associated with the image frames, and atomic switching is performed when the parameter is updated; Perform spatial domain filtering on a vector processing unit, and adaptively select filtering parameters according to image statistics; Estimate the relative displacement of left and right views based on correlation degree measurement, and align the pair of synchronized images to a unified pixel grid through sub-pixel interpolation, and output the pair of registered images; Use a preprocessing pipeline with fixed processing timing, and associate and store the pair of registered images, corresponding timestamps and the parameter version in a ring buffer structure for sequential reading by the multi-task network module.
4. The system of claim 1, wherein, The shared feature extraction layer is configured to: Perform hierarchical convolution feature coding with weight sharing on left and right views of the pair of registered images to obtain multi-scale feature representation; Fuse and unify the detail features and semantic features of different scales in a tensor format based on cross-layer connection to form the shared feature map. The shared feature map is provided to the depth estimation branch and the semantic segmentation branch simultaneously, and a feature index corresponding to a position of the shared feature map is output to be called by the multi-task coupling unit.
5. The system of claim 1, wherein, The depth estimation branch is configured to: construct a cost representation based on the shared feature map within an adaptive disparity range, the adaptive disparity range being dynamically determined at a pixel or region level according to a scene distance distribution or image statistics; weight and aggregate the class prior delivered by the shared feature map and the multi-task coupling unit with the cost representation corresponding to the attention weight, and apply a masking constraint to an unmatchable region; perform a learnable sub-pixel refinement on the aggregation result, restore continuous disparity by inverse residual convolution, and convert the continuous disparity into the depth map according to an imaging model; reuse the shared feature map cache in an inference stage and perform the semantic segmentation branch in parallel to output the depth map in alignment with a data structure of the output interface module.
6. The system of claim 1, wherein, The semantic segmentation branch is configured to: perform context modeling on the shared feature map based on a multi-scale dilated convolution array to form a semantic feature map; perform differential decoding and refinement on a semantic boundary corresponding region according to the geometric prior provided by the multi-task coupling unit; perform adaptive re-weighting according to a class distribution in a model training or updating stage, and select pixels satisfying a dynamic threshold to participate in parameter updating based on online hard sample mining; generate the semantic mask by upsampling and pixel alignment of the refined semantic feature map and output the semantic mask.
7. The system of claim 5 or 6, wherein, The multi-task coupling unit is configured to: calculate a consistency measure of edge features of the depth estimation branch and boundary features of the semantic segmentation branch in a boundary neighborhood of the shared feature map, and use the consistency measure as a training constraint term to constrain boundary alignment of the two branches; generate a geometric prior tensor containing planarity and normal vector information from the depth estimation branch, input the geometric prior tensor into a decoding layer of the semantic segmentation branch, and fuse the geometric prior tensor with a feature tensor of the decoding layer to form a geometric-constrained semantic feature; map a class prior tensor generated from the semantic segmentation branch to a weight of a pixel and disparity combination, weight each item of the cost representation of the depth estimation branch, and set a weight of a pixel and disparity combination item incompatible with the class prior to zero; jointly optimize the training constraint term, the geometric prior fusion term, and the class prior weighting term in a training stage, and perform the fusion and weighting according to fixed parameters obtained by training in an inference stage.
8. The system of claim 1, wherein, The output interface module is configured to: associate the depth map and the semantic mask with corresponding time stamps, calibration parameter identifiers, and model version identifiers in fields, and frame write into a payload buffer according to a preset data structure; map the fields in the payload buffer into real-time cyclic data objects according to an interface mapping table, configure a communication period and a priority, and establish a corresponding relationship with the real-time control interface; associate the payload buffer with the industrial real-time communication interface, and send according to the communication period under a hardware interrupt trigger. Carrying check code and frame sequence number in sending frame for sequence and integrity control, and carrying updated identification with frame and performing atomic switching at frame boundary when parameter or model version is updated; Switching sending path to standby physical link when detecting that main link meets timeout threshold, and keeping continuous readability of the load buffer during switching.
9. The system of claim 1, wherein, The system further comprises a power supply and alarm module for: Receiving a direct current input power supply and generating multiple mutually electrically isolated stabilized outputs; Monitoring voltage, current and temperature of the direct current input and each stabilized output, determining overvoltage, overcurrent, undervoltage or overtemperature according to preset threshold, and outputting alarm signals to the multi-task network module and the output interface module; When detecting that the voltage of the input power supply drops to a threshold, driving an energy buffer unit to maintain short-time power supply, outputting an abnormal identification to the output interface module, and issuing a controlled shutdown signal to the image acquisition module, the fusion preprocessing module and the multi-task network module; Providing a good power supply signal and a reset control signal to establish the power-on and power-off timing of the image acquisition module, the fusion preprocessing module, the multi-task network module and the output interface module.
10. A visual perception method based on the end-to-end multi-task stereo depth and semantic fusion visual perception system according to any one of claims 1 to 9, characterized in that, Comprise: Acquiring a pair of synchronized images by binocular imaging under a unified time base, and associating a timestamp and a calibration parameter identification for the pair of synchronized images; Performing geometric correction and epipolar alignment on the pair of synchronized images according to the calibration parameters, executing noise suppression, and completing stereo registration based on correlation metrics and sub-pixel interpolation to obtain a pair of registered images; Performing hierarchical convolutional coding with weight sharing on left and right views of the pair of registered images, and forming a shared feature map through cross-layer fusion; Based on adaptive disparity range, constructing a cost representation on the shared feature map, performing weighted aggregation on the cost representation, and generating continuous disparity through learnable sub-pixel refinement to generate a depth map, generating semantic features based on multi-scale semantic parsing, and decoding and outputting a semantic mask; Calculating a consistency metric of depth edges and semantic boundaries in the boundary neighborhood of the shared feature map, inputting geometric priors to semantic decoding, and using class priors to perform item-by-item weighting and incompatible item shielding on the cost representation to jointly correct the boundary regions of the depth map and the semantic mask; The depth map and the semantic mask are encapsulated with the corresponding timestamp, calibration parameter identification and model version identification according to a preset data structure, and output to the real-time control interface of the external execution control system through the industrial real-time communication interface for calling.
Citation Information
Patent Citations
Real-time binocular depth estimation method and device based on semantic feature fusion
CN116630391A
Semantic segmentation and stereo matching method and framework based on multi-task joint learning
CN117710453A
Robot three-dimensional environment sensing method and device based on deep visual learning
CN120564156A
Systems and methods for depth estimation using semantic features
US20210090280A1
Cited By
Three-dimensional map surveying and mapping method, device and equipment based on unmanned aerial vehicle and medium
CN121655470A