Visual navigation method and equipment of medical body-equipped robot and medium
By preprocessing, aligning and weighting fusion of multimodal data, combined with neural architecture search and delay regularization optimization of neural network architecture, it solves multiple challenges of visual navigation of medical embodied robots in medical environments, and achieves efficient and accurate visual navigation capabilities.
Patent Information
- Application Number
- CN202510533520.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Medical embossed robots face problems such as multimodal data processing, complex environmental adaptability, real-time and resource limitations, as well as high accuracy and robustness requirements in medical environments, making it difficult for traditional visual navigation technology to meet the high standards of medical applications.
By obtaining multimodal data for preprocessing, alignment and weighting fusion, a unified feature vector is generated; an initial model is built and the neural network architecture is generated for visual navigation through neural architecture search and delay regularization optimization; finally, a unified feature vector is input into the optimized model to realize visual navigation of medical embossed robots.
It realizes efficient and accurate visual navigation in multimodal and dynamic environments, reduces the computing burden, enhances adaptability to complex medical scenarios, and ensures stable path navigation and rapid response of medical embodied robots.
Smart Images

Figure CN120036934A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical embodied robots, and specifically relates to a visual navigation method, device, and medium for a medical embodied robot. Background Art
[0002] The field of embodied intelligence has developed rapidly, and the application of embodied intelligent robots in medicine has gradually increased. Especially in scenarios such as surgical assistance, precise diagnosis, rehabilitation care, and telemedicine, such robots have played an important role. Medical embodied robots need to have high-precision visual navigation capabilities to achieve autonomous movement, environmental perception, and task execution. For example, in the operating room environment, the robot not only needs to accurately identify the position information of surgical tools and organs, but also needs to adjust the operation path in real time to avoid interference with doctors and other equipment, ensuring the safety and efficiency of the surgical process. Therefore, the medical environment poses higher requirements for the visual navigation of robots in terms of accuracy, real-time performance, and safety.
[0003] Traditional visual navigation technologies face the following main challenges in medical scenarios: Multi-modal data processing: The medical environment involves different types of data (such as real-time images, CT / MRI scans, and patient physiological data). The effective fusion and real-time processing of these data are crucial for the navigation and diagnostic capabilities of the robot. In the processing of medical images, especially CT and MRI scans, it is necessary to fuse multi-modal data to improve the accuracy of diagnosis.
[0004] Complex environment adaptability: There are various interference factors in the medical scenario, such as light changes, occlusions, and dynamic objects (such as the movement of medical staff), which will affect the visual navigation effect of the robot.
[0005] Real-time performance and resource limitations: Medical embodied robots need to respond quickly, and visual navigation tasks require high computing resources. Therefore, there is an urgent need for an efficient and lightweight neural architecture.
[0006] High-precision and robustness requirements: The error tolerance in medical application scenarios is extremely low, and the robot needs to have high-precision and robust visual navigation capabilities to ensure safety and effectiveness.
[0007] In summary, there is an urgent need for a visual navigation method, device, and medium for a medical embodied robot to solve the problems in the prior art. Summary of the Invention
[0008] The purpose of the present invention is to provide a visual navigation method, device, and medium for a medical embodied robot. The specific technical solutions are as follows: A visual navigation method for a medical embodied robot includes the following steps: S1. Obtain a unified feature vector: Obtain multimodal data, perform preprocessing, alignment, and weighted fusion on the multimodal data to obtain a unified feature vector; S2. Construct an initial model: Use neural architecture search to generate a neural network architecture for visual navigation. On the basis of the neural network architecture, add lightweight convolution and attention mechanisms to obtain an initial model; S3. Optimize the initial model: Construct an optimization objective function based on latency regularization, and optimize the initial model based on the optimization objective function to obtain a visual navigation model; S4. Complete visual navigation: Input the unified feature vector into the visual navigation model to achieve visual navigation of the medical embodied robot.
[0009] Optionally, in S1, the multimodal data includes real-time camera images, medical images, and patient physiological data.
[0010] Optionally, in S1, perform preprocessing on the multimodal data. The preprocessing includes adaptive histogram equalization and median filtering. The adaptive histogram equalization is used to enhance contrast, and the median filtering is used to denoise.
[0011] Optionally, in S1, perform alignment on the multimodal data, including spatial alignment and temporal alignment; The steps of spatial alignment are as follows: Feature point extraction: Extract key feature points from real-time camera images and medical images; for real-time camera images, use an edge detection algorithm to find recognizable boundaries or reference points in the environment as key feature points; for medical images, use an edge detection algorithm or a feature extraction algorithm to extract anatomical landmark points as key feature points; Feature point matching: Use a matching algorithm to match the key feature points in real-time camera images and medical images to obtain pairs of matched key feature points; Estimate the spatial transformation matrix: Use the matched pairs of key feature points to calculate a spatial transformation matrix applicable to the two modalities, and estimate the spatial transformation matrix parameters by the least squares method; Image resampling and registration: According to the calculated spatial transformation matrix, resample the medical image so that the medical image is presented in the same coordinate system as the real-time camera image to complete spatial alignment; The steps of temporal alignment are as follows: Determine the sampling frequency: Obtain the sampling frequencies of each modality data; Timestamp alignment: Label each modality data with an accurate timestamp; Interpolation processing: Apply an interpolation algorithm to low-frequency data to make the time step of low-frequency data the same as that of high-frequency data; Time synchronization processing: For high-frequency data, downsample according to the sampling period of low-frequency data to ensure that the time points of each modal data correspond one by one and complete time alignment.
[0012] Optionally, in S1, perform weighted fusion on multi-modal data, and the expression is as follows: ; Among them, represents the unified feature vector, represents the aligned real-time camera image, represents the weight coefficient of the real-time camera image, represents the aligned medical image, represents the weight coefficient of the medical image, represents the aligned patient physiological data, represents the weight coefficient of the patient physiological data; , , are adaptively adjusted through model training.
[0013] Optionally, in S2, the initial model includes a feature extraction layer, an attention mechanism, a multi-task adaptation layer, and a path planning module set in sequence; The feature extraction layer is a depthwise separable convolutional layer for feature extraction; The attention mechanism is a channel-level attention mechanism for processing the extracted features; The multi-task adaptation layer includes multiple branch networks, and each branch network corresponds to a different type of task; The path planning module outputs path planning based on the unified feature vector to complete visual navigation.
[0014] Optionally, in S3, the expression of the optimization objective function is as follows: ; Among them, represents the error between the model prediction value and the true label value , representing the loss function; represents the prediction output of the model for the input under the parameters and the neural network architecture ; represents the input data of the sample, represents the total delay of the neural network architecture, represents the neural network architecture, represents the total number of samples, represents the regularization coefficient.
[0015] Optionally, the total latency of the neural network architecture in S3 The expression is as follows: ; in, Indicates candidate operations The average execution time of Indicates candidate operations In the neural network architecture The probability of selection under It is calculated by the softmax function, and the calculation expression is as follows: ; in, Indicates the candidate operation in the architecture parameter The relevant weights, Represents the normalization factor, which is used to normalize the probability and the sum of candidate operations.
[0016] Additionally, a computer device includes a memory and a processor; The memory is used to store a computer program executable on the processor; The processor is used to implement the steps of the visual navigation method as described above when executing the computer program.
[0017] In addition, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the visual navigation method described above are implemented.
[0018] The application of the technical solution of the present invention has the following beneficial effects: (1) The present invention discloses a visual navigation method for a medical embodied robot. Unlike the convolutional neural network (CNN) which has a fixed structure and is difficult to cope with dynamic and multimodal scenarios, the method of the present invention ensures real-time geometric and temporal consistency by spatially aligning and temporally aligning multimodal data. In addition, the method of the present invention fuses multimodal data through an adaptive weighting mechanism, further enhancing the adaptability to dynamic changes, enabling the medical embodied robot to effectively extract key features from multi-scale and multimodal data, and more accurately adapt to the real-time needs of complex medical scenarios.
[0019] (2) Traditional multimodal data fusion methods have high computational complexity and are prone to misjudgment in real-time applications. The method of the present invention discloses spatial alignment and temporal alignment strategies to accurately synchronize data from each modality, reduce the computational burden, and shorten data processing time. The method of the present invention can effectively avoid feature conflicts and ensure feature fusion and rapid response in complex dynamic environments.
[0020] (3)The method of the present invention discloses an optimization objective function based on latency regularization, selects a low-latency and efficient neural network architecture through latency optimization control, and reduces the computational overhead of the model. Compared with traditional neural architecture search (NAS), the method of the present invention reduces latency without compromising performance, ensuring that in a multi-task real-time medical scenario, a medical embodied robot can achieve stable path navigation and fast response.
[0021] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the drawings to further elaborate on the present invention in detail. Description of the Drawings
[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a flowchart of the steps of the visual navigation method in the preferred embodiment of the present invention. Detailed Embodiments
[0024] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the following will further elaborate on the present invention in detail in conjunction with the drawings and specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0025] As Figure 1 shown, this embodiment provides a visual navigation method for a medical embodied robot, including the following steps: S1. Obtain a unified feature vector: Obtain multi-modal data, perform preprocessing, alignment, and weighted fusion on the multi-modal data to obtain a unified feature vector; S2. Construct an initial model: Use neural architecture search to generate a neural network architecture for visual navigation, and on the basis of the neural network architecture, add lightweight convolution and attention mechanisms to obtain an initial model; S3. Optimize the initial model: Construct an optimization objective function based on latency regularization, and optimize the initial model based on the optimization objective function to obtain a visual navigation model; S4. Complete visual navigation: Input the unified feature vector into the visual navigation model to achieve visual navigation of the medical embodied robot.
[0026] Specifically, in S1, the multimodal data includes real-time camera images, medical images, and patient physiological data. In this embodiment, the real-time camera images are used for obstacle detection and path navigation to help the robot plan a safe path in a dynamic environment; the medical images include CT images or MRI images, which are used to identify specific parts such as target organs, lesions, and surgical instruments to improve the accuracy of diagnosis and navigation; the patient physiological data includes heart rate, blood pressure, and respiratory rate, etc., which are used as auxiliary judgments to help the robot identify the patient's current condition and make adaptive adjustments.
[0027] Specifically, in S1, the multimodal data is preprocessed. The preprocessing includes adaptive histogram equalization and median filtering. The adaptive histogram equalization is used to enhance the contrast, and the median filtering is used to denoise.
[0028] Specifically, the method of adaptive histogram equalization (CLAHE) is as follows: Since medical images usually have low contrast, CLAHE is used to enhance the contrast of the images to make the details clearer. When processing each local area, CLAHE limits the improvement of the contrast to prevent global over-enhancement from causing image noise. The CLAHE enhancement formula is as follows: ; where and respectively represent the minimum gray value and the maximum gray value of the original image, represents the original pixel value, represents the pixel value after image enhancement, represents the minimum pixel value.
[0029] In this embodiment, CLAHE can enhance the details of low-contrast regions while retaining the overall features of the image, which is helpful for the accuracy in subsequent recognition tasks.
[0030] Optionally, in S1, the multimodal data is aligned, including spatial alignment and temporal alignment; The steps of spatial alignment are as follows: Feature point extraction: Extract key feature points from the real-time camera images and medical images; for the real-time camera images, use an edge detection algorithm to find recognizable boundaries or reference points in the environment as key feature points; for medical images, use an edge detection algorithm or a feature extraction algorithm to extract anatomical landmark points as key feature points; Feature point matching: Use a matching algorithm to match the key feature points in the real-time camera images and medical images to obtain pairs of matched key feature points; Estimate the spatial transformation matrix: Using the paired key feature points that are matched, calculate the spatial transformation matrix applicable to the two modalities, and estimate the parameters of the spatial transformation matrix by the least squares method; Image resampling and registration: According to the calculated spatial transformation matrix, resample the medical image so that the medical image is presented in the same coordinate system as the real-time camera image, and complete the spatial alignment; The steps for time alignment are as follows: Determine the sampling frequency: Obtain the sampling frequencies of the data of each modality; Time stamp alignment: Label each modality data with an accurate time stamp; Interpolation processing: Apply an interpolation algorithm to the low-frequency data (such as patient physiological data) to make the time step of the low-frequency data the same as that of the high-frequency data (such as the real-time camera image); The interpolation method in this embodiment includes at least one of the following: Linear interpolation: Suitable for relatively small differences in sampling frequencies; Spline Interpolation: When the data sampling is sparse and high-precision alignment is required, cubic spline interpolation can be used to smooth the data; The interpolation formula for physiological data is: ; where and are two adjacent time points, is the data value after interpolation.
[0031] Time synchronization processing: For high-frequency data, downsample according to the sampling period of the low-frequency data to ensure that the time points of the data of each modality correspond one by one, and complete the time alignment.
[0032] It should be noted that this embodiment also includes verifying and optimizing the alignment result after spatial alignment or time alignment. Among them, the verification of spatial alignment is performed by superimposing the two images for inspection, and the verification of time alignment is performed by checking the time alignment effect of the data after interpolation.
[0033] Optionally, in S1, weighted fusion of multi-modal data is performed, and the expression is as follows: ; where represents the unified feature vector, represents the real-time camera image after alignment, represents the weight coefficient of the real-time camera image, represents the medical image after alignment, represents the weight coefficient of the medical image, Represents the aligned patient physiological data, represents the weight coefficient of the patient physiological data; , , and is adaptively adjusted through model training.
[0034] Optionally, in S2, the initial model includes a feature extraction layer, an attention mechanism, a multi-task adaptation layer, and a path planning module arranged in sequence; The feature extraction layer is a depthwise separable convolution layer for feature extraction; It should be noted that for a depthwise separable convolution operation, its calculation can be divided into two steps: depth convolution and pointwise convolution: Depth convolution: Each channel independently performs convolution to extract spatial features without cross-channel operation. The depth convolution calculation formula is: ; where: represents the output feature map after depth convolution, represents the depth convolution kernel, which operates independently for each channel; represents the pixel value at position in the input feature map; Pointwise Convolution: Uses 1×1 convolution operation to perform convolution across channels for fusing information from different channels. The pointwise convolution calculation formula is: ; where, represents the output feature map after pointwise convolution, is the pointwise convolution kernel, which is used to fuse the features of each channel of the depth convolution. Compared with traditional convolution, the computational complexity of depthwise separable convolution is significantly reduced.
[0035] The attention mechanism is a channel-level attention mechanism (SE module) for processing the extracted features; The multi-task adaptation layer includes multiple branch networks, and each branch network corresponds to a different type of task. In this embodiment, the branch networks include an instrument recognition branch, a lesion localization branch, a path planning branch, etc. This embodiment can further divide the branch networks according to the usage scenario. By assigning different tasks to different branch networks, the differentiated requirements in the medical scenario are realized.
[0036] It should be noted that in this embodiment, the path planning branch uses deep reinforcement learning (DRL) to continuously interact with the environment and perform feedback learning, enabling the robot to gradually master the optimal path planning strategy. In DRL, the path planning of the robot can be represented as a Markov decision, and the settings of the state, action, and reward are as follows: State s: Represents the current state of the robot in the surgical environment, including information such as position, direction, the positions of obstacles in the environment, and the target position.
[0037] Action a: Represents the moving directions or behaviors that the robot can choose in state s, such as forward, left, right, etc.
[0038] Reward r: The robot will receive a reward every time it completes an action. The size of the reward depends on whether the action helps to approach the target. For example, if the robot successfully avoids obstacles and moves forward in the direction of the target, the reward is positive; if it hits an obstacle, the reward is negative.
[0039] The path planning module outputs path planning based on the unified feature vector to complete visual navigation.
[0040] Optionally, in S3, the expression of the optimization objective function is as follows: ; Where, Represents the model prediction value And the true label value The error between them represents the loss function; Represents the prediction output of the model for the input Under the parameters And the neural network architecture ; Represents the input data of the sample, Represents the total latency of the neural network architecture, Represents the neural network architecture, Represents the total number of samples, Represents the regularization coefficient.
[0041] Optionally, in S3, the expression of the total latency Of the neural network architecture is as follows: ; Where, Represents the average execution time of the candidate operation ; Represents the selection probability of the candidate operation Under the neural network architecture , which is calculated through the softmax function, and the calculation expression is as follows: ; Among them, represents the weight related to the candidate operation in the architecture parameters, represents the normalization factor used to normalize the probabilities of candidate operations.
[0042] Furthermore, in this embodiment, the DARTS (Differentiable Architecture Search) search space is preferably adopted, and its search space setting and architecture parameter optimization are as follows: The search space of DARTS contains multiple candidate operations (such as convolution, pooling, skip connection, etc.), and different combinations of these operations will affect the accuracy and latency of the model. The optimization of architecture parameters is achieved by continuousizing the search space, and this process is divided into multiple steps to ensure that the generated architecture meets the adaptability in a multi-task environment.
[0043] Search space design: Convolution operation: including different convolution kernel sizes and strides.
[0044] 3x3 convolution: suitable for fine feature extraction and helpful for tasks such as lesion detection.
[0045] 5x5 convolution: used for feature extraction with a larger receptive field and suitable for the recognition of global features.
[0046] Pooling operation: including max pooling and average pooling.
[0047] Max pooling: strengthens the highlighted parts in the feature map and is suitable for highlighting important features in lesion detection tasks.
[0048] Average pooling: smooths the feature map, reduces noise, and helps in the extraction of background information.
[0049] Stride selection: stride 1 maintains the original resolution, and stride 2 is used for downsampling to reduce the computational amount.
[0050] Neural network architecture parameter optimization: DARTS controls the operation selection through architecture parameters. The architecture parameters of each layer are initially initialized to a uniform distribution and are gradually optimized during the search process. The specific steps are as follows: 1) Initialize architecture parameters: At the beginning of the search, randomly initialize the architecture parameter α of each candidate operation.
[0051] 2) Calculate the operation selection probability: For each node, DARTS calculates the selection probability of the operation through the softmax function: ; 3) Continuousization of mixed operations: DARTS continuousizes the discrete search space through mixed operations. The mixed operation formula is as follows: ; Among them, represents the operation on the input output; represents the output of the mixed operation between the and th nodes, including the weighted sum of all candidate operations. represents the probability of selecting the operation .
[0052] Gradient optimization: DARTS adopts a dual optimization process to optimize the weight parameters and architecture parameters respectively.
[0053] Model weight optimization: Use the standard gradient descent algorithm to update the model weights to reduce the loss function.
[0054] Architecture parameter optimization: Calculate the gradient of the architecture parameters through chain derivation and update the architecture parameters to select a better operation combination.
[0055] In addition, this embodiment also provides a computer device, including a memory and a processor; The memory is used to store a computer program that can run on the processor; The processor is used to implement the steps of the above visual navigation method when executing the computer program.
[0056] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments that can complete specific functions, and this instruction segment is used to describe the execution process of the computer program in the computer device.
[0057] The computer device can be a computing device such as a mobile phone, a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor and a memory. For example, the computer device may further include input / output devices, network access devices, a bus, etc.
[0058] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer device and connects all parts of the computer device through various interfaces and lines.
[0059] The memory can be used to store the computer program and / or module. The processor realizes the computer program by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0060] Among them, if the modules / units integrated in the computer device are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0061] In addition, an embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-described visual navigation method are implemented.
[0062] The method of this embodiment discloses a visual navigation method for a medical embodied robot. Different from the fixed structure of the convolutional neural network (CNN) which is difficult to handle dynamic and multimodal scenarios, the method of this embodiment ensures real-time geometric and temporal consistency by spatially aligning and temporally aligning multimodal data. In addition, the method of this embodiment fuses multimodal data through an adaptive weighting mechanism, further enhancing the adaptability to dynamic changes, enabling the medical embodied robot to effectively extract key features from multi-scale and multimodal data and more precisely adapt to the real-time requirements of complex medical scenarios.
[0063] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0064] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A visual navigation method for a medical embodied robot, characterized in that: The steps include: S1. Obtaining a unified feature vector: obtaining multimodal data, preprocessing, aligning and weighted fusion of the multimodal data, and obtaining a unified feature vector; S2. Build the initial model: Generate a neural network architecture for visual navigation using neural architecture search. Add lightweight convolution and attention mechanisms to the neural network architecture to get the initial model. S3. Optimize the initial model: construct an optimization objective function based on delay regularization, optimize the initial model based on the optimization objective function, and obtain a visual navigation model; S4. Complete visual navigation: input the unified feature vector into the visual navigation model to realize the visual navigation of the medical embodied robot.
2. The visual navigation method according to claim 1, characterized in that: In S1, the multimodal data includes real-time camera images, medical images and patient physiological data.
3. The visual navigation method according to claim 2, characterized in that: In S1, the multimodal data is preprocessed, and the preprocessing includes adaptive histogram equalization and median filtering. The adaptive histogram equalization is used to enhance contrast, and the median filtering is used to remove noise.
4. The visual navigation method according to claim 3, characterized in that: In S1, the multimodal data are aligned, including spatial alignment and temporal alignment; The steps for spatial alignment are as follows: Feature point extraction: Extract key feature points from real-time camera images and medical images. For real-time camera images, edge detection algorithms are used to find recognizable boundaries or reference points in the environment as key feature points. For medical images, edge detection algorithms or feature extraction algorithms are used to extract anatomical landmarks as key feature points. Feature point matching: Use matching algorithms to match key feature points in real-time camera images and medical images to obtain matched key feature point pairs; Estimate the spatial transformation matrix: Use the matched key feature point pairs to calculate the spatial transformation matrix applicable to the two modes, and estimate the spatial transformation matrix parameters by the least squares method; Image resampling and registration: Resample the medical image based on the calculated spatial transformation matrix so that the medical image is presented in the same coordinate system as the real-time camera image to complete spatial alignment; The steps for time alignment are as follows: Determine the sampling frequency: obtain the sampling frequency of each modal data; Timestamp alignment: annotate each modality data with an accurate timestamp; Interpolation processing: Apply interpolation algorithms to low-frequency data to make the time step of low-frequency data the same as that of high-frequency data; Time synchronization processing: For high-frequency data, down-sampling is performed according to the sampling period of low-frequency data to ensure that the time points of each modal data correspond one to one and complete time alignment.
5. The visual navigation method according to claim 4, characterized in that: In S1, multimodal data is weighted fused, and the expression is as follows: ; in, represents the unified eigenvector, represents the aligned real-time camera image, represents the weight coefficient of the real-time camera image, represents the aligned medical image, represents the weight coefficient of medical imaging, represents the aligned patient physiological data, Represents the weight coefficient of the patient's physiological data; , , Adaptive adjustment through model training.
6. The visual navigation method according to claim 5, characterized in that: In S2, the initial model includes the feature extraction layer, attention mechanism, multi-task adaptation layer and path planning module which are set in sequence; The feature extraction layer is a depth-separable convolutional layer, which is used for feature extraction; The attention mechanism is a channel-level attention mechanism, which is used to process the extracted features; The multi-task adaptation layer includes multiple branch networks, each branch network corresponds to a different type of task; The path planning module outputs path planning based on the unified feature vector to complete visual navigation.
7. The visual navigation method according to claim 6, characterized in that: In S3, the expression of the optimization objective function is as follows: ; in, Represents the model prediction value and the true label value The error represents the loss function; Indicates that the model has parameters and neural network architecture Next pair input The predicted output of represents the input data of the sample, represents the total latency of the neural network architecture, represents the neural network architecture, represents the total number of samples, represents the regularization coefficient.
8. The visual navigation method according to claim 7, characterized in that: Total latency of neural network architectures in S3 The expression is as follows: ; in, Indicates candidate operations The average execution time of Indicates candidate operations In the neural network architecture The probability of selection under It is calculated by the softmax function, and the calculation expression is as follows: ; in, Indicates the candidate operation in the architecture parameter The relevant weights, Represents the normalization factor, which is used to normalize the probability and the sum of candidate operations.
9. A computer device, characterized in that: including memory and processor; The memory is used to store a computer program executable on the processor; The processor is configured to implement the steps of the visual navigation method according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the visual navigation method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Kinect-based space teleoperation robot control system and method thereof
CN103302668A
Time delay calculation method for navigation positioning equipment and three-dimensional perspective equipment
CN115844533A
Artificial intelligence-based thyroid nodule surgical operation planning and navigation system
CN119655879A
Teleoperation control method and system for alimentary canal catheterization robot
CN119679520A
Medical image recognition processing system and method based on multi-modal image fusion
CN119887773A
Cited By
Unmanned aerial vehicle body cognition alignment method based on man-machine cooperation
CN120803003A
Space-time trajectory splicing method and device of robot, electronic equipment and medium
CN121670621A
Methods, devices, electronic equipment and media for spatiotemporal trajectory stitching of robots
CN121670621B