A space station in-station equipment interaction method and system based on scene semantic segmentation
Patent Information
- Application Number
- CN202510316725.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-03-18
AI Technical Summary
[0006]为了解决现有技术存在的场景语义分割感受野覆盖不够密集,难以捕捉连续多尺度上下文,且复杂场景下易受背景噪声干扰的技术问题,本发明实施例提供了一种基于场景语义分割的空间站站内设备交互方法及系统
[0019] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
Smart Images

Figure CN120495646B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a method and system for interaction of equipment within a space station based on scene semantic segmentation. Background Technology
[0002] The confined and cramped nature of the space station means that air circulation relies on a ventilation system. Astronauts' breathing, sweat, and moisture generated by equipment operation make equipment prone to oxidation and mold growth, seriously threatening equipment performance and lifespan, affecting the safe and stable operation of the space station, as well as the health and work efficiency of astronauts. To ensure the long-term reliable operation of the space station, astronauts need to perform regular maintenance on the equipment. New methods for interacting with space station equipment have become a current research hotspot. Early equipment inspection was based on traditional image processing methods, which typically used computer vision and image processing techniques such as edge detection, color segmentation, and shape matching.
[0003] In recent years, deep learning technology has made significant progress in device recognition. Deep learning models can learn complex feature representations from large amounts of data, improving the accuracy and robustness of device recognition. This study uses a DeepLabV3+ network for semantic segmentation of the target device, employs a Visual Background Extractor (ViBe) combined with grayscale histograms for bi-Gaussian fitting, and uses a RetinaNet network for object detection.
[0004] DeepLabV3+, as a classic semantic segmentation network, has achieved remarkable results in complex scene segmentation tasks by introducing atrous spatial pyramid pooling (ASPP) modules and encoder-decoder structures. However, it still has some limitations. The original DeepLabV3+ network has limited ability to fuse contextual information of multi-scale features, especially when restoring edge details and segmenting small objects, where information loss is prone to occur. The backbone network based on the Xception neural network architecture or Residual Neural Network (ResNet) architecture has a large number of parameters, making it difficult to deploy on mobile devices or resource-constrained devices. The dilated convolution of the traditional atrous spatial pyramid pooling (ASPP) uses discrete interval sampling, resulting in insufficient receptive field coverage and difficulty in capturing continuous multi-scale context. The importance of feature channels and spatial locations is not dynamically weighted, making it susceptible to background noise interference in complex scenes.
[0005] In existing technologies, there is a lack of an accurate and efficient method for interaction between devices within a space station based on scene semantic segmentation. Summary of the Invention
[0006] To address the technical problems of insufficient receptive field coverage in existing scene semantic segmentation technologies, making it difficult to capture continuous multi-scale context and susceptible to background noise interference in complex scenes, this invention provides a method and system for interaction of equipment within a space station based on scene semantic segmentation. The technical solution is as follows:
[0007] On the one hand, a method for interaction between devices within a space station based on scene semantic segmentation is provided. This method is implemented by devices interacting within the space station and includes:
[0008] Collect images of equipment inside the space station to obtain training data;
[0009] A semantic segmentation model for the image to be trained is constructed based on the model structure of the DeepLabv3+ network.
[0010] Based on the image segmentation loss function, the training data is used to train the image semantic segmentation model to obtain the image semantic segmentation model.
[0011] Acquire the device image of the current scene; input the device image of the current scene into the image semantic segmentation model to obtain the device segmentation image;
[0012] Obtain interaction requirements; based on the space station equipment database, retrieve space station equipment information in the UI interface according to the interaction requirements and equipment segmentation images.
[0013] On the other hand, a space station in-station equipment interaction system based on scene semantic segmentation is provided. This system is applied to the space station in-station equipment interaction method based on scene semantic segmentation. The system includes a camera, electronic devices, and a touch screen, wherein:
[0014] The camera is used to acquire images of the device in the current scene;
[0015] The electronic device is used to collect images of equipment inside the space station and obtain training data; to construct a semantic segmentation model for the image to be trained based on the model structure of the DeepLabv3+ network; to train the semantic segmentation model for the image to be trained using the training data based on the image segmentation loss function, and to obtain the image semantic segmentation model; and to input the current scene equipment image into the image semantic segmentation model to obtain the equipment segmentation image.
[0016] The touchscreen is used to acquire interaction requests; based on the space station equipment database, and according to the interaction requests and equipment segmentation images, it retrieves space station equipment information in the UI interface.
[0017] On the other hand, a space station intra-station equipment interaction device is provided, the space station intra-station equipment interaction device comprising: a processor; a memory, the memory storing computer-readable instructions, the computer-readable instructions being executed by the processor to implement any of the above-described space station intra-station equipment interaction methods based on scene semantic segmentation.
[0018] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement any of the above-described methods for interaction of equipment within a space station based on scene semantic segmentation.
[0019] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0020] This invention proposes a method for interaction between devices within a space station based on scene semantic segmentation. A channel-spatial dual attention module is introduced into the encoder part of the DeepLabV3+ network to improve the model's robustness to complex scenes. MobileNetV2 is used to replace the original backbone network Xception, utilizing its inverted residual structure and linear bottleneck layer to reduce the number of parameters and computational cost. The original ASPP module is replaced with DenseASPP, and by stacking densely connected dilated convolutional layers, a denser and wider receptive field is constructed, enhancing the ability to model the context of small targets and multi-scale objects. This invention provides an accurate and efficient method for interaction between devices within a space station based on scene semantic segmentation. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a space station intra-station equipment interaction method based on scene semantic segmentation provided by an embodiment of the present invention;
[0023] Figure 2 This is a block diagram of a space station intra-station equipment interaction system based on scene semantic segmentation provided in an embodiment of the present invention;
[0024] Figure 3 This is a schematic diagram of the structure of an interactive device for equipment within a space station, provided in an embodiment of the present invention. Detailed Implementation
[0025] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0026] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0027] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0028] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0029] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0030] This invention provides a method for interaction between devices within a space station based on scene semantic segmentation. This method can be implemented by an interaction device within the space station, which can be a terminal or a server. Figure 1 The flowchart shown is for a space station intra-site equipment interaction method based on scene semantic segmentation. The processing flow of this method may include the following steps:
[0031] S1. Collect images of equipment inside the space station to obtain training data.
[0032] In one feasible implementation, this invention collects a large number of images of instruments and components within the space station and annotates each component with semantic tags. Annotation data is generated using tools such as Labelme, in common formats such as Pascal VOC. Basic image operations are implemented using open-source computer vision libraries (OpenCV). The acquired images are preprocessed, including image cropping, scaling, and denoising, to improve image quality. The training data images have a resolution of 3840×2160.
[0033] S2. Construct a semantic segmentation model for the image to be trained based on the model structure of the DeepLabv3+ network.
[0034] In one feasible implementation, a deep learning model for semantic segmentation (DeepLab Version 3 Plus, DeeplabV3+) introduces a novel Encoder-Decoder architecture. The core of its encoder is a deep convolutional neural network (DCNN) with dilated convolutions, which can employ commonly used classification networks. At the end of the encoder part, DeeplabV3+ introduces an atrous spatial pyramid pooling (ASPP) module with dilated convolutions to achieve segmentation capabilities for multi-scale objects.
[0035] ASPP performs 1x1 convolutions, 3x3 convolutions with dilation rates of 6, 12, and 18 on the input feature maps, followed by global average pooling. The feature maps are then fused and compressed into 256 channels by a 1x1 convolution. ASPP can extract and distinguish feature information from targets at different scales, effectively segmenting multi-scale targets. The goal of this module is to capture contextual information at different scales to better understand the semantic information of the image.
[0036] Compared to DeepLabv3, DeepLabv3+ introduces a Decoder module, which performs necessary downsampling operations on the input image, further fusing low-level and high-level features. During feature map restoration, low-level features are fused to restore the boundary information of the target part. The feature map restoration adopts a linear interpolation method, which ultimately improves the accuracy of network segmentation.
[0037] S3. Based on the image segmentation loss function, use the training data to train the image semantic segmentation model to obtain the image semantic segmentation model.
[0038] The image semantic segmentation model includes an encoder and a decoder.
[0039] The backbone network of the encoder is the lightweight convolutional neural network MobileNetV2;
[0040] The multi-scale feature extraction module of the encoder is the DenseASPP module with dense void spatial pyramid pooling;
[0041] The image semantic segmentation model employs a CBAM module that combines channel attention and spatial attention mechanisms.
[0042] In one feasible implementation, the classic DeeplabV3+ network uses the Xception network as its backbone. The Xception model has a relatively complex architecture, powerful expressive capabilities, and is suitable for complex tasks and large-scale datasets. However, due to its relatively large model size, it requires more computational resources and storage space, and the training time is usually long, making it less suitable for resource-constrained devices.
[0043] Therefore, the classic DeeplabV3+ network structure was improved by replacing the backbone network with the lightweight convolutional neural network MobileNetV2. Compared to the Xception network, MobileNetV2 has a lower model size and computational complexity, making it suitable for mobile devices and embedded systems. Due to its lightweight design, MobileNetV2 also boasts faster inference speeds.
[0044] MobileNet primarily replaces ordinary convolutions with depthwise convolutions and introduces two hyperparameters: the width factor and the resolution factor. This allows for flexible control over the size of the network model. Through optimization of the network structure, it achieves higher accuracy than most neural networks with fewer parameters and less computation.
[0045] MobileNetv2 adds two important modules to its foundation, introducing a linear bottleneck relationship between network layers; the bottleneck blocks are connected using residual connections. The design of the bottleneck blocks enables the effective encoding of feature information at both the input and output ends of the model, while the inner layers of the network encapsulate the transformation of information from lower layers (such as object edges and contour styles) to higher-level abstract representations.
[0046] MobileNetv2 adds point convolutions before depthwise separable convolutions, allowing for flexible adjustment of channels. Activation functions can effectively increase the nonlinear representation of the network model in high-dimensional space, but destroy features in low-dimensional feature space. This is the design of the inverse residual module: the input first undergoes 1x1 convolution for channel expansion, then 3x3 depthwise convolution, and finally 1x1 point convolution to compress the channels to a lower dimension. The whole process is "expansion-convolution-compression".
[0047] The Convolutional Block Attention Module (CBAM) is one of the most widely used attention mechanisms in deep learning. Its purpose is to improve model performance by focusing on important parts of an image. CBAM processes both channel and spatial dimensions sequentially. First, the channel attention module focuses on "which channels are important," and then the spatial attention module focuses on "where" an informational part is located. This dual attention mechanism allows CBAM to comprehensively capture key information from features.
[0048] Traditional Atrous Spatial Pyramid Pooling (ASPP) utilizes parallel dilated convolutions, extracting features using dilated convolutions with different dilation rates. It can capture different information at multiple scales through convolutions with different receptive fields.
[0049] However, as the dilation rate increases (especially when it exceeds 24), the effective weights of dilated convolution decrease, and the feature extraction capability also declines. Therefore, this invention uses a densely connected atrous spatial pyramid pooling (DenseASPP) module to replace the traditional ASPP module. DenseASPP can better capture semantic information in images.
[0050] The image segmentation loss function is calculated by weighting multiple image semantic category loss functions; the category loss function is the cross-entropy loss function.
[0051] In one feasible implementation, this invention divides the training dataset into a training set, a validation set, and a test set to ensure the model's generalization ability. Model hyperparameters (such as learning rate and batch size) are adjusted, and transfer learning techniques are used to improve training efficiency. Training is performed on an Nvidia Jetson Xavier NX development board, and the loss value and accuracy during training are monitored, with the optimal model saved.
[0052] During training, for each pixel of the image, the cross-entropy loss between its predicted and true classes is calculated. Cross-entropy loss is commonly used in classification problems and is suitable for pixel-by-pixel classification tasks in image semantic segmentation. To address the class imbalance problem, different weights are assigned to different classes. By setting weights for different classes, the loss contributions of each class can be balanced, preventing any one class from having an excessive impact on the overall loss.
[0053] S4. Obtain the current scene device image; input the current scene device image into the image semantic segmentation model to obtain the device segmentation image.
[0054] Optionally, the current scene device image is input into the image semantic segmentation model to obtain a device segmentation image, including:
[0055] Preprocess the device image of the current scene to obtain the processed device image;
[0056] Based on the processed device image, preliminary feature extraction is performed through the backbone network to obtain the first feature map;
[0057] Based on the first feature map, the CBAM module is used to perform feature enhancement to obtain the third feature map;
[0058] The third feature map is input into the DenseASPP module for multi-scale feature extraction to obtain a multi-scale dense feature map.
[0059] Based on the first feature map, resolution recovery is performed according to the multi-scale dense feature map to obtain the device segmentation image.
[0060] In one feasible implementation, after the camera captures an image, it transmits it to an electronic device in real time for processing. A preprocessing module is used to perform operations such as cropping, scaling, and noise reduction on the image to improve image quality.
[0061] DenseASPP employs a densely connected and cascaded dilated convolutional layer design. The output of each dilated convolutional layer is not only passed to the next layer, but also connected to all subsequent unvisited layers, forming a dense feature propagation path. The dilation rate of each layer in DenseASPP increases progressively, and the final output is a dense feature map with multiple dilation rates and multiple scales, as shown in Equation (1):
[0062] (1);
[0063] in, Indicates the expansion rate of the l-th layer; ( ) indicates dilated convolution; This represents the feature map formed by connecting the outputs of all previous layers. DenseASPP not only accumulates multi-scale features stepwise through dense connections, but also determines the receptive field by multiple consecutive dilated convolutional layers, avoiding the performance degradation caused by a single high dilation rate.
[0064] In the decoder section, high-level features are fused with low-level features. Low-level features, derived from the first feature map, contain more spatial detail information, helping to recover the spatial details of the image. Bilinear interpolation or deconvolution operations are used to upsample the feature map to the resolution of the original image to generate the final segmentation result. Upsampling progressively restores the image resolution, resulting in a more refined segmentation.
[0065] Optionally, based on the first feature map, feature enhancement is performed using the CBAM module to obtain a third feature map, including:
[0066] Based on a multilayer perceptron, channel attention is calculated according to the first feature map to obtain a channel attention map;
[0067] Based on the channel attention map, channel feature enhancement is performed on the first feature map to obtain the second feature map;
[0068] Based on a 7×7 convolutional layer, spatial attention is calculated using the second feature map to obtain a spatial attention map;
[0069] Based on the channel attention map and the spatial attention map, feature enhancement is performed on the first feature map to obtain the third feature map.
[0070] In one feasible implementation, channel attention applies global average pooling and global max pooling to each channel of the input feature map to obtain two distinct feature descriptions (an average and a maximum). These two descriptions are processed by a multilayer perceptron with shared weights, and the output feature map is then element-wise summed and a sigmoid function is applied to generate a channel attention map M. c The shape is consistent with the number of input channels. The formula is as follows (2):
[0071] (2);
[0072] Where MLP stands for Multilayer Perceptron; AvgPool represents average pooling; MaxPool represents max pooling; F is the first feature map; and σ represents the Sigmoid activation function.
[0073] The spatial attention map, after applying channel attention, is used to generate two 2D feature maps through average pooling and max pooling along the channel direction. These two feature maps are stacked along the channel dimension, passed through a 7×7 convolutional layer, and a spatial attention map M is generated using the sigmoid function. s The formula is as follows (3):
[0074] (3);
[0075] Here, Conv represents convolution operation; Concat represents concatenation operation.
[0076] The third feature map is obtained by combining the first feature map F with the channel attention map and the spatial attention map through element-wise multiplication, as shown in equation (4):
[0077] (4);
[0078] With this structure, CBAM can effectively enhance the network's response to important features in images, improving the model's performance on various visual tasks, especially in image recognition and segmentation in complex scenes.
[0079] S5. Obtain interaction requirements; Based on the space station equipment database, retrieve space station equipment information in the UI interface according to the interaction requirements and equipment segmentation images.
[0080] Optionally, based on the space station equipment database, and according to interaction requirements and equipment segmentation images, space station equipment information can be retrieved in the UI interface, including:
[0081] Information is extracted from the segmented image of the device to obtain device image information;
[0082] Based on the space station equipment database, data queries are performed according to equipment image information to obtain detailed equipment information;
[0083] Based on interaction requirements, the space station equipment information is retrieved through the UI interface using equipment image information and detailed equipment information.
[0084] In one feasible implementation, the present invention calls preset equipment images from the space station equipment database based on the equipment segmentation image and generates detailed equipment information of the image with annotations.
[0085] The UI displays annotated images, and users can click on areas of interest to view detailed information. Manually tapping an annotated area on the touchscreen brings up a 3D model of the component in that area, along with relevant information. Through the 3D model and text descriptions, users can learn about the component's name, function, and operation.
[0086] The interactive UI used in this invention is developed based on the Tkinter UI framework, supporting interactive functions such as exploded view operation, zooming in, zooming out, and rotation of 3D models of equipment within the space station. This invention creates 3D models for each industrial instrument and its parts, using SketchUp software to design and export the 3D model files.
[0087] This invention proposes a method for interaction between devices within a space station based on scene semantic segmentation. A channel-spatial dual attention module is introduced into the encoder part of the DeepLabV3+ network to improve the model's robustness to complex scenes. MobileNetV2 is used to replace the original backbone network Xception, utilizing its inverted residual structure and linear bottleneck layer to reduce the number of parameters and computational cost. The original ASPP module is replaced with DenseASPP, and by stacking densely connected dilated convolutional layers, a denser and wider receptive field is constructed, enhancing the ability to model the context of small targets and multi-scale objects. This invention provides an accurate and efficient method for interaction between devices within a space station based on scene semantic segmentation.
[0088] Figure 2 This is a block diagram illustrating an in-station device interaction system based on scene semantic segmentation, according to an exemplary embodiment. This system is used for in-station device interaction methods based on scene semantic segmentation. (Refer to...) Figure 2 The system includes a camera 210, electronic devices 220, and a touchscreen 230, wherein:
[0089] Camera 210 is used to acquire images of the device in the current scene;
[0090] Electronic device 220 is used to collect images of equipment inside the space station and obtain training data; a semantic segmentation model for the image to be trained is constructed based on the model structure of the DeepLabv3+ network; the semantic segmentation model for the image to be trained is trained using the training data based on the image segmentation loss function to obtain the image semantic segmentation model; the current scene equipment image is input into the image semantic segmentation model to obtain the equipment segmentation image;
[0091] Touchscreen 230 is used to obtain interaction requests; based on the space station equipment database, it retrieves space station equipment information in the UI interface according to the interaction requests and equipment segmentation images.
[0092] The image semantic segmentation model includes an encoder and a decoder.
[0093] The backbone network of the encoder section is a lightweight convolutional neural network, MobileNetV2.
[0094] The multi-scale feature extraction module of the encoder is the DenseASPP module with dense void space pyramid pooling.
[0095] The image semantic segmentation model adopts a CBAM module that combines channel attention and spatial attention mechanisms.
[0096] The image segmentation loss function is calculated by weighting multiple image semantic category loss functions; the category loss function is the cross-entropy loss function.
[0097] Optionally, the electronic device 220 is further used for:
[0098] Preprocess the device image of the current scene to obtain the processed device image;
[0099] Based on the processed device image, preliminary feature extraction is performed through the backbone network to obtain the first feature map;
[0100] Based on the first feature map, the CBAM module is used to perform feature enhancement to obtain the third feature map;
[0101] The third feature map is input into the DenseASPP module for multi-scale feature extraction to obtain a multi-scale dense feature map.
[0102] Based on the first feature map, resolution recovery is performed according to the multi-scale dense feature map to obtain the device segmentation image.
[0103] Optionally, the electronic device 220 is further used for:
[0104] Based on a multilayer perceptron, channel attention is calculated according to the first feature map to obtain a channel attention map;
[0105] Based on the channel attention map, channel feature enhancement is performed on the first feature map to obtain the second feature map;
[0106] Based on a 7×7 convolutional layer, spatial attention is calculated using the second feature map to obtain a spatial attention map;
[0107] Based on the channel attention map and the spatial attention map, feature enhancement is performed on the first feature map to obtain the third feature map.
[0108] Optionally, the touchscreen 230 is further used for:
[0109] Information is extracted from the segmented image of the device to obtain device image information;
[0110] Based on the space station equipment database, data queries are performed according to equipment image information to obtain detailed equipment information;
[0111] Based on interaction requirements, the space station equipment information is retrieved through the UI interface using equipment image information and detailed equipment information.
[0112] This invention proposes a method for interaction between devices within a space station based on scene semantic segmentation. A channel-spatial dual attention module is introduced into the encoder part of the DeepLabV3+ network to improve the model's robustness to complex scenes. MobileNetV2 is used to replace the original backbone network Xception, utilizing its inverted residual structure and linear bottleneck layer to reduce the number of parameters and computational cost. The original ASPP module is replaced with DenseASPP, and by stacking densely connected dilated convolutional layers, a denser and wider receptive field is constructed, enhancing the ability to model the context of small targets and multi-scale objects. This invention provides an accurate and efficient method for interaction between devices within a space station based on scene semantic segmentation.
[0113] Figure 3 This is a structural schematic diagram of an inter-station equipment interaction device provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the inter-station equipment interaction devices within the space station may include the aforementioned Figure 2 The illustrated space station intra-station equipment interaction system is based on scene semantic segmentation. Optionally, the space station intra-station equipment interaction device 310 may include a first processor 2001.
[0114] Optionally, the space station's intra-station equipment interaction device 310 may also include a memory 2002 and a transceiver 2003.
[0115] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.
[0116] The following is combined Figure 3 A detailed introduction to each component of the space station's internal equipment interaction device 310:
[0117] The first processor 2001 is the control center of the space station's intra-station equipment interaction device 310. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0118] Optionally, the first processor 2001 can perform various functions of the space station intra-station equipment interaction device 310 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0119] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 3 CPU0 and CPU1 are shown in the diagram.
[0120] In a specific implementation, as one example, the space station's intra-station equipment interaction device 310 may also include multiple processors, for example... Figure 3 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0121] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0122] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be accessed through the interface circuit of the space station intra-station equipment interaction device 310. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0123] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0124] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 3 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0125] Optionally, the transceiver 2003 can be integrated with the first processor 2001, or it can exist independently and interact with the interface circuit of the space station's internal equipment interaction device 310 via the interface circuit. Figure 3 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0126] It should be noted that, Figure 3 The structure of the space station intra-station equipment interaction device 310 shown in the figure does not constitute a limitation on the router. The actual knowledge structure identification device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0127] Furthermore, the technical effects of the space station intra-station equipment interaction device 310 can be referenced from the technical effects of the space station intra-station equipment interaction method based on scene semantic segmentation described in the above method embodiments, and will not be repeated here.
[0128] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0129] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0130] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable system. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0131] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0132] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0133] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0134] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0135] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, systems, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0136] In the embodiments provided by this invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.
[0137] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0138] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0139] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0140] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for interaction of equipment within a space station based on scene semantic segmentation, characterized in that, The method includes: Collect images of equipment inside the space station to obtain training data; A semantic segmentation model for the image to be trained is constructed based on the model structure of the DeepLabv3+ network. Based on the image segmentation loss function, the training data is used to train the image semantic segmentation model to obtain the image semantic segmentation model. The image semantic segmentation model includes an encoder and a decoder. The backbone network of the encoder section is a lightweight convolutional neural network, MobileNetV2. The multi-scale feature extraction module of the encoder is the DenseASPP module with dense void space pyramid pooling. The image semantic segmentation model adopts a CBAM module that combines channel attention and spatial attention mechanisms; Obtain the device image of the current scene; input the device image of the current scene into the image semantic segmentation model to obtain the device segmentation image, including: Preprocess the device image of the current scene to obtain the processed device image; Based on the processed device image, preliminary feature extraction is performed through the backbone network to obtain the first feature map; Based on the first feature map, the CBAM module is used to perform feature enhancement to obtain the third feature map; The third feature map is input into the DenseASPP module for multi-scale feature extraction to obtain a multi-scale dense feature map. Based on the first feature map, resolution restoration is performed according to the multi-scale dense feature map to obtain the device segmentation image; Obtain interaction requirements; based on the space station equipment database, retrieve space station equipment information in the UI interface according to the interaction requirements and equipment segmentation images.
2. The space station intra-station equipment interaction method based on scene semantic segmentation according to claim 1, characterized in that, The image segmentation loss function is obtained by weighted calculation of multiple image semantic category loss functions; the category loss function is the cross-entropy loss function.
3. The space station intra-station equipment interaction method based on scene semantic segmentation according to claim 1, characterized in that, The step of using the CBAM module to perform feature enhancement based on the first feature map to obtain the third feature map includes: Based on a multilayer perceptron, channel attention is calculated according to the first feature map to obtain a channel attention map; Based on the channel attention map, channel feature enhancement is performed on the first feature map to obtain the second feature map; Based on a 7×7 convolutional layer, spatial attention is calculated using the second feature map to obtain a spatial attention map; Based on the channel attention map and the spatial attention map, feature enhancement is performed on the first feature map to obtain the third feature map.
4. The space station intra-station equipment interaction method based on scene semantic segmentation according to claim 1, characterized in that, The method based on the space station equipment database retrieves space station equipment information from the UI interface according to interaction requirements and equipment segmentation images, including: Information is extracted from the segmented image of the device to obtain device image information; Based on the space station equipment database, data queries are performed according to equipment image information to obtain detailed equipment information; Based on interaction requirements, the space station equipment information is retrieved through the UI interface using equipment image information and detailed equipment information.
5. A space station intra-station equipment interaction system based on scene semantic segmentation, wherein the space station intra-station equipment interaction system based on scene semantic segmentation is used to implement the space station intra-station equipment interaction method based on scene semantic segmentation as described in any one of claims 1-4, characterized in that, The system includes a camera, electronic devices, and a touchscreen, wherein: The camera is used to acquire images of the device in the current scene; The electronic device is used to collect images of equipment inside the space station and obtain training data; to construct a semantic segmentation model for the image to be trained based on the model structure of the DeepLabv3+ network; to train the semantic segmentation model for the image to be trained using the training data based on the image segmentation loss function, and to obtain the image semantic segmentation model; and to input the current scene equipment image into the image semantic segmentation model to obtain the equipment segmentation image. The touchscreen is used to acquire interaction requests; based on the space station equipment database, and according to the interaction requests and equipment segmentation images, it retrieves space station equipment information in the UI interface.
6. An inter-station equipment interaction device for a space station, characterized in that, The space station's internal equipment interaction devices include: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Lightweight algorithm for auricular point segmentation
CN117036379A