Image semantic segmentation method and device, equipment and storage medium

By inputting target pose data and images into the image semantic segmentation network, the problem of inaccurate image semantic segmentation under large poses is solved, and higher semantic segmentation accuracy is achieved.

CN120219730APending Publication Date: 2025-06-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311793690.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In a large posture, the image quality collected by the camera on the XR device is poor, resulting in inaccurate image semantic segmentation results.

Method used

By obtaining the target pose data and the target image and inputting it into the image semantic segmentation network, the network can process the image based on the pose data, thereby improving the accuracy of the semantic segmentation results.

Benefits of technology

By introducing pose data, the problem of inaccurate image semantic segmentation results under large poses is solved, and the accuracy of image semantic segmentation is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219730A_ABST
    Figure CN120219730A_ABST
Patent Text Reader

Abstract

The invention provides an image semantic segmentation method and apparatus, a device and a storage medium. The method comprises the steps of obtaining target attitude data and a target image; and inputting the target attitude data and the target image into an image semantic segmentation network to obtain a semantic segmentation result for the target image output by the image semantic segmentation network. According to the invention, the accuracy of the image semantic segmentation result can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of image processing technology, and in particular, to an image semantic segmentation method, apparatus, device, and storage medium. Background Art

[0002] Image semantic segmentation technology is a technology that divides an image into regions with different semantic information and labels corresponding semantic tags for each region. For example, after performing image semantic segmentation on a certain room image, semantic tags can be added to the objects in the room image, such as the ceiling, beam, lamp, wall, floor, door, and window of the room, and it can be applied to fields such as map making, autonomous driving, and extended reality (XR).

[0003] When the image semantic segmentation technology is applied to XR devices, since XR devices are often worn on the user's head in a head-mounted form, during the free movement of the user's head, the posture of the XR device may change greatly. And the image quality of the images captured by the camera on the XR device is poor in such a large posture. Then, when using the image semantic segmentation technology to perform semantic segmentation on this image, semantic segmentation errors are likely to occur, resulting in inaccurate semantic segmentation results. Summary of the Invention

[0004] The embodiments of the present application provide an image semantic segmentation method, apparatus, device, and storage medium, which can improve the accuracy of the image semantic segmentation result.

[0005] In a first aspect, the embodiments of the present application provide an image semantic segmentation method, including:

[0006] Obtaining target pose data and a target image;

[0007] Inputting the target pose data and the target image into an image semantic segmentation network to obtain a semantic segmentation result of the image semantic segmentation network for the target image.

[0008] In a second aspect, the embodiments of the present application provide an image semantic segmentation apparatus, including:

[0009] An obtaining module, configured to obtain target pose data and a target image;

[0010] A processing module, configured to input the target pose data and the target image into an image semantic segmentation network to obtain a semantic segmentation result of the image semantic segmentation network for the target image.

[0011] In a third aspect, the embodiments of the present application provide an electronic device, including:

[0012] A processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the image semantic segmentation method as described in the embodiments of the first aspect.

[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium for storing a computer program, and the computer program causes a computer to execute the image semantic segmentation method as described in the embodiments of the first aspect.

[0014] In a fifth aspect, an embodiment of the present application provides a computer program product containing program instructions. When the program instructions run on an electronic device, the electronic device is caused to execute the image semantic segmentation method as described in the embodiments of the first aspect.

[0015] The technical solution disclosed in the embodiments of the present application inputs target pose data and a target image into an image semantic segmentation network, and the image semantic segmentation network processes the target image based on the target pose data to obtain a semantic segmentation result of the target image, thereby solving the problem of inaccurate image semantic segmentation results in large poses by introducing pose data, and thus improving the accuracy of image semantic segmentation. Description of the Drawings

[0016] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1a It is a schematic diagram of an image captured from an upward view by a camera on an XR device in a large pose;

[0018] Figure 1b It is a schematic diagram of an image captured from a normal view by a camera on an XR device;

[0019] Figure 2 It is a training flow chart of an image semantic segmentation network provided by an embodiment of the present application;

[0020] Figure 3 It is a structural image of an image semantic segmentation network to be trained provided by an embodiment of the present application;

[0021] Figure 4 It is a training flow chart of an image semantic segmentation network to be trained provided by an embodiment of the present application;

[0022] Figure 5Schematic diagram of another image semantic segmentation network to be trained provided by an embodiment of the present application;

[0023] Figure 6 Flowchart of training another image semantic segmentation network to be trained provided by an embodiment of the present application;

[0024] Figure 7 Schematic diagram of a specific image semantic segmentation network provided by an embodiment of the present application;

[0025] Figure 8 Schematic diagram of a head-mounted device when the electronic device is a head-mounted device provided by an embodiment of the present application;

[0026] Figure 9 Flowchart of an image semantic segmentation method provided by an embodiment of the present application;

[0027] Figure 10 Schematic block diagram of an image semantic segmentation device provided by an embodiment of the present application;

[0028] Figure 11 Schematic block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. According to the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0032] In the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more, that is, at least two. "At least one" means one or more.

[0033] To facilitate the understanding of the embodiments of the present application, before describing each embodiment of the present application, some concepts involved in all embodiments of the present application are first explained appropriately as follows:

[0034] 1) Virtual Reality (VR for short), a technology for creating and experiencing virtual worlds, which determines the generation of a virtual environment. It is a multi-source information (the virtual reality mentioned in this article includes at least visual perception, and may also include auditory perception, tactile perception, motion perception, and even taste perception, olfactory perception, etc.). It realizes the fusion of the virtual environment, an interactive three-dimensional dynamic visual scene and the simulation of entity behavior, enabling users to immerse themselves in the simulated virtual reality environment and realizing applications in various virtual environments such as maps, games, videos, education, medical treatment, simulation, collaborative training, sales, assisting manufacturing, maintenance and repair.

[0035] 2) VR device, a terminal that realizes the virtual reality effect, usually provided in the form of glasses, a head-mounted display (HMD for short), or contact lenses, for realizing visual perception and other forms of perception. Of course, the form in which the VR device is realized is not limited to this, and it can be further miniaturized or enlarged according to actual needs.

[0036] Optionally, the VR devices described in the embodiments of the present application may include, but are not limited to, the following types:

[0037] 2.1) PC-based virtual reality (PCVR) device, which uses the PC to perform relevant calculations and data output for virtual reality functions, and the external PC-based virtual reality device realizes the virtual reality effect using the data output by the PC.

[0038] 2.2) Mobile virtual reality devices support setting up a mobile terminal (such as a smartphone) in various ways (such as a head-mounted display with a dedicated card slot). Through a wired or wireless connection with the mobile terminal, the mobile terminal performs relevant calculations for virtual reality functions and outputs data to the mobile virtual reality device. For example, watch virtual reality videos through the APP of the mobile terminal.

[0039] 2.3) All-in-one virtual reality devices have a processor for performing relevant calculations for virtual functions, thus having independent virtual reality input and output functions and not requiring connection to a PC or mobile terminal, with high degrees of freedom in use.

[0040] 3) Augmented Reality (AR): A technology that, during the process of a camera capturing an image, calculates in real time the camera pose parameters of the camera in the real world (or three-dimensional world, physical world), and adds virtual elements to the image captured by the camera according to the camera pose parameters. Virtual elements include, but are not limited to: images, videos, and three-dimensional models. The goal of AR technology is to interact by superimposing the virtual world on the real world on the screen.

[0041] 4) Mixed Reality (MR): By presenting virtual scene information in a real-world scene, an interactive feedback information loop is established between the real world, the virtual world, and the user to enhance the sense of reality of the user experience. For example, an analog scene that integrates sensory inputs created by a computer (such as virtual objects) with sensory inputs from a physical set or their representations. In some MR scenes, the sensory inputs created by the computer can adapt to changes in the sensory inputs from the physical set. Additionally, some electronic systems for presenting MR scenes can monitor the orientation and / or position relative to the physical set so that virtual objects can interact with real objects (i.e., physical elements from the physical set or their representations). For example, the system can monitor movement so that a virtual plant appears stationary relative to a physical building.

[0042] 5) XR refers to combining the real and the virtual through a computer to create a virtual environment for human-computer interaction. XR is also a general term for multiple technologies such as VR, AR, and MR. By integrating the visual interaction technologies of the three, it brings a "sense of immersion" of seamless conversion between the virtual world and the real world to users.

[0043] 6) A virtual scene is a virtual scene displayed (or provided) when an application runs on an electronic device. The virtual scene can be a simulation environment of the real world, a semi-simulated and semi-fictional virtual scene, or a purely fictional virtual scene. The virtual scene can be any one of a two-dimensional virtual scene, a 2.5D virtual scene, or a three-dimensional virtual scene. The embodiments of the present application do not limit the dimension of the virtual scene. For example, the virtual scene can include the sky, land, ocean, etc. The land can include environmental elements such as deserts, cities, etc. The user can control the virtual object to move in the virtual scene. It should be understood that the above virtual scene can also be referred to as a virtual space.

[0044] 7) A virtual object is an object that interacts in a virtual scene, is controlled by a user or a robot program (for example, a robot program based on artificial intelligence), and can be an object that is stationary, moves, and performs various behaviors in the virtual scene, such as various characters in a game.

[0045] Considering that when applying the image semantic segmentation technology to an XR device, since the XR device is often worn on the user's head in a head-mounted form, during the free movement of the user's head, the posture of the XR device may change greatly. And the image quality of the camera on the XR device under such a large posture is poor. Then, when using the image semantic segmentation technology to perform semantic segmentation on the image captured by the corresponding camera, it is easy to have semantic segmentation errors, resulting in inaccurate semantic segmentation results. For example, when the user's head makes a large posture of looking up during the process of wearing the XR device, the image captured by the camera on the XR device for this large posture is an image from a looking-up perspective, specifically as Figure 1a shown, while the image captured by the camera under a normal perspective should be as Figure 1b shown. Then, when using the traditional semantic segmentation technology to perform semantic segmentation on the Figure 1a looking-up perspective image shown, the table in the Figure 1b normal perspective image shown will be recognized as the ceiling, resulting in a large error in the segmentation result.

[0046] To solve the above technical problems, the inventive concept of the present application is: when performing semantic segmentation on a target image, by obtaining the target pose data corresponding to the target image and inputting the target pose data and the target image into an image semantic segmentation network together, the image semantic segmentation network can perform semantic segmentation on the target image based on the relationship between the device pose used by the user and the image semantic segmentation, thereby improving the accuracy of the image semantic segmentation.

[0047] The technical solutions of the present application will be described in detail below through some embodiments. The embodiments described below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0048] First, the training process of the image semantic segmentation network in this application will be specifically described.

[0049] Figure 2 FIG. is a training flowchart of an image semantic segmentation network provided by an embodiment of this application. The training execution entity of this image semantic segmentation network can be a device with a model training function, such as an image semantic segmentation device. In some alternative embodiments, this device can be composed of hardware and / or software and can be integrated into an electronic device. In this application, the electronic device can be, but is not limited to: tablet computers, personal desktop computers, laptop computers, XR devices, and other wearable devices, etc. This application does not impose any restrictions on the type of electronic device.

[0050] As Figure 2 shown, the method can include the following steps:

[0051] S101, obtain training pose data and a training image corresponding to the training pose data, where the training image includes semantic segmentation information.

[0052] The above-mentioned training pose data and the training image corresponding to the training pose data can be understood as the pose data of the electronic device at any moment and the image at that moment. For example, if the training pose data is the pose data of the electronic device at the T1-th moment, then the training image is the image at the T1-th moment.

[0053] The above-mentioned pose data can be understood as the rotation angle of the electronic device.

[0054] The semantic segmentation information in the above-mentioned training image can be pre-annotated based on the training image, where the pre-annotation can be manual annotation or automatic annotation based on relevant software or algorithms. There are no restrictions on the annotation method of the semantic segmentation information in the training image here.

[0055] In this application, the electronic device can be optionally any device worn on the user's head and capable of displaying an image screen to the user, such as a VR device, an AR device, an MR device, or a head-mounted display (HMD) under an XR device that can simulate a virtual scene. This application does not impose any restrictions on this.

[0056] The pose of the above-mentioned electronic device at any moment can obtain pose data through the pose measurement device in the electronic device, and then obtain the attitude data from the pose data. In some embodiments, the pose measurement device can be selected as any device or apparatus capable of collecting the pose information of the electronic device, such as a pose sensor, an inertial measurement unit (IMU), and other devices including a gyroscope and an accelerometer, etc. This application places no restrictions on the pose measurement device as long as it can obtain the pose information of the electronic device.

[0057] In this application, obtaining the pose data of the electronic device through the pose measurement device can be taking the origin of the pose measurement device as the origin of the world coordinate system to obtain the pose data of the electronic device relative to the world coordinate system, that is, the pose data obtained by the pose measurement device is the pose data of the electronic device in the world coordinate system.

[0058] Normally, the pose data of the electronic device relative to the world coordinate system can be denoted as T, and wherein, T can represent a 4×4 transformation matrix, the rotation part is a 3×3 rotation matrix, which can be denoted as R; the translation part is a 3×1 translation vector, which can be denoted as t. Then, obtaining the attitude data from the pose data in this application can be understood as obtaining the rotation matrix R from the pose data T.

[0059] To facilitate the image semantic segmentation network to learn the attitude prior of the electronic device, this application can convert the above-mentioned attitude data represented as the rotation matrix R into the quaternion form. It should be understood that the quaternion is a compact expression for representing the rotation attitude.

[0060] As an alternative implementation, when the attitude data of the rotation matrix R is then the four components corresponding to each element in the rotation matrix R in the quaternion form can be calculated by the following formula:

[0061]

[0062] wherein, q0, q1, q2, and q3 represent the four components corresponding to each element in the rotation matrix R, sqrt represents the square root calculation, sign represents the sign function, and Each element in the above-mentioned rotation matrix R can be represented as r ij This r ij represents the element in the i-th row and j-th column of the rotation matrix R.

[0063] When the electronic device is a head-mounted device worn on the user's head, a camera for collecting the external environment is provided on the head-mounted device, so that when the user uses the head-mounted device for an immersive experience, the user can perceive the external real environment based on the environmental image collected by the camera, thereby using the electronic device safely, ensuring personal safety, and avoiding danger. Therefore, the image corresponding to the electronic device in any posture in this application can be understood as the environmental image collected by the camera on the electronic device when the electronic device is in any posture.

[0064] Optionally, the above training posture data and training image can be a training posture data and a training image in a training sample set. Among them, the training posture data is specifically the posture data represented by quaternions. It should be understood that if the image semantic segmentation network can process the rotation matrix R, then the training posture data in this application can also be the rotation matrix R in the pose data T. For the convenience of description, the following will take the training posture data in the form of quaternions as an example for illustration.

[0065] It should be understood that the above training sample set can be obtained by acquiring multiple posture data and the image corresponding to each posture data based on different electronic devices, and using these posture data and the corresponding images as the training posture data and training images. Then, all the training posture data and their corresponding training images are combined to form a training sample set.

[0066] In some embodiments, considering that each training posture data and training image in the training sample set have the same training process for the image semantic segmentation network. For example, in each training process, a training posture data and a training image corresponding to the training posture are input into the image semantic segmentation network. After the training of the training posture data and the training image is completed, the next training posture data and training image are input and the training starts again. Therefore, for the convenience of describing the technical solution of this application, this application will take a training posture data and a training image as an example to illustrate the training process of the image semantic segmentation network in the following embodiments.

[0067] S102, input the training posture data and the training image into the image semantic segmentation network to be trained, and train the image semantic segmentation network to be trained to obtain the image semantic segmentation network.

[0068] Among them, the image semantic segmentation network to be trained can be understood as an image semantic segmentation network whose parameters have not been optimized. That is, the image semantic segmentation network in the initial state.

[0069] In some alternative embodiments, the training pose data and the training images are input into an image semantic segmentation network to be trained, so that the image semantic segmentation network to be trained performs semantic segmentation processing on the training images based on the training pose data and outputs a prediction result. Then, by comparing the prediction result with the semantic segmentation information of the training images, the difference (or error) between the prediction result output by the image semantic segmentation network and the semantic segmentation information of the training images is determined. If the prediction result is quite different from the semantic segmentation information of the training images, it indicates that the semantic segmentation accuracy of the image semantic segmentation network is relatively poor. Then, based on the difference between the prediction result and the semantic segmentation information of the training images, the image semantic segmentation network is trained again, and the above operations are repeated until the difference between the prediction result output by the image semantic segmentation network and the semantic segmentation information of the training images is less than a preset value.

[0070] When the difference between the prediction result output by the image semantic segmentation network for the last time and the semantic segmentation information of the training images is less than the preset value, it indicates that the semantic segmentation result of the image semantic segmentation network after the latest training is relatively accurate. At this time, the image semantic segmentation network after the latest training can be determined as the final image semantic segmentation network.

[0071] The above preset value can be flexibly set according to the network training requirements. For example, if a high accuracy and small error of the semantic segmentation result output by the network are required, then the preset value can be set smaller; on the contrary, the preset value can be set larger. The present application does not impose any restrictions on this.

[0072] The technical solution provided by the embodiments of the present application trains an image semantic segmentation network by using the pose data of an electronic device and the images corresponding to the pose data as training samples, so that the trained image semantic segmentation network can learn the relationship between the electronic device pose and the image semantic segmentation. Furthermore, based on the relationship between the electronic device pose and the image semantic segmentation, semantic segmentation is performed on the images to solve the problem of inaccurate image semantic segmentation results of the electronic device in large poses by introducing pose priors, thereby improving the accuracy of image semantic segmentation.

[0073] In some alternative embodiments, as Figure 3 shown, the image semantic segmentation network to be trained in the present application includes a feature extraction module and a semantic segmentation module, where the output of the feature extraction module is connected to the input of the semantic segmentation module.

[0074] The above feature extraction module is used to extract features from the input training pose data to obtain training pose features.

[0075] The above semantic segmentation module is used to extract features from the input training image to obtain a feature map, and fuse the feature map and the training pose features output by the feature extraction module to obtain a predicted semantic segmentation result.

[0076] Based on the network structure shown in Figure 3 as shown in Figure 4 above, S102 may include the following steps: S102-1 to S102-3:

[0077] S102-1, input the training pose data into the feature extraction module to obtain the training pose features output by the feature extraction module.

[0078] In this application, the feature extraction module can be optionally any network that supports processing the training pose data represented by quaternions into a one-dimensional feature vector. Such as, a one-dimensional convolutional network, a fully connected network, a Transformer network, and a recurrent neural network (RNN), etc. This application places no restrictions on the feature extraction module as long as it can process the training pose data represented by quaternions into a one-dimensional feature vector.

[0079] That is, the pose features output by the above feature extraction module are one-dimensional pose feature vectors.

[0080] It should be noted that when the above feature extraction module is a one-dimensional convolutional network, the one-dimensional convolutional network may include multiple convolutional layers.

[0081] S102-2, input the training pose features and the training image into the semantic segmentation module, and through the semantic segmentation module, fuse the training pose features and the training image to obtain a predicted semantic segmentation result.

[0082] As Figure 3 shown, in this application, the output end of the feature extraction module included in the image semantic segmentation network is connected to the input end of the semantic segmentation module. Then, the one-dimensional pose feature vector output by the feature extraction module will be input into the semantic segmentation module as input data of the semantic segmentation module, and the training image corresponding to the one-dimensional pose feature vector will also be input into the semantic segmentation module as input data, so that the semantic segmentation module can perform semantic segmentation on the training image based on the training pose features and the training image. Among them, when the semantic segmentation module performs semantic segmentation on the training image based on the training pose features and the training image, it can first extract features from the training image to obtain a feature map of the training image. Then, splice the feature map and the one-dimensional pose feature vector to achieve the fusion of the pose features and the feature map. Then, perform semantic segmentation on the splicing result to obtain a predicted semantic segmentation result, so as to learn the relationship between the pose of the electronic device and semantic segmentation.

[0083] S102-3: Based on the predicted semantic segmentation result and the semantic segmentation information of the training image, train the image semantic segmentation network to obtain the image semantic segmentation network.

[0084] Compare the predicted semantic segmentation result output by the image semantic segmentation network after the above training with the semantic segmentation information of the training image to determine whether the semantic segmentation of the trained image semantic segmentation network for the training image is accurate. If it is determined after comparison that the difference (or error) between the predicted semantic segmentation result and the semantic segmentation information of the training image is greater than or equal to the preset threshold, it indicates that the image semantic segmentation network has not been trained yet. At this time, based on the difference between the predicted semantic segmentation result and the semantic segmentation information of the training image, perform backpropagation training on the image semantic segmentation network, and then determine the difference between the new predicted semantic segmentation result output by the image semantic segmentation network and the semantic segmentation information of the training image again. If the difference between the new predicted semantic segmentation result and the semantic segmentation information of the training image is still greater than the preset threshold, repeat the above backpropagation training and result comparison operations until the stop condition is met. If it is determined after the latest comparison that the difference between the predicted semantic segmentation result and the semantic segmentation information of the training image is less than the preset threshold, it indicates that the image semantic segmentation network has been trained, and at this time, determine the trained image semantic segmentation network as the final image semantic segmentation network.

[0085] In some alternative embodiments, determining the difference between the predicted semantic segmentation result and the semantic segmentation information of the training image can also be achieved by using methods such as the backpropagation algorithm, cross-entropy loss function, gradient descent method, cost function, or error function, as long as the difference between the predicted semantic segmentation result and the semantic segmentation information of the training image can be determined. The present application does not impose any restrictions on the specific determination method.

[0086] In some alternative embodiments, the above-mentioned backpropagation training of the image semantic segmentation network based on the difference between the predicted semantic segmentation result and the semantic segmentation information of the training image specifically involves adjusting or optimizing the parameters in the image semantic segmentation network, such as adjusting or optimizing the parameters in the feature extraction module and the semantic segmentation module.

[0087] The above-mentioned stop condition can be that the difference between the predicted semantic segmentation result and the semantic segmentation information of the training image is less than the preset threshold, or the number of training times reaches the preset number of times.

[0088] In this application, if a higher-accuracy semantic segmentation result is desired, the preset threshold can be set smaller or the number of training times can be set larger. Conversely, the preset threshold can be set relatively larger or the number of training times can be set relatively smaller. Specifically, it can be flexibly adjusted according to the actual semantic segmentation accuracy requirements, and this application does not impose any restrictions on this.

[0089] Exemplarily, assume that the training image is image XX, and the semantic segmentation information of this image XX is that region a is a table and region b is the floor, and the training pose data is YY. Then, inputting the training pose data YY into the feature extraction module obtains a one-dimensional pose feature vector yy. At this time, input the one-dimensional pose feature vector yy and the image XX into the semantic segmentation module, so that the semantic segmentation module performs semantic segmentation on the image XX based on the correlation between the one-dimensional feature vector yy and the image XX to obtain a predicted semantic segmentation result that region a in the image XX is the ceiling and region b is the floor. By comparing the semantic segmentation information: region a is a table, region b is the floor, and the predicted semantic segmentation result: region a is the ceiling, region b is the floor, it is determined that the semantic segmentation information is different from the predicted semantic segmentation result, indicating that the semantic segmentation result of this image is not well-trained. At this time, the image semantic segmentation network can be continuously trained based on the difference between the semantic segmentation information and the predicted semantic segmentation result until the training end condition is met.

[0090] As Figure 5 shown, the image semantic segmentation network to be trained in this application further includes a feature processing module. Among them, the input end of the feature processing module is connected to the output end of the feature extraction module, and the output end of the feature processing module is connected to the input end of the semantic segmentation module.

[0091] As Figure 6 shown, the above S102-2 includes: S102-4 to S102-5:

[0092] S102-4, input the training pose feature output by the feature extraction module into the feature processing module to normalize the training pose feature through the feature processing module to obtain a normalized training pose feature.

[0093] S102-5, input the normalized training pose feature and the training image into the semantic segmentation module, and the semantic segmentation module performs feature fusion on the normalized training pose feature and the training image to obtain a predicted semantic segmentation result.

[0094] Considering that the feature extraction module extracts features from the training pose data represented by quaternions, the scales or change amplitudes of the feature vectors in the obtained one-dimensional pose feature vectors are inconsistent. Therefore, in order to obtain feature vectors under the same scale or the same change amplitude, the present application can perform normalization processing on the one-dimensional pose feature vectors (training pose features) output by the feature extraction module. Then, the normalized pose features are input into the semantic segmentation module, so that the semantic segmentation module can simplify the semantic segmentation operation based on the normalized pose features and the training images, and improve the speed of obtaining the predicted semantic segmentation result.

[0095] To facilitate understanding of the image semantic segmentation network training solution provided by the embodiments of the present application, the network structure of the image semantic segmentation network in the present application will be specifically described below in combination with a specific example. As Figure 7 shown, the image semantic segmentation network includes: a feature extraction module, a feature processing module, and a semantic segmentation module.

[0096] Among them, the feature extraction module is used to process the quaternion pose data of the electronic device into a one-dimensional pose feature vector. And this feature extraction module can include multiple convolutional layers, specifically referring to Figure 7 .

[0097] The feature processing module is used to process the one-dimensional pose feature vector into training pose features under the same scale or the same change amplitude, that is, the normalized training pose feature vector.

[0098] The semantic segmentation module is used to extract features from the training image to obtain a feature map. Then, the normalized training pose feature vector is fused with the feature map to obtain a fusion result. Then, the semantic segmentation result of the image is obtained based on the fusion result.

[0099] As Figure 7 shown, the network structure of this semantic segmentation module is an encoder-decoder structure.

[0100] Among them, when encoding, the encoder may include 4 downsampling convolutional layers. The first downsampling convolutional layer first converts the normalized training pose features from 1D to a 2D image to obtain a pose image. For example, if the normalized training pose feature is 1, then converting 1 to a 2D image is specifically 1*(H*W) = H*W, where H represents the height of the image and W represents the width of the image. Then, downsampling convolution is performed on the pose image and the training image, and the convolution result is output to the following three downsampling convolutional layers. The convolution result is processed by the following 3 downsampling convolutional layers to output a fused feature map, so that the image semantic segmentation network can learn the relationship between the pose data of the electronic device and image semantic segmentation. Thus, when the semantic segmentation network performs semantic segmentation on the image, it can fully consider the influence of the electronic device pose on semantic segmentation to improve the accuracy of the semantic segmentation result. When decoding, the decoder may include 4 upsampling convolutional layers, and the fused feature image dimension is restored through the 4 upsampling convolutional layers to output the semantic segmentation result.

[0101] It should be understood that the number of downsampling convolutional layers in the above encoder is the same as the number of upsampling convolutional layers in the decoder. Moreover, the number of convolutional layers can be flexibly set according to the actual semantic segmentation accuracy. For example, if a higher-accuracy semantic segmentation result is required, the number of convolutional layers can be set to be more. If a lower-accuracy semantic segmentation result is required, the number of convolutional layers can be set to be less. The present application does not impose any restrictions on this.

[0102] In some alternative embodiments, considering that U-Net is a classic semantic segmentation network architecture, the semantic segmentation module in the present application is optionally set as U-Net. It should be understood that U-Net has an encoder-decoder structure and realizes the fusion of information through skip connections. The main feature of U-Net is to establish skip connections between the layers corresponding to the encoder in the decoder part to retain more spatial information. Specifically, the encoder part of U-Net gradually reduces the size and dimension of the feature map through convolution and pooling operations to extract semantic information. The decoder part gradually restores the size of the feature map through upsampling and convolution operations, and fuses the feature map of the corresponding layer of the encoder with the feature map of the decoder. Such a design can effectively utilize features of different scales to improve the accuracy of segmentation. U-Net performs well in semantic segmentation tasks, especially suitable for small data sets and scenarios where boundary details are important.

[0103] Considering that there are many networks with an encoder-decoder structure, the semantic segmentation module in this application can also be set to other networks with an encoder-decoder structure, such as the basic convolutional model (InternImage), Mask2Former, Fully Convolutional Networks (FCN), SegNet, DeepLab, SPNet (Pyramid Scene Parsing Network), Mask R-CNN, HRNet (High-Resolution Network), BiSeNet (Bilateral Segmentation Network), DANet (Dual Attention Network), DeeplabV3, etc. This application does not impose any restrictions on this.

[0104] In this application, by combining the pose of the electronic device with traditional semantic segmentation algorithms, the traditional semantic segmentation algorithms can learn the relationship between the pose of the electronic device and semantic segmentation, and an image semantic segmentation network is trained by introducing pose priors. Therefore, when performing semantic segmentation on an image based on this image semantic segmentation network, the influence of the pose of the electronic device on semantic segmentation can be fully considered, thereby improving the accuracy of the semantic segmentation results.

[0105] After introducing the training process of the image semantic segmentation network in detail, the process of performing semantic segmentation on an image based on the image semantic segmentation network will be specifically introduced below. It should be understood that the image semantic segmentation method provided in this application can be applied to an electronic device, and a camera can be set on the electronic device. In this application, the electronic device is preferably a head-mounted device, and the structure of the head-mounted device can be as Figure 8 shown. Among them, the image semantic segmentation network trained in the foregoing embodiment is applied to the electronic device.

[0106] Figure 9 It is a flowchart of an image semantic segmentation method provided by an embodiment of this application. As Figure 9 shown, the method may include the following steps:

[0107] S201, obtain target pose data and a target image.

[0108] S202, input the target pose data and the target image into the image semantic segmentation network, and obtain the semantic segmentation result of the target image output by the image semantic segmentation network.

[0109] Among them, the image semantic segmentation network is trained based on the training embodiment of the foregoing image semantic segmentation network.

[0110] The above-mentioned image semantic segmentation network can be as follows Figure 7 As shown, the image semantic segmentation network includes: a feature extraction module, a feature processing module, and a semantic segmentation module. Among them, the network structure of the semantic segmentation module is an encoder-decoder structure. For details, please refer to Figure 7 .

[0111] In some alternative embodiments, the electronic device obtains real-time pose information through a pose measurement device, obtains the real-time attitude from the real-time pose information, and determines the real-time attitude as the target attitude data. At the same time, the electronic device can also collect an environmental image corresponding to the real-time pose through a camera on itself, and determine the environmental image as the target image. Then, the electronic device converts the target attitude data into a quaternion form. After that, the target attitude data in quaternion form and the target image are used as input data and input into the image semantic segmentation network, so that the feature extraction module in the image semantic segmentation network extracts features from the target attitude data in quaternion form to obtain target attitude features, and inputs the target attitude features into the feature processing module in the image semantic segmentation network, so that the feature processing module performs normalization processing on the target attitude features to obtain the normalized target attitude features. The feature processing module outputs the normalized target attitude features to the semantic segmentation module in the image semantic segmentation network. The semantic segmentation module extracts features from the target image input by the electronic device to obtain a feature map, then performs feature fusion on the feature map and the normalized target attitude features, and performs semantic segmentation on the target image based on the fused features to obtain a semantic segmentation result.

[0112] In this application, when the semantic segmentation module performs feature fusion on the feature map and the normalized target attitude features, it can splice the feature map and the normalized target attitude features to obtain a splicing result, so as to realize the fusion of the target attitude features and the feature map. Furthermore, the semantic segmentation module performs semantic segmentation on the splicing result to obtain a semantic segmentation result for the target image.

[0113] In some alternative embodiments, considering that the real-time pose information obtained by the pose measurement device may have errors due to its own defects, such as cumulative errors or other noises in the real-time pose information, resulting in errors in the real-time pose information. Therefore, after obtaining the real-time pose information in this application, optionally, the real-time pose is input into a pose optimization module, and the pose optimization module corrects the real-time pose information to obtain the corrected real-time pose. Then, an accurate real-time attitude is obtained from the corrected real-time pose, which can further improve the accuracy of image semantic segmentation.

[0114] In this application, the pose optimization module can be any device or apparatus capable of correcting pose information, such as a Kalman filter, an extended Kalman filter, etc.

[0115] In some alternative embodiments, the present application may further be to input the real-time pose into the pose optimization module before inputting the real-time pose into the feature extraction module, so as to correct the real-time pose through the pose optimization module. At this time, the pose optimization module may be an optional quaternion filter.

[0116] In some alternative implementation scenarios, considering the difference in the data acquisition frequencies of the pose measurement device and the camera on the electronic device. Therefore, before performing image semantic segmentation on the environmental image collected by the camera on the electronic device, the present application may first calibrate the pose measurement device and the camera to determine the time difference (delay time) between the data collected by the pose measurement device and the camera. Thus, when the electronic device performs image semantic segmentation on the environmental image collected by the camera, the pose data and the environmental image collected synchronously can be aligned based on the delay time. Then, the target pose data and the target image after the alignment process are input into the image semantic segmentation network to obtain the semantic segmentation result of the target image output by the image semantic segmentation network.

[0117] In the present application, to determine the delay time between the data collected by the pose measurement device and the camera, the pose measurement device may be controlled to collect multiple pose data within a preset time period, and the camera may be controlled to collect multiple environmental images. A first motion trajectory is obtained based on the multiple pose data, and a second motion trajectory is obtained based on the multiple environmental images using the simultaneous localization and mapping (SLAM) technology. The first motion trajectory and the second motion trajectory are placed in the same spatial coordinate system, and the first time corresponding to the peak of the first motion trajectory and the second time corresponding to the peak of the second motion trajectory are respectively determined. The difference between the first time and the second time is determined, and this difference is determined as the time difference between the data collected by the pose measurement device and the camera. Among them, the preset time period can be flexibly set according to the calibration requirements, such as 3 minutes, 5 minutes, etc., and the present application does not impose any restrictions on this.

[0118] Based on the above delay time, the alignment processing of the pose data and the environmental image at the same timestamp can be understood as follows: when the data acquisition frequency of the pose measurement device is lower than that of the camera, it means that the data acquisition period of the pose measurement device for one time is longer than that of the camera for one time. Then, the above delay time can be added to the time corresponding to the pose data collected by the pose measurement device, so that the pose data collected by the pose measurement device is aligned with the image data collected by the camera in time. When the data acquisition frequency of the pose measurement device is higher than that of the camera, it means that the data acquisition period of the pose measurement device for one time is shorter than that of the camera for one time. Then, the above delay time can be added to the time corresponding to the image collected by the camera, so that the image data collected by the camera is aligned with the pose data collected by the pose measurement device in time.

[0119] That is to say, after obtaining the target pose and the target image in this application, it further includes aligning the target pose and the target image.

[0120] It can be understood that in this application, by inputting the pose data of the electronic device and the environmental image corresponding to the pose data into the image semantic segmentation network, when the image semantic segmentation network performs semantic segmentation on the environmental image, it can fully consider the pose of the electronic device corresponding to the environmental image, so as to perform semantic segmentation on the environmental image based on the prior of the electronic device pose, thereby improving the accuracy of image semantic segmentation.

[0121] The technical solution disclosed in the embodiments of this application inputs the target pose data and the target image into the image semantic segmentation network, and processes the target image through the image semantic segmentation network based on the target pose data to obtain the semantic segmentation result of the target image, thereby solving the problem of inaccurate image semantic segmentation results under large poses by introducing pose data, and improving the accuracy of image semantic segmentation.

[0122] Next, refer to the attached Figure 10 , and describe an image semantic segmentation device proposed in the embodiments of this application. Figure 10 It is a schematic block diagram of an image semantic segmentation device provided in the embodiments of this application.

[0123] As Figure 10 shown, the image semantic segmentation device 300 includes: an acquisition module 310 and a processing module 320.

[0124] Among them, the acquisition module 310 is used to acquire target pose data and a target image;

[0125] A processing module 320, configured to input the target pose data and the target image into an image semantic segmentation network, and obtain a semantic segmentation result of the target image output by the image semantic segmentation network.

[0126] An optional implementation manner of the embodiment of the present application, the apparatus 300 further includes:

[0127] An alignment module, configured to perform alignment processing on the target pose data and the target image;

[0128] Correspondingly, the processing module 320 is specifically configured to: input the target pose data and the target image after alignment processing into the image semantic segmentation network.

[0129] An optional implementation manner of the embodiment of the present application, the image semantic segmentation network includes: a feature extraction module and a semantic segmentation module, and the processing module 320 includes:

[0130] A feature extraction unit, configured to input the target pose data into the feature extraction module, and obtain target pose features output by the feature extraction module;

[0131] A feature fusion unit, configured to input the target pose features and the target image into the semantic segmentation module, and perform feature fusion on the target pose features and the target image through the semantic segmentation module to obtain a semantic segmentation result of the target image.

[0132] An optional implementation manner of the embodiment of the present application, the feature fusion unit is specifically configured to: extract a feature map of the target image through the semantic segmentation module, splice the feature map and the target pose features, and obtain a semantic segmentation result of the target image based on the splicing result.

[0133] An optional implementation manner of the embodiment of the present application, the network structure of the semantic segmentation module is an encoder-decoder structure.

[0134] An optional implementation manner of the embodiment of the present application, the image semantic segmentation network further includes: a feature processing module, and the processing module 320 further includes:

[0135] A feature processing unit, configured to input the target pose features into the feature processing module, and obtain normalized target pose features output by the feature processing module;

[0136] Correspondingly, the processing module 320 is specifically configured to: input the normalized target pose features and the target image into the image semantic segmentation network, and obtain a semantic segmentation result of the target image output by the image semantic segmentation network.

[0137] An alternative implementation of the embodiment of the present application, the apparatus 300 further includes:

[0138] A data acquisition module, configured to acquire training pose data and a training image corresponding to the training pose data, where the training image includes semantic segmentation information;

[0139] A network training module, configured to input the training pose data and the training image into an image semantic segmentation network to be trained, and train the image semantic segmentation network to be trained to obtain the image semantic segmentation network.

[0140] An alternative implementation of the embodiment of the present application, the network training module is specifically configured to: input the training pose data into the feature extraction module to obtain training pose features output by the feature extraction module; input the training pose features and the training image into the semantic segmentation module, and perform feature fusion on the training pose features and the training image through the semantic segmentation module to obtain a predicted semantic segmentation result; and train the image semantic segmentation network based on the predicted semantic segmentation result and the semantic segmentation information of the training image.

[0141] An alternative implementation of the embodiment of the present application, the network training module is further configured to: determine an error between the predicted semantic segmentation result and the semantic segmentation information of the training image; and iteratively optimize the parameters of the feature extraction module and the semantic segmentation module based on the error until a stop condition is met.

[0142] An alternative implementation of the embodiment of the present application, the network training module includes:

[0143] A training feature processing unit, configured to input the training pose features into the feature processing module to obtain normalized training pose features output by the feature processing module;

[0144] Correspondingly, the network training module is further configured to:

[0145] Input the normalized training pose features and the training image into an image semantic segmentation network to be trained, and train the image semantic segmentation network to be trained.

[0146] It should be understood that the apparatus embodiment and the foregoing method embodiment can correspond to each other, and similar descriptions can refer to the method embodiment. To avoid repetition, it will not be elaborated here. Specifically, Figure 10 The illustrated apparatus 300 can execute Figure 9 The corresponding method embodiment, and the foregoing and other operations and / or functions of each module in the apparatus 300 are respectively for implementing Figure 9 The corresponding processes in each method, and for the sake of brevity, will not be elaborated here.

[0147] In the above, the device 300 of the embodiments of the present application has been described from the perspective of functional modules in combination with the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, can also be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiments in the first aspect of the embodiments of the present application can be completed by the integrated logic circuit in the hardware in the processor and / or instructions in the form of software. The steps of the method in the first aspect disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or can be executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the method embodiments in the first aspect above.

[0148] Figure 11 This is a schematic block diagram of an electronic device provided by the embodiments of the present application. As Figure 11 shown, the electronic device 400 may include:

[0149] A memory 410 and a processor 420. The memory 410 is used to store a computer program and transmit the program code to the processor 420. In other words, the processor 420 can call and run the computer program from the memory 410 to implement the image semantic segmentation method in the embodiments of the present application.

[0150] For example, the processor 420 can be used to execute the above image semantic segmentation method embodiments according to the instructions in the computer program.

[0151] In some embodiments of the present application, the processor 420 may include, but is not limited to:

[0152] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and so on.

[0153] In some embodiments of the present application, the memory 410 includes, but is not limited to:

[0154] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be Read-Only Memory (ROM), Programmable ROM (PROM), Erasable PROM (EPROM), Electrically Erasable PROM (EEPROM), or flash memory. The volatile memory can be Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double DataRate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), synch link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).

[0155] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 410 and executed by the processor 420 to complete the image semantic segmentation method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0156] As Figure 11 shown, the electronic device 400 may further include:

[0157] A transceiver 430, which may be connected to the processor 420 or the memory 410.

[0158] Among them, the processor 420 can control the transceiver 430 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices. The transceiver 430 may include a transmitter and a receiver. The transceiver 430 may further include an antenna, and the number of antennas may be one or more.

[0159] It should be understood that the various components in the electronic device are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0160] The present application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the image semantic segmentation method in the above method embodiment.

[0161] The embodiment of the present application also provides a computer program product including program instructions. When the program instructions run on an electronic device, the electronic device executes the image semantic segmentation method in the above method embodiment.

[0162] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, a data center, etc. that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0163] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0164] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0165] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of this application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0166] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.

[0167] As mentioned above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. An image semantic segmentation method, characterized in that, Including: Obtain target pose data and a target image; Input the target pose data and the target image into an image semantic segmentation network to obtain a semantic segmentation result of the target image output by the image semantic segmentation network.

2. The method according to claim 1, wherein The method further includes: Perform alignment processing on the target pose data and the target image; Correspondingly, the inputting the target pose data and the target image into the image semantic segmentation network includes: Input the target pose data and the target image after alignment processing into the image semantic segmentation network.

3. The method according to claim 1, wherein The image semantic segmentation network includes a feature extraction module and a semantic segmentation module. The inputting the target pose data and the target image into the image semantic segmentation network to obtain a semantic segmentation result of the target image output by the image semantic segmentation network includes: Input the target pose data into the feature extraction module to obtain target pose features output by the feature extraction module; Input the target pose features and the target image into the semantic segmentation module, and perform feature fusion on the target pose features and the target image through the semantic segmentation module to obtain a semantic segmentation result of the target image.

4. The method according to claim 3, wherein The performing feature fusion on the target pose features and the target image through the semantic segmentation module to obtain a semantic segmentation result of the target image includes: Extract a feature map of the target image through the semantic segmentation module, splice the feature map and the target pose features, and obtain a semantic segmentation result of the target image based on the splicing result.

5. The method according to claim 3, wherein The network structure of the semantic segmentation module is an encoder-decoder structure.

6. The method according to claim 3, characterized in that The image semantic segmentation network further includes a feature processing module. The feature processing module is used to perform normalization processing on the target pose features. The method further includes: Input the target pose features into the feature processing module to obtain normalized target pose features output by the feature processing module; Correspondingly, the inputting the target pose data and the target image into the image semantic segmentation network to obtain a semantic segmentation result of the target image output by the image semantic segmentation network includes: Input the normalized target pose features and the target image into the image semantic segmentation network to obtain a semantic segmentation result of the target image output by the image semantic segmentation network.

7. The method according to any one of claims 1-6, characterized in that, The method further includes: Obtain training pose data and a training image corresponding to the training pose data. The training image includes semantic segmentation information; Input the training pose data and the training image into an image semantic segmentation network to be trained, and train the image semantic segmentation network to be trained to obtain the image semantic segmentation network.

8. The method according to claim 7, wherein The inputting the training pose data and the training image into the image semantic segmentation network to be trained and training the image semantic segmentation network to be trained includes: Input the training pose data into the feature extraction module to obtain training pose features output by the feature extraction module; Input the training pose features and the training image into the semantic segmentation module, and perform feature fusion on the training pose features and the training image through the semantic segmentation module to obtain a predicted semantic segmentation result; Train the image semantic segmentation network based on the predicted semantic segmentation result and the semantic segmentation information of the training image.

9. The method according to claim 8, wherein The training of the image semantic segmentation network based on the predicted semantic segmentation result and the semantic segmentation information of the training image includes: Determine the error between the predicted semantic segmentation result and the semantic segmentation information of the training image; Iteratively optimize the parameters of the feature extraction module and the semantic segmentation module based on the error until the stopping condition is met.

10. The method according to claim 8, wherein The method further includes: Input the training pose features into the feature processing module to obtain the normalized training pose features output by the feature processing module ; Correspondingly, the input of the training pose data and the training image into the image semantic segmentation network to be trained for training the image semantic segmentation network to be trained includes: Input the normalized training pose features and the training image into the image semantic segmentation network to be trained for training the image semantic segmentation network to be trained.

11. An image semantic segmentation device, characterized in that, It includes: An acquisition module for acquiring target pose data and a target image; A processing module for inputting the target pose data and the target image into the image semantic segmentation network to obtain a semantic segmentation result of the target image output by the image semantic segmentation network.

12. An electronic device, characterized in that, It includes: A processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causes a computer to execute the method according to any one of claims 1 to 10.