Image region processing method and apparatus
By acquiring device posture and posture correction methods, the problem of difficulty in distinguishing ceilings, walls, and floors when shooting at close range is solved, improving the accuracy of image region recognition, especially in indoor scenes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2021-09-30
- Publication Date
- 2026-05-15
AI Technical Summary
In image recognition, ceilings, walls, and floors are difficult to distinguish when photographed at close range, and existing deep learning algorithms and traditional edge detection algorithms are not very accurate in this situation.
By acquiring the target image and the terminal device's posture, and performing initial recognition using a deep learning model, the recognition results are corrected based on the device's posture. This includes comparing the terminal's elevation and depression angles with angle thresholds and adjusting the region probability of pixels to improve recognition accuracy.
It improves the accuracy of ceiling, wall and floor area recognition, especially when shooting at close range, enhancing the precision of image area recognition.
Smart Images

Figure CN115908792B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to an image region processing method and apparatus. Background Technology
[0002] Deep learning algorithms based on convolutional neural networks can achieve end-to-end learning and have good performance, showing broad application prospects in the field of image recognition. Among these applications, the recognition of indoor areas such as ceilings, walls, and floors is one of the research directions in image recognition.
[0003] In close-up shots, ceilings, walls, and floors appear very similar due to the lack of reference points, making it difficult for deep learning algorithms to distinguish between them. Traditional algorithms based on edge detection and spatial geometry rely on boundaries for plane segmentation, resulting in smooth planes. However, these methods are prone to failure in some videos or images with blurred edges.
[0004] It is evident that the accuracy of image region recognition in images needs to be improved. Summary of the Invention
[0005] This disclosure provides an image region processing method and apparatus to overcome the problem of low accuracy in image region recognition.
[0006] In a first aspect, embodiments of this disclosure provide an image region processing method, including:
[0007] Acquire the target image and the device posture of the terminal when capturing the target image;
[0008] The target image is subjected to image region recognition to obtain an initial recognition result, wherein the image region includes at least one of the ceiling region, wall region and floor region;
[0009] The initial recognition result is corrected based on the device posture to obtain the corrected recognition result.
[0010] In a second aspect, embodiments of this disclosure provide an image processing apparatus, comprising:
[0011] An acquisition unit is used to acquire a target image and the device posture of the terminal when capturing the target image;
[0012] The recognition unit is used to perform image region recognition on the target image to obtain an initial recognition result, wherein the image region includes at least one of the ceiling region, wall region and floor region;
[0013] The correction unit is used to correct the initial recognition result according to the device posture to obtain the corrected recognition result.
[0014] Thirdly, embodiments of this disclosure provide an electronic device, including: at least one processor and a memory;
[0015] The memory stores computer-executed instructions;
[0016] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the image region processing method as described in the first aspect or various possible designs of the first aspect.
[0017] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image region processing method described in the first aspect or various possible designs of the first aspect.
[0018] Fifthly, according to one or more embodiments of the present disclosure, a computer program product is provided, the computer program product comprising computer execution instructions that, when executed by a processor, implement the image region processing method as described in the first aspect or various possible designs of the first aspect.
[0019] The image region processing method and apparatus provided in this embodiment perform image region recognition on a target image to obtain an initial recognition result. The image region includes at least one of a ceiling region, a wall region, and a floor region. The initial recognition result is then corrected based on the device posture to obtain a corrected recognition result. Therefore, by initially recognizing the image region and then correcting the initial recognition result based on the corresponding device posture, the accuracy of image region recognition is improved. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram illustrating an application scenario to which the embodiments of this disclosure are applicable;
[0022] Figure 2 Flowchart of the image region processing method provided in the embodiments of this disclosure Figure 1 ;
[0023] Figure 3 Flowchart of the image region processing method provided in the embodiments of this disclosure Figure 2 ;
[0024] Figure 4 This is a structural block diagram of the image region processing device provided in the embodiments of this disclosure;
[0025] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0027] When identifying image regions such as ceilings, walls, and floors in an image, there are several methods:
[0028] Method 1 involves using a deep learning model based on convolutional neural networks to identify image regions. Deep learning models can be trained end-to-end and have good performance. However, in close-up shots, due to the lack of reference objects, the ceiling, wall, and ground regions are highly similar, making it difficult for deep learning models to distinguish between them.
[0029] Method two involves identifying image regions using algorithms based on edge detection and spatial geometry. However, this method has lower accuracy in identifying image regions for videos or images with blurred edges.
[0030] To improve the accuracy of image region recognition, this disclosure provides an image region processing method and apparatus. In this method, considering that the device posture of the terminal in capturing the image has a significant impact on the positional distribution of ceiling, wall, and floor regions in the image, after performing image region recognition, the image region recognition results are corrected based on the device posture when the image was captured, thereby improving the accuracy of recognizing ceiling, wall, and floor regions in the image.
[0031] refer to Figure 1 , Figure 1 This is a schematic diagram illustrating the application scenarios to which the embodiments of this disclosure apply.
[0032] Figure 1The application scenario shown is an image processing scenario. In this scenario, the device involved includes an image processing device 101, used for image region recognition. The image processing device 101 can be either a terminal or a server. Figure 1 Take the Sino-Singapore server as an example.
[0033] Optionally, the devices involved in the application scenario also include an image capturing device 102, used to capture images and send the captured images to an image processing device 101.
[0034] Among them, the image capturing device 102 is a terminal with camera function, such as: camera, handheld device with camera (e.g., smartphone, tablet), computing device with camera (e.g., personal computer (PC)), wearable device with camera (e.g., smartwatch), smart home device with camera.
[0035] In this configuration, the image processing device 101 and the image capturing device 102 are the same device. For example, the image processing device 101 and the image capturing device 102 are the same smartphone, which captures images and performs real-time or non-real-time image region recognition on the images. Alternatively, the image processing device 101 and the image capturing device 102 are different devices. For example, the image capturing device 102 sends the captured image to the image processing device 101 via a network, and the image processing device 101 performs image scene recognition on the image. For example, the smartphone sends the captured image to a server, and the server performs image region recognition.
[0036] Optionally, the image is a scene image of an indoor scene. Therefore, based on the image region processing method provided in this disclosure embodiment, the accuracy of image region recognition for indoor scene images where ceiling, wall, and floor areas are difficult to identify can be improved.
[0037] For example, the image region processing method provided in this disclosure can be applied to electronic devices, such as terminals and servers. The terminal can be a personal digital assistant (PDA) device, a handheld device (e.g., a smartphone, tablet), a computing device (e.g., a personal computer), an in-vehicle device, a wearable device (e.g., a smartwatch, a smart bracelet), or a smart home device (e.g., a smart display device). The server can be a distributed server, a centralized server, a cloud server, etc.
[0038] refer to Figure 2 , Figure 2 Flowchart of the image region processing method provided in the embodiments of this disclosure Figure 1 .like Figure 2As shown, the image region processing method includes:
[0039] S201. Obtain the target image and the device posture of the terminal when capturing the target image.
[0040] Device posture includes device angle (also known as device orientation). In the same scene, different device postures of the terminal will result in different image content. The target image can be an image captured in real time by the terminal, an image stored in a database, or an image input by the user. The target image can also be a video frame from a target video, which can be a video captured in real time by the terminal, a video stored in a database, or a video input by the user.
[0041] In one example, the target image captured in real time by the terminal and the device posture of the terminal when capturing the target image are obtained in order to perform real-time image region recognition on the target image.
[0042] In another example, the target image and the device posture of the terminal when capturing the target image are obtained from an online or offline database to perform image region recognition on the target image pre-stored in the database. The target image is obtained in the following ways: randomly or sequentially from the database, or by obtaining a user-specified target image from the database.
[0043] In another example, the target image input by the user and the device posture of the terminal when the target image was captured are obtained to perform image region recognition on the target image input by the user.
[0044] S202. Perform image region recognition on the target image to obtain an initial recognition result. The image region includes at least one of the ceiling region, wall region, and floor region.
[0045] The initial recognition result may include the image region identified in the target image. Specifically, the initial recognition result includes at least one of the identified ceiling region, wall region, and floor region.
[0046] In this embodiment, an image recognition model can be used to identify at least one image region among the ceiling region, wall region, and floor region of the target image, resulting in an identified target image containing at least one of the ceiling region, wall region, and floor region. The image recognition model is a deep learning model for image region recognition, such as a convolutional neural network.
[0047] Optionally, before performing image region recognition on the target image, preprocessing operations are performed on the target image to ensure that the target image meets the requirements of the deep learning model for input data and to improve the image quality of the target image. The preprocessing operations include one or more of the following: resizing, cropping, flipping, and image enhancement operations. Image enhancement operations include enhancement of one or more aspects of image contrast, image saturation, and image tone.
[0048] As an example, the preprocessing of the target image includes: first, randomly scaling the target image within a preset multiple range; then, randomly cropping the scaled target image to the target size; next, randomly flipping the cropped target image horizontally; and finally, performing data enhancement on the contrast, saturation, and hue of the flipped target image.
[0049] S203. Correct the initial recognition result according to the device posture to obtain the corrected recognition result.
[0050] In this embodiment, because the ceiling, wall, and floor areas are relatively similar in planar shape, especially the ceiling and floor areas, recognition errors may occur when the target image is identified using an image recognition model. For example, the ceiling area may be misidentified as the floor area, and vice versa. Considering that the device posture when the terminal captures the target image has a significant impact on the image position distribution of the ceiling, wall, and floor areas in the target image, after obtaining the initial recognition result, the misidentified areas in the initial recognition result can be corrected based on the device posture when the target image was captured, thereby improving the accuracy of image region recognition.
[0051] For example, one or more factors, such as the device's distance from the ground and tilt angle, during the capture of a target image can affect whether the terminal can capture the ceiling, walls, and ground, thus affecting whether the target image contains ceiling, wall, or ground areas. If the device's posture determines that the terminal cannot capture the ground during the capture of the target image, then the target image should not contain ground areas. In this case, if the initial recognition result includes the ground areas identified in the target image, then the ground areas in the initial recognition result are misidentified areas.
[0052] Optionally, when correcting misidentified areas in the initial recognition results, the misidentified areas can be corrected based on the device posture of the terminal when capturing the target image. The correction can be made to identify the most likely image area among the remaining two regions (ceiling, wall, and ground) excluding the misidentified area, thereby improving the accuracy of image region recognition. For example, if the ground area is identified as a misidentified area in the initial recognition results, and the device posture when capturing the target image determines that the most likely image area in the target image is the ceiling area, then the ground area identified in the initial recognition results will be corrected to the ceiling area. Similarly, if the device posture when capturing the target image determines that the most likely image area in the target image is the wall area, then the ground area identified in the initial recognition results will be corrected to the wall area.
[0053] Optionally, when correcting misidentified regions in the initial recognition results, the target image can be re-identified using the image recognition model used in the initial recognition or a more complex deep learning model. Based on the recognition results obtained from the re-recognition, the misidentified regions in the initial recognition results can be corrected, thereby improving the accuracy of image region recognition.
[0054] In this embodiment of the disclosure, based on the initial recognition result obtained by performing image region recognition on the target image through an image recognition model, the initial recognition result is corrected based on the shooting posture of the terminal when capturing the image. This solves the problem of low recognition accuracy when recognizing one or more of the ceiling region, wall region, and ground region of the image, and improves the accuracy of image region recognition.
[0055] In some embodiments, the device posture of the terminal includes the terminal's elevation angle and / or depression angle. The elevation angle and depression angle refer to the angle between the shooting direction (also understood as the line of sight) of the camera on the terminal and the horizontal line. When the shooting direction of the camera on the terminal is above the horizontal line, the angle between the shooting direction of the camera on the terminal and the horizontal line is the elevation angle of the terminal; when the shooting direction of the camera on the terminal is below the horizontal line, the angle between the shooting direction of the camera on the terminal and the horizontal line is the depression angle of the terminal. Based on this, refer to Figure 3 , Figure 3 Flowchart of the image region processing method provided in the embodiments of this disclosure Figure 2 .like Figure 3 As shown, the image region processing method includes:
[0056] S301. Acquire the target image and the device posture of the terminal when capturing the target image. The device posture of the terminal includes the elevation angle and / or the depression angle of the terminal.
[0057] In this embodiment, the elevation and / or depression angles of the terminal when capturing the target image can be obtained through sensors on the terminal. These sensors can be angle sensors, gravity sensors, etc. For example, when the terminal is a mobile phone, the sensor is an inertial measurement unit (IMU) within the terminal.
[0058] S302. Perform image region recognition on the target image to obtain an initial recognition result. The image region includes at least one of the ceiling region, wall region, and floor region.
[0059] The implementation principle and technical effect of S302 can be referred to the aforementioned embodiments, and will not be repeated here.
[0060] S303. Compare the terminal's elevation angle and / or depression angle with the angle threshold.
[0061] The angle thresholds include an angle threshold for comparison with the terminal's elevation angle and an angle threshold for comparison with the terminal's depression angle. The angle thresholds for comparison with the terminal's elevation angle and the angle thresholds for comparison with the terminal's depression angle can be the same or different.
[0062] In this embodiment, the terminal's elevation angle can be compared with the corresponding angle threshold to obtain a comparison result; and / or, the terminal's depression angle can be compared with the corresponding angle threshold to obtain a comparison result.
[0063] S304. Based on the comparison results, the initial recognition results are corrected to obtain the corrected recognition results.
[0064] In this embodiment, based on the comparison result of the terminal's elevation angle and the corresponding angle threshold, and / or based on the comparison result of the terminal's depression angle and the corresponding angle threshold, it is determined whether there is a misidentified area in the initial recognition result. If so, the misidentified area is corrected to obtain the corrected recognition result.
[0065] In some embodiments, the angle threshold for comparison with the terminal's elevation angle includes a first threshold. In this case, one possible implementation of S304 includes: if the comparison result between the terminal's elevation angle and the first threshold is that the terminal's elevation angle is greater than the first threshold, then the ground area in the initial recognition result is re-recognized as the ceiling area to obtain a corrected recognition result.
[0066] Specifically, if the terminal's elevation angle exceeds a first threshold, it indicates that the camera's shooting direction is tilted significantly towards the ceiling, preventing the terminal from capturing the ground area. In this case, if the initial recognition result includes the identified ground area, it is determined that the ground area in the initial recognition result is a misidentified area. Considering that the planar similarity between the ground area and the ceiling area is higher than that between the ground area and the wall area, the ground area in the initial recognition result is re-identified as the ceiling area, resulting in a corrected recognition result. This improves the accuracy of image region recognition.
[0067] In some embodiments, the angle threshold for comparison with the terminal's elevation angle includes a third threshold, which is greater than the first threshold. In this case, another possible implementation of S304 includes: if the comparison result between the terminal's elevation angle and the third threshold is that the terminal's elevation angle is greater than the third threshold, then the ground area and wall area in the initial recognition result are re-recognized as ceiling area to obtain a corrected recognition result.
[0068] Specifically, the third threshold is larger than the first threshold. If the terminal's elevation angle is greater than the third threshold, it indicates that the camera's shooting direction is tilted more severely towards the ceiling, preventing the terminal from capturing the ground and wall areas. In this case, if the initial recognition result includes the identified ground and / or wall areas, then the ground and wall areas in the initial recognition result are determined to be misidentified areas. The ground and wall areas in the initial recognition result are then identified as ceiling areas, resulting in a corrected recognition result. This improves the accuracy of image region recognition.
[0069] In some embodiments, the angle threshold for comparison with the terminal's tilt angle includes a second threshold, wherein the second threshold may be equal to the first threshold or may be different from the first threshold. In this case, another possible implementation of S304 includes: if the comparison result between the terminal's tilt angle and the second threshold shows that the terminal's tilt angle is greater than the second threshold, then the ceiling area in the initial recognition result is re-identified as a ground area, resulting in a corrected recognition result.
[0070] Specifically, if the terminal's tilt angle exceeds the second threshold, it indicates that the camera's shooting direction is tilted significantly towards the ground, preventing the terminal from capturing the ceiling area. In this case, if the initial recognition result includes the identified ceiling area, it is determined that the ceiling area in the initial recognition result is a misidentified area. Considering that the planar similarity between the ceiling area and the ground area is higher than that between the ceiling area and the wall area, the ceiling area in the initial recognition result is re-identified as the ground area, resulting in a corrected recognition result. This improves the accuracy of image region recognition.
[0071] In some embodiments, the angle threshold compared with the terminal's tilt angle includes a fourth threshold, which is greater than the second threshold. The fourth threshold can be equal to the third threshold or have a different value than the third threshold. In this case, another possible implementation of S304 includes: if the comparison result between the terminal's tilt angle and the fourth threshold is that the terminal's tilt angle is greater than the fourth threshold, then the wall area and ceiling area in the initial recognition result are re-identified as ground areas to obtain a corrected recognition result.
[0072] Specifically, the fourth threshold is larger than the second threshold. If the terminal's tilt angle is greater than the fourth threshold, it indicates that the camera's shooting direction is tilted more severely towards the ground, preventing the terminal from capturing the ceiling and wall areas. In this case, if the initial recognition result includes the identified ceiling and / or wall areas, then the ceiling and wall areas in the initial recognition result are determined to be misidentified areas, and they are reclassified as ground areas, resulting in a corrected recognition result. This improves the accuracy of image region recognition.
[0073] In some embodiments, the initial recognition result may include the region probabilities corresponding to multiple pixels in the target image, where the region probabilities include at least one of ceiling probability, wall probability, and ground probability. Specifically, the ceiling probability corresponding to a pixel is the probability that the region where the pixel is located is the ceiling region; the wall probability corresponding to a pixel is the probability that the region where the pixel is located is the wall region; and the ground probability corresponding to a pixel is the probability that the region where the pixel is located is the ground region. During image region recognition, the region probabilities corresponding to multiple pixels in the target image can be obtained through an image recognition model.
[0074] At this point, another possible implementation of S304 includes: adjusting the region probabilities corresponding to multiple pixels of the target image in the initial recognition result based on the comparison result between the terminal's elevation angle and / or depression angle and the angle threshold; and obtaining the corrected recognition result based on the adjusted region probabilities corresponding to the multiple pixels of the target image. The angle threshold may include one or more angle thresholds for comparison with the terminal's elevation angle and one or more angle thresholds for comparison with the terminal's depression angle, and is not limited to the first, second, third, and fourth thresholds mentioned above.
[0075] Specifically, if the terminal's elevation angle is greater than the angle threshold, it indicates that the probability of the terminal capturing the ceiling is higher than the probability of capturing the ground. In this case, the probability of multiple pixels in the target image corresponding to the ceiling can be increased, and / or the probability of multiple pixels in the target image corresponding to the ground can be decreased. Similarly, if the terminal's depression angle is greater than the angle threshold, it indicates that the probability of the terminal capturing the ceiling is lower than the probability of capturing the ground. In this case, the probability of multiple pixels in the target image corresponding to the ceiling can be decreased, and / or the probability of multiple pixels in the target image corresponding to the ground can be increased. In addition to adjusting the ceiling and / or ground probabilities, the probability of multiple pixels in the target image corresponding to walls can also be adjusted.
[0076] After adjusting the region probabilities, the region where a pixel is located can be determined based on the maximum probability value among the region probabilities corresponding to pixels in the target image. For example, if the ceiling has the highest probability among the region probabilities corresponding to a pixel, then the region where the pixel is located is the ceiling region. Furthermore, based on the regions where multiple pixels are located in the target image, the image regions in the target image can be determined.
[0077] Optionally, there can be multiple angle thresholds, and different angle thresholds can correspond to different probability adjustment amounts.
[0078] The larger the angle threshold, the larger the corresponding probability adjustment. Specifically:
[0079] Multiple angle thresholds are used for comparison with the terminal's elevation angle, and different angle thresholds correspond to different probability adjustment amounts. For example, if the terminal's elevation angle is greater than 45 degrees, the probability of the ceiling corresponding to multiple pixels in the target image is increased by 10% and the probability of the ground corresponding to multiple pixels is decreased by 10%; if the terminal's elevation angle is greater than 60 degrees, the probability of the ceiling corresponding to multiple pixels in the target image is increased by 20% and the probability of the ground corresponding to multiple pixels is decreased by 20%.
[0080] And / or, there are multiple angle thresholds used for comparison with the terminal's tilt angle, and different angle thresholds correspond to different probability adjustment amounts. Examples are not provided here.
[0081] Optionally, if the terminal's elevation angle is greater than the first threshold, considering that the terminal cannot capture the ground area at this time, the ground probability corresponding to multiple pixels in the target image is reduced to 0, and the ground probability before pixel adjustment is added to the ceiling probability corresponding to the pixel.
[0082] Optionally, if the terminal's elevation angle is greater than the third threshold, considering that the terminal cannot capture the ground and wall areas, the ground probability and wall probability corresponding to multiple pixels in the target image are reduced to 0, and the ceiling probability corresponding to the pixel is adjusted to 100%.
[0083] Optionally, if the terminal's tilt angle is greater than the second threshold, considering that the terminal cannot capture the ceiling area at this time, the ceiling probability corresponding to multiple pixels in the target image is reduced to 0, and the ceiling probability before pixel adjustment is added to the ground probability corresponding to the pixel.
[0084] Optionally, if the terminal's tilt angle is greater than the fourth threshold, considering that the terminal cannot capture the ceiling and wall areas, the ceiling and wall probabilities corresponding to multiple pixels in the target image are reduced to 0, and the ground probability corresponding to the pixel is adjusted to 100%.
[0085] The first threshold, second threshold, third threshold and fourth threshold can be referred to in the aforementioned embodiments.
[0086] Optionally, when the region probability corresponding to a pixel includes the ceiling probability, wall probability, and ground probability, the sum of the ceiling probability, wall probability, and ground probability for the same pixel is 1. In this case, while increasing the ceiling probability, it is necessary to decrease the ground probability, and the wall probability can be further decreased.
[0087] Therefore, by adjusting the region probability corresponding to the pixel in the target image based on the comparison results, the flexibility and accuracy of image region correction are improved, and the accuracy of image region recognition is enhanced.
[0088] In this embodiment of the disclosure, based on the initial recognition result obtained by performing image region recognition on the target image through an image recognition model, the initial recognition result is corrected based on the comparison result of the elevation angle and / or depression angle of the terminal when capturing the image with the angle threshold. This solves the problem of low recognition accuracy when recognizing one or more of the ceiling region, wall region, and ground region of the image, and effectively improves the accuracy of image region recognition.
[0089] Based on any of the foregoing embodiments, optionally, the image recognition model is a deep learning model trained by model distillation. This can improve the recognition accuracy of the image recognition model while reducing its size. In particular, a lightweight image recognition model can be trained, which is convenient for deployment on various terminals. Real-time recognition of image regions on images and / or videos can be achieved on the terminal. For example, a user can walk around while shooting a video. While shooting the video, the terminal uses the image region processing method provided in any embodiment to identify the image regions of each video frame, effectively improving the user experience.
[0090] Optionally, when training the image recognition model using model distillation, the teacher model is trained multiple times using the training data and the loss function of the teacher model to obtain a well-trained teacher model. Then, the student model is trained multiple times using the training data, the well-trained teacher model, and the loss function of the student model to obtain a well-trained student model. The well-trained student model is then selected as the image recognition model for image region recognition. The teacher model has a larger model size than the student model, and both the teacher and student models are deep learning models.
[0091] The training data may include multiple training images, which may be pre-labeled with at least one of the following regions: ground, ceiling, and wall. Therefore, when training the teacher model, the difference between the image region recognition result output by the teacher model and the image region labeled on the training image can be determined based on the loss function of the teacher model. Then, the model parameters of the teacher model can be adjusted based on this difference.
[0092] In training the student model, training images can be input into both the student model and the teacher model. Based on the loss function of the student model, the differences between the image region recognition results output by the student model and the image region recognition results output by the teacher model, as well as the image regions marked on the training images, are determined. Then, based on these differences, the model parameters of the student model are adjusted.
[0093] Optionally, the main network structure of the teacher model adopts deeplab v3. By using deeplab v3, the image region recognition accuracy of the teacher model can be improved, thereby improving the image region recognition accuracy of the student model.
[0094] Optionally, the teacher model can use the binary cross-entropy (BCE) loss function to improve the training effect of the teacher model.
[0095] Optionally, the loss function of the teacher model can also be obtained by weighted summation of the BCE loss function and the Regional Mutual Information (RMI) loss function to improve model performance and reduce the occurrence of missed segments in the teacher model.
[0096] Furthermore, in the weighted summation, the BCE loss function and the RMI loss function have the same weight ratio.
[0097] Optionally, the main network structure of the student model (i.e., the image recognition model) adopts the Ghostnet network structure. Ghostnet is a lightweight network structure, which is easy to deploy on lightweight devices. Therefore, using Ghostnet as the main network structure of the student model helps to reduce the model size, facilitates the deployment of the trained student model to user terminals, increases the range of devices the model can be used on, and improves the user experience.
[0098] Optionally, the loss function of the student model can include a weighted loss function obtained by combining the BCE and RMI loss functions, as well as a distillation loss function. The weighted loss function is used to determine the difference between the image region recognition result output by the student model and the image region labeled on the training image, while the distillation loss function is used to determine the difference between the image region recognition result output by the student model and the image region recognition result output by the teacher model. This reduces the chance of the student model missing segments and guides the training of the student model through the trained teacher model, thereby effectively improving the performance of the student model.
[0099] Furthermore, the distillation loss function can employ the Kullback-Leibler (KL) divergence loss function. This KL divergence loss function improves the image segmentation metrics achievable using the student model. In other words, the KL divergence loss function enables the student model to achieve better image segmentation results. One example of an image segmentation metric is the Mean Intersection over Union (MIoU).
[0100] Optionally, based on any of the foregoing embodiments, after obtaining the corrected recognition result, the corrected recognition result can be displayed on the target image. When the executing entity is a server, the server can send the corrected recognition result to the user terminal, which then displays the corrected recognition result on the target image. When the executing entity is a terminal, the terminal can display the image region recognition result corresponding to the video frame or image in real time on the video frame or image while capturing the video or image. Thus, the user can intuitively see the various regions identified on the target image, improving the user experience.
[0101] Furthermore, during display, different image areas can be marked on the target image using different colors to improve the display effect.
[0102] Corresponding to the image region processing method in the above embodiments, Figure 4 This is a structural block diagram of an image region processing device provided in an embodiment of this disclosure. For ease of explanation, only the parts relevant to the embodiments of this disclosure are shown. (Refer to...) Figure 4 The image region processing device includes: an acquisition unit 401, a recognition unit 402, and a correction unit 403.
[0103] The acquisition unit 401 is used to acquire the target image and the device posture of the terminal when capturing the target image;
[0104] The recognition unit 402 is used to perform image region recognition on the target image to obtain an initial recognition result. The image region includes at least one of the ceiling region, wall region, and floor region.
[0105] The correction unit 403 is used to correct the initial recognition result according to the device posture to obtain the corrected recognition result.
[0106] In one embodiment of this disclosure, the device posture includes the terminal's elevation angle and / or the terminal's depression angle. The correction unit 403 is further configured to: compare the terminal's elevation angle and / or the terminal's depression angle with an angle threshold; and correct the initial recognition result based on the comparison result to obtain the corrected recognition result.
[0107] In one embodiment of this disclosure, the correction unit 403 is further configured to: if the comparison result is that the elevation angle of the terminal is greater than the first threshold, then re-identify the ground area in the initial identification result as the ceiling area to obtain the corrected identification result.
[0108] In one embodiment of this disclosure, the correction unit 403 is further configured to: if the comparison result is that the tilt angle of the terminal is greater than the second threshold, then re-identify the ceiling area in the initial identification result as the ground area to obtain the corrected identification result.
[0109] In one embodiment of this disclosure, the correction unit 403 is further configured to: if the comparison result is that the elevation angle of the terminal is greater than the third threshold, then re-identify the ground area and wall area in the initial identification result as the ceiling area to obtain the corrected identification result, wherein the third threshold is greater than the first threshold.
[0110] In one embodiment of this disclosure, the correction unit 403 is further configured to: if the comparison result is that the tilt angle of the terminal is greater than the fourth threshold, then re-identify the wall area and ceiling area in the initial identification result as the ground area to obtain the corrected identification result, wherein the fourth threshold is greater than the second threshold.
[0111] In one embodiment of this disclosure, the initial recognition result includes the region probabilities corresponding to multiple pixels of the target image, and the region probabilities include at least one of ceiling probability, wall probability, and ground probability. The correction unit 403 is further configured to: adjust the region probabilities corresponding to multiple pixels of the target image in the initial recognition result according to the comparison result; and obtain the corrected recognition result according to the adjusted region probabilities corresponding to multiple pixels of the target image.
[0112] In one embodiment of this disclosure, the recognition unit 402 is further configured to: perform image region recognition on the target image using an image recognition model to obtain an initial recognition result, wherein the image recognition model is a deep learning model trained by model distillation.
[0113] In one embodiment of this disclosure, the image region processing device further includes a display unit 404 for displaying the corrected recognition result on the target image.
[0114] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
[0115] refer to Figure 5 The diagram illustrates a structural schematic of an electronic device 500 suitable for implementing embodiments of the present disclosure. The electronic device 500 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0116] like Figure 5As shown, the electronic device 500 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0117] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0118] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0119] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0120] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0121] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.
[0122] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0123] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0124] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0125] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0126] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0127] In a first aspect, according to one or more embodiments of the present disclosure, an image region processing method is provided, comprising: acquiring a target image and a device posture of a terminal when capturing the target image; performing image region recognition on the target image to obtain an initial recognition result, wherein the image region includes at least one of a ceiling region, a wall region, and a ground region; and correcting the initial recognition result according to the device posture to obtain a corrected recognition result.
[0128] According to one or more embodiments of this disclosure, the device posture includes the elevation angle and / or the depression angle of the terminal. The step of correcting the initial recognition result based on the device posture to obtain a corrected recognition result includes: comparing the elevation angle and / or the depression angle of the terminal with an angle threshold; and correcting the initial recognition result based on the comparison result to obtain the corrected recognition result.
[0129] According to one or more embodiments of this disclosure, the step of correcting the initial identification result based on the comparison result to obtain the corrected identification result includes: if the comparison result is that the elevation angle of the terminal is greater than a first threshold, then the ground area in the initial identification result is re-identified as the ceiling area to obtain the corrected identification result.
[0130] According to one or more embodiments of this disclosure, the step of correcting the initial identification result based on the comparison result to obtain the corrected identification result includes: if the comparison result is that the tilt angle of the terminal is greater than a second threshold, then the ceiling area in the initial identification result is re-identified as the ground area to obtain the corrected identification result.
[0131] According to one or more embodiments of this disclosure, the step of correcting the initial identification result based on the comparison result to obtain the corrected identification result includes: if the comparison result is that the elevation angle of the terminal is greater than a third threshold, then the ground area and wall area in the initial identification result are re-identified as ceiling area to obtain the corrected identification result, wherein the third threshold is greater than the first threshold.
[0132] According to one or more embodiments of this disclosure, the step of correcting the initial identification result based on the comparison result to obtain the corrected identification result includes: if the comparison result is that the tilt angle of the terminal is greater than a fourth threshold, then the wall area and ceiling area in the initial identification result are re-identified as ground areas to obtain the corrected identification result, wherein the fourth threshold is greater than the second threshold.
[0133] According to one or more embodiments of this disclosure, the initial recognition result includes region probabilities corresponding to multiple pixels of the target image, wherein the region probabilities include at least one of ceiling probability, wall probability, and ground probability. The step of correcting the initial recognition result based on a comparison result to obtain the corrected recognition result includes: adjusting the region probabilities corresponding to multiple pixels of the target image in the initial recognition result based on the comparison result; and obtaining the corrected recognition result based on the adjusted region probabilities corresponding to the multiple pixels of the target image.
[0134] According to one or more embodiments of this disclosure, the step of performing image region recognition on the target image to obtain an initial recognition result includes: performing image region recognition on the target image using an image recognition model to obtain the initial recognition result, wherein the image recognition model is a deep learning model trained by model distillation.
[0135] According to one or more embodiments of this disclosure, after obtaining the corrected recognition result, the method further includes: displaying the corrected recognition result on the target image.
[0136] Secondly, according to one or more embodiments of this disclosure, an image region processing device is provided, comprising: an acquisition unit for acquiring a target image and a device posture of a terminal when capturing the target image; an identification unit for performing image region identification on the target image to obtain an initial identification result, wherein the image region includes at least one of a ceiling region, a wall region, and a ground region; and a correction unit for correcting the initial identification result according to the device posture to obtain a corrected identification result.
[0137] According to one or more embodiments of this disclosure, the device attitude includes the elevation angle and / or the depression angle of the terminal, and the correction unit is further configured to: compare the elevation angle and / or the depression angle of the terminal with an angle threshold; and correct the initial recognition result based on the comparison result to obtain the corrected recognition result.
[0138] According to one or more embodiments of this disclosure, the correction unit is further configured to: if the comparison result is that the elevation angle of the terminal is greater than a first threshold, then re-identify the ground area in the initial identification result as the ceiling area to obtain the corrected identification result.
[0139] According to one or more embodiments of this disclosure, the correction unit is further configured to: if the comparison result is that the tilt angle of the terminal is greater than a second threshold, then re-identify the ceiling area in the initial identification result as the ground area to obtain the corrected identification result.
[0140] According to one or more embodiments of this disclosure, the correction unit is further configured to: if the comparison result is that the elevation angle of the terminal is greater than a third threshold, then re-identify the ground area and wall area in the initial identification result as the ceiling area to obtain the corrected identification result, wherein the third threshold is greater than the first threshold.
[0141] According to one or more embodiments of this disclosure, the correction unit is further configured to: if the comparison result is that the tilt angle of the terminal is greater than a fourth threshold, then re-identify the wall area and ceiling area in the initial identification result as ground areas to obtain the corrected identification result, wherein the fourth threshold is greater than the second threshold.
[0142] According to one or more embodiments of this disclosure, the initial recognition result includes region probabilities corresponding to multiple pixels of the target image, the region probabilities including at least one of ceiling probability, wall probability, and ground probability, and the correction unit is further configured to: adjust the region probabilities corresponding to multiple pixels of the target image in the initial recognition result according to the comparison result; and obtain the corrected recognition result according to the adjusted region probabilities corresponding to multiple pixels of the target image.
[0143] According to one or more embodiments of this disclosure, the recognition unit is further configured to: perform image region recognition on the target image using an image recognition model to obtain the initial recognition result, wherein the image recognition model is a deep learning model trained by model distillation.
[0144] According to one or more embodiments of this disclosure, the image region processing device further includes: a display unit for displaying the corrected recognition result on the target image.
[0145] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory;
[0146] The memory stores computer-executed instructions;
[0147] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the image region processing method as described in the first aspect or various possible designs of the first aspect.
[0148] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when a processor executes the computer-executable instructions, the image region processing method described in the first aspect or various possible designs of the first aspect is implemented.
[0149] Fifthly, according to one or more embodiments of the present disclosure, a computer program product is provided, the computer program product comprising computer execution instructions that, when executed by a processor, implement the image region processing method as described in the first aspect or various possible designs of the first aspect.
[0150] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0151] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0152] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image region processing method, comprising: Acquire a target image and the device posture of the terminal when capturing the target image, wherein the device posture includes the elevation angle and / or the depression angle of the terminal; The target image is subjected to image region recognition to obtain an initial recognition result. The image region includes at least one of a ceiling region, a wall region, and a ground region. The initial recognition result includes the region probability corresponding to multiple pixels of the target image. The region probability includes at least one of a ceiling probability, a wall probability, and a ground probability. Compare the elevation angle and / or depression angle of the terminal with an angle threshold; The initial identification result is corrected based on the comparison result to obtain the corrected identification result; The step of correcting the initial recognition result based on the comparison result to obtain the corrected recognition result includes: Based on the comparison results, the region probabilities corresponding to multiple pixels in the target image are adjusted in the initial recognition results; The corrected recognition result is obtained based on the adjusted region probabilities corresponding to multiple pixels in the target image.
2. The image region processing method according to claim 1, wherein correcting the initial recognition result based on the comparison result to obtain a corrected recognition result further includes: If the comparison result indicates that the elevation angle of the terminal is greater than the first threshold, then the ground area in the initial recognition result is re-identified as the ceiling area to obtain the corrected recognition result.
3. The image region processing method according to claim 2, wherein correcting the initial recognition result based on the comparison result to obtain a corrected recognition result further includes: If the comparison result indicates that the tilt angle of the terminal is greater than the second threshold, then the ceiling area in the initial recognition result is re-recognized as the ground area to obtain the corrected recognition result.
4. The image region processing method according to claim 2, wherein correcting the initial recognition result based on the comparison result to obtain a corrected recognition result further includes: If the comparison result indicates that the elevation angle of the terminal is greater than the third threshold, then the ground area and wall area in the initial recognition result are re-identified as the ceiling area to obtain the corrected recognition result, wherein the third threshold is greater than the first threshold.
5. The image region processing method according to claim 3, wherein correcting the initial recognition result based on the comparison result to obtain a corrected recognition result further includes: If the comparison result shows that the tilt angle of the terminal is greater than the fourth threshold, then the wall area and ceiling area in the initial recognition result are re-recognized as ground areas to obtain the corrected recognition result, wherein the fourth threshold is greater than the second threshold.
6. The image region processing method according to any one of claims 1 to 4, wherein performing image region recognition on the target image to obtain an initial recognition result includes: The target image is used to identify image regions by an image recognition model to obtain the initial recognition result. The image recognition model is a deep learning model trained by model distillation.
7. The image region processing method according to any one of claims 1 to 4, further comprising, after obtaining the corrected recognition result: The corrected recognition result is displayed on the target image.
8. An image region processing device, comprising: An acquisition unit is used to acquire a target image and the device posture of the terminal when capturing the target image, wherein the device posture includes the elevation angle and / or the depression angle of the terminal; The recognition unit is used to perform image region recognition on the target image to obtain an initial recognition result. The image region includes at least one of a ceiling region, a wall region, and a ground region. The initial recognition result includes the region probabilities corresponding to multiple pixels of the target image. The region probabilities include at least one of a ceiling probability, a wall probability, and a ground probability. The correction unit is used to compare the elevation angle and / or depression angle of the terminal with an angle threshold. The initial identification result is corrected based on the comparison result to obtain the corrected identification result; The correction unit is specifically used to adjust the region probabilities corresponding to multiple pixels of the target image in the initial recognition result according to the comparison result; and to obtain the corrected recognition result according to the adjusted region probabilities corresponding to multiple pixels of the target image.
9. An electronic device, comprising: At least one processor and memory; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the image region processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the image region processing method as described in any one of claims 1 to 7.
11. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the image region processing method as described in any one of claims 1 to 7.