Smart home external visual perception method and device based on semantic segmentation
By using a semantic segmentation-based method for external visual perception of smart homes, the problems of insufficient static structure recognition and low edge computing efficiency in outdoor environment identification of smart home systems are solved, achieving efficient and accurate semantic segmentation and improved user experience.
Patent Information
- Application Number
- CN202510944010.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-31
AI Technical Summary
Existing smart home systems lack the ability to recognize static structures in outdoor environments, cannot achieve pixel-level boundary extraction and type recognition, and have low inference efficiency on edge computing devices, resulting in an inadequate user experience.
We adopt a semantic segmentation-based method for external visual perception of smart homes. We extract features through morphological enhancement, shallow and deep separable convolution, and mid-to-deep channel compression strategies. Combined with upsampling fusion and adaptive encoding, we use a hybrid loss mechanism for correction to achieve efficient and accurate semantic segmentation.
It achieves accurate identification of small-scale occluded targets, supports pixel-level boundary extraction and type recognition, and improves user experience and inference efficiency.
Smart Images

Figure CN120876335A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this application relate to the field of image processing, and more particularly to a method, apparatus, device, and computer-readable storage medium for smart home external visual perception based on semantic segmentation. Background Technology
[0002] As smart home systems are increasingly deployed outdoors and at the edge, the system's ability to understand its surroundings has become a key element in enhancing user experience. For example, smart doorbell cameras, home perimeter security cameras, and unattended access control devices all require a certain level of external environment recognition capabilities.
[0003] Current outdoor sensing technologies have the following shortcomings: 1. Home vision systems mostly focus on detecting "people / vehicles / animals" and lack the ability to recognize static structures in the environment such as "buildings, roads, and fences"; 2. Traditional edge detection, optical flow variation, or YOLO target detection methods do not support pixel-level boundary extraction and type recognition, making it difficult to achieve environmental modeling; 3. Unable to identify small-scale occluded targets (such as distant building outlines, slopes, and platforms), resulting in misjudgments of height; 4. On terminal devices with limited edge computing capabilities, model inference efficiency is low, making it impossible to deploy complex segmentation models; 5. Lack of visual feedback capabilities, users cannot clearly perceive what the device "sees".
[0004] In summary, how to construct an external environment recognition technology solution that can run on edge devices and has efficient and accurate semantic segmentation capabilities is an urgent problem to be solved. Summary of the Invention
[0005] According to embodiments of this application, a smart home external visual perception solution based on semantic segmentation is provided, which can operate on edge devices, has efficient and accurate semantic segmentation capabilities, and significantly improves the user experience.
[0006] In a first aspect of this application, a method for external visual perception of smart homes based on semantic segmentation is provided. The method includes: Obtain the original image to be processed; The original image is morphologically enhanced to obtain an enhanced image; The enhanced image is subjected to semantic segmentation processing to obtain a segmented image; The segmented image is corrected using a predefined hybrid loss mechanism to obtain the target image.
[0007] Furthermore, the morphological enhancement of the original image to obtain the enhanced image includes: The original image is converted into an image in a three-channel visual tensor format. The transformed image is enhanced using a predefined composite enhancement function to obtain an enhanced image.
[0008] Furthermore, the composite enhancement function includes:
[0009] Where V is an image in three-channel visual tensor format; M stands for reverse mirror; For angle transformation; Light disturbance; Random occlusion; It is an elastic deformation.
[0010] Further, the semantic segmentation processing of the enhanced image to obtain a segmented image includes: Based on shallow depth-separable convolution and mid-to-deep channel compression strategies, feature extraction is performed on the enhanced image to obtain feature data; The image features are analyzed hierarchically using upsampling fusion to obtain the analyzed data; The parsed data is subjected to confidence output and adaptive encoding to obtain a segmented image.
[0011] Furthermore, the feature extraction of the enhanced image based on the shallow depth-separable convolution and mid-to-deep channel compression strategy yields feature data including:
[0012] in, Features that have undergone attention adjustment; For global evaluation pooling; For activation functions; Feature compression path:
[0013] in, This refers to a convolutional block that includes an activation function.
[0014] Furthermore, the upsampling fusion method is used to perform hierarchical parsing of the image features to obtain parsed data, including:
[0015] in, This is either a transposed convolution or a nearest neighbor interpolation; For splicing channels; It is a 3×3 convolutional module of the j-th layer.
[0016] Furthermore, the predefined hybrid loss mechanism includes:
[0017] in, A predefined regional overlap index; As weight; Loss function: =-
[0018] Where N is the total number of pixels; The true label of the i-th pixel; Predict the probability of the model for the i-th pixel; α is the prospect weighting factor; It is a very small value.
[0019] In a second aspect of this application, a smart home external visual sensing device based on semantic segmentation is provided. The device includes: The acquisition module is used to acquire the original image to be processed; The enhancement module is used to perform morphological enhancement on the original image to obtain an enhanced image; The segmentation module is used to perform semantic segmentation processing on the enhanced image to obtain a segmented image; The correction module is used to correct the segmented image using a predefined hybrid loss mechanism to obtain the target image.
[0020] In a third aspect of this application, an electronic device is provided. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described above.
[0021] In a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to the first aspect of this application.
[0022] The semantic segmentation-based smart home external visual perception method provided in this application involves: acquiring an original image to be processed; performing morphological enhancement on the original image to obtain an enhanced image; performing semantic segmentation on the enhanced image to obtain a segmented image; and correcting the segmented image through a predefined hybrid loss mechanism to obtain a target image. This method can run on edge devices, has efficient and accurate semantic segmentation capabilities, and significantly improves the user experience.
[0023] It should be understood that the description in the Summary Section is not intended to limit the key or essential features of the embodiments of this application, nor is it intended to restrict the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0024] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 A flowchart of a semantic segmentation-based smart home external visual perception method according to an embodiment of this application; Figure 2 This is a schematic diagram of a segmented image according to an embodiment of this application; Figure 3 A block diagram of a semantic segmentation-based smart home external visual sensing device according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of a terminal device or server suitable for implementing the embodiments of this application. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0026] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0027] Figure 1A flowchart of a smart home external visual perception method based on semantic segmentation according to an embodiment of the present disclosure is shown. The method includes: S110, Obtain the original image to be processed.
[0028] In some embodiments, the raw image to be processed can be acquired through smart home devices. For example, outdoor environmental images can be captured through smart home devices such as cameras and doorbells.
[0029] S120, perform morphological enhancement on the original image to obtain an enhanced image.
[0030] In some embodiments, the original image may be normalized to a size of 512x512. Morphological enhancement is then performed on the normalized image to improve the model's generalization ability, resulting in an enhanced image.
[0031] Specifically, the normalized image is converted into a three-channel visual tensor format:
[0032] The transformed image is enhanced using the following predefined composite enhancement function to obtain the enhanced image:
[0033] Where V is an image in three-channel visual tensor format; M stands for flip mirror (horizontal / vertical flip). For angle transformation, the transformation range ; For illumination perturbation, the pixel is multiplied by a coefficient. ; Random occlusion; For elastic deformation, used to simulate local elastic deformation; .
[0034] S130, perform semantic segmentation processing on the enhanced image to obtain a segmented image.
[0035] In some embodiments, the enhanced image can be semantically segmented through feature extraction -> hierarchical parsing -> confidence output -> adaptive encoding -> visualization to obtain a segmented image.
[0036] Specifically, feature extraction includes: Based on shallow depth-separable convolution and mid-to-deep channel compression strategies, feature extraction is performed on the enhanced image to obtain feature data:
[0037] in, These are attention-adjusted features used to improve the network's responsiveness to key channel features and enhance its representation capabilities. For global evaluation pooling; that is, perform global average pooling to obtain a channel description vector [C,1,1][C, 1, 1][C,1,1] (number of channels, height, width); For activation functions; This is a dimension reduction transformation (channel compression) used to compress and fuse image information, reducing the number of parameters; For dimensionality restoration (channel recovery), it is used to reproject the fused information back into the channel space; Feature compression path:
[0038] in, For convolutional blocks that include activation functions; To replace MaxPool with convolutions with a stride of 2; Hierarchical resolution includes: The image features are analyzed hierarchically using upsampling fusion to obtain the following parsed data:
[0039] in, This is either a transposed convolution or a nearest neighbor interpolation; For splicing channels; It is a 3×3 convolutional module of the j-th layer; like This indicates that the decoder input is initialized to the deepest output of the encoder; Confidence outputs include: The output can be trusted in the following way:
[0040] This represents the probability that a pixel belongs to the target. For activation functions; H is the height of the pixel (512); W is the width of pixels (512); This is the last layer feature map of the decoder, with the same spatial dimensions as the input image (based on z-adjusted feature fusion to ensure the accuracy of the confidence output). The probability that each pixel is in the foreground; Adaptive encoding and visualization include: Define an adaptive threshold:
[0041] in, The confidence level is the mean. Standard deviation; To control sensitivity; Mask generation yields a segmented image, such as... Figure 2 As shown:
[0042] That is, the semantic segmentation region is obtained.
[0043] S140, the segmented image is corrected using a predefined hybrid loss mechanism to obtain the target image.
[0044] In some embodiments, the segmented image can be corrected through a predefined hybrid loss mechanism to improve the accuracy of foreground edge and small target recognition, thereby obtaining the target image.
[0045] The hybrid loss mechanism includes:
[0046] in, A predefined regional overlap index; As weight; This is a binary classification divergence loss function used to control the bias caused by imbalanced foreground and background samples: =-
[0047] Where N is the total number of pixels; The true label of the i-th pixel; Predict the probability of the model for the i-th pixel; α is the prospect weighting factor; It is a very small value.
[0048] According to the embodiments of this disclosure, the following technical effects are achieved: A multi-level feature fusion design is adopted to preserve shallow edge information and combine it with deep contextual features, significantly improving spatial reconstruction accuracy. A custom hybrid loss mechanism is used to improve the accuracy of foreground edge and small target recognition.
[0049] In summary, the semantic segmentation-based smart home external visual perception method disclosed herein can accurately identify environmental targets (including small-scale occluded targets), supports pixel-level boundary extraction and type recognition, and has high inference efficiency. While improving the ability to understand the surrounding environment, it also significantly enhances the user experience.
[0050] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0051] The above is an introduction to the method embodiments. The following describes the solution described in this application through device embodiments.
[0052] Figure 3 A block diagram of a smart home external visual perception method apparatus 300 based on semantic segmentation according to an embodiment of this application is shown, as follows: Figure 3 The following are included: The acquisition module 310 is used to acquire the original image to be processed; Enhancement module 320 is used to perform morphological enhancement on the original image to obtain an enhanced image; The segmentation module 330 is used to perform semantic segmentation processing on the enhanced image to obtain a segmented image; The correction module 340 is used to correct the segmented image through a predefined hybrid loss mechanism to obtain the target image.
[0053] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0054] Figure 4 A schematic diagram of a terminal device or server suitable for implementing embodiments of this application is shown.
[0055] like Figure 4As shown, the terminal device or server includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage section 408 into random access memory (RAM) 403. The RAM 403 also stores various programs and data required for the operation of the terminal device or server. The CPU 401, ROM 402, and RAM 403 are interconnected via bus 404. An input / output (I / O) interface 405 is also connected to bus 404.
[0056] The following components are connected to I / O interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to I / O interface 405 as needed. A removable medium 411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 410 as needed so that computer programs read from it can be installed into storage section 408 as needed.
[0057] Specifically, according to embodiments of this application, the above method flow steps can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit (CPU) 401, it performs the functions defined in the system of this application.
[0058] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0059] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0060] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be located in a processor. The names of these units or modules do not, in certain circumstances, constitute a limitation on the unit or module itself.
[0061] In another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the methods described in this application.
[0062] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions claimed in this application.
Claims
1. A method for external visual perception of smart homes based on semantic segmentation, characterized in that, include: Obtain the original image to be processed; The original image is morphologically enhanced to obtain an enhanced image; The enhanced image is subjected to semantic segmentation processing to obtain a segmented image; The segmented image is corrected using a predefined hybrid loss mechanism to obtain the target image.
2. The method according to claim 1, characterized in that, The morphological enhancement of the original image to obtain the enhanced image includes: The original image is converted into an image in a three-channel visual tensor format. The transformed image is enhanced using a predefined composite enhancement function to obtain an enhanced image.
3. The method according to claim 2, characterized in that, The composite enhancement function includes: Where V is an image in three-channel visual tensor format; M stands for reverse mirror; For angle transformation; Light disturbance; Random occlusion; It is an elastic deformation.
4. The method according to claim 3, characterized in that, The semantic segmentation process performed on the enhanced image to obtain the segmented image includes: Based on shallow depth-separable convolution and mid-to-deep channel compression strategies, feature extraction is performed on the enhanced image to obtain feature data; The image features are analyzed hierarchically using upsampling fusion to obtain the analyzed data; The parsed data is subjected to confidence output and adaptive encoding to obtain a segmented image.
5. The method according to claim 4, characterized in that, The feature extraction based on the shallow depth-separable convolution and mid-to-deep channel compression strategy yields feature data including: in, Features that have undergone attention adjustment; For global evaluation pooling; For activation functions; Feature compression path: in, It is a convolutional block that includes activation functions.
6. The method according to claim 5, characterized in that, The upsampling fusion method is used to perform hierarchical parsing of the image features to obtain parsed data, including: in, This is either a transposed convolution or a nearest neighbor interpolation; For splicing channels; It is a 3×3 convolutional module of the j-th layer.
7. The method according to claim 1, characterized in that, The predefined hybrid loss mechanism includes: in, A predefined regional overlap index; As weight; Loss function: =- Where N is the total number of pixels; The true label of the i-th pixel; Predict the probability of the model for the i-th pixel; α is the prospect weighting factor; It is a very small value.
8. A smart home external visual perception device based on semantic segmentation, characterized in that, include: The acquisition module is used to acquire the original image to be processed; The enhancement module is used to perform morphological enhancement on the original image to obtain an enhanced image; The segmentation module is used to perform semantic segmentation processing on the enhanced image to obtain a segmented image; The correction module is used to correct the segmented image using a predefined hybrid loss mechanism to obtain the target image.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method as described in any one of claims 1 to 7.