Visual identification system for intelligent electric appliance

Through the visual recognition system for smart appliances, low-light enhancement and multi-scale feature fusion technology are used to solve the problem of electrical identification accuracy in different brands and lighting changes, especially in low-light environments, which improves the recognition effect, providing reliable technical support for smart home equipment management.

CN119942083AInactive Publication Date: 2025-05-06SHENZHEN JUNHUI TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510108334.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing smart electrical identification system is insufficient in recognition accuracy and efficiency when facing changes in different brands, models and lighting conditions, especially in low-light environments.

Method used

The visual recognition system for smart electrical appliances is adopted, including target electrical image acquisition, low-light environment detection, low-light enhancement, multi-scale feature extraction and feature fusion. Low-light enhancement is performed through the Retinex-Net model, electrical object detection is performed using the YOLO model, and feature fusion is performed in combination with the hollow pyramid network and multiple attention structure, and finally electrical type recognition is performed through the support vector machine.

Benefits of technology

It improves the accuracy and efficiency of electrical identification, especially in low-light environments, and optimizes the identification effect, providing reliable technical support for the management and control of smart home equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942083A_ABST
    Figure CN119942083A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual identification, and particularly discloses a visual identification system for an intelligent electric appliance, which is characterized in that a target electric appliance image acquired by a camera is acquired, and whether the target electric appliance image is a low-light environment image is determined based on the brightness distribution characteristics of the target electric appliance image; then, after determining that the target electric appliance image is a low-light environment image, performing low-light enhancement and electric appliance target detection on the target electric appliance image to obtain an electric appliance ROI image; then, extracting shallow-layer features and deep-layer features of the ROI image of the electric appliance, and performing feature fusion based on semantic guidance to obtain multi-scale joint perception coding features of the image of the target electric appliance; and finally, determining a type label of the target electric appliance based on the multi-scale joint perception coding characteristics of the target electric appliance image. Therefore, the accuracy of electric appliance identification is improved, optimization is carried out especially for a low-light environment, and reliable technical support is provided for management and control of smart home equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual recognition technology, and more specifically, to a visual recognition system for smart appliances. Background Art

[0002] With the continuous development of smart home technology, visual recognition systems for smart appliances are playing an increasingly important role. As more and more devices are connected to the home network, users expect to be able to automatically manage and control these devices through visual recognition technology. For example, they can automatically detect and identify various appliances in the home to provide personalized services or manage energy consumption. However, the smart appliance recognition solutions currently on the market still have many limitations, especially in terms of recognition accuracy in a variety of appliance types and complex scenarios.

[0003] Traditional smart appliance recognition methods mainly rely on pre-set template matching or classification algorithms based on simple features (such as color and shape). This method may be effective for specific types of appliances, but its generalization ability is obviously insufficient when faced with a large number of appliances of different brands, models, and with high similarity in appearance. In addition, due to the changing placement and angle of appliances, coupled with the influence of ambient lighting conditions, it is difficult for traditional methods to stably capture the key features of appliances, thus affecting the accuracy and efficiency of recognition. Especially in low-light environments, the quality of image acquisition will drop significantly, resulting in the loss of a lot of detailed information, making it almost impossible for traditional appliance recognition methods to work properly.

[0004] To solve the above problems, deep learning algorithms have gradually attracted attention, but most existing deep learning solutions focus on high performance under ideal conditions and still face many challenges in actual deployment. For example, most pre-trained models are not optimized specifically for low-light environments, so they may encounter performance bottlenecks in real-world applications. Moreover, most models only consider feature extraction at a single scale, which limits the model's effective recognition of appliances of different sizes and proportions.

[0005] Therefore, an optimized visual recognition system for smart appliances is expected. Summary of the invention

[0006] The present application provides a visual recognition system for smart appliances, which not only improves the accuracy of appliance recognition, but also is optimized specifically for low-light environments, providing reliable technical support for the management and control of smart home devices.

[0007] In a first aspect, a visual recognition system for a smart appliance is provided, comprising: A target appliance image acquisition module is used to acquire the target appliance image captured by the camera; An image environment determination module, configured to determine whether the target appliance image is a low-light environment image based on a brightness distribution feature of the target appliance image; A low-light enhancement module, configured to, after determining that the target appliance image is a low-light environment image, perform low-light enhancement on the target appliance image using a low-light enhancement model to obtain an enhanced target appliance image; An electrical appliance target detection module, used for performing electrical appliance target detection on the enhanced target electrical appliance image to obtain an electrical appliance ROI image; A multi-scale feature extraction module, used to extract multi-scale features of the appliance ROI image to obtain shallow features of the target appliance image and deep features of the target appliance image; A feature fusion module, used for performing feature fusion based on semantic guidance on the shallow features of the target appliance image and the deep features of the target appliance image to obtain multi-scale joint perceptual coding features of the target appliance image; The type label determination module is used to determine the type label of the target appliance based on the multi-scale joint perceptual coding features of the target appliance image.

[0008] In the above-mentioned visual recognition system for smart appliances, the low-light enhancement model is a Retinex-Net model.

[0009] In the above-mentioned visual recognition system for intelligent electrical appliances, the low-light enhancement module is used to: Decomposing the target electrical appliance image to obtain an illumination component and a reflection component; Correcting the illumination component and the reflection component respectively to obtain a corrected illumination component and a corrected reflection component; The corrected illumination component and the corrected reflection component are fused to obtain the enhanced target appliance image.

[0010] In the above-mentioned intelligent electrical appliance visual recognition system, the electrical appliance target detection module is used to: The YOLO model is used to perform electrical appliance target detection on the enhanced target electrical appliance image to obtain the electrical appliance ROI image.

[0011] In the above-mentioned visual recognition system for intelligent electrical appliances, the multi-scale feature extraction module is used to: The appliance ROI image is input into an image multi-scale feature extractor based on a dilated pyramid network to obtain a shallow feature map of a target appliance image and a deep feature map of a target appliance image as shallow features of the target appliance image and deep features of the target appliance image.

[0012] In the above-mentioned visual recognition system for intelligent electrical appliances, the feature fusion module includes: An upsampling unit, configured to upsample the target appliance image deep feature map to obtain an upsampled target appliance deep feature map, wherein the upsampled target appliance deep feature map has the same size as the target appliance image shallow feature map; The feature joint perception unit is used to perform target electrical appliance image deep and shallow feature joint perception on the target electrical appliance image shallow feature and the target electrical appliance image deep feature based on the semantic information field between the upsampled target electrical appliance deep feature map and the target electrical appliance image shallow feature map to obtain the target electrical appliance image multi-scale joint perception coding feature map as the target electrical appliance image multi-scale joint perception coding feature.

[0013] In the above-mentioned visual recognition system for intelligent electrical appliances, the feature joint perception unit includes: A field modulation subunit is used to calculate the target electrical appliance image semantic information field between the upsampled target electrical appliance deep feature map and the target electrical appliance image shallow feature map, and perform field modulation on the upsampled target electrical appliance deep feature map based on the target electrical appliance image semantic information field to obtain a field modulated target electrical appliance deep feature map; The target appliance image multi-scale joint perception encoding subunit is used to input the target appliance image shallow feature map and the field modulated target appliance deep feature map into a feature joint perception module based on a multiple attention structure to obtain the target appliance image multi-scale joint perception encoding feature map.

[0014] In the above-mentioned visual recognition system for intelligent electrical appliances, the field modulation subunit is used to: The upsampled target appliance deep feature map and the target appliance image shallow feature map are feature-connected and then input into a semantic information field predictor based on gated convolution to obtain the target appliance image semantic information field; The upsampled target electrical appliance deep feature map is mapped to the target electrical appliance image semantic information field to obtain the field domain modulated target electrical appliance deep feature map.

[0015] In the above-mentioned visual recognition system for intelligent electrical appliances, the type label determination module is used to: The multi-scale joint perceptual coding feature map of the target electrical appliance image is input into a classifier-based electrical appliance identifier to obtain a recognition result, where the recognition result is a type label of the target electrical appliance.

[0016] The present application provides a visual recognition system for smart appliances, which obtains a target appliance image captured by a camera, and determines whether the target appliance image is a low-light environment image based on the brightness distribution characteristics of the target appliance image; then, after determining that the target appliance image is a low-light environment image, the target appliance image is subjected to low-light enhancement and appliance target detection to obtain an appliance ROI image; then, the shallow features and deep features of the appliance ROI image are extracted and feature fusion based on semantic guidance is performed to obtain the multi-scale joint perception coding features of the target appliance image; finally, based on the multi-scale joint perception coding features of the target appliance image, the type label of the target appliance is determined. In this way, not only the accuracy of appliance recognition is improved, but also it is optimized especially for low-light environments, providing reliable technical support for the management and control of smart home devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings of the embodiments of the present application are briefly introduced below. Obviously, the drawings described below only relate to some embodiments of the present application, and are not intended to limit the present application.

[0018] Figure 1 This is a schematic block diagram of a visual recognition system for smart appliances according to an embodiment of the present application.

[0019] Figure 2 Schematic diagram of data flow of a visual recognition system for smart appliances according to an embodiment of the present application.

[0020] Figure 3 This is a schematic block diagram of a feature fusion module in a visual recognition system for smart appliances according to an embodiment of the present application.

[0021] Figure 4 This is a schematic block diagram of a feature joint perception unit in a visual recognition system for smart appliances according to an embodiment of the present application.

[0022] Figure 5 This is a schematic flowchart of a visual recognition method for smart appliances according to an embodiment of the present application. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work also fall within the scope of protection of the present application.

[0024] In response to the above technical problems, the technical concept of the present application is as follows: it obtains the target appliance image captured by the camera, and determines whether the target appliance image is a low-light environment image based on the brightness distribution characteristics of the target appliance image; then, after determining that the target appliance image is a low-light environment image, the target appliance image is subjected to low-light enhancement and appliance target detection to obtain an appliance ROI image; then, the shallow features and deep features of the appliance ROI image are extracted and feature fusion based on semantic guidance is performed to obtain the multi-scale joint perception coding features of the target appliance image; finally, based on the multi-scale joint perception coding features of the target appliance image, the type label of the target appliance is determined. In this way, not only the accuracy of appliance recognition is improved, but also it is optimized especially for low-light environments, providing reliable technical support for the management and control of smart home devices.

[0025] Specifically, in the technical solution of this application, if Figure 1 and Figure 2 As shown, the visual recognition system for intelligent electrical appliances includes: a target electrical appliance image acquisition module 10, which is used to acquire the target electrical appliance image collected by the camera; an image environment determination module 20, which is used to determine whether the target electrical appliance image is a low-light environment image based on the brightness distribution characteristics of the target electrical appliance image; a low-light enhancement module 30, which is used to perform low-light enhancement on the target electrical appliance image using a low-light enhancement model after determining that the target electrical appliance image is a low-light environment image to obtain an enhanced target electrical appliance image; an electrical appliance target detection module 40, which is used to perform electrical appliance target detection on the enhanced target electrical appliance image to obtain an electrical appliance ROI image; a multi-scale feature extraction module 50, which is used to extract multi-scale features of the electrical appliance ROI image to obtain shallow features of the target electrical appliance image and deep features of the target electrical appliance image; a feature fusion module 60, which is used to perform semantic-guided feature fusion on the shallow features of the target electrical appliance image and the deep features of the target electrical appliance image to obtain multi-scale joint perceptual coding features of the target electrical appliance image; and a type label determination module 70, which is used to determine the type label of the target electrical appliance based on the multi-scale joint perceptual coding features of the target electrical appliance image.

[0026] Exemplarily, in the target appliance image acquisition module 10, the target appliance image captured by the camera is acquired. It should be understood that in a smart home environment, automatically capturing and identifying appliances in the home can greatly facilitate users to manage and control these devices, such as remote monitoring, timer switches, etc., without the need for users to manually input or mark appliance information, thereby reducing the operational burden and improving the overall convenience and intelligence level. At the same time, the location of appliances in a home environment may change frequently. Using a camera to capture the target appliance image and identify it can capture the latest layout, so that the system can update information in a timely manner.

[0027] Exemplarily, in the image environment determination module 20, based on the brightness distribution characteristics of the target appliance image, it is determined whether the target appliance image is a low-light environment image. It should be understood that considering that images captured in low-light environments often have problems such as low contrast, high noise, and blurred details, these problems will directly lead to a decrease in the accuracy of target detection and classification tasks. Therefore, based on the brightness distribution characteristics of the target appliance image, it is determined whether the target appliance image is a low-light environment image. By identifying low-light environment images in advance, a specially designed enhancement algorithm can be activated to improve image quality and ensure the accuracy of subsequent processing. In addition, not all captured images need to be processed for low-light enhancement. The relevant modules are only started when it is confirmed to be a low-light environment. This can save computing resources and improve the overall efficiency of the system.

[0028] In one embodiment, based on the brightness distribution characteristics of the target electrical appliance image, determining whether the target electrical appliance image is a low-light environment image includes: first, performing standardized preprocessing on the collected target electrical appliance image, such as resizing, normalization, etc., to ensure that all images have the same size and pixel value range for uniform comparison. Then, constructing a brightness histogram of the image, which shows the distribution of the number of pixels at each brightness level in the image. For images under normal lighting conditions, the brightness distribution is usually relatively uniform; while in a low-light environment, most of the pixels of the image are concentrated in a lower brightness area, that is, the histogram shows a left-biased trend. Extracting key statistical features from the brightness histogram, such as average brightness, standard deviation, percentile, etc., these statistics can help quantify the overall brightness level of the image and its degree of change. In particular, if the average brightness is lower than a preset threshold, it is preliminarily determined that it may be in a low-light environment. Furthermore, the brightness distribution of local areas in the image can also be examined, and by dividing multiple sub-areas and calculating their brightness statistics respectively, it can be understood whether there are some specific parts that are particularly dark, which helps to distinguish the impact of global low light and local shadows.

[0029] Exemplarily, in the low-light enhancement module 30, after determining that the target appliance image is a low-light environment image, the target appliance image is low-light enhanced using a low-light enhancement model to obtain an enhanced target appliance image. It should be understood that the quality of the image can be significantly improved through low-light enhancement processing, so that the enhanced image has higher contrast and clearer details, thereby ensuring that the subsequent target detection module (such as the YOLO model) can more accurately locate and identify the appliance. In addition, the low-light enhancement model can suppress noise while improving brightness, ensuring that useful information in the image will not be masked by noise, and high-quality input images mean that the multi-scale feature extraction module can capture more detailed features, thereby improving the accuracy of the final classification results.

[0030] In one embodiment, the low-light enhancement model is a Retinex-Net model. This model is based on the Retinex theory and is intended to simulate how the human visual system perceives color and brightness, and to maintain the true color of objects even under different lighting conditions. In one embodiment, the low-light enhancement module is used to: decompose the target electrical appliance image to obtain an illumination component and a reflection component; respectively correct the illumination component and the reflection component to obtain a corrected illumination component and a corrected reflection component; and fuse the corrected illumination component and the corrected reflection component to obtain the enhanced target electrical appliance image. Specifically: First, the target electrical appliance image is decomposed to obtain an illumination component and a reflection component. It should be understood that this decomposition method is based on the Retinex theory, which believes that the brightness of the image is the result of the combined effect of illumination and reflection. In this way, the influence of illumination in the image can be better understood and processed. Then, the illumination component and the reflection component are corrected respectively to obtain a corrected illumination component and a corrected reflection component, wherein the illumination component is adjusted by a nonlinear transformation or filter to improve the overall brightness of the image while avoiding overexposure; the detail information of the reflection component, especially the edge and texture, is enhanced by image processing techniques such as sharpening filtering to restore the structural features in the image. Finally, the corrected illumination component and the corrected reflection component are fused to obtain the enhanced target electrical appliance image. In a specific embodiment, in order to ensure that the enhanced target electrical appliance image is bright while retaining the original details and colors, a weighted average method is used for fusion. At the same time, considering that different regions may require different degrees of enhancement, a nonlinear fusion method can be used. For example, the weight distribution on each pixel is dynamically adjusted based on local contrast or gradient information. This not only avoids the problems caused by global consistent enhancement, but also better adapts to complex illumination changes.

[0031] Exemplarily, in the electrical appliance target detection module 40, electrical appliance target detection is performed on the enhanced target electrical appliance image to obtain an electrical appliance ROI image. It should be understood that, considering that a direct comprehensive analysis of the enhanced target electrical appliance image will introduce a large amount of background noise and irrelevant information, these additional data not only increase the computational burden, but may also lead to false detection or missed detection. Electrical appliance target detection is used to obtain an electrical appliance ROI (region of interest) image to ensure that subsequent processing can be focused on key electrical appliance areas, thereby improving the accuracy and efficiency of recognition. And electrical appliance target detection can also significantly reduce background interference and unnecessary information, making subsequent processing more accurate. At the same time, focusing processing resources on the electrical appliance ROI avoids the computational burden brought by global image analysis and improves the operating efficiency of the system. In one embodiment, the electrical appliance target detection module is used to: use the YOLO model to perform electrical appliance target detection on the enhanced target electrical appliance image to obtain the electrical appliance ROI image.

[0032] Exemplarily, in the multi-scale feature extraction module 50, the multi-scale features of the appliance ROI image are extracted to obtain the shallow features of the target appliance image and the deep features of the target appliance image. It should be understood that the appliance image contains multi-scale information ranging from macroscopic appearance structure to microscopic texture details. For example, in an image of a refrigerator, macroscopic scale information such as the overall outline of the refrigerator, the number and position of the doors, etc. can be presented, and microscopic features such as the texture of the surface material and the design of the control panel can also be observed. These different scales of information are very important for accurately identifying the type and brand of appliances.

[0033] In one embodiment, the multi-scale feature extraction module is used to: input the appliance ROI image into an image multi-scale feature extractor based on a dilated pyramid network to obtain a shallow feature map of a target appliance image and a deep feature map of a target appliance image as the shallow features of the target appliance image and the deep features of the target appliance image. It should be understood that the appliance ROI image is input into an image multi-scale feature extractor based on a dilated pyramid network to expand the receptive field without reducing the image resolution, thereby capturing features at different scales, and obtaining a shallow feature map of the target appliance image and a deep feature map of the target appliance image. In particular, dilated convolution is the core operation of the dilated pyramid network, which can increase the receptive field of the convolution kernel without increasing the number of parameters and the amount of calculation. This enables the network to obtain information in a wider area, thereby better understanding the relationship between different areas in the appliance image. For example, when analyzing the handle design in a refrigerator image, the dilated convolution can simultaneously consider the material information and shape characteristics around the handle, which helps to more accurately identify a specific model of refrigerator. By constructing a pyramid structure, the model can extract features at different levels, each level corresponding to a different scale. This hierarchical feature extraction method can adapt to the multi-scale characteristics of electrical appliance images, encode features at different scales separately, and provide rich materials for subsequent comprehensive analysis.

[0034] Exemplarily, in the feature fusion module 60, the shallow features of the target electrical appliance image and the deep features of the target electrical appliance image are subjected to feature fusion based on semantic guidance to obtain the multi-scale joint perceptual coding features of the target electrical appliance image. It should be understood that, considering that the shallow features usually contain local information such as edges and textures, while the deep features can reflect the overall structure and shape of the object. By combining these two features, the target electrical appliance can be described more comprehensively, thereby improving the generalization ability of the classifier in the face of various complex situations. For example, in a low-light environment, some details may be lost, but if the global information in the deep features can be effectively utilized, it can still help accurately identify the type of electrical appliance. Therefore, in the technical solution of the present application, the shallow features of the target electrical appliance image and the deep features of the target electrical appliance image are subjected to feature fusion based on semantic guidance to comprehensively utilize the shallow features and the deep features. At the same time, by introducing semantic information as a guide for feature fusion, it means that which types of features should be more emphasized or suppressed can be dynamically adjusted according to the needs of specific application scenarios. This is particularly important for improving performance under specific tasks. For example, when identifying electrical appliances, the semantic information field can be used to enhance the feature expressions related to the electrical appliances, while weakening the influence of irrelevant background. Such modulation makes the final feature map not only contain rich visual information, but also have stronger semantic interpretation, which is convenient for subsequent processing steps. Semantic-level feature modulation helps improve the model's understanding of complex scenes, especially in low-light environments to distinguish similar objects or identify objects under occlusion, enhances the sensitivity of deep features to specific semantic areas, and further improves the quality of feature expression.

[0035] In one embodiment, Figure 3 As shown, the feature fusion module 60 includes: an upsampling unit 61, which is used to upsample the deep feature map of the target electrical appliance image to obtain an upsampled target electrical appliance deep feature map, wherein the upsampled target electrical appliance deep feature map and the shallow feature map of the target electrical appliance image have the same size; a feature joint perception unit 62, which is used to perform target electrical appliance image deep and shallow feature joint perception on the shallow feature of the target electrical appliance image and the deep feature of the target electrical appliance image based on the semantic information field between the upsampled target electrical appliance deep feature map and the shallow feature map of the target electrical appliance image to obtain a target electrical appliance image multi-scale joint perception coding feature map as the target electrical appliance image multi-scale joint perception coding feature.

[0036] Exemplarily, in the upsampling unit 61, the target appliance image deep feature map is upsampled to obtain an upsampled target appliance deep feature map. Here, the upsampling process for the target appliance image deep feature map is not only to simply increase the spatial resolution of the feature map, but more importantly, to restore some detail information lost in the deep network. By using methods such as transposed convolution (deconvolution) or interpolation, fine-grained structural features can be reconstructed, which is crucial for appliance recognition. Specifically, the process can be expressed as: in, represents the deep feature map of the target appliance image, represents the upsampling operation, Represents the deep feature map of the upsampled target appliance.

[0037] In one embodiment, Figure 4 As shown, the feature joint perception unit 62 includes: a field modulation subunit 621, which is used to calculate the target electrical appliance image semantic information field between the upsampled target electrical appliance deep feature map and the target electrical appliance image shallow feature map, and perform field modulation on the upsampled target electrical appliance deep feature map based on the target electrical appliance image semantic information field to obtain a field modulated target electrical appliance deep feature map; a target electrical appliance image multi-scale joint perception encoding subunit 622, which is used to input the target electrical appliance image shallow feature map and the field modulated target electrical appliance deep feature map into a feature joint perception module based on a multiple attention structure to obtain the target electrical appliance image multi-scale joint perception encoding feature map.

[0038] In one embodiment, the field modulation subunit is used to: perform feature connection on the upsampled target appliance deep feature map and the target appliance image shallow feature map, and then input the feature map into a semantic information field predictor based on gated convolution to obtain the target appliance image semantic information field. Specifically, the process can be expressed as follows: in, represents the shallow feature map of the target appliance image, The convolution kernel is The convolutional layer, is point convolution, represents the feature connection operation, Represents the semantic information field of the target electrical appliance image.

[0039] The upsampled target appliance deep feature map is mapped to the target appliance image semantic information field to obtain the field domain modulated target appliance deep feature map. Specifically, the process can be expressed as follows: in, It means point multiplication by position. Represents the deep feature map of the field modulated target electrical appliance.

[0040] Specifically, when these concatenated features are fed into a semantic information field predictor based on gated convolution, it allows the construction of a representation that reflects the global and local semantic structure of the entire image. As the core component of the predictor, gated convolution can adaptively adjust the importance of feature channels by introducing additional learning parameters to ensure that the feature response in each region is more consistent with its true semantic meaning. For example, for a photo of a refrigerator, the semantic information field can help highlight key parts such as door handles and trademarks while suppressing the influence of irrelevant background, thereby improving the accuracy of recognition. Next, the upsampled deep feature map is mapped to the semantic information field, that is, "field modulation" is performed. This process is actually a high-level feature enhancement method that dynamically adjusts each element on the original deep feature map according to the semantic context. Specifically, field modulation will enhance the feature expression of those areas that are associated with or more important for appliance recognition based on the guidance provided by the semantic information field, while weakening other unimportant parts. This has two main benefits: first, it enhances the sensitivity of deep features to specific semantic areas (such as key parts of appliances); second, it improves the quality of feature expression, so that the true form and characteristics of objects can be better maintained even in low-light environments or in the presence of noise interference. The final deep feature map of the field-modulated target appliance not only retains the original high-level semantic information, but also incorporates the high-level understanding provided by the semantic information field.

[0041] Exemplarily, in the target appliance image multi-scale joint perceptual coding subunit 622, the target appliance image shallow feature map and the field modulated target appliance deep feature map are input into the feature joint perception module based on the multiple attention structure to obtain the target appliance image multi-scale joint perceptual coding feature map. It should be understood that in order to fully utilize the advantages of shallow and deep features, the target appliance image shallow feature map and the field modulated target appliance deep feature map are input into the feature joint perception module based on the multiple attention structure to obtain the target appliance image multi-scale joint perceptual coding feature map. Figure 1At the same time, they are sent to the feature joint perception module based on the multiple attention structure. The multiple attention structure is designed to simulate the way the human visual system processes information, that is, to flexibly allocate attention resources according to different task requirements. In this module, the shallow feature map of the target appliance image and the deep feature map of the field-modulated target appliance complement each other to form a multi-level feature representation system. The multiple attention mechanism allows the model to dynamically adjust the focus according to the requirements of different tasks and automatically learn the feature combination that is most beneficial to the current task. The multi-scale joint perception encoding feature map of the target appliance image finally generated not only contains rich low-level visual information, but also contains high-level semantic understanding, which provides strong support for accurate appliance recognition. Specifically, the process can be expressed by the formula: in, For point-by-point convolution processing, is the global mean pooling operation, For channel-by-channel convolution processing, is the batch normalization layer, is the Sigmoid function, Optimize the expression feature map for the target appliance's deep attention. Optimize the shallow attention expression feature map for the target appliance. For positional addition, Represents a multi-scale joint perceptual coding feature map of the target appliance image.

[0042] In summary, the semantic-guided feature fusion method is explained. Compared with the traditional visual processing architecture, it introduces upsampling and feature connection to enable low-level and high-level features to interact at the same scale, solves the problem of multi-scale information fusion, and improves the model's ability to understand complex scenes. In addition, by constructing a semantic information field and modulating deep features accordingly, it ensures that the feature expression not only contains rich detail information, but also is consistent with the semantic structure in the image. This semantic-level feature modulation helps to improve the model's sensitivity to specific objects or regions.

[0043] Exemplarily, in the type label determination module 70, the type label of the target electrical appliance is determined based on the multi-scale joint perceptual coding features of the target electrical appliance image. In one embodiment, the type label determination module is used to: input the multi-scale joint perceptual coding feature map of the target electrical appliance image into an electrical appliance identifier based on a classifier to obtain a recognition result, and the recognition result is the type label of the target electrical appliance. It should be understood that through the previous multi-scale joint perceptual coding, the present application has obtained a multi-scale joint perceptual coding feature map of the target electrical appliance image that contains both rich details and high-level semantic understanding. The multi-scale joint perceptual coding feature map of the target electrical appliance image not only combines the advantages of shallow and deep features, but also optimizes and modulates through semantic guidance to ensure the ability to understand complex scenes and robustness. The multi-scale joint perceptual coding feature map of the target electrical appliance image is input into the electrical appliance identifier, and its rich information content can be fully utilized to make the final classification decision. The task of the electrical appliance identifier is to select the category that best meets the current feature description from many possible electrical appliance categories, so as to give an accurate recognition result. Since the multi-scale joint perceptual coding feature map of the target appliance image already contains multi-level information from local details to global structure, it can provide a more comprehensive and reliable basis for the classifier, improving the accuracy and reliability of recognition.

[0044] Specifically, in one embodiment, the classifier-based electrical appliance identifier uses a support vector machine (SVM). SVM is a supervised learning model that is widely used in pattern recognition and classification tasks. It divides sample points of different categories by finding an optimal hyperplane so that samples in the same category are clustered together as much as possible, while the distance between different categories is maximized. For the electrical appliance identification task, SVM can find the best classification boundary based on the feature vector after the multi-scale joint perceptual coding feature map is expanded, thereby achieving efficient distinction between various types of electrical appliances.

[0045] In a preferred example, inputting the target appliance image multi-scale joint perceptual coding feature map into a classifier-based appliance identifier to obtain an identification result includes: Expanding the target appliance image multi-scale joint perceptual coding feature map features into a target appliance image multi-scale joint perceptual coding feature vector; Based on the feature value relationship between the multi-scale joint perceptual coding feature vector of the target appliance image Distance and Distance matrix obtained by multi-scale joint perceptual coding of target appliance image and target appliance image multi-scale joint perceptual encoding two distance matrices ; Calculate the weighted sum of the first distance matrix of the multi-scale joint perceptual coding of the target electrical appliance image and the second distance matrix of the multi-scale joint perceptual coding of the target electrical appliance image to obtain the joint distance matrix of the multi-scale joint perceptual coding of the target electrical appliance image ,in, represents positional addition, It means point multiplication by position. and represents different weighting hyperparameters; Determine each eigenvalue of the multi-scale joint perceptual coding joint distance matrix of the target appliance image , and compose the target appliance image multi-scale joint perceptual coding joint distance eigenvector ; The target appliance image multi-scale joint perceptual coding feature vector as a row vector is matrix-multiplied with the target appliance image multi-scale joint perceptual coding-distance matrix to obtain the target appliance image multi-scale joint perceptual coding-distance query vector ,in, represents the multi-scale joint perceptual coding feature vector of the target appliance image, Represents matrix multiplication; The target electrical appliance image multi-scale joint perceptual coding two-distance matrix is ​​matrix-multiplied with the target electrical appliance image multi-scale joint perceptual coding feature vector autocorrelation matrix to obtain the target electrical appliance image multi-scale joint perceptual coding two-distance correlation matrix ,in, Represents the transpose of a vector; After matrix multiplication of the target appliance image multi-scale joint perceptual coding distance query vector and the target appliance image multi-scale joint perceptual coding distance correlation matrix, further dot multiplication is performed with the target appliance image multi-scale joint perceptual coding joint distance eigenvector to obtain an optimized target appliance image multi-scale joint perceptual coding feature vector ; The optimized multi-scale joint perceptual coding feature vector of the target electrical appliance image is input into an electrical appliance identifier based on a classifier to obtain a recognition result.

[0046] Here, when the shallow feature map of the target electrical appliance image and the deep feature map of the target electrical appliance image respectively represent the shallow and deep semantic coding features of the electrical appliance ROI image, when performing feature joint perception based on semantic information field guided migration, the differences in the receptive field and feature expression granularity of the image features will cause the sparse fine-grained feature union of the multi-scale joint perception coding feature vector of the target electrical appliance image, thereby reducing the accuracy of the recognition result obtained by the input classifier-based electrical appliance identifier due to the lack of classification reasoning degree.

[0047] Therefore, the first distance matrix and the second distance matrix of the multi-scale joint perceptual coding feature vector of the target electrical appliance image are used as the fine-grained metric association cluster representation of the multi-scale joint perceptual coding feature vector of the target electrical appliance image, and the dynamic programming of the relationship between association clusters of different association clusters is performed on the multi-scale joint perceptual coding feature vector of the target electrical appliance image and the self-association representation of the multi-scale joint perceptual coding feature vector of the target electrical appliance image, respectively, to simulate the sparse activation based on neuron clusters of the association system, and the fine-grained predictable sparsity of the multi-scale joint perceptual coding feature vector of the target electrical appliance image is coordinated with the intrinsic representation of the metric association cluster of the first distance matrix and the second distance matrix, so as to avoid the lack of classification reasoning degree affected by insufficient feature joint perception caused by sparsity, and improve the accuracy of the recognition result obtained by inputting the multi-scale joint perceptual coding feature vector of the target electrical appliance image into the classifier-based electrical appliance identifier.

[0048] In summary, according to the embodiment of the present application, the visual recognition system for smart appliances is explained, which obtains the target appliance image captured by the camera, and determines whether the target appliance image is a low-light environment image based on the brightness distribution characteristics of the target appliance image; then, after determining that the target appliance image is a low-light environment image, the target appliance image is subjected to low-light enhancement and appliance target detection to obtain an appliance ROI image; then, the shallow features and deep features of the appliance ROI image are extracted and feature fusion based on semantic guidance is performed to obtain the multi-scale joint perception coding features of the target appliance image; finally, based on the multi-scale joint perception coding features of the target appliance image, the type label of the target appliance is determined. In this way, not only the accuracy of appliance recognition is improved, but also it is optimized especially for low-light environments, providing reliable technical support for the management and control of smart home devices.

[0049] Figure 5 Schematic flow chart of the visual recognition method for smart appliances according to the embodiment of the present application. Figure 5As shown, the visual recognition method for smart appliances includes: S1, acquiring a target appliance image captured by a camera; S2, determining whether the target appliance image is a low-light environment image based on the brightness distribution characteristics of the target appliance image; S3, after determining that the target appliance image is a low-light environment image, using a low-light enhancement model to perform low-light enhancement on the target appliance image to obtain an enhanced target appliance image; S4, performing appliance target detection on the enhanced target appliance image to obtain an appliance ROI image; S5, extracting multi-scale features of the appliance ROI image to obtain shallow features of the target appliance image and deep features of the target appliance image; S6, performing semantic-guided feature fusion on the shallow features of the target appliance image and the deep features of the target appliance image to obtain multi-scale joint perceptual coding features of the target appliance image; S7, determining the type label of the target appliance based on the multi-scale joint perceptual coding features of the target appliance image.

[0050] Here, those skilled in the art can understand that the specific operations of each step in the above-mentioned smart appliance visual recognition method have been referred to above. Figures 1 to 4 The description of the visual recognition system for smart appliances has been introduced in detail, and therefore, its repeated description will be omitted.

[0051] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the devices, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains at least one executable instruction for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0052] The present application uses specific words to describe the embodiments of the present application. For example, "first / second embodiment", "one embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of the present application. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned in different positions in this specification does not necessarily refer to the same embodiment. In addition, some features, structures or characteristics in one or more embodiments of the present application can be appropriately combined.

[0053] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or extremely formal sense, unless explicitly defined as such herein.

[0054] The above is an explanation of the present application and should not be considered as a limitation thereof. Although several exemplary embodiments of the present application are described, it will be readily understood by those skilled in the art that many modifications may be made to these exemplary embodiments without departing from the novel teachings and advantages of the present application. Therefore, all of these modifications are intended to be included within the scope of the present application. It should be understood that the foregoing is an explanation of the present application and should not be considered to be limited to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the definition of the present application. The present application is defined by its description and equivalents.

Claims

1. A visual recognition system for intelligent electrical appliances, characterized in that: include: A target appliance image acquisition module is used to acquire the target appliance image captured by the camera; An image environment determination module, configured to determine whether the target appliance image is a low-light environment image based on a brightness distribution feature of the target appliance image; A low-light enhancement module, configured to, after determining that the target appliance image is a low-light environment image, perform low-light enhancement on the target appliance image using a low-light enhancement model to obtain an enhanced target appliance image; An electrical appliance target detection module, used for performing electrical appliance target detection on the enhanced target electrical appliance image to obtain an electrical appliance ROI image; A multi-scale feature extraction module, used to extract multi-scale features of the appliance ROI image to obtain shallow features of the target appliance image and deep features of the target appliance image; A feature fusion module, used for performing feature fusion based on semantic guidance on the shallow features of the target appliance image and the deep features of the target appliance image to obtain multi-scale joint perceptual coding features of the target appliance image; The type label determination module is used to determine the type label of the target appliance based on the multi-scale joint perceptual coding features of the target appliance image.

2. The visual recognition system for intelligent electrical appliances according to claim 1, characterized in that: The low-light enhancement model is a Retinex-Net model.

3. The visual recognition system for intelligent electrical appliances according to claim 2, characterized in that: The low light enhancement module is used to: Decomposing the target electrical appliance image to obtain an illumination component and a reflection component; Correcting the illumination component and the reflection component respectively to obtain a corrected illumination component and a corrected reflection component; The corrected illumination component and the corrected reflection component are fused to obtain the enhanced target appliance image.

4. The visual recognition system for intelligent electrical appliances according to claim 3, characterized in that: The electrical appliance target detection module is used to: The YOLO model is used to perform electrical appliance target detection on the enhanced target electrical appliance image to obtain the electrical appliance ROI image.

5. The visual recognition system for intelligent electrical appliances according to claim 4, characterized in that: The multi-scale feature extraction module is used to: The appliance ROI image is input into an image multi-scale feature extractor based on a dilated pyramid network to obtain a shallow feature map of a target appliance image and a deep feature map of a target appliance image as shallow features of the target appliance image and deep features of the target appliance image.

6. The visual recognition system for intelligent electrical appliances according to claim 5, characterized in that: The feature fusion module comprises: An upsampling unit, configured to upsample the target appliance image deep feature map to obtain an upsampled target appliance deep feature map, wherein the upsampled target appliance deep feature map has the same size as the target appliance image shallow feature map; The feature joint perception unit is used to perform target electrical appliance image deep and shallow feature joint perception on the target electrical appliance image shallow feature and the target electrical appliance image deep feature based on the semantic information field between the upsampled target electrical appliance deep feature map and the target electrical appliance image shallow feature map to obtain the target electrical appliance image multi-scale joint perception coding feature map as the target electrical appliance image multi-scale joint perception coding feature.

7. The visual recognition system for intelligent electrical appliances according to claim 6, characterized in that: The feature joint perception unit includes: A field modulation subunit is used to calculate the target electrical appliance image semantic information field between the upsampled target electrical appliance deep feature map and the target electrical appliance image shallow feature map, and perform field modulation on the upsampled target electrical appliance deep feature map based on the target electrical appliance image semantic information field to obtain a field modulated target electrical appliance deep feature map; The target appliance image multi-scale joint perception encoding subunit is used to input the target appliance image shallow feature map and the field modulated target appliance deep feature map into a feature joint perception module based on a multiple attention structure to obtain the target appliance image multi-scale joint perception encoding feature map.

8. The visual recognition system for intelligent electrical appliances according to claim 7, characterized in that: The field modulation subunit is used for: The upsampled target appliance deep feature map and the target appliance image shallow feature map are feature-connected and then input into a semantic information field predictor based on gated convolution to obtain the target appliance image semantic information field; The upsampled target electrical appliance deep feature map is mapped to the target electrical appliance image semantic information field to obtain the field domain modulated target electrical appliance deep feature map.

9. The visual recognition system for intelligent electrical appliances according to claim 8, characterized in that: The type label determination module is used to: The multi-scale joint perceptual coding feature map of the target electrical appliance image is input into a classifier-based electrical appliance identifier to obtain a recognition result, where the recognition result is a type label of the target electrical appliance.

Citation Information

Patent Citations

  • Small target detection method and system under low visibility

    CN110807384A

  • Visual feature and semantic feature fusion-based document layout identification method and system

    CN114187595A

  • Low-light image enhancement method based on deep convolution attention and multi-scale feature fusion

    CN116091357A

  • Multi-station intelligent logistics conveying system

    CN117292193A

  • Image processing method and system for mine lamp, storage medium and mine lamp

    CN118941455A