Full-scene smart home supervision method and system based on graphic image recognition
Through the combination of graphic image recognition technology and photonic crystal waveguide array, the phase distribution of nano-level units of the metasurface array is dynamically adjusted, the multi-spectral home environment perception matrix is generated, and the speed-level parallel calculation is carried out, which solves the problem of limited perception capabilities of traditional smart home supervision systems and realizes efficient monitoring and abnormal detection of full-scene smart home supervision.
Patent Information
- Application Number
- CN202510364190.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional smart home supervision systems rely on a single type of sensor, have limited perception capabilities, and are inefficient in multi-spectral data processing, making it difficult to effectively monitor in dynamic environments.
The full-scene smart home supervision method based on graphic image recognition is adopted, and the phase distribution of nano-level units of the metasurface array is dynamically adjusted to generate a multi-spectral home environment perception matrix, and a photonic crystal waveguide array is used for parallel light-speed calculation, combining light-weight YOLO-Q model and quantum attention enhancement for full-scene risk analysis, and a structured alarm data packet is generated.
It realizes efficient monitoring of complex environments, improves data acquisition accuracy and processing speed, enhances adaptability to dynamic environments, and provides a safer, more comfortable and efficient living environment.
Smart Images

Figure CN120279377A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart home, in particular to an all-scenario intelligent home supervision method and system based on graphic image recognition. Background Art
[0002] In recent years, with the rapid development of Internet of Things (IoT) technology and artificial intelligence (AI), smart home supervision has gradually become an important part of modern life. Early smart home supervision mainly relied on simple sensor networks to collect environmental parameters. However, with the progress of technology, especially the application of image recognition technology and multispectral imaging technology, smart home supervision has been further optimized and extended. These new technologies not only improve the accuracy and scope of data collection, but also enhance the real-time monitoring ability of complex environmental parameters. In particular, the introduction of cutting-edge technologies such as quantum computing and photonic crystal waveguide arrays provides new possibilities for the efficient parallel computing of smart home supervision.
[0003] Disadvantages: First, traditional smart home supervision usually relies on a single type of sensor for environmental monitoring, which limits its perception ability of complex environmental variables. Second, although existing image recognition technologies can identify specific objects or scenes, they are still insufficient in processing and analyzing multispectral data in a dynamic environment. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides an all-scenario intelligent home supervision method based on graphic image recognition, which solves the problems of limited perception ability of a single type of sensor and low efficiency of multispectral data processing in the prior art.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides an all-scenario intelligent home supervision method based on graphic image recognition, which includes: collecting environmental parameters in real time and performing preprocessing, dynamically adjusting the phase distribution of nano-scale units of a metasurface array, and generating a multispectral home environment perception matrix; performing light-speed-level parallel computing on the multispectral home environment perception matrix through a photonic crystal waveguide array to generate an all-scenario feature matrix and a photon confidence parameter; inputting the all-scenario feature matrix into a lightweight YOLO-Q model, performing all-scenario risk analysis through quantum attention enhancement, and performing all-scenario anomaly detection based on the photon confidence parameter to generate a structured alarm data packet; performing emergency disposal, environmental adjustment and parameter optimization according to the structured alarm data packet.
[0008] As a preferred solution of the full-scenario intelligent home supervision method based on graphic image recognition of the present invention, wherein: dynamically adjusting the phase distribution of the nano-scale units of the metasurface array and generating a multi-spectral home environment perception matrix, the specific steps are as follows,
[0009] Dynamically adjust the phase distribution of the metasurface array nanostructure according to environmental parameters;
[0010] Separate the incident light into three independent bands according to the adjusted metasurface array;
[0011] Obtain three-channel images including visible light, near-infrared, and short-wave infrared through spectral separation and synchronous imaging;
[0012] Generate a multi-spectral home environment perception matrix through data normalization processing.
[0013] As a preferred solution of the full-scenario intelligent home supervision method based on graphic image recognition of the present invention, wherein: the preprocessing is to perform spatio-temporal alignment on environmental parameters using a data fusion algorithm to eliminate noise and instantaneous interference.
[0014] As a preferred solution of the full-scenario intelligent home supervision method based on graphic image recognition of the present invention, wherein: performing light-speed-level parallel computing on the multi-spectral home environment perception matrix through a photonic crystal waveguide array to generate a full-scenario feature matrix and a photon confidence parameter, the specific steps are as follows,
[0015] Map the multi-spectral home environment perception matrix into an optical signal through light intensity mapping and optical convolution calculation;
[0016] Convert the optical signal into a full-scenario feature matrix through graphene photoelectric detection and dynamic range compression;
[0017] Generate a confidence parameter through full-scenario credibility evaluation.
[0018] As a preferred solution of the full-scenario intelligent home supervision method based on graphic image recognition of the present invention, wherein: inputting the full-scenario feature matrix into a lightweight YOLO-Q model and performing full-scenario risk analysis through quantum attention enhancement, the specific steps are as follows,
[0019] Input the full-scenario feature matrix into a lightweight YOLO-Q model to extract a home supervision feature matrix;
[0020] Perform full-scenario risk analysis on the home supervision feature matrix through a photon computing acceleration layer to generate the channel values of the home supervision feature matrix.
[0021] As a preferred solution of the full-scenario intelligent home supervision method based on graphic image recognition according to the present invention, wherein: anomaly detection is performed based on photon confidence parameters to generate a structured alarm data packet, and the specific steps are as follows.
[0022] Combined with photon confidence parameters, anomaly detection is performed on the home supervision feature matrix, and an anomaly score is calculated.
[0023] A structured alarm data packet is generated according to the anomaly score, and anomaly information is extracted.
[0024] As a preferred solution of the full-scenario intelligent home supervision method based on graphic image recognition according to the present invention, wherein: emergency disposal, environmental adjustment and parameter optimization are performed according to the structured alarm data packet, and the specific steps are as follows.
[0025] An emergency disposal instruction is generated from the anomaly information in the alarm data packet, and the environmental parameters are dynamically adjusted.
[0026] According to the confidence level and severity in the alarm data packet, the metasurface parameters are optimized.
[0027] In a second aspect, the present invention provides a full-scenario intelligent home supervision system based on graphic image recognition, including a dynamic regulation module for real-time collecting environmental parameters through a micro sensor array, preprocessing the collected environmental parameters, dynamically adjusting the phase distribution of the nano-units of the metasurface array through the environmental parameters, and generating a multi-spectral data matrix; a photon computing module for performing light-speed-level parallel computing on the multi-spectral data matrix through a photonic crystal waveguide array to generate a full-scenario feature matrix and photon confidence parameters; an anomaly detection module for inputting the full-scenario feature matrix into a lightweight YOLO-Q model, enhancing quantum attention through a photon computing acceleration layer, and performing anomaly detection based on the photon confidence parameters to generate a structured alarm data packet; a response and disposal module for performing emergency disposal, environmental adjustment and parameter optimization according to the structured alarm data packet.
[0028] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and wherein: when the computer program is executed by the processor, any step of the full-scenario intelligent home supervision method based on graphic image recognition as described in the first aspect of the present invention is implemented.
[0029] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program is executed by the processor, any step of the full-scenario intelligent home supervision method based on graphic image recognition as described in the first aspect of the present invention is implemented.
[0030] The beneficial effects of the present invention are as follows: By combining hypersurface array-assisted multispectral fusion perception and topological photonic computing-accelerated anomaly detection, the comprehensive performance of smart home supervision is improved, and true full-scenario intelligent home supervision is realized. The data acquisition accuracy, processing speed, and adaptability to complex environmental changes are improved, providing users with a safer, more comfortable, and efficient living environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0032] Figure 1 It is a flowchart of the full-scenario intelligent home supervision method based on graphic image recognition in Embodiment 1.
[0033] Figure 2 It is a flowchart of the optimization of photonic computing in Embodiment 1.
[0034] Figure 3 It is a flowchart of quantum-enhanced anomaly detection in Embodiment 1.
[0035] Figure 4 It is a flowchart of the optimization of intelligent response in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be made in conjunction with the accompanying drawings of the specification.
[0037] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention, but the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention, so the present invention is not limited by the specific embodiments disclosed below.
[0038] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that excludes other embodiments.
[0039] Embodiment 1, referring to Figures 1 to 4 , is the first embodiment of the present invention. This embodiment provides a full-scenario intelligent home supervision method based on graphic image recognition, including the following steps:
[0040] S1: Collect environmental parameters in real time and perform preprocessing, dynamically adjust the phase distribution of the nano-scale units of the metasurface array, and generate a multi-spectral home environment perception matrix;
[0041] S1.1: Monitor environmental parameters including light, temperature, humidity, and smoke concentration in real time;
[0042] It should be noted that a micro-sensor array distributed in a honeycomb layout is integrated around the metasurface camera to form a full-scenario monitoring network, including:
[0043] Light detection: Based on the principle of the photoelectric effect, measure and capture abnormal changes in indoor lighting through a silicon-based photodiode;
[0044] Temperature and humidity detection: Use a capacitive sensing element to monitor the comfort level of the home environment by detecting changes in the dielectric constant;
[0045] Smoke concentration detection: Adopt laser scattering technology to infer the smoke concentration by detecting the scattered light intensity of particulate matter in the kitchen air, improving the fire warning ability.
[0046] Use a data fusion algorithm (such as Kalman filtering) to perform spatio-temporal alignment on environmental parameters and eliminate noise and instantaneous interference.
[0047] S1.2: Dynamically adjust the phase distribution of the nanostructure of the metasurface array according to environmental parameters;
[0048] It should be noted that the metasurface is composed of a titanium dioxide (TiO2) nano-column array, and adaptive imaging of the home scene is performed by changing the height and diameter of the nano-columns;
[0049] Each nano-column is connected to a piezoelectric ceramic actuator, and after receiving a control signal, it fine-tunes the physical height (accuracy ±5nm);
[0050] The control signal is generated in parallel by the FPGA, supporting independent parameter adjustment of 256×256 sub-blocks, with a refresh frequency of 100Hz.
[0051] Dynamically adjust the focal length parameter according to the light change rate and smoke concentration, and automatically increase or decrease the imaging focal length;
[0052] It should be noted that the light intensity and smoke concentration are collected in real time through environmental sensors, and the partial derivative of the light intensity with respect to time is calculated to obtain the light change rate; the light change rate is multiplied by the light change rate compensation coefficient, and the smoke concentration is multiplied by the smoke concentration compensation coefficient to obtain the compensation value; the compensation value is added to the reference focal length to generate a dynamic focal length parameter. When the light change rate increases or the smoke concentration rises, the dynamic focal length parameter automatically increases to enhance the imaging clarity; when the light change rate decreases or the smoke concentration decreases, the dynamic focal length parameter automatically decreases to optimize the imaging range, thereby realizing the adaptive adjustment of the focal length parameter according to the environmental conditions.
[0053] The preset value of the focal length parameter is 5 cm, which is optimized based on the target object distance in a typical home environment (such as objects within the range of 3 - 5 m).
[0054] When the light intensity change rate exceeds 1000 lux / s, for every increase of 1000 lux / s in the rate change, the focal length increases or decreases by 0.5 cm (the direction is determined by whether the light is increasing or decreasing).
[0055] For example: If the light intensity increases by 5000 lux within 0.1 second, then the focal length compensation amount Δf = 50,000 / 1000 × 0.5 cm = 25 cm.
[0056] For every increase of 200 μg / m 3 in the smoke concentration (S), the focal length increases by 0.3 cm to extend the depth of field of short - wave infrared (SWIR) and compensate for the light attenuation caused by smoke scattering.
[0057] The environmental compensation coefficient maps the temperature and humidity to the phase noise compensation amount through a non - linear function;
[0058] It should be noted that the combined influence of temperature and humidity is quantified into a coefficient within the range of [0, 1] through the Sigmoid function:
[0059] When T > 50 °C or H > 80%, the environmental compensation coefficient approaches 1, activating the maximum phase noise compensation (0.2π);
[0060] When T = 25 °C and H = 50%, the environmental compensation coefficient is 0.5, and the compensation amount is the reference value (0.1π).
[0061] The focal length parameter and the environmental compensation coefficient are set according to the environmental parameters, and the phase distribution is calculated in combination with the metasurface geometric characteristics. The expression is:
[0062]
[0063] Among them, φ(x, y, λ, t) is the phase distribution of the metasurface at coordinates (x, y), wavelength λ, and time t. λ is the wavelength of the incident light, f(t) is the dynamically adjusted focal length parameter at time t, (x, y) are the spatial coordinates on the metasurface plane, k(t) is the environmental compensation coefficient at time t, Δφ is the preset phase noise compensation amount, x is the direction coordinate parallel to the horizontal reference edge of the metasurface, y is the metasurface plane direction coordinate perpendicular to the x-axis, f0 = 5 is the reference focal length, α = 0.02 is the light intensity change rate compensation coefficient, β = 0.1 is the smoke concentration compensation coefficient, S(t) is the smoke concentration at time t, L(t) is the light intensity at time t, is the partial derivative of the light intensity L(t) with respect to time t;
[0064] It should be noted that f(t) 2 +x 2 +y 2 The physical meaning of is to calculate the square of the distance from the point on the metasurface plane to the focus, and based on the focal length formula in geometric optics, it is used to describe the optical path difference from the metasurface to the focus.
[0065] S1.3: Separate the incident light into three independent bands according to the wavelength built into the metasurface. Through spectral separation and synchronous imaging, synchronously output a three-channel image including visible light, near-infrared, and short-wave infrared, and generate a multi-spectral home environment perception matrix through data normalization processing;
[0066] Furthermore, use a multi-layer dielectric film beam splitter prism to separate the incident light into three independent channels according to the wavelength. The three-channel CMOS sensor adopts a global shutter and a hardware trigger mechanism to ensure that the time synchronization deviation is less than 10 microseconds;
[0067] It should be noted that after the incident light enters the multi-layer dielectric film beam splitter prism, through the multi-layer interference effect of the dielectric film, lights of different wavelengths are accurately separated and respectively reflected or transmitted to three independent channels, realizing the separation of visible light, near-infrared, and short-wave infrared spectra; the optical signals of each channel are received by the corresponding three-channel CMOS sensor. The three-channel CMOS sensor adopts the global shutter technology to ensure that all pixels are exposed at the same moment, avoiding motion blur; at the same time, the hardware trigger mechanism controls the acquisition time of the three-channel CMOS sensor to ensure that the time synchronization deviation between the three channels is less than 10 microseconds, thus realizing the high-precision synchronous acquisition of multi-spectral data.
[0068] The wavelength of visible light is 400 - 700 nm, which is used to capture the state of doors and windows and human activities;
[0069] The wavelength of near-infrared is 850 nm, which is used to penetrate curtains to monitor abnormal behaviors at night;
[0070] The wavelength of short-wave infrared is 1450nm, which is used to detect wall heat leakage and hidden fire sources.
[0071] Data standardization processing normalizes the light intensity of each channel, and attaches a high-precision timestamp (UTC time, error ±1 microsecond) and a spatial coordinate mapping table to each frame of data, generating a multi-spectral home environment perception matrix.
[0072] S2: Perform light-speed-level parallel computing on the multi-spectral home environment perception matrix through a photonic crystal waveguide array to generate a full-scene feature matrix and a photon confidence parameter;
[0073] S2.1: Separate the optical signals of the three channels of visible light, near-infrared, and short-wave infrared;
[0074] It should be noted that a hexagonal lattice photonic crystal (lattice constant 450 nanometers) is used, with air holes (diameter 80 - 200 nanometers) periodically arranged in a silicon substrate to form three independent waveguide channels, forming a three-channel processing core for smart homes;
[0075] S2.2: Map the multi-spectral home environment perception matrix into an optical signal through light intensity mapping and optical convolution calculation;
[0076] Furthermore, light intensity mapping converts the converted light intensity values of each channel of the multi-spectral home environment perception matrix into laser power, and generates an optical signal through a silicon-based microring modulator;
[0077] It should be noted that each channel of the multi-spectral home environment perception matrix first collects the light intensity values in the environment, and then maps these light intensity values into corresponding laser power values through a preset conversion algorithm.
[0078] The laser power value is transmitted as an input signal to the silicon-based microring modulator. The silicon-based microring modulator dynamically adjusts its resonance characteristics according to the received laser power value, thereby modulating the phase and intensity of the incident light. The modulated optical signal is generated through the output port of the silicon-based microring modulator, and finally an optical signal output that matches the environmental light intensity distribution is formed. This process realizes a seamless conversion from multi-spectral environment perception to optical signal generation, ensuring the accurate transmission and efficient utilization of light intensity information.
[0079] Optical convolution calculation performs edge enhancement on the optical signal through a 5×5 optical Gabor filter built into the photonic crystal waveguide, which is used to retain the dynamic range characteristics of the home scene;
[0080] Furthermore, according to the spatial frequency characteristics of the optical signal, the wavelength, direction, and bandwidth parameters of the 5×5 optical Gabor filter are designed to ensure that it can effectively extract the edge information in the scene.
[0081] It should be noted that the wavelength is selected according to the typical width of the edge features in the home scene, such as 3 - 5 pixels; the direction is set according to the main direction of the edges in the scene, such as 0°, 45°, 90°, 135°; the bandwidth is selected according to the complexity of the edge features in the scene, such as 1.0 - 1.5.
[0082] According to the parameters of the 5×5 optical Gabor filter, through the periodic refractive index distribution of the photonic crystal waveguide, the spatial frequency response characteristics of the filter are realized to ensure that the convolution operation is completed during the propagation of the optical signal.
[0083] The input optical signal is introduced into the photonic crystal waveguide. When the optical signal propagates in the waveguide, it undergoes convolution calculation through the built-in 5×5 optical Gabor filter to extract the high-frequency edge information in the scene.
[0084] During the convolution calculation process, the 5×5 optical Gabor filter enhances the edges of the optical signal, and at the same time, through the refractive index distribution characteristics of the photonic crystal waveguide, the dynamic range characteristics of the home scene are retained to avoid signal distortion.
[0085] The convolved optical signal is output from the photonic crystal waveguide to generate an edge-enhanced optical signal, providing high-precision and high-dynamic range scene feature data for subsequent processing.
[0086] Filter parameters: standard deviation of Gaussian envelope 1.2, center frequency 0.25 cycles / μm, phase shift 0 radians, optimized for door frame edges and human contours.
[0087] S2.3: Convert the optical signal into a full-scene feature matrix through graphene photoelectric detection and dynamic range compression.
[0088] It should be noted that the incident optical signal is first received by the graphene photodetector. The high sensitivity and wide spectral response characteristics of the graphene material ensure that the optical signal is efficiently converted into photocurrent. The photocurrent is amplified and subjected to analog-to-digital conversion to generate initial light intensity data. Subsequently, the dynamic range compression algorithm processes the initial light intensity data, compressing the high-dynamic range light intensity values to a low-dynamic range through non-linear mapping or logarithmic transformation to avoid data overexposure or underexposure. Finally, the compressed light intensity data is organized into a full-scene feature matrix, and each element in the matrix corresponds to the light intensity information at a specific position in the scene, providing a high-precision and highly consistent data basis for subsequent analysis and processing.
[0089] The specific steps for the dynamic range compression algorithm to process the initial light intensity data are as follows:
[0090] Based on the statistical results of historical data, preset the target light intensity value and the light intensity threshold.
[0091] Calculate the average light intensity value of the initial light intensity data and compare it with the preset target light intensity value to obtain a difference value;
[0092] If the absolute value of the difference exceeds the preset light intensity threshold, start dynamic range compression adjustment;
[0093] Calculate the initial mapping coefficient. Add the target light intensity value and the threshold to obtain a sum value. The initial mapping coefficient is the ratio of the average light intensity value to the sum value;
[0094] Dynamically adjust the mapping coefficient through the proportional relationship between the current light intensity value and the adjacent smaller light intensity value to ensure that the compression ratio in different light intensity intervals adapts to the scene requirements;
[0095] The steps for dynamically adjusting the mapping coefficient are as follows: Extract the current light intensity value (the first preset light intensity value) and the adjacent smaller light intensity value (the second preset light intensity value) from the light intensity value table. Use the difference between the current light intensity value and the adjacent smaller light intensity value as a benchmark, calculate the quotient of the current light intensity value and the initial mapping coefficient, and dynamically adjust the mapping coefficient corresponding to the current light intensity value according to the proportional relationship; Adjust the initial mapping coefficient according to the proportional relationship (such as 1.25). For example, if the proportional relationship indicates that the current light intensity value is relatively high, appropriately increase the mapping coefficient (such as from 0.8 to 1.0) to ensure the retention of details in the high-light area; if the proportional relationship indicates that the current light intensity value is relatively low, appropriately decrease the mapping coefficient (such as from 0.8 to 0.6) to ensure the retention of details in the low-light area. Ensure that the compression ratio smoothly transitions between the dark area and the bright area and retains scene details. Through the adjusted mapping coefficient, perform non-linear mapping on the light intensity data to compress the high dynamic range to the low dynamic range and avoid overexposure or underexposure;
[0096] Form a compression curve based on the initial mapping coefficient and the dynamically adjusted mapping coefficient;
[0097] Perform non-linear mapping on each light intensity value based on the compression curve to compress the high dynamic range light intensity values to the low dynamic range;
[0098] The compressed light intensity data is organized into a full-scene feature matrix according to spatial coordinates, where each element corresponds to the normalized light intensity value of the position in the scene;
[0099] Furthermore, the graphene-boron nitride heterojunction detector converts the optical signal into photocurrent for capturing weak home risk signals;
[0100] Perform logarithmic-linear hybrid compression on the photocurrent, retain key details such as nighttime intrusion and smoke penetration, and quantize them to 8-bit digital values (0 - 255) to adapt to the input specification of the YOLO-Q model;
[0101] S2.4: Through full-scene credibility assessment, fuse the signal-to-noise ratio and error to generate a confidence parameter;
[0102] Further, calculate the optical domain signal-to-noise ratio (signal mean / noise standard deviation), combine it with the local consistency error of the full-scene feature matrix, and fuse the mapped signal-to-noise ratio and the error through the Sigmoid function to generate a confidence parameter.
[0103] It should be noted that the calculation of the optical domain signal-to-noise ratio is achieved by dividing the mean value of the optical signal output by the photonic crystal waveguide array by the noise standard deviation, where the signal mean value comes from the average value of the photocurrent output by the graphene photodetector, and the noise standard deviation is obtained by statistically analyzing the fluctuation characteristics of the optical signal during transmission;
[0104] The local consistency error of the full-scene feature matrix is obtained by calculating the differences between adjacent pixels in the feature matrix, and is used to evaluate the continuity and stability of the scene features;
[0105] Calculate the differences between adjacent pixels in the feature matrix: For each pixel in the feature matrix, calculate the difference value between it and the pixels in its four-neighborhood or eight-neighborhood; usually in the form of absolute value or squared difference; then normalize the difference values to eliminate the influence of overall brightness or contrast; finally, statistically analyze the normalized difference values of all pixels, and take their average value or median as the local consistency error. The smaller this error value is, the higher the continuity and stability of the scene features, and vice versa, it indicates that there are discontinuous or unstable regions in the scene features.
[0106] Take the optical domain signal-to-noise ratio and the local consistency error as inputs, and perform normalized fusion through the Sigmoid function to generate a photon confidence parameter, whose value range is from 0 to 1, and is used to quantify the reliability and accuracy of the optical signal in scene perception.
[0107] S3: Input the full-scene feature matrix into the lightweight YOLO-Q model, perform full-scene risk analysis through quantum attention enhancement, and perform full-scene anomaly detection based on the photon confidence parameter to generate a structured alarm data packet;
[0108] S3.1: Input the full-scene feature matrix into the lightweight YOLO-Q model to extract the home supervision feature matrix;
[0109] It should be noted that the construction steps of the lightweight YOLO-Q model are as follows: perform channel pruning on the standard YOLOv4 model, remove redundant convolutional layers and feature channels, and retain the key feature extraction ability; quantize 32-bit floating-point weights and activation values into 8-bit fixed-point numbers to reduce the model storage and calculation overhead; in the multi-spectral feature fusion module, introduce a weighted fusion mechanism to fuse the features of the visible light, near-infrared, and short-wave infrared three channels; optimize the structure of the detection head part using 1×1 convolution and depthwise separable convolution;
[0110] The lightweight YOLO-Q model can directly receive an 8-bit quantized full-scene feature matrix as input, perform object detection through a backbone network, a multi-spectral feature fusion module, and a detection head, and output the object category, confidence, and bounding box coordinates to meet the requirements of the multi-spectral home environment perception scenario.
[0111] The training process of the lightweight YOLO-Q model The training process includes dataset preparation, model initialization, and optimized training of the multi-spectral feature fusion module and the detection head. It uses a multi-task loss function and quantization-aware training technology for joint optimization, and avoids overfitting through iterative training and early stopping strategies. Finally, the model performance is verified on the test set to ensure that it meets the requirements of the multi-spectral home environment perception scenario.
[0112] Normalize the full-scene feature matrix to ensure that the input data is within the range of [0, 1]. Input the normalized full-scene feature matrix into the lightweight YOLO-Q model, and use the quantum attention mechanism to optimize the calculation efficiency.
[0113] S3.2: Perform full-scene risk analysis on the home supervision feature matrix through the photon computing acceleration layer to generate the channel values of the home supervision feature matrix;
[0114] Furthermore, the photon computing acceleration layer uses a quantum dot array to achieve parallel computing, enhances the home supervision feature matrix, and enhances the output of the photon computing acceleration layer through the quantum attention mechanism to generate the channel values of the home supervision feature matrix;
[0115] It should be noted that the photon computing acceleration layer utilizes the parallel computing ability of the quantum dot array to enhance the features of the home supervision feature matrix. The quantum dot array maps the input feature matrix to a high-dimensional quantum state space through its unique quantum confinement effect to achieve rapid extraction and enhancement of feature information; the quantum attention mechanism adaptively enhances key features by calculating the feature weights output by the quantum dot array. The specific process includes generating a quantum state feature map, calculating the feature correlation matrix, normalizing the weights through the Softmax function, and finally fusing the weighted feature map with the original feature map to generate the channel values of the home supervision feature matrix.
[0116] S3.3: Perform anomaly detection on the home supervision feature matrix based on the photon confidence parameter to generate a structured alarm data packet;
[0117] Furthermore, in combination with the photon confidence parameter, perform anomaly detection on the home supervision feature matrix through the Gaussian weighting mechanism and the smoothing factor, and calculate the anomaly score based on the photon confidence parameter, the enhanced value of the home supervision feature matrix, the Gaussian distribution center position, and the smoothing factor. The expression is:
[0118]
[0119] Among them, D(a, b, d) is the anomaly score of the d-th channel at position (a, b) in the home supervision feature matrix, ρ is the photon confidence parameter, G(a, b, d) is the value of the d-th channel at position (a, b) in the home supervision feature matrix, (a c , b c ) is the central position of the home supervision feature matrix, ω is the standard deviation of the Gaussian distribution, ζ is the smoothing factor, H is the height of the target feature matrix, W is the width of the home supervision feature matrix, a is the row index of the home supervision feature matrix, b is the column index of the home supervision feature matrix, G(i, j, d) is the value of the d-th channel at position (i, j) in the enhanced home supervision feature matrix, i is the row index of the enhanced home supervision feature matrix, and j is the column index of the enhanced home supervision feature matrix.
[0120] S3.4: Generate a structured alarm data packet based on the anomaly score;
[0121] Furthermore, set an anomaly threshold according to historical data. When the anomaly score is greater than the anomaly threshold, extract the anomaly information of the abnormal area, including the abnormal position, time, confidence, and severity, and generate a structured alarm data packet.
[0122] It should be noted that the anomaly threshold is set according to historical data, and the value range of the anomaly threshold is from 0.5 to 1.0. The specific value is determined by statistically calculating the average value and standard deviation of historical anomaly scores, and the default value is 0.5;
[0123] When the anomaly score is greater than the anomaly threshold, extract the anomaly information of the abnormal area, including the abnormal position, time, confidence, and severity. Among them, the abnormal position is determined by the coordinates of the anomaly score in the feature matrix, the time is the timestamp when the anomaly is detected, the confidence is the value of the photon confidence parameter, and the severity is divided into three levels: low (difference ≤ 0.2), medium (0.2 < difference ≤ 0.5), and high (difference > 0.5) according to the difference between the anomaly score and the anomaly threshold;
[0124] Integrate the abnormal position, time, confidence, and severity into a structured alarm data packet for subsequent analysis and processing.
[0125] S4: Perform emergency handling, environmental adjustment, and parameter optimization according to the structured alarm data packet;
[0126] S4.1: Trigger the emergency handling mechanism according to the alarm data packet;
[0127] Furthermore, according to the anomaly type (such as fire, smoke, intrusion, etc.) in the alarm data packet, classify and trigger different emergency handling mechanisms to generate emergency handling instructions, including:
[0128] Level 1 Alarm: Fire / Intrusion, triggering audible and visual alarms;
[0129] Level 2 Alarm: Water Leakage / Gas Leakage, closing the solenoid valve;
[0130] Level 3 Alarm: Abnormal temperature and humidity, starting environmental regulation;
[0131] Associate with data in the same area in the past 24 hours to evaluate the severity.
[0132] S4.2: Dynamically adjust environmental parameters according to the abnormal information in the alarm data packet;
[0133] It should be noted that, combined with the abnormal type, location and severity, targeted adjustment is carried out on key devices in the home environment;
[0134] For lighting equipment, adjust the lighting intensity near the abnormal location according to the severity. When the abnormal severity is high, increase it to the maximum brightness and trigger audible and visual alarms. When the abnormal severity is medium, set it to medium brightness and flash for reminder. When the abnormal severity is low, slightly increase the brightness;
[0135] For air conditioning equipment, turn off the air conditioner in case of fire or smoke abnormality to prevent the spread of fire or smoke, and adjust the operation mode in case of abnormal temperature and humidity to restore environmental comfort;
[0136] For security equipment, start the real-time recording and alarm functions of the surveillance camera in case of intrusion abnormality;
[0137] For energy equipment, close the solenoid valve to cut off the energy supply in case of gas leakage or water leakage abnormality; Through the above methods, dynamically adjust environmental parameters combined with abnormal information to achieve multiple goals of risk control, environmental restoration and energy optimization.
[0138] S4.3: Optimize the metasurface parameters according to the confidence level and severity in the alarm data packet.
[0139] It should be noted that, based on the confidence level and severity, adjust the focal length parameter of the metasurface. For abnormalities with high confidence level and high severity, adjust the focal length parameter to a short focal length to enhance the ability to capture local details. For abnormalities with low confidence level and low severity, adjust the focal length parameter to a long focal length to expand the monitoring range;
[0140] Adjust the environmental compensation coefficient according to the change of environmental parameters. When the light intensity or the refractive index of the medium changes significantly, correct the phase distribution through the environmental compensation coefficient to ensure the imaging accuracy;
[0141] Generate control signals in parallel through FPGA with the optimized metasurface parameters, and adjust the height and diameter of the titanium dioxide nanocolumn array in real time to achieve adaptive imaging in the home scene and improve the accuracy and reliability of abnormal detection.
[0142] This embodiment also provides an all-scenario intelligent home supervision system based on graphic image recognition, including: a dynamic regulation module, which is used to collect environmental parameters in real time, preprocess them, dynamically adjust the phase distribution of the nano-scale units of the metasurface array, and generate a multi-spectral home environment perception matrix; a photon computing module, which is used to perform light-speed-level parallel computing on the multi-spectral home environment perception matrix through a photonic crystal waveguide array to generate an all-scenario feature matrix and a photon confidence parameter; an anomaly detection module, which is used to input the all-scenario feature matrix into a lightweight YOLO-Q model, perform all-scenario risk analysis through quantum attention enhancement, and perform all-scenario anomaly detection based on the photon confidence parameter to generate a structured alarm data packet; a response handling module, which is used to perform emergency handling, environmental regulation, and parameter optimization according to the structured alarm data packet.
[0143] This embodiment also provides a computer device applicable to the case of the all-scenario intelligent home supervision method based on graphic image recognition, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the all-scenario intelligent home supervision method based on graphic image recognition proposed in the above embodiment.
[0144] This computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a carrier network, NFC (near-field communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0145] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the full-scenario intelligent home supervision method based on graphic image recognition proposed in the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk or optical disk.
[0146] In summary, the present invention combines multi-spectral fusion perception assisted by a metasurface array and anomaly detection accelerated by topological photonic computing, improves the comprehensive performance of the smart home system, and realizes full-scenario intelligent home supervision in a true sense. It improves the data acquisition accuracy, processing speed, and adaptability to complex environmental changes, providing a safer, more comfortable, and efficient living environment for users.
[0147] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A full-scenario intelligent home supervision method based on graphic and image recognition, characterized in that: including, real-time collecting environmental parameters and preprocessing them, dynamically adjusting the phase distribution of nano-scale units of the metasurface array, and generating a multi-spectral home environment perception matrix; performing light-speed level parallel computing on the multi-spectral home environment perception matrix through a photonic crystal waveguide array to generate a full-scene feature matrix and photon confidence parameters; inputting the full-scene feature matrix into a lightweight YOLO-Q model, performing full-scene risk analysis through quantum attention enhancement, and performing full-scene anomaly detection based on the photon confidence parameters to generate a structured alarm data packet; performing emergency handling, environmental adjustment, and parameter optimization according to the structured alarm data packet.
2. The full-scenario intelligent home supervision method based on graphic image recognition according to claim 1, characterized in that: The step of dynamically adjusting the phase distribution of nano-scale units of the metasurface array and generating a multi-spectral home environment perception matrix is as follows. Dynamically adjusting the phase distribution of the nanostructure of the metasurface array according to environmental parameters; Separating the incident light into three independent bands according to the adjusted metasurface array; Obtaining three-channel images including visible light, near-infrared, and short-wave infrared through spectral separation and synchronous imaging; Generating a multi-spectral home environment perception matrix through data normalization processing.
3. The full-scenario intelligent home supervision method based on graphic image recognition according to claim 1, characterized in that: The preprocessing is to use a data fusion algorithm to perform spatio-temporal alignment on environmental parameters to eliminate noise and instantaneous interference.
4. The full-scenario intelligent home supervision method based on graphic image recognition according to claim 2, characterized in that: The step of performing light-speed level parallel computing on the multi-spectral home environment perception matrix through a photonic crystal waveguide array to generate a full-scene feature matrix and photon confidence parameters is as follows. Mapping the multi-spectral home environment perception matrix into an optical signal through light intensity mapping and optical convolution calculation; Converting the optical signal into a full-scene feature matrix through graphene photoelectric detection and dynamic range compression; Generating confidence parameters through full-scene credibility evaluation.
5. The full-scenario intelligent home supervision method based on graphic image recognition according to claim 4, characterized in that: The step of inputting the full-scene feature matrix into a lightweight YOLO-Q model and performing full-scene risk analysis through quantum attention enhancement is as follows. Inputting the full-scene feature matrix into a lightweight YOLO-Q model to extract a home supervision feature matrix; Performing full-scene risk analysis on the home supervision feature matrix through a photon computing acceleration layer to generate the channel values of the home supervision feature matrix.
6. The full-scenario intelligent home supervision method based on graphic image recognition according to claim 5, characterized in that: The step of performing anomaly detection based on the photon confidence parameters to generate a structured alarm data packet is as follows. Combining the photon confidence parameters to perform anomaly detection on the home supervision feature matrix and calculating the anomaly score; Generating a structured alarm data packet according to the anomaly score and extracting the anomaly information.
7. The full-scenario intelligent home supervision method based on graphic image recognition according to claim 6, characterized in that: The step of performing emergency handling, environmental adjustment, and parameter optimization according to the structured alarm data packet is as follows. Generating an emergency handling instruction from the anomaly information in the alarm data packet and dynamically adjusting the environmental parameters; Optimizing the metasurface parameters according to the confidence and severity in the alarm data packet.
8. A full-scenario intelligent home supervision system based on graphic image recognition, based on the full-scenario intelligent home supervision method based on graphic image recognition according to any one of claims 1 to 7, characterized in that: including a dynamic regulation module, a photon computing module, an anomaly detection module, and a response handling module. The dynamic regulation module is used to collect environmental parameters in real time and perform preprocessing, dynamically adjust the phase distribution of nano-scale units of the metasurface array, and generate a multi-spectral home environment perception matrix; The photon computing module is used to perform light-speed level parallel computing on the multi-spectral home environment perception matrix through a photonic crystal waveguide array to generate a full-scene feature matrix and photon confidence parameters; Anomaly detection module, which is used to input the full-scenario feature matrix into the lightweight YOLO-Q model, perform full-scenario risk analysis through quantum attention enhancement, and conduct full-scenario anomaly detection based on photon confidence parameters to generate structured alarm data packets; Response handling module, which is used to perform emergency handling, environmental adjustment and parameter optimization according to the structured alarm data packets.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the full-scenario intelligent home supervision method based on graphic image recognition according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the full-scenario intelligent home supervision method based on graphic image recognition according to any one of claims 1 to 7.