Target equipment behavior control and joint training method and device

By modulating the optical signal with optical components and combining the detection model, the image quality and accuracy of the control results are improved, the accuracy of the water-filled shutdown control strategy in the prior art is solved, and efficient and safe water resource utilization is achieved.

CN119924696APending Publication Date: 2025-05-06SHPHOTONICS LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202412000334.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When using the water-filled control strategy, the image quality is easily affected by various factors, resulting in a decrease in the accuracy of the analysis results and control results.

Method used

The optical signal is modulated through optical elements, filter background light and reflected interference, improve the image signal-to-noise ratio, and use the detection model to generate the detection results of the target image to control the behavior of the target device.

Benefits of technology

It improves image quality, improves the accuracy of detection and control results, and ensures efficient utilization of water resources and the safety of target equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119924696A_ABST
    Figure CN119924696A_ABST
Patent Text Reader

Abstract

The invention provides a target equipment behavior control and joint training method and device, and relates to the artificial intelligence fields of deep learning, computer vision, image processing, optical design and the like. The target equipment behavior control method comprises the steps that a to-be-processed target image is acquired, the target image is a to-be-detected area image generated by a photosensitive element according to an acquired second optical signal, and the second optical signal is an optical signal obtained after an optical element modulates an acquired first optical signal, the first optical signal is an optical signal that light emitted by the light source irradiates an area to be detected and then enters the optical element through reflection; a detection result corresponding to the target image is generated through the detection model, the behavior of the target equipment is controlled according to the detection result, and the to-be-detected area is opposite to a material outlet in the target equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the fields of deep learning, computer vision, image processing, and optical design, and specifically to a method and apparatus for target device behavior control and joint training. Background Art

[0002] At present, water dispensers are widely used in different places such as homes and offices. In order to reduce the waste of water resources, a water-full-stop control strategy is usually adopted, that is, if it is determined that the water in the cup is full, the water will stop flowing. Summary of the invention

[0003] The present disclosure provides a method and apparatus for target device behavior control and joint training.

[0004] A target device behavior control method, comprising:

[0005] Acquire a target image to be processed, wherein the target image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal, wherein the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, and the first light signal is a light signal emitted by a light source and then reflected by the light entering the optical element after irradiating the area to be detected;

[0006] A detection result corresponding to the target image is generated by using a detection model, and the behavior of the target device is controlled according to the detection result, and the area to be detected is opposite to a material outlet in the target device.

[0007] A joint training method of a detection model and an optical element, comprising:

[0008] Acquire an initial image, wherein the initial image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal, wherein the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, and the first light signal is a light signal emitted by a light source and then reflected by the light entering the optical element after irradiating the area to be detected;

[0009] For each initial image, generate corresponding training samples respectively;

[0010] The detection model and the optical element are jointly trained using the training samples.

[0011] A target device behavior control device comprises: a first acquisition module and a detection control module;

[0012] The first acquisition module is used to acquire a target image to be processed, wherein the target image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, and the first light signal is a light signal emitted by a light source and then reflected by the light entering the optical element after irradiating the area to be detected;

[0013] The detection control module is used to generate a detection result corresponding to the target image using a detection model, and control the behavior of the target device according to the detection result. The area to be detected is opposite to a material outlet in the target device.

[0014] A target device comprises the target device behavior control device as described above.

[0015] A joint training device for a detection model and an optical element, comprising: a second acquisition module, a sample generation module and a model training module;

[0016] The second acquisition module is used to acquire an initial image, wherein the initial image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by the optical element modulating the acquired first light signal, and the first light signal is a light signal emitted by the light source and then reflected by the light entering the optical element after irradiating the area to be detected;

[0017] The sample generation module is used to generate corresponding training samples for each initial image;

[0018] The model training module is used to jointly train the detection model and the optical element using the training samples.

[0019] An electronic device, comprising:

[0020] at least one processor; and

[0021] a memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0023] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.

[0024] A computer program product comprises a computer program / instruction, wherein the computer program / instruction implements the method described above when executed by a processor.

[0025] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0027] Figure 1 A flowchart of an embodiment of the target device behavior control method disclosed in the present invention;

[0028] Figure 2 It is a schematic diagram of the composition structure of the detection device disclosed in the present invention;

[0029] Figure 3 It is a schematic diagram of the position of the detection device disclosed in the present invention;

[0030] Figure 4 This is a schematic diagram of the judgment logic of the three test results disclosed in this disclosure;

[0031] Figure 5 A flowchart of an embodiment of a joint training method of a detection model and an optical element disclosed in the present invention;

[0032] Figure 6 A flowchart of an embodiment of the iterative training method disclosed in the present invention;

[0033] Figure 7 Schematic diagram of the structure of the target device behavior control device embodiment 700 of the present disclosure;

[0034] Figure 8 It is a schematic diagram of the composition structure of an embodiment 800 of the joint training device for the detection model and the optical element disclosed in the present invention;

[0035] Fig. 9 A schematic block diagram of an electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0036] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0037] In addition, it should be understood that the term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0038] Figure 1 Flow chart of an embodiment of the target device behavior control method disclosed in the present invention. Figure 1 As shown, the following specific implementation methods are included.

[0039] In step 101, a target image to be processed is obtained. The target image is an image of the area to be detected generated by the photosensitive element according to the obtained second light signal. The second light signal is a light signal obtained by modulating the obtained first light signal by the optical element. The first light signal is a light signal emitted by a light source and then reflected into the optical element after irradiating the area to be detected.

[0040] In step 102, a detection result corresponding to the target image is generated using the detection model, and the behavior of the target device is controlled according to the detection result. The area to be detected is opposite to the material outlet in the target device.

[0041] Traditional water-full stop control mainly relies on images collected by ordinary visual cameras, that is, the collected images are analyzed to determine whether the water in the water cup is full, and when it is full, the water dispenser is controlled to stop discharging water, etc. However, the collected images may be affected by various factors, resulting in poor image quality, which in turn reduces the accuracy of the analysis results and correspondingly reduces the accuracy of the control results of the target device.

[0042] By adopting the scheme described in the above method embodiment, optical elements can be used to effectively modulate the light signal, filter out background stray light and reflection interference at the optical level, improve the image signal-to-noise ratio, etc., and then generate a target image based on the modulated light signal, thereby improving the image quality, that is, optimizing the imaging quality, and then improving the accuracy of subsequent detection results and control results. Moreover, a detection model can be used to generate detection results corresponding to the target image, thereby utilizing the powerful reasoning ability of the detection model, thereby further improving the accuracy of the detection results, and further improving the accuracy of the control results, etc.

[0043] In practical applications, the scheme described in the above method embodiment can be executed in real time. In addition, the target device can be various devices such as a water dispenser, a water purifier, a coffee machine, a beverage machine, and an ice cream machine, which has wide applicability.

[0044] In some embodiments of the present disclosure, the light source, the optical element, and the photosensitive element may all be located in the detection device, and the detection device may further include: a high-transmittance sheet, and the first light signal may be a light signal that is emitted by the light source and then irradiates the area to be detected, and then is reflected and passes through the high-transmittance sheet before entering the optical element. The high-transmittance sheet may be used to protect the internal components of the detection device, and may be used to weaken the influence of light in other bands other than the light source, thereby improving the accuracy of the detection results.

[0045] Accordingly, Figure 2 Schematic diagram of the structure of the detection device disclosed in the present invention. Figure 2 As shown, the detection device may include a light source, a high-transmittance sheet, an optical element, and a photosensitive element.

[0046] In some embodiments of the present disclosure, the optical element may include: a metasurface element. In addition, the metasurface element may include at least two partitions, each partition may include a structural unit set, each structural unit set may include at least two nanostructure units, and each structural unit set is used to generate a second optical signal.

[0047] A metasurface is an artificial layered material with a size less than or approximately equal to the wavelength. It can be regarded as the two-dimensional counterpart of a metamaterial. It can modulate the polarization, phase, amplitude, frequency, propagation mode, angular momentum and other characteristics of electromagnetic waves through sub-wavelength metastructure units on the surface, thereby achieving beam shaping, beam deflection, superlens, super holography, optical rotation, and anti-reflection and anti-reflection properties.

[0048] In the scheme disclosed in the present invention, the metasurface element can be divided into multiple partitions, each of which can be configured with a structural unit set composed of multiple nanostructure units. The structural unit set of each partition can express an imaging function, that is, generate a second light signal respectively. In this way, the photosensitive element can generate a target image based on the multiple second light signals obtained.

[0049] By dividing the metasurface element into multiple partitions, it can focus on different areas of the image separately. For example, assuming that there are multiple partitions including partition 1, partition 2 and partition 3, taking a water dispenser as an example, partition 1 can focus light on the mouth of the water cup, partition 2 can focus light on the bottom of the water cup, and partition 3 can focus light on the water surface. Accordingly, based on the multiple image information corresponding to the multiple partitions, the detection model can better perceive the characteristics of the mouth of the water cup, the bottom of the water cup and the water surface, and thus obtain more accurate detection results.

[0050] In addition to being a metasurface element, the optical element in the solution disclosed in the present invention may also be a refractive optical element, a diffractive optical element, or a scattering medium element, etc. The above description takes a metasurface element as an example.

[0051] In some embodiments of the present disclosure, the area to be detected may include: a container placement area, the center of the container placement area may be opposite to the material outlet of the target device, the detection device may be located within a predetermined area centered on the material outlet of the target device, and the metasurface element may be inclined at a predetermined angle to the plane where the container placement area is located. The specific values ​​of the predetermined area range and the predetermined angle may be determined according to actual needs, and are not limited in the scheme described in the present disclosure. For example, the predetermined angle may be 10 degrees or 15 degrees, etc. In addition, in actual applications, the container placement area may be part of the target device or not.

[0052] Taking a water dispenser as an example, the container placement area refers to the water receiving area of ​​the water dispenser, the container refers to a container that can hold water, such as a cup, the material refers to water, and the material outlet refers to the water outlet of the water dispenser.

[0053] Figure 3 Schematic diagram of the position of the detection device described in the present disclosure. Figure 3 As shown, the detection device can be located around the water outlet of the water dispenser, and the metasurface element and the plane where the water receiving area is located can be inclined at an angle of about 10 degrees, so that the metasurface element can better obtain the required optical signal, thereby improving the accuracy of subsequent detection results.

[0054] The light source may be a monochromatic light source of a specific wavelength, or a broadband light source (i.e., a band light source) covering multiple wavelengths. The monochromatic light source of a specific wavelength may be, for example, a 850 nanometer (nm) wavelength light source or a 940 nanometer (nm) wavelength light source, and the band light source may be, for example, a light source covering wavelengths from 850nm to 940nm. Accordingly, the high-transmittance sheet may be a sheet that is transparent to the light emitted by the light source. Alternatively, the light source may also include a 940nm wavelength light source and other visible light sources (such as natural light sources), etc., so as to be able to effectively cope with various complex environments and improve processing performance. The light emitted by the light source irradiates the water receiving area, and the light enters the optical element as a first light signal after being reflected and passing through the high-transmittance sheet. The optical element modulates the first light signal to generate a second light signal and outputs it to the photosensitive element. The photosensitive element generates a target image through photoelectric conversion processing.

[0055] After acquiring the target image, it can also be formatted. The formatting process may include converting the target image into a standard grayscale image and normalizing the size to a predetermined size, etc. The predetermined size may be 224*224. The formatted target image can then be input into the detection model to obtain the output detection result.

[0056] In some embodiments of the present disclosure, the detection results may include: a first detection result, a second detection result, and a third detection result, wherein the first detection result indicates that there is no obstacle in the target path, and there is a container in the area to be detected, and the material in the container is not full, and the target path is the path of the material flowing from the material outlet of the target device into the container, and the second detection result indicates that there is no obstacle in the target path, and there is no container in the area to be detected, or, the second detection result indicates that there is no obstacle in the target path, and there is a container in the area to be detected, and the material in the container is full, and the third detection result indicates that there is an obstacle in the target path.

[0057] The detection model may include a network input layer, a convolution layer, and a classification layer. The network input layer can be used to obtain a formatted target image, the convolution layer can be used to extract features such as edges and textures of the image, and the classification layer can be used to generate and output detection results based on the extracted features.

[0058] Based on the target image, the detection model can determine whether there are obstacles in the target path, whether there are containers in the area to be detected, and whether the materials in the container are full. Therefore, in theory, the detection model can output 8 (2*2*2) detection results, but in order to better match the actual needs, the solution described in this disclosure defines 3 detection results, namely the first detection result, the second search result, and the third detection result. In addition, taking the water dispenser as an example, the obstacle can be a hand.

[0059] Assume that the first test result is represented by 0, the second test result is represented by 1, and the third test result is represented by 2. Figure 4 Schematic diagram of the judgment logic of the three detection results described in this disclosure. Figure 4 As shown, if it is determined that there is an obstacle in the target path, the detection result output by the detection model may be 2; if it is determined that there is no obstacle in the target path and it is determined that there is no container in the area to be detected, the detection result output by the detection model may be 1; if it is determined that there is no obstacle in the target path and it is determined that there is a container in the area to be detected and it is determined that the material in the container is full, the detection result output by the detection model may also be 1; if it is determined that there is no obstacle in the target path and it is determined that there is a container in the area to be detected and it is determined that the material in the container is not full, the detection result output by the detection model may be 0.

[0060] It can be seen that the above processing method can efficiently and accurately generate various required detection results, thus laying a good foundation for subsequent processing.

[0061] According to the detection result, the behavior of the target device can be controlled. In some embodiments of the present disclosure, in response to determining that the detection result is the second detection result, a first instruction can be sent to the control center in the target device to instruct the control center to stop the outflow of the material, and in response to determining that the detection result is the third detection result, a second instruction can be sent to the control center in the target device to instruct the control center to issue an alarm and / or control the outflow of the material.

[0062] Taking a water dispenser as an example, if the detection result is determined to be the second detection result, the control center can be used to control the water dispenser to stop discharging water. If the detection result is determined to be the third detection result, an alarm can be issued through the control center, and the water dispenser can be controlled by the control center to stop discharging water.

[0063] Based on the above introduction, it can be seen that the solution disclosed in the present invention supports multi-task detection, that is, it can realize water full detection and hand intrusion detection at the same time, and can effectively control the water dispenser according to the detection results, such as avoiding misoperation or safety hazards caused by hand intrusion, thereby improving safety, and can efficiently control water output and stop water output, thereby reducing the waste of water resources. In addition, the solution disclosed in the present invention adopts a combination of soft and hard implementation methods, that is, using metasurface elements to improve image quality, and using detection models to optimize detection results, thereby providing a more efficient and reliable solution for the realization of smart homes.

[0064] The detection model and the optical element may be obtained through pre-training. The above mainly describes the model reasoning process. The following will describe the training process of the detection model and the optical element in conjunction with the embodiments.

[0065] Figure 5 Flow chart of an embodiment of the joint training method of the detection model and the optical element disclosed in the present invention. Figure 5 As shown, the following specific implementation methods are included.

[0066] In step 501, an initial image is acquired. The initial image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal. The second light signal is a light signal obtained by modulating the acquired first light signal by the optical element. The first light signal is a light signal emitted by a light source and then reflected into the optical element after irradiating the area to be detected.

[0067] In step 502, corresponding training samples are generated for each initial image.

[0068] In step 503, the detection model and the optical element are jointly trained using the training samples.

[0069] By adopting the scheme described in the above method embodiment, end-to-end joint training of the detection model and the optical element can be achieved. The detection model can perform detection based on the initial image generated with the aid of the optical element, and the detection model and the optical element can be updated based on the detection results output by the detection model, thereby ultimately determining the configuration combination with the best cooperation effect between the two.

[0070] The detection model can be a convolutional neural network (CNN) model, such as a residual network (ResNet) model, an efficient convolutional neural network architecture (EfficientNet) model, or a lightweight deep convolutional neural network architecture (ShuffleNet) model, etc. The specific model to be used can be determined according to the actual device performance requirements.

[0071] The initial image can be generated in the same way as the target image, that is, the light emitted by the light source irradiates the area to be detected, and the light enters the optical element as a first light signal after being reflected and passing through the high-transmittance sheet. The optical element modulates the first light signal to generate a second light signal and outputs it to the photosensitive element, and the photosensitive element generates the initial image through photoelectric conversion processing. Among them, the light source, the high-transmittance sheet, the optical element and the photosensitive element can all be located in the detection device.

[0072] In addition, in some embodiments of the present disclosure, the area to be inspected may include: a container placement area, the center of the container placement area may be opposite to the material outlet of the target device, and the container is used to hold the material flowing into the container from the material outlet of the target device. The method of obtaining the initial image may include: respectively obtaining initial images generated under M different imaging conditions, where M is a positive integer greater than 1, and the specific value may be determined according to actual needs, wherein the difference between any two different imaging conditions includes one or all of the following: different light exposure conditions for the area to be inspected and different background environments where the target device is located.

[0073] Different light irradiation conditions may refer to using different light sources, such as using a 940nm wavelength light source, or using a 940nm wavelength light source and other visible light sources at the same time, and the visible light source can be divided into low visible light and high visible light, etc. In addition, taking a water dispenser as an example, the background environment where the target device is located may refer to placing the water dispenser on a desktop of different materials, such as a marble desktop, a wooden desktop, a glass desktop, and a composite material desktop.

[0074] Through the above processing, the types of training samples generated subsequently can be enriched so that they can cover different application scenarios, thereby improving the training effect of the detection model.

[0075] For each acquired initial image, a corresponding training sample can be generated. In some embodiments of the present disclosure, for each initial image, the following processing can be performed: obtaining a label corresponding to the initial image, formatting the initial image, determining the formatted initial image as a sample image, and using the label corresponding to the sample image to form a training sample. Through the formatting process, the format of the sample image can be standardized, so as to facilitate the detection model to process, etc.

[0076] There is no restriction on how to obtain the label corresponding to the initial image, for example, the label can be obtained by manual annotation. In addition, in some embodiments of the present disclosure, the label may include: a first detection result, a second detection result, and a third detection result, the first detection result indicating that there are no obstacles in the target path, and there is a container in the area to be detected, and the material contained in the container is not full, the target path is the path of the material flowing from the material outlet of the target device into the container, the second detection result indicating that there are no obstacles in the target path, and there is no container in the area to be detected, or the second detection result indicating that there are no obstacles in the target path, and there is a container in the area to be detected, and the material contained in the container is full, and the third detection result indicating that there are obstacles in the target path.

[0077] Based on the initial image, it can be determined whether there are obstacles in the target path, whether there are containers in the area to be detected, and whether the materials in the container are full. Therefore, in theory, 8 (2*2*2) detection results can be obtained, but in order to better match the actual needs, 3 detection results are defined in the scheme described in the present disclosure, namely the first detection result, the second search result, and the third detection result. Correspondingly, the label can be the first detection result, the second detection result, or the third detection result. Among them, the first detection result can be represented by 0, the second detection result can be represented by 1, and the third detection result can be represented by 2.

[0078] Each training sample may be saved in a standard data format, such as storing the sample image in a Joint Photographic Experts Group (JPG) file format and storing the corresponding label in a comma-separated values ​​(CSV) format.

[0079] In addition, in practical applications, the obtained training samples can also be divided into training set, validation set and test set, among which the number of training samples in the training set can account for 80% of the total training samples, the number of training samples in the validation set can account for 10% of the total training samples, and the number of training samples in the test set can also account for 10% of the total training samples.

[0080] When jointly training the detection model and the optical element, in some embodiments of the present disclosure, the detection result output by the detection model can be obtained for a sample image, and a loss value (loss) can be determined based on the detection result and the label corresponding to the sample image. Furthermore, in response to determining that the loss value satisfies a predetermined condition, the training process can be terminated, and in response to determining that the loss value does not satisfy the predetermined condition, the optical element and / or the detection model can be updated based on the loss value.

[0081] In addition, in some embodiments of the present disclosure, for a sample image, a method for obtaining a detection result output by a detection model may include: performing data enhancement processing on the sample image, inputting the sample image after data enhancement processing into the detection model, and obtaining an output detection result.

[0082] The above process can be iterated continuously until the preset conditions are met, thereby ending the training process.

[0083] Accordingly, Figure 6 Flow chart of an embodiment of the iterative training method described in the present disclosure. Figure 6 As shown, the following specific implementation methods are included.

[0084] In step 601, data enhancement processing is performed on the acquired sample image.

[0085] For example, data augmentation processing may refer to adding random noise to a sample image or performing random brightness changes to improve the generalization capability of the detection model.

[0086] In step 602, the sample image after data enhancement processing is input into the detection model to obtain an output detection result.

[0087] In step 603, a loss value is determined according to the detection result and the label corresponding to the sample image.

[0088] When determining the loss value, the loss function used may be a cross entropy loss function or the like.

[0089] In step 604 , it is determined whether the loss value satisfies a predetermined condition. If so, the process ends; otherwise, step 605 is executed.

[0090] The specific conditions of the predetermined conditions can be determined according to actual needs.

[0091] In step 605 , the optical element and / or the detection model is updated according to the loss value, and then step 601 is repeated.

[0092] Among them, an adaptive moment estimation (ADAM) algorithm can be used as an optimizer to update the model parameters of the detection model, and the model parameters may include convolution kernel weights and fully connected layer weights.

[0093] In some embodiments of the present disclosure, the optical element may be a metasurface element. In addition, the metasurface element may include at least two partitions, each partition may include a structural unit set, each structural unit set may include at least two nanostructure units, and each structural unit set is used to generate a second optical signal.

[0094] Accordingly, in some embodiments of the present disclosure, updating the metasurface element according to the loss value may include: adjusting the nanostructure units in at least one structure unit set according to the loss value.

[0095] In the solution disclosed in the present invention, the metasurface element can be divided into multiple partitions, each partition can be respectively configured with a structural unit set consisting of multiple nanostructure units, and the structural unit set of each partition can express an imaging function, that is, generate a second light signal respectively, so that the photosensitive element can generate an initial image based on the acquired multiple second light signals. When the metasurface element needs to be updated, the nanostructure units in one or more of the structural unit sets can be adjusted according to actual needs.

[0096] In addition, during the training stage, the output of the metasurface element can be determined by the optical response or optical model of the metasurface determined by existing optical simulation methods, that is, the metasurface element can be a simulated optical response relationship, and its output can be determined by calculation. In addition, the output of the metasurface element can also be determined by constructing a physical metasurface for experiment.

[0097] After training is completed, the trained detection model can be directly used for Figure 1 The model shown is inferred and can generate target images using the trained optical elements.

[0098] Alternatively, in some embodiments of the present disclosure, the trained detection model may be quantized, and then the quantized detection model may be used to perform Figure 1 Reasoning for the model shown.

[0099] For example, the trained detection model can be quantized to 8-bit integer (INT8) to reduce the detection model's demand for computing resources, improve detection efficiency, and reduce inference time.

[0100] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the described order of actions, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure. In addition, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0101] The above is an introduction to the method embodiment. The following is a further explanation of the scheme disclosed in the present invention through an apparatus embodiment.

[0102] Figure 7 FIG. 7 is a schematic diagram of the structure of the target device behavior control device embodiment 700 described in the present disclosure. Figure 7 As shown, it includes: a first acquisition module 701 and a detection control module 702.

[0103] The first acquisition module 701 is used to acquire a target image to be processed. The target image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal. The second light signal is a light signal obtained by modulating the acquired first light signal by the optical element. The first light signal is a light signal emitted by a light source and then reflected into the optical element after irradiating the area to be detected.

[0104] The detection control module 702 is used to generate a detection result corresponding to the target image using the detection model, and control the behavior of the target device according to the detection result. The area to be detected is opposite to the material outlet in the target device.

[0105] In some embodiments of the present disclosure, the light source, the optical element and the photosensitive element can all be located in the detection device, and the detection device can also include: a high-transmittance thin sheet, and the first light signal can be a light signal emitted by the light source, which is then reflected by the high-transmittance thin sheet and enters the optical element after irradiating the area to be detected.

[0106] In some embodiments of the present disclosure, the optical element may include: a metasurface element. In addition, the metasurface element may include at least two partitions, each partition may include a structural unit set, each structural unit set may include at least two nanostructure units, and each structural unit set is used to generate a second optical signal.

[0107] In some embodiments of the present disclosure, the area to be detected may include: a container placement area, the center of the container placement area may be opposite to the material outlet of the target device, the container is used to hold the material flowing into the container from the material outlet in the target device, the detection device may be located within a predetermined area of ​​the target device centered on the material outlet, and the metasurface element may be inclined at a predetermined angle to the plane where the container placement area is located.

[0108] In addition, in some embodiments of the present disclosure, the detection results may include: a first detection result, a second detection result, and a third detection result, wherein the first detection result indicates that there is no obstacle in the target path, and there is a container in the area to be detected, and the material in the container is not full, and the target path is the path of the material flowing from the material outlet of the target device into the container, the second detection result indicates that there is no obstacle in the target path, and there is no container in the area to be detected, or, the second detection result indicates that there is no obstacle in the target path, and there is a container in the area to be detected, and the material in the container is full, and the third detection result indicates that there is an obstacle in the target path.

[0109] According to the detection result, the behavior of the target device can be controlled. In some embodiments of the present disclosure, in response to determining that the detection result is the second detection result, the detection control module 702 can send a first instruction to the control center in the target device to instruct the control center to stop the outflow of the material, and in response to determining that the detection result is the third detection result, send a second instruction to the control center in the target device to instruct the control center to issue an alarm and / or control the outflow of the material.

[0110] The present disclosure also discloses a target device, which may include the target device behavior control device as described above.

[0111] Figure 8 FIG. 8 is a schematic diagram of the structure of the combined training device for the detection model and the optical element according to the present disclosure. Figure 8 As shown, it includes: a second acquisition module 801, a sample generation module 802 and a model training module 803.

[0112] The second acquisition module 801 is used to acquire an initial image, which is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal. The second light signal is a light signal obtained by modulating the acquired first light signal by the optical element. The first light signal is a light signal emitted by the light source and then reflected into the optical element after irradiating the area to be detected.

[0113] The sample generation module 802 is used to generate corresponding training samples for each initial image.

[0114] The model training module 803 is used to jointly train the detection model and the optical element using training samples.

[0115] In some embodiments of the present disclosure, the area to be inspected may include: a container placement area, the center of the container placement area may be opposite to the material outlet of the target device, and the method of obtaining the initial image may include: respectively obtaining initial images generated under M different imaging conditions, M is a positive integer greater than 1, wherein the difference between any two different imaging conditions includes one or all of the following: different light exposure conditions for the area to be inspected and different background environments where the target device is located.

[0116] For each acquired initial image, the sample generation module 802 may generate a corresponding training sample. In some embodiments of the present disclosure, the sample generation module 802 may perform the following processing for each initial image: obtain a label corresponding to the initial image, format the initial image, determine the formatted initial image as a sample image, and use the label corresponding to the sample image to form a training sample.

[0117] In some embodiments of the present disclosure, the label may include: a first detection result, a second detection result, and a third detection result, the first detection result indicating that there is no obstacle in the target path, and there is a container in the area to be detected, and the material in the container is not full, the target path is the path of the material flowing from the material outlet of the target device into the container, the second detection result indicating that there is no obstacle in the target path, and there is no container in the area to be detected, or, the second detection result indicating that there is no obstacle in the target path, and there is a container in the area to be detected, and the material in the container is full, and the third detection result indicating that there is an obstacle in the target path.

[0118] When jointly training the detection model and the optical element, in some embodiments of the present disclosure, the model training module 803 can obtain the detection result output by the detection model for a sample image, and can determine the loss value based on the detection result and the label corresponding to the sample image. Furthermore, in response to determining that the loss value satisfies a predetermined condition, the training process can be terminated; in response to determining that the loss value does not satisfy the predetermined condition, the optical element and / or the detection model can be updated according to the loss value.

[0119] In addition, in some embodiments of the present disclosure, the model training module 803 may first perform data enhancement processing on the sample image, and then input the sample image after the data enhancement processing into the detection model to obtain an output detection result.

[0120] In some embodiments of the present disclosure, the optical element may be a metasurface element. In addition, the metasurface element may include at least two partitions, each partition may include a structural unit set, each structural unit set may include at least two nanostructure units, and each structural unit set is used to generate a second optical signal.

[0121] Accordingly, in some embodiments of the present disclosure, the model training module 803 may adjust the nanostructure units in at least one structure unit set according to the loss value.

[0122] Furthermore, after jointly training the detection model and the optical element using the training samples, the model training module 803 can also perform quantization processing on the trained detection model.

[0123] The specific working processes of the above-mentioned device embodiments can refer to the relevant descriptions in the above-mentioned method embodiments and will not be repeated here.

[0124] The scheme disclosed in the present invention can be applied to the field of artificial intelligence, especially to the fields of deep learning, computer vision, image processing, and optical design. Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.

[0125] In addition, in the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0126] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0127] Fig. 9 A schematic block diagram of an electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0128] like Fig. 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0129] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0130] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI, Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP, Digital Signal Processing), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the methods described in the present disclosure may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the method described in the present disclosure in any other appropriate manner (for example, by means of firmware).

[0131] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0132] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0133] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, Electronically Programmable Read-Only Memory), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM, Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0134] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0135] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0136] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0137] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0138] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A target device behavior control method, comprising: Acquire a target image to be processed, wherein the target image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal, wherein the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, and the first light signal is a light signal emitted by a light source and then reflected by the light entering the optical element after irradiating the area to be detected; A detection result corresponding to the target image is generated by using a detection model, and the behavior of the target device is controlled according to the detection result, and the area to be detected is opposite to a material outlet in the target device.

2. The method according to claim 1, wherein: The optical element includes: a metasurface element.

3. The method according to claim 2, wherein: The metasurface element includes at least two partitions, each partition includes a structure unit set, each structure unit set includes at least two nanostructure units, and each structure unit set is used to generate one second optical signal.

4. The method according to claim 2 or 3, wherein: The area to be detected includes: a container placement area, the center of which is opposite to the material outlet of the target device; The metasurface element and the plane where the container placement area is located are inclined at a predetermined angle.

5. The method according to claim 1, 2 or 3, wherein: The test results include: a first test result, a second test result and a third test result; The first detection result indicates that there is no obstacle in the target path, there is a container in the area to be detected, and the container is not full of material, and the target path is the path of the material flowing from the material outlet into the container; The second detection result indicates that there is no obstacle in the target path and there is no container in the area to be detected, or the second detection result indicates that there is no obstacle in the target path and there is a container in the area to be detected, and the container is full of materials; The third detection result indicates that there is an obstacle in the target path.

6. The method according to claim 5, wherein: The controlling the behavior of the target device according to the detection result includes: In response to determining that the detection result is the second detection result, sending a first instruction to a control center in the target device to instruct the control center to control the material to stop flowing out; In response to determining that the detection result is the third detection result, a second instruction is sent to the control center in the target device to instruct the control center to issue an alarm and / or control the material to stop flowing out.

7. A joint training method for a detection model and an optical element, comprising: Acquire an initial image, wherein the initial image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal, wherein the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, and the first light signal is a light signal emitted by a light source and then reflected by the light entering the optical element after irradiating the area to be detected; For each initial image, generate corresponding training samples respectively; The detection model and the optical element are jointly trained using the training samples.

8. The method according to claim 7, wherein: The area to be detected includes: a container placement area, the center of which is opposite to the material outlet of the target device; The obtaining of the initial image comprises: respectively obtaining initial images generated under M different imaging conditions, where M is a positive integer greater than 1; The difference between any two different imaging conditions includes one or all of the following: different light irradiation conditions for the area to be detected, and different background environments where the target device is located.

9. The method according to claim 7, wherein: The generating corresponding training samples for each initial image includes: For each initial image, the following processing is performed respectively: Obtaining a label corresponding to the initial image, the label comprising: a first detection result, a second detection result, or a third detection result, the first detection result indicating that there is no obstacle in the target path, and there is a container in the area to be detected, and the material contained in the container is not full, the target path is a path for the material to flow from the material outlet into the container, the second detection result indicating that there is no obstacle in the target path, and there is no container in the area to be detected, or the second detection result indicating that there is no obstacle in the target path, and there is a container in the area to be detected, and the material contained in the container is full, and the third detection result indicating that there is an obstacle in the target path; Performing formatting processing on the initial image, and determining the formatted initial image as a sample image; The training samples are formed using the sample images and the labels.

10. The method according to any one of claims 7 to 9, wherein: The optical element includes: a metasurface element.

11. The method according to claim 10, wherein: The metasurface element includes at least two partitions, each partition includes a structure unit set, each structure unit set includes at least two nanostructure units, and each structure unit set is used to generate one second optical signal.

12. The method according to claim 11, wherein: The joint training of the detection model and the optical element using the training samples includes: updating the model parameters of the detection model according to the determined loss value, and / or adjusting the nanostructure unit in at least one structure unit set according to the loss value.

13. A target device behavior control device, comprising: A first acquisition module and a detection control module; The first acquisition module is used to acquire a target image to be processed, wherein the target image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by modulating the acquired first light signal by the optical element, and the first light signal is a light signal emitted by a light source and then reflected by the light entering the optical element after irradiating the area to be detected; The detection control module is used to generate a detection result corresponding to the target image using a detection model, and control the behavior of the target device according to the detection result. The area to be detected is opposite to a material outlet in the target device.

14. The device according to claim 13, wherein: The optical element includes: a metasurface element.

15. The device according to claim 14, wherein: The metasurface element includes at least two partitions, each partition includes a structure unit set, each structure unit set includes at least two nanostructure units, and each structure unit set is used to generate one second optical signal.

16. The device according to claim 13, wherein: The test results include: a first test result, a second test result and a third test result; The first detection result indicates that there is no obstacle in the target path, there is a container in the area to be detected, and the container is not full of material, and the target path is the path of the material flowing from the material outlet into the container; The second detection result indicates that there is no obstacle in the target path and there is no container in the area to be detected, or the second detection result indicates that there is no obstacle in the target path and there is a container in the area to be detected, and the container is full of materials; The third detection result indicates that there is an obstacle in the target path.

17. The device according to claim 16, wherein: In response to determining that the detection result is the second detection result, the detection control module sends a first instruction to the control center in the target device, used to instruct the control center to control the material to stop flowing out. In response to determining that the detection result is the third detection result, the detection control module sends a second instruction to the control center in the target device, used to instruct the control center to issue an alarm and / or control the material to stop flowing out.

18. A target device, comprising the target device behavior control device according to any one of claims 13 to 17.

19. A joint training device for a detection model and an optical element, comprising: A second acquisition module, a sample generation module, and a model training module; The second acquisition module is used to acquire an initial image, wherein the initial image is an image of the area to be detected generated by the photosensitive element according to the acquired second light signal, the second light signal is a light signal obtained by the optical element modulating the acquired first light signal, and the first light signal is a light signal emitted by the light source and then reflected by the light entering the optical element after irradiating the area to be detected; The sample generation module is used to generate corresponding training samples for each initial image; The model training module is used to jointly train the detection model and the optical element using the training samples.

20. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.

21. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 12.

22. A computer program product, comprising a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 12 is implemented.

Citation Information

Cited By

  • Home safety protection and joint training method and device

    CN120673045A

  • Collaborative optimization and task processing method and device and optical encryption system

    CN121000828A

  • Joint training and image processing method, device and system

    CN121330429A