Image data processing method, device and medium
By extracting frames with significant changes in the target pixel area from the display case image data and performing frame extraction based on the moving distance, the problems of large image data volume and low accuracy in existing technologies are solved, achieving efficient data processing and accurate item recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING GENKI FOREST BEVERAGE CO LTD
- Filing Date
- 2022-11-17
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies struggle to effectively reduce the amount of image data while ensuring accuracy when performing frame extraction on image data collected from display cases, resulting in high data processing costs and low efficiency.
By acquiring multiple images from the display case, the image after the cabinet door is unlocked is determined, and frames with significant changes in the target pixel area are extracted from these images. Frames are extracted based on the movement distance of the changed area to obtain images of the user's limb operation trajectory.
While reducing the amount of image data, it improved the accuracy of identifying user-operated items, reduced data processing costs, and improved the user experience.
Smart Images

Figure CN115731144B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of control technology, specifically to an image data processing method, device, and medium. Background Technology
[0002] In recent years, to facilitate user access to product information, merchants and businesses have often placed items in display cases, simultaneously storing and showcasing goods. When users need to remove or add items, they can open the display case themselves. In this scenario, merchants and businesses need to monitor the items in the display cases, identify users moving items in or out, and process payments accordingly. Summary of the Invention
[0003] This disclosure provides an image data processing method, apparatus, medium, and product.
[0004] In a first aspect, this disclosure provides an image data processing method.
[0005] Specifically, the image data processing method includes:
[0006] Acquire multiple images to be processed from the display case, arranged according to the acquisition time, and identify multiple cabinet door unlocking images whose acquisition time is after the cabinet door unlocking time from among the multiple images to be processed;
[0007] Multiple first frame images to be extracted are determined from multiple cabinet door unlocking images. The number of target pixels in the target image region of the first frame images to be extracted is greater than or equal to a first target pixel number threshold. The target pixel is a pixel whose pixel value difference with the pixel value of the corresponding position pixel in the cabinet door unlocking background image is greater than or equal to a pixel difference threshold. The target image region includes at least a part of the item entrance and exit of the display cabinet.
[0008] In each first frame to be extracted image, at least one change region is determined, and the movement distance of the change region corresponding to any two adjacent first frames to be extracted images is obtained. The change region is the region where the proportion of target pixels is greater than or equal to the target pixel proportion threshold and the number of target pixels is greater than or equal to the second target pixel number threshold.
[0009] Multiple first-frame images are extracted based on the movement distance of the changed region to obtain multiple first-frame images. The movement distance of the corresponding changed region in any two adjacent first-frame images belongs to the movement distance interval of the changed region.
[0010] In one embodiment of this disclosure, obtaining the movement distance of the changed region corresponding to any two adjacent first frame images to be extracted from a plurality of first frame images to be extracted includes:
[0011] Calculate the similarity between any pair of changed regions located in any two adjacent first frames to be extracted based on the pixel values of the corresponding pixels in the changed regions.
[0012] Obtain the positions of the changing regions in any two adjacent first frames to be extracted;
[0013] The movement distance of the changed region corresponding to any two adjacent first frames to be extracted is calculated based on the position of any pair of changed regions in any two adjacent first frames to be extracted, where the similarity is greater than or equal to the similarity threshold.
[0014] In one embodiment of this disclosure, obtaining the position corresponding to the changed region in any two adjacent first frames to be extracted includes:
[0015] The position corresponding to the changed region in any two adjacent first frames to be extracted is obtained based on the image position of the pixels in the changed region in any two adjacent first frames to be extracted.
[0016] Alternatively, the position corresponding to the changed region in any two adjacent first frames to be extracted can be obtained based on the depth information of the pixels in the changed region in any two adjacent first frames to be extracted.
[0017] In one embodiment of this disclosure, the movement distance of the changed regions corresponding to any two adjacent first frames to be extracted is calculated based on the positions of any pair of changed regions corresponding to any two adjacent first frames to be extracted, with a similarity greater than or equal to a similarity threshold, including:
[0018] In response to the fact that the change regions located in any two adjacent first frames to be extracted, and whose similarity is greater than or equal to the similarity threshold, include only one pair of change regions, the movement distance of the change regions corresponding to any two adjacent first frames to be extracted is calculated based on the position of the pair of change regions.
[0019] Alternatively, in response to the fact that the change regions located in any two adjacent first frames to be extracted, and whose similarity is greater than or equal to the similarity threshold, include multiple pairs of change regions, the movement distance of the change region corresponding to any two adjacent first frames to be extracted is calculated based on the position of the pair of change regions with the largest position change among the multiple pairs of change regions.
[0020] In one embodiment of this disclosure, the method further includes:
[0021] Multiple second frames to be extracted are identified from multiple cabinet door unlocking images. The acquisition time of the second frames to be extracted is earlier than the acquisition time of multiple first frames to be extracted, and / or the acquisition time of the second frames to be extracted is later than the acquisition time of multiple first frames to be extracted.
[0022] Multiple second-frame images are extracted based on the acquisition time to obtain multiple second-frame images. The time difference between the acquisition times of any two adjacent second-frame images belongs to the second acquisition time difference interval.
[0023] In one embodiment of this disclosure, the method further includes:
[0024] The number of frames obtained by subtracting the number of first-frame images from the number of first-frame images obtained;
[0025] In response to a frame extraction number being less than or equal to a preset frame extraction number, multiple first-extracted frame images are interpolated to obtain multiple interpolated frame images.
[0026] In one embodiment of this disclosure, the number of interpolated frames obtained by subtracting the number of first extracted frames from the number of interpolated frames is greater than or equal to the difference between the preset number of extracted frames and the number of extracted frames.
[0027] In one embodiment of this disclosure, the method further includes:
[0028] Obtain cabinet door opening indication information, which is used to indicate that the cabinet door of the display cabinet has been opened;
[0029] In response to the cabinet door opening indication, acquire multiple images to be processed.
[0030] Secondly, this disclosure provides an image data processing apparatus.
[0031] Specifically, the image data processing device includes:
[0032] The image data acquisition module is configured to acquire multiple images to be processed from the display case, arranged according to the acquisition time, and to identify multiple cabinet door unlocking images from the multiple images to be processed whose acquisition time is after the cabinet door unlocking time.
[0033] The frame extraction image determination module is configured to determine multiple first frame extraction images from multiple cabinet door unlocking images, wherein the number of target pixels in the target image region of the first frame extraction image is greater than or equal to a first target pixel quantity threshold, the target pixel is a pixel whose pixel value difference with the pixel value of the corresponding position pixel in the cabinet door unlocking background image is greater than or equal to a pixel difference threshold, and the target image region includes at least a portion of the item entrance / exit of the display cabinet.
[0034] The movement distance acquisition module is configured to determine at least one change region in each first frame to be extracted image, and acquire the movement distance of the change region corresponding to any two adjacent first frames to be extracted images in multiple first frames to be extracted images, wherein the change region is the region where the proportion of target pixels is greater than or equal to the target pixel proportion threshold, and the number of target pixels is greater than or equal to the second target pixel number threshold.
[0035] The frame extraction module is configured to extract frames from multiple first images to be extracted based on the movement distance of the changed region, so as to obtain multiple first extracted images. The movement distance of the corresponding changed region in any two adjacent first extracted images belongs to the movement distance interval of the changed region.
[0036] Thirdly, this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method as described in any embodiment of the first aspect.
[0037] Fourthly, this disclosure provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the method described in any embodiment of the first aspect.
[0038] Fifthly, this disclosure provides a computer program product including computer instructions that, when executed by a processor, implement the method as described in any embodiment of the first aspect.
[0039] The technical solutions provided in this disclosure may have the following beneficial effects:
[0040] In the technical solution provided in this disclosure, multiple images to be processed are acquired from the display case, and multiple cabinet door unlocking images acquired after the cabinet door unlocking time are determined from the multiple images to be processed. Multiple first frame-extracting images are then determined from the multiple cabinet door unlocking images. Since the number of target pixels in the target image region of the first frame-extracting image is greater than or equal to a first target pixel quantity threshold, and a target pixel is a pixel whose pixel value difference with the pixel value of the corresponding position in the cabinet door unlocking background image is greater than or equal to a pixel difference threshold, and the target image region includes at least a portion of the item entrance / exit of the display case, the first frame-extracting images are highly likely to include a user's moving limb that may be operating on items in the item storage area of the display case. At least one change region is determined in each first frame-extracting image, and the movement distance of the change region corresponding to any two adjacent first frame-extracting images is obtained. The change region is the target image. The region where the proportion of pixels is greater than or equal to the target pixel proportion threshold and the number of target pixels is greater than or equal to the second target pixel number threshold can be understood as the region where the user's limbs are in motion. Multiple first-frame images are then extracted based on the movement distance of the changed region to obtain multiple first-frame images. Since the movement distance of the corresponding changed region in any two adjacent first-frame images falls within the range of the changed region movement distance, it is possible to minimize the amount of data in the multiple first-frame images while ensuring that the multiple first-frame images obtained from the extraction also provide a relatively coherent movement trajectory of the user's limbs that may be manipulating items in the display case's storage area. This reduces the amount of image data to be processed and helps improve the accuracy of determining which items are moved out or into the display case by the user's limbs based on the multiple first-frame images, thereby reducing data processing costs and improving the user experience.
[0041] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0042] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken in conjunction with the accompanying drawings. In the drawings:
[0043] Figure 1 A schematic structural block diagram of a display cabinet according to an embodiment of the present disclosure is shown;
[0044] Figure 2 A schematic structural block diagram of a motherboard according to an embodiment of the present disclosure is shown;
[0045] Figure 3 A schematic structural block diagram of a control panel according to an embodiment of the present disclosure is shown;
[0046] Figure 4 A schematic structural block diagram of a power management module according to an embodiment of the present disclosure is shown.
[0047] Figure 5 A flowchart illustrating an image data processing method according to an embodiment of the present disclosure is shown;
[0048] Figure 6 A schematic structural diagram of a display case according to an embodiment of the present disclosure is shown;
[0049] Figure 7 A schematic top view of a display case according to one embodiment of the present disclosure is shown;
[0050] Figure 8 A schematic structural diagram of a display case according to an embodiment of the present disclosure is shown;
[0051] Figure 9 A schematic structural block diagram of an image data processing apparatus according to an embodiment of the present disclosure is shown;
[0052] Figure 10 A schematic structural block diagram of an electronic device according to an embodiment of the present disclosure is shown;
[0053] Figure 11 This is a schematic diagram of the structure of a computer system suitable for implementing an image data processing method according to an embodiment of the present disclosure. Detailed Implementation
[0054] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily implement them. Furthermore, for clarity, portions unrelated to the description of the exemplary embodiments have been omitted from the drawings.
[0055] In this disclosure, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of features, figures, steps, behaviors, components, parts or combinations thereof disclosed in this specification, and do not preclude the possibility of the presence or addition of one or more other features, figures, steps, behaviors, components, parts or combinations thereof.
[0056] It should also be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0057] As mentioned above, with the development of technology and the improvement of people's living standards, merchants or enterprises no longer simply place goods on shelves. Instead, to facilitate users' understanding of product information, they can place items in display cases, thus simultaneously storing and displaying the goods. When users need to remove items from the display case or put items into it, they can open the display case themselves and perform the corresponding operations.
[0058] In recent years, the number of display cases put into operation has gradually increased. During the use of display cases, merchants or enterprises generally need to monitor the items in the display cases to identify items that are moved out or into the display cases, and settle accounts based on the above information.
[0059] In one embodiment, when an item is moved out of or into a user's display case, image data collected by the display case can be acquired and uploaded to a server or cloud. The server or cloud can then perform image recognition on the image data and determine the item being moved out of or into the display case based on the image recognition results.
[0060] However, in this approach, the image data collected by the display cases is often quite large. Therefore, uploading this data to the server or cloud consumes significant bandwidth and is slow, affecting the efficiency of identifying items being moved in or out of the display case. Therefore, it is typically necessary to perform frame-by-frame processing on the image data collected by the display cases to obtain smaller, processed image data before uploading it to the server or cloud.
[0061] The inventors of this application discovered that when extracting frames from image data collected from a display case, if a large number of images are extracted, the data volume of the extracted image data is not significantly different from the total data volume of the collected image data. Extracting the frame data still consumes a significant amount of bandwidth, resulting in slow upload speeds. Furthermore, processing the extracted image data still requires substantial processing resources, thus increasing data processing costs and reducing the efficiency of settlement based on the data processing results. Conversely, if a small number of images are extracted from the collected images, the image recognition results obtained from the extracted image data often have low accuracy. This reduces the accuracy of determining whether items have been moved out or into the display case based on the image recognition results, compromising the reliability of settlement based on these results.
[0062] Therefore, how to minimize the amount of image data obtained from frame extraction while ensuring a high accuracy rate in identifying items being moved out or into display cases based on image data obtained from frame extraction is an increasingly urgent problem that needs to be solved.
[0063] In view of the above-mentioned defects, in one embodiment of this disclosure, an image data processing method is proposed. This method acquires multiple images to be processed from a display case, and identifies multiple unlocked images of the display case door acquired after the door unlocking time from among the multiple images to be processed. From these unlocked images, multiple first frame-extraction images are identified. Since the number of target pixels in the target image region of the first frame-extraction images is greater than or equal to a first target pixel quantity threshold, and a target pixel is a pixel whose pixel value difference with the corresponding pixel value in the unlocked background image is greater than or equal to a pixel difference threshold, and the target image region includes at least a portion of the display case's item entrance / exit area, the first frame-extraction images are highly likely to include a user's moving limbs that may be manipulating items in the display case's storage area. At least one change region is identified in each first frame-extraction image, and the movement distance of the change region corresponding to any two adjacent first frame-extraction images is obtained. The changing region is defined as the area where the proportion of target pixels is greater than or equal to a target pixel proportion threshold, and the number of target pixels is greater than or equal to a second target pixel number threshold. Therefore, the changing region can be understood as the area where the user's limbs are in motion. Multiple first-frame images are extracted based on the movement distance of the changing region to obtain multiple first-frame images. Since the movement distance of the corresponding changing region in any two adjacent first-frame images falls within the changing region movement distance range, it is possible to minimize the amount of data in the multiple first-frame images while ensuring that the multiple first-frame images obtained from the extraction also provide a relatively coherent movement trajectory of the user's limbs that may be manipulating items in the display case's storage area. This reduces the amount of image data to be processed and helps improve the accuracy of determining which items are moved out or into the display case by the user's limbs based on the multiple first-frame images, thereby reducing data processing costs and improving the user experience.
[0064] The image data processing method provided in this application embodiment can be applied to display cabinets, which can have a temperature control function. The temperature control function can be a cooling function, such as a refrigerated display cabinet, a frozen display cabinet, a refrigerator, a wine cabinet, a cosmetic preservation cabinet, etc.; the temperature control function can also be a heating function, such as a warm cabinet, a heated display cabinet, a hot beverage cabinet, etc. This application embodiment does not limit the specific type of display cabinet.
[0065] For example, Figure 1 A schematic structural block diagram of a display cabinet according to an embodiment of the present disclosure is shown, such as... Figure 1As shown, the display case 100 may include a compressor 11, a condenser 12, a throttling element 13, and an evaporator 14. The compressor 11, condenser 12, throttling element 13, and evaporator 14 are connected by pipes filled with refrigerant to form a closed pipeline, which constitutes a refrigeration system or heating system capable of circulating refrigerant.
[0066] The compressor refers to a driven fluid machine used to upgrade low-pressure refrigerant to high-pressure refrigerant. The compressor can draw in low-temperature, low-pressure gaseous refrigerant, and after the refrigerant is compressed by the piston driven by the motor, it discharges high-temperature, high-pressure gaseous refrigerant to provide power for the refrigeration cycle. The compressor can include reciprocating compressors, screw compressors, rotary compressors, scroll compressors and centrifugal compressors, etc. The specific type of compressor is not limited in the embodiments of this application.
[0067] A condenser is a heat exchanger used to exchange heat between the refrigerant inside the condenser and the air outside, thereby releasing heat. Specifically, a condenser may include a long pipe for containing the refrigerant, typically made of a metal with high thermal conductivity such as copper, and often coiled into a spiral shape. Furthermore, to improve the heat exchange efficiency of the condenser, heat sinks with excellent thermal conductivity can be installed on the pipes to increase the heat dissipation area, thereby accelerating the heat exchange rate and improving efficiency. Additionally, installing a fan or blower matched to the condenser can increase the airflow around it, further accelerating heat exchange and improving efficiency.
[0068] A throttling element is used to reduce the pressure of room-temperature, high-pressure liquid refrigerant, transforming it into a low-temperature, low-pressure gaseous refrigerant. This element can also be called a throttling device or a regulating valve, and may include expansion valves, capillary tubes, etc. Furthermore, the throttling element controls the flow rate of the refrigerant passing through it, preventing excessive or insufficient flow. If the flow rate is too high, the refrigerant exiting the element will still contain liquid refrigerant, which can cause liquid slugging and damage to the compressor. Conversely, if the flow rate is too low, insufficient refrigerant will enter the compressor, reducing its efficiency.
[0069] An evaporator is a heat exchanger used to exchange heat between the refrigerant inside the evaporator and the air outside the condenser, thereby achieving heat absorption. Specifically, an evaporator may include a long pipe for containing the refrigerant. This pipe is typically made of a metal with high thermal conductivity, such as copper, and is often coiled into a spiral shape. Furthermore, to improve the heat exchange efficiency of the condenser, heat sinks with excellent thermal conductivity can be installed on the pipe to increase the heat dissipation area, thereby accelerating the heat exchange rate and improving efficiency. Additionally, a fan or blower matched to the evaporator can be used to increase the airflow around the evaporator, further accelerating heat exchange and improving efficiency.
[0070] Refrigerant, also known as coolant or refrigerant fluid, refers to the medium through which energy conversion is completed in a refrigeration or heating system. Refrigerants are typically substances that readily undergo reversible phase changes (e.g., absorbing heat to become a gas and releasing heat to become a liquid). Through these reversible phase changes, refrigerants can transfer heat. Specifically, a gaseous refrigerant releases heat and becomes a liquid when pressurized, and absorbs heat when the high-pressure liquid is depressurized and becomes a gas. Refrigerants can include ammonia, air, water, brine, and Freon (also known as chlorofluorocarbons or fluorocarbons), among which Freon can include chlorofluoromethane, dichlorofluoromethane, trifluoromethane, tetrafluoroethane, and dichlorofluoroethane.
[0071] When the display case is equipped with a refrigeration function, low-temperature, low-pressure gaseous refrigerant flows from the evaporator into the compressor. The compressor compresses the low-temperature, low-pressure gaseous refrigerant, causing the high-temperature, high-pressure gaseous refrigerant to flow into the condenser. The high-temperature, high-pressure gaseous refrigerant exchanges heat with the outside air through the condenser, cooling it into a normal-temperature, high-pressure liquid refrigerant. This liquid refrigerant then flows into a throttling element, which restricts the flow, causing the refrigerant exiting the element to become a low-temperature, low-pressure liquid refrigerant. This low-temperature, low-pressure liquid refrigerant flows into the evaporator, where it exchanges heat with the outside air, evaporating and vaporizing into a low-temperature, low-pressure gaseous refrigerant to absorb heat. In this system, outside air can be introduced into the storage area of the display case through the evaporator, and outside air can be introduced into the outside of the display case through the condenser, thereby transferring the heat in the storage area of the display case to the outside of the display case and cooling the storage area.
[0072] When the display case is equipped with a heating function, low-temperature, low-pressure vaporous refrigerant flows from the condenser into the compressor. The compressor compresses the low-temperature, low-pressure vaporous refrigerant, causing the high-temperature, high-pressure vaporous refrigerant to flow into the evaporator. The high-temperature, high-pressure vaporous refrigerant exchanges heat with the outside air through the evaporator, cooling it into a normal-temperature, high-pressure liquid refrigerant. This liquid refrigerant then flows into a throttling element, which throttles the flow, causing the refrigerant exiting the element to become a low-temperature, low-pressure liquid refrigerant. This low-temperature, low-pressure liquid refrigerant flows into the condenser, where it exchanges heat with the outside air, evaporating and vaporizing into a low-temperature, low-pressure gaseous refrigerant to absorb heat. The system allows outside air to be introduced into the storage area of the display case, and outside air to be introduced into the outside of the display case, thereby transferring heat from outside the display case to the storage area to heat the storage area.
[0073] In one embodiment of this application, the display cabinet includes a cabinet body and a cabinet door, wherein a control board and a power management module may be installed in the cabinet body, and a main board may be installed in the cabinet door.
[0074] In one embodiment of this application, Figure 2 A schematic structural block diagram of a motherboard according to an embodiment of the present disclosure is shown, such as... Figure 2 As shown, the motherboard 200 includes a processor 201, random access memory 202, flash memory 203, wireless LAN Bluetooth module 204, gyroscope 205, pressure sensor 206, microphone 207, speaker 208, camera 209, and cellular communication module 210.
[0075] A processor may include one or more processing units, such as an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a memory, a video codec, a digital signal processor, a baseband processor, and / or a neural network processor. The different processing units may be independent devices or integrated into one or more processors.
[0076] The image signal processor (Image Signal Processor) processes data fed back from the camera. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then transmitted to the Image Signal Processor for processing, transforming it into a visible image. The Image Signal Processor can also perform algorithmic optimizations on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the Image Signal Processor can be integrated into the camera itself.
[0077] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, DSPs can be used to perform Fourier transforms on frequency energy.
[0078] Video codecs are used to compress or decompress digital video. A display case can support one or more video codecs. This allows the display case to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0079] Neural network computing processors, by drawing inspiration from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, can rapidly process input information and continuously learn on their own. These processors can enable applications such as intelligent cognition in display cases, including image recognition, facial recognition, speech recognition, and text understanding.
[0080] In some embodiments, the processor may include one or more interfaces. Interfaces may include integrated circuit interfaces, integrated circuit built-in audio interfaces, pulse code modulation interfaces, universal asynchronous transceiver interfaces, mobile industry processor interfaces, universal input / output interfaces, user identity module interfaces, and / or universal serial bus interfaces, etc.
[0081] Random access memory 202 can be used to store computer executable program code, which includes instructions and data. Processor 201 executes various functional applications and data processing of the display case by running the instructions stored in random access memory 202. Random access memory 202 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of the display case (such as audio data, image data, etc.).
[0082] Flash memory 203 can be used to expand the storage capacity of the display case. Flash memory 203 can communicate with processor 201 via the flash memory interface to achieve data storage functionality. For example, it can store music, video, and other files in the flash memory.
[0083] A minimal system can be constructed using processor 201, random access memory 202, and flash memory 203 to provide the system operating environment.
[0084] The Bluetooth LAN module 204 can provide wireless communication solutions for display cases, including wireless LAN, Bluetooth, GPS, FM, NNHF, and infrared technologies. The Bluetooth LAN module 204 can be one or more devices integrating at least one communication processing module. The Bluetooth LAN module 204 receives electromagnetic waves via an antenna, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to the processor 201. The Bluetooth LAN module 204 can also receive signals to be transmitted from the processor 201, perform frequency modulation and amplification, and then convert them into electromagnetic waves for radiation via the antenna. In one embodiment of this application, the Bluetooth LAN module can communicate with a user's terminal.
[0085] The cellular communication module 210 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use in display cases. The cellular communication module 210 may include at least one filter, switch, power amplifier, low-noise amplifier, etc. The cellular communication module 210 can receive electromagnetic waves via an antenna, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The cellular communication module 210 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the cellular communication module 210 may be housed in the processor 201. In some embodiments, at least some functional modules of the cellular communication module 210 and at least some modules of the processor 201 may be housed in the same device. In one embodiment of this application, the cellular communication module 210 can communicate with a cloud server of an image data processing service provider.
[0086] Through the Bluetooth wireless LAN module 204 and the cellular communication module 210, the display case can communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSMA), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband CDMA, Time Division CDMA, and Long Term Evolution (LTE).
[0087] The gyroscope 205 can be used to determine the real-time attitude of the display case doors.
[0088] The pressure sensor 206 is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor 206 can be disposed on the display screen. There are many types of pressure sensors 206, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive materials. When a force is applied to the pressure sensor 206, the capacitance between the electrodes changes, and the pressure intensity is determined based on the change in capacitance. When a touch operation is applied to the display screen, the touch operation intensity is detected by the pressure sensor 206, and the touch position can also be calculated based on the detection signal from the pressure sensor 206. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the beverage selection app icon, a command to view specific beverage information is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the beverage selection app icon, a command to purchase a beverage is executed.
[0089] Microphone 207, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 207, inputting the sound signal into microphone 207. A display case can be equipped with at least one microphone 207. In some embodiments, the display case can be equipped with two microphones 207, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, the display case can be equipped with three, four, or more microphones 207, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions. In one embodiment of this application, the microphone 207 can be used to collect the sound of the display case during operation.
[0090] Speaker 208, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The display case can play music or announcements via speaker 208.
[0091] Camera 209 is used to capture images, including still images and moving images (i.e., video). An object is projected onto a photosensitive element by an optical image generated through a lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an image signal processor (ISS) for conversion into a digital image signal. The ISS outputs the digital image signal to a digital signal processor (DSP) for further processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, the display case may include one or more cameras 209. In one embodiment of this application, the camera 209 may have a self-heating function to ensure that its lens does not fog up.
[0092] In one embodiment of this application, Figure 3 A schematic structural block diagram of a control panel according to an embodiment of the present disclosure is shown, such as... Figure 3 As shown, the control board 300 includes a power input interface 301, a power output interface 302, a metering chip 303, a microcontroller chip 304, a real-time clock chip, a light switch interface 305, a temperature control switch interface 306, an evaporator fan interface 307, a compressor interface 308, a condenser fan interface 309, a temperature sensor interface 310, a communication interface 311, and a power interface 312.
[0093] The metering chip 303, also known as the power sensor, acquires voltage, current, real-time power, and average power data. A real-time clock chip maintains the time for the microcontroller chip 304. The light switch interface 305 receives control signals from the display case's light switch. The temperature control switch interface 306 receives control signals from the display case's temperature control switch. The evaporator fan interface 307 sends control signals to the evaporator fan to control its operation. The compressor interface 308 sends control signals to the compressor to control its operation. The condenser fan interface 309 sends control signals to the condenser fan to control its operation. The temperature sensor interface 310 receives temperature data from one or more temperature sensors to determine the temperature at one or more locations within the display case.
[0094] In one embodiment of this application, Figure 4 A schematic structural block diagram of a power management module according to an embodiment of the present disclosure is shown, such as... Figure 4As shown, the power management module 400 includes an AC-to-DC conversion module 401, a charging management module 402, and a battery 403. The power management module 400 supplies power to the motherboard and control board and manages the charging and discharging of the battery. The power management module 400 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 400 may also be located within the processor.
[0095] In one embodiment of this application, the display case further includes a display screen. The display case implements its display function through a graphics processor, a display screen, and an application processor. The graphics processor is a microprocessor for image processing, connected to the display screen and the application processor. The graphics processor is used to perform mathematical and geometric calculations and for graphics rendering. The processor may include one or more graphics processors that execute program instructions to generate or modify display information.
[0096] The display screen is used to display still images, videos, etc. The display screen includes a display panel. The display panel can be a liquid crystal display, an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode, a quantum dot light-emitting diode, etc. In some embodiments, the display case may include one or more display screens.
[0097] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the display cabinet. In other embodiments of this application, the display cabinet may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0098] Figure 5 A flowchart illustrating an image data processing method according to an embodiment of the present disclosure is shown, such as... Figure 5 As shown, the image data processing method includes the following steps S101-S104:
[0099] In step S101, multiple images to be processed are acquired from the display case and arranged according to the acquisition time, and multiple cabinet door unlocking images whose acquisition time is after the cabinet door unlocking time are identified from the multiple images to be processed.
[0100] In one embodiment of this disclosure, acquiring multiple images to be processed captured by a display case can be achieved by receiving multiple images to be processed sent by an image acquisition device on the display case, other devices, or other systems, or by reading multiple images to be processed pre-stored in the display case. These multiple images to be processed can be captured by an image acquisition device in the display case, such as a camera, or by other image acquisition devices or systems corresponding to the display case, such as security cameras. The images in the multiple images to be processed can be understood as including all or part of the images of the display case, or as including all or part of the images of the items in the display case.
[0101] In one embodiment of this disclosure, the multiple images to be processed arranged according to the acquisition time can be understood as images acquired by the display case according to the corresponding sampling frequency, or as samples obtained from the video acquired by the display case according to the corresponding sampling frequency.
[0102] In step S102, multiple first frame images to be extracted are determined from multiple cabinet door unlocking images;
[0103] Wherein, the number of target pixels in the target image region of the first frame to be extracted is greater than or equal to the first target pixel quantity threshold, the target pixel is a pixel whose pixel value difference between the pixel value and the pixel value of the corresponding position pixel in the cabinet door unlocking background image is greater than or equal to the pixel difference threshold, and the target image region includes at least a part of the item entrance and exit of the display cabinet.
[0104] It should be noted that the cabinet door unlocking background image can be pre-stored or obtained from other devices or systems; alternatively, the cabinet door unlocking background image can be the Nth cabinet door unlocking image among multiple cabinet door unlocking images. When the multiple cabinet door unlocking images are acquired at an image acquisition rate of 24 frames per second, the value of N can be greater than 0 and less than or equal to 24. The cabinet door unlocking background image can be understood as an image in the target image area that does not include the user's limbs or where the probability of this situation occurring is low.
[0105] In one embodiment of this disclosure, the item entry / exit of the display case can be understood as an entrance / exit connecting the item storage area of the display case with an entrance / exit outside the display case. Through the item entry / exit, items outside the display case can be moved into the item storage area of the display case, or items can be moved out of the item storage area of the display case. Based on multiple images to be processed that include at least a portion of the item entry / exit of the display case, items moved out of or into the display case can be determined for settlement. The settlement result can be used to operate on the corresponding account, or it can be used to determine at least one of the quantity, type, and location of items stored in the display case after the items are moved out or moved in.
[0106] In one embodiment of this disclosure, the target image region can be a region with at least one of the following predefined dimensions and shape: for example, the target image region can be a rectangle or a circle, or the target image region can be a polygon with the number of sides belonging to a preset range, the side length belonging to a preset range, and the included angle between the sides belonging to a preset range.
[0107] It should be noted that, within the display case, since the speed at which the cabinet door opens or closes may be close to the speed of a user's limb movement, and considering that the first frame to be extracted can be understood as an image captured when the user's limb moves near the item entrance / exit of the display case, to avoid reducing the accuracy of the determined first frame to be extracted due to door movement, the image acquisition device for capturing the unlocked door image can be reasonably configured, or the range of the target image area can be reasonably set to ensure that the cabinet door is not included in the target image area. For example, Figure 6 A schematic structural diagram of a display cabinet according to an embodiment of the present disclosure is shown. Figure 7 A schematic top view of a display case according to one embodiment of the present disclosure is shown, such as... Figure 6 as well as Figure 7 As shown, the display case includes a cabinet body 501, a cabinet door 502, and a camera 503. The cabinet body 501 includes a display area 511 and an item access point 521. The display area 511 is connected to the outside of the cabinet body 501 through the item access point 521. The cabinet door 502 is rotatably or slidably connected to the cabinet body 501 and is used to open or close the item access point 521. The camera 503 is located on the top of the cabinet body and is used to capture images of the item access point 521. Based on multiple images captured by the camera 503, multiple images to be processed can be obtained. It should be noted that the display case may also include multiple cameras that capture images of items entering and exiting from different directions. For example... Figure 8 A schematic structural diagram of a display case according to an embodiment of the present disclosure is shown, such as... Figure 8As shown, the display case includes a first camera 516, a second camera 526, a third camera 536, a fourth camera 546, and a fifth camera 556. The first camera 516 captures images of the goods entrance / exit 521 from a first direction 5161, the second camera 526 captures images of the goods entrance / exit 521 from a second direction 5162, the third camera 536 captures images of the goods entrance / exit 521 from a third direction 5163, the fourth camera 546 captures images of the goods entrance / exit 521 from a fourth direction 5164, and the fifth camera 556 captures images of the goods entrance / exit 521 from a fifth direction 5165. The first direction 5161, the second direction 5162, the third direction 5163, the fourth direction 5164, and the fifth direction 5165 are all different. Multiple images to be processed can be obtained from multiple images captured by at least one of the first camera 516, the second camera 526, the third camera 536, the fourth camera 546, and the fifth camera 556.
[0108] In one embodiment of this disclosure, the target image region of the first frame to be extracted can be understood as recognizing the first frame to be extracted based on a pre-acquired target image region algorithm, and determining the target image region of the first frame to be extracted based on the target image region recognition result; or it can be understood as acquiring a pre-trained target image region model, taking the first frame to be extracted as input, inputting the target image region model to obtain the target image region recognition result, and determining the target image region of the first frame to be extracted based on the target image region recognition result.
[0109] In one embodiment of this disclosure, the pixel values and pixel coordinates of one or more pixels in the target image area of the cabinet door unlocking background image can be obtained, and the pixel values of corresponding pixels (i.e., pixels in the same position) in multiple cabinet door unlocking images can be obtained based on the pixel coordinates. The target pixel in each cabinet door unlocking image can be determined according to the obtained pixel values.
[0110] In one embodiment of this disclosure, the pixel difference threshold can be understood as being stored in advance in a display case, or it can be understood as being obtained from other devices or systems.
[0111] In step S103, at least one change region is determined in each first frame to be extracted image, and the movement distance of the change region corresponding to any two adjacent first frames to be extracted images is obtained.
[0112] The variable region is the region where the proportion of target pixels is greater than or equal to the target pixel proportion threshold, and the number of target pixels is greater than or equal to the second target pixel number threshold.
[0113] In one embodiment of this disclosure, the proportion of target pixels is greater than or equal to the target pixel proportion threshold, which can be understood as the ratio of the number of target pixels in the change region to the total number of pixels in the change region being greater than or equal to the target pixel proportion threshold.
[0114] In one embodiment of this disclosure, the target pixel ratio threshold and the second target pixel quantity threshold can be pre-stored in a display case or obtained from other devices or systems.
[0115] In one embodiment of this disclosure, the movement distance of the changed region corresponding to any two adjacent first frame images to be extracted among multiple first frame images to be extracted can be understood as calculating the position of the changed region in each of the two adjacent first frame images to be extracted, and calculating the movement distance of the changed region corresponding to any two adjacent first frame images to be extracted among multiple first frame images to be extracted based on the position of the changed region corresponding to any two adjacent first frame images to be extracted. The position of the changed region can be understood as the position of the center of the changed region, or it can be understood as the position obtained by averaging the positions of all pixels in the changed region. The position of a pixel can be understood as the two-dimensional image position of the image in which the pixel is located, and in this case, the movement distance of the changed region can be understood as a two-dimensional distance in units of pixels; the position of a pixel can also be understood as the three-dimensional spatial position of the object surface position corresponding to the pixel obtained based on the depth information corresponding to the pixel, and in this case, the movement distance of the changed region can be understood as a three-dimensional distance in three-dimensional space. The depth information corresponding to the pixel can be obtained based on a depth image acquisition device (e.g., a binocular vision camera).
[0116] In step S104, multiple first frames to be extracted are extracted based on the moving distance of the changed area to obtain multiple first extracted frames.
[0117] Among them, the movement distance of the corresponding changed region in any two adjacent first frame images in multiple first frame images belongs to the movement distance interval of the changed region.
[0118] In one embodiment of this disclosure, the corresponding change regions in two adjacent first frame images can be understood as determining two change regions in the two adjacent first frame images whose position changes fall within a preset range of position changes. Alternatively, the pixel values and positions of pixels in the change regions of the two adjacent first frame images can be substituted into a pre-acquired corresponding change region algorithm for calculation, and the two change regions can be determined as corresponding change regions in the two adjacent first frame images based on the calculation results. The corresponding change regions in the two adjacent first frame images can be understood as regions in the two adjacent first frame images used to display the same limb of the same user.
[0119] In one embodiment of this disclosure, frame extraction is performed on multiple first images to be extracted based on the change region movement distance. This can be understood as follows: One first image to be extracted is selected as the first extracted image from the multiple first images to be extracted. Another first image to be extracted is selected, with its sampling time located after the sampling time of the first extracted image. The first change region movement distance corresponding to this other first image to be extracted and the first extracted image is obtained. When the first change region movement distance falls within a change region movement distance interval, the other first image to be extracted is also selected as the first extracted image. Then, another first image to be extracted is selected, with its sampling time located after the sampling time of the next first image to be extracted. The second change region movement distance corresponding to this second first image to be extracted and the next first image to be extracted is obtained. When the second change region movement distance falls within a change region movement distance interval, the other first image to be extracted is also selected as the first extracted image. This process is repeated until multiple first extracted images are obtained.
[0120] Alternatively, the frame extraction can be performed on multiple first frame images based on the movement distance of the changed region. This can also be understood as taking multiple first frame images and the movement distance of the changed region corresponding to two adjacent first frame images as input, inputting them into a pre-acquired first frame extraction model, and obtaining multiple first frame images output by the first frame extraction model.
[0121] In the technical solution provided in this disclosure, multiple images to be processed are acquired from the display case, and multiple cabinet door unlocking images acquired after the cabinet door unlocking time are determined from the multiple images to be processed. Multiple first frame-extracting images are then determined from the multiple cabinet door unlocking images. Since the number of target pixels in the target image region of the first frame-extracting image is greater than or equal to a first target pixel quantity threshold, and a target pixel is a pixel whose pixel value difference with the pixel value of the corresponding position in the cabinet door unlocking background image is greater than or equal to a pixel difference threshold, and the target image region includes at least a portion of the item entrance / exit of the display case, the first frame-extracting images are highly likely to include a user's moving limb that may be operating on items in the item storage area of the display case. At least one change region is determined in each first frame-extracting image, and the movement distance of the change region corresponding to any two adjacent first frame-extracting images is obtained. The change region is the target image. The region where the proportion of pixels is greater than or equal to the target pixel proportion threshold and the number of target pixels is greater than or equal to the second target pixel number threshold can be understood as the region where the user's limbs are in motion. Multiple first-frame images are then extracted based on the movement distance of the changed region to obtain multiple first-frame images. Since the movement distance of the corresponding changed region in any two adjacent first-frame images falls within the range of the changed region movement distance, it is possible to minimize the amount of data in the multiple first-frame images while ensuring that the multiple first-frame images obtained from the extraction also provide a relatively coherent movement trajectory of the user's limbs that may be manipulating items in the display case's storage area. This reduces the amount of image data to be processed and helps improve the accuracy of determining which items are moved out or into the display case by the user's limbs based on the multiple first-frame images, thereby reducing data processing costs and improving the user experience.
[0122] In one embodiment of this disclosure, obtaining the movement distance of the changed region corresponding to any two adjacent first frame images to be extracted from a plurality of first frame images to be extracted includes:
[0123] Calculate the similarity between any pair of changed regions located in any two adjacent first frames to be extracted based on the pixel values of the corresponding pixels in the changed regions.
[0124] Obtain the positions of the changing regions in any two adjacent first frames to be extracted;
[0125] The movement distance of the changed region corresponding to any two adjacent first frames to be extracted is calculated based on the position of any pair of changed regions in any two adjacent first frames to be extracted, where the similarity is greater than or equal to the similarity threshold.
[0126] In one embodiment of this disclosure, the similarity between any pair of changing regions located in any two adjacent first frames to be extracted is calculated based on the pixel values of corresponding pixels in the changing regions. This can be understood as substituting the pixel values of corresponding pixels in the changing regions of the two adjacent first frames into a pre-acquired similarity algorithm to obtain the similarity; or, it can be understood as using the pixel values of corresponding pixels in the changing regions of the two adjacent first frames as input to a pre-acquired phase velocity model to obtain the similarity output by the similarity model. A pair of changing regions with a similarity greater than or equal to a similarity threshold can be understood as a pair of changing regions used to display the same object (e.g., a user's limb) in the two images.
[0127] In the technical solution provided in this disclosure, the similarity between any pair of changing regions located in any two adjacent first frame images to be extracted is calculated based on the pixel values of the corresponding pixels in the changing regions. The positions of the changing regions in any two adjacent first frame images to be extracted are obtained. The movement distance of the changing regions corresponding to any two adjacent first frame images to be extracted is calculated based on the similarity being greater than or equal to the similarity threshold and the positions corresponding to any pair of changing regions located in any two adjacent first frame images to be extracted. This ensures that the movement distance of the changing regions can accurately reflect the distance of the user's limb movement in the two adjacent first frame images to be extracted. This helps to ensure that the multiple first frame images obtained based on the subsequent frame extraction step can also obtain a relatively coherent movement trajectory of the user's limbs that may be operating on the items in the storage area of the display case.
[0128] In one embodiment of this disclosure, obtaining the position corresponding to the changed region in any two adjacent first frames to be extracted includes:
[0129] The position corresponding to the changed region in any two adjacent first frames to be extracted is obtained based on the image position of the pixels in the changed region in any two adjacent first frames to be extracted.
[0130] Alternatively, the position corresponding to the changed region in any two adjacent first frames to be extracted can be obtained based on the depth information of the pixels in the changed region in any two adjacent first frames to be extracted.
[0131] In one embodiment of this disclosure, the image position of a pixel can be understood as the position of the pixel in the corresponding first frame to be extracted image. The depth information of a pixel can be understood as indicating the distance between the object surface displayed by the pixel and the image acquisition device when the first frame to be extracted image is acquired. It should be noted that, in order to obtain the depth information of a pixel, multiple images to be processed can be acquired by an image acquisition device with depth acquisition function (such as a binocular vision image acquisition device).
[0132] In the technical solution provided in this disclosure, the position corresponding to the changed region in any two adjacent first frames to be extracted is obtained by obtaining the image position of the pixels in the changed region in any two adjacent first frames to be extracted, or by obtaining the position corresponding to the changed region in any two adjacent first frames to be extracted by obtaining the depth information of the pixels in the changed region in any two adjacent first frames to be extracted, which can improve the accuracy of the obtained position corresponding to the changed region.
[0133] In one embodiment of this disclosure, the movement distance of the changed regions corresponding to any two adjacent first frames to be extracted is calculated based on the positions of any pair of changed regions corresponding to any two adjacent first frames to be extracted, with a similarity greater than or equal to a similarity threshold, including:
[0134] In response to the fact that the change regions located in any two adjacent first frames to be extracted, and whose similarity is greater than or equal to the similarity threshold, include only one pair of change regions, the movement distance of the change regions corresponding to any two adjacent first frames to be extracted is calculated based on the position of the pair of change regions.
[0135] Alternatively, in response to the fact that the change regions located in any two adjacent first frames to be extracted, and whose similarity is greater than or equal to the similarity threshold, include multiple pairs of change regions, the movement distance of the change region corresponding to any two adjacent first frames to be extracted is calculated based on the position of the pair of change regions with the largest position change among the multiple pairs of change regions.
[0136] In the technical solution provided in this disclosure, considering that in reality, there may be only one user moving an item out of or into a display case at the same time, or there may be multiple users moving items out of or into a display case at the same time, the first image to be extracted may include only one limb of one user, or it may include two limbs of one user, or multiple limbs of multiple users. In order to ensure that the number of limbs in the first image to be extracted does not affect the tracking of limbs or objects picked up by limbs based on multiple first images to be extracted, while minimizing the number of first images obtained after extraction, the method is to respond that the change regions located in any two adjacent first images to be extracted, and whose similarity is greater than or equal to the similarity threshold, include only one pair of change regions. The method calculates the movement distance of the changed region corresponding to any two adjacent first frame images to be extracted based on the position of a pair of changed regions. Alternatively, in response to the fact that the changed regions located in any two adjacent first frame images to be extracted and whose similarity is greater than or equal to the similarity threshold include multiple pairs of changed regions, the method calculates the movement distance of the changed region corresponding to any two adjacent first frame images to be extracted based on the position of the pair of changed regions with the largest position change among the multiple pairs of changed regions. This ensures that regardless of how many hands are present in the first frame images to be extracted, the multiple first frame images obtained based on the extraction can stably track the fastest moving limb in the multiple first frame images to be extracted. This helps to improve the accuracy of determining items moved out or into the display case based on multiple first frame images and improves the user experience.
[0137] In one embodiment of this disclosure, the method further includes:
[0138] Multiple second frames to be extracted are identified from multiple cabinet door unlocking images. The acquisition time of the second frames to be extracted is earlier than the acquisition time of multiple first frames to be extracted, and / or the acquisition time of the second frames to be extracted is later than the acquisition time of multiple first frames to be extracted.
[0139] Multiple second-frame images are extracted based on the acquisition time to obtain multiple second-frame images. The time difference between the acquisition times of any two adjacent second-frame images belongs to the second acquisition time difference interval.
[0140] In the technical solution provided in this disclosure, considering that in multiple cabinet door unlocking images, the probability of a user's limbs appearing at the item entrance / exit of the display cabinet is low in images acquired earlier than the acquisition time of multiple first-to-be-extracted frame images, and in images acquired later than the acquisition time of multiple first-to-be-extracted frame images, the probability of an item being moved out or into the display cabinet by the user through their limbs in the second-to-be-extracted frame image is also low. Therefore, there is no need to track the user's limbs in the second-to-be-extracted frame image. By determining multiple second-to-be-extracted frame images from multiple cabinet door unlocking images, and extracting frames from multiple second-to-be-extracted frame images according to the acquisition time, multiple second-to-be-extracted frame images can be obtained. This can minimize the amount of image data that needs to be processed while recording the entire image information after the user unlocks the cabinet door.
[0141] In one embodiment of this disclosure, the method further includes:
[0142] The number of frames obtained by subtracting the number of first-frame images from the number of first-frame images obtained;
[0143] In response to a frame extraction number being less than or equal to a preset frame extraction number, multiple first-extracted frame images are interpolated to obtain multiple interpolated frame images.
[0144] In one embodiment of this disclosure, frame interpolation of multiple first-frame images can be performed by substituting the multiple first-frame images into a pre-acquired algorithm to obtain multiple interpolated frames. For example, the interpolation algorithm can be used to calculate the interpolation image based on any two adjacent first-frame images to obtain the interpolation image to be interpolated, and then the interpolation image to be interpolated is inserted into the two adjacent first-frame images to obtain multiple interpolated frames including the multiple first-frame images. Alternatively, a pre-trained interpolation model can be obtained, and the multiple first-frame images can be used as input to obtain the multiple interpolated frames output by the interpolation model.
[0145] In the technical solution provided in this disclosure, considering that when the user's hand moves too fast, the number of frames may be small, and it may not be possible to reliably track the user's hand trajectory based on multiple first frame images. Therefore, the number of frames is obtained by subtracting the number of first frame images from the number of first frame images. In response to the number of frames being less than or equal to the preset number of frames, multiple first frame images are interpolated to obtain multiple interpolated frames. This ensures that the number of interpolated frames is large, and the user's hand trajectory can be reliably tracked based on multiple interpolated frames. This helps to improve the accuracy of determining items being moved out or into the display case based on multiple interpolated frames, thus improving the user experience.
[0146] In one embodiment of this disclosure, the number of interpolated frames obtained by subtracting the number of first extracted frames from the number of interpolated frames is greater than or equal to the difference between the preset number of extracted frames and the number of extracted frames.
[0147] In the technical solution provided in this disclosure, by limiting the number of supplementary frames obtained by subtracting the number of multiple first extracted frames from the number of supplementary frames, the number of supplementary frames is greater than or equal to the difference between the preset number of extracted frames and the number of extracted frames. This allows for a more convenient setting of the number of images used to determine the items being moved out or into the display case, thus improving the user experience.
[0148] In one embodiment of this disclosure, the method further includes:
[0149] Obtain cabinet door opening indication information, which is used to indicate that the cabinet door of the display cabinet has been opened;
[0150] In response to a cabinet door opening indication, multiple images to be processed are acquired.
[0151] In one embodiment of this disclosure, obtaining cabinet door opening instruction information can be understood as receiving cabinet door opening instruction information sent by the cabinet door locking device of the display cabinet, or it can be understood as receiving cabinet door opening instruction information sent by other devices or systems.
[0152] In the technical solution provided in this disclosure, by acquiring cabinet door opening indication information and responding to the cabinet door opening indication information, multiple images to be processed are acquired. This can minimize the amount of data in the acquired images to be processed without affecting the accuracy of determining whether items are moved out or into the display case, thereby reducing processing costs.
[0153] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein.
[0154] Figure 9 A schematic structural block diagram of an image data processing apparatus according to an embodiment of the present disclosure is shown. This image data processing apparatus can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 9 As shown, the image data processing device includes:
[0155] The image data acquisition module is configured to acquire multiple images to be processed from the display case, arranged according to the acquisition time, and to identify multiple cabinet door unlocking images from the multiple images to be processed whose acquisition time is after the cabinet door unlocking time.
[0156] The frame extraction image determination module is configured to determine multiple first frame extraction images from multiple cabinet door unlocking images, wherein the number of target pixels in the target image region of the first frame extraction image is greater than or equal to a first target pixel quantity threshold, the target pixel is a pixel whose pixel value difference with the pixel value of the corresponding position pixel in the cabinet door unlocking background image is greater than or equal to a pixel difference threshold, and the target image region includes at least a portion of the item entrance / exit of the display cabinet.
[0157] The movement distance acquisition module is configured to determine at least one change region in each first frame to be extracted image, and acquire the movement distance of the change region corresponding to any two adjacent first frames to be extracted images in multiple first frames to be extracted images, wherein the change region is the region where the proportion of target pixels is greater than or equal to the target pixel proportion threshold, and the number of target pixels is greater than or equal to the second target pixel number threshold.
[0158] The frame extraction module is configured to extract frames from multiple first images to be extracted based on the movement distance of the changed region, so as to obtain multiple first extracted images. The movement distance of the corresponding changed region in any two adjacent first extracted images belongs to the movement distance interval of the changed region.
[0159] The above technical solution involves acquiring multiple images to be processed from the display case, identifying multiple unlocked images of the cabinet door acquired after the door unlocking time, and then identifying multiple first frame-extracting images from these unlocked images. Since the number of target pixels in the target image region of the first frame-extracting image is greater than or equal to a first target pixel quantity threshold (a target pixel is defined as a pixel whose pixel value difference from the corresponding pixel value in the cabinet door unlocking background image is greater than or equal to a pixel difference threshold), and the target image region includes at least a portion of the display case's item access area, the first frame-extracting images are highly likely to include a user's moving limbs that may be interacting with items in the display case's storage area. At least one change region is identified in each first frame-extracting image, and the movement distance of the change region corresponding to any two adjacent first frame-extracting images is obtained. The change region is defined as the ratio of the target pixel value to the change region's movement distance. For example, the area where the target pixel ratio is greater than or equal to the target pixel ratio threshold and the number of target pixels is greater than or equal to the second target pixel number threshold can be understood as the area where the user's limbs are in motion. Multiple first-frame images are then extracted based on the movement distance of the changed area to obtain multiple first-frame images. Since the movement distance of the corresponding changed area in any two adjacent first-frame images falls within the range of the changed area movement distance, it is possible to minimize the amount of data in the multiple first-frame images while ensuring that the multiple first-frame images obtained from the extraction also provide a relatively coherent movement trajectory of the user's limbs that may be manipulating items in the display case's storage area. This reduces the amount of image data that needs to be processed and helps improve the accuracy of determining which items are moved out or into the display case by the user's limbs based on the multiple first-frame images, thereby reducing data processing costs and improving the user experience.
[0160] This disclosure also discloses an electronic device. Figure 10 A schematic structural block diagram of an electronic device according to an embodiment of the present disclosure is shown, such as... Figure 10 As shown, the electronic device includes a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the above method steps.
[0161] Figure 11 This is a schematic diagram of the structure of a computer system suitable for implementing an image data processing method according to an embodiment of the present disclosure. For example... Figure 11As shown, the computer system includes a processing unit that can execute various processes described above based on a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer system. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0162] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard disks; and communication sections including network interface cards such as LAN cards and modems. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as needed. The processing unit can be implemented as a CPU, GPU, TPU, FPGA, NPU, etc.
[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0164] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The described units or modules can also be located in a processor, and the names of these units or modules do not necessarily constitute a limitation on the unit or module itself.
[0165] In another aspect, this disclosure also provides a computer-readable storage medium, which may be a computer-readable storage medium included in the apparatus described in the above embodiments; or it may be a standalone computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs that are used by one or more processors to perform the methods described in this disclosure.
[0166] In addition, this disclosure also provides a computer program product storing a computer program that, when executed by a processor, enables the processor to at least implement the methods provided in the foregoing embodiments.
[0167] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
Claims
1. An image data processing method, characterized in that, The method includes: Acquire multiple images to be processed from the display case, arranged according to the acquisition time, and identify multiple cabinet door unlocking images from the multiple images to be processed whose acquisition time is after the cabinet door unlocking time; Multiple first frame images to be extracted are determined from the multiple cabinet door unlocking images, wherein the number of target pixels in the target image region of the first frame images to be extracted is greater than or equal to a first target pixel number threshold, the target pixel is a pixel whose pixel value difference with the pixel value of the corresponding position pixel in the cabinet door unlocking background image is greater than or equal to a pixel difference threshold, and the target image region includes at least a part of the item entrance and exit of the display cabinet. In each first frame to be extracted image, at least one change region is determined, and the movement distance of the change region corresponding to any two adjacent first frames to be extracted images is obtained. The change region is the region where the proportion of the target pixel is greater than or equal to the target pixel proportion threshold, and the number of the target pixels is greater than or equal to the second target pixel number threshold. The multiple first frames to be extracted are extracted according to the moving distance of the changed region to obtain multiple first frame images. The moving distance of the corresponding changed region in any two adjacent first frame images in the multiple first frame images belongs to the moving distance interval of the changed region. The step of obtaining the movement distance of the changed region corresponding to any two adjacent first frames to be extracted from the plurality of first frames to be extracted includes: The similarity between any pair of changing regions located in any two adjacent first frames to be extracted is calculated based on the pixel values of the corresponding pixels in the changing regions. Obtain the position of the changing region in any two adjacent first frames to be extracted; The movement distance of the changed regions corresponding to any two adjacent first frames to be extracted is calculated based on the similarity being greater than or equal to the similarity threshold and located in the positions of any pair of changed regions in any two adjacent first frames to be extracted.
2. The image data processing method according to claim 1, characterized in that, The step of obtaining the position corresponding to the changed region in any two adjacent first frames to be extracted includes: The position corresponding to the changed region in any two adjacent first frame images to be extracted is obtained based on the image position of the pixels in the changed region in any two adjacent first frame images to be extracted. Alternatively, the position corresponding to the changed region in any two adjacent first frames to be extracted can be obtained based on the depth information of the pixels in the changed region in any two adjacent first frames to be extracted.
3. The image data processing method according to claim 1, characterized in that, The step of calculating the movement distance of the changed regions corresponding to any two adjacent first frames to be extracted, based on the similarity being greater than or equal to a similarity threshold and located at the positions of any pair of changed regions in any two adjacent first frames to be extracted, includes: In response to the fact that the change regions located in any two adjacent first frames to be extracted and whose similarity is greater than or equal to the similarity threshold include only one pair of change regions, the movement distance of the change regions corresponding to the pair of change regions is calculated based on the position of the change regions. Alternatively, in response to the fact that the change regions located in any two adjacent first frames to be extracted and whose similarity is greater than or equal to the similarity threshold include multiple pairs of change regions, the movement distance of the change region corresponding to the pair of change regions with the largest position change in the multiple pairs of change regions is calculated based on the position of the pair of change regions corresponding to the pair of adjacent first frames to be extracted.
4. The image data processing method according to claim 1, characterized in that, The method further includes: Multiple second frames to be extracted are determined from the multiple cabinet door unlocking images. The acquisition time of the second frames to be extracted is earlier than the acquisition time of the multiple first frames to be extracted, and / or the acquisition time of the second frames to be extracted is later than the acquisition time of the multiple first frames to be extracted. The plurality of second frames to be extracted are extracted according to the acquisition time to obtain a plurality of second extracted frames. The time difference between the acquisition times of any two adjacent second extracted frames in the plurality of second extracted frames belongs to the second acquisition time difference interval.
5. The image data processing method according to any one of claims 1-4, characterized in that, The method further includes: The number of the multiple first frame-sampling images is obtained as the frame-sampling number; In response to the number of frames being less than or equal to a preset number of frames, the multiple first frame-sampling images are interpolated to obtain multiple interpolated images.
6. The image data processing method according to claim 5, characterized in that, The number of interpolated frames obtained by subtracting the number of first extracted frames from the number of interpolated frames is greater than or equal to the difference between the preset number of extracted frames and the number of extracted frames.
7. The image data processing method according to any one of claims 1-4, characterized in that, The method further includes: Obtain cabinet door opening indication information, which is used to indicate that the cabinet door of the display cabinet is opened; In response to the cabinet door opening instruction, the multiple images to be processed are acquired.
8. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method of any one of claims 1-6.
9. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-6.
Citation Information
Patent Citations
Automatic vending method
CN108492451A
Image data processing method and image data processing apparatus
CN110415295A
Image detection method and device, computer equipment and storage medium
CN111523347A