Image data processing method, device, medium and product

By recognizing hand movements in display cases and performing frame extraction based on the distance the hands move, the problem of large image data volume and low accuracy in existing technologies is solved, achieving efficient and low-cost item monitoring.

CN115830708BActive Publication Date: 2026-05-05BEIJING GENKI FOREST BEVERAGE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING GENKI FOREST BEVERAGE CO LTD
Filing Date
2022-11-17
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies struggle to minimize image data volume while ensuring accuracy when performing frame extraction on image data collected from display cases, resulting in high data processing costs and low efficiency.

Method used

By acquiring image data collected from the display case, the movement of the hand in the target image area is identified. Frames are extracted based on the distance of hand movement to obtain multiple framed target images, ensuring the continuity of the hand movement trajectory and reducing the amount of data.

Benefits of technology

While reducing the amount of image data, it improves the accuracy of determining whether an item is moved in or out, reduces data processing costs, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115830708B_ABST
    Figure CN115830708B_ABST
Patent Text Reader

Abstract

This disclosure provides an image data processing method, apparatus, medium, and product. The method includes: acquiring image data to be processed collected from a display case; identifying hands in target image regions of each of the multiple images to be processed; determining multiple target images from the multiple images to be processed based on the hand recognition results, and acquiring the hand movement distances corresponding to any two adjacent target images; and performing frame extraction on the multiple target images based on the hand movement distances to obtain multiple frame-extracted target images. This technical solution can ensure a relatively coherent hand movement trajectory in the image data obtained from frame extraction while minimizing the number of frame-extracted target images. This reduces the amount of image data to be processed and helps improve the accuracy of determining items being moved out or into the display case, thereby reducing data processing costs and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of control technology, specifically to an image data processing method, device, medium, and product. Background Technology

[0002] In recent years, to facilitate user access to product information, merchants and businesses have often placed items in display cases, simultaneously storing and showcasing goods. When users need to remove or add items, they can open the display case themselves. In this scenario, merchants and businesses need to monitor the items in the display cases, identify users moving items in or out, and process payments accordingly. Summary of the Invention

[0003] This disclosure provides an image data processing method, apparatus, medium, and product.

[0004] In a first aspect, this disclosure provides an image data processing method.

[0005] Specifically, the image data processing method includes:

[0006] Acquire image data to be processed collected from the display case. The image data to be processed includes multiple images to be processed arranged according to the collection time.

[0007] The system identifies hands in the target image region of each of multiple images to be processed, including the item entrances and exits of display cases.

[0008] Based on the hand recognition results, multiple target images are determined from multiple images to be processed, and the hand movement distance corresponding to any two adjacent target images is obtained. The target image is the image to be processed that includes the hand in the corresponding target image region.

[0009] Multiple target images are extracted by frame-by-frame based on the hand movement distance to obtain multiple extracted target images. The hand movement distances corresponding to any two adjacent extracted target images belong to the hand movement distance interval.

[0010] In one embodiment of this disclosure, the method further includes:

[0011] The number of frames extracted is obtained by subtracting the number of multiple target images from the number of target images to be extracted.

[0012] In response to a frame extraction number being less than or equal to a preset frame extraction number, multiple extracted target images are interpolated to obtain multiple interpolated target images.

[0013] In one embodiment of this disclosure, the number of interpolated frames obtained by subtracting the number of extracted frames from the number of interpolated target images is greater than or equal to the difference between the preset number of extracted frames and the number of extracted frames.

[0014] In one embodiment of this disclosure, obtaining the hand movement distance corresponding to any two adjacent target images includes:

[0015] In response to the fact that each of the two adjacent target images contains only one hand in its respective target image region, the distance that a hand moves in the two adjacent target images is determined as the hand movement distance corresponding to the two adjacent target images.

[0016] Alternatively, in response to the fact that each of the target image regions of any two adjacent target images includes multiple hands, the target hand among the multiple hands that moves the longest distance in any two adjacent target images is determined, and the distance that the target hand moves in any two adjacent target images is determined as the hand movement distance corresponding to any two adjacent target images.

[0017] In one embodiment of this disclosure, the method further includes:

[0018] Obtain cabinet door opening indication information, which is used to indicate that the cabinet door of the display cabinet has been opened;

[0019] In response to the cabinet door opening indication, acquire the image data to be processed.

[0020] In one embodiment of this disclosure, the method further includes:

[0021] Based on the hand recognition results, multiple first frame images to be extracted are determined from multiple images to be processed. The first frame images to be extracted are images to be processed where the corresponding target image area does not include the hand and the hand is located outside the item storage area of ​​the display case.

[0022] Based on the acquisition time, multiple first frames to be extracted are extracted to obtain multiple first extracted frames. The time difference between the acquisition times of any two adjacent first extracted frames belongs to the first acquisition time difference interval.

[0023] In one embodiment of this disclosure, the method further includes:

[0024] Based on the hand recognition results, multiple second frames to be extracted are determined from multiple images to be processed. The second frames to be extracted are images to be processed where the corresponding target image area does not include the hand, and the hand is located in the item storage area of ​​the display case.

[0025] Multiple second-frame images are extracted based on the acquisition time to obtain multiple second-frame images. The time difference between the acquisition times of any two adjacent second-frame images belongs to the second acquisition time difference interval.

[0026] Secondly, this disclosure provides an image data processing apparatus.

[0027] Specifically, the image data processing device includes:

[0028] The image data acquisition module is configured to acquire image data to be processed collected by the display case. The image data to be processed includes multiple images to be processed arranged according to the acquisition time.

[0029] The hand recognition module is configured to recognize hands in the target image region of each of multiple images to be processed, including the item entrances and exits of the display case.

[0030] The movement distance acquisition module is configured to determine multiple target images from multiple images to be processed based on the hand recognition results, and to acquire the hand movement distance corresponding to any two adjacent target images among the multiple target images. The target image is the image to be processed that includes the hand in the corresponding target image region.

[0031] The frame extraction module is configured to extract frames from multiple target images based on the hand movement distance to obtain multiple frame-extracted target images. The hand movement distances corresponding to any two adjacent frame-extracted target images in the multiple frame-extracted target images belong to the hand movement distance interval.

[0032] Thirdly, this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method as described in any embodiment of the first aspect.

[0033] Fourthly, this disclosure provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the method described in any embodiment of the first aspect.

[0034] Fifthly, this disclosure provides a computer program product including computer instructions that, when executed by a processor, implement the method as described in any embodiment of the first aspect.

[0035] The technical solutions provided in this disclosure may have the following beneficial effects:

[0036] In the technical solution provided in this disclosure, by acquiring the image data to be processed collected by the display case, the hand in the target image region of each of the multiple images to be processed is identified, multiple target images are determined in the multiple images to be processed based on the hand recognition results, and the hand movement distance corresponding to any two adjacent target images in the multiple target images is obtained. The multiple target images are then framed based on the hand movement distance to obtain multiple framed target images. The target image is the image to be processed that includes the hand in the target image region, which includes the item entrance / exit of the display case. Therefore, by focusing on the user's hand movement trajectory in the target image, it can be determined whether the user's hand has moved an item out of or into the display case through the item entrance / exit. By limiting the hand movement distance between any two adjacent target images obtained from multiple frame-by-frame extraction to a specific range, it is possible to ensure a relatively coherent hand movement trajectory based on the image data obtained from frame extraction while minimizing the amount of data in multiple target images. This reduces the amount of image data to be processed and improves the accuracy of determining whether an item is moved out or into the display case based on multiple target images, thereby reducing data processing costs and improving the user experience.

[0037] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0038] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments, taken in conjunction with the accompanying drawings. In the drawings:

[0039] Figure 1 A schematic structural block diagram of a display cabinet according to an embodiment of the present disclosure is shown;

[0040] Figure 2 A schematic structural block diagram of a motherboard according to an embodiment of the present disclosure is shown;

[0041] Figure 3 A schematic structural block diagram of a control panel according to an embodiment of the present disclosure is shown;

[0042] Figure 4 A schematic structural block diagram of a power management module according to an embodiment of the present disclosure is shown.

[0043] Figure 5 A flowchart illustrating an image data processing method according to an embodiment of the present disclosure is shown;

[0044] Figure 6 A schematic structural diagram of a display case according to an embodiment of the present disclosure is shown;

[0045] Figure 7 A schematic top view of a display case according to one embodiment of the present disclosure is shown;

[0046] Figure 8 A schematic structural diagram of a display case according to an embodiment of the present disclosure is shown;

[0047] Figure 9 A schematic structural block diagram of an image data processing apparatus according to an embodiment of the present disclosure is shown;

[0048] Figure 10 A schematic structural block diagram of an electronic device according to an embodiment of the present disclosure is shown;

[0049] Figure 11 This is a schematic diagram of the structure of a computer system suitable for implementing an image data processing method according to an embodiment of the present disclosure. Detailed Implementation

[0050] In the following, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings to enable those skilled in the art to readily implement them. Furthermore, for clarity, portions unrelated to the description of the exemplary embodiments have been omitted from the drawings.

[0051] In this disclosure, it should be understood that terms such as “comprising” or “having” are intended to indicate the presence of features, figures, steps, behaviors, components, parts or combinations thereof disclosed in this specification, and do not preclude the possibility of the presence or addition of one or more other features, figures, steps, behaviors, components, parts or combinations thereof.

[0052] It should also be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0053] As mentioned above, with the development of technology and the improvement of people's living standards, merchants or enterprises no longer simply place goods on shelves. Instead, to facilitate users' understanding of product information, they can place items in display cases, thus simultaneously storing and displaying the goods. When users need to remove items from the display case or put items into it, they can open the display case themselves and perform the corresponding operations.

[0054] In recent years, the number of display cases put into operation has gradually increased. During the use of display cases, merchants or enterprises generally need to monitor the items in the display cases to identify items that are moved out or into the display cases, and settle accounts based on the above information.

[0055] In one embodiment, when an item is moved out of or into a user's display case, image data collected by the display case can be acquired and uploaded to a server or cloud. The server or cloud can then perform image recognition on the image data and determine the item being moved out of or into the display case based on the image recognition results.

[0056] However, in this approach, the image data collected by the display cases is often quite large. Therefore, uploading this data to the server or cloud consumes significant bandwidth and is slow, affecting the efficiency of identifying items being moved in or out of the display case. Therefore, it is typically necessary to perform frame-by-frame processing on the image data collected by the display cases to obtain smaller, processed image data before uploading it to the server or cloud.

[0057] The inventors of this application discovered that when extracting frames from image data collected from a display case, if a large number of images are extracted, the data volume of the extracted image data is not significantly different from the total data volume of the collected image data. Extracting the frame data still consumes a significant amount of bandwidth, resulting in slow upload speeds. Furthermore, processing the extracted image data still requires substantial processing resources, thus increasing data processing costs and reducing the efficiency of settlement based on the data processing results. Conversely, if a small number of images are extracted from the collected images, the image recognition results obtained from the extracted image data often have low accuracy. This reduces the accuracy of determining whether items have been moved out or into the display case based on the image recognition results, compromising the reliability of settlement based on these results.

[0058] Therefore, how to minimize the amount of image data obtained from frame extraction while ensuring a high accuracy rate in identifying items being moved out or into display cases based on image data obtained from frame extraction is an increasingly urgent problem that needs to be solved.

[0059] In view of the above-mentioned defects, in one embodiment of this disclosure, an image data processing method is proposed. The method acquires image data to be processed collected by a display case, identifies the hand in the target image region of each of the multiple images to be processed, determines multiple target images in the multiple images to be processed based on the hand recognition results, obtains the hand movement distance corresponding to any two adjacent target images in the multiple target images, and performs frame extraction on the multiple target images based on the hand movement distance to obtain multiple frame-extracted target images. The target image is the image to be processed that includes the hand in the target image region, which includes the item entrance / exit of the display case. Therefore, by focusing on the user's hand movement trajectory in the target image, it can be determined whether the user's hand has moved an item out of or into the display case through the item entrance / exit. By limiting the hand movement distance between any two adjacent target images obtained from multiple frame-by-frame extraction to a specific range, it is possible to ensure a relatively coherent hand movement trajectory based on the image data obtained from frame extraction while minimizing the amount of data in multiple target images. This reduces the amount of image data to be processed and improves the accuracy of determining whether an item is moved out or into the display case based on multiple target images, thereby reducing data processing costs and improving the user experience.

[0060] The image data processing method provided in this application embodiment can be applied to display cabinets, which can have a temperature control function. The temperature control function can be a cooling function, such as a refrigerated display cabinet, a frozen display cabinet, a refrigerator, a wine cabinet, a cosmetic preservation cabinet, etc.; the temperature control function can also be a heating function, such as a warm cabinet, a heated display cabinet, a hot beverage cabinet, etc. This application embodiment does not limit the specific type of display cabinet.

[0061] For example, Figure 1 A schematic structural block diagram of a display cabinet according to an embodiment of the present disclosure is shown, such as... Figure 1 As shown, the display case 100 may include a compressor 11, a condenser 12, a throttling element 13, and an evaporator 14. The compressor 11, condenser 12, throttling element 13, and evaporator 14 are connected by pipes filled with refrigerant to form a closed pipeline, which constitutes a refrigeration system or heating system capable of circulating refrigerant.

[0062] The compressor refers to a driven fluid machine used to upgrade low-pressure refrigerant to high-pressure refrigerant. The compressor can draw in low-temperature, low-pressure gaseous refrigerant, and after the refrigerant is compressed by the piston driven by the motor, it discharges high-temperature, high-pressure gaseous refrigerant to provide power for the refrigeration cycle. The compressor can include reciprocating compressors, screw compressors, rotary compressors, scroll compressors and centrifugal compressors, etc. The specific type of compressor is not limited in the embodiments of this application.

[0063] A condenser is a heat exchanger used to exchange heat between the refrigerant inside the condenser and the air outside, thereby releasing heat. Specifically, a condenser may include a long pipe for containing the refrigerant, typically made of a metal with high thermal conductivity such as copper, and often coiled into a spiral shape. Furthermore, to improve the heat exchange efficiency of the condenser, heat sinks with excellent thermal conductivity can be installed on the pipes to increase the heat dissipation area, thereby accelerating the heat exchange rate and improving efficiency. Additionally, installing a fan or blower matched to the condenser can increase the airflow around it, further accelerating heat exchange and improving efficiency.

[0064] A throttling element is used to reduce the pressure of room-temperature, high-pressure liquid refrigerant, transforming it into a low-temperature, low-pressure gaseous refrigerant. This element can also be called a throttling device or a regulating valve, and may include expansion valves, capillary tubes, etc. Furthermore, the throttling element controls the flow rate of the refrigerant passing through it, preventing excessive or insufficient flow. If the flow rate is too high, the refrigerant exiting the element will still contain liquid refrigerant, which can cause liquid slugging and damage to the compressor. Conversely, if the flow rate is too low, insufficient refrigerant will enter the compressor, reducing its efficiency.

[0065] An evaporator is a heat exchanger used to exchange heat between the refrigerant inside the evaporator and the air outside the condenser, thereby achieving heat absorption. Specifically, an evaporator may include a long pipe for containing the refrigerant. This pipe is typically made of a metal with high thermal conductivity, such as copper, and is often coiled into a spiral shape. Furthermore, to improve the heat exchange efficiency of the condenser, heat sinks with excellent thermal conductivity can be installed on the pipe to increase the heat dissipation area, thereby accelerating the heat exchange rate and improving efficiency. Additionally, a fan or blower matched to the evaporator can be used to increase the airflow around the evaporator, further accelerating heat exchange and improving efficiency.

[0066] Refrigerant, also known as coolant or refrigerant fluid, refers to the medium through which energy conversion is completed in a refrigeration or heating system. Refrigerants are typically substances that readily undergo reversible phase changes (e.g., absorbing heat to become a gas and releasing heat to become a liquid). Through these reversible phase changes, refrigerants can transfer heat. Specifically, a gaseous refrigerant releases heat and becomes a liquid when pressurized, and absorbs heat when the high-pressure liquid is depressurized and becomes a gas. Refrigerants can include ammonia, air, water, brine, and Freon (also known as chlorofluorocarbons or fluorocarbons), among which Freon can include chlorofluoromethane, dichlorofluoromethane, trifluoromethane, tetrafluoroethane, and dichlorofluoroethane.

[0067] When the display case is equipped with a refrigeration function, low-temperature, low-pressure gaseous refrigerant flows from the evaporator into the compressor. The compressor compresses the low-temperature, low-pressure gaseous refrigerant, causing the high-temperature, high-pressure gaseous refrigerant to flow into the condenser. The high-temperature, high-pressure gaseous refrigerant exchanges heat with the outside air through the condenser, cooling it into a normal-temperature, high-pressure liquid refrigerant. This liquid refrigerant then flows into a throttling element, which restricts the flow, causing the refrigerant exiting the element to become a low-temperature, low-pressure liquid refrigerant. This low-temperature, low-pressure liquid refrigerant flows into the evaporator, where it exchanges heat with the outside air, evaporating and vaporizing into a low-temperature, low-pressure gaseous refrigerant to absorb heat. In this system, outside air can be introduced into the storage area of ​​the display case through the evaporator, and outside air can be introduced into the outside of the display case through the condenser, thereby transferring the heat in the storage area of ​​the display case to the outside of the display case and cooling the storage area.

[0068] When the display case is equipped with a heating function, low-temperature, low-pressure vaporous refrigerant flows from the condenser into the compressor. The compressor compresses the low-temperature, low-pressure vaporous refrigerant, causing the high-temperature, high-pressure vaporous refrigerant to flow into the evaporator. The high-temperature, high-pressure vaporous refrigerant exchanges heat with the outside air through the evaporator, cooling it into a normal-temperature, high-pressure liquid refrigerant. This liquid refrigerant then flows into a throttling element, which throttles the flow, causing the refrigerant exiting the element to become a low-temperature, low-pressure liquid refrigerant. This low-temperature, low-pressure liquid refrigerant flows into the condenser, where it exchanges heat with the outside air, evaporating and vaporizing into a low-temperature, low-pressure gaseous refrigerant to absorb heat. The system allows outside air to be introduced into the storage area of ​​the display case, and outside air to be introduced into the outside of the display case, thereby transferring heat from outside the display case to the storage area to heat the storage area.

[0069] In one embodiment of this application, the display cabinet includes a cabinet body and a cabinet door, wherein a control board and a power management module may be installed in the cabinet body, and a main board may be installed in the cabinet door.

[0070] In one embodiment of this application, Figure 2 A schematic structural block diagram of a motherboard according to an embodiment of the present disclosure is shown, such as... Figure 2 As shown, the motherboard 200 includes a processor 201, random access memory 202, flash memory 203, wireless LAN Bluetooth module 204, gyroscope 205, pressure sensor 206, microphone 207, speaker 208, camera 209, and cellular communication module 210.

[0071] A processor may include one or more processing units, such as an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a memory, a video codec, a digital signal processor, a baseband processor, and / or a neural network processor. The different processing units may be independent devices or integrated into one or more processors.

[0072] The image signal processor (Image Signal Processor) processes data fed back from the camera. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then transmitted to the Image Signal Processor for processing, transforming it into a visible image. The Image Signal Processor can also perform algorithmic optimizations on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the Image Signal Processor can be integrated into the camera itself.

[0073] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, DSPs can be used to perform Fourier transforms on frequency energy.

[0074] Video codecs are used to compress or decompress digital video. A display case can support one or more video codecs. This allows the display case to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0075] Neural network computing processors, by drawing inspiration from the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, can rapidly process input information and continuously learn on their own. These processors can enable applications such as intelligent cognition in display cases, including image recognition, facial recognition, speech recognition, and text understanding.

[0076] In some embodiments, the processor may include one or more interfaces. Interfaces may include integrated circuit interfaces, integrated circuit built-in audio interfaces, pulse code modulation interfaces, universal asynchronous transceiver interfaces, mobile industry processor interfaces, universal input / output interfaces, user identity module interfaces, and / or universal serial bus interfaces, etc.

[0077] Random access memory 202 can be used to store computer executable program code, which includes instructions and data. Processor 201 executes various functional applications and data processing of the display case by running the instructions stored in random access memory 202. Random access memory 202 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of the display case (such as audio data, image data, etc.).

[0078] Flash memory 203 can be used to expand the storage capacity of the display case. Flash memory 203 can communicate with processor 201 via the flash memory interface to achieve data storage functionality. For example, it can store music, video, and other files in the flash memory.

[0079] A minimal system can be constructed using processor 201, random access memory 202, and flash memory 203 to provide the system operating environment.

[0080] The Bluetooth LAN module 204 can provide wireless communication solutions for display cases, including wireless LAN, Bluetooth, GPS, FM, NNHF, and infrared technologies. The Bluetooth LAN module 204 can be one or more devices integrating at least one communication processing module. The Bluetooth LAN module 204 receives electromagnetic waves via an antenna, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to the processor 201. The Bluetooth LAN module 204 can also receive signals to be transmitted from the processor 201, perform frequency modulation and amplification, and then convert them into electromagnetic waves for radiation via the antenna. In one embodiment of this application, the Bluetooth LAN module can communicate with a user's terminal.

[0081] The cellular communication module 210 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use in display cases. The cellular communication module 210 may include at least one filter, switch, power amplifier, low-noise amplifier, etc. The cellular communication module 210 can receive electromagnetic waves via an antenna, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The cellular communication module 210 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via the antenna. In some embodiments, at least some functional modules of the cellular communication module 210 may be housed in the processor 201. In some embodiments, at least some functional modules of the cellular communication module 210 and at least some modules of the processor 201 may be housed in the same device. In one embodiment of this application, the cellular communication module 210 can communicate with a cloud server of an image data processing service provider.

[0082] Through the Bluetooth wireless LAN module 204 and the cellular communication module 210, the display case can communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSMA), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband CDMA, Time Division CDMA, and Long Term Evolution (LTE).

[0083] The gyroscope 205 can be used to determine the real-time attitude of the display case doors.

[0084] The pressure sensor 206 is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor 206 can be disposed on the display screen. There are many types of pressure sensors 206, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive materials. When a force is applied to the pressure sensor 206, the capacitance between the electrodes changes, and the pressure intensity is determined based on the change in capacitance. When a touch operation is applied to the display screen, the touch operation intensity is detected by the pressure sensor 206, and the touch position can also be calculated based on the detection signal from the pressure sensor 206. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the beverage selection app icon, a command to view specific beverage information is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the beverage selection app icon, a command to purchase a beverage is executed.

[0085] Microphone 207, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 207, inputting the sound signal into microphone 207. A display case can be equipped with at least one microphone 207. In some embodiments, the display case can be equipped with two microphones 207, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, the display case can be equipped with three, four, or more microphones 207, enabling sound signal collection, noise reduction, sound source identification, and directional recording, among other functions. In one embodiment of this application, the microphone 207 can be used to collect the sound of the display case during operation.

[0086] Speaker 208, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The display case can play music or announcements via speaker 208.

[0087] Camera 209 is used to capture images, including still images and moving images (i.e., video). An object is projected onto a photosensitive element by an optical image generated through a lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an image signal processor (ISS) for conversion into a digital image signal. The ISS outputs the digital image signal to a digital signal processor (DSP) for further processing. The DSP converts the digital image signal into image signals in standard formats such as RGB and YUV. In some embodiments, the display case may include one or more cameras 209. In one embodiment of this application, the camera 209 may have a self-heating function to ensure that its lens does not fog up.

[0088] In one embodiment of this application, Figure 3 A schematic structural block diagram of a control panel according to an embodiment of the present disclosure is shown, such as... Figure 3 As shown, the control board 300 includes a power input interface 301, a power output interface 302, a metering chip 303, a microcontroller chip 304, a real-time clock chip, a light switch interface 305, a temperature control switch interface 306, an evaporator fan interface 307, a compressor interface 308, a condenser fan interface 309, a temperature sensor interface 310, a communication interface 311, and a power interface 312.

[0089] The metering chip 303, also known as the power sensor, acquires voltage, current, real-time power, and average power data. A real-time clock chip maintains the time for the microcontroller chip 304. The light switch interface 305 receives control signals from the display case's light switch. The temperature control switch interface 306 receives control signals from the display case's temperature control switch. The evaporator fan interface 307 sends control signals to the evaporator fan to control its operation. The compressor interface 308 sends control signals to the compressor to control its operation. The condenser fan interface 309 sends control signals to the condenser fan to control its operation. The temperature sensor interface 310 receives temperature data from one or more temperature sensors to determine the temperature at one or more locations within the display case.

[0090] In one embodiment of this application, Figure 4 A schematic structural block diagram of a power management module according to an embodiment of the present disclosure is shown, such as... Figure 4 As shown, the power management module 400 includes an AC-to-DC conversion module 401, a charging management module 402, and a battery 403. The power management module 400 supplies power to the motherboard and control board and manages the charging and discharging of the battery. The power management module 400 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 400 may also be located within the processor.

[0091] In one embodiment of this application, the display case further includes a display screen. The display case implements its display function through a graphics processor, a display screen, and an application processor. The graphics processor is a microprocessor for image processing, connected to the display screen and the application processor. The graphics processor is used to perform mathematical and geometric calculations and for graphics rendering. The processor may include one or more graphics processors that execute program instructions to generate or modify display information.

[0092] The display screen is used to display still images, videos, etc. The display screen includes a display panel. The display panel can be a liquid crystal display, an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode, a quantum dot light-emitting diode, etc. In some embodiments, the display case may include one or more display screens.

[0093] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the display cabinet. In other embodiments of this application, the display cabinet may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0094] Figure 5 A flowchart illustrating an image data processing method according to an embodiment of the present disclosure is shown, such as... Figure 5 As shown, the image data processing method includes the following steps S101-S104:

[0095] In step S101, the image data to be processed is acquired from the display case.

[0096] The image data to be processed includes multiple images arranged according to the acquisition time.

[0097] In one embodiment of this disclosure, acquiring multiple images to be processed can be achieved by receiving multiple images sent by other devices or systems, or by reading multiple images to be processed pre-stored in a display case. These multiple images to be processed can be captured by an image acquisition device in the display case, such as a camera, or by other image acquisition devices or systems corresponding to the display case, such as security cameras. The images in the multiple images to be processed can be understood as including all or part of the images of the display case, or as including all or part of the images of the items in the display case.

[0098] In one embodiment of this disclosure, the multiple images to be processed arranged according to the acquisition time can be understood as images acquired by the display case according to the corresponding sampling frequency, or as samples obtained from the video acquired by the display case according to the corresponding sampling frequency.

[0099] In step S102, the hand is identified in the target image region of each of the multiple images to be processed.

[0100] The target image area includes the entrances and exits of the display cases.

[0101] In one embodiment of this disclosure, the item entry / exit of the display case can be understood as an entrance / exit connecting the item storage area of ​​the display case with an entrance / exit outside the display case. Through the item entry / exit, items outside the display case can be moved into the item storage area of ​​the display case, or items can be moved out of the item storage area of ​​the display case. Based on multiple images to be processed that include the item entry / exit of the display case, items moved out of or into the display case can be determined for settlement. The settlement result can be used to operate on the corresponding account, or it can be used to determine at least one of the quantity, type, and location of items stored in the display case after the items are moved out or moved in.

[0102] For example, Figure 6 A schematic structural diagram of a display cabinet according to an embodiment of the present disclosure is shown. Figure 7 A schematic top view of a display case according to one embodiment of the present disclosure is shown, such as... Figure 6 as well as Figure 7 As shown, the display case includes a cabinet body 501, a cabinet door 502, and a camera 503. The cabinet body 501 includes a display area 511 and an item access point 521. The display area 511 is connected to the outside of the cabinet body 501 through the item access point 521. The cabinet door 502 is rotatably or slidably connected to the cabinet body 501 and is used to open or close the item access point 521. The camera 503 is located on the top of the cabinet body and is used to capture images of the item access point 521. Based on multiple images captured by the camera 503, multiple images to be processed can be obtained. It should be noted that the display case may also include multiple cameras that capture images of items entering and exiting from different directions. For example... Figure 8 A schematic structural diagram of a display case according to an embodiment of the present disclosure is shown, such as... Figure 8 As shown, the display case includes a first camera 516, a second camera 526, a third camera 536, a fourth camera 546, and a fifth camera 556. The first camera 516 captures images of the goods entrance / exit 521 from a first direction 5161, the second camera 526 captures images of the goods entrance / exit 521 from a second direction 5162, the third camera 536 captures images of the goods entrance / exit 521 from a third direction 5163, the fourth camera 546 captures images of the goods entrance / exit 521 from a fourth direction 5164, and the fifth camera 556 captures images of the goods entrance / exit 521 from a fifth direction 5165. The first direction 5161, the second direction 5162, the third direction 5163, the fourth direction 5164, and the fifth direction 5165 are all different. Multiple images to be processed can be obtained from multiple images captured by at least one of the first camera 516, the second camera 526, the third camera 536, the fourth camera 546, and the fifth camera 556.

[0103] In one embodiment of this disclosure, the target image region of the image to be processed can be understood as recognizing the image to be processed based on a pre-acquired target image region algorithm, and determining the target image region of the image to be processed based on the target image region recognition result; or it can be understood as acquiring a pre-trained target image region model, taking the image to be processed as input, inputting the target image region model to obtain the target image region recognition result, and determining the target image region of the image to be processed based on the target image region recognition result.

[0104] In one embodiment of this disclosure, recognizing a hand in a target image region of an image to be processed can be understood as acquiring a pre-trained hand recognition model, and inputting the image to be processed or its target image region into the hand recognition model to obtain hand recognition information output by the model. Based on this hand recognition information, the hand in the target image region of the image to be processed is determined. The hand recognition model can be pre-stored in a display case or acquired from other devices or systems. The hand recognition model can be a neural network (NN) model, a convolutional neural network (CNN) model, or a long short-term memory (LSTM) model, etc. Recognizing a hand in the target image region of an image to be processed can also be understood as recognizing a hand in the target image region of the image to be processed based on an algorithm obtained from a pre-trained model.

[0105] In step S103, multiple target images are determined from multiple images to be processed based on the hand recognition results, and the hand movement distance corresponding to any two adjacent target images is obtained.

[0106] The target image is the image to be processed, which includes the corresponding target image region, specifically the hand.

[0107] In one embodiment of this disclosure, the hand movement distance can be understood as the actual distance or as the pixel distance.

[0108] In one embodiment of this disclosure, obtaining the hand movement distance corresponding to two adjacent target images can be understood as obtaining the position of the hand in the target image region of each of the two adjacent target images, and calculating the hand movement distance corresponding to these two adjacent target images based on the position of the hand in the respective target image region of the two adjacent target images. Specifically, obtaining the position of the hand in the target image region of the target image can be understood as obtaining the hand position based on the hand recognition result, or it can be understood as inputting the target image region of the target image into a pre-acquired hand position model to obtain the hand position information output by the hand position model that indicates the position of the hand.

[0109] Alternatively, obtaining the hand movement distance corresponding to two adjacent target images can also be understood as inputting two adjacent target images into a pre-acquired hand movement distance model to obtain the hand movement distance output by the hand movement distance model.

[0110] In step S104, multiple target images are framed according to the hand movement distance to obtain multiple framed target images.

[0111] Among them, the hand movement distance corresponding to any two adjacent target frames in multiple extracted target images belongs to the hand movement distance interval.

[0112] In one embodiment of this disclosure, frame extraction from multiple target images based on hand movement distance can be understood as follows: First, any target image is selected as the frame extraction target image from the multiple target images. Then, another target image is selected whose sampling time is after the sampling time of the image to be framed. A first hand movement distance is obtained between this other target image and the image to be framed. If the first hand movement distance falls within a hand movement distance range, this other target image is also selected as the frame extraction target image. Next, another target image is selected whose sampling time is after the sampling time of the next frame extraction target image. A second hand movement distance is obtained between this second target image and the next frame extraction target image. If the second hand movement distance falls within a hand movement distance range, this second target image is also selected as the frame extraction target image. This process is repeated until multiple images to be framed and the frame extraction target image are obtained.

[0113] Alternatively, frame extraction from multiple target images based on hand movement distance can be understood as taking multiple target images and the hand movement distances corresponding to two adjacent target images as input, inputting them into a pre-acquired first frame extraction model, and obtaining multiple frame-extracted target images output by the first frame extraction model.

[0114] In the technical solution provided in this disclosure, by acquiring the image data to be processed collected by the display case, the hand in the target image region of each of the multiple images to be processed is identified, multiple target images are determined in the multiple images to be processed based on the hand recognition results, and the hand movement distance corresponding to any two adjacent target images in the multiple target images is obtained. The multiple target images are then framed based on the hand movement distance to obtain multiple framed target images. The target image is the image to be processed that includes the hand in the target image region, which includes the item entrance / exit of the display case. Therefore, by focusing on the user's hand movement trajectory in the target image, it can be determined whether the user's hand has moved an item out of or into the display case through the item entrance / exit. By limiting the hand movement distance between any two adjacent target images obtained from multiple frame-by-frame extraction to a hand movement distance range, it is possible to ensure a relatively coherent hand movement trajectory based on the image data obtained from frame extraction while minimizing the amount of data in multiple target images. This reduces the amount of image data to be processed and improves the accuracy of determining whether an item is moved out or into the display case based on multiple target images, thereby reducing data processing costs and improving the user experience.

[0115] In one embodiment of this disclosure, the method further includes:

[0116] The number of frames extracted is obtained by subtracting the number of multiple target images from the number of target images to be extracted.

[0117] In response to a frame extraction number being less than or equal to a preset frame extraction number, multiple extracted target images are interpolated to obtain multiple interpolated target images.

[0118] In one embodiment of this disclosure, frame interpolation of multiple extracted target images can be performed by obtaining multiple interpolated target images based on a pre-acquired algorithm. For example, any two adjacent extracted target images can be processed using an interpolation algorithm to obtain an image to be interpolated, which is then inserted into the two adjacent images. Alternatively, a pre-trained interpolation model can be obtained, and the multiple extracted target images can be used as input to obtain multiple interpolated target images output by the model.

[0119] In the technical solution provided in this disclosure, considering that when the user's hand moves too fast, the number of frames may be small, and it may not be possible to reliably track the user's hand trajectory based on multiple framed target images. Therefore, by obtaining the number of frames obtained by subtracting the number of multiple framed target images from the number of multiple target images, and responding to the number of frames being less than or equal to a preset number of frames, the multiple framed target images are interpolated to obtain multiple interpolated target images. This ensures that the number of interpolated target images is large, and the user's hand trajectory can be reliably tracked based on the multiple interpolated target images. This helps to improve the accuracy of determining items being moved out or into the display case based on the multiple interpolated target images, and improves the user experience.

[0120] In one embodiment of this disclosure, the number of interpolated frames obtained by subtracting the number of extracted frames from the number of interpolated target images is greater than or equal to the difference between the preset number of extracted frames and the number of extracted frames.

[0121] In the technical solution provided in this disclosure, by limiting the number of interpolated frames obtained by subtracting the number of extracted frames from the number of interpolated target images, and ensuring that the number of interpolated frames is greater than or equal to the difference between the preset number of extracted frames and the number of extracted frames, the number of images used to determine the items being moved out or into the display case can be set more conveniently, thereby improving the user experience.

[0122] In one embodiment of this disclosure, obtaining the hand movement distance corresponding to any two adjacent target images includes:

[0123] In response to the fact that each of the two adjacent target images contains only one hand in its respective target image region, the distance that a hand moves in the two adjacent target images is determined as the hand movement distance corresponding to the two adjacent target images.

[0124] Alternatively, in response to the fact that each of the target image regions of any two adjacent target images includes multiple hands, the target hand among the multiple hands that moves the longest distance in any two adjacent target images is determined, and the distance that the target hand moves in any two adjacent target images is determined as the hand movement distance corresponding to any two adjacent target images.

[0125] In the technical solution provided in this disclosure, considering that in reality, there may be only one user moving items out of or into the display case at any given time, or multiple users moving items out of or into the display case at the same time, the target image area may include only one hand or multiple hands simultaneously. To ensure that the number of hands in the target image area does not affect the tracking of hands based on multiple frame-dropped target images while minimizing the number of multiple frame-dropped target images obtained after frame-dropping, the tracking is limited to when the target image area of ​​any two adjacent target images each includes only one hand. This limits the movement of a hand between any two adjacent target images. The distance is defined as the hand movement distance between any two adjacent target images. Each of the target image regions in any two adjacent target images contains multiple hands. The target hand with the longest movement distance between any two adjacent target images is identified, and the movement distance of the target hand between any two adjacent target images is defined as the hand movement distance between any two adjacent target images. This ensures that regardless of how many hands exist in the target image region at the same time, the fastest moving hand in the target image region can be stably tracked based on multiple frame-sampling target images. This helps improve the accuracy of determining items being moved out or into the display case based on multiple frame-sampling target images and improves the user experience.

[0126] In one embodiment of this disclosure, the method further includes:

[0127] Obtain cabinet door opening indication information, which is used to indicate that the cabinet door of the display cabinet has been opened;

[0128] In response to the cabinet door opening indication, acquire the image data to be processed.

[0129] In one embodiment of this disclosure, obtaining cabinet door opening instruction information can be understood as receiving cabinet door opening instruction information sent by the cabinet door locking device of the display cabinet, or it can be understood as receiving cabinet door opening instruction information sent by other devices or systems.

[0130] In the technical solution provided in this disclosure, by acquiring cabinet door opening indication information and responding to the cabinet door opening indication information to acquire image data to be processed, the amount of image data to be processed can be minimized without affecting the accuracy of determining the items being moved out or into the display cabinet, thereby reducing processing costs.

[0131] In one embodiment of this disclosure, the method further includes:

[0132] Based on the hand recognition results, multiple first frame images to be extracted are determined from multiple images to be processed. The first frame images to be extracted are images to be processed where the corresponding target image area does not include the hand and the hand is located outside the item storage area of ​​the display case.

[0133] Based on the acquisition time, multiple first frames to be extracted are extracted to obtain multiple first extracted frames. The time difference between the acquisition times of any two adjacent first extracted frames belongs to the first acquisition time difference interval.

[0134] In the technical solution provided in this disclosure, considering that when the target image area does not include the hand and the hand is located outside the item storage area of ​​the display case, it can be understood that the user is in a state where the display case door has just been opened and no other action has been taken, or it can be understood that the user has completed the operation of moving items out or into the display case and is about to close the display case door. In either of the above situations, the user cannot move items out or into the display case. There is no need to track the trajectory of the user's hand based on the image. Instead, multiple first frames to be extracted are determined from multiple images to be processed based on the hand recognition results. Frames are extracted from the multiple first frames to be extracted based on the acquisition time to obtain multiple first frame images. The time difference between the acquisition times of any two adjacent first frame images in the multiple first frame images belongs to the first acquisition time difference interval. This can minimize the amount of data in the multiple first frame images and reduce the processing cost of further processing the multiple first frame images without affecting the accuracy of determining the items moved out or into the display case.

[0135] In one embodiment of this disclosure, the method further includes:

[0136] Based on the hand recognition results, multiple second frames to be extracted are determined from multiple images to be processed. The second frames to be extracted are images to be processed where the corresponding target image area does not include the hand, and the hand is located in the item storage area of ​​the display case.

[0137] Based on the acquisition time, multiple second-frame images are extracted to obtain multiple second-frame images. The time difference between the acquisition times of any two adjacent second-frame images belongs to the second acquisition time difference interval.

[0138] In the technical solution provided in this disclosure, considering that when the target image area does not include the hand, and the hand is located within the item storage area of ​​the display case, it can be understood that the user is in a state where the display case door is open and the user has reached into the item storage area of ​​the display case, taking or putting down items in the item storage area. Under the above situation, the items taken or put down by the user in the multiple second frame images to be extracted are not clear, and it is impossible to determine the items moved out or into the display case based on the multiple second frame images to be extracted. Therefore, by determining multiple second frame images to be extracted from the multiple images to be processed based on the hand recognition result, and extracting frames from the multiple second frame images to be extracted according to the acquisition time to obtain multiple second frame images, the amount of data of multiple second frame images can be reduced as much as possible without affecting the accuracy of determining the items moved out or into the display case, thereby reducing the processing cost of further processing multiple first frame images.

[0139] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein.

[0140] Figure 9 A schematic structural block diagram of an image data processing apparatus according to an embodiment of the present disclosure is shown. This image data processing apparatus can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 9 As shown, the image data processing device includes:

[0141] The image data acquisition module is configured to acquire image data to be processed collected by the display case. The image data to be processed includes multiple images to be processed arranged according to the acquisition time.

[0142] The hand recognition module is configured to recognize hands in the target image region of each of multiple images to be processed, including the item entrances and exits of the display case.

[0143] The movement distance acquisition module is configured to determine multiple target images from multiple images to be processed based on the hand recognition results, and to acquire the hand movement distance corresponding to any two adjacent target images among the multiple target images. The target image is the image to be processed that includes the hand in the corresponding target image region.

[0144] The frame extraction module is configured to extract frames from multiple target images based on the hand movement distance to obtain multiple frame-extracted target images. The hand movement distances corresponding to any two adjacent frame-extracted target images in the multiple frame-extracted target images belong to the hand movement distance interval.

[0145] The above technical solution acquires image data to be processed collected by the display case, identifies the hand in the target image region of each of the multiple images to be processed, determines multiple target images in the multiple images to be processed based on the hand recognition results, obtains the hand movement distance corresponding to any two adjacent target images in the multiple target images, and performs frame extraction on the multiple target images based on the hand movement distance to obtain multiple frame-extracted target images. The target image is the image to be processed that includes the hand in the target image region, which includes the item entrance / exit of the display case. Therefore, by focusing on the user's hand movement trajectory in the target image, it can be determined whether the user's hand has moved an item out of or into the display case through the item entrance / exit. By limiting the hand movement distance between any two adjacent target images obtained from multiple frame-by-frame extraction to a hand movement distance range, it is possible to ensure a relatively coherent hand movement trajectory based on the image data obtained from frame extraction while minimizing the amount of data in multiple target images. This reduces the amount of image data to be processed and improves the accuracy of determining whether an item is moved out or into the display case based on multiple target images, thereby reducing data processing costs and improving the user experience.

[0146] This disclosure also discloses an electronic device. Figure 10 A schematic structural block diagram of an electronic device according to an embodiment of the present disclosure is shown, such as... Figure 10 As shown, the electronic device includes a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the above method steps.

[0147] Figure 11 This is a schematic diagram of the structure of a computer system suitable for implementing an image data processing method according to an embodiment of the present disclosure. For example... Figure 11 As shown, the computer system includes a processing unit that can execute various processes described above based on a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer system. The processing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0148] The following components are connected to the I / O interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard disks; and communication sections including network interface cards such as LAN cards and modems. The communication sections perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as needed. The processing unit can be implemented as a CPU, GPU, TPU, FPGA, NPU, etc.

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0150] The units or modules described in the embodiments of this disclosure can be implemented in software or hardware. The described units or modules can also be located in a processor, and the names of these units or modules do not necessarily constitute a limitation on the unit or module itself.

[0151] In another aspect, this disclosure also provides a computer-readable storage medium, which may be a computer-readable storage medium included in the apparatus described in the above embodiments; or it may be a standalone computer-readable storage medium not assembled into a device. The computer-readable storage medium stores one or more programs that are used by one or more processors to perform the methods described in this disclosure.

[0152] In addition, this disclosure also provides a computer program product storing a computer program that, when executed by a processor, enables the processor to at least implement the methods provided in the foregoing embodiments.

[0153] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

Claims

1. An image data processing method, characterized in that, The method includes: Acquire image data to be processed collected from the display case, the image data to be processed including multiple images to be processed arranged according to the collection time; The hand is identified in the target image region of each of the multiple images to be processed, and the target image region includes the item entrance / exit of the display case; Based on the hand recognition results, multiple target images are determined from the multiple images to be processed, and the hand movement distance corresponding to any two adjacent target images is obtained. The target image is the image to be processed whose corresponding target image region includes a hand. Obtaining the hand movement distance corresponding to any two adjacent target images includes: in response to the fact that each of the two adjacent target images contains only one hand in its target image region, determining the distance moved by that one hand in the two adjacent target images as the hand movement distance corresponding to the two adjacent target images; or, in response to the fact that each of the two adjacent target images contains multiple hands in its target image region, determining the target hand with the longest movement distance in the two adjacent target images, and determining the distance moved by that target hand in the two adjacent target images as the hand movement distance corresponding to the two adjacent target images. The multiple target images are framed according to the hand movement distance to obtain multiple framed target images. The hand movement distances corresponding to any two adjacent framed target images in the multiple framed target images belong to the hand movement distance interval.

2. The image data processing method according to claim 1, characterized in that, The method further includes: The number of frames extracted is obtained by subtracting the number of the multiple target images from the number of the extracted target images. In response to the number of frames being less than or equal to a preset number of frames, the multiple target images with frames being extracted are interpolated to obtain multiple target images with interpolated frames.

3. The image data processing method according to claim 2, characterized in that, The number of frames obtained by subtracting the number of frames extracted from the number of target images for multiple frame-filling is greater than or equal to the difference between the preset number of frames extracted and the number of frames extracted.

4. The image data processing method according to any one of claims 1-3, characterized in that, The method further includes: Obtain cabinet door opening indication information, which is used to indicate that the cabinet door of the display cabinet is opened; In response to the cabinet door opening indication information, the image data to be processed is acquired.

5. The image data processing method according to claim 4, characterized in that, The method further includes: Based on the hand recognition result, multiple first frame images to be extracted are determined from the multiple images to be processed. The first frame images to be extracted are images to be processed where the corresponding target image area does not include the hand and the hand is located outside the item storage area of ​​the display case. According to the acquisition time, the multiple first frames to be extracted are extracted to obtain multiple first extracted frames. The time difference between the acquisition times of any two adjacent first extracted frames in the multiple first extracted frames belongs to the first acquisition time difference interval.

6. The image data processing method according to claim 5, characterized in that, The method further includes: Based on the hand recognition result, multiple second frame images to be extracted are determined from the multiple images to be processed. The second frame images to be extracted are images whose corresponding target image areas do not include the hand, and whose hand is located in the item storage area of ​​the display case. According to the acquisition time, the multiple second frames to be extracted are extracted to obtain multiple second frame-extracted images. The time difference between the acquisition times of any two adjacent second frame-extracted images belongs to the second acquisition time difference interval.

7. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the method of any one of claims 1-6.

8. A computer-readable storage medium storing computer instructions thereon, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-6.

9. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Commodity pick-and-place recognition device and commodity pick-and-place recognition method

    CN113052020A

  • Inventory management device for product showcase using artificial intelligence-based image recognition, management method therefor, and refrigerator including same

    KR102429774B1