Image recognition methods, apparatus, devices and computer-readable storage media
By setting the unit duration based on the number of object frames on the smart refrigerator's large screen, and combining voice and focus events to filter target object frames and perform deduplication, the problem of low efficiency in recognizing people on the smart refrigerator's large screen has been solved, achieving a more efficient and accurate recognition effect.
Patent Information
- Application Number
- CN202011221473.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-04
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2040-11-04
AI Technical Summary
Existing smart refrigerator screens are inefficient and prone to errors when recognizing people on the screen, especially when multiple celebrity images appear simultaneously, resulting in excessively long recognition times and high error rates.
By determining the number of object frames in the current page and setting the unit duration, the target object frames are filtered out by combining external voice information and focus events. The target object images are deduplicated to narrow the recognition range, reduce the workload of the recognition task, and the recognition is performed using a cloud server to display the recognition results.
This improves the efficiency of the smart refrigerator's large screen in recognizing people on the screen, reduces resource waste and time consumption, and enhances recognition accuracy and user experience.
Smart Images

Figure CN112348077B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, and in particular to an image recognition method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] As users demand a higher level of experience during the cooking process, smart refrigerator screens are becoming increasingly common in homes. These screens not only meet basic audio-visual needs but also cater to users' entertainment needs, such as following celebrities. Users can conveniently browse entertainment information on the smart refrigerator screen during breaks in cooking. For example, the screen can display images of celebrities on the main page. When a user clicks or swipes, the smart refrigerator recognizes the celebrity's image, retrieves their corresponding information, and then displays the relevant celebrity profile page. However, traditionally, when a user opens a celebrity profile page, recognizing multiple celebrities on a single main page can be time-consuming. The more celebrities recognized, the greater the chance of error, leading to the technical problem of low recognition efficiency for people appearing on a page in existing smart refrigerator screens. Summary of the Invention
[0003] The main objective of this invention is to provide an image recognition method that aims to solve the technical problem of low recognition efficiency of people appearing on the screen of an existing smart refrigerator.
[0004] To achieve the above objectives, the present invention provides an image recognition method, which is applied to a smart terminal with a screen, and the method includes:
[0005] Determine the number of object frames in the current page of the smart terminal, and set the unit duration according to the number of object frames, wherein each object frame contains an object image;
[0006] Acquire external voice information and / or monitor focus events triggered within each object frame within the unit time period, and filter out target object frames from the object frames based on the external voice information and / or the focus events;
[0007] The target object image contained within the target object bounding box is obtained, and the target object image is deduplicated to identify the deduplicated target object image and obtain the target object recognition result.
[0008] Optionally, the step of acquiring external voice information and / or monitoring focus events triggered within each object frame within the unit time period, and filtering target object frames from the object frames based on the external voice information and / or the focus events, includes:
[0009] Within the specified unit duration, external voice information is acquired. When a preset keyword exists in the external voice information, the object frame pointed to by the external voice information is determined and recorded as a voice event of the object frame pointed to by the external voice.
[0010] Within the specified unit duration, focus events triggered within the scope of each of the specified object frames are monitored, wherein the focus events include click events and / or gaze events;
[0011] According to the preset event weight conversion rules, the focus event and the voice event are converted into corresponding weights to obtain the total weight corresponding to each object box;
[0012] When the sum of the weights reaches a preset weight threshold, the object box corresponding to the sum of the weights is taken as the target object box.
[0013] Optionally, obtaining the target object image contained within the target object frame and performing deduplication processing on the target object image includes:
[0014] Obtain pixel information for each target object image, and obtain the pixel difference between any two target object frames based on the pixel information;
[0015] When the pixel difference meets the preset pixel difference condition, the two corresponding target object images are taken as similar image pairs, and any one of the target object images in the similar image pairs is retained as the deduplicated target object image.
[0016] Optionally, obtaining pixel information for each of the target object images and obtaining the pixel difference between any two target object bounding boxes based on the pixel information includes:
[0017] Obtain the horizontal pixel value in the horizontal direction and the vertical pixel value in the vertical direction of each target object image, and use the horizontal pixel value and the vertical pixel value as the pixel information;
[0018] Obtain the difference in horizontal pixel values and the difference in vertical pixel values between any two images of the target object, and use the difference in horizontal pixel values and the difference in vertical pixel values as the pixel difference;
[0019] Before the pixel difference meets the preset pixel difference condition, the method further includes:
[0020] Determine whether both the horizontal pixel difference and the vertical pixel difference are less than a preset pixel difference threshold;
[0021] If so, the pixel difference is determined to satisfy the preset pixel difference condition.
[0022] Optionally, determining the number of object frames on the current page of the smart terminal and setting a unit duration based on the number of object frames, wherein each object frame contains an object image, including:
[0023] Upon receiving a target recognition instruction, the number of object boxes in the current page of the smart terminal is determined based on the target recognition instruction;
[0024] The number of object frames is used as the timing value, and the unit duration is set in combination with the preset timing unit.
[0025] Optionally, the target object image includes a target person image.
[0026] The process of identifying the deduplicated target object image and obtaining the target object identification result includes:
[0027] The image of the target person is uploaded to a cloud server so that the cloud server can recognize the image of the target person.
[0028] The system receives the person recognition result of the target person image fed back by the cloud server, and uses it as the target object recognition result.
[0029] Optionally, after obtaining the target object image contained within the target object bounding box, performing deduplication processing on the target object image, and recognizing the deduplicated target object image to obtain the target object recognition result, the method further includes:
[0030] Based on the target object identification results, target object description information is generated and displayed on the current page.
[0031] Furthermore, to achieve the above objectives, the present invention also provides an image recognition device, the image recognition device comprising:
[0032] The duration setting module is used to determine the number of object boxes in the current page of the smart terminal, and set the unit duration according to the number of object boxes, wherein the object box contains an object image;
[0033] The target determination module is used to acquire external voice information and / or listen to focus events triggered within each object frame within the unit time period, and filter out target object frames from the object frames based on the external voice information and / or the focus events.
[0034] The target recognition module is used to acquire the target object image contained within the target object box, perform deduplication processing on the target object image, and then recognize the deduplicated target object image to obtain the target object recognition result.
[0035] Optionally, the target determination module includes:
[0036] The voice event recording unit is used to acquire external voice information within the unit duration, determine the object frame pointed to by the external voice information when a preset keyword exists in the external voice information, and record it as a voice event of the object frame pointed to by the external voice.
[0037] A focus event listening unit is used to listen for focus events triggered within the scope of each of the object frames within the unit duration, wherein the focus events include click events and / or gaze events;
[0038] The event weight conversion unit is used to convert the focus event and the voice event into corresponding weights according to the preset event weight conversion rules, so as to obtain the total weight of each object box.
[0039] The target object filtering unit is used to select the object box corresponding to the total weight as the target object box when the total weight reaches a preset weight threshold.
[0040] Optionally, the target recognition module includes:
[0041] A pixel difference acquisition unit is used to acquire pixel information of each of the target object images and obtain the pixel difference between any two target object boxes based on the pixel information.
[0042] The target image deduplication unit is used to, when the pixel difference meets the preset pixel difference condition, treat the corresponding two target object images as similar image pairs, and retain any one of the target object images in the similar image pairs as the deduplicated target object image.
[0043] Optionally, the pixel difference acquisition unit is further configured to:
[0044] Obtain the horizontal pixel value in the horizontal direction and the vertical pixel value in the vertical direction of each target object image, and use the horizontal pixel value and the vertical pixel value as the pixel information;
[0045] Obtain the difference in horizontal pixel values and the difference in vertical pixel values between any two images of the target object, and use the difference in horizontal pixel values and the difference in vertical pixel values as the pixel difference;
[0046] Before the pixel difference meets the preset pixel difference condition, the method further includes:
[0047] Determine whether both the horizontal pixel difference and the vertical pixel difference are less than a preset pixel difference threshold;
[0048] If so, the pixel difference is determined to satisfy the preset pixel difference condition.
[0049] Optionally, the duration setting module includes:
[0050] The number determination unit is used to determine the number of object boxes in the current page of the smart terminal based on the target recognition instruction when a target recognition instruction is received.
[0051] The duration determination unit is used to set the unit duration by taking the number of object frames as the timing value and combining it with a preset timing unit.
[0052] Optionally, the target object image includes a target person image.
[0053] The target recognition module includes:
[0054] The target image uploading unit is used to upload the target person image to the cloud server so that the cloud server can recognize the target person image;
[0055] The recognition result receiving unit is used to receive the person recognition result of the target person image fed back by the cloud server, and use it as the target object recognition result.
[0056] Optionally, the image recognition device further includes:
[0057] The introductory information display module is used to generate introductory information about the target object based on the target object identification result, and to display the introductory information about the target object on the current page.
[0058] In addition, to achieve the above objectives, the present invention also provides an image recognition device, the image recognition device comprising: a memory, a processor, and an image recognition program stored in the memory and executable on the processor, wherein the image recognition program, when executed by the processor, performs the steps of the method described above.
[0059] In addition, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an image recognition program, which, when executed by a processor, implements the steps of the method described above.
[0060] This invention provides an image recognition method, apparatus, device, and computer-readable storage medium. The image recognition method sets a unit duration based on the actual number of object frames, making the unit duration more suitable for the current scenario and avoiding being set too long or too short. By acquiring external voice information and / or listening to focus events within a unit time, and filtering target object frames based on this, it is possible to select the target object frames that the user currently wants to recognize according to the actual situation, narrowing the recognition range and reducing the workload of the recognition task. Furthermore, by deduplicating the target object images, the workload of the recognition task is further reduced, avoiding the waste of resources and time caused by repeatedly recognizing the same content. Therefore, it can improve the recognition efficiency of the deduplicated target object images, thereby solving the technical problem of low recognition efficiency of people appearing on the screen of existing smart refrigerators. Attached Figure Description
[0061] Figure 1 This is a schematic diagram of the image recognition device structure in the hardware operating environment involved in the embodiments of the present invention;
[0062] Figure 2 This is a flowchart illustrating the first embodiment of the image recognition method of the present invention;
[0063] Figure 3 This is a schematic diagram of the functional modules of the image recognition device of the present invention.
[0064] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0065] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0066] like Figure 1 As shown, Figure 1 This is a schematic diagram of the image recognition device structure in the hardware operating environment involved in the embodiments of the present invention.
[0067] In this embodiment of the invention, the image recognition device is a terminal with an image capturing device, preferably a smart TV.
[0068] like Figure 1As shown, the image recognition device may include: a processor 1001, such as a CPU; a communication bus 1002; a user interface 1003; a network interface 1004; and a memory 1005. The communication bus 1002 is used to enable communication between these components. The optional user interface 1003 may include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM or a stable, non-volatile memory. Alternatively, the memory 1005 may be a storage device independent of the aforementioned processor 1001.
[0069] Those skilled in the art will understand that Figure 1 The image recognition device structure shown does not constitute a limitation on the image recognition device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0070] like Figure 1 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an image recognition program.
[0071] exist Figure 1 In the image recognition device shown, the network interface 1004 is mainly used to connect to the backend server and communicate data with it; the user interface 1003 is mainly used to connect to the client (user terminal) and communicate data with it; and the processor 1001 can be used to call the image recognition program stored in the memory 1005 and perform the following operations:
[0072] Determine the number of object frames in the current page of the smart terminal, and set the unit duration according to the number of object frames, wherein each object frame contains an object image;
[0073] Acquire external voice information and / or monitor focus events triggered within each object frame within the unit time period, and filter out target object frames from the object frames based on the external voice information and / or the focus events;
[0074] The target object image contained within the target object bounding box is obtained, and the target object image is deduplicated to identify the deduplicated target object image and obtain the target object recognition result.
[0075] Further, the step of acquiring external voice information and / or monitoring focus events triggered within each object frame within the unit time period, and filtering target object frames from the object frames based on the external voice information and / or the focus events, includes:
[0076] Within the specified unit duration, external voice information is acquired. When a preset keyword exists in the external voice information, the object frame pointed to by the external voice information is determined and recorded as a voice event of the object frame pointed to by the external voice.
[0077] Within the specified unit duration, focus events triggered within the scope of each of the specified object frames are monitored, wherein the focus events include click events and / or gaze events;
[0078] According to the preset event weight conversion rules, the focus event and the voice event are converted into corresponding weights to obtain the total weight corresponding to each object box;
[0079] When the sum of the weights reaches a preset weight threshold, the object box corresponding to the sum of the weights is taken as the target object box.
[0080] Further, the step of obtaining the target object image contained within the target object frame and performing deduplication processing on the target object image includes:
[0081] Obtain pixel information for each target object image, and obtain the pixel difference between any two target object frames based on the pixel information;
[0082] When the pixel difference meets the preset pixel difference condition, the two corresponding target object images are taken as similar image pairs, and any one of the target object images in the similar image pairs is retained as the deduplicated target object image.
[0083] Further, the step of acquiring pixel information of each of the target object images and obtaining the pixel difference between any two target object bounding boxes based on the pixel information includes:
[0084] Obtain the horizontal pixel value in the horizontal direction and the vertical pixel value in the vertical direction of each target object image, and use the horizontal pixel value and the vertical pixel value as the pixel information;
[0085] Obtain the difference in horizontal pixel values and the difference in vertical pixel values between any two images of the target object, and use the difference in horizontal pixel values and the difference in vertical pixel values as the pixel difference;
[0086] Before the pixel difference meets the preset pixel difference condition, the method further includes:
[0087] Determine whether both the horizontal pixel difference and the vertical pixel difference are less than a preset pixel difference threshold;
[0088] If so, the pixel difference is determined to satisfy the preset pixel difference condition.
[0089] Further, the step of determining the number of object frames in the current page of the smart terminal and setting a unit duration based on the number of object frames, wherein each object frame contains an object image, including:
[0090] Upon receiving a target recognition instruction, the number of object boxes in the current page of the smart terminal is determined based on the target recognition instruction;
[0091] The number of object frames is used as the timing value, and the unit duration is set in combination with the preset timing unit.
[0092] Furthermore, the target object image includes a target person image.
[0093] The process of identifying the deduplicated target object image and obtaining the target object identification result includes:
[0094] The image of the target person is uploaded to a cloud server so that the cloud server can recognize the image of the target person.
[0095] The system receives the person recognition result of the target person image fed back by the cloud server, and uses it as the target object recognition result.
[0096] Further, after acquiring the target object image contained within the target object frame, performing deduplication processing on the target object image, recognizing the deduplicated target object image, and obtaining the target object recognition result, the processor 1001 can call the image recognition program stored in the memory 1005 and perform the following operations:
[0097] Based on the target object identification results, target object description information is generated and displayed on the current page.
[0098] Based on the above hardware structure, various embodiments of the image recognition method of the present invention are proposed.
[0099] As users demand a higher level of experience during the cooking process, smart refrigerator screens are becoming increasingly common in homes. These screens not only meet basic audio-visual needs but also cater to users' entertainment needs, such as following celebrities. Users can conveniently browse entertainment information on the smart refrigerator screen during breaks in cooking. For example, the screen can display images of celebrities on the main page. When a user clicks or swipes, the smart refrigerator recognizes the celebrity's image, retrieves their corresponding information, and then displays the relevant celebrity profile page. However, traditionally, when a user opens a celebrity profile page, recognizing multiple celebrities on a single main page can be time-consuming. The more celebrities recognized, the greater the chance of error, leading to the technical problem of low recognition efficiency for people appearing on a page in existing smart refrigerator screens.
[0100] To address the aforementioned technical problems, this invention provides an image recognition method. This method sets the unit duration based on the actual number of object frames, making the unit duration more suitable for the current scenario and avoiding setting it too long or too short. By acquiring external voice information and / or monitoring focus events within a unit time, and filtering target object frames based on this information, the method can select the target object frames that the user currently wants to recognize based on the actual situation, narrowing the recognition range and reducing the workload of the recognition task. Furthermore, by deduplicating the target object images, the workload of the recognition task is further reduced, avoiding the waste of resources and time caused by repeatedly recognizing the same content. Therefore, it can improve the recognition efficiency of the deduplicated target object images, thereby solving the technical problem of low recognition efficiency of people appearing on the screen of existing smart refrigerators.
[0101] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the image recognition method.
[0102] The first embodiment of the present invention provides an image recognition method, which is applied to a smart terminal with a screen. The image recognition method includes:
[0103] Step S10: Determine the number of object frames in the current page of the smart terminal, and set the unit duration according to the number of object frames, wherein the object frame contains an object image;
[0104] In this embodiment, the method is applied to a smart terminal with a screen, typically a smart home appliance with a screen, such as a smart refrigerator, smart washing machine, or smart TV. For ease of description, a smart refrigerator will be used as an example below. The current page is the content page displayed by the smart refrigerator when providing functional or entertainment services to the user. The page may include images of various types of elements, such as images of people and objects. The object image refers to the images of people, objects, and other page content that may appear on the page. The object frame is a graphic frame that defines the scope of an object. The specific shape can be flexibly set according to actual needs, such as rectangle, circle, or polygon. The object frame can be hidden or displayed on the page. The number of objects in the object frame can be one or more; generally, one object frame defines one object image. The unit duration is the period of filtering and recognition, and the starting point of the unit duration is the moment the user enters the current page. The smart refrigerator's large-screen system can directly detect the position and number of object frames contained in the current page. The unit duration can be set based on the number of object boxes contained on the current page. Specifically, it can be set by directly using the number of boxes as the duration value and then assigning an appropriate time unit, or it can be set by performing some calculations based on the number of boxes and using the result as the duration value and then assigning an appropriate time unit, or other methods. It can be set flexibly according to the actual situation.
[0105] Step S20: Acquire external voice information and / or listen to focus events triggered within each object frame within the unit time period, and filter out target object frames from the object frames based on the external voice information and / or the focus events;
[0106] In this embodiment, external voice information refers to the user's voice feedback on the content on the page within a unit of time. External voice information is typically acquired through a built-in or external microphone in the smart refrigerator. Focus events include click events, gaze events, and touch events. A click event occurs when the user clicks on a specific location on the current page using a remote control, a corresponding button on the smart refrigerator, or a corresponding touchscreen button on the smart refrigerator. The click location is counted as a click event for the object frame within which the click occurs. A gaze event occurs when the smart refrigerator captures the user's gaze on the current page using its built-in or external camera. Analysis determines which object frame the user is currently gazing at, and if the gaze duration reaches a preset gaze threshold, each gaze event is counted as a gaze event for that object frame. A touch event occurs when the user touches a location on the current page using the smart refrigerator's touchscreen; each touch is counted as a touch event for the corresponding object frame. The target object box is a subset of object boxes selected from all object boxes on the current page. In practice, all object boxes on the current page may be eligible as target object boxes, only a portion may be eligible, or there may be no target object boxes at all. The target object box can be selected based solely on external audio information, solely on focus events, or a combination of both.
[0107] Step S30: Obtain the target object image contained within the target object frame, perform deduplication processing on the target object image, and then identify the deduplicated target object image to obtain the target object identification result.
[0108] In this embodiment, the target object image is the object image framed by the target object bounding box. The deduplication strategy for the target object image can be to randomly select one of the target object images with the same theme and retain it, or to select the one with the highest or lowest pixel count among the target object images with the same theme and retain it, or other deduplication strategies, which can be flexibly formulated according to actual needs. The target object recognition result is the relevant information obtained after recognizing the target object image, which may specifically include the descriptive content and related information of the target object image.
[0109] As a specific implementation, when a user activates the large screen and selects the entertainment function, the smart refrigerator displays an entertainment page featuring five celebrity images, each corresponding to a character frame. The large screen system calculates twice the number of character frames as the timing duration, assigning seconds as the time unit, i.e., setting 10 seconds as the unit duration. Within 10 seconds of the user opening the page, the system monitors the user's gaze, click, and touch events, and acquires the user's voice information during these 10 seconds. If the user clicks on a character in a frame and says "I don't know" or "I don't recognize," this is recorded as a click event and voice event for that character frame. Based on the information acquired during these 10 seconds, the system filters these five character frames. If three character frames are selected as target character frames, the system begins to check for duplicate characters within these three frames, specifically using facial recognition technology. If duplicate characters exist, one is retained; otherwise, no deduplication is performed. If no duplicates are found at this time, the celebrities in these three figure frames can be identified, and the identification results can be displayed in the corresponding positions on the page for users to view.
[0110] In this embodiment, the number of object frames in the current page of the smart terminal is determined, and a unit duration is set according to the number of object frames. Each object frame contains an object image. Within the unit duration, external voice information is acquired and / or focus events triggered within the range of each object frame are monitored. Target object frames are then selected from the object frames based on the external voice information and / or the focus events. The target object image contained within the target object frame is acquired, and the target object image is deduplicated. The deduplicated target object image is then recognized to obtain the target object recognition result. Through the above methods, this invention sets the unit duration based on the actual number of object frames, making the unit duration more suitable for the current scenario and avoiding setting it too long or too short. By acquiring external voice information and / or listening to focus events within a unit time and filtering out target object frames based on this, it is possible to select the target object frames that the user currently wants to recognize according to the actual situation, narrowing the recognition range and reducing the workload of the recognition task. By deduplicating the target object images, the workload of the recognition task is further reduced, avoiding the waste of resources and time caused by repeatedly recognizing the same content. Therefore, it can improve the recognition efficiency of the deduplicated target object images, thereby solving the technical problem of low recognition efficiency of people appearing on the page in existing smart refrigerator screens.
[0111] Furthermore, based on the above Figure 2 The first embodiment shown presents a second embodiment of the image recognition method of the present invention. In this embodiment, step S20 includes:
[0112] Within the specified unit duration, external voice information is acquired. When a preset keyword exists in the external voice information, the object frame pointed to by the external voice information is determined and recorded as a voice event of the object frame pointed to by the external voice.
[0113] Within the specified unit duration, focus events triggered within the scope of each of the specified object frames are monitored, wherein the focus events include click events and / or gaze events;
[0114] According to the preset event weight conversion rules, the focus event and the voice event are converted into corresponding weights to obtain the total weight corresponding to each object box;
[0115] When the sum of the weights reaches a preset weight threshold, the object box corresponding to the sum of the weights is taken as the target object box.
[0116] In this embodiment, the preset keywords can be negative words with similar meanings such as "don't know," "don't remember," or "not clear," or interrogative sentences such as "who is this?" The preset event weight conversion rule is a mapping table that records the same or different weights corresponding to different types of events. In addition, for the same type of event with different frequencies, weights with varying values can be further set. For example, for 1-3 click events, each click event corresponds to a weight of 3; for 3-5 click events, each click event can be set to a weight of 4. The weight threshold can be flexibly set according to actual needs. For example, if for the same object frame, 3 click events and 1 voice event are detected within a unit of time, and the 1-3 click events each correspond to a weight of 3, and the voice event corresponds to a weight of 5, then the total weight of the object frame within this unit of time is 14. Assuming the preset weight threshold is 10, the large screen system of the smart refrigerator can determine this object frame as the target object frame.
[0117] Further, the step of obtaining the target object image contained within the target object frame and performing deduplication processing on the target object image includes:
[0118] Obtain pixel information for each target object image, and obtain the pixel difference between any two target object frames based on the pixel information;
[0119] When the pixel difference meets the preset pixel difference condition, the two corresponding target object images are taken as similar image pairs, and any one of the target object images in the similar image pairs is retained as the deduplicated target object image.
[0120] In this embodiment, the large-screen system of the smart refrigerator acquires the pixel information of the currently determined target object image, such as pixel value size, color level information, or grayscale values after converting it to a grayscale image. The system compares the pixel differences between each target object image and then determines whether they meet preset pixel difference conditions. If the system determines that the conditions are met, it treats the two target object images that meet the conditions as similar image pairs and selects one of them as the target object image that can be retained.
[0121] Further, the step of acquiring pixel information of each of the target object images and obtaining the pixel difference between any two target object bounding boxes based on the pixel information includes:
[0122] Obtain the horizontal pixel value in the horizontal direction and the vertical pixel value in the vertical direction of each target object image, and use the horizontal pixel value and the vertical pixel value as the pixel information;
[0123] Obtain the difference in horizontal pixel values and the difference in vertical pixel values between any two images of the target object, and use the difference in horizontal pixel values and the difference in vertical pixel values as the pixel difference;
[0124] Before the pixel difference meets the preset pixel difference condition, the method further includes:
[0125] Determine whether both the horizontal pixel difference and the vertical pixel difference are less than a preset pixel difference threshold;
[0126] If so, the pixel difference is determined to satisfy the preset pixel difference condition.
[0127] In this embodiment, pixel information includes horizontal pixel values and vertical pixel values. Horizontal pixel values are the pixel values of the target object bounding box in the horizontal direction, and vertical pixel values are the pixel values of the target object bounding box in the vertical direction. The system acquires the horizontal and vertical pixel values of each target object bounding box, and also acquires the differences in horizontal and vertical pixel values between any two target object bounding boxes. Similarity is determined only when the differences in both horizontal and vertical pixel values of two target object images are less than a preset pixel value threshold, thus satisfying the preset pixel difference condition; otherwise, the preset pixel difference condition is not met.
[0128] In this embodiment, the target object box is further filtered by combining voice events and focus events, making the filtering results more in line with the user's actual needs and the judgment of the user's intent more accurate; the deduplication process is made simple and easy by selecting any duplicate target object image; and the accuracy of the similarity judgment operation is improved by obtaining the horizontal pixel difference and the vertical pixel difference and judging whether they meet the preset conditions.
[0129] Furthermore, based on the above Figure 2 The first embodiment shown is followed by a third embodiment of the image recognition method of the present invention. In this embodiment, step S10 includes:
[0130] Upon receiving a target recognition instruction, the number of object boxes in the current page of the smart terminal is determined based on the target recognition instruction;
[0131] The number of object frames is used as the timing value, and the unit duration is set in combination with the preset timing unit.
[0132] In this embodiment, the target recognition command can be sent by the user to the smart refrigerator's large screen via a physical object or touchscreen button on the smart refrigerator, or a remote control. Upon receiving the command, the system begins detecting the number of object frames present on the current display page. After determining the number, the system directly uses this number as the timing value and assigns an appropriate time unit to obtain the final unit duration. For example, if the system detects 5 object frames on the current page, it sets the timing duration to 5 seconds.
[0133] Furthermore, the target object image includes a target person image.
[0134] The process of identifying the deduplicated target object image and obtaining the target object identification result includes:
[0135] The image of the target person is uploaded to a cloud server so that the cloud server can recognize the image of the target person.
[0136] The system receives the person recognition result of the target person image fed back by the cloud server, and uses it as the target object recognition result.
[0137] In this embodiment, the system uses a cloud server to perform image recognition. The system uploads the deduplicated image of the target object, such as a celebrity image, to the cloud server and simultaneously sends a recognition request. Upon receiving the request, the cloud server identifies the corresponding celebrity's information and sends the recognition result back to the smart refrigerator.
[0138] Furthermore, after step S30, the following steps are also included:
[0139] Based on the target object identification results, target object description information is generated and displayed on the current page.
[0140] In this embodiment, the large-screen system of the smart refrigerator integrates and arranges the currently acquired target object recognition results, displaying them in the corresponding positions of the person frames on the page. For example, the celebrity's name, gender, age, place of origin, and works introduction are arranged and displayed below the person frames for user browsing.
[0141] In this embodiment, the number of object frames is directly used as the duration of each unit of time, making the setting of the unit of time both practical and easy to implement, without requiring excessive computing power from the system. By using a cloud server to recognize the target object image, the accuracy of the recognition results is ensured. By displaying the introductory information generated based on the recognition results on the current page, users can view it immediately, thereby improving the user experience.
[0142] The present invention also provides an image recognition device.
[0143] The image recognition device includes:
[0144] The duration setting module 10 is used to determine the number of object boxes in the current page of the smart terminal and set the unit duration according to the number of object boxes, wherein the object box contains an object image;
[0145] The target determination module 20 is used to acquire external voice information and / or listen to focus events triggered within the scope of each object frame within the unit time period, and filter out target object frames from the object frames based on the external voice information and / or the focus events.
[0146] The target recognition module 30 is used to acquire the target object image contained within the target object frame, perform deduplication processing on the target object image, and then recognize the deduplicated target object image to obtain the target object recognition result.
[0147] The present invention also provides an image recognition device.
[0148] The image recognition device includes a processor, a memory, and an image recognition program stored in the memory and executable on the processor, wherein when the image recognition program is executed by the processor, it implements the steps of the image recognition method as described above.
[0149] The method implemented when the image recognition program is executed can be referred to in various embodiments of the image recognition method of the present invention, and will not be repeated here.
[0150] The present invention also provides a computer-readable storage medium.
[0151] The present invention provides a computer-readable storage medium storing an image recognition program, which, when executed by a processor, implements the steps of the image recognition method described above.
[0152] The method implemented when the image recognition program is executed can be referred to in various embodiments of the image recognition method of the present invention, and will not be repeated here.
[0153] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0154] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0155] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause an image recognition device to execute the methods described in the various embodiments of the present invention.
[0156] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. An image recognition method, characterized in that, The method is applied to a smart terminal with a screen, and the method includes: Determine the number of object frames in the current page of the smart terminal, and set the unit duration according to the number of object frames, wherein the object frame contains an object image; Acquire external voice information and / or monitor focus events triggered within each object frame within the unit time period, and filter out target object frames from the object frames based on the external voice information and / or the focus events; Pixel information of each target object image is obtained, and the pixel difference between any two target object frames is obtained based on the pixel information. The pixel information is the horizontal pixel value in the horizontal direction and the vertical pixel value in the vertical direction of each target object image. The pixel difference is the difference between the horizontal pixel value and the vertical pixel value of any two target object images. When the pixel difference meets the preset pixel difference condition, the corresponding two target object images are taken as similar image pairs, and any one of the target object images in the similar image pairs is retained as the deduplicated target object image, so as to identify the deduplicated target object image and obtain the target object identification result.
2. The method as described in claim 1, characterized in that, The step of acquiring external voice information and / or monitoring focus events triggered within each object frame within the unit time period, and filtering target object frames from the object frames based on the external voice information and / or the focus events, includes: Within the specified unit duration, external voice information is acquired. When a preset keyword exists in the external voice information, the object frame pointed to by the external voice information is determined and recorded as a voice event of the object frame pointed to by the external voice. Within the specified unit duration, focus events triggered within the scope of each of the specified object frames are monitored, wherein the focus events include click events and / or gaze events; According to the preset event weight conversion rules, the focus event and the voice event are converted into corresponding weights to obtain the total weight corresponding to each object box; When the sum of the weights reaches a preset weight threshold, the object box corresponding to the sum of the weights is taken as the target object box.
3. The method as described in claim 1, characterized in that, The step of acquiring pixel information for each target object image and obtaining the pixel difference between any two target object bounding boxes based on the pixel information includes: Obtain the horizontal pixel value in the horizontal direction and the vertical pixel value in the vertical direction of each target object image, and use the horizontal pixel value and the vertical pixel value as the pixel information; Obtain the difference in horizontal pixel values and the difference in vertical pixel values between any two images of the target object, and use the difference in horizontal pixel values and the difference in vertical pixel values as the pixel difference; Before the pixel difference meets the preset pixel difference condition, the method further includes: Determine whether both the horizontal pixel difference and the vertical pixel difference are less than a preset pixel difference threshold; If so, the pixel difference is determined to satisfy the preset pixel difference condition.
4. The method as described in claim 1, characterized in that, The step involves determining the number of object frames on the current page of the smart terminal, and setting a unit duration based on the number of object frames. Each object frame contains an object image, including: Upon receiving a target recognition instruction, the number of object boxes in the current page of the smart terminal is determined based on the target recognition instruction; The number of object frames is used as the timing value, and the unit duration is set in combination with the preset timing unit.
5. The method as described in claim 1, characterized in that, The target object image includes a target person image. The process of identifying the deduplicated target object image and obtaining the target object identification result includes: The image of the target person is uploaded to a cloud server so that the cloud server can recognize the image of the target person. The system receives the person recognition result of the target person image fed back by the cloud server, and uses it as the target object recognition result.
6. The method according to any one of claims 1-5, characterized in that, After obtaining the target object image contained within the target object bounding box, performing deduplication processing on the target object image, recognizing the deduplicated target object image, and obtaining the target object recognition result, the method further includes: Based on the target object identification results, target object description information is generated and displayed on the current page.
7. An image recognition device, characterized in that, The image recognition device includes: The duration setting module is used to determine the number of object boxes on the current page of the smart terminal, and set the unit duration according to the number of object boxes, wherein the object box contains an object image; The target determination module is used to acquire external voice information and / or listen to focus events triggered within each object frame within the unit time period, and filter out target object frames from the object frames based on the external voice information and / or the focus events. The target recognition module is used to acquire pixel information of each target object image, obtain the pixel difference between any two target object frames based on the pixel information, wherein the pixel information is the horizontal pixel value in the horizontal direction and the vertical pixel value in the vertical direction of each target object image, and the pixel difference is the difference between the horizontal pixel value and the vertical pixel value of any two target object images. When the pixel difference satisfies a preset pixel difference condition, the corresponding two target object images are regarded as similar image pairs, and any target object image in the similar image pair is retained as the deduplicated target object image, so as to recognize the deduplicated target object image and obtain the target object recognition result.
8. An image recognition device, characterized in that, The image recognition device includes: a memory, a processor, and an image recognition program stored in the memory and executable on the processor, wherein the image recognition program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an image recognition program, which, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Video processing method and related products
CN108958592A
Video map recognition method, device, terminal and storage medium
CN109034115A