Image Recognition Method, Apparatus, Device, Storage Medium, and Program Product
By integrating image recognition capabilities in the screen projector, the recognition and information superposition of screen projectors is achieved, which solves the problem that existing screen projectors cannot meet user information needs and provides a richer content display and interactive experience.
Patent Information
- Application Number
- CN202110633591.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-06-07
AI Technical Summary
The existing screen projector only implements simple screen projection function and cannot meet users' rich information needs.
By integrating image recognition capabilities in the screen projector, in response to the user's recognition operation, obtain the image to be recognized, perform object recognition, and superimpose the recognition results on the current screen projection video clip and project the screen onto the display screen.
It provides users with the ability to identify any content played on the screen, enhancing the richness and interactivity of content information.
Smart Images

Figure CN113221846B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the fields of computer vision and deep learning technology, and particularly relates to an image recognition method, apparatus, device, storage medium, and program product. Background Art
[0002] Currently, there are a variety of screen mirroring devices on the market, but they only implement screen mirroring technology. The main purpose of screen mirroring is to watch online videos more conveniently. With the increasing richness of network content, simple screen mirroring can no longer meet users' information needs. Summary of the Invention
[0003] Embodiments of the present disclosure propose an image recognition method, apparatus, device, storage medium, and program product.
[0004] In a first aspect, embodiments of the present disclosure propose an image recognition method, including: in response to detecting an identification operation during screen mirroring, obtaining an image to be recognized of the screen mirroring device; obtaining an identification result of an object in the image to be recognized; superimposing the identification result on the current screen mirroring video segment of the screen mirroring device, and screen mirroring it to a display screen connected to the screen mirroring device.
[0005] In a second aspect, embodiments of the present disclosure propose an image recognition apparatus, including: a first obtaining module configured to obtain an image to be recognized of the screen mirroring device in response to detecting an identification operation during screen mirroring; a second obtaining module configured to obtain an identification result of an object in the image to be recognized; a first superimposing module configured to superimpose the identification result on the current screen mirroring video segment of the screen mirroring device, and screen mirror it to a display screen connected to the screen mirroring device.
[0006] In a third aspect, embodiments of the present disclosure propose a server, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation manner of the first aspect.
[0007] In a fourth aspect, embodiments of the present disclosure propose a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in any implementation manner of the first aspect.
[0008] In a fifth aspect, embodiments of the present disclosure propose a computer program product including a computer program, and the computer program implements the method described in any implementation manner of the first aspect when executed by a processor.
[0009] The image recognition method, device, equipment, storage medium, and program product provided by the embodiments of the present disclosure add image recognition capabilities to the screen mirroring device, enabling users to use the image recognition capabilities to recognize any content played through the screen mirroring device during the screen mirroring process, thereby presenting richer content information to users.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 is an exemplary system architecture diagram to which the present disclosure can be applied;
[0013] Figure 2 is a flowchart of the first embodiment of the image recognition method according to the present disclosure;
[0014] Figure 3 is a flowchart of the second embodiment of the image recognition method according to the present disclosure;
[0015] Figure 4 is a flowchart of the third embodiment of the image recognition method according to the present disclosure;
[0016] Figure 5 is a flowchart of the fourth embodiment of the image recognition method according to the present disclosure;
[0017] Figure 6 is a flowchart of the fifth embodiment of the image recognition method according to the present disclosure;
[0018] Figure 7 is a schematic structural diagram of an embodiment of the image recognition device according to the present disclosure;
[0019] Figure 8 is a block diagram of an electronic device for implementing the image recognition method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0021] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The present disclosure will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0022] Figure 1 An exemplary system architecture 100 is shown which can apply embodiments of the image recognition method or the image recognition device of the present disclosure.
[0023] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a screen mirroring device 102, and a display device 103. There is a communication connection between the terminal device 101 and the screen mirroring device 102, and the connection method can be, for example, a wireless network connection. The terminal device 101 can send the video played thereon to the screen mirroring device 102 through the wireless network. There is a communication connection between the screen mirroring device 102 and the display device 103, and the connection method can be, for example, a data cable connection. The screen mirroring device 102 can mirror the video played on the terminal device 101 to the display device 103 for display.
[0024] The terminal device 101 can be various electronic devices including a display screen, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on. The display device 103 can also be various electronic devices including a display screen, including but not limited to laptop computers, desktop computers, and televisions, and so on. Generally, the size of the display screen of the display device 103 is larger than that of the terminal device 101.
[0025] It should be noted that the image recognition method provided by the embodiments of the present disclosure is generally executed by the screen mirroring device 102. Correspondingly, the image recognition device is generally disposed in the screen mirroring device 102.
[0026] It should be understood that Figure 1 the numbers of the terminal device, the screen mirroring device, and the display device in
[0027] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, screen mirroring devices, and display devices. Figure 2 Continuing to refer to
[0028] Step 201: In response to detecting an identification operation during screen mirroring, obtain the image to be identified of the screen mirroring device.
[0029] In this embodiment, the screen mirroring device can mirror the video played on the terminal device to the display screen for display. If an identification operation is detected during screen mirroring, the screen mirroring device can obtain the image to be identified.
[0030] Here, the screen mirroring device is equipped with image recognition capabilities and is also called an AI (Artificial Intelligence) screen mirroring device. The appearance of the screen mirroring device is not significantly different from that of a traditional screen mirroring device, and it is composed of a hardware main body, a power supply cable, and a video signal output cable. The modules included in the screen mirroring device mainly include a screen mirroring module, a network communication module, an AI recognition module, and a power supply module. Different from the traditional screen mirroring device, a physical image recognition button can be set on the screen mirroring device. When the user sees an object on the video screen displayed on the display screen that they want to identify, they can press this physical image recognition button. At this time, the screen mirroring device detects the user's identification operation. In addition, a virtual image recognition button can also be set on the screen mirroring application installed on the terminal device. When the user sees an object on the video screen displayed on the display screen that they want to identify, they can click this virtual image recognition button. At this time, the terminal device can send an identification instruction to the screen mirroring device. When receiving the identification instruction, the screen mirroring device detects the user's identification operation. In addition, the user can also initiate an identification instruction by performing specific operations on the terminal device. Specific operations can include, but are not limited to, shaking, specific gestures, and so on.
[0031] Here, the image to be identified can be an image in the current screen mirroring video segment of the screen mirroring device. The screen mirroring device can select a frame of image from the current screen mirroring video segment as the image to be identified. Among them, the current screen mirroring video segment can be the video segment mirrored within a preset number of milliseconds before and after the moment when the identification operation is detected. For example, obtain the current screen mirroring image of the screen mirroring device as the image to be identified, so as to quickly locate the image to be identified. Among them, the current screen mirroring image can be the image mirrored at the moment when the identification operation is detected.
[0032] Step 202: Obtain the recognition result of the object in the image to be identified.
[0033] In this embodiment, the screen mirroring device can obtain the recognition result of the object in the image to be identified.
[0034] Generally, the image to be recognized may contain more than one object. In the case where the image to be recognized contains multiple objects, at least some of the objects can be recognized. For example, all the objects in the image to be recognized are recognized. For another example, only the objects of a specific category in the image to be recognized are recognized. Among them, the specific category can be preset, including but not limited to at least one of people, animals, items, buildings, etc. The recognition result of an object is usually the name of the object. For example, when the object is a celebrity, its recognition result can be the name of the celebrity. For another example, when the object is a car, its recognition result can be the brand and model of the car.
[0035] Step 203, superimpose the recognition result on the current screen-casting video segment of the screen-caster and cast it onto the display screen connected to the screen-caster.
[0036] In this embodiment, the screen-caster can superimpose the recognition result of the object in the image to be recognized on the current screen-casting video segment of the screen-caster and cast it onto the display screen linked to the screen-caster.
[0037] Generally, the screen-caster can first determine the position of the object on the current screen-casting video segment, and then superimpose the recognition result of this object near the position of the object. In this way, when the user watches the video on the display screen, they can conveniently obtain the recognition result of this object. In practical applications, the recognition result of the object will not always be displayed on the current playing video segment. For example, if this object no longer appears on the current playing video segment, its recognition result automatically disappears. For another example, if the display duration of the recognition result of this object exceeds a preset duration (such as 5 seconds), its recognition result automatically disappears.
[0038] In some optional implementation manners of this embodiment, in order to enable the user to obtain richer information. When the screen-caster superimposes the recognition result of the object, it can also superimpose the associated information of the object. Specifically, first, based on the recognition result, obtain the associated information of the object in the image to be recognized; then superimpose the associated information on the current screen-casting video segment of the screen-caster. Among them, the associated information can be information related to the object, including but not limited to the introduction information, promotional information, purchase link, etc. of the object. For example, when the object is a celebrity, its associated information can be the encyclopedia information of the celebrity. For another example, when the object is a car, its associated information can be the promotional video of the car.
[0039] The image recognition method provided by the embodiments of the present disclosure adds image recognition capabilities to the screen-caster, enabling the user to use the image recognition capabilities to recognize any content played through the screen-caster during the screen-casting process, thereby presenting richer content information to the user.
[0040] Further refer to Figure 3 , Figure 3Flow 300 of the second embodiment of the image recognition method according to the present disclosure is shown. The image recognition method includes the following steps:
[0041] Step 301, in response to detecting an identification operation during the screen mirroring process, obtain the current screen mirroring image of the screen mirroring device, as well as a preset number of frames of images before and after the current screen mirroring image, to obtain an image set.
[0042] In this embodiment, the screen mirroring device can mirror the video played on the terminal device to the display screen for display. If an identification operation is detected during the screen mirroring process, the screen mirroring device can obtain the current screen mirroring image, as well as a preset number of frames of images before and after the current screen mirroring image, to obtain an image set.
[0043] Step 302, select an image to be recognized from the image set.
[0044] In this embodiment, the screen mirroring device can select an image to be recognized from the image set.
[0045] Generally, a video has 24 frames per second, and there is a certain delay between the moment when the user sees the video frame to be recognized and the moment when the recognition operation is performed. Here, the image to be recognized is selected from the current screen mirroring image and a preset number of frames of images before and after it, rather than directly using the current screen mirroring image as the image to be recognized. This can eliminate the influence of the delay and select a more suitable video frame.
[0046] In some optional implementation manners of this embodiment, the screen mirroring device can obtain the clarity of the images in the image set; based on the clarity, select an image to be recognized from the image set. Thus, an image to be recognized with high clarity can be selected, and further improve the recognition accuracy.
[0047] Step 303, obtain the recognition result of the object in the image to be recognized.
[0048] Step 304, superimpose the recognition result on the current screen mirroring video segment of the screen mirroring device, and mirror it to the display screen connected to the screen mirroring device.
[0049] In this embodiment, the specific operations of steps 303-304 have been introduced in detail in steps 202-203 of the Figure 2 shown embodiment, and will not be elaborated here.
[0050] From Figure 3 it can be seen that compared with Figure 2Compared with the corresponding embodiment, the image recognition method in this embodiment emphasizes the step of selecting the image to be recognized. Thus, the solution described in this embodiment takes into account the delay between the moment when the user sees the video frame to be recognized and the moment when the recognition operation is performed, and selects the image to be recognized from the current screen-cast image and a preset number of frames of images before and after it, which can eliminate the influence of the delay and select a more appropriate video frame.
[0051] Further referring to Figure 4 , Figure 4 FIG. 400 shows a flowchart of a third embodiment of the image recognition method according to the present disclosure. The image recognition method includes the following steps:
[0052] Step 401, in response to detecting a recognition operation during the screen-casting process, obtain the current screen-cast image of the screen-caster and a preset number of frames of images before and after the current screen-cast image to obtain an image set.
[0053] In this embodiment, the specific operation of step 401 has been described in detail in step 301 of the embodiment shown in Figure 3 and will not be elaborated here.
[0054] Step 402, obtain the main category of the images in the image set.
[0055] In this embodiment, for each frame of image in the image set, the screen-caster can obtain the main category of the image. Among them, the main category can be the category of the object in the image, including but not limited to people, animals, items, buildings, etc.
[0056] In some optional implementation manners of this embodiment, the screen-caster can input the images in the image set into a pre-trained main object recognition model to obtain the main category of the images in the image set. Using the main object recognition model to recognize the main category of the image improves the recognition accuracy and efficiency. Among them, the main object recognition model can be used to recognize the main category of the image and is pre-trained by using a training sample set through a deep learning method. Here, the training samples in the training sample set can be sample images labeled with the main category.
[0057] Step 403, select the image to be recognized from the image set based on the main category priority.
[0058] In this embodiment, the screen-caster can select the image to be recognized from the image set based on the main category priority. Usually, the screen-caster will select the image with a higher main category priority. For example, the priority of people is higher than that of animals, the priority of animals is higher than that of items, and the priority of items is higher than that of buildings. The screen-caster will preferentially select the image with people.
[0059] Step 404, obtain the recognition result of the object in the image to be recognized.
[0060] Step 405: superimpose the recognition result on the current screen-cast video segment of the screen-casting device, and then cast it onto the display screen connected to the screen-casting device.
[0061] In this embodiment, the specific operations of steps 404-405 have been described in detail in steps 303-304 of the embodiment shown in Figure 3 and will not be elaborated here.
[0062] As can be seen from Figure 4 compared with the corresponding embodiment, the image recognition method in this embodiment highlights the step of selecting the image to be recognized. Therefore, the solution described in this embodiment can select a video frame containing an object more suitable for recognition based on the priority of the subject category. Figure 3
[0063] Further referring to Figure 5 , Figure 5 shows a flowchart 500 of a fourth embodiment of the image recognition method according to the present disclosure. The image recognition method includes the following steps:
[0064] Step 501: In response to detecting a recognition operation during the screen-casting process, obtain the image to be recognized of the screen-casting device.
[0065] In this embodiment, the specific operation of step 501 has been described in detail in step 201 of the embodiment shown in Figure 2 and will not be elaborated here.
[0066] Step 502: Input the image to be recognized into a pre-trained image recognition model to obtain a recognition result.
[0067] In this embodiment, the screen-casting device can input the image to be recognized into a pre-trained image recognition model to obtain a recognition result.
[0068] Generally, when the computing power of the screen-casting device is sufficient, an image recognition model can be stored thereon to achieve local image recognition. There is no need to transmit the image and the recognition result, thereby improving the image recognition efficiency. The image recognition model can be used to recognize objects in an image and is pre-trained by using a training sample set through a deep learning method. Here, the training samples in the training sample set can be sample images labeled with object names.
[0069] Step 503: superimpose the recognition result on the current screen-cast video segment of the screen-casting device, and then cast it onto the display screen connected to the screen-casting device.
[0070] In this embodiment, the specific operation of step 503 has been described in detail in step 203 of the embodiment shown in Figure 2 and will not be elaborated here.
[0071] As can be seen from Figure 5 , compared with the corresponding embodiment Figure 2 , in this embodiment, the image recognition steps are highlighted in the image recognition method. Thus, when the computing power of the screen mirroring device is sufficient, an image recognition model can be stored thereon, so as to realize local image recognition. There is no need to transmit the image and the recognition result, thereby improving the image recognition efficiency.
[0072] Further referring to Figure 6 , Figure 6 shows a flowchart 600 of a fifth embodiment of the image recognition method according to the present disclosure. The image recognition method includes the following steps:
[0073] Step 601, in response to detecting a recognition operation during screen mirroring, obtain the image to be recognized of the screen mirroring device.
[0074] In this embodiment, the specific operation of step 601 has been described in detail in step 201 of the embodiment Figure 2 shown, and will not be elaborated here.
[0075] Step 602, send the image to be recognized to the cloud.
[0076] In this embodiment, the screen mirroring device can send the image to be recognized to the cloud. The cloud can input the image to be recognized into a pre-trained image recognition model to obtain the recognition result of the object in the image to be recognized.
[0077] Generally, the cloud can store the image recognition model to realize image recognition in the cloud. There is no need for local image recognition on the screen mirroring device, thereby reducing the computing power requirement for the screen mirroring device and further reducing the cost of the screen mirroring device. The image recognition model can be used to recognize objects in images and is pre-trained using a training sample set by a deep learning method. Here, the training samples in the training sample set can be sample images labeled with object names.
[0078] Step 603, receive the recognition result sent by the cloud.
[0079] In this embodiment, the cloud can send the recognition result of the object in the image to be recognized to the screen mirroring device. In this way, the screen mirroring device receives the recognition result of the object in the image to be recognized.
[0080] Step 604, superimpose the recognition result on the current screen mirroring video segment of the screen mirroring device and screen mirror it to the display screen connected to the screen mirroring device.
[0081] In this embodiment, the specific operation of step 604 has been described in detail in step 203 of the embodiment Figure 2 shown, and will not be elaborated here.
[0082] FromFigure 6 It can be seen that compared with the corresponding embodiments, the image recognition method in this embodiment highlights the image recognition steps. Thus, the solution described in this embodiment performs image recognition in the cloud without the need for local image recognition by the wireless display device, thereby reducing the computing power requirements for the wireless display device and further reducing the cost of the wireless display device. Figure 2 As shown in FIG. 2, compared with the corresponding embodiments, the image recognition method in this embodiment highlights the image recognition steps. Thus, the solution described in this embodiment performs image recognition in the cloud without the need for local image recognition by the wireless display device, thereby reducing the computing power requirements for the wireless display device and further reducing the cost of the wireless display device.
[0083] Further referring to Figure 7 FIG. 3, as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image recognition apparatus. The apparatus embodiment corresponds to the method embodiment shown in Figure 2 FIG. 3, and the apparatus can be specifically applied to various electronic devices.
[0084] As shown in Figure 7 FIG. 4, the image recognition apparatus 700 in this embodiment may include: a first acquisition module 701, a second acquisition module 702, and a first superimposition module 703. Among them, the first acquisition module 701 is configured to acquire an image to be recognized of the wireless display device in response to detecting a recognition operation during the screen mirroring process; the second acquisition module 702 is configured to acquire a recognition result of an object in the image to be recognized; the first superimposition module 703 is configured to superimpose the recognition result on the current screen mirroring video segment of the wireless display device and project it onto a display screen connected to the wireless display device.
[0085] In this embodiment, for the image recognition apparatus 700: the specific processing of the first acquisition module 701, the second acquisition module 702, and the first superimposition module 703 and the technical effects brought by them can be respectively referred to the relevant descriptions of steps 201-203 in the corresponding embodiments, which will not be elaborated here. Figure 2 For the image recognition apparatus 700 in this embodiment, the specific processing of the first acquisition module 701, the second acquisition module 702, and the first superimposition module 703 and the technical effects brought by them can be respectively referred to the relevant descriptions of steps 201-203 in the corresponding embodiments, which will not be elaborated here.
[0086] In some alternative implementation manners of this embodiment, the first acquisition module 701 is further configured to: acquire the current screen mirroring image of the wireless display device as the image to be recognized.
[0087] In some alternative implementation manners of this embodiment, the first acquisition module 701 includes: an acquisition sub-module configured to acquire the current screen mirroring image of the wireless display device and a preset number of frames of images before and after the current screen mirroring image to obtain an image set; a selection sub-module configured to select the image to be recognized from the image set.
[0088] In some alternative implementation manners of this embodiment, the selection sub-module is further configured to: acquire the clarity of the images in the image set; and select the image to be recognized from the image set based on the clarity.
[0089] In some alternative implementation manners of this embodiment, the selection sub-module includes: an obtaining unit configured to obtain the main category of the images in the image set; a selection unit configured to select the images to be recognized from the image set based on the main category priority.
[0090] In some alternative implementation manners of this embodiment, the obtaining unit is further configured to: input the images in the image set into a pre-trained main body recognition model to obtain the main category of the images in the image set.
[0091] In some alternative implementation manners of this embodiment, the second obtaining module 702 is further configured to: input the image to be recognized into a pre-trained image recognition model to obtain a recognition result.
[0092] In some alternative implementation manners of this embodiment, the second obtaining module 702 is further configured to: send the image to be recognized to the cloud; receive the recognition result sent by the cloud.
[0093] In some alternative implementation manners of this embodiment, the image recognition device 700 further includes: a third obtaining module configured to obtain the association information of the object in the image to be recognized based on the recognition result; a second superimposing module configured to superimpose the association information on the current projection video segment of the projector.
[0094] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0095] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0096] Figure 8 A schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0097] Such as Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 802 or computer programs loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of device 800 can also be stored. The computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0098] Multiple components in device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, optical disk, etc.; and a communication unit 809, such as a network card, modem, wireless communication transceiver, etc. The communication unit 809 allows device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0099] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as an image recognition method. For example, in some embodiments, the image recognition method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the image recognition method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the image recognition method by any other appropriate means (e.g., by means of firmware).
[0100] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0101] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0102] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0103] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0104] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0105] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0106] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided in this disclosure can be achieved, and no limitation is imposed herein.
[0107] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. An image recognition method, including: In response to detecting an identification operation during screen mirroring, obtaining an image to be recognized of the screen mirroring device; Obtaining an identification result of an object in the image to be recognized; Overlaying the identification result on the current screen mirroring video segment of the screen mirroring device and projecting it onto a display screen connected to the screen mirroring device; wherein, obtaining the image to be recognized of the screen mirroring device includes: Obtaining the current screen mirroring image of the screen mirroring device and a preset number of frames of images before and after the current screen mirroring image to obtain an image set; Selecting the image to be recognized from the image set; wherein, selecting the image to be recognized from the image set includes: Obtaining the main category of the images in the image set; Selecting the image to be recognized from the image set based on the priority of the main category.
2. The method according to claim 1, wherein, obtaining the image to be recognized of the screen mirroring device includes: Obtaining the current screen mirroring image of the screen mirroring device as the image to be recognized.
3. The method according to claim 1, wherein, selecting the image to be recognized from the image set includes: Obtaining the clarity of the images in the image set; Selecting the image to be recognized from the image set based on the clarity.
4. The method according to claim 1, wherein, obtaining the main category of the images in the image set includes: Inputting the images in the image set into a pre-trained main body recognition model to obtain the main category of the images in the image set.
5. The method according to claim 1, wherein, obtaining the identification result of the object in the image to be recognized includes: Inputting the image to be recognized into a pre-trained image recognition model to obtain the identification result.
6. The method according to claim 1, wherein, obtaining the identification result of the object in the image to be recognized includes: Sending the image to be recognized to the cloud; Receiving the identification result sent by the cloud.
7. The method according to any one of claims 1-6, wherein, after overlaying the identification result on the current screen mirroring video segment of the screen mirroring device, further including: Based on the identification result, obtaining the associated information of the object in the image to be recognized; Overlaying the associated information on the current screen mirroring video segment of the screen mirroring device.
8. An image recognition device, including: A first acquisition module configured to obtain an image to be recognized of the screen mirroring device in response to detecting an identification operation during screen mirroring; A second acquisition module configured to obtain an identification result of an object in the image to be recognized; A first overlay module configured to overlay the identification result on the current screen mirroring video segment of the screen mirroring device and project it onto a display screen connected to the screen mirroring device; wherein, the first acquisition module includes: An acquisition sub-module configured to obtain the current screen mirroring image of the screen mirroring device and a preset number of frames of images before and after the current screen mirroring image to obtain an image set; A selection sub-module configured to select the image to be recognized from the image set; wherein, the selection sub-module includes: An acquisition unit, configured to acquire the main category of the images in the image set; A selection unit, configured to select the image to be recognized from the image set based on the main category priority.
9. The apparatus according to claim 8, wherein, the first acquisition module is further configured to: acquire the current screen-cast image of the screen-casting device as the image to be recognized.
10. The apparatus according to claim 8, wherein, the selection sub-module is further configured to: acquire the clarity of the images in the image set; select the image to be recognized from the image set based on the clarity.
11. The apparatus according to claim 8, wherein, the acquisition unit is further configured to: input the images in the image set into a pre-trained main body recognition model to obtain the main category of the images in the image set.
12. The apparatus according to claim 8, wherein, the second acquisition module is further configured to: input the image to be recognized into a pre-trained image recognition model to obtain the recognition result.
13. The apparatus according to claim 8, wherein, the second acquisition module is further configured to: send the image to be recognized to the cloud; receive the recognition result sent by the cloud.
14. The apparatus according to any one of claims 8-13, wherein, the apparatus further includes: a third acquisition module, configured to acquire the association information of the object in the image to be recognized based on the recognition result; a second superimposing module, configured to superimpose the association information onto the current screen-cast video segment of the screen-casting device.
15. A server, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-7.
16. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method according to any one of claims 1-7.
17. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Device and method for recognizing garbage types based on images
CN111582336A
Image recognition method and device, equipment, storage medium and program product
CN113239890A