An image recognition method, server, and display device
By using voiceprint-assisted recognition technology, combined with media asset databases and preset face databases, the problem of environmental influences on celebrity face recognition in smart terminals has been solved, improving recognition accuracy and recall rate, and enhancing user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 青岛聚看云科技有限公司
- Filing Date
- 2021-08-02
- Publication Date
- 2026-05-26
Smart Images

Figure CN115690866B_ABST
Abstract
Description
[0001] This application claims priority to Chinese Patent Application No. 202110836627.0, filed on July 23, 2021, entitled "An Image Recognition Method, Server and Display Device", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of display technology, and in particular to an image recognition method, server, and display device. Background Technology
[0003] A smart terminal is a television product that enables two-way human-computer interaction and integrates multiple functions such as audio-visual entertainment and data. To meet the diverse needs of users, operators are committed to developing various convenient functions to enhance the user experience of smart terminals.
[0004] Currently, a crucial application of image retrieval on smart terminals is celebrity facial recognition. For instance, users can take screenshots of videos displayed on their smart terminals, which then upload the screenshots to a server for identification. However, the celebrity facial recognition function on smart terminals is affected by factors such as ambient lighting, actor makeup, facial angles, and facial expressions during video playback. This can lead to inaccurate facial feature extraction, resulting in failure to recognize the person or incorrect identification, thus reducing the accuracy of facial recognition and impacting the user experience. Summary of the Invention
[0005] This application provides an image recognition method, server, and display device, which can be used to solve the technical problem of reduced user experience due to low accuracy of face recognition in the prior art.
[0006] To address the aforementioned technical problems, the embodiments of this application disclose the following technical solutions:
[0007] In a first aspect, some embodiments of this application disclose a server, the server comprising a media asset database, a preset face database, and a recognition processor, wherein the media asset database is used to store audio track voiceprint data of videos, and the recognition processor is configured to:
[0008] The device receives an image recognition request sent by a display device. The image recognition request includes a target image, a video identifier, and a playback time point. The target image is a screenshot of the video corresponding to the video identifier that is being played by the display device. The playback time point is the moment when the video is playing when the screenshot is taken.
[0009] The feature matching confidence score is obtained by comparing the face data corresponding to the person object and the target person name in the target image.
[0010] If the feature matching confidence of the target person's name is greater than the first preset threshold, the person information corresponding to the target person's name will be directly output as the recognition result.
[0011] If the feature matching confidence of the target person's name is not greater than a first preset threshold, the person information corresponding to the target person's name is output as the recognition result when the target person's name is in the person list, based on the video identifier, the playback time point, and the list of people determined by the media asset database; otherwise, an unrecognized identifier is output. The media asset database stores the person time correspondence between people in the video and the time when the person's voiceprint appears.
[0012] As can be seen from the above technical solutions, the first aspect of this application provides a server, which includes a media asset database, a preset face database, and a recognition processor. The recognition processor is configured to: receive an image recognition request sent by a display device, the image recognition request including a target image, a video identifier, and a playback time point; obtain the feature matching confidence level of the face data corresponding to the person object in the target image and the target person's name, wherein the feature matching confidence level is obtained by comparing the person object in the target image with the preset face database; if the feature matching confidence level of the target person's name is greater than a first preset threshold, directly add the corresponding face data to the target face database. The system outputs the person information of the target person as the recognition result; if the feature matching confidence of the target person is not greater than a first preset threshold, the system outputs the person information corresponding to the target person as the recognition result when the target person is in the list of persons determined by the video identifier, the playback time point, and the media asset database; otherwise, it outputs an unrecognized identifier; wherein, the media asset database stores the person time correspondence between the person in the video and the time when the person's voiceprint appears; voiceprint-assisted recognition improves the accuracy and recall rate of person object recognition, thereby enhancing the user experience.
[0013] Secondly, in some embodiments of this application, a display device is disclosed, including a display and a controller; the display is used to display a user interface; the controller is configured to:
[0014] Receive image recognition request;
[0015] In response to the image recognition request, during video playback, the target image, video identifier, and playback time point of the user interface are determined to identify the human object in the target image;
[0016] Display the recognition results of the person / object.
[0017] As can be seen from the above technical solutions, the second aspect of this application provides a display device, including a display and a controller; the display is used to display a user interface; the controller is configured to receive an image recognition request; in response to the image recognition request, during video playback, the controller determines the target image, video identifier, and playback time point of the user interface to identify a person in the target image; the controller displays the recognition result of the person; and the controller improves the accuracy and recall rate of person recognition through voiceprint-assisted recognition, thereby enhancing the user experience.
[0018] Thirdly, in some embodiments of this application, a display device is disclosed, including a display and a controller; the display is used to display a user interface; the controller is configured to:
[0019] Receive image recognition request;
[0020] In response to the image recognition request, when not playing a video, the target image of the user interface is determined to identify the human figures in the target image;
[0021] Display the recognition results of the person / object.
[0022] As can be seen from the above technical solutions, the third aspect of this application provides a display device, including a display and a controller; the display is used to display a user interface; the controller is configured to receive an image recognition request; in response to the image recognition request, when not playing video, determine a target image of the user interface to identify a person in the target image; display the recognition result of the person, and identify the person in the target image through the target image when not playing video.
[0023] Fourthly, in some embodiments of this application, an image recognition method is disclosed, the image recognition method comprising the following steps:
[0024] The device receives an image recognition request sent by a display device. The image recognition request includes a target image, a video identifier, and a playback time point. The target image is a screenshot of the video corresponding to the video identifier that is being played by the display device. The playback time point is the moment when the video is playing when the screenshot is taken.
[0025] The feature matching confidence score is obtained by comparing the face data corresponding to the person object and the target person name in the target image.
[0026] If the feature matching confidence of the target person's name is greater than the first preset threshold, the person information corresponding to the target person's name will be directly output as the recognition result.
[0027] If the feature matching confidence of the target person's name is not greater than a first preset threshold, the person information corresponding to the target person's name is output as the recognition result when the target person's name is in the person list, based on the video identifier, the playback time point, and the list of people determined by the media asset database; otherwise, an unrecognized identifier is output. The media asset database stores the person time correspondence between people in the video and the time when the person's voiceprint appears.
[0028] As can be seen from the above technical solutions, the third aspect of this application provides an image recognition method, comprising: receiving an image recognition request sent by a display device, the image recognition request including a target image, a video identifier, and a playback time point, wherein the target image is obtained by taking a screenshot of a video corresponding to the video identifier played by the display device, and the playback time point is the time when the screenshot is taken; obtaining the feature matching confidence level of the face data corresponding to the person object in the target image and the target person name, wherein the feature matching confidence level is obtained by comparing the person object in the target image with a preset face database; if the feature matching confidence level of the target person name is large... At a first preset threshold, the person information corresponding to the target name is directly output as the recognition result; if the feature matching confidence of the target name is not greater than the first preset threshold, the person information corresponding to the target name is output as the recognition result when the target name is in the person list determined by the video identifier, the playback time point, and the media asset database; otherwise, an unrecognized identifier is output; wherein, the media asset database stores the person time correspondence between the person in the video and the time when the person's voiceprint appears; voiceprint-assisted recognition improves the accuracy and recall rate of person object recognition, thereby enhancing the user experience. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a schematic diagram illustrating the operational scenario between the display device and the control device in an embodiment of this application;
[0032] Figure 2 This is a hardware configuration block diagram of the control device 100 in the embodiments of this application;
[0033] Figure 3 This is a hardware configuration block diagram of the display device 200 in this embodiment of the application;
[0034] Figure 4 This is a schematic diagram of the software configuration of the display device 200 in an embodiment of this application;
[0035] Figure 5 This is a schematic diagram of the device homepage displayed in an embodiment of this application;
[0036] Figure 6 This is a flowchart of an image recognition method according to an embodiment of this application;
[0037] Figures 7a-7b This is a schematic diagram of a user interface in an embodiment of this application;
[0038] Figure 8 This is a flowchart of yet another image recognition method in the embodiments of this application. Detailed Implementation
[0039] To make the objectives and implementation methods of this application clearer, the exemplary implementation methods of this application will be clearly and completely described below with reference to the accompanying drawings of the exemplary embodiments of this application. Obviously, the exemplary embodiments described are only some embodiments of this application, and not all embodiments.
[0040] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0041] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.
[0042] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.
[0043] The term "module" refers to any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0044] Figure 1 This is a schematic diagram illustrating the operational scenario between the display device and the control unit according to the embodiment. Figure 1 As shown, the user can operate the display device 200 through the smart device 300 or the control device 100.
[0045] In some embodiments, the control device 100 may be a remote control. Communication between the remote control and the display device includes infrared protocol communication, Bluetooth protocol communication, and other short-range communication methods, controlling the display device 200 wirelessly or via wired means. Users can control the display device 200 by inputting user commands through buttons on the remote control, voice input, control panel input, etc.
[0046] In some embodiments, a smart device 300 (such as a mobile terminal, tablet computer, computer, laptop computer, etc.) may also be used to control the display device 200. For example, an application running on the smart device may be used to control the display device 200.
[0047] In some embodiments, the display device 200 can also be controlled in ways other than the control device 100 and the smart device 300. For example, it can be controlled by directly receiving the user's voice commands through a module configured inside the display device 200 for acquiring voice commands, or it can be controlled by receiving the user's voice commands through a voice control device set outside the display device 200.
[0048] In some embodiments, the display device 200 also communicates with the server 400. The display device 200 may communicate via a local area network (LAN), wireless local area network (WLAN), and other networks. The server 400 may provide various content and interactive features to the display device 200. The server 400 may be a cluster or multiple clusters, and may include one or more types of servers.
[0049] Figure 2 An exemplary block diagram of the configuration of the control device 100 according to an exemplary embodiment is shown. Figure 2 As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input operation commands and convert the operation commands into commands that the display device 200 can recognize and respond to, thus acting as an intermediary for interaction between the user and the display device 200.
[0050] Figure 3A hardware configuration block diagram of a display device 200 according to an exemplary embodiment is shown.
[0051] In some embodiments, the display device 200 includes at least one of a tuner 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface.
[0052] In some embodiments, the controller includes a processor, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first to an nth interface for input / output.
[0053] In some embodiments, the display 260 includes a display screen component for presenting an image, a driving component for driving image display, a component for receiving image signals from the controller output, and a user control UI interface for displaying video content, image content, menu control interface, and user control UI interface.
[0054] In some embodiments, the display 260 may be a liquid crystal display, an OLED display, or a projection display, and may also be a projection device and a projection screen.
[0055] In some embodiments, the communicator 220 is a component used to communicate with external devices or servers according to various communication protocol types. For example, the communicator may include at least one of a Wi-Fi module, a Bluetooth module, a wired Ethernet module, other network communication protocol chips or near-field communication protocol chips, and an infrared receiver. The display device 200 can establish the transmission and reception of control signals and data signals with the external control device 100 or the server 400 through the communicator 220.
[0056] In some embodiments, the user interface can be used to receive control signals from the control device 100 (e.g., an infrared remote control).
[0057] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.
[0058] In some embodiments, the external device interface 240 may include, but is not limited to, one or more interfaces such as: High Definition Multimedia Interface (HDMI), analog or data high-definition component input interface (component), composite video input interface (CVBS), USB input interface (USB), RGB port, etc. It may also be a composite input / output interface formed by multiple interfaces mentioned above.
[0059] In some embodiments, the tuner 210 receives broadcast television signals via wired or wireless reception and demodulates audio and video signals, such as EPG data signals, from a plurality of wireless or wired broadcast television signals.
[0060] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0061] In some embodiments, the controller 250 controls the operation of the display device and responds to user operations via various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200. For example, in response to receiving a user command to select a UI object to display on the display 260, the controller 250 can perform operations related to the object selected by the user command.
[0062] In some embodiments, the object can be any of the optional objects, such as a hyperlink, an icon, or other operable controls. Operations related to the selected object include: displaying links to hyperlinked pages, documents, images, etc., or performing operations corresponding to the program associated with the icon.
[0063] In some embodiments, the controller includes at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processing unit (GPU), RAM (random access memory), ROM (read-only memory), a first to an nth interface for input / output, a communication bus, etc.
[0064] A CPU (CPU) processor is used to execute operating system and application instructions stored in memory, as well as various interactive instructions received from external input, to execute various applications, data, and content, ultimately for the display and playback of various audio and video content. A CPU processor can include multiple processors, such as a main processor and one or more sub-processors.
[0065] In some embodiments, a graphics processor is used to generate various graphical objects, such as icons, operation menus, and graphics displayed based on user input commands. The graphics processor includes an arithmetic logic unit (ALU) that performs calculations based on various user-input interactive commands and displays various objects according to display attributes; it also includes a renderer that renders the various objects obtained from the ALU, and the rendered objects are used to display on a monitor.
[0066] In some embodiments, the video processor is configured to receive external video signals and perform video processing such as decompression, decoding, scaling, noise reduction, frame rate conversion, resolution conversion, and image synthesis according to the standard encoding and decoding protocol of the input signals, so as to obtain a signal that can be directly displayed or played on the display device 200.
[0067] In some embodiments, the video processor includes a demultiplexing module, a video decoding module, an image compositing module, a frame rate conversion module, and a display formatting module. The demultiplexing module demultiplexes the input audio and video data streams. The video decoding module processes the demultiplexed video signal, including decoding and scaling. The image compositing module, such as an image synthesizer, overlays and blends a GUI signal generated by a graphics generator based on user input or its own generation with the scaled video image to generate a displayable image signal. The frame rate conversion module converts the input video frame rate. The display formatting module modifies the received frame rate-converted video output signal to conform to a display format, such as outputting RGB data signals.
[0068] In some embodiments, the audio processor is configured to receive external audio signals, and according to the standard codec protocol of the input signals, perform decompression and decoding, as well as noise reduction, digital-to-analog conversion, and amplification processing, to obtain a sound signal that can be played in a speaker.
[0069] In some embodiments, the user can input user commands through a graphical user interface (GUI) displayed on the display 260, and the user input interface receives the user input commands through the GUI. Alternatively, the user can input user commands by inputting specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.
[0070] In some embodiments, a "user interface" is the medium through which an application or operating system interacts and exchanges information with a user, converting information between its internal form and a form acceptable to the user. A common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an icon, window, control, or other interface element displayed on the screen of an electronic device. Controls can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.
[0071] In some embodiments, the display device's system may include a kernel, a command interpreter (shell), a file system, and applications. The kernel, shell, and file system together form the basic operating system structure, allowing users to manage files, run programs, and use the system. Upon power-up, the kernel starts, activates the kernel space, abstracts hardware, initializes hardware parameters, and runs and maintains virtual memory, the scheduler, signals, and inter-process communication (IPC). After the kernel starts, the shell and user applications are loaded. Applications are compiled into machine code after startup, forming a process.
[0072] See Figure 4 In some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the Android runtime and system library layer (referred to as the "System Runtime Layer"), and the kernel layer.
[0073] In some embodiments, at least one application runs in the application layer. These applications may be Windows programs, system settings programs, or clock programs that come with the operating system; they may also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0074] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.
[0075] like Figure 4As shown, the application framework layer in this embodiment includes managers, content providers, etc., wherein the managers include at least one of the following modules: ActivityManager, which interacts with all activities running in the system; LocationManager, which provides access to system location services for system services or applications; PackageManager, which retrieves various information related to application packages currently installed on the device; NotificationManager, which controls the display and clearing of notification messages; and WindowManager, which manages icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0076] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling display window changes (e.g., shrinking the display window, shaking the display, distorting the display, etc.).
[0077] In some embodiments, the system runtime library layer provides support for the upper layer, namely the framework layer. When the framework layer is used, the Android operating system runs the C / C++ libraries contained in the system runtime library layer to implement the functions that the framework layer needs to perform.
[0078] In some embodiments, the kernel layer is a layer between hardware and software. For example... Figure 4 As shown, the kernel layer includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver.
[0079] In some embodiments, when displaying video content, a user can input image recognition commands to control the display device 200 to capture part or all of the currently displayed content. For example, while watching a video program, a user can input image recognition commands to control the display device to take a screenshot of the currently displayed video frame, obtain a target image, and identify objects in the target image.
[0080] A screenshot refers to capturing a portion or all of the content currently displayed on a monitor to obtain a target image. User input used to trigger a screenshot can be in the form of keystrokes, voice input, or gestures.
[0081] In some embodiments, user input that triggers a screenshot can also enable facial recognition of celebrities in a video, or user input that directly triggers image recognition can enable facial recognition of celebrities in a video. However, in actual applications, video playback is affected by factors such as scene lighting, actors' makeup, facial angles, and facial expressions, which reduces the accuracy and recall of object recognition in the target image determined by the screenshot.
[0082] In some embodiments, the same person may have different makeup in different videos or even within the same video; similarly, the same person may have different makeup in everyday photos and stills from videos. The scenes in videos can be bright or dark for the sake of the plot. Celebrities may also make exaggerated expressions in videos, or the angle of the object in the target image may be too large when taking screenshots. A combination of one or more of the above situations can lead to discrepancies in facial feature extraction, resulting in failure to recognize the face or incorrect recognition.
[0083] To address the aforementioned problems, some embodiments of this application provide a server, which includes a media asset database, a preset face database, and a recognition processor. The recognition processor is configured to receive an image recognition request sent by a display device. The image recognition request includes a target image, a video identifier, and a playback time point. The target image is a screenshot of a video corresponding to the video identifier played on the display device, and the playback time point is the moment the video was playing when the screenshot was taken. The processor obtains the feature matching confidence score of the face data corresponding to the person in the target image and the target person's name, wherein the feature matching confidence score is obtained by comparing the person in the target image with the preset face database.
[0084] In some embodiments, if the feature matching confidence of the target person's name is greater than a first preset threshold, the person information corresponding to the target person's name is directly output as the recognition result; if the feature matching confidence of the target person's name is not greater than the first preset threshold, based on the video identifier, the playback time point, and the person list determined by the media asset database, and when the target person's name is in the person list, the person information corresponding to the target person's name is output as the recognition result; otherwise, an unrecognized identifier is output; wherein, the media asset database stores the person-time correspondence between people in the video and the time when their voiceprints appear. For the recognition of objects in the target image, facial recognition or voiceprint recognition can be used to improve the accuracy and recall rate of recognition.
[0085] In some trial cases, such as Figure 5As shown, this is the homepage of the display device. If the task object to be identified is further identified through voiceprint recognition, it can only be done in a video playback scenario. Therefore, in non-video playback scenarios, the image recognition request only contains the target image and does not include video identifiers and playback time points. The server can determine whether the scene to be identified is a video playback scenario or a non-video playback scenario based on the specific information contained in the received image recognition request. Alternatively, the server can add information such as whether it is a video playback scenario or whether voiceprint recognition is possible to the image recognition request, so as to provide the server with a choice of image recognition methods.
[0086] In one embodiment, for Figure 5 The homepage shown can identify people in screenshots by sending an image recognition request. It can send an image recognition request for the target image only. The service area can determine the people and their information in the homepage screenshot based on the target image. It can present the name and similarity score, or it can present the result as "looks like so-and-so" and the corresponding similarity score.
[0087] In one embodiment, the server may be Elastic Search (ES), a Lucene-based search server that provides a distributed, multi-user, full-text search engine that can easily enable large amounts of data to be searched, analyzed, and explored.
[0088] In some embodiments, the media asset database is used to store audio track voiceprint data of videos. It can perform voiceprint analysis on publicly available video data offline, as well as on newly released or updated video data, continuously expanding the media asset database and providing the server with complete database resources.
[0089] In some embodiments, the media asset database has independent audio track voiceprint data for different videos. Each video's audio track voiceprint data has a unique corresponding video identifier. The audio track voiceprint data may also include the person's name, the person's voiceprint characteristics, timeline parameters, and person-time tags. The person-time tags represent the correspondence between the person and the time their voiceprint appears. For example, for the video with the identifier "66198523", it includes the person Yang and the times Yang's voiceprint appears (8-75 seconds, 214-231 seconds), and the person Liu and the times Liu's voiceprint appears (12-24 seconds, 159-178 seconds). In some embodiments, the above information mapping relationship is shown in Table 1.
[0090]
[0091] Table 1
[0092] In some embodiments, the person and time tag of Yang's voiceprint can have a unique first identifier, and the person and time tag of Liu's voiceprint can have a unique second identifier, and the first identifier is different from the second identifier.
[0093] In some embodiments, the corresponding first identifier and / or second identifier can be retrieved by retrieving the playback time point, thereby determining the corresponding character name and thus determining the character list.
[0094] In some embodiments, the characters and / or a list of characters that appear can be determined based on the timeline data corresponding to the playback time point.
[0095] In some embodiments, the list of characters can be determined based on the mapping relationship between playback time and character names.
[0096] In some embodiments, the preset face database includes the characteristic information of celebrities and their corresponding introductions, and may also include keywords matching the celebrities and related content expressions and / or content sources or types.
[0097] In some embodiments, in response to a user's input image recognition request, the display device sends the target image and video information to the server. The video information may include a video identifier and playback time point. The server's recognition processor performs facial feature recognition on the target image, that is, obtains the facial feature information of the person in the target image, compares it with a preset face database, and obtains the feature matching confidence of the person contained in the target image.
[0098] The server can retrieve the corresponding person from the media asset database based on the playback time. Assuming the media asset database provides complete and comprehensive media assets, the corresponding person can be accurately identified by the playback time.
[0099] In other embodiments, the display device can identify objects in the target image, i.e., identify facial information, and then send the identified facial region image and video information to the server. The server's recognition processor completes the recognition of the object based on the facial region image sent by the display device and obtains the feature matching confidence of the object.
[0100] In some embodiments, such as Figure 6As shown, the feature matching confidence scores of the facial data corresponding to the person object and multiple names in the target image are obtained; the name corresponding to the highest feature matching confidence score is determined as the target person name. If the feature matching confidence score is greater than a first preset threshold, the accurate recognition result of the corresponding person object is determined, and the person information corresponding to the target person name is directly output as the recognition result and sent to the display device for display. If the feature matching confidence score is not greater than the first preset threshold, the person list determined according to the video identifier, the playback time point, and the media asset database, and when the target person name is in the person list, the person information corresponding to the target person name is output as the recognition result. The person object with the highest current matching confidence score is obtained, and the audio track voiceprint data of the corresponding video in the media asset database is requested according to the video identifier in the video information. The person list corresponding to the playback time point of the screenshot is compared. If the person object with the highest current matching confidence score is in the person list, the accurate recognition result of the corresponding person object is determined. If the corresponding person object is not in the person list, the corresponding person is not recognized. For example, if feature recognition technology for the target image determines that the confidence level of a facial feature in the screenshot matches the corresponding feature in a preset facial database greater than a first preset threshold (e.g., 85%), then the corresponding person is identified, and their information can be retrieved for display. If the confidence level of a facial feature in the screenshot does not exceed the first preset threshold, then the person with the highest matching confidence is retrieved. Based on the video identifier, the corresponding audio track and voiceprint data from the media asset database are obtained. A list of people is generated based on the playback time. If the person is in the list, the accurate identification result is confirmed; otherwise, no person is identified. For another example, if the person identified through the preset facial database is "Liu Mou," and the corresponding feature matching confidence is 81%, which is the highest confidence level for that person, then if the list of people obtained based on the playback time includes "Qiao Mou" and "Liu Mou," then the person is identified as "Liu Mou." If the list includes "Yang Mou" and "Wang Mou," then no person can be identified, and there may be no identification result.
[0101] In some embodiments, obtaining the feature matching confidence of the face data corresponding to the person object and the target person name in the target image may involve obtaining the feature matching confidence of the face data corresponding to the person object and multiple person names in the target image; and determining the person name corresponding to the highest feature matching confidence as the target person name.
[0102] In some embodiments, in obtaining the feature matching confidence of the face data corresponding to the person object and the target person name in the target image, the feature matching confidence of the face data corresponding to the person object and multiple person names in the target image is obtained; according to the feature matching confidence in descending order, the person name corresponding to the highest feature matching confidence is determined as the target person name, and the person name corresponding to the feature matching confidence that is excluding the highest feature matching confidence and is greater than a second preset threshold is determined as the secondary person name, wherein the second preset threshold is less than the first preset threshold.
[0103] If the feature matching confidence of the target person's name is not greater than a first preset threshold, based on the video identifier, the playback time point, and the list of people determined by the media asset database, the target person's name and the secondary person's name are sequentially determined in descending order of feature matching confidence. If the target person's name is in the list, the person information corresponding to the target person's name is output as the recognition result. If the target person's name is not in the list but the secondary person's name is in the list, the person information corresponding to the secondary person's name is output as the recognition result. Otherwise, an unrecognized identifier is output.
[0104] Based on face recognition, the addition of voiceprint data from the audio track at the corresponding playback time point improves the accuracy and recall rate of person recognition in image recognition, and enhances the operability of the application.
[0105] In some embodiments, displaying the screen of the currently playing content in the playback window of the currently playing content display area can be either continuing to play the current video content, taking a screenshot of the target image displayed after the current video content has been played, or continuing to play the current video content at the top of the display while displaying the screenshot target image area and the output results sent by the display server at the bottom of the display.
[0106] For example, in Figure 7a The screenshot of the video playback interface shown includes five objects: A, B, C, D, and E, with the corresponding video identifier being 66198545. The screenshot corresponds to a playback time of 5 minutes and 40 seconds. Assume that, after comparison with a preset face database, object A might correspond to "Qiao" with a 95% confidence level of feature matching; object B might correspond to "Jiang" with a 72% confidence level of feature matching, or "Zhang" with a 57% confidence level of feature matching; object C might correspond to "Wang" with an 87% confidence level of feature matching; object D might correspond to "Yang" with a 75% confidence level of feature matching; and object E might correspond to "Liu" with an 81% confidence level of feature matching, or "Cao" with a 72% confidence level of feature matching.
[0107] In the above example, if the first preset threshold is set to 85%, the feature matching confidence of the target image obtained by screenshotting is used to determine whether it is greater than the first preset threshold. The people objects with a confidence level greater than the first preset threshold are A and C, meaning the people identified in the target image are "Qiao" and "Wang". The people objects with a confidence level not greater than the first preset threshold are B, D, and E. For person object B, the one with the highest feature matching confidence level, 72%, "Jiang", is selected for voiceprint recognition. In the media asset database, in the video with video encoding 66198545, the list of people corresponding to the playback time point 5 minutes and 40 seconds is "Jiang, Qiao". Since "Jiang" falls into the list of people corresponding to this time, it can be determined that person object B in the target image is "Jiang". In some embodiments, as shown in Table 2.
[0108]
[0109] Table 2
[0110] In some embodiments, such as Figure 7b As shown, the screenshot display interface can include a continuing video, and the face can be a face in the screenshot, a face from the encyclopedia, or both. The display method can be to show the screenshot information, the corresponding face information, and the encyclopedia information at the bottom of the screen.
[0111] In some real-time examples, because there is no data in the audio track voiceprint data at the time of screenshotting (i.e., no character is speaking lines at the time of screenshotting, resulting in voice asynchrony), the character corresponding to the time axis parameter at the screenshot time may be empty in the preset mapping relationship, or the preset mapping relationship may not include that time axis parameter. A list of characters within a preset time range (e.g., 20 seconds) before and after the playback time point can be obtained according to the request. For example, in the above example, when no data is obtained, for character objects D and E, the one with the highest feature matching confidence can be selected: for object D, "Yang Mou" with a feature matching confidence of 75%; for object E, "Liu Mou" with a feature matching confidence of 81%. Voiceprint recognition is performed on both. At the playback time point of 5 minutes and 40 seconds, there is no corresponding character information in the character list. This time, the list of characters "Jiang Mou, Liu Mou, Yang Mou, Qiao Mou" corresponding to the time range of 5 minutes and 20 seconds to 6 minutes of playback time is obtained. At this point, it can be determined that character object D in the target image is "Yang Mou" and E is "Liu Mou". See Table 3 for some embodiments.
[0112]
[0113] Table 3
[0114] In some embodiments, even if a character corresponds to a timeline parameter at a given playback time point, the names of characters within a preset time range (e.g., 20 seconds) before and after the playback time point are retrieved to generate a character list. This is because videos often use camera cuts, and sometimes the screen shows someone listening to someone else, rather than the person speaking. However, the plot of a video is continuous, and the person currently shown listening to someone else may be having a conversation with someone else within the preset time range before and after. Therefore, retrieving the character list by mapping a period of time before and after the playback time point can make the results more accurate.
[0115] like Figure 7a As shown, based on the video identifier in the video information, the system requests the person and time identifier in the audio track voiceprint data of the corresponding video in the media asset database. By using the playback time point and the person names corresponding to the time point before and after the time point, the system obtains the corresponding person list. If the person with the highest matching confidence is in the person list, the system determines the accurate recognition result of the corresponding person. If the corresponding person is not in the person list, the corresponding person has not been recognized.
[0116] In one embodiment, in obtaining the feature matching confidence score of the facial data corresponding to the target person's name, the feature matching confidence scores of the person object in the target image and the facial data corresponding to multiple names are obtained; according to the feature matching confidence scores from high to low, the person's name corresponding to the highest feature matching confidence score is determined as the target person's name, and the person's name corresponding to the feature matching confidence score other than the highest feature matching confidence score and greater than a second preset threshold is determined as a secondary person's name, wherein the second preset threshold is less than the first preset threshold. Figure 8 As shown, if the feature matching confidence level is not greater than the first preset threshold, based on the video identifier, the playback time point, and the list of people determined by the media asset database, the target person name and the secondary person name are sequentially determined in descending order of feature matching confidence level. If the target person name is in the list, the person information corresponding to the target person name is output as the recognition result; if the target person name is not in the list but the secondary person name is in the list, the person information corresponding to the secondary person name is output as the recognition result; otherwise, an unrecognized identifier is output.
[0117] for Figure 7aIn the example described, if the first preset threshold is set to 85% and the second preset threshold is set to 70%, the objects with a feature matching confidence level greater than the first preset threshold are A and C, i.e., the identified objects in the target image are "Qiao" and "Wang". The objects with a feature matching confidence level not greater than the first preset threshold are B, D, and E. For object E, all possible objects with a feature matching confidence level between the first and second preset thresholds (i.e., 70%-85%) are selected. The target name "Liu" with a feature matching confidence level of 81% and the secondary name "Cao" with a feature matching confidence level of 72% both satisfy this condition. The criteria are then checked sequentially. The list of characters within a preset time range around 5 minutes and 40 seconds into the playback time is "Jiang, Liu, Yang, Qiao". Character object E is determined to be "Liu". For character object B, all possible characters with a feature matching confidence level between 70% and 85% are identified as "Jiang" with a feature matching confidence level of 72%. Another possible character, "Zhang", has a feature matching confidence level of 57%, which is outside the 70%-85% range and is not considered for further determination. The list of characters "Jiang, Liu, Yang, Qiao" within the preset time range around 5 minutes and 40 seconds into the playback time is compared, and character object B is determined to be "Jiang". See Table 4 for some embodiments.
[0118]
[0119]
[0120] Table 4
[0121] In some embodiments, the higher the first preset threshold is set, the higher the accuracy and recall are. The secondary names determined by the second threshold in voiceprint recognition further improve the accuracy and recall.
[0122] In some embodiments, the target image received by the server may be a complete screenshot or it may contain information about the person or object after analysis and processing by the display device.
[0123] In some embodiments, after the server completes image recognition, it returns the recognition results corresponding to all objects in the target image, namely the accurately recognized results and the unrecognized results, to the display device. The display device retains the accurately recognized results and discards the data of the unrecognized objects.
[0124] In other embodiments, after the server completes image recognition, it returns the accurate recognition result to the display device and deletes the unrecognized result directly.
[0125] In some embodiments, for unidentified results, the result of their feature matching confidence can be regarded as a similar identification result, retained, and returned to the display device. For similar identification results, they can be displayed in association with similarity indication information so that users can understand the degree of similarity between each identification result and the corresponding object, as well as the accuracy differences between each identification result.
[0126] This application provides a display device in some embodiments, comprising a display and a controller. The display is used to display a user interface, and the controller is configured to receive an input image recognition request. In response to the image recognition request, during video playback, the controller determines a target image, video identifier, and playback time point of the user interface to identify a person in the target image. The controller then sends the target image, video identifier, and playback time point to a server, so that the server outputs the person information corresponding to the target person's name identified from the target image and the video information as a recognition result to the display device, displaying the accurate recognition result of the person.
[0127] In some implementations, the controller is configured to receive an input image recognition request; in response to the image recognition request, determine the target image, video identifier, and playback time point in the user interface to identify human objects in the target image, identify human objects in the target image, and send the identified human object information and video information to the server so that the server can provide feedback on the accurate recognition result corresponding to the object based on the human object information and video information.
[0128] This application provides a display device in some embodiments, comprising a display and a controller. The display is used to display a user interface, and the controller is configured to receive an input image recognition request. In response to the image recognition request, when not playing video, the controller determines a target image of the user interface to identify a person in the target image. The controller then sends the target image to a server, so that the server outputs the person information corresponding to the target person's name identified from the target image and the video information as a recognition result to the display device, displaying the accurate recognition result of the person.
[0129] As can be seen from the above technical solutions, this application discloses an image recognition method, a server, and a display device. The display device sends an image recognition request to the server. The server receives the image recognition request, which includes the target image, video identifier, and playback time point. It obtains the feature matching confidence of the facial data corresponding to the person in the target image and the target person's name. If the feature matching confidence of the target person's name is greater than a first preset threshold, the person information corresponding to the target person's name is directly output as the recognition result. If the feature matching confidence of the target person's name is not greater than the first preset threshold, the person information corresponding to the target person's name is output as the recognition result based on the video identifier, playback time point, and a list of people determined by the media asset database, and if the target person's name is in the list of people, the person information corresponding to the target person's name is output as the recognition result. Otherwise, an unrecognized identifier is output. Voiceprint-assisted recognition improves the accuracy and recall rate of person recognition, thereby enhancing the user experience.
[0130] Similar parts between the embodiments provided in this application can be referred to mutually. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods extended from the solution of this application without creative effort shall fall within the scope of protection of this application.
Claims
1. A server, characterized in that, include: Media asset database, used to store audio track and voiceprint data for videos; Pre-set face database; The recognition processor is configured to: The device receives an image recognition request sent by a display device. The image recognition request includes a target image, a video identifier, and a playback time point. The target image is a screenshot of the video corresponding to the video identifier that is being played by the display device. The playback time point is the moment when the video is playing when the screenshot is taken. The feature matching confidence score is obtained by comparing the face data corresponding to the person object and the target person name in the target image. If the feature matching confidence of the target person's name is greater than the first preset threshold, the person information corresponding to the target person's name will be directly output as the recognition result. If the feature matching confidence of the target person's name is not greater than a first preset threshold, the person information corresponding to the target person's name is output as the recognition result when the target person's name is in the person list, based on the video identifier, the playback time point, and the list of people determined by the media asset database; otherwise, an unrecognized identifier is output. The media asset database stores the person time correspondence between people in the video and the time when the person's voiceprint appears.
2. The server according to claim 1, characterized in that, In the step of obtaining the feature matching confidence score of the face data corresponding to the person object and the target person name in the target image, the recognition processor is further configured to: Obtain the feature matching confidence score of the face data corresponding to the person object and multiple names in the target image; The name corresponding to the highest feature matching confidence is identified as the target name.
3. The server according to claim 1, characterized in that, In the step of obtaining the feature matching confidence score of the face data corresponding to the person object and the target person name in the target image, the recognition processor is further configured to: Obtain the feature matching confidence score of the face data corresponding to the person object and multiple names in the target image; According to the feature matching confidence level from high to low, the name corresponding to the highest feature matching confidence level is determined as the target name, and the name corresponding to the feature matching confidence level other than the highest feature matching confidence level and the feature matching confidence level greater than the second preset threshold is determined as the secondary name, wherein the second preset threshold is less than the first preset threshold.
4. The server according to claim 3, characterized in that, If the feature matching confidence of the target person's name is not greater than a first preset threshold, the recognition processor is further configured to: based on the video identifier, the playback time point, and the list of people determined by the media asset database. The target name and the secondary name are determined sequentially in descending order of feature matching confidence. If the target name is in the list of people, the person information corresponding to the target name is output as the recognition result. If the target name is not in the list of people but the secondary name is in the list of people, the person information corresponding to the secondary name is output as the recognition result. Otherwise, an unrecognized identifier is output.
5. The server according to claim 1 or 4, characterized in that, The media asset database includes constantly updated audio track and voiceprint data; Each audio track's voiceprint data includes a unique video identifier, as well as a person and time tag, which represents the person and the time when their voiceprint appears.
6. The server according to claim 5, characterized in that, In the step of determining the list of people based on the video identifier, the playback time point, and the media asset database, the recognition processor is further configured to: The video in the media asset database is determined based on the video identifier; Based on the audio track voiceprint data of the corresponding video at the playback time point, determine the list of characters corresponding to the playback time point.
7. The server according to claim 1 or 4, characterized in that, In the step of determining the list of people based on the video identifier, the playback time point, and the media asset database, the recognition processor is further configured to: The video in the media asset database is determined based on the video identifier; The identification time period is determined based on the playback time point, and the identification time period is a preset time range before and after the playback time point; Based on the voiceprint data of the audio track in the corresponding video during the identified time period, a list of people corresponding to the identified time period is determined.
8. A display device, characterized in that, include: A monitor is used to display the user interface. The controller is configured as follows: Receive image recognition request; In response to the image recognition request, during video playback, the target image, video identifier, and playback time point of the user interface are determined to identify the human object in the target image. The target image is obtained by taking a screenshot of the video corresponding to the video identifier played on the display device. The playback time point is the moment when the screenshot is taken. The target image, video identifier, and playback time point are sent to the server so that the server compares the person in the target image with a preset face database to obtain the feature matching confidence of the face data corresponding to the person in the target image and the target person's name. The system receives and displays the identification results of the person object sent by the server. Specifically, if the feature matching confidence of the target person's name is greater than a first preset threshold, the person information corresponding to the target person's name is directly output as the identification result. If the feature matching confidence of the target person's name is not greater than the first preset threshold, the system outputs the person information corresponding to the target person's name as the identification result when the target person's name is in the person list, based on the video identifier, the playback time point, and the person list determined by the media asset database. Otherwise, an unidentified identifier is output. The media asset database stores the person-time correspondence between people in the video and the time when their voiceprints appear.
9. An image recognition method, characterized in that, include: The device receives an image recognition request sent by a display device. The image recognition request includes a target image, a video identifier, and a playback time point. The target image is a screenshot of the video corresponding to the video identifier that is being played by the display device. The playback time point is the moment when the video is playing when the screenshot is taken. The feature matching confidence score is obtained by comparing the face data corresponding to the person object and the target person name in the target image. If the feature matching confidence of the target person's name is greater than the first preset threshold, the person information corresponding to the target person's name will be directly output as the recognition result. If the feature matching confidence of the target person's name is not greater than a first preset threshold, the person information corresponding to the target person's name is output as the recognition result when the target person's name is in the person list, based on the video identifier, the playback time point, and the list of people determined by the media asset database; otherwise, an unrecognized identifier is output. The media asset database stores the person time correspondence between people in the video and the time when the person's voiceprint appears.