Display device and feature recognition method
By using query and feature parameters to perform image recognition in the display device and generating a result display view, the problem of low efficiency in multi-target object recognition is solved and the user interaction experience is improved.
Patent Information
- Application Number
- PCT/CN2025/083271
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2025-03-18
- Publication Date
- 2025-10-02
AI Technical Summary
In a display device, when a screenshot image contains multiple target objects, the server returns multiple image recognition results, and the user needs to select from multiple results, resulting in reduced image recognition efficiency and affecting the user interaction experience.
The display device obtains the image to be identified, uses query parameters and feature parameters to perform image recognition, and generates a result display view. The display position and status of the feature label and details are set according to the arrangement parameters to highlight the feature information of the specific object.
The efficiency of image recognition is improved, and users can directly view the highlighted specific object information on the display device, which reduces the selection process and improves the user interaction experience.
Smart Images

Figure CN2025083271_02102025_PF_FP_ABST
Abstract
Description
Display device and feature recognition method
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure claims priority to Chinese patent applications filed with the State Intellectual Property Office of China on March 29, 2024, with application number 202410381279.6; filed with the State Intellectual Property Office of China on April 29, 2024, with application number 202410529035.8; and filed with the State Intellectual Property Office of China on May 17, 2024, with application number 202410619867.9, the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0003] The present disclosure relates to the technical field of display devices, and in particular to a display device and a feature recognition method. Background Art
[0004] Display devices are intelligent devices that can present user interfaces and support user interaction. For example, smart TVs are based on Internet application technologies, feature open operating systems and chips, and boast an open application platform. They enable two-way human-computer interaction and integrate multiple functions, including audio, video, entertainment, and data, to meet diverse and personalized user needs. Display devices can interact with servers through built-in applications to implement various functions. For example, display devices can implement voice recognition through voice applications and image recognition through image recognition applications.
[0005] Taking the image recognition function as an example, users can input control instructions to the image recognition application of the display device through a specific method, causing the image recognition application to take a screenshot of the display device, upload the screenshot image generated by the screenshot to the server, and then display the image recognition results returned by the server. However, when the screenshot image includes multiple target objects, the server will also return multiple image recognition results to the display device. In this way, the user needs to select the recognition result they want from the multiple image recognition results, which reduces the overall efficiency of image recognition in the display device and affects the user's interactive experience. Summary of the Invention
[0006] In a first aspect, an embodiment of the present disclosure provides a display device, which may include: a display, which may be configured to display a user interface; a memory, which may be configured to store computer instructions and data associated with the display device; at least one processor, connected to the display and the memory, and configured to execute computer instructions so that the display device performs: in response to an image recognition instruction, obtaining an image to be recognized, wherein the image to be recognized includes an object to be recognized, and the image recognition instruction includes a query parameter and a feature parameter, wherein the query parameter is used to characterize the object type of the object to be recognized; the feature parameter is used to characterize the feature attribute of the object to be recognized; and performing a processing on the image to be recognized according to the query parameter extracted from the image recognition instruction. Image recognition to obtain a recognition result, the recognition result including an identification object that meets the query parameters and at least one of an arrangement parameter, a feature label and feature details of the identification object, the arrangement parameter being used to characterize the positional relationship of the identification object in the image to be identified; the feature label being used to characterize the identification information of the identification object, and the feature details being associated information queried based on the feature label; generating a result display view based on the recognition result, the result display view including a feature label and feature details, the display position of the feature label being set according to the arrangement parameter, and the display status of the feature details being highlighted according to the feature parameter; and controlling the display to display the result display view.
[0007] In a second aspect, an embodiment of the present disclosure also provides a feature recognition method, which may include: in response to an image recognition instruction, obtaining an image to be recognized, the image to be recognized includes an object to be recognized, the image recognition instruction includes a query parameter and a feature parameter, the query parameter is used to characterize the object type of the recognition object; the feature parameter is used to characterize the feature attribute of the recognition object; performing image recognition on the image to be recognized according to the query parameter extracted from the image recognition instruction to obtain a recognition result, the recognition result includes the recognition object that meets the query parameter and at least one of the arrangement parameters, feature labels and feature details of the recognition object, the arrangement parameter is used to characterize the positional relationship of the recognition object in the image to be recognized; the feature label is used to characterize the recognition information of the recognition object, and the feature details are associated information based on the feature label query; generating a result display view according to the recognition result, the result display view including the feature label and feature details, the display position of the feature label is set according to the arrangement parameter, and the display status of the feature details is highlighted according to the feature parameter; controlling the display to display the result display view. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG1 is a schematic diagram of a usage scenario of a display device according to some embodiments;
[0009] FIG2 is a diagram illustrating interaction between a display device and a server according to some embodiments;
[0010] FIG3 is a block diagram of a hardware configuration of a display device according to some embodiments;
[0011] FIG4 is an architectural diagram illustrating a communication connection between a server and a display device according to some embodiments;
[0012] FIG5 is a schematic diagram of software configuration in a display device according to some embodiments;
[0013] FIG6 is a flow chart of interaction between a display device and a server according to some embodiments;
[0014] FIG7 is an architectural diagram of a speech recognition module according to some embodiments;
[0015] FIG8 is an architectural diagram of an image recognition module according to some embodiments;
[0016] FIG9 is a schematic flow chart of a feature recognition method according to some embodiments;
[0017] FIG10 is a schematic diagram of a voice recognition input interface according to some embodiments;
[0018] FIG11 is a schematic diagram of a smart view interface according to some embodiments;
[0019] FIG12 is a schematic diagram of a process for obtaining an image to be recognized according to some embodiments;
[0020] FIG13 is a schematic diagram of a process for extracting query parameters and feature parameters according to some embodiments;
[0021] FIG14 is a schematic diagram of a process for generating a result display view according to some embodiments;
[0022] FIG15 is a diagram showing a display effect of a result display view according to some embodiments;
[0023] FIG16 is a diagram showing another result display view according to some embodiments;
[0024] FIG17 is a diagram showing the effect of the first result after screening based on inherent attributes according to some embodiments;
[0025] FIG18 is a diagram showing the effect of a second method of displaying the results of the inherent attribute screening according to some embodiments;
[0026] FIG19 is a diagram showing the effect of a third method of displaying the results of the inherent attribute screening according to some embodiments;
[0027] FIG20 is a diagram showing the effect of displaying results based on multiple feature words according to some embodiments;
[0028] FIG21 is a diagram showing the display effect of the result based on the third feature word according to some embodiments;
[0029] FIG22 is a schematic diagram of a process of performing feature recognition by a server according to some embodiments;
[0030] FIG23 is a schematic diagram of a target coordinate system according to some embodiments;
[0031] FIG24 is a schematic diagram of position coordinates of a first object according to some embodiments;
[0032] FIG25 is another schematic flow chart of a feature recognition method according to some embodiments;
[0033] FIG26 is a schematic diagram of a process of image recognition according to some embodiments;
[0034] FIG27 is a schematic diagram of a process for extracting a second object according to some embodiments;
[0035] FIG28 is a flow chart of a method for displaying target recognition results according to some embodiments;
[0036] FIG29 is a schematic diagram of a flow chart of semantic recognition performed by a display device according to some embodiments;
[0037] FIG30 is a diagram showing the display effect of target recognition results according to some embodiments;
[0038] FIG31 is a diagram showing a display effect of a user interface according to some embodiments;
[0039] FIG32 is a display effect diagram of a recognition result according to some embodiments;
[0040] FIG33 is a display effect diagram of another recognition result according to some embodiments. DETAILED DESCRIPTION
[0041] In the embodiments of the present disclosure, the display device 200 generally refers to a device capable of displaying images and processing data. For example, the display device 200 includes but is not limited to smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.
[0042] Figure 1 illustrates an operational scenario between a display device and a control device. As shown in Figure 1, a user can operate a display device 200 through touch control, a mobile terminal 300, and a control device 100. The control device 100 receives user input and converts it into control commands that the display device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus, a controller, or the like.
[0043] The mobile terminal 300 can function as a control device for performing human-computer interaction between a user and the display device 200. The mobile terminal 300 can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, enabling connection and communication via a network communication protocol, enabling one-to-one control operations and data communication. Audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 for synchronized display.
[0044] In some embodiments, the mobile terminal 300 or other electronic device can also simulate the functions of the control device 100 by running an application that controls the display device 200. Continuing with FIG1 , the display device 200 can also communicate data with the server 400 through various communication methods. For example, the display device 200 can be connected to a local area network (LAN), a wireless local area network (WLAN), or other networks.
[0045] As shown in Figure 2, the display device 200 can communicate with the server 400 during use to achieve data interaction. The server 400 can be a server that provides various services. For example, a background server that provides support for the voice data collected by the display device 200. The background server can analyze and process the received voice and other data, and feed back the processing results (such as endpoint information) to the display device 200. It can also be an audio and video data server that feeds back the audio and video data to be played based on the audio and video request sent by the terminal. In some embodiments, the server 400 can be a server cluster or a plurality of server clusters, and can include one or more types of servers. For example, voice servers, audio and video servers, etc. are located in the same server cluster.
[0046] In some embodiments, the server 400 can be deployed on a remote server to perform data storage, processing, and analysis. For example, the display device 200 can send a data packet to the server 400 based on its built-in application. After receiving the data packet sent by the display device 200, the server 400 parses the request type and content in the data packet, processes the data packet accordingly, and then sends the generated response data packet back to the application of the display device 200, thereby enabling data exchange between the display device 200 and the server 400.
[0047] In order to implement data interaction between the display device 200 and the server 400, the display device 200 needs to establish a communication connection with the server 400. For example, the display device 200 and the server 400 can both be connected to the Internet, and the interactive data is transmitted between the display device 200 and the server 400 through the Internet transmission protocol.
[0048] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart TV, Internet Protocol television (IPTV), etc.
[0049] Figure 3 is a hardware configuration block diagram of the display device 200 in Figure 1. As shown in Figure 3, the display device 200 may include at least one of a tuner and demodulator 210, a communication device 220, a detector 230, a device interface 240, a processor 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0050] In some embodiments, detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 may include a light receiver, such as a sensor for collecting ambient light intensity; or an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or a sound collector, such as a microphone, for receiving external sounds.
[0051] In some embodiments, the display 260 includes a display component for presenting images and a driver component for driving image display. The display 260 is configured to receive image signals output from the processor 250 for display. For example, the display 260 can be configured to display video content, image content, menu control interface components, and user control UI interfaces.
[0052] In some embodiments, the communication device 220 is a component used to communicate with an external device or server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 depending on the supported communication methods. For example, if the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including WiFi functionality. If the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including Bluetooth functionality.
[0053] The communication device 220 can establish a communication connection between the display device 200 and an external device or server 400 via a wireless or wired connection. A wired connection can connect the display device 200 to an external device via a data cable, an interface, or other components. A wireless connection can connect the display device 200 to an external device via a wireless signal or wireless network. The display device 200 can establish a connection with an external device directly or indirectly through a gateway, router, or connection device.
[0054] In some embodiments, the display device 200 and the server 400 each need to be provided with a component for establishing a communication connection. Specifically, as shown in FIG4 , the display device 200 may be provided with a communication device 220, while the server 400 may be provided with a communication module 410. The communication device 220 and the communication module 410 may simultaneously support at least one of the same communication methods to establish a communication connection. For example, the communication device 220 on the display device 200 includes a fiber optic interface, allowing the display device 200 to connect to the network via the fiber optic interface. Simultaneously, the communication module 410 on the server 400 also includes a fiber optic interface and can also connect to the network via the fiber optic interface, thereby achieving a communication connection between the display device 200 and the server 400.
[0055] It should be noted that the display device 200 and the server 400 may also establish a communication connection relationship using other connection methods, such as wired broadband, wireless local area network, cellular network, Bluetooth, infrared, radio frequency communication, etc.
[0056] In some embodiments, the connection between the display device 200 and the server 400 can be a "many-to-one" relationship, that is, multiple display devices 200 can establish communication connections with the same server 400, allowing the server 400 to provide services for multiple display devices 200. The connection between the display device 200 and the server 400 can also be a "many-to-many" relationship, that is, multiple display devices 200 can establish communication connections with multiple servers 400, allowing the multiple servers 400 to provide different services to the display devices 200. Obviously, in some application scenarios, the connection between the display device 200 and the server 400 can also be a "one-to-one" relationship, that is, one server 400 specifically provides services for one display device 200.
[0057] To provide services for the display device 200, in some embodiments, the server 400 may further include a storage module 420. The storage module 420 may store various resource data, information files, control programs, and recognition models, and may also be used to regularly back up key data sent by the display device 200 to prevent data loss. For example, the storage module 420 may store files, data, virtual machine images, and the like uploaded by the display device 200.
[0058] In some embodiments, the storage module 420 of the server 400 can provide fast I / O access to support database operations and ensure efficient data processing and querying. The storage module 420 supports multiple access methods, such as the Network File System (NFS), the File Transfer Protocol (FTP), and the Common Internet File System (CIFS), enabling different applications or display devices 200 to access or share data simultaneously, thereby enabling functions such as document sharing and resource access. That is, as the user interacts, the display device 200 can obtain different data from the storage module 420 of the server 400 to implement specific functions.
[0059] In some embodiments, the processor 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and may include first to nth interfaces for input / output. The processor 250 controls the operation of the display device and responds to user operations through various software control programs stored in a memory. The processor 250 controls the overall operation of the display device 200.
[0060] In some embodiments, the processor 250 and the tuner / demodulator 210 may be located in different separate devices, that is, the tuner / demodulator 210 may also be located in an external device of the main device where the processor 250 is located, such as an external set-top box.
[0061] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0062] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may further be provided with an external audio output terminal, through which the audio output device may be connected to the display device 200 to output the sound of the display device 200.
[0063] In some embodiments, the user input interface 280 may be configured to receive instructions from a user.
[0064] In some embodiments, the detector may also be used to trigger the display device to generate corresponding operation instructions based on the detected sound or image signal.
[0065] To facilitate user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can control the display device to provide a user interface. For example, the operating system can directly control the display device to provide a user interface, or it can provide a user interface by running an application program. The operating system also allows the user to interact with the display device 200.
[0066] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.
[0067] The operating system can be divided into different modules or layers according to the functions implemented. Taking the Android system as an example, as shown in Figure 5, in some embodiments, the system is divided into four layers, from top to bottom: the application layer (referred to as the "application layer"), the application framework layer (referred to as the "framework layer"), the system library layer and the kernel layer.
[0068] It should be noted that the above example is only a simple division of the operating system functions and does not limit the specific operating system form of the display device 200 in the embodiment of the present disclosure. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.
[0069] In some embodiments, the application layer of the display device 200 can be configured with a voice application, and the display device 200 can respond to the voice command input by the user based on the data interaction between the voice application and the server 400. That is, the display device 200 can collect the voice (sound signal) input by the user through the sound collector in the detector 220, and then convert the analog electrical signal corresponding to the voice into a digital signal, and then send the converted digital signal as a voice command to the voice application, so that the display device 200 processes it based on the voice application. In some embodiments, the control device 100 can receive and process the voice, generate a control signal, and then send it to the display device 200. That is, the display device 200 can receive the control signal generated by the remote control and send the control signal as a voice command to the voice application.
[0070] For example, the display device 200 collects the user's voice input through a sound collector and generates an analog electrical signal. The analog-to-digital converter (ADC) in the sound collector then samples the continuous analog point signal, that is, it periodically measures the voltage of the analog point signal and converts it into discrete values. The ADC then quantizes these discrete values and maps them into a discrete numerical range. Finally, the ADC uses an encoder to convert the quantized values into a binary digital signal and passes the binary digital signal (voice data) to the voice application for processing.
[0071] In order to facilitate the distinction of the user's voice, in some embodiments, the control device 100 of the display device 200 is configured with a specific voice function button. When the user presses the voice function button, the display device 200 collects the user's input voice data based on the detector 220 or the sound collector in the control device 100 to generate a voice command. Alternatively, in some embodiments, the display device 200 itself is configured with a specific voice function triggering method. For example, the display device 200 is connected to an external or built-in camera, and the camera collects the user's image in real time. When the user makes a specific gesture or action, the display device 200 collects the user's input voice data based on the sound collector as a voice command.
[0072] In some embodiments, the display device 200 may include a built-in voice interaction system for performing voice recognition on user input voice data, generating interaction commands based on the voice recognition results, and executing corresponding interaction controls in response to the interaction commands. Specifically, the voice data collected by the sound collector of the display device 200 may be fed into the voice interaction system, which then analyzes and processes the data and presents a corresponding interaction interface or interface transition through the voice application.
[0073] In some embodiments, after receiving the voice instruction, the voice application of the display device 200 sends the voice instruction to the server 400 , so that the server 400 performs analysis processing on the voice instruction and sends the analysis processing result to the display device 200 .
[0074] FIG6 is a flowchart of the interaction between a display device and a server according to some embodiments. As shown in FIG6 , the display device 200 may include a voice application 210 and an image recognition application 220, and the server 400 may include a voice recognition module 430 and an image recognition module 440. The interaction process between the display device and the server may include the following steps:
[0075] S601: The voice application 210 sends a voice command to the voice recognition module 430;
[0076] S602: The speech recognition module 430 performs semantic recognition on the speech instruction;
[0077] S603: The speech recognition module 430 returns the semantic recognition result to the speech application 210;
[0078] S604: The voice application 220 performs corresponding processing on the semantic recognition process;
[0079] After receiving the voice command input by the user, the voice application of the display device 200 sends the voice command to the voice recognition module 430 of the server 400. The voice recognition module 430 then performs semantic recognition on the voice command and sends the obtained semantic recognition result back to the voice application of the display device 200, so that the voice application of the display device 200 performs corresponding processing based on the semantic recognition result.
[0080] S605: The voice application 210 notifies the image recognition application 220 of the recognition result;
[0081] S606: The image recognition application 220 executes the image recognition instruction;
[0082] S607: The image recognition application 220 takes a screenshot of the image to be recognized;
[0083] S608: The image recognition application 220 sends an image recognition request carrying the image to be recognized to the image recognition module 440;
[0084] S609: The image recognition module 440 performs image recognition on the image to be recognized;
[0085] S610: The image recognition module 440 returns the image recognition result to the image recognition application 220.
[0086] In some embodiments, the voice recognition module 430 may only perform text recognition corresponding to the voice command and send the text recognition to the display device 200. The display device 200 performs semantic recognition itself or performs voice recognition through the server 400 to obtain the recognition result.
[0087] In some embodiments, the text recognized by the speech recognition module 430 can be directly used as the recognition result.
[0088] In some embodiments, the display device 200 can obtain the speech recognition result alone or in collaboration with the server 400. It only needs to obtain the recognition result to proceed to the next step.
[0089] FIG7 is an architecture diagram of a speech recognition module according to some embodiments. As shown in FIG7 , the server 400 may include a speech recognition module 430, which includes a recognition unit 431, a semantic understanding unit 432, a management unit 433, a language generation unit 434, and a speech synthesis unit 435. The recognition unit 431 is deployed with a speech recognition service for recognizing speech data in speech instructions as text; the semantic understanding unit 432 is deployed with a semantic understanding service for performing semantic analysis on text; the management unit 433 is deployed with a business instruction management service for providing business instructions; the language generation unit 434 is deployed with a language generation service (NLG) for converting instructions for the display device 200 to execute into text language; and the speech synthesis unit 435 is deployed with a text-to-speech (TTS) service for processing the text language corresponding to the instruction and sending it to the speaker of the display device 200 for broadcast. In one embodiment, the architecture shown in FIG7 may include multiple physical service devices deployed with different business services, or one or more physical service devices may be combined with one or more functional services.
[0090] In some embodiments, the speech recognition service deployed by the recognition unit 431 can perform noise reduction processing and feature extraction on the audio of the query statement. Among them, the noise reduction processing may include steps such as removing echoes and environmental noise. The speech understanding service deployed by the semantic understanding unit 432 can use acoustic models and language models to perform natural language understanding on the identified candidate text and associated context information, parse the text into structured, machine-readable information, business domain, intent, word slots and other information to express semantics, etc., so as to obtain an executable intent to determine the intent confidence score. The semantic understanding unit 432 then selects one or more candidate executable intents based on the determined intent confidence score. The business instruction management service deployed by the management unit 433 can give query results based on the query instructions issued by the semantic understanding unit 432, as well as execute the actions required for the user's final request, and feedback the display device execution instructions corresponding to the query results.
[0091] In some embodiments, the application layer of the display device 200 may also be configured with an image recognition application, and the display device 200 may respond to image recognition instructions input by the user based on the data interaction between the image recognition application and the server 400. The image recognition instruction may be a control instruction input by the user based on a specific operation, or may be a voice instruction with image recognition intent input by the user. For example, the user may trigger the image recognition application of the display device 200 to execute the screenshot action and image recognition program by using a button representing the screenshot function in the control device 100; or the user may type text with image recognition intent in a specific input box to trigger the image recognition application to execute the screenshot action and image recognition program; or the user may speak a voice with image recognition intent within the voice detection range of the display device 200 to trigger the image recognition application to execute the screenshot action and image recognition program.
[0092] In some embodiments, when the image recognition instruction is a voice instruction, the display device 200 receives the voice instruction based on a voice application and sends the voice instruction to the server 400. The server 400 performs semantic recognition on the voice instruction and sends the semantic recognition result obtained to the display device 200. The voice recognition result includes an instruction to call the image recognition application. After receiving the semantic recognition result, the display device 200 calls the image recognition application to execute the screenshot action and image recognition program.
[0093] In some embodiments, the image recognition instruction may further specify an image to be recognized. For example, a user may select any image in a file management application. At this point, the display device 200 may display a "Recognize Image" function option in the image viewing interface. After the user selects the "Recognize Image" function option, the display device 200 may launch the image recognition application and perform image recognition using the selected image as the image to be recognized.
[0094] In some embodiments, the above methods are only exemplary descriptions and do not limit the image recognition scenarios. In the embodiments of the present disclosure, triggering the display device 200 to execute the screenshot and image recognition program can also be other scenarios.
[0095] In some embodiments, after receiving the image recognition instruction, the image recognition application of the display device 200 obtains the image to be recognized, and then sends the image to be recognized and the image recognition request to the server 400 to obtain the image recognition result of the image to be recognized. For example, as shown in Figure 6, the server 400 includes an image recognition module 440. After the image recognition application of the display device 200 takes a screenshot of the user interface in response to the image recognition instruction, the image to be recognized and the image recognition request generated by the screenshot are sent to the image recognition module 440 of the server 400. The image recognition module 440 performs image recognition on the image to be recognized based on the image recognition algorithm, and sends the obtained image recognition result back to the image recognition application of the display device 200, so that the image recognition application can display the image recognition result.
[0096] In some embodiments, the image to be recognized sent by the image recognition application in the embodiments of the present disclosure may be an image generated by the display device 200 through screenshot, or may be a media resource specified in the display device 200. This disclosure does not limit this.
[0097] To implement the above-mentioned image recognition function, in some embodiments, as shown in FIG8 , the display device 200 or the server 400 may include an image recognition module. Taking the server 400 as an example, the server 400 includes an image recognition module 440, which includes an acquisition unit 441, a preprocessing unit 442, a feature extraction unit 443, a matching unit 444, and a result output unit 445. Among them, the acquisition unit 441 is used to acquire the image to be identified sent by the display device 200; the preprocessing unit 442 is deployed with an image preprocessing service, which is used to remove noise from the image to be identified, enhance the image quality of the image to be identified, and convert the image to be identified into a form suitable for analysis, such as image scaling, cropping, smoothing filtering, binarization and other steps; the feature extraction unit 443 is deployed with an image feature extraction service, which is used to use a specific algorithm to extract key features (shape, texture, color distribution, etc.) from the preprocessed image, and encode the extracted features into numerical or vector form; the matching unit 440 is deployed with a pattern matching service, which is used to match the extracted features with standard patterns or templates pre-stored in a database, and calculate the similarity or distance measurement between the features to obtain the standard pattern that best matches the image to be identified; the result output unit 445 is deployed with an image recognition result output service, which is used to determine the category or identification of the image to be identified based on the result of pattern matching, and then output the recognition result in text, label or other form according to the determined type and identification for use by the display device 200.
[0098] For example, in the case of person recognition, the image to be recognized, generated by a screenshot of the display device 200, contains a person. The image recognition application of the display device 200 sends the image to be recognized and a recognition request to the image recognition module 440 of the server 400. Upon receiving the image to be recognized, the image recognition module 440 first preprocesses the image to be recognized, including steps such as adjusting the image size, determining the location information to be recognized, cropping, grayscaling, and denoising, thereby highlighting the person's features and reducing the computational effort. Following preprocessing, the image recognition module 440 of the server 400 extracts key features from the image to be recognized, including facial contours, the shape and position of the eyes, nose, and mouth, and skin texture, through edge detection, template matching, and deep learning. The extracted features are then compared with person features pre-stored in a database, and the similarity or distance between the features is calculated to identify the person who best matches the image to be recognized. Based on the comparison results, the person's associated information is obtained as the person recognition result. The person recognition result may include data such as the person's associated tag, name, gender, height, weight, and ID. The server 400 then returns the identified person recognition result to the image recognition application of the display device 200, so that the display device 200 displays the person recognition result based on the image recognition application.
[0099] It should be noted that the architectures shown in Figures 6-8 are merely examples and do not limit the scope of protection of this disclosure. In the embodiments of this disclosure, other architectures may also be used to implement similar functions. For example, some of the unit architectures of the server 400 described above may be provided in the display device 200; accordingly, all or part of the processes of the voice recognition function and image recognition function described above may be performed by the display device 200, and will not be described in detail here.
[0100] For ease of description, in some embodiments of this disclosure, the voice application used to implement the voice recognition function is referred to as the first application, and the application used to implement the image recognition function is referred to as the second application. The first application and the second application can be two independent applications, or two independent application modules integrated into a single application, respectively used to implement the voice recognition function and the image recognition function.
[0101] In some application scenarios, the image to be identified on the display device 200 may include multiple identification objects of the same object type (such as people, animals, objects, places, etc.). For example, the object type of the identification object is a person. When the image to be identified includes two or more people, the image recognition result returned by the server 400 to the display device 200 will include the recognition results of multiple people. Accordingly, the display device 200 will display the recognition results of all people in the user interface. At this time, the user needs to select the desired person recognition result from the multiple person recognition results displayed on the display device 200, which will consume a certain amount of time, resulting in a decrease in the overall efficiency of image recognition in the display device 200, affecting the user's interactive experience.
[0102] In some embodiments, the display device 200 or the server 400 can output the image recognition result based on the feature parameters in the image recognition instruction and the arrangement parameters in the image to be recognized, so that the feature labels in the result display view presented by the display device 200 conform to the arrangement rules of the recognition object in the image to be recognized, and the feature details conform to the feature parameters in the image recognition instruction. That is, for the display device 200, the processor 250 of the display device 200 can run a feature recognition method so that the display device 200 can display the recognition results according to the feature recognition method. Similarly, for the server 400, the control module 450 (not shown in the figure) of the server 400 can also run a feature recognition method so that the server 400 can obtain the image recognition result according to the feature recognition method.
[0103] In order to implement the feature recognition method, the display device 200 should include at least a display 260 and a processor 250. The display 260 can be configured to display a user interface. The processor 250 can execute the steps shown in FIG9 to perform feature recognition on the image to be recognized. As shown in FIG9, the following steps may be included:
[0104] S901: In response to an image recognition instruction, an image to be recognized is acquired.
[0105] The image recognition instruction is used to trigger the display device 200 to perform an image recognition function. For example, after receiving the image recognition instruction, the display device 200 can run a voice interaction application and an image recognition application. The image recognition instruction can be input in different ways depending on the interaction methods supported by different display devices 200.
[0106] In some embodiments, the image recognition instruction includes a field that is set to distinguish other voice instructions to represent the image recognition function. The display device can perform the image recognition function after determining the presence of the field. For example, the field can be "Nb.01".
[0107] In some embodiments, the image recognition instruction includes query parameters and feature parameters, and the display device can perform the image recognition function after determining that the query parameters are present.
[0108] In some embodiments, image recognition instructions can be input through voice interaction. For example, as shown in Figure 10, when the display device 200 displays a user interface such as a homepage interface, a media resource details interface, or a channel interface, the user can input voice instructions such as "Who is this person?", "Who is this actor?", "Who is the person on the left?", "Who is the person on the right?", "Who is the third person on the left?", "Who is this star?", "Who is this man?", "Who is this woman?", etc. Afterwards, the display device 200 can receive image recognition instructions. These image recognition instructions include at least voice data representing image recognition intent, and the voice data representing image recognition intent can be pre-set according to the language scenario.
[0109] In some embodiments, the voice data representing the image recognition intention may be "who is it".
[0110] In some embodiments, the voice data representing the image recognition intention may be "Who is it / ah".
[0111] In some embodiments, the voice data representing the image recognition intention may be "who".
[0112] In some embodiments, the voice data representing the image recognition intention may be “what is / are”.
[0113] In some embodiments, the voice data representing the image recognition intention may be field parameters containing preset vocabulary in the voice instruction generated based on the preset voice data, such as "TUSOU" or "Nb.02".
[0114] In some embodiments, the voice data representing the image recognition intent may be referred to as query parameters.
[0115] In some embodiments, the display device 200 may send the received voice data to the server, so that the server processes the voice data and then feeds it back to the display device.
[0116] In some embodiments, after receiving the voice data, the display device 200 may process the voice data itself.
[0117] In some embodiments, after receiving the voice data, the display device 200 may first perform voice recognition processing on the voice data to determine whether the voice data contains query parameters, and then determine whether to perform subsequent image recognition based on the voice recognition processing results. That is, after receiving the voice data, the display device 200 may also input the voice data into a voice recognition model, obtain the recognition text results output by the voice recognition model, and then extract the query parameters from the recognition text results. If the recognition text results include query parameters, an image recognition instruction is generated, and in response to the image recognition instruction, the image to be recognized is obtained. For example, the user inputs the voice data content of "Who is this person?" According to the voice recognition results, it can be determined that the current recognition text results include the query word "who is," then the display device 200 can generate an image recognition instruction, and in response to the image recognition instruction, take a screenshot based on the currently displayed user interface to obtain the image to be recognized.
[0118] If the recognized text result does not include a query parameter, the recognized text result is sent to the voice control module to determine the control intent of the current voice data, and to generate and respond to a voice control instruction. For example, if the user inputs voice data containing "Please play media asset A," and voice recognition determines that the voice data does not contain a query term (or trigger term) representing an object type, the display device 200 can determine that the control intent of the current voice data is not a feature query. Therefore, the display device 200 can skip the step of obtaining the image to be recognized and directly respond to the voice data by playing media asset A.
[0119] In some embodiments, the image recognition instruction may include, in addition to the query parameter, a feature parameter, which is used to enable the display device 200 or the server 400 to identify the features of the object to be recognized in the image to be recognized. The image recognition instruction may also not include the feature parameter, and the display device 200 or the server 400 may identify the features of the object to be recognized while performing image recognition on the object to be recognized in the image to be recognized.
[0120] In some embodiments, image recognition instructions can also be input through key interaction. For example, as shown in Figure 11, the display device 200 can display a file management application interface or a gallery interface, in which image files stored on the display device 200 can be displayed. The user can control the focus cursor by using the direction keys on the control device 100 to move the focus cursor and select an image for image display, forming an image display interface. The image display interface may include an "intelligent image recognition" control. After the user controls the focus cursor to select this control, the image recognition instruction can be input.
[0121] In some embodiments, the image recognition instruction can also be input along with the screenshot interaction operation. For example, the user can input the screenshot instruction in the form of a combination of the "menu key" and the "up arrow key". The display device 200 can respond to the screenshot instruction and take a screenshot of the currently displayed user interface to generate a screenshot image. While generating the screenshot image, the display device 200 can display the screenshot image through the screenshot display window. The "Smart Image Recognition" control can also be installed in the screenshot display window. After the user controls the focus cursor to select the control, the image recognition instruction can be input.
[0122] In some embodiments, the user can manipulate the cursor to trigger accurate identification of the target in the location area by placing the cursor in different positions, or can adjust the cursor position in the server feedback of multiple results to control the highlighting of the introduction of the target in the location area in the feedback results.
[0123] In some embodiments, after receiving the image recognition instruction, the display device 200 may obtain the image to be recognized. Depending on the method used to input the image recognition instruction, the display device 200 may use different methods to obtain the image to be recognized. For example, when the display device 200 is displaying the homepage interface, after the user inputs the image recognition instruction through a voice interaction instruction such as "Who is the person on the left?", the display device 200 may take a screenshot of the currently displayed homepage interface and obtain the image to be recognized through the screenshot.
[0124] Similarly, when the display device 200 displays the image display interface, after the user selects the "Smart Image Recognition" control through key operation and enters the image recognition instruction, the display device 200 can obtain the storage address of the currently displayed image, and obtain the image file through the storage address, thereby obtaining the image to be recognized.
[0125] In some embodiments, the method for acquiring the image to be recognized can be determined based on a combination of circumstances, such as the interaction method, the current display interface, and the current operating state. For example, when the display device 200 displays an image display interface, and the user inputs a voice command such as "Who is this person?", then, because the currently displayed interface is an image display interface and is associated with specific image information, the display device 200 does not need to perform a screenshot, but can directly acquire an image file based on the image information of the displayed image as the image to be recognized.
[0126] Therefore, the process shown in Figure 12 can be used to extract the image to be identified. Figure 12 is a schematic diagram of the process of obtaining the image to be identified according to some embodiments. As shown in Figure 12, the following steps may be included:
[0127] S1201: Detecting image information from an image recognition instruction;
[0128] S1202: If no image information is detected, calling the screenshot process;
[0129] S1203: taking a screenshot of the currently displayed user interface through a screenshot process to generate an image to be recognized;
[0130] S1204: If image information is detected, extract the image to be identified based on the image information.
[0131] In the process shown in FIG12 , the display device 200 can detect image information from the image recognition instruction. If no image information is detected in the image recognition instruction, a screenshot process is called, and a screenshot of the currently displayed user interface is taken through the screenshot process to generate an image to be recognized. If image information is extracted in the image recognition instruction, the image to be recognized is extracted based on the image information. The acquired image to be recognized may include an object to be recognized. The object to be recognized is a target in the image to be recognized that meets certain characteristics. For example, the object to be recognized may be a portrait target, a text target, a specific object target, an animal target, a symbolic pattern target, etc. in the image to be recognized.
[0132] It should be noted that when the display device 200 obtains the image to be recognized by taking a screenshot, the display device 200 can determine the execution time of the screenshot based on the acquisition time of the image recognition instruction. For image recognition instructions input in different ways, the display device 200 can determine different screenshot execution times based on different acquisition times.
[0133] In some embodiments, the display device 200 can classify the image recognition instructions according to the duration of the interaction action corresponding to the image recognition instruction, that is, the image recognition instruction whose interaction action duration is greater than or equal to the preset duration threshold is a first-category recognition instruction, and the image recognition instruction whose interaction action duration is less than the preset duration threshold is a second-category recognition instruction. For example, when the image recognition instruction is input through voice interaction, since the voice data is input through speaking behavior, the input process of the voice data consumes time, that is, the duration of the voice interaction action. The duration of the voice interaction action is greater than the preset duration threshold, therefore, the display device 200 can determine that the current image recognition instruction is a first-category recognition instruction, and can set the screenshot time to the end time of the voice data input according to the duration of the voice interaction action.
[0134] However, since the screen content in the user interface will change dynamically, that is, the content of the image to be recognized obtained at different screenshot times is different, therefore, in some embodiments, the display device 200 can also take a screenshot based on the reception time of the voice data, that is, the display device 200 can take a screenshot when the user starts to input voice data, so as to reduce the difference in user interface content caused by the time difference between the start and end of voice input, so as to reduce the impact of content differences on recognition results.
[0135] The voice input start time can be the start time of the entire voice data, or the time when the voice interaction content audio begins to appear in the voice data. For example, the control device 100 supporting the display device 200 is provided with a voice button, and the user inputs voice data by pressing the voice button. The display device 200 can start the screenshot process when the user presses the voice button to obtain the image to be recognized. For another example, if the user presses the voice button at time T1 and starts speaking at time T2, the display device 200 can detect the waveform change point in the audio signal and determine to start the screenshot process at time T2 when the voice waveform appears to obtain the image to be recognized.
[0136] When the display initiates a screenshot based on the start time of voice input, since some voice data is not used to perform the image recognition function, the display device 200 can run a screenshot process in the background. That is, after the screenshot process is initiated to capture the image to be recognized, the display device 200 can first temporarily store the captured image in a cache space, then perform voice detection on the voice data. If the voice data includes query parameters, the image to be recognized is extracted from the cache space for subsequent image recognition processing.
[0137] For example, if a user presses the voice button on control device 100 at time T1 and enters voice data such as "Who is the person on the left?", display device 200 can initiate a screenshot process at time T1, obtain a screenshot of the current user interface, and store it in cache as image P1 to be recognized. Furthermore, after receiving the complete voice data, display device 200 uses the voice recognition model to detect keywords contained in the voice data. Since the voice data includes the query parameter "who is", display device 200 can retain image P1 to be recognized in cache for subsequent feature recognition processing.
[0138] If the voice data does not include query parameters, the image to be recognized stored in the cache space can be deleted. For example, the user presses the voice button on the control device 100 at time T1 and inputs voice data with the content of "play next song", then the display device 200 can start the screenshot process at time T1, obtain a screenshot of the current user interface, and store it in the cache space as the image to be recognized P1. However, by detecting the keywords in the voice data, it is determined that the voice data does not contain query parameters such as "who, what, what are", so the display device 200 can determine that the current user's voice command does not have the intention of recognizing the image. At this time, the display device 200 can delete the image to be recognized P1 in the cache space.
[0139] For some display devices 200 that support far-field voice interaction, the screenshot process can also be started according to the input time of the wake-up word. That is, in some embodiments, the display device 200 can detect the wake-up word in the voice data in real time when receiving the voice data. When the voice data includes the wake-up word, the display device 200 can start the screenshot process to obtain the image to be identified at the moment the wake-up word is detected. For example, the user inputs voice data with the content "Hi! Xiao J, who is the person on the left?" Among them, "Hi! Xiao J" is the voice wake-up word of the display device 200, then the display device 200 can start the screenshot process to obtain the image to be identified when it detects the user input "Hi! Xiao J" at time T3.
[0140] S902: Extracting query parameters and feature parameters from the image recognition instruction.
[0141] In some embodiments, when the image recognition instruction includes a field for distinguishing other voice instructions to represent the image recognition function, the display device may execute step S902 after acquiring the image to be recognized.
[0142] In some embodiments, the display device may obtain the image to be recognized by extracting the query parameters in the image recognition instruction. Then, after obtaining the image to be recognized, step S902 does not need to be performed.
[0143] In some embodiments, after receiving an image recognition command input by a user, the display device 200 can parse, recognize, and respond to the image recognition command. For example, after a user inputs voice data such as "Who is this person?" through voice interaction, the display device 200 can first perform pre-processing on the voice data, including segmentation, extraction, noise reduction, and analog-to-digital conversion. The display device then uses the voice-to-text tool in the voice interaction system to convert the voice data into audio text. Furthermore, the display device 200 can obtain keyword information from the voice data through voice recognition algorithms such as word segmentation and semantic recognition.
[0144] In some embodiments, after obtaining keyword information, the display device 200 may filter the keyword information to determine the query parameters and feature parameters contained in the image recognition instruction. The query parameters are used to characterize the object type of the identified object, i.e., the query parameters enable the display device 200 or the server 400 to identify the object to be identified based on the specified object type; and the feature parameters are used to characterize the characteristic attributes of the identified object.
[0145] In some embodiments, the display device 200 can determine the query parameters and feature parameters in the image recognition instruction based on the part of speech of the keyword. That is, as shown in Figure 13, the display device 200 can obtain the image recognition instruction, convert the image recognition instruction into an instruction text, and perform word segmentation on the instruction text according to a preset vocabulary to obtain a keyword set. The keyword set includes at least one keyword, and the keyword has a preset part of speech. The preset part of speech includes a first type of part of speech and a second type of part of speech. The first type of part of speech is used for keywords that represent object types; the second type of part of speech is used for keywords that represent feature attributes. Therefore, the display device 200 can determine the query parameters and feature parameters in the keyword set based on the part of speech of the keyword.
[0146] In some embodiments, since the query parameter may be a query word (or trigger word) representing the object type, the first part of speech may be a pronoun or a combination of a noun, verb, etc. and a pronoun. The query parameter is used to trigger the screenshot operation and the image recognition program. For example, the query parameter may be at least one of the keywords that can represent the object type corresponding to the identified object, such as "who," "what," or "where."
[0147] Some of the above query parameters can be represented by some predetermined character strings.
[0148] In some embodiments, since the feature parameter may be a feature word representing a feature attribute, the second part of speech may be an adjective or a noun. The feature parameter is used to represent a keyword representing a feature attribute of an identified object. For example, the feature parameter may be at least one of the keywords representing a feature attribute of an identified object, such as "left side," "above," "third from the left," "fat," "tall," "left," "female," and the like.
[0149] In some embodiments, the query parameters may be represented by some predetermined character strings.
[0150] In some embodiments, the query term includes a keyword for characterizing an object type. For example, the object type of the identified object may be a person, an item, a place, an animal, or the like.
[0151] In some embodiments, the feature words include keywords used to characterize the characteristic attributes of the identified object. For example, the characteristic attributes of the identified object can be the orientation attributes of the identified object in the image to be identified (up, down, left, right, etc.), the relative orientation attributes between the identified objects, the fixed attributes associated with the identified objects (such as height, fatness, age, gender, etc.), the appearance attributes of the identified object (such as color, appearance of a person, appearance of an animal, etc.), and the action attributes of the identified object (such as standing, lying down, etc.).
[0152] In some embodiments, the feature attribute corresponding to the feature parameter can adopt the corresponding attribute detection method according to the attribute type. That is, in some embodiments, when the feature attribute is an orientation attribute, the feature parameter includes an orientation parameter, and the orientation parameter includes a physical orientation relative to the display range, such as left, right, top, bottom, top left, bottom left, top right, bottom right, center, etc. The display device 200 can detect the orientation of the identification object relative to the image to be identified. For example, with the center point of the image as the reference, the image is divided into four quadrants by the horizontal axis and the vertical axis. The first quadrant represents the right and top orientations of the image to be identified; the second quadrant represents the left and top orientations of the image to be identified; the third quadrant represents the left and bottom orientations of the image to be identified; and the fourth quadrant represents the right and bottom orientations of the image to be identified.
[0153] In some embodiments, when the feature attribute is a relative orientation attribute, the display device 200 can detect the relative positional relationship between the objects to be identified in the image to be identified. For example, the image to be identified includes two objects to be identified, object A and object B. The display device 200 can obtain the position coordinates of the objects to be identified in the image, namely A(x1, y1) and B(x2, y2). At this time, the display device 200 can determine the relative positional relationship between the two objects to be identified by comparing the position coordinates of the two objects to be identified. That is, if x1>x2, object A is located to the right of object B, and if y1>y2, object A is located above object B.
[0154] In some embodiments, after receiving an image recognition instruction input via voice interaction, the display device 200 or server 400 may generate a voice text and parse the semantic recognition result of the image recognition instruction based on the voice text. The semantic recognition result may include query parameters obtained based on the query term and feature parameters obtained based on the feature term. That is, the query parameter is the identifier or character corresponding to the query term, and the feature parameter is the identifier or character corresponding to the feature term. The corresponding control action response is then executed based on the semantic recognition result.
[0155] In some embodiments, the display device 200 obtains a target voice instruction collected by a sound collector based on a first application, and the target voice instruction is "Who is the person on the left?" The display device 200 calls the voice recognition module through the first application, and the voice recognition module performs text conversion and semantic understanding on it, and obtains the query word "who is" used to characterize the type of the recognition object, and the feature word "left" used to characterize the characteristics of the recognition object. Then, based on the parsed query word, it is determined that the target voice instruction is a voice instruction for triggering the screenshot and image recognition program. The query parameter "1" is used to identify the type of the query object as a person based on the query word, and the feature parameter "left" is generated through the feature word. Then, a command for calling the second application is generated through the query parameter, and the command includes the query parameter. The first application of the display device 200 then executes the corresponding program based on the command and query parameter for calling the second application. The display device 200 can generate an image recognition request and / or obtain an image to be recognized based on the semantic recognition result.
[0156] It is understood that the query parameters and feature parameters described above are merely exemplary parameter types, and the query parameters and feature parameters of the present disclosure may also be other characters, such as foreign characters, numeric characters, Chinese character labels, or a combination of one or two of the above characters. In some embodiments, the server 400 and the display device 200 may identify the query parameters and feature word parameters by setting specific flags in the data packet.
[0157] It should be noted that, during the execution of the feature recognition method, the display device 200 and the server 400 may execute specific program steps according to the configuration of the application corresponding to the method. Therefore, in response to the image recognition instruction, the steps of obtaining the image to be recognized (S901) and extracting the query parameters and feature parameters from the image recognition instruction (S902) may be performed entirely by the display device 200, or may be performed by the display device 200 and the server 400 in steps or stages.
[0158] In some embodiments, the display device 200 may be configured with a voice processing module. After receiving voice data, the display device 200 may call the voice processing module to perform text recognition on the voice data to obtain text recognition data. The display device 200 may then extract query parameters and feature parameters from the text recognition data, and perform a screenshot based on the query parameters to obtain an image to be recognized.
[0159] In some embodiments, the display device 200 can be connected to the server 400 via the communication device 220. The server 400 can be configured with a text recognition module. For example, the server 400 is a KD server with a voice-to-text recognition function. After receiving the voice data, the display device 200 can pre-process the voice data. The processed voice data is then sent to the server 400. The server 400 then performs text recognition based on the voice data uploaded by the display device 200 to generate recognized text data. The generated recognized text data is then sent to the display device 200, so that the display device 200 detects query parameters and feature parameters based on the recognized text data and performs feature recognition in the manner described in the above embodiments.
[0160] In some embodiments, for a display device 200 that is not equipped with a voice processing module, the server 400 can be connected through the communication device 220. The server 400 can be equipped with a text recognition module and a voice processing module. For example, the server 400 includes a KD server with a text recognition function and an HX voice server with a voice processing function. After receiving the voice data and pre-processing it, the display device 200 first sends the pre-processed voice data to the KD server. The KD server performs text recognition on the voice data, obtains the recognized text data and feeds it back to the display device 200. The display device 200 then sends a recognition request to the HX server based on the recognized text data. The HX server then performs feature extraction based on the text recognition data to determine the query parameters and feature parameters in the text recognition data. The HX voice server then sends the query parameters and feature parameters obtained by recognition to the display device 200, so that the display device 200 can perform a screenshot based on the query parameters and perform subsequent image recognition.
[0161] It should be noted that, for server 400, multiple functional modules or multiple servers can transmit relevant data. Therefore, after obtaining text recognition data, the server with text recognition functionality can directly transmit the text recognition data to the server with voice processing functionality. For example, the KD server performs text recognition on voice data, obtains recognized text data, and transmits it to the HX server. The HX server then performs feature extraction based on the text recognition data to determine query parameters and feature parameters in the text recognition data. The HX voice server then transmits the query parameters and feature parameters obtained by recognition to the display device 200.
[0162] In some embodiments, server 400 may have both text recognition and voice processing functions, that is, server 400 may be a voice server. After receiving voice data, display device 200 may upload the voice data to the voice server for voice recognition. After the voice server recognizes the text recognition data, it recognizes query parameters and feature parameters based on the text and then sends the recognized query parameters and feature parameters to display device 200. Display device 200 initiates a screenshot based on the query parameters and performs image recognition based on the image to be recognized and the feature parameters obtained from the screenshot.
[0163] S903: Perform image recognition on the image to be recognized according to the query parameters to obtain a recognition result.
[0164] In some embodiments, the display device performs image recognition on the image to be recognized according to the query parameters and obtains a recognition result.
[0165] In some embodiments, the display device sends the query parameters and the image to be recognized to a server, so that the server performs image recognition on the image to be recognized, and obtains a recognition result from the server.
[0166] In some embodiments, after extracting the query parameters and feature parameters, the display device 200 may perform image recognition on the image to be identified according to the query parameters to detect a feature target that matches the query parameters from the image to be identified. For example, if the user inputs the image recognition instruction "Who is the person on the left?", the corresponding query parameter is "who is it?". Therefore, the display device 200 may determine that the type of the identification object is a portrait based on the query parameters and detect the portrait target in the image to be identified based on the portrait features.
[0167] In some embodiments, image recognition can be performed by the server 400. That is, the server 400 can be configured with an image recognition model. After the display device 200 obtains the image to be recognized and extracts the query parameters and feature parameters from the image recognition instruction, it can generate an image recognition request based on the image to be recognized, the query parameters, and the feature parameters, and send the image recognition request to the server 400. The server 400 then calls the image recognition model according to the image recognition request and performs image recognition on the image to be recognized based on the image recognition model to obtain a recognition result.
[0168] In some embodiments, the display device 200 can generate a first image recognition request based on the query parameters and the feature parameters. The first image recognition request and the image to be recognized are then sent to the server 400, so that the server 400 responds to the first image recognition request, performs image recognition on the image to be recognized according to the query parameters, and performs screening on the recognition objects obtained by image recognition according to the feature parameters to determine the target recognition objects that meet the feature parameters and generate a first recognition result. The first recognition result includes the recognition objects that meet the query parameters, the target recognition objects that meet the feature parameters, and the arrangement parameters of the recognition objects. The display device 200 then receives the first recognition result fed back by the server 400 to display the recognition result.
[0169] In some embodiments, the recognition results obtained by performing image recognition can include various forms. In some embodiments, the recognition results include the identified objects that meet the query parameters and at least one of the following: arrangement parameters, feature labels, and feature details of the identified objects. The arrangement parameters are used to characterize the positional relationship of the identified objects in the image to be recognized; the feature labels are used to characterize the identification information of the identified objects; and the feature details are associated information queried based on the feature labels.
[0170] In some embodiments, when the display device 200 displays the homepage interface, the display device 200 extracts facial feature data from the image to be identified based on the query parameter "Who is it?" and determines feature tags based on the specific portrait patterns that meet the facial feature data, such as the names of the characters corresponding to the target person, such as "Actor 1, Actor 2, Actor 3." Simultaneously, the position of the identified object in the image to be identified is marked based on the position of the facial feature image to generate arrangement parameters, i.e., Actor 1 is located on the left side of the image, Actor 2 is located in the middle of the image, and Actor 3 is located on the right side of the image. The display device 200 can also extract feature details from the feature information database based on the determined feature tags, such as introduction information about the relevant characters.
[0171] In some embodiments, the result obtained by image recognition may also include all the information of the identified objects, and the display device 200 then filters the object information obtained by recognition through query parameters and feature parameters to obtain the recognition result. The display device 200 may generate a second image recognition request based on the query parameters. The second image recognition request and the image to be recognized are then sent to the server 400, so that the server 400 responds to the second image recognition request and performs image recognition on the image to be recognized according to the query parameters to generate a second recognition result. The second recognition result includes the identified objects that meet the query parameters and the arrangement parameters of the identified objects. The display device 200 then receives the second recognition result fed back by the server to display the recognition result.
[0172] In some embodiments, in response to the user input of voice data "Who is the person on the left?", the display device 200 may take a screenshot of the current user interface and, based on the query parameter "who is it" and the feature parameter "left", as well as the current screenshot image, generate an image recognition request and send it to the server 400. After the display device 200 sends the image recognition request to the server 400, the server 400 may, in response to the image recognition request, input the image to be recognized into the image recognition model. After the image recognition model detects the human targets in the image to be recognized, it may output the recognition objects "actor A, actor B..." in the image, as well as the position coordinates "A(x1, y1), B(x2, y2)..." of the recognition objects in the image. The server 400 then sends the model output data to the display device 200, which then filters the output data according to the query parameters and feature parameters to obtain a recognition result that meets the "who is it" and the feature parameter "left", that is, the human target information with the smallest x value in the position coordinates.
[0173] In some embodiments, the display device 200 can also perform image recognition on the image to be recognized through an image recognition model. That is, the display device 200 can call a pre-trained image recognition model. The image recognition model is a neural network model obtained by training a training image, and the training image is provided with a result label, and the result label is used to characterize the position of the recognized object in the training image. The image to be recognized is then input into the image recognition model to detect the recognized object in the image to be recognized and the position of the recognized object in the image to be recognized through the image recognition model. The feature label is then queried based on the recognized object, and the arrangement parameters are generated based on the position of the recognized object in the image to be recognized.
[0174] For example, after receiving an image recognition command asking "Who is the person on the left?", the display device 200 can respond to the image recognition command by invoking the image recognition model. The image recognition model can detect whether the image contains a human target, as well as the location and specific person of the human target. After the image to be recognized is input into the image recognition model, the image recognition model can detect the human target in the image based on the classification algorithm of the neural network model and output the person's name and location in the image.
[0175] In order to obtain feature details, the display device 200 can also perform information matching based on the feature tag and extract the feature details corresponding to the feature tag. In some embodiments, the feature detail information can be stored in the server 400, and the display device 200 can obtain the object type to which the identified object belongs. If the object type is the target type corresponding to the query parameter, the identified object is marked as an identified object to be displayed. The feature tag of the identified object to be displayed is then extracted, and a detail query request is generated based on the feature tag of the identified object to be displayed. The display device 200 then sends the detail query request to the server 400 to obtain the feature details corresponding to the feature tag to be displayed.
[0176] For example, if the image recognition instruction input by the user is "Who is this woman?", then based on the query parameter "who is", the type of the recognition object is determined to be a portrait, and at the same time, based on the feature parameter "female", the target type is determined to be "female". After performing image recognition based on the image recognition model and obtaining "portrait 1, portrait 2, portrait 3", the identified recognition objects are screened according to the target type, and it is determined that the character target corresponding to "portrait 2" is a female character, so portrait 2 can be determined as the recognition object to be displayed. The feature label corresponding to portrait 2 is then extracted, that is, the character's name is "actor 2". At this time, the display device 200 can generate a detail query request containing "actor 2" based on the feature label and send it to the server 400. After receiving the detail query request, the server 400 can parse the feature label "actor 2" in the detail query request and query the database for the actor introduction text related to "actor 2" to obtain the feature details.
[0177] In some embodiments, the above-mentioned image recognition process can be performed on a server.
[0178] S904: Generate a result display view based on the recognition result.
[0179] In some embodiments, after obtaining the recognition results, the display device 200 may generate a result display view based on the recognition results. The result display view is used to display the result information obtained during the image recognition process. Specifically, the result display view includes a feature label and feature details. The display position of the feature label is set according to the arrangement parameters, and the display status of the feature details is set according to the feature parameters.
[0180] In some embodiments, the arrangement parameter may be a position parameter of the identified object in the image.
[0181] For example, the result display view may include label display controls for "Actor 1, Actor 2, Actor 3, Actor 4, Actor 5, Actor 6," as well as a details display area below the label display controls. Because detailed information is also displayed in the form of text or images, it occupies a large display area. Therefore, in the result display view, only the feature details corresponding to some feature labels can be displayed. Furthermore, when displaying feature details, only partial text can be displayed, and a "More Information" control can be used to support further viewing of the displayed content.
[0182] FIG14 is a schematic diagram of a process for generating a result display view according to some embodiments. As shown in FIG14 , to generate a result display view, the display device 200 may perform the following steps to improve the display effect:
[0183] S1401: After obtaining the recognition result, call the result display template;
[0184] S1402: Adding a label control to the result display template according to the recognition result;
[0185] S1403: Setting the display position and / or display order of the label control according to the layout parameters;
[0186] S1404: Determine the focus label control according to the characteristic parameters;
[0187] S1405: Add feature details corresponding to the focus label control in the details display area associated with the focus label control.
[0188] In order to generate a result display view, the display device 200 can call the result display template after obtaining the recognition result through the process shown in Figure 14 above, and add a label control to the result display template according to the recognition result. Among them, the label control is used to display the feature label of the recognized object. Then, the display position and / or display order of the label control is set according to the layout parameters, and the focus label control is determined according to the feature parameters. The focus label control is the label control that meets the feature parameters. The focus identifier is set on the focus label control, and the feature details corresponding to the focus label control are added to the detail display area associated with the focus label control.
[0189] For example, as shown in FIG15 , after a user enters the image recognition command "Who is the person on the left?" through voice interaction, display device 200 extracts the query parameter "who" and the feature parameter "left" from the image recognition command. Then, through image recognition, the recognition results include the feature tags "Actor 1, Actor 2, Actor 3, Actor 4, Actor 5, Actor 6," as well as the feature details of each tag, namely, the personal profile text information of each actor (Actor 1, Actor 2, Actor 3, Actor 4, Actor 5, Actor 6).
[0190] Therefore, the display device 200 can first call the result display template and then add six label controls to the result display template, respectively used to display the names of actors 1 to 6. It then obtains the arrangement parameters of the recognition targets in the image to be recognized, that is, from left to right, actor 1, actor 2, actor 3, actor 4, actor 5, and actor 6. At this point, the display device 200 can adjust the display order of the added label controls according to the arrangement parameters.
[0191] After adding the label control, the display device 200 can also extract the feature parameter "left" from the image recognition request, and according to the positional relationship of the label controls in the result display view, determine that the focus label control is the "Actor 1" label control located on the left, and then add a focus mark to the "Actor 1" label control so that the focus cursor selects the "Actor 1" label control. At the same time, the display device 200 also sets the personal profile text information of "Actor 1" to a visible state and displays it in the feature details display area to present the image recognition results.
[0192] In some embodiments, when adjusting the display order of the added label controls according to the arrangement parameters, the display device 200 can further adjust the display order in combination with the feature parameters in the image recognition instruction. That is, when the feature parameters include feature words representing orientation attributes, the display device 200 first determines the first arrangement direction according to the orientation attributes represented by the feature parameters, and sets the first arrangement order of the label controls according to the first arrangement direction. For example, if the voice data content received by the display device 200 is "Who is the person on the left?", then according to the feature parameter "left" in the voice data, it can be determined that the first arrangement direction is the horizontal direction. Therefore, when setting the label display order, it can first be sorted in the horizontal direction based on the horizontal coordinate of the recognized object, and a result display view can be generated and displayed in the user interface.
[0193] In some embodiments, after displaying the result display interface, the display device 200 may further receive voice data and determine a second arrangement direction based on characteristic parameters in the voice data. When the voice data received again by the display device 200 includes characteristic parameters of the same type as those in the voice data received previously, the second arrangement direction is the same as the first arrangement direction. The display device 200 does not need to adjust the arrangement order and still displays the label controls in the display order determined by the first arrangement direction. For ease of description, in this embodiment, the characteristic parameters in the previously received voice data are referred to as first characteristic parameters, and the characteristic parameters in the again received voice data are referred to as second characteristic parameters.
[0194] In some embodiments, as shown in Figure 16, after the display device 200 displays the result display view generated based on the horizontal sorting of the horizontal coordinates of the recognized objects, it receives voice data with the content of "Who is the person on the right?" It can be seen that the second feature parameter "right" and the first feature parameter "left" contained therein are both directional words in the horizontal direction. Therefore, the display device 200 does not need to readjust the display order of the label controls, and still displays the label controls in the order previously determined, but moves the focus mark to the label control on the right and displays the detailed information corresponding to the label control on the right.
[0195] When the voice data received again by the display device 200 includes feature parameters of a different type from those in the voice data received last time, the display device 200 can reset the display order of the label controls according to the second feature parameters. For example, after the display device 200 displays the result display view generated based on the horizontal sorting of the horizontal coordinates of the identified object, it receives voice data with the content of "Who is the person above?" It can be seen that the second feature parameter "above" contained therein belongs to the directional word in the vertical direction, while the first feature parameter "left" belongs to the directional word in the horizontal direction. Therefore, the first feature parameter and the second feature parameter are not the same type of feature directions. The display device 200 needs to determine the second arrangement direction according to the second feature parameter, that is, the second arrangement direction is the vertical direction perpendicular to the first arrangement direction. The display device 200 then rearranges the display order of the label controls according to the vertical coordinates of the identified object based on the second arrangement direction, and sets the focus mark on the label control with the smallest vertical coordinate value.
[0196] In some embodiments, when the feature parameters include feature words of inherent attributes such as gender, the display device 200 can also filter the recognition results according to the attribute words of the inherent attributes, and add label controls according to the filtered recognition results. For example, if the voice data content received by the display device 200 is "Who is this man?", after obtaining the image recognition results "actor 1, actor 2, actor 3, actor 4, actor 5, actor 6" according to the image recognition method in the above embodiment, the display device 200 can read the inherent attributes of the recognition objects respectively to filter out the recognition objects "actor 1, actor 3, actor 4" whose inherent attribute is "male". Therefore, the display device 200 can add label controls corresponding to "actor 1, actor 3, actor 4" in the result display view to generate a result display view that only contains male actors, as shown in Figure 17.
[0197] In some embodiments, for the image to be identified whose image recognition result is "Actor 1, Actor 2, Actor 3, Actor 4, Actor 5, Actor 6", when the voice data content received by the display device 200 is "Who is this woman", the inherent attributes of the identification object can be read to filter out the identification objects "Actor 2, Actor 5, Actor 6" whose inherent attribute is "female". Therefore, the display device 200 can add label controls corresponding to "Actor 2, Actor 5, Actor 6" in the result display view to generate a result display view that only contains female actors, as shown in Figure 18.
[0198] In some embodiments, when the feature parameters include feature words with inherent attributes such as gender, the display device 200 can also add label controls for all recognized objects in the result display view, and indicate the recognized objects that meet the inherent attributes through focus marks. For example, for an image to be recognized whose image recognition result is "actor 1, actor 2, actor 3, actor 4, actor 5, actor 6", when the voice data content received by the display device 200 is "Who is this woman", the display device 200 can add 6 label controls to the result display interface according to the recognition result, and then determine the recognized objects "actor 2, actor 5, actor 6" with the inherent attribute of "female" by reading the inherent attributes of each recognized object, and set the focus mark on a label control among the recognized objects "actor 2, actor 5, actor 6", as shown in Figure 19.
[0199] It should be noted that when there are multiple recognition objects that meet the characteristic parameters, the display device 200 can set the focus mark according to the default order. For example, if the display device 200 receives voice data with the content "Who is this woman?" and the recognition objects with the inherent attribute "female" include "Actor 2, Actor 5, Actor 6", the display device 200 can set the focus mark on the label control corresponding to "Actor 2" in the default order from left to right, as shown in Figure 18.
[0200] In some embodiments, the default order can be set based on different application scenarios of the display device 200, that is, the default order can be different in different application environments. For example, some display devices 200 can set the default order to be from left to right; some display devices 200 can set the default order to be from top to bottom; some display devices 200 can set the default order to be from near to far distance from the center of the image; some display devices 200 can set the default order to be from large to small area occupied by the identified object in the image, etc.
[0201] In some embodiments, the display device 200 may also set a focus mark in combination with other feature words in the voice data. That is, the feature parameters extracted by the display device 200 from the image recognition instruction may include a first feature word and a second feature word. The display device 200 may then use the first feature parameter and the second feature parameter to filter the recognition objects, respectively, to determine the recognition objects that meet both the first feature word and the second feature word, and set a focus mark for the label control corresponding to the determined recognition object.
[0202] In some embodiments, for the image to be identified whose image recognition result is "actor 1, actor 2, actor 3, actor 4, actor 5, actor 6", the voice data content received by the display device 200 is "Who is the girl on the right", then the feature parameters of the voice data include two feature words, namely the first feature word "right" and the second feature word "girl". Then, the display device 200 can first determine the identification objects "actor 2, actor 5, actor 6" whose inherent attribute is "female" based on the feature word "girl". Then, based on the feature word "right", determine the identification object with the largest horizontal coordinate, namely "actor 6". As the identification object meets the feature parameters, set the focus mark on the label control corresponding to "actor 6", as shown in Figure 20.
[0203] In some embodiments, the feature parameters may also include a third feature word for characterizing the arrangement order. For example, for the image to be identified whose image recognition result is "Actor 1, Actor 2, Actor 3, Actor 4, Actor 5, Actor 6", the voice data content received by the display device 200 is "Who is the second male on the left?" Then the display device 200 can first filter out the identification objects "Actor 1, Actor 3, Actor 4" with the inherent attribute of "male" based on the feature word "male". Then, the screening direction is determined based on the feature word "left", and the identification object "Actor 3" with the second smallest horizontal coordinate value among "Actor 1, Actor 3, Actor 4" is determined as the identification object that meets the feature parameters based on the feature word "second". Therefore, the display device 200 can set the focus mark on the label control corresponding to "Actor 3", as shown in Figure 21.
[0204] It should be noted that the method in which the display device 200 selects and recognizes objects based on feature parameters in the above embodiment is merely an example. In actual applications, corresponding filtering rules can be set based on the different types of feature words contained in the voice data. For example, directional words such as up, down, left, right, center, upper left, and lower right can be used to replace the feature parameters in the voice data, and the display device 200 can then select and recognize objects based on specific directional relationships.
[0205] In some embodiments, the feature words in the feature parameters can also be pre-set with a screening priority. For example, the priority of the feature words is "inherent attribute, orientation attribute, and order attribute" from high to low. If the voice data content received by the display device 200 is "Who is the girl on the left?", then even if the first identification object on the left in the arrangement parameters is male, the focus mark is set on the first female identification object on the left, "Actor 2", in the result display view.
[0206] In some embodiments, the image recognition instruction input by the user may only include query parameters and no feature parameters. For example, the user input voice message "Who is this?" only includes the query parameter "who" and does not include content that identifies the feature attributes of the object to be recognized. To this end, after obtaining a keyword set, the display device 200 can traverse the parts of speech of the keywords in the keyword set; if the keyword set includes keywords of the first part of speech but does not include keywords of the second part of speech, image recognition is performed on the image to be recognized according to the query parameters to obtain a recognition result of the object to be recognized that meets the query parameters. A default layout template is then obtained, and a result display view is generated based on the default layout template and the recognition result.
[0207] For example, after the user inputs the voice "Who is this" without a feature parameter, the display device 200 can determine the segmentation result of the keyword set "this / who is" based on the word segmentation algorithm, and based on the word segmentation result, determine that the keyword set corresponding to the image recognition instruction includes the two keywords "this" and "who is". Then, by detecting the part of speech of each keyword in the keyword set, the keyword part of speech can be determined. Since the current keyword set does not include keywords with the part of speech of adjectives, that is, it does not include feature parameters, the display device 200 can identify the human target in the image to be identified according to the query parameters, and after obtaining the image recognition result, it does not filter the recognition result and directly generates the result display view according to the default layout template.
[0208] In some embodiments, the result display view may also include an object image for displaying the recognized object. For example, after the display device 200 displays the label control and label details according to the methods provided in the above embodiments, it may also add a recognition background image based on the image to be recognized in the result display view. Then, based on the position of the recognized object in the image, a rectangular frame representing the recognized object is added to the recognition background image to generate an object image.
[0209] When the image recognition process is performed by the server 400, the display device 200 may also generate a result display view in different ways based on the different recognition results fed back by the server 400. In some embodiments, after receiving the first recognition result, the display device 200 may parse the recognition object, the target recognition object, and the arrangement parameters of the recognition objects from the first recognition result. This allows the display device 200 to subsequently set the position of the feature label corresponding to the recognition object in the result display view and the position of the focus mark in the result display view according to the target recognition object based on the arrangement parameters.
[0210] In some embodiments, after receiving the second recognition result, the display device 200 may parse the recognition object and the arrangement parameters of the recognition object from the second recognition result. The recognition objects may be screened based on the characteristic parameters to determine target recognition objects that meet the characteristic parameters, so that the position of the characteristic label corresponding to the target recognition object in the result display view may be set based on the arrangement parameters. Furthermore, the position of the focus mark in the result display view may be determined based on the target recognition object.
[0211] S905: Control the display to display a result presentation view.
[0212] After the result display view is generated, the display device 200 may display the result display view, wherein the result display view may be an independent result display interface or a result display pop-up window that is displayed on the user interface.
[0213] In some embodiments, the display device 200 can also broadcast the image recognition results through the voice interaction system while displaying the result display view. That is, as shown in Figure 15, after the display device 200 generates the result display view through the above-mentioned feature recognition method, it can display the result display view in the form of a pop-up window, and broadcast the focus label control and the displayed feature details information through voice. For example, after the user inputs the image recognition instruction "Who is the person on the left" through voice interaction, the display device 200 can display a label control containing 6 feature labels "Actor 1, Actor 2, Actor 3, Actor 4, Actor 5, Actor 6", and the focus cursor is located on the "Actor 1" control, and the personal profile of "Actor 1" is displayed in the feature details display area. While displaying the above-mentioned result display view, the display device 200 also broadcasts through the voice system "The person on the left is Actor 1, Actor 1, male, born in ×× in 1989, a well-known actor, his representative works include "Movie A", "Movie B"... "
[0214] To facilitate user interaction, the displayed result display view may also include voice interaction text, such as the user's input voice text and the voice text broadcast by the voice interaction system. The voice interaction text can be located in the voice interaction area of the result display view and presented in the form of a dialogue. For example, the user's input voice text is located in the first line of the text display area, and the voice broadcast text is located in the second line of the text display area.
[0215] In some embodiments, the display device 200 also supports further user interaction during the display of the result display view. That is, after displaying the result display view, the display device 200 can obtain the recognition interaction instruction input by the user based on the result display view, and extract the change parameter from the recognition interaction instruction. Among them, the change parameter is a feature parameter in the recognition interaction instruction that is different from the image recognition instruction. Then, a new focus label control is determined based on the change parameter, that is, the label control that meets the change parameter is re-determined based on the interaction instruction. Then, the focus identifier is moved to the new focus label control, and the feature details corresponding to the new focus label control are added to the detail display area associated with the new focus label control.
[0216] For example, after the display device 200 responds to the image recognition instruction "Who is the person on the left?", during the display result display view, it can receive an interactive instruction input by the user with the content "Who is the person on the right?". The display device 200 can parse the query parameter "who" and the feature parameter "right" in the interactive instruction. Since the input interactive instruction also queries the same type of identification object, that is, both are portraits, the display device 200 can skip the step of image recognition to obtain the recognition result. And by comparing the feature parameters in the two instructions, it can be determined that the changed parameter is "right". At this time, as shown in Figure 16, the display device 200 can directly determine the new focus label control "Actor 6" according to the changed parameter "right", and set the feature details corresponding to "Actor 6", that is, the personal profile information of Actor 6, to a visible state, and set the feature details corresponding to Actor 1 to a hidden state, thereby moving the focus mark to the "Actor 6" label control, and displaying the personal profile information of Actor 6 in the details display area.
[0217] It can be seen that further interaction based on the result display view can enable the display device 200 to skip the image recognition step when the user inputs an image recognition instruction with the same query parameters, and directly change the focus label control and display the feature details information according to the previous image recognition result, which can improve the display efficiency of the image recognition results and enhance the user experience.
[0218] In some embodiments, the result display view can be set to a display duration. If the user does not continue to input interactive instructions related to image recognition within the set display duration, the result display view will automatically be canceled after displaying the set display duration. If the user inputs interactive instructions related to image recognition, the timer can be reset based on the input time of the interactive instructions to ensure that the result display view continues to be displayed, supporting further user interaction.
[0219] For example, if the display duration of the result display view is set to 20 seconds, the display device 200 can receive user interaction instructions within the 20 seconds of displaying the result display view. When the user enters an interaction instruction such as "Who is the person on the right?", the display device 200 can respond to the interaction instruction by re-determining the focus label control and updating the feature details, and continue to receive user interaction instructions within the next 20 seconds. If the user does not enter an interaction instruction within 20 seconds, the display device 200 will cancel the display of the result display view after displaying it for 20 seconds and continue to display the user interface before the image recognition interaction.
[0220] It can be seen from the above implementation methods that the display device 200 in the above embodiment can run the feature recognition method, and after inputting the image recognition instruction, set the display effect of the recognition result according to the query parameters and feature parameters in the image recognition instruction, so that the feature labels in the result display view presented by the display device 200 conform to the arrangement rules of the recognition objects in the image to be recognized, and the feature details conform to the feature parameters in the image recognition instruction, which is convenient for users to perform interactive control and improves the image recognition efficiency of the display device 200.
[0221] In some embodiments, some embodiments of the present disclosure further provide a feature recognition method applied to a display device, the method comprising:
[0222] In response to an image recognition instruction, an image to be recognized is acquired, wherein the image to be recognized includes an object to be recognized, and the image recognition instruction includes a query parameter and a feature parameter, wherein the query parameter is used to characterize the object type of the object to be recognized; and the feature parameter is used to characterize the feature attribute of the object to be recognized;
[0223] performing image recognition on the image to be recognized according to the query parameters extracted from the image recognition instruction to obtain a recognition result, the recognition result including an object to be recognized that meets the query parameters and at least one of an arrangement parameter, a feature label, and feature details of the object to be recognized, the arrangement parameter being used to characterize a positional relationship of the object to be recognized in the image to be recognized; the feature label being used to characterize recognition information of the object to be recognized; and the feature details being associated information queried based on the feature label;
[0224] generating a result display view according to the recognition result, wherein the result display view includes a feature label and feature details, wherein a display position of the feature label is set according to the arrangement parameter, and a display state of the feature details is highlighted according to the feature parameter;
[0225] The display is controlled to display the result presentation view.
[0226] In some embodiments, due to the low computing power of some display devices 200, the natural language processing and image recognition speeds are slow. Therefore, in some embodiments, the feature recognition process can also be performed by a server 400 that establishes a communication connection with the display device 200. That is, the server 400 provided by the present disclosure may include: a storage module 410, a communication module 420, and a control module 450. The storage module 410 is configured to store feature tags and feature details, and the communication module 420 is configured to establish a communication connection with the display device.
[0227] FIG22 is a flow chart of the process of performing feature recognition by the server 400 according to some embodiments. As shown in FIG22 , the process may include the following steps:
[0228] S2201: Obtaining an image recognition instruction and an image to be recognized sent by a display device;
[0229] The image to be identified includes an identification object; the image recognition instruction includes a query parameter and a feature parameter, the query parameter is used to characterize the object type of the identification object; the feature parameter is used to characterize the feature attributes of the identification object;
[0230] S2202: performing image recognition on the image to be recognized according to the query parameters to obtain a recognition result;
[0231] The recognition result includes the identified object that meets the query parameters and at least one of the arrangement parameters, feature labels, and feature details of the identified object. The arrangement parameters are used to characterize the positional relationship of the identified object in the image to be identified; the feature labels are used to characterize the identification information of the identified object; and the feature details are associated information based on the feature label query.
[0232] S2203: Send recognition results;
[0233] By sending the image recognition result to the real device 200, the display device 200 generates and displays a result display view based on the recognition result. The result display view includes feature labels and feature details. The display position of the feature label is set according to the arrangement parameters, and the display status of the feature details is highlighted according to the feature parameters.
[0234] During the feature recognition process, the display device 200 can receive an image recognition instruction input by the user, such as a voice instruction. And in response to the image recognition instruction, the image to be recognized is obtained by screenshot or file extraction, and the image recognition instruction and the image to be recognized are sent to the server 400. After receiving the image recognition instruction sent by the display device 200, the server 400 can parse the semantic recognition result of the image recognition instruction. Among them, the semantic recognition result may include query parameters obtained based on the query words, and feature parameters obtained based on the feature words. That is, the query parameters are the identifiers or characters corresponding to the query words, and the feature parameters are the identifiers or characters corresponding to the feature words. The semantic recognition result is then sent to the display device 200, so that the display device 200 executes the corresponding program according to the semantic recognition result.
[0235] In some embodiments, the display device 200 obtains an image recognition instruction collected by a sound collector based on a first application, and the image recognition instruction is "Who is the person on the left?" The display device 200 sends the instruction to the voice recognition module 430 of the server 400 through the first application. The voice recognition module 430 performs text conversion and semantic understanding on it, and obtains the query word "who" used to characterize the type of the recognition object and the feature word "left" used to characterize the characteristics of the recognition object. Then, based on the query word parsed above, the server 400 determines that the image recognition instruction is a voice instruction for triggering the screenshot and image recognition program. Based on the query word new word query parameter "1", it is used to identify the type of the query object as a person, and at the same time, the feature parameter "left" is generated through the feature word. The query parameter is then used to generate a command to call the image recognition model, which includes the query parameter. The server 400 then inputs the image to be recognized into the image recognition model and obtains the image recognition result based on the image recognition model.
[0236] In some embodiments, after obtaining the image recognition result, the server 400 can extract the feature parameters in the image recognition instruction and filter the image recognition result according to the feature parameters to generate a result display view. For example, if the image recognition instruction is "Who is the person on the left?", the query word is "who," the feature word is "left," the type of identification object represented by the query parameter is a person, and the feature attribute represented by the feature parameter is the orientation attribute "left." The server 400 then parses all the person objects in the image to be identified and obtains the associated feature tags of each person object. Then, based on the layout parameters of the identification object in the image to be identified, the position coordinates of each person object are determined to filter out the leftmost person object and the name (feature tag) corresponding to the person object.
[0237] In some embodiments, when a feature word represents an orientation feature, the feature parameter includes an orientation parameter, which is used to represent the relative position of the identification object in the image to be identified (such as left, right, top, bottom, upper left, lower left, upper right, lower right, center, etc.). After parsing the identification object in the image to be identified, the server 400 obtains a target coordinate system for image recognition. The origin of the target coordinate system can be located at the geometric center point of the user interface of the display device 200, or at any boundary vertex in the user interface of the display device 200; as shown in Figure 23, corresponding to the image to be identified, the origin of the target coordinate system is located at the geometric center point of the image to be identified, or at a boundary vertex of the image to be identified. Then, based on the target coordinate system, the position coordinates of the identification object in the image to be identified are detected. The position coordinates include coordinate values in a first direction and coordinate values in a second direction, where the first direction is perpendicular to the second direction, such as the first direction being the horizontal axis and the second direction being the vertical axis. After the server 400 detects the position coordinates of the identification object in the image to be identified, it records the position coordinates, such as (x, y), in the associated information of the identification object.
[0238] In some embodiments, while the server 400 performs image recognition on the image to be recognized according to the query parameters, it also detects the position of the recognition object in the image to be recognized, and records the position coordinates as associated information in the server 400 for the server 400 to perform subsequent comparison and processing.
[0239] In some embodiments, when the feature word contains the semantics of left, the relative position represented by its orientation parameter is left, such as the feature words such as left, left side, and left direction; when the feature word contains the semantics of right, the relative position represented by its orientation parameter is right, such as the feature words such as right, right side, and right direction; when the feature word contains the semantics of up, the relative position represented by its orientation parameter is up, such as the feature words such as right, right side, right side, and top; when the feature word contains the semantics of down, the relative position represented by its orientation parameter is down, such as down, below, below, and bottom. When a feature word contains the semantics of upper left, the relative position represented by its position parameter is upper left, such as upper left, upper left corner and other feature words; when a feature word contains the semantics of upper right, the relative position represented by its position parameter is upper right, such as upper right, upper right corner and other feature words; when a feature word contains the semantics of lower left, the relative position represented by its position parameter is lower left, such as lower left, lower left corner and other feature words; when a feature word contains the semantics of lower right, the relative position represented by its position parameter is lower right, such as lower right, lower right corner and other feature words.
[0240] In some embodiments, the position coordinates of the identified object recorded by the server 400 can be the position coordinates of any pixel point of the identified object in the image to be identified, or can be the coordinates of the edge pixel points of the identified object, such as the lower left corner, the upper right corner, the leftmost pixel point, the rightmost pixel point, the top pixel point, the bottom pixel point, the pixel point of the geometric center, etc.
[0241] In some embodiments, the screening process can be performed by comparing the coordinates of each target object. For example, the X coordinates of each target object can be used to determine which is on the left and which is on the right, and the Y coordinates of each target object can be used to determine which is on the top and which is on the bottom.
[0242] In some embodiments, the screening process can also be performed by determining an initial range based on the first characteristic parameter of the characteristic parameters, and then continuing to screen again within the initial range based on another characteristic parameter. Alternatively, the screening can be performed simultaneously based on multiple characteristic parameters. For example, a preliminary screening is performed based on the characteristic word "male", and then a second screening is performed based on the characteristic parameters of "left side" or "first on the left".
[0243] In order to extract the recognition results more accurately, in some embodiments, the server 400 can also record the position coordinates of the identified object according to the orientation parameters. When the orientation parameters are different, the position coordinates recorded by the server 400 are also different, and the position coordinates can be consistent with the relative position represented by the orientation parameters. For example, when the relative position represented by the orientation parameters is the upper left, the server 400 records the coordinates of the upper left pixel of the identified object; when the relative position represented by the orientation parameters is the lower left, the server 400 records the coordinates of the lower left pixel of the identified object; when the relative position represented by the orientation parameters is the left, the server 400 records the coordinates of the leftmost pixel of the identified object. In some embodiments, when the relative position represented by the orientation parameters is other relative positions, the principle of the server 400 recording the position coordinates is the same as the above principle, which will not be repeated here.
[0244] In some embodiments, when the feature parameters include orientation parameters and the origin of the target coordinate system is located at the center of the object to be identified, the server 400 may further calculate a cropping region of the image to be identified based on the relative position represented by the feature parameters, and crop the image to be identified based on the cropping region to preserve the image content of the cropped region in the image to be identified. Image recognition is then performed on the cropped image to be identified, thereby reducing the amount of data computation performed by the server 400.
[0245] In some embodiments, when the feature word is "left," the feature parameters include orientation parameters, and the origin of the target coordinate system is located at the geometric center of the image to be identified, the server 400 calculates that the cropped area of the image to be identified is the image portion in the second and third quadrants based on the relative position "left" represented by the feature parameters, and the server 400 crops out the image portion of the image to be identified that is located in the first and fourth quadrants, and retains the image portion in the second and third quadrants. For another example, when the feature word is "upper left," the feature parameters include orientation parameters, and the origin of the target coordinate system is located at the geometric center of the image to be identified, the server 400 calculates that the cropped area of the image to be identified is the image portion in the second quadrant based on the relative position "upper left" represented by the feature parameters, and the server 400 crops out the image portion of the image to be processed that is located in the first, third, and fourth quadrants, and retains the image portion in the second quadrant.
[0246] In some embodiments, when the server 400 calculates the cropping area of the image to be identified, it obtains the coordinate critical value of each quadrant and determines the image portion of the image to be identified in each quadrant based on the coordinate critical value of each quadrant, so that the server 400 can crop the image to be identified. The coordinate critical value is used to represent the coordinate value characteristics of the position coordinates of the pixel points corresponding to each quadrant. For example, when the origin of the target coordinate system is located at the geometric center point of the image to be identified, the coordinate critical value of the first quadrant is x>0 and y>0, the coordinate critical value of the second quadrant is x<0 and y>0, the coordinate critical value of the third quadrant is x<0 and y<0, and the coordinate critical value of the fourth quadrant is x>0 and y<0.
[0247] To facilitate the extraction of target recognition results, in some embodiments, the relative position represented by the orientation parameter includes at least one of left, right, top, bottom, and center, the first direction represents the horizontal direction, and the second direction represents the vertical direction. The server 400 also analyzes the critical trend corresponding to the orientation parameter. The critical trend includes the maximum or minimum end of the first direction, and / or the maximum and minimum ends of the second direction. For example, when the first direction is the x-direction and the second direction is the y-direction, and the origin of the target coordinate system is located in the lower left corner of the user interface, when the physical range is left, the critical trend is the minimum value in the x-direction; when the physical range is top, the critical trend is the maximum value in the y-direction; when the physical range is top left, the critical trend is the minimum value in the x-direction and the maximum value in the y-direction. Then, the coordinate values of the recognition object are read from the layout parameters, and the recognition object whose coordinate value is the target value is extracted from the recognition object to determine the focus feature label. The target coordinate value is the coordinate value of the recognition object at the critical trend, that is, the recognition object whose coordinate value matches the critical trend.
[0248] In some embodiments, the origin of the target coordinate system is located in the lower left corner of the image to be identified, the object type represented by the query parameter is a person, and the relative position represented by the feature parameter is on the left. As shown in Figure 24, the server 400 performs image recognition on the image to be identified, identifies all the people it includes, and records the position coordinates (x1, y1) of person 1, the position coordinates (x2, y2) of person two, and the position coordinates (x3, y3) of person three. According to the relative position represented by the orientation parameter and the target coordinate system, the critical trend is analyzed to be the minimum value in the x direction. Then, the server 400 traverses the coordinate values of the x coordinates in the associated information corresponding to each identified object, and extracts the identified object with the largest coordinate value in the x coordinate. x3>x1>x2, so the person three corresponding to (x3, y3) is filtered out as the focus feature label, and then the associated information corresponding to person three is extracted as the target recognition result.
[0249] In some embodiments, the server 400 may also obtain a boundary set of the image to be identified in the target coordinate system. The boundary set includes the left boundary, right boundary, upper boundary, and lower boundary of the image to be identified. When determining the focus feature label, the server 400 selects a reference boundary based on the physical range and calculates the relative distance between the identified object and the reference boundary based on the position coordinates of the identified object. The identified object whose relative distance matches the relative position (relative position represented by the orientation parameter) is then extracted to determine the focus feature label.
[0250] In some embodiments, the relative position represented by the orientation parameter is left, and the reference boundary can be the left boundary. The server 400 calculates the relative distance between each identified object and the reference boundary and extracts the feature tag of the identified object with the smallest relative distance as the focus feature tag. For another example, if the physical range is left, and the reference boundary can be the right boundary, the server 400 calculates the relative distance between each identified object and the reference boundary and extracts the identified object with the largest relative distance as the focus feature tag.
[0251] In some embodiments, when the relative position represented by the orientation parameter includes two or more items, the server 400 further reads a first coordinate value of the identified object in a first direction and a second coordinate value of the identified object in a second direction. A first judgment value is generated based on the critical trend and the first coordinate value, and a second judgment value is generated based on the critical trend and the second coordinate value. The sum of the first judgment value and the second judgment value is calculated, and the identified object with the largest sum is extracted from the identified objects to determine the focus feature label.
[0252] In some embodiments, the server 400 can perform normalization from left to right and from bottom to top on the coordinate values of the first direction and the second direction based on the width and height of the image to be identified, and generate a first judgment value and a second judgment value based on the normalized coordinate values.
[0253] In some embodiments, when the relative position represented by the orientation parameter is the upper left, the server 400 reads the normalized coordinate value of the identified object in the x-direction and the normalized coordinate value of the identified object in the y-direction. The critical trend is the minimum end of the first direction and the minimum end of the second direction. The server 400 calculates the inverse value of the first coordinate value and the second coordinate value (i.e., the value after 1-coordinate value) as the first judgment value and the second judgment value. The sum of the first judgment value and the second judgment value is then calculated, and the identified object with the largest sum is extracted from the identified objects to determine the focus feature label.
[0254] In some embodiments, when the relative position characterized by the orientation parameter is the lower left, the server 400 reads the normalized coordinate value of the recognition object in the x direction and the normalized coordinate value of the recognition object in the y direction. The critical trend is the minimum end of the first direction and the maximum end of the second direction. The server 400 will calculate the opposite value of the first coordinate value (i.e., 1 - the normalized value of the coordinate value) as the first judgment value, and the normalized coordinate value of the second coordinate value as the second judgment value. Then calculate the sum value of the first judgment value and the second judgment value, and extract the recognition object with the largest sum value in the recognition objects to determine the focus feature label.
[0255] In some embodiments, the sum value calculation for other relative positions is based on the same principle as the above embodiments, except that it is for the comparison between recognition objects rather than the comparison with a specified range. Exemplarily, taking the x - axis increasing from left to right and the y - axis decreasing from top to bottom as an example. When the position coordinates of person one are (x1, y1), the position coordinates of person two are (x2, y2), and the position coordinates of person three are (x3, y3), in some embodiments, x1 < x2 < x3, person one is on the left and person three is on the right. In some embodiments, y1 < y2 < y3, person one is on the top and person three is on the bottom. In some embodiments, person one is female, person two is male, and person three is male. The server 400 first determines that the males are person two and person three, and then according to the corresponding abscissa x2 < x3, obtains that the second boy on the left is person 3.
[0256] In some embodiments, when the feature word includes an inherent attribute, the feature parameter includes a basic parameter, and the basic parameter is used to characterize the inherent attribute of the recognition object, such as color attribute, recognition object name, etc. After the server 400 analyzes the recognition object and obtains the associated information of the recognition object, it queries the associated information representing the inherent attribute in the associated information, so as to determine the focus feature label in the recognition object according to the associated information of the inherent attribute.
[0257] In some embodiments, the feature parameter includes a basic parameter, and the basic parameter is used to characterize the inherent attribute corresponding to the recognition object. The server 400 also analyzes the target inherent attribute corresponding to the basic parameter and then reads the reference value of the target inherent attribute. For example, taking the object type represented by the query parameter as a person, the inherent attribute can include the gender of the person, the name of the person, the skin color of the person, etc., and the target inherent attribute can be a combination of one or more of the above; when the target inherent attribute is gender, the reference value can be "male / female". After reading the reference value of the target inherent attribute, extract the recognition object in which the inherent attribute matches the reference value to determine the focus feature label.
[0258] In some embodiments, when the object type represented by the query parameter is an item and the inherent attribute represented by the feature parameter is red, server 400 may obtain the color value corresponding to red (e.g., RGB value) and use the color value corresponding to red as a reference value. Next, the server 400 queries the associated information for the color value corresponding to each item and calculates the similarity between the item color value and the reference value. Items whose color values match the reference value (which can be the most similar or completely identical) are then extracted to determine the focal feature label.
[0259] In some embodiments, the feature attributes include extrinsic parameters, which represent the appearance attributes and / or action attributes (standing, lying down, etc.) of the recognition object. For example, when the object type represented by the query parameter is a person, the appearance attributes may be weight, long hair, short hair, wearing glasses, etc.; when the object type represented by the query parameter is an item, the appearance attributes may be shape, etc. The server 400 obtains the target template corresponding to the extrinsic parameter and calculates the similarity between the target template and the recognition object. The recognition object with the highest similarity among the recognition objects is then extracted as the second object. For example, when the appearance attribute represented by the extrinsic parameter is a triangle, the target template obtained by the server 400 includes templates of various triangles, and then the similarity between each recognition object and the target template is calculated separately.
[0260] In some embodiments, server 400 may also detect the width of the object to be identified in the image to be identified and record the detected width in the associated information. Thus, when determining the focal feature label, server 400 can extract the object that matches the characteristic attribute represented by the feature parameter based on the width of the object to be identified. For example, if the query parameter represents the object type as a person, and the characteristic parameter represents the appearance attribute as thin, server 400 can query the associated information for the width values corresponding to each person and extract the person with the smallest width value to determine the focal feature label.
[0261] In some embodiments, when the object type represented by the query parameter is a person, when the server 400 records the width of the identified object in the image to be identified, it can detect the face area of the person based on the image recognition algorithm, and then record the width of the face area as the width corresponding to the identified object.
[0262] In some embodiments, the characteristic parameters may include a combination of one or more of the above parameters.
[0263] In some embodiments, when the feature parameters include multiple parameters, server 400 may also preset weights for the different parameters. When determining a focal feature label, server 400 may calculate a comprehensive score for the identified object based on the weights corresponding to the feature parameters, and then extract the identified object with the highest comprehensive score to determine the focal feature label. The implementation logic of the feature parameters is the same as described above and will not be further elaborated here.
[0264] In some embodiments, after generating the result display view, the server 400 can encapsulate the interface parameters or interface data corresponding to the result display view into a response data packet, and send the response data packet to the display device 200. After receiving the response data packet, the display device 200 parses the interface parameters or interface data from the response data packet, and forms a result display view based on the parsed data for display. In some embodiments, after receiving the response data packet, the display device 200 can also generate a prompt corresponding to the result display view based on the result display view, and display the prompt in the result display view. For example, when the result type specified by the query parameter is a person type, the display device 200 can generate a prompt related to the result display view, such as "What are his related works?", "Recommend a few of his movies", etc.
[0265] In some embodiments, the server 400 may complete the screening of the recognition results and generate rendering data corresponding to the result display view so that the display device displays the result display view. That is, the control module of the server 400 is further configured to:
[0266] Obtaining an image recognition instruction and an image to be recognized sent by a display device, wherein the image to be recognized includes an object to be recognized, and the image recognition instruction includes a query parameter and a feature parameter, wherein the query parameter is used to characterize the object type of the object to be recognized; and the feature parameter is used to characterize the feature attribute of the object to be recognized;
[0267] Perform image recognition on the image to be identified according to the query parameters to obtain a recognition result;
[0268] Generate a result display view based on the recognition result. The result display view includes feature labels and feature details. The display position of the feature labels is set according to the arrangement parameters, and the display status of the feature details is set according to the feature parameters.
[0269] The rendering data corresponding to the result display view is sent to the display device, so that the display device displays the result display view.
[0270] The difference between this embodiment and the above-mentioned server 400 embodiment is that, in this embodiment, the server 400 can perform image recognition on the image to be recognized, and after obtaining the recognition result, generate rendering data corresponding to the result display view on the server 400 end, and send the rendering data to the display device 200, so that the display device 200 displays the result display view.
[0271] For example, server 400 may be configured with an image recognition model. Upon receiving voice data such as "Who is the second boy on the left?", display device 200 may invoke the voice recognition model to identify the query parameter "who" and the feature parameters "left," "second," and "boy" from the voice data. Based on the identified query parameters, the system then invokes a screenshot process to take a screenshot of the currently displayed user interface to obtain the image to be recognized. The system then generates an image recognition request and sends the image recognition request and the image to be recognized to server 400.
[0272] After receiving the image recognition request, the server 400 can respond to the image recognition request by calling the image recognition model to perform image recognition on the image to be recognized to obtain a recognition result containing portrait features, and then send the recognition result to the display device 200 so that the display device 200 generates a result display view based on the recognition result.
[0273] In some embodiments, when server 400 feeds back the recognition results to display device 200, it can directly feed back the entire recognition results for the image. For example, in response to an image recognition request sent by display device 200 based on voice data asking "Who is the second boy on the left?", server 400 can perform image recognition on the image to be recognized based only on the query parameter "who is" to identify the portrait targets contained in the image and the location information of the portrait targets in the image. All recognized portrait targets and their location information are then sent to display device 200. Display device 200 then determines the portrait targets that meet the requirements of "left," "second," and "boy" based on the feature parameters to generate a result display view and display it.
[0274] In some embodiments, the server 400 may also filter the recognition results based on the feature parameters and provide feedback on the filtered recognition results. That is, after identifying and obtaining the human portrait target contained in the image and the position information of the human portrait target in the image, the server 400 may filter the human portrait target and position information based on the feature parameters to generate an image recognition result. The recognition result includes the objects that meet the query parameters, as well as the arrangement parameters, feature labels, and feature details of the objects. The arrangement parameters are used to characterize the positional relationship of the objects in the image to be recognized; the feature labels are used to characterize the recognition information of the objects; and the feature details are associated information based on the feature label query.
[0275] For example, the server 400 performs image recognition on the image to be identified based on the query parameter "who is it" to identify the portrait target contained in the image and the position information of the portrait target in the image. It can then filter the identified portrait targets based on the feature parameters "left", "second", and "boy" to determine the portrait target that meets the feature parameters of "left", "second", and "boy", namely, "Actor 3", and then send the recognition result containing "Actor 3" to the display device 200, so that the display device 200 generates a result display view containing "Actor 3" based on the recognition result.
[0276] In some embodiments, the server 400 may perform the recognition processing of the image recognition instruction, and the display device 200 may perform the image recognition and generate the result display view. That is, the control module of the server 400 may also be configured as follows:
[0277] Receiving a voice recognition request sent by a display device, where the voice recognition request includes voice data collected by the display device;
[0278] performing speech recognition on the speech data to generate speech text data;
[0279] The invention provides a method for transmitting voice and text data to a display device, so that the display device first generates an image recognition instruction based on the voice and text data, and obtains an image to be recognized in response to the image recognition instruction, wherein the image to be recognized includes an object to be recognized; secondly, extracts query parameters and feature parameters from the image recognition instruction; the query parameters are used to characterize the object type of the recognition object; the feature parameters are used to characterize the feature attributes of the recognition object; thirdly, performs image recognition on the image to be recognized according to the query parameters to obtain a recognition result, wherein the recognition result includes the recognition object that meets the query parameters and the arrangement parameters, feature labels, and feature details of the recognition object; the arrangement parameters are used to characterize the positional relationship of the recognition object in the image to be recognized; the feature labels are used to characterize the recognition information of the recognition object, and the feature details are associated information based on the feature label query; and finally, generates a result display view based on the recognition result, wherein the result display view includes the feature labels and feature details, the display position of the feature labels is set according to the arrangement parameters, and the display status of the feature details is set according to the feature parameters. Finally, the display is controlled to display the result display view.
[0280] As can be seen, the difference between this embodiment and the aforementioned server 400 embodiment is that in this embodiment, server 400 is only responsible for performing recognition on the language data, while image recognition, result display view generation, and display are all performed by display device 200. Based on this, server 400 can have a built-in voice processing module, and display device 200 can have a built-in image processing module and view generation module.
[0281] For example, the server 400 may be configured with a speech recognition model. After receiving speech data such as "Who is the second boy on the left?", the display device 200 generates a speech recognition request based on the speech data and sends the speech recognition request to the server 400. In response to the speech recognition request, the server 400 calls the speech recognition model to recognize the speech data, obtains a speech recognition result, and sends it to the display device 200.
[0282] In some embodiments, when the server 400 feeds back the speech recognition result to the display device 200, it can also feed back the recognition text of the speech data. For example, for the speech data with the content of "Who is the second boy on the left", the recognition text of "left / second / boy / who is" can be obtained, and the recognition text can be fed back to the display device 200. The display device 200 then extracts the query parameter "who is" and the feature parameters "left", "second", and "boy" from the recognition text. Then, the screenshot process is called according to the recognized query parameters, and a screenshot is taken of the currently displayed user interface to obtain the image to be recognized. Then, the image recognition result and the result display view are generated in the manner in the above embodiment, which will not be repeated here.
[0283] In some embodiments, when the server 400 feeds back the voice recognition results to the display device 200, it can also feed back the query parameters and feature parameters identified in the voice data. For example, for voice data with the content "Who is the second boy on the left", the server 400 can first obtain the recognition text of "left / second / boy / who is" through the voice recognition model, and then extract the query parameter "who is" and the feature parameters "left", "second", and "boy" from the recognition text based on the part of speech of the keywords in the recognition text or the preset vocabulary. The extracted query parameters and feature parameters are then sent to the display device 200, so that the display device 200 generates and displays the result display view in the manner described in the above embodiment.
[0284] In some embodiments, the server 400 can perform both speech recognition and image recognition. That is, the control module of the server 400 is further configured to:
[0285] Receiving a voice recognition request sent by a display device, where the voice recognition request includes voice data collected by the display device;
[0286] performing speech recognition on the speech data to generate speech text data;
[0287] Sending voice and text data to a display device, so that the display device first generates an image recognition instruction based on the voice and text data, and in response to the image recognition instruction, obtains an image to be recognized, where the image to be recognized includes an object to be recognized; extracts a query parameter and a feature parameter from the image recognition instruction, where the query parameter is used to represent an object type of the object to be recognized; and the feature parameter is used to represent a feature attribute of the object to be recognized; and generates an image recognition request based on the image to be recognized;
[0288] receiving an image recognition request sent by a display device, where the image recognition request includes an image to be recognized;
[0289] Perform image recognition on the image to be identified to obtain a recognition result. The recognition result includes the identified objects that meet the query parameters, as well as the arrangement parameters, feature labels, and feature details of the identified objects. The arrangement parameters are used to characterize the positional relationship of the identified objects in the image to be identified; the feature labels are used to characterize the identification information of the identified objects, and the feature details are the associated information based on the feature label query.
[0290] The recognition result is sent to the display device so that the display device 200 generates and displays a result display view according to the recognition result. The result display view includes a feature label and feature details. The display position of the feature label is set according to the arrangement parameter, and the display status of the feature details is set according to the feature parameter.
[0291] It can be seen that the difference between this embodiment and the above embodiments is that, in this embodiment, the server 400 can perform voice recognition and image recognition respectively, and the display device 200 performs voice data collection, generation of image recognition instructions, acquisition of images to be recognized, and generation and display of result display views, etc.
[0292] It should be understood that the other action steps involved in the above examples can also be divided and processed by the display device 200 and the server 400 according to the actual device layout. For example, steps such as the generation of image recognition instructions, the acquisition of images to be recognized, and the generation of result display views can also be performed by the server 400. Similarly, voice recognition and image recognition can also be performed by the display device. Other execution subject change schemes that can be associated by those skilled in the art based on the above embodiments also fall within the scope of protection of this disclosure. In addition, some functions can also be performed jointly by the display device 200 and the server 400 to maximize the processing capabilities of the display device 200 and the server 400, which will not be repeated here.
[0293] It can be seen from the above embodiments that the display device of the embodiment of the present disclosure can obtain the image to be identified in response to the image recognition instruction. And extract the query parameters and feature parameters from the image recognition instruction. Then perform image recognition on the image to be identified according to the query parameters to obtain the recognition result, and generate a result display view based on the recognition result, so as to display the result display view. Among them, the result display view includes feature labels and feature details, the display position of the feature labels is set according to the arrangement parameters of the identification object in the image to be identified, and the display status of the feature details is set according to the feature parameters in the image recognition instruction. After the image recognition instruction is input, the method can set the display effect of the recognition result according to the query parameters and feature parameters in the image recognition instruction, so that the feature labels in the result display view presented by the display device conform to the arrangement rules of the identification object in the image to be identified, and the feature details conform to the feature parameters in the image recognition instruction, so as to facilitate the user to perform interactive control, so as to solve the problem of low overall efficiency of image recognition of the display device.
[0294] Based on the above application scenario, in order to improve the efficiency of image recognition, the server 400 of the embodiment of the present disclosure may further execute the steps shown in FIG. 25 to perform feature recognition on the image to be recognized. As shown in FIG. 25 , the steps may include:
[0295] S2501: Receive an image to be recognized and an image recognition request from a display device.
[0296] Among them, the image to be identified is an image generated by the display device 200 taking a screenshot of the user interface. The image to be identified includes an identification object. The image recognition request includes query parameters and feature parameters. The query parameters are used to characterize the object type of the identification object; the feature parameters are used to characterize the feature attributes of the identification object.
[0297] After establishing a communication connection with the display device 200, the server 400 can receive various data packets sent by the display device 200. The data packets may include image data, voice data, and various data requests. After receiving the data packets sent by the display device 200, the server 400 performs corresponding processing and analysis on the data packets and returns the generated response data packets to the display device 200. Accordingly, the response data packets may include storage addresses, semantic recognition results, image recognition results, etc.
[0298] In some embodiments, the server 400 may receive a target voice command sent by the display device 200, wherein the target voice command includes a query word (or trigger word) and a feature word. The query word may be used to trigger the screenshot operation and the image recognition program, such as keywords that can represent the type of the object corresponding to the identified object, such as "who", "what", "where", etc.; the feature word is used to represent the characteristic attributes of the identified object, such as keywords that can represent the characteristic attributes of the identified object, such as "fat", "tall", "left", "female", etc.
[0299] That is, in some embodiments, the query words include keywords for characterizing object types. For example, the object type of the identified object may be a person, an item, a place, an animal, or the like. Feature words include keywords for characterizing characteristic attributes of the identified object. For example, the characteristic attributes of the identified object may be the orientation attribute of the identified object in the image to be identified, fixed attributes associated with the identified object (such as height, fatness, age, gender, etc.), appearance attributes of the identified object (such as color, appearance of a person, appearance of an animal), and action attributes of the identified object (such as standing, lying down, etc.).
[0300] After receiving the target voice command from display device 200, server 400 analyzes the semantic recognition results of the target voice command. The semantic recognition results may include query parameters obtained based on the query terms and feature parameters obtained based on the feature terms. Specifically, the query parameters are the identifiers or characters corresponding to the query terms, and the feature parameters are the identifiers or characters corresponding to the feature terms. The semantic recognition results are then sent to display device 200, which executes the corresponding program based on the semantic recognition results.
[0301] For example, the display device 200 acquires a target voice command collected by a sound collector based on the first application. The target voice command is "Who is the person on the left?" The display device 200 sends the command to the voice recognition module 430 of the server 400 through the first application. The voice recognition module 430 performs text conversion and semantic understanding on the command, and derives the query term "who" used to characterize the type of the recognized object and the feature term "left" used to characterize the characteristics of the recognized object. Then, based on the parsed query term, the server 400 determines that the target voice command is a voice command for triggering the screenshot and image recognition program. Based on the query term, the query parameter "1" is used to identify the type of the query object as a person, and the feature parameter "left" is generated based on the feature term. The query parameter is then used to generate a command to call the second application, which includes the query parameter. Finally, the command to call the second application and the query parameter are sent as a response data packet to the first application of the display device 200, causing the first application to execute the corresponding program based on the response data packet.
[0302] It is understood that the query parameters and feature parameters described above are merely exemplary parameter types, and the query parameters and feature parameters of the present disclosure may also be other characters, such as English characters, numeric characters, Chinese character labels, or a combination of one or two of the above characters. In some embodiments, the server 400 and the display device 200 may identify the query parameters and feature word parameters by setting specific flags in the data packet.
[0303] In some embodiments, after the display device 200 receives the semantic recognition result corresponding to the target voice command sent by the server 400, it parses the voice recognition result based on the first application, and responds to the command to call the second application in the semantic recognition result, sends a call notification including query parameters and feature parameters to the second application as a content recognition instruction. Among them, the content recognition instruction is a control instruction for triggering the second application to execute the screenshot and image recognition program. The second application of the display device 200 then responds to the content recognition instruction to take a screenshot of the user interface to generate an image to be recognized. The image to be recognized is then sent to the server 400 based on the second application.
[0304] In some embodiments, the storage module 420 of the server 400 is configured to store image data. After receiving the image to be recognized from the display device 200, the control module 430 of the server 400 stores the image to be recognized at the target address of the storage module 420, i.e., the network address of the image to be recognized in the storage module 420. The target address is then sent to the display device 200, so that the display device 200 generates an image recognition request based on the target address, query parameters, and feature parameters.
[0305] For example, the display device 200 can use the target address of the image to be recognized in the storage module 420 as an input parameter, and send an image recognition request to the image recognition module 440 of the server 400 based on the query parameters and feature parameters, so as to call the server 400 to perform an image recognition program on the image to be recognized based on the query parameters and feature parameters. The data packet of the image recognition request includes the target address, query parameters, and feature parameters.
[0306] After generating an image recognition request, the display device 200 sends the image recognition request to the server 400. In some embodiments, after receiving the image recognition request, the server 400 responds to the image recognition request and obtains the image to be recognized from the storage module 420 according to the target address, so as to perform subsequent image recognition procedures on the image to be recognized.
[0307] Since the image recognition request uses the target address as an input parameter and simultaneously includes query parameters obtained from the query term and feature parameters obtained from the feature term, the query term is used to characterize the object type of the recognized object, and the feature term is used to characterize the feature attributes of the recognized object. Accordingly, the query parameter can be used to characterize the object type of the recognized object (such as a person, object, place, or animal); the feature parameter can be used to characterize the feature attributes of the recognized object, such as at least one of the aforementioned action attributes, external attributes, fixed attributes, and orientation attributes. In other words, the query parameter is used to specify the object type of the recognized object to be recognized in this image recognition, and the feature parameter is used as a filtering condition to filter the objects identified by the query parameter.
[0308] In some embodiments, the query parameters and feature parameters can be fixed or dynamically changed during the transmission process. That is, the query parameters and feature parameters in the semantic recognition results sent by server 400 to the first application, the query parameters and feature parameters sent by the first application to the second application, and the query parameters and feature parameters sent by the second application to server 400 can be the same characters or different characters, but the information represented by the query parameters and feature parameters is fixed. The query parameters are used to represent the object type of the identified object; the feature parameters are used to represent the characteristic attributes of the identified object.
[0309] S2502: In response to the image recognition request, perform image recognition on the image to be recognized based on the query parameters and the feature parameters to generate a target recognition result.
[0310] After receiving the image recognition request sent by the display device 200, the server 400 can obtain the image to be recognized from the storage module 420 according to the target address included in the image recognition request, and then perform image recognition (image recognition program) on the image to be recognized based on the query parameters and feature parameters.
[0311] In some embodiments, the server 400 performs image recognition (image recognition program) on the image to be identified according to the query parameters to generate an initial recognition result. The initial recognition result matches the object type represented by the query parameters. For example, if the query term is "who" and the object type represented by the query parameters is people, the initial recognition result is the recognition result of all people in the image to be identified.
[0312] In some embodiments, when the display device 200 server 400 sends an image recognition request, the query parameters can be encapsulated into the header of the image recognition request (data packet); or, the result type specified by the query parameters can be specified in the image recognition request. For example, when the result type is a person, a "type" field is included in the JSON or XML payload of the image recognition request, and its value is set to "1" representing the person.
[0313] FIG26 is a schematic diagram of a process of image recognition according to some embodiments. The server 400 may execute the process shown in FIG26 to perform image recognition. As shown in FIG26 , the process may include the following steps:
[0314] S2601: Analyzing a first object in an image to be recognized based on an image recognition algorithm;
[0315] S2602: Query the associated information of the first object to obtain an initial recognition result;
[0316] In some embodiments, server 400 analyzes a first object in an image to be identified based on an image recognition algorithm. The type of the first object matches the object type represented by the query parameter. Related information about the first object is then queried to obtain an initial recognition result. For example, the object type represented by the query parameter is a person. After server 400 matches the ID of the first object using the image recognition algorithm, it can query related information such as the first object's corresponding gender, height, occupation, weight, age, the first object's position in the image to be identified, and the first object's width and height in the image to be identified using the first object's ID.
[0317] As shown in the unit architecture of Figure 8, in some embodiments, the display device 200 can extract key features from the image to be identified based on the query parameters and retrieve standard patterns or standard templates pre-stored in the database according to the query parameters to match the key features. For example, when the query parameters represent the object type of a person, the extracted key features include facial contours, the shape and position of the eyes, nose, and mouth, and skin texture. The standard patterns and standard templates are pre-stored person features in the database.
[0318] It is understandable that when the object type represented by the query parameter is other types, the standard mode and standard module obtained by the server 400 are the modes or templates corresponding to other types, and the principle is the same as the result type of the above-mentioned person.
[0319] In some embodiments, after the server 400 performs image processing on the image to be recognized and generates an initial recognition result, it filters the initial recognition result based on the feature parameters to extract recognition results that match the feature attributes represented by the feature parameters. In other words, the target recognition result is the recognition result in the initial recognition result that matches the feature attributes represented by the feature parameters.
[0320] S2603: extracting a second object from the first object based on the association information between the feature parameter and the first object;
[0321] S2604: Acquire association information of the second object to generate a target recognition result.
[0322] In some embodiments, server 400 extracts a second object from the first object based on association information between the feature parameters and the first object; wherein the second object is an identification object including target association information, and the attribute characteristics represented by the target association information match the attribute characteristics represented by the feature parameters. The server then obtains the association information of the second object to generate a target recognition result.
[0323] For example, the query term is "who," the feature term is "left," the query parameter represents the object type "person," and the feature parameter represents the orientation attribute "left." Server 400 parses all the human objects (first objects) in the image to be recognized and obtains associated information for each human object (initial recognition result). Based on the position coordinates of each human object in the associated information, server 400 selects the leftmost human object (second object) and its associated information (target recognition result).
[0324] In some embodiments, when a feature word represents an orientation feature, the feature parameter includes an orientation parameter, which is used to represent the relative position of the identified object in the image to be identified (e.g., left, right, top, bottom, top left, bottom left, top right, bottom right, center, etc.). After parsing the first object in the image to be identified, the server 400 obtains a target coordinate system for image recognition. The origin of the target coordinate system can be located at the geometric center point 11A of the user interface of the display device 200, or at any boundary vertex 11B in the user interface of the display device 200; as shown in FIG. 23 , corresponding to the image to be identified, the origin of the target coordinate system is located at the geometric center point 11a of the image to be identified, or at a boundary vertex 11b of the image to be identified. Then, based on the target coordinate system, the position coordinates of the first object in the image to be identified are detected. The position coordinates include coordinate values in a first direction and coordinate values in a second direction, with the first direction being perpendicular to the second direction, such as the x-direction and the y-direction. After detecting the position coordinates of the first object in the image to be identified, the server 400 records the position coordinates in the associated information of the first object, such as (x, y).
[0325] That is, when the server 400 performs image recognition on the image to be recognized according to the query parameters, it also detects the position of the first object in the image to be recognized and records the position coordinates as associated information in the server 400 for the server 400 to perform subsequent comparison and processing.
[0326] In order to extract the target recognition results more accurately, in some embodiments, the server 400 can also record the position coordinates of the first object according to the orientation parameters. When the orientation parameters are different, the position coordinates recorded by the server 400 are also different, and the position coordinates can be consistent with the relative position represented by the orientation parameters. For example, when the relative position represented by the orientation parameter is the upper left, the server 400 records the coordinates of the upper left pixel of the first object; when the relative position represented by the orientation parameter is the lower left, the server 400 records the coordinates of the lower left pixel of the first object; when the relative position represented by the orientation parameter is the left, the server 400 records the coordinates of the leftmost pixel of the first object. It can be understood that when the relative position represented by the orientation parameter is other relative positions, the principle of the server 400 recording the position coordinates is the same as the above principle, which will not be repeated here.
[0327] In some embodiments, when the feature parameters include orientation parameters and the origin of the target coordinate system is located at the center of the object to be identified, server 400 may further calculate a cropping region of the image to be identified based on the relative position represented by the feature parameters, and crop the image to be identified based on the cropping region to retain the image content of the cropped region in the image to be identified. Image recognition is then performed on the cropped image to be identified. That is, steps S2502-S2503 are performed on the image content to be identified that remains after cropping, thereby reducing the amount of data computation required by server 400.
[0328] For example, when the feature word is "left", the feature parameters include orientation parameters, and the origin of the target coordinate system is located at the geometric center of the image to be identified, the server 400 calculates that the cropping area of the image to be identified is the image portion in the second and third quadrants based on the relative position "left" represented by the feature parameters, and the server 400 crops out the image portion of the image to be identified that is located in the first and fourth quadrants, and retains the image portion in the second and third quadrants. For another example, when the feature word is "upper left", the feature parameters include orientation parameters, and the origin of the target coordinate system is located at the geometric center of the image to be identified, the server 400 calculates that the cropping area of the image to be identified is the image portion in the second quadrant based on the relative position "upper left" represented by the feature parameters, and the server 400 crops out the image portion of the image to be processed that is located in the first, third, and fourth quadrants, and retains the image portion in the second quadrant.
[0329] In some embodiments, when the server 400 calculates the cropping area of the image to be identified, it obtains the coordinate critical value of each quadrant and determines the image portion of the image to be identified in each quadrant based on the coordinate critical value of each quadrant, so that the server 400 can crop the image to be identified. The coordinate critical value is used to represent the coordinate value characteristics of the position coordinates of the pixel points corresponding to each quadrant. For example, when the origin of the target coordinate system is located at the geometric center point of the image to be identified, the coordinate critical value of the first quadrant is x>0 and y>0, the coordinate critical value of the second quadrant is x<0 and y>0, the coordinate critical value of the third quadrant is x<0 and y<0, and the coordinate critical value of the fourth quadrant is x>0 and y<0.
[0330] To facilitate extracting the target recognition result, the server 400 may execute the process shown in FIG. 27 to determine the second object. As shown in FIG. 27 , the process may include the following steps:
[0331] S2701: Analyze the critical trend corresponding to the azimuth parameter;
[0332] In some embodiments, the relative position represented by the orientation parameter includes at least one of left, right, up, down, and center, the first direction represents the horizontal direction, and the second direction represents the vertical direction. Server 400 also analyzes the critical trend corresponding to the orientation parameter; the critical trend includes the maximum or minimum end of the first direction and / or the maximum and minimum ends of the second direction.
[0333] For example, when the first direction is the x direction and the second direction is the y direction, and the origin of the target coordinate system is located in the lower left corner of the user interface, when the physical range is left, the critical trend is the minimum value in the x direction; when the physical range is top, the critical trend is the maximum value in the y direction; when the physical range is top left, the critical trend is the minimum value in the x direction and the maximum value in the y direction.
[0334] S2702: Read the coordinate value of the first object from the association information of the first object;
[0335] S2703: extracting the identified object whose coordinate value is the target value from the first object as the second object;
[0336] Among them, the target coordinate value is the coordinate value of the first object in the critical trend, that is, the second object is the first object whose coordinate value matches the critical trend. For example, the origin of the target coordinate system is located in the lower left corner of the image to be identified, the object type represented by the query parameter is a person, and the relative position represented by the feature parameter is on the left. As shown in Figure 24 above, the server 400 performs image recognition on the image to be identified, identifies all the people it includes, and records the position coordinates (x1, y1) of person 1, the position coordinates (x2, y2) of person 2, and the position coordinates (x3, y3) of person 3. According to the relative position represented by the orientation parameter and the target coordinate system, the critical trend is analyzed to be the minimum value in the x direction. Then, the server 400 traverses the coordinate values of the x coordinates in the associated information corresponding to each first object, and extracts the first object with the largest coordinate value in the x coordinate. x3>x1>x2, so the person 3 corresponding to (x3, y3) is filtered out as the second object, and then the associated information corresponding to person 3 is extracted as the target recognition result.
[0337] In some embodiments, the server 400 may also obtain a boundary set of the image to be identified in the target coordinate system. The boundary set includes the left boundary, right boundary, upper boundary, and lower boundary of the image to be identified. When extracting the second object, the server 400 selects a reference boundary based on the physical range and calculates the relative distance between the first object and the reference boundary based on the position coordinates of the first object. The server 400 then extracts the first object whose relative distance matches the relative position (relative position represented by the orientation parameter) to filter out the second object.
[0338] For example, if the relative position represented by the orientation parameter is left, the reference boundary may be the left boundary. Server 400 calculates the relative distance between each first object and the reference boundary and extracts the first object with the smallest relative distance as the second object. For another example, if the physical range is left, the reference boundary may be the right boundary. Server 400 calculates the relative distance between each first object and the reference boundary and extracts the first object with the largest relative distance as the second object.
[0339] In some embodiments, when the relative position represented by the orientation parameter includes two or more items, the server 400 further reads a first coordinate value of the first object in the first direction and a second coordinate value of the first object in the second direction. A first judgment value is generated based on the critical trend and the first coordinate value, and a second judgment value is generated based on the critical trend and the second coordinate value. The sum of the first judgment value and the second judgment value is calculated, and the identified object with the largest sum is extracted from the first object as the second object.
[0340] In some embodiments, the server 400 can perform normalization from left to right and from bottom to top on the coordinate values of the first direction and the second direction based on the width and height of the image to be identified, and generate a first judgment value and a second judgment value based on the normalized coordinate values.
[0341] In some implementations, when the relative position represented by the orientation parameter is the upper left, the server 400 reads the normalized x- and y-coordinate values of the first object. The critical trend is the minimum end of the first direction and the minimum end of the second direction. The server 400 calculates the inverse of the first and second coordinate values (i.e., 1-the normalized value) as the first and second judgment values. The sum of the first and second judgment values is then calculated, and the identified object with the largest sum is extracted from the first object as the second object.
[0342] When the relative position represented by the orientation parameter is lower left, the server 400 reads the normalized coordinate value of the first object in the x-direction and the normalized coordinate value of the first object in the y-direction. The critical trend is the minimum end of the first direction and the maximum end of the second direction. The server 400 calculates the opposite value of the first coordinate value (i.e., 1-the normalized value of the coordinate value) as the first judgment value and the normalized coordinate value of the second coordinate value as the second judgment value. The sum of the first judgment value and the second judgment value is then calculated, and the identified object with the largest sum value is extracted from the first object as the second object.
[0343] It is understandable that the calculation principle of the sum values of other relative positions is the same as that of the above embodiment, and will not be elaborated here.
[0344] In some embodiments, when the feature word includes inherent attributes, the feature parameter includes a basic parameter, which is used to characterize the inherent attributes of the identified object, such as color attributes, the name of the identified object, etc. After parsing the first object and obtaining the associated information of the first object, the server 400 searches the associated information for associated information characterizing the inherent attributes, thereby extracting the second object from the first object based on the associated information of the inherent attributes.
[0345] That is, in some embodiments, the feature parameters include basic parameters, and the basic parameters are used to characterize the inherent attributes corresponding to the identified object. The server 400 also parses the target inherent attributes corresponding to the basic parameters, and then reads the reference value of the target inherent attributes. For example, taking the object type represented by the query parameter as a person, the inherent attributes may include the person's gender, the person's name, the person's skin color, etc., and the target inherent attributes may be a combination of one or more of the above; when the target inherent attribute is gender, the reference value may be "male / female". After reading the reference value of the target inherent attribute, the identified object whose inherent attribute matches the reference value in the first object is extracted as the second object.
[0346] For example, if the object type represented by the query parameter is an item and the inherent attribute represented by the feature parameter is red, server 400 can obtain the color value corresponding to red (such as an RGB value) and use the color value corresponding to red as a reference value. Next, server 400 searches the associated information for the color value corresponding to each item and calculates the similarity between the item color value and the reference value. Then, server 400 extracts items whose color values match the reference value (which can be the most similar or completely identical) as the second filtered object.
[0347] In some embodiments, the feature attributes include extrinsic parameters, which represent the appearance attributes and / or action attributes (standing, lying down, etc.) of the identification object. For example, when the object type represented by the query parameter is a person, the appearance attributes can be weight, long hair, short hair, wearing glasses, etc.; when the object type represented by the query parameter is an object, the appearance attributes can be shape, etc. The server 400 obtains the target template corresponding to the extrinsic parameter and calculates the similarity between the target template and the first object. The identification object with the highest similarity in the first object is then extracted as the second object. For example, when the appearance attribute represented by the extrinsic parameter is a triangle, the target template obtained by the server 400 includes templates of various triangles, and then the similarity between each first object and the target template is calculated separately.
[0348] In some embodiments, server 400 may also detect the width of the first object in the image to be identified and record the detected width in the association information. This allows server 400 to extract the second object based on the width of the first object and the second object that matches the characteristic attribute represented by the feature parameter. For example, if the query parameter represents the object type as a person and the characteristic parameter represents the appearance attribute as thin, server 400 may query the association information for the width values corresponding to each person and extract the person with the smallest width value as the second object.
[0349] In some embodiments, when the object type represented by the query parameter is a person, when the server 400 records the width of the first object in the image to be identified, it can detect the face area of the person based on the image recognition algorithm, and then record the width of the face area as the width corresponding to the first object.
[0350] It is understood that the characteristic parameters may include a combination of one or more of the above parameters. In some embodiments, when the characteristic parameters include multiple parameters, the server 400 may also preset weights for the different parameters. When extracting the second object, the server 400 may calculate the comprehensive score of the first object based on the weights corresponding to the characteristic parameters, and then extract the first object with the highest comprehensive score as the second object. The implementation logic of the characteristic parameters is the same as the above principle and will not be elaborated here.
[0351] S2503: Send the target recognition result to the display device.
[0352] After server 400 obtains the target recognition result, it sends the target recognition result to display device 200, such as to the second application of display device 200. In this way, upon receiving the target recognition result sent by server 400, display device 200 can immediately display the target recognition result. Because server 400 has already screened the image recognition results of the image to be recognized based on the feature parameters, the target recognition result displayed by display device 200 can be more closely aligned with the user's recognition intent, thereby improving the recognition efficiency of image recognition.
[0353] Based on the above image recognition process performed on the image to be recognized, the display device 200 of the embodiment of the present disclosure may further perform the process shown in FIG28 to perform image recognition. As shown in FIG28 , the process may include the following steps:
[0354] S2801: In response to a content recognition instruction, take a screenshot of the user interface to generate an image to be recognized.
[0355] Among them, the content recognition instruction includes query parameters and feature parameters, the image to be recognized includes an object to be recognized, the object type of the recognized object in the target recognition result matches the object type represented by the query parameters, the query parameters are used to represent the object type of the recognized object; the feature parameters are used to represent the feature attributes of the recognized object.
[0356] In some embodiments, the display device 200 may further include a detector 220 configured to collect voice commands input by the user. FIG29 is a flow diagram of a display device performing semantic recognition according to some embodiments. The display device 200 may perform semantic recognition by executing the flow shown in FIG29 , which may include the following steps:
[0357] S2901: Acquire the target voice command collected by the detector;
[0358] The target voice command includes voice data, and the voice data includes query words and feature words.
[0359] S2902: Sending the target voice command to the server;
[0360] The display device 200 sends the target voice instruction to the server 400 so that the server 400 parses the semantic recognition result of the target voice instruction, where the semantic recognition result includes query parameters obtained based on the query words and feature parameters obtained based on the feature words.
[0361] S2903: Receive the semantic recognition result returned by the server;
[0362] S2904: Generate a content recognition instruction based on the semantic recognition result;
[0363] The content recognition instruction is a control instruction for triggering the second application to execute a screenshot and image recognition program.
[0364] In some embodiments, the display device 200 includes a first application, that is, the application layer of the display device 200 is configured with the first application. The display device 200 receives the target voice instruction based on the first application and then sends the voice instruction to the server 400. The semantic recognition result returned by the server 400 is received, and a content recognition instruction is generated based on the voice recognition result. The content recognition instruction is then sent to the second application so that the second application responds to the content recognition instruction. That is, the display device 200 performs steps S2901-S2904 based on the first application and then sends the content recognition instruction generated in step S2904 to the second application.
[0365] After generating the content recognition instruction, the display device 200 also responds to the content recognition instruction and takes a screenshot of the user interface to generate the image to be recognized. To reduce the computational complexity of image recognition, in some embodiments, when the feature parameters include orientation parameters and the origin of the target coordinate system used by the display device 200 is located at the geometric center of the user interface, the display device 200 can also calculate a screenshot range based on the relative position represented by the orientation parameters and the coordinate thresholds of each quadrant of the target coordinate system, and take a screenshot of the user interface based on the screenshot range.
[0366] For example, if the feature word is "left" and the relative position represented by the orientation parameter is "left", the display device 200 will calculate the screenshot range as the user interface in the second and third quadrants. For another example, if the feature word is "upper left" and the relative position represented by the orientation parameter is "left" and "upper", the display device 200 will calculate the screenshot range as the user interface in the second quadrant.
[0367] S2802: Send the image to be recognized to the server.
[0368] After the display device 200 captures a screenshot to generate an image to be recognized, the image to be recognized is sent to the server 400 for storage, so that the server 400 can perform subsequent image recognition on the image to be recognized.
[0369] That is, in some embodiments, after the display device 200 sends the image to be recognized to the server 400, it receives a target address returned by the server 400. The target address is the storage address of the image to be recognized on the server. An image recognition request is then generated based on the target address, query parameters, and feature parameters, allowing the server 400 to perform image recognition on the image to be recognized based on the above data.
[0370] S2803: Send an image recognition request to the server.
[0371] After the display device 200 generates an image recognition request, it sends the image recognition request to the server 400, so that the server 400 can respond to the image recognition request and call the corresponding module to perform image recognition on the image to be recognized, that is, steps S2501-S2504 described in the above embodiment, which will not be repeated here.
[0372] S2804: Receive the target recognition result returned by the server based on the image recognition request.
[0373] The display device 200 sends an image recognition request for an image to be recognized to the server 400, and the server 400 returns a corresponding target recognition result to the display device 200 in response to the image recognition request. The object type of the recognized object in the target recognition result matches the object type represented by the query parameter, and the feature attributes of the recognized object in the target recognition result match the feature attributes represented by the feature parameters.
[0374] S2805: Control the display to display the target recognition result.
[0375] After receiving the image recognition result returned by the server 400, the display device 200 controls the display 260 to display the target recognition result, as shown in Figure 30. Similarly, because the server 400 has already screened the image recognition results of the image to be recognized based on the feature parameters, the target recognition result displayed by the display device 200 can be closer to the user's recognition intention, thereby improving the recognition efficiency of the image recognition.
[0376] In some embodiments, the display device 200 includes a second application, that is, the application layer of the display device 200 is configured with the second application. The display device 200 receives a content recognition instruction sent by the first application based on the second application, and in response to the content recognition instruction, takes a screenshot of the user interface and sends the image to be recognized to the server 400. The display device 200 then sends an image recognition request to the server 400. The display device 200 then receives the target recognition result returned by the server 400 and controls the display 260 to display the target recognition result. That is, the display device 200 receives the content recognition instruction sent by the first application based on the second application and executes steps SS2801-S2805 described in the above embodiment, which will not be described in detail here.
[0377] It should be noted that the above embodiment is described with the first application (voice application) and the second application (image recognition application) as two independent applications. When the first application and the second application are applications integrated into the same application program, the logic executed is the same as the above logic.
[0378] In some embodiments, when the target recognition result is displayed, the display device 200 may also generate a prompt corresponding to the target recognition result based on the target recognition result, and display the prompt in the display view of the target recognition result. For example, as shown in FIG30 , when the result type specified by the query parameter is a person type, the display device 200 may generate a prompt related to the target recognition result, such as "What are his related works?" or "Recommend several of his movies."
[0379] In some embodiments, when displaying the target recognition result, the display device 200 can also receive an object acquisition instruction input by the user, and in response to the object acquisition instruction, send an acquisition request for the first object to the server 400, so that the server 400 sends the associated information of the first object to the display device 200 based on the acquisition request for the first object. After receiving the associated information of the first object, the display device 200 controls the display 260 to display the associated information of the first object for the user to switch and select. In this way, when the user interface includes recognition objects corresponding to multiple query parameters, if the user wants to obtain information about other recognition objects of the same type or the feature parameters input by the user do not match the currently displayed recognition result (such as the user description is not accurate enough), the user does not need to enter the voice command again, which can improve the convenience of image recognition.
[0380] In some embodiments, the object acquisition instruction may be based on an acquisition control input. The acquisition control may be a subview in the target recognition result display interface that can acquire the focus of the display device 200. When the display device 200 detects that the focus is on the acquisition control and detects a click event on the acquisition control, the object acquisition instruction is generated in response to the click event of the acquisition control.
[0381] For example, as shown in the user interface of Figure 31, the user interface includes recognition objects of multiple character types, namely object 1, object 2, object 3, object 4 and object 5. After the user enters the voice command of "who do you know on the right", the display device 200 interacts with the server 400, and the server 400 returns the associated information of object 5 (i.e., the target recognition result) that meets the above voice command to the display device 200. Correspondingly, the display device 200 displays the associated information of object 5, as shown in Figure 32, and the acquisition control 1801 is displayed in the upper right position of the target recognition result display area. At this time, if the user wants to obtain the associated information of object 4, the focus can be moved to the acquisition control 1801 by the control device 100 (such as a remote control, etc.), and the acquisition control 1801 is selected by the control device 100 to generate a click event for the acquisition control 1801. After detecting the click event of the acquisition control 1801, the display device 200 responds to the click event and sends a request to obtain the character type identification object to the server 400, and then displays the associated information returned by the server 400 (the associated information of other character type identification objects, i.e., object 1-object 4), displaying the interface effect as shown in Figure 33.
[0382] It should be noted that all or part of the programs executed by the server 400 in the above embodiment can also be executed by the display device 200, which will not be described in detail here.
[0383] Based on the above server 400, some embodiments of the present disclosure further provide a feature recognition method, which can be applied to the server 400 provided in the above embodiment. As shown in FIG25 , the method includes the following procedural steps:
[0384] S2501: Receive an image to be recognized and an image recognition request sent by a display device, where the image to be recognized is an image generated by the display device performing a screenshot of a user interface, and the image recognition request includes query parameters and feature parameters; the query parameters are used to specify a result type when performing image recognition on the image to be recognized; the feature parameters are used to specify a recognition range for performing image recognition on the image to be recognized, and the recognition range includes a physical range and / or a feature type range;
[0385] S2502: In response to the image recognition request, perform image recognition on the image to be recognized based on the query parameters and the feature parameters to generate a target recognition result, wherein the object type of the recognized object in the target recognition result matches the object type represented by the query parameters, and the feature attributes of the recognized object in the target recognition result match the feature attributes represented by the feature parameters;
[0386] S2503: Send the target recognition result to the display device, so that the display device displays the target recognition result.
[0387] Based on the above-mentioned display device 200, some embodiments of the present disclosure further provide a method for displaying target recognition results, which can be applied to the display device 200 provided in the above-mentioned embodiment. As shown in FIG28 , the method includes the following procedural steps:
[0388] S2801: In response to a content recognition instruction, taking a screenshot of the user interface to generate an image to be recognized; the content recognition instruction includes a query parameter and a feature parameter, the query parameter being used to specify a result type when performing image recognition on the image to be recognized; the feature parameter being used to specify a recognition range for performing image recognition on the image to be recognized, the recognition range including a physical range and / or a feature type range;
[0389] S2802: Sending the image to be recognized to a server;
[0390] S2803: Sending an image recognition request including the query parameter and the feature parameter to the server;
[0391] S2804: Receive a target recognition result returned by the server based on the image recognition request; the target recognition result is feature information included in the initial recognition result within the recognition range corresponding to the feature parameter, and the initial recognition result includes feature information corresponding to the result type specified by the query parameter;
[0392] S2805: Control the display to display the target recognition result.
[0393] In the above process, the server 400 of the embodiment of the present disclosure can receive the image to be identified and the image recognition request sent by the display device 200. The image to be identified is an image generated by the display device 200 based on the screenshot of the user interface, and the image recognition request includes query parameters and feature parameters. The query parameters are used to specify the result type when performing image recognition on the image to be identified, and the feature parameters are used to specify the recognition range of image recognition performed on the image to be identified. Then, in response to the image recognition request, image recognition is performed on the image to be identified according to the query parameters, and the target recognition result is extracted from the recognition result generated by the image recognition based on the feature parameters. The target recognition result is then sent to the display device 200 so that the display device displays the target recognition result. By combining the query parameters and the feature parameters, the method can perform more accurate image recognition on the image to be identified, thereby feeding back image recognition results that are more in line with the user's intention.
[0394] Those skilled in the art will clearly understand that the technology in the embodiments of the present disclosure can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the implementation methods in the embodiments of the present disclosure, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment of the present disclosure or certain parts of the embodiments.
[0395] Finally, it should be noted that the above embodiments are only used to illustrate the implementation methods of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the implementation methods described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding implementation methods to deviate from the scope of the implementation methods of the embodiments of the present disclosure.
[0396] For ease of explanation, the above description has been made with reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations are possible. The above embodiments are selected and described to better explain the principles and practical applications, so that those skilled in the art can better utilize the embodiments and various different variations of the embodiments suitable for specific use considerations.
Claims
1. A display device, comprising: a display configured to display a user interface; a memory configured to store computer instructions and data associated with the display device; At least one processor, connected to the display and the memory, configured to execute computer instructions to cause the display device to perform: In response to an image recognition instruction, an image to be recognized is acquired, wherein the image to be recognized includes an object to be recognized, and the image recognition instruction includes a query parameter and a feature parameter, wherein the query parameter is used to characterize an object type of the object to be recognized; The characteristic parameters are used to characterize the characteristic attributes of the identified object; performing image recognition on the image to be recognized according to the query parameters extracted from the image recognition instruction to obtain a recognition result, the recognition result including an object to be recognized that meets the query parameters and at least one of an arrangement parameter, a feature label, and feature details of the object to be recognized, the arrangement parameter being used to characterize a positional relationship of the object to be recognized in the image to be recognized; The feature tag is used to represent the identification information of the identification object, and the feature details are related information queried based on the feature tag; generating a result display view according to the recognition result, wherein the result display view includes a feature label and feature details, wherein a display position of the feature label is set according to the arrangement parameter, and a display state of the feature details is highlighted according to the feature parameter; The display is controlled to display the result presentation view.
2. The display device according to claim 1, wherein the at least one processor is specifically configured to execute computer instructions to enable the display device to perform the step of acquiring the image to be recognized: If the image information is not detected in the image recognition instruction, calling a screenshot process, and performing a screenshot of the currently displayed user interface through the screenshot process to generate the image to be recognized; If the image information is detected in the image recognition instruction, the image to be recognized is extracted according to the image information.
3. The display device according to claim 1, further comprising: a communication device configured to establish a communication connection with a server; The at least one processor is specifically configured to execute computer instructions to enable the display device to perform image recognition on the image to be recognized according to the query parameter extracted from the image recognition instruction to obtain a recognition result: generating a first image recognition request according to the query parameter and the feature parameter; Sending the first image recognition request and the image to be recognized to a server, so that the server performs image recognition on the image to be recognized according to the query parameters, and screens the recognition objects obtained by the image recognition according to the feature parameters to determine the target recognition objects that meet the feature parameters; Receive a first recognition result fed back by the server; wherein the first recognition result includes an identification object that meets the query parameters, a target identification object that meets the feature parameters, and arrangement parameters of the identification object.
4. The display device according to claim 3, wherein the at least one processor is specifically configured to execute computer instructions to enable the display device to generate a result display view according to the recognition result: parsing the identification object, the target identification object, and arrangement parameters of the identification object from the first identification result; According to the arrangement parameters, the position of the feature label corresponding to the identified object in the result display view is set; The position of the focus mark in the result display view is displayed according to the target recognition object.
5. The display device according to claim 1 , wherein the at least one processor is specifically configured to execute computer instructions to cause the display device to perform image recognition on the image to be recognized according to the query parameter extracted from the image recognition instruction to obtain a recognition result: generating a second image recognition request according to the query parameters; sending the second image recognition request and the image to be recognized to a server, so that the server performs image recognition on the image to be recognized according to the query parameters; Receive the second recognition result fed back by the server; wherein, The second recognition result includes the recognition objects that meet the query parameters and the arrangement parameters of the recognition objects.
6. The display device according to claim 5, wherein the at least one processor is specifically configured to execute computer instructions to enable the display device to generate a result display view according to the recognition result: parsing the identified object and arrangement parameters of the identified object from the second recognition result; Screening the identification objects according to the characteristic parameters to determine target identification objects that meet the characteristic parameters; According to the arrangement parameters, the position of the feature label corresponding to the target recognition object in the result display view is set; The position of the focus mark in the result display view is displayed according to the target recognition object.
7. The display device according to claim 1, wherein the at least one processor is specifically configured to execute computer instructions to enable the display device to generate a result display view according to the recognition result: Call result display template; Adding a label control to the result display template, wherein the label control is used to display the feature label of the identified object; Setting the display position and / or display order of the label control according to the arrangement parameters; Determining a focus label control according to the characteristic parameters; wherein the focus label control is a label control that meets the characteristic parameters; The focus identifier is set on the focus label control, and feature details corresponding to the focus label control are added to the detail display area associated with the focus label control.
8. The display device according to claim 7, wherein the at least one processor is further configured to execute computer instructions to cause the display device to: Obtaining a recognition interaction instruction input by a user based on the result display view; Extract the change parameter from the identification interaction instruction; wherein, The change parameter is a characteristic parameter in the recognition interaction instruction that is different from the image recognition instruction; Determining a new focus label control according to the change parameters; wherein the new focus label control is a label control that meets the change parameters; The focus identifier is moved to the new focus label control, and feature details corresponding to the new focus label control are added to the detail display area associated with the new focus label control.
9. The display device according to claim 3, wherein the at least one processor is further configured to execute computer instructions to cause the display device to: In response to the content recognition instruction, a screenshot of the user interface is executed to generate an image to be recognized; the image to be recognized includes an object to be recognized, and the content recognition instruction includes a query parameter and a feature parameter, the query parameter is used to represent the object type of the object to be recognized; the feature parameter is used to represent the feature attribute of the object to be recognized; Sending the image to be identified to a server; Sending an image recognition request including the query parameter and the feature parameter to the server; receiving a target recognition result returned by the server based on the image recognition request; wherein the object type of the recognized object in the target recognition result matches the object type represented by the query parameter, and the feature attributes of the recognized object in the target recognition result match the feature attributes represented by the feature parameters; Control the display to display the target recognition result.
10. The display device according to claim 9, wherein the server generates the target recognition result in the following manner: receiving an image to be recognized and an image recognition request sent by the display device, wherein the image to be recognized is an image generated by the display device by taking a screenshot of a user interface, the image to be recognized includes an object to be recognized, and the image recognition request includes a query parameter and a feature parameter; The query parameter is used to characterize the object type of the identified object; The characteristic parameters are used to characterize the characteristic attributes of the identified object; In response to the image recognition request, image recognition is performed on the image to be recognized based on the query parameters and the feature parameters to generate a target recognition result; wherein the object type of the recognized object in the target recognition result matches the object type represented by the query parameters, and the feature attributes of the recognized object in the target recognition result match the feature attributes represented by the feature parameters.
11. The display device according to claim 10, wherein the server generates the target recognition result by: parsing a first object in the image to be identified based on an image recognition algorithm, where the first object matches an object type represented by the query parameter; Querying the associated information of the first object; Extracting a second object from the first object based on association information between the feature parameter and the first object, where the second object is an identification object including target association information, and attribute features represented by the target association information match the attribute features represented by the feature parameter; Acquire the associated information of the second object to generate the target recognition result.
12. The display device according to claim 11, wherein the characteristic parameters include: An orientation parameter for characterizing the relative position of the identification object in the image to be identified; After parsing the first object in the image to be identified, the server obtains a target coordinate system for image recognition, detects the position coordinates of the first object in the image to be identified based on the target coordinate system, and records the position coordinates in the associated information of the first object; wherein the position coordinates include coordinate values in a first direction and coordinate values in a second direction, and the first direction is perpendicular to the second direction.
13. The display device according to claim 12, wherein the relative position represented by the orientation parameter comprises at least one of left, right, up, down, and center, the first direction represents a horizontal direction, and the second direction represents a vertical direction; The server determines the second object by: Analyze the critical trend corresponding to the orientation parameter; wherein, The critical trend includes: a maximum value end or a minimum value end of the coordinate value in a first direction, and / or a maximum value end or a minimum value end of the coordinate value in a second direction, the first direction being perpendicular to the second direction; Reading the coordinate value of the first object from the associated information of the first object; Extract the identified object whose coordinate value is the target coordinate value from the first object as the second object; wherein the target coordinate value is the coordinate value of the first object at the critical trend.
14. The display device according to claim 13, wherein the server determines the second object by: If the relative position represented by the orientation parameter includes two or more items, reading a first coordinate value of the first object in the first direction, and reading a second coordinate value of the first object in the second direction; generating a first judgment value according to the critical trend and the first coordinate value, and generating a second judgment value according to the critical trend and the second coordinate value; The sum of the first judgment value and the second judgment value is calculated, and the identification object with the largest sum is extracted from the first objects as the second object.
15. The display device according to claim 12, wherein the origin of the target coordinate system is located at the center of the image to be recognized; and the server generates the target recognition result by: Calculating a cropping area of the image to be identified according to the relative position represented by the feature parameters; cropping the image to be identified according to the cropping area to retain image content of the image to be identified in the cropping area; Image recognition is performed on the cropped image to be recognized to generate the target recognition result.
16. The display device according to claim 11, wherein the characteristic parameters include: Extrinsic parameters for characterizing appearance attributes and / or action attributes of the identified object; The server determines the second object by: Obtaining a target template corresponding to the extrinsic parameters; Calculating the similarity between the target template and the first object; The identified object with the highest similarity among the first objects is extracted as the second object.
17. The display device according to claim 11, wherein the characteristic parameters include: Basic parameters for characterizing the inherent attributes corresponding to the identified object; The server determines the second object by: parsing the target inherent attributes corresponding to the basic parameters; Reading a reference value of the inherent attribute of the target; An identified object whose inherent attribute matches the reference value in the first object is extracted as the second object.
18. A feature recognition method, applied to the display device according to any one of claims 1 to 17, the method comprising: In response to an image recognition instruction, an image to be recognized is acquired, wherein the image to be recognized includes an object to be recognized, and the image recognition instruction includes a query parameter and a feature parameter, wherein the query parameter is used to characterize an object type of the object to be recognized; The characteristic parameters are used to characterize the characteristic attributes of the identified object; performing image recognition on the image to be recognized according to the query parameters extracted from the image recognition instruction to obtain a recognition result, the recognition result including an object to be recognized that meets the query parameters and at least one of an arrangement parameter, a feature label, and feature details of the object to be recognized, the arrangement parameter being used to characterize a positional relationship of the object to be recognized in the image to be recognized; The feature tag is used to represent the identification information of the identification object, and the feature details are related information queried based on the feature tag; generating a result display view according to the recognition result, wherein the result display view includes a feature label and feature details, wherein a display position of the feature label is set according to the arrangement parameter, and a display state of the feature details is highlighted according to the feature parameter; The display is controlled to display the result presentation view.
Citation Information
Patent Citations
Method, device and system for displaying character information in video based on artificial intelligence
CN107105340A
Display device and image content recognition method
CN112580625A
Display device, server and character introduction display method
CN113727162A
Display device and character recognition display method
CN114945102A
Display device, server and feature recognition method
CN118540548A