Information query method, display device, storage medium and program product

By obtaining multimodal information in the information retrieval system, conducting comprehensive understanding and analysis and modal sample database search, the problem of excessive search results in the existing system is solved, and more accurate search results and improved user experience is achieved.

CN120011579APending Publication Date: 2025-05-16HISENSE ELECTRONIC TECH (WUHAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411885491.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When existing information retrieval systems handle user queries, the search results are too broad and it is difficult to provide accurate results.

Method used

By obtaining multimodal information, conducting comprehensive understanding and analysis, generating demand text that meets user needs, and retrieving corresponding modal samples in the modal sample database. Demand the requirement text and extract information based on modal samples to generate accurate query results.

Benefits of technology

The comprehensive understanding and analysis of multimodal input information is realized, which avoids the broadness and lack of targeting of the generation of text description information in traditional solutions, and improves the accuracy and user experience of the search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011579A_ABST
    Figure CN120011579A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the display technology, and provides an information query method, a display device, a storage medium and a program product, the information query method comprises the following steps: based on at least one type of modal information used for query, obtaining a demand text used for representing a user query demand; querying at least one modal sample matched with the at least one modal information in a modal sample database; for each modal sample, performing demand decomposition on the demand text based on the modal sample to obtain sub-demand texts for the modal sample; based on the sub-demand text for the modal sample, performing information extraction on the modal sample to obtain an information extraction result corresponding to the modal sample; and according to the information extraction result corresponding to the at least one modal sample, generating a query result corresponding to the at least one modal information. According to the method, the retrieval result can be simplified, and the accuracy of the retrieval result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to natural language processing technology, and more specifically, to an information query method, a display device, a storage medium, and a program product. Background Art

[0002] As the market demand for diversified television functions grows, an information retrieval function has been developed. The information retrieval function uses a Retrieval-Augmented Generation (RAG) system to retrieve information.

[0003] In traditional solutions, RAG systems perform information retrieval based on the query statements entered by users, and then summarize and analyze all the search results for integrated output. However, the text descriptions corresponding to the search results may be too broad, and users cannot obtain more accurate search results.

[0004] Therefore, how to simplify the search results and improve the accuracy of the search results is still an urgent problem to be solved. Summary of the invention

[0005] The embodiment of the present application provides a display device that can simplify search results during information query and improve the accuracy of search results.

[0006] In a first aspect, an embodiment of the present application provides an information query method, which is applied to a display device as described in the second aspect below; the method includes:

[0007] Based on at least one modality information used for querying, obtaining a requirement text for representing the user's query requirement;

[0008] In a modality sample database, searching for at least one modality sample matching the at least one modality information;

[0009] For each modality sample, based on the modality sample, the requirement text is decomposed to obtain a sub-requirement text for the modality sample;

[0010] Based on the sub-requirement text for the modal sample, extract information from the modal sample to obtain an information extraction result corresponding to the modal sample;

[0011] According to the information extraction result corresponding to the at least one modality sample, a query result corresponding to the at least one modality information is generated.

[0012] The method provided in this embodiment obtains at least one modal information for query, then generates a demand text that meets the user's needs after a comprehensive understanding and analysis of the input at least one modal information, and retrieves the corresponding retrieved modal sample in the database according to the at least one modal information. Based on the modal sample, the demand text is decomposed, and information is extracted from the modal sample according to the sub-demand text obtained by the decomposition to obtain the corresponding information extraction result. Finally, all the information extraction results are integrated to generate the query result. In this process, the comprehensive understanding and analysis of multimodal input information is realized, and targeted information extraction is performed on the modal sample according to the sub-demand text, avoiding the broadness and lack of targetedness of the traditional solution in which the text description information of the database information is generated in advance. The method provided in this embodiment fully mines the information about user needs in multimodal data, which is of great significance for improving the utilization effect of data in the database, improving the accuracy of the reply and user experience. More specifically, it can extract key information that better meets user needs, so as to simplify the search results and improve the accuracy of the search results.

[0013] In one embodiment, the modal sample database stores a correspondence between modal samples and sample feature vectors corresponding to the modal samples; the step of searching the modal sample database for at least one modal sample matching the at least one modal information comprises:

[0014] For each type of modal information, feature extraction is performed on the modal information to obtain a corresponding feature vector;

[0015] Merge the eigenvectors corresponding to each modal information together to obtain an integrated vector;

[0016] At least one modality sample in the modality sample database that matches the at least one modality information is determined according to the similarity between the integrated vector and the sample feature vector corresponding to the modality sample in the modality sample database.

[0017] The display device provided in this embodiment realizes the comprehensive understanding and analysis of multimodal information, solves the problem that the format and content of different modal information input cannot be unified, and makes it easier to retrieve corresponding modal samples.

[0018] In one embodiment, the demand text is obtained by inputting the prompt word and the at least one modal information into the first large model and then outputting it; the processor performs feature extraction on the modal information to obtain a corresponding feature vector, and is configured as follows:

[0019] In the process of processing the modal information by the first large model, the feature vector output by the fully connected layer of the first large model is used as the feature vector corresponding to the modal information.

[0020] In one embodiment, the step of performing requirement decomposition on the requirement text based on the modal sample to obtain a sub-requirement text for the modal sample includes:

[0021] The modal sample and the requirement text are input into the second large model, and the sub-requirement text for the modal sample is output.

[0022] The display device provided in this embodiment can decompose the demand text to obtain a sub-demand text that is more in line with user needs and targeted at modal samples. The sub-demand text is more targeted and accurate. Information extraction based on the sub-demand text can extract key information that is more in line with user needs, thereby achieving the purpose of simplifying the search results and improving the accuracy of the search results.

[0023] In one embodiment, the step of generating a query result corresponding to the at least one modality information according to the information extraction result corresponding to the at least one modality sample includes:

[0024] The information extraction result corresponding to the at least one modal sample is input into the third largest model, and the query result corresponding to the at least one modal information is output.

[0025] In one embodiment, the step of inputting the information extraction result corresponding to the at least one modality sample into the third model and outputting the query result corresponding to the at least one modality information includes:

[0026] The prompt words of the third largest model are established, and the information extraction result corresponding to the at least one modal sample and the requirement text are input into the third largest model to obtain the query result corresponding to the at least one modal information output by the third largest model.

[0027] In a second aspect, the present application provides a display device, including:

[0028] At least one processor, the processor being configured to:

[0029] Based on at least one modality information used for querying, obtaining a requirement text for representing the user's query requirement;

[0030] In a modality sample database, searching for at least one modality sample matching the at least one modality information;

[0031] For each modality sample, based on the modality sample, the requirement text is decomposed to obtain a sub-requirement text for the modality sample;

[0032] Based on the sub-requirement text for the modal sample, extract information from the modal sample to obtain an information extraction result corresponding to the modal sample;

[0033] According to the information extraction result corresponding to the at least one modality sample, a query result corresponding to the at least one modality information is generated.

[0034] The display device provided in this embodiment obtains at least one modal information for query, and then generates a demand text that meets the user's needs after a comprehensive understanding and analysis of the input at least one modal information, and retrieves the corresponding retrieved modal sample in the database according to the at least one modal information. Based on the modal sample, the demand text is decomposed, and information is extracted from the modal sample according to the sub-demand text obtained by the decomposition to obtain the corresponding information extraction result. Finally, all the information extraction results are integrated to generate the query result. In this process, a comprehensive understanding and analysis of multimodal input information is achieved, and targeted information extraction is performed on the modal sample according to the sub-demand text, avoiding the broadness and lack of targetedness of the traditional solution in which text description information is generated for the database information in advance. The display device provided in this embodiment fully mines the information about user needs in the multimodal data, which is of great significance for improving the utilization effect of data in the database, improving the accuracy of the reply and the user experience. More specifically, it can extract key information that better meets the user's needs, so as to simplify the search results and improve the accuracy of the search results.

[0035] In one embodiment, the modality sample database stores a correspondence between modality samples and sample feature vectors corresponding to the modality samples; the processor executes in the modality sample database to query at least one modality sample matching the at least one modality information, and is configured to:

[0036] For each type of modal information, feature extraction is performed on the modal information to obtain a corresponding feature vector;

[0037] Merge the eigenvectors corresponding to each modal information together to obtain an integrated vector;

[0038] At least one modality sample in the modality sample database that matches the at least one modality information is determined according to the similarity between the integrated vector and the sample feature vector corresponding to the modality sample in the modality sample database.

[0039] The display device provided in this embodiment realizes the comprehensive understanding and analysis of multimodal information, solves the problem that the format and content of different modal information input cannot be unified, and makes it easier to retrieve corresponding modal samples.

[0040] In one embodiment, the demand text is obtained by inputting the prompt word and the at least one modal information into the first large model and then outputting it; the processor performs feature extraction on the modal information to obtain a corresponding feature vector, and is configured as follows:

[0041] In the process of processing the modal information by the first large model, the feature vector output by the fully connected layer of the first large model is used as the feature vector corresponding to the modal information.

[0042] In one embodiment, the processor performs requirement decomposition on the requirement text based on the modal sample to obtain a sub-requirement text for the modal sample, and is configured to:

[0043] The modal sample and the requirement text are input into the second large model, and the sub-requirement text for the modal sample is output.

[0044] The display device provided in this embodiment can decompose the demand text to obtain a sub-demand text that is more in line with user needs and targeted at modal samples. The sub-demand text is more targeted and accurate. Information extraction based on the sub-demand text can extract key information that is more in line with user needs, thereby achieving the purpose of simplifying the search results and improving the accuracy of the search results.

[0045] In one embodiment, the processor executes the information extraction result corresponding to the at least one modality sample to generate the query result corresponding to the at least one modality information, and is configured to:

[0046] The information extraction result corresponding to the at least one modal sample is input into the third largest model, and the query result corresponding to the at least one modal information is output.

[0047] In one embodiment, the processor executes inputting the information extraction result corresponding to the at least one modality sample into the third model, and outputting the query result corresponding to the at least one modality information, and is configured as follows:

[0048] The prompt words of the third largest model are established, and the information extraction result corresponding to the at least one modal sample and the requirement text are input into the third largest model to obtain the query result corresponding to the at least one modal information output by the third largest model.

[0049] In a third aspect, the present application provides a computer-readable storage medium having computer-executable instructions stored thereon. When the computer-executable instructions are executed by a processor, the method described in the first aspect is implemented.

[0050] In a fourth aspect, the present application provides a computer program product, comprising a computer program, which implements the method described in the first aspect when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the implementation methods in the embodiments of the present application or the related technologies, the following is a brief introduction to the drawings required for use in the embodiments or the related technology descriptions. Obviously, the drawings described below are some embodiments of the present application, and a person skilled in the art can also obtain other drawings based on these drawings.

[0052] Figure 1 A schematic diagram of a display device provided by an embodiment of the present application;

[0053] Figure 2 Another schematic diagram of a display device provided by an embodiment of the present application;

[0054] Figure 3 Another schematic diagram of a display device provided by an embodiment of the present application;

[0055] Figure 4 Another schematic diagram of a display device provided by an embodiment of the present application;

[0056] Figure 5 A schematic diagram of a flow chart of an information query method provided by an embodiment of the present application;

[0057] Figure 6 Another flowchart of an information query method provided by an embodiment of the present application;

[0058] Figure 7 Another flowchart of an information query method provided by an embodiment of the present application;

[0059] Figure 8 Another schematic diagram of a flow chart of an information query method provided by an embodiment of the present application;

[0060] Fig. 9 Another flowchart of an information query method provided by an embodiment of the present application;

[0061] Fig.10 A schematic diagram of an information query device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0062] The following embodiments are described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementations described in the following embodiments do not represent all implementations consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application as detailed in the claims.

[0063] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.

[0064] The terms "first", "second", "third", etc. in the specification and claims of this application and the above drawings are used to distinguish similar or similar objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances.

[0065] The terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0066] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware and / or software code that is capable of performing the functions associated with that element.

[0067] In the embodiment of the present application, the display device 200 generally refers to a device with image display and data processing capabilities. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.

[0068] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG. 1 , the user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.

[0069] The mobile terminal 300 can be used as a control device to perform human-computer interaction between the user and the display device 200. The mobile terminal 300 can also be used as a communication device to establish a communication connection with the display device 200 and perform data interaction. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, and achieve connection and communication through a network communication protocol to achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve a synchronous display function.

[0070] like Figure 1 As also shown in FIG. 4 , the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate and connect through a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0071] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), and the like.

[0072] Figure 2 Some embodiments of the present application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.

[0073] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0074] In some embodiments, the detector 230 is used to collect signals of the external environment or external interaction. For example, the detector 230 includes a light receiver, a sensor for collecting the intensity of ambient light; or, the detector 230 includes an image collector, such as a camera, which can be used to collect external environment scenes, user attributes or user interaction gestures; or, the detector 230 includes a sound collector, such as a microphone, etc., for receiving external sounds.

[0075] In some embodiments, the display 260 includes a display function component for presenting a picture, and a driving component for driving an image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, and components of a menu control interface and a user control UI interface.

[0076] In some embodiments, the communication device 220 is a component for communicating with an external device or server 400 according to various communication protocol types. The display device 200 may be provided with a plurality of communication devices 220 according to different supported communication modes. For example, when the display device 200 supports wireless network communication, the display device 200 may be provided with a communication device 220 including a WiFi function. When the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including a Bluetooth function.

[0077] The communication device 220 can enable the display device 200 to communicate with the external device or server 400 by wireless or wired connection. Among them, the wired connection can connect the display device 200 with the external device through components such as data cables and interfaces. The wireless connection can connect the display device 200 with the external device through wireless signals or wireless networks. The display device 200 can establish a connection relationship with the external device directly, or indirectly establish a connection relationship through a gateway, a router, a connection device, etc.

[0078] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first interface to an nth interface for input / output. The controller 250 controls the operation of the display device and responds to the user's operation through various software control programs stored in the memory. The controller 250 controls the overall operation of the display device 200.

[0079] In some embodiments, the controller 250 and the tuner-demodulator 210 may be located in different separate devices, that is, the tuner-demodulator 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0080] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).

[0081] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or may be an external audio output device of the display device 200. In particular, for the external audio output device of the display device 200, the display device 200 may also be provided with an external audio output terminal, and the audio output device may be connected to the display device 200 through the external audio output terminal to output the sound of the display device 200.

[0082] In some embodiments, the user input interface 280 may be used to receive instructions from a user.

[0083] Figure 3 Some embodiments of the present application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.

[0084] The control device 100 is configured to control the display device 200 , and can receive user input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200 .

[0085] In some embodiments, the control device 100 may be a smart device, for example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.

[0086] In some embodiments, Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .

[0087] The controller 110 includes a processor 112, a RAM 113, a ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between the internal components and the external and internal data processing functions.

[0088] The communication interface 130 implements communication of control signals and data signals with the display device 200 under the control of the controller 110. The communication interface 130 may include at least one of other near field communication modules such as a WiFi chip 131, a Bluetooth module 132, and an NFC module 133.

[0089] The user input / output interface 140 , wherein the input interface includes at least one of other input interfaces such as a microphone 141 , a touch panel 142 , a sensor 143 , and a button 144 .

[0090] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with a communication interface 130, such as a WiFi, Bluetooth, NFC or other module, and can encode the user input command through the WiFi protocol, Bluetooth protocol, or NFC protocol and send it to the display device 200.

[0091] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.

[0092] The power supply 180 is used to provide operating power support for each component of the control device 100 under the control of the controller.

[0093] In order to perform user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling hardware resources and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.

[0094] It should be noted that the operating system can be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0095] The operating system can be divided into different modules or layers according to the functions implemented, such as Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (Applications) layer (referred to as "application layer"), the application framework layer (Application Framework) layer (referred to as "framework layer"), the system library layer and the kernel layer.

[0096] In some embodiments, the application layer is used to provide services and interfaces for applications so that the display device 200 can run applications and interact with users based on the applications. At least one application can be run in the application layer, and these applications can be window programs, system settings programs, clock programs, etc. that come with the operating system; they can also be applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the above examples.

[0097] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications. The application framework layer includes some predefined functions. The application framework layer is equivalent to a processing center that determines the actions that applications in the application layer take. Applications can access system resources and obtain system services during execution through the API interface.

[0098] like Figure 4As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to the application package currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0099] In some embodiments, the activity manager is used to manage the life cycle of each application and the usual navigation back function, such as controlling the exit, opening, and back of the application. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, capturing the screen, and controlling the display window changes, for example, reducing the display window, shaking the display, distorting the display, etc.

[0100] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.

[0101] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4 As shown, the kernel layer may be configured with hardware drivers, and the drivers included in the kernel layer may be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFl driver, USB driver, HDMl driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0102] In some embodiments, the display device 200 includes the processor 112 as described above. The processor 112 is configured to execute instructions. Specifically, the processor 112 is configured to execute the information query method described below.

[0103] In some embodiments, the processor 112 is configured to: obtain a requirement text for characterizing a user query requirement based on at least one modal information used for query; query at least one modal sample matching the at least one modal information in a modal sample database; for each modal sample, perform requirement decomposition on the requirement text based on the modal sample to obtain a sub-requirement text for the modal sample; perform information extraction on the modal sample based on the sub-requirement text for the modal sample to obtain an information extraction result corresponding to the modal sample; and generate a query result corresponding to the at least one modal information based on the information extraction result corresponding to the at least one modal sample.

[0104] Among them, modality refers to different information carriers, such as text, pictures, audio and video. These different information carriers can be regarded as different modalities for transmitting and expressing information. At least one modality information can include text, pictures, videos, audio, knowledge graphs, etc.

[0105] Exemplarily, after receiving at least one modal information, a multimodal large language model can be used to understand different modal information, and then generate a demand text based on the understood modal information. For example, a user inputs multimodal information, and the multimodal information includes a picture and text, in which actor A is shown, and the text is such as "What classic film and television shots does this actor have, and what are the top five TV series with the highest film and television ratings?" After receiving the multimodal information, the multimodal large model understands and analyzes the picture and text. For example, the multimodal large model understands and analyzes the picture, and the result of the understanding analysis is that the actor in the picture is A. Combined with the text and the result of the understanding analysis (the actor is A), the final output demand text can be "Please list the classic film and television shots of actor A, and analyze the film and television ratings of all TV series he has played, and output the names of the top five TV series". Among them, the multimodal large model is an advanced artificial intelligence model that can process and understand multiple types of information. Specifically, the multimodal large model is a model that combines multimodal information such as text, images, videos, and audio for training. This large multimodal model can be replaced by other artificial intelligence models as long as they can understand different modal information.

[0106] The modal sample database stores multimodal samples and is a multimodal database. Modal samples can be understood as modal data. The modal data in the multimodal database may have many formats, such as text, pictures, videos, voice, knowledge graphs, etc. The queried modal sample has the same modality as the at least one modal information. For example, if the at least one modal information is a picture, the queried at least one modal sample matching the at least one modal information is also a picture.

[0107] For example, when building and retrieving a database, it may include data in various modalities or formats, such as text, pictures, videos, voice, knowledge graphs, user device logs, etc.

[0108] Before using the at least one modal information to query in the modal sample database, it is necessary to convert the at least one modal information into a format that can be recognized by the modal sample database. For example, if the modal sample database is a vector database, it is necessary to convert the at least one modal information into a feature vector (or vector data), and then query in the modal sample database based on the feature vector. Among them, a vector database is a database system specifically used to store, manage and query high-dimensional vector data. Vector databases perform particularly well in efficient query and similarity search of unstructured data such as images, text, and audio.

[0109] In order to make the final query result more in line with user needs, the required text needs to be obtained and the key information in the modal sample needs to be extracted. That is, after the modal samples of different modalities are retrieved, it is necessary to extract special information from the modal samples of each modality according to user needs.

[0110] Based on the modal sample, the requirement text is decomposed, which can be understood as combining the requirement text with the modal sample to obtain a more specific sub-requirement text. The sub-requirement text is for the modal sample and is used to extract key information in the modal sample that meets the user's needs.

[0111] For example, one of the modal samples queried is video, such as some video footage of actor A in TV series B. The requirement text is "Please list several classic video footages of actor A, analyze the video ratings of all TV series he has acted in, and output the names of the top five TV series." Combining the queried video footage of actor A in TV series B and the requirement text, the sub-requirement text can be obtained as "Is there a role played by actor A in this video, and is it a classic shot of this role?"

[0112] Exemplarily, based on the modal sample, the requirement text is decomposed to obtain the sub-requirement text for the modal sample by rewriting the requirement text in combination with the modal sample into a query statement for the modal sample. Then, based on the rewritten query statement, key information that meets the user's needs is extracted from the modal sample.

[0113] The sub-requirement text is used to extract information that meets the user's needs from the modal sample. For example, the sub-requirement text is "Is there a role played by actor A in this video? Is it a classic shot of this role?" After extracting information from the modal sample based on the sub-requirement text, the information extraction result obtained can be "This video contains the role of actor A, which is a classic shot in this TV series."

[0114] Exemplarily, the modal sample can be understood as modal data, and the way to extract information from the modal sample can be to determine through the sub-requirement text whether the modal data contains the information that the sub-requirement text wants to know. For example, the information that the sub-requirement text wants to know includes whether there is a role for actor A, whether the role is a classic shot, etc., then it is necessary to query the information such as actor A, actor A's role, and the role's shot in the modal data.

[0115] Exemplarily, the modal sample and the sub-requirement text may also be input into the multimodal large model, and the output of the multimodal large model may be the information extraction result corresponding to the modal sample.

[0116] The query result is generated by combining the corresponding information extraction results of each modality sample. The query result can be understood as a reply message to the user. The reply message is, for example, "The classic film and television shots of actor A are as follows, and the five TV series with the highest ratings he has starred in are "TV Series B", "TV Series C", "TV Series D", "TV Series E", and "TV Series F"".

[0117] The query result may be displayed to the user by controlling a display screen to display the query result, announcing the query result by voice, or directly sending the query result to a user terminal, etc., which is not limited in this embodiment.

[0118] Exemplarily, when generating the query result, it can be generated according to the data format support of the specific user interaction device. For example, the query result corresponding to the TV device can contain data in different formats such as text, video, and picture. For devices such as smart speakers, the generated query result only contains text data.

[0119] In summary, the display device provided in this embodiment obtains at least one modal information for query, and then generates a demand text that meets the user's needs after a comprehensive understanding and analysis of the input at least one modal information, and retrieves the corresponding retrieved modal sample in the database according to the at least one modal information. Based on the modal sample, the demand text is decomposed, and information is extracted from the modal sample according to the sub-demand text obtained by the decomposition to obtain the corresponding information extraction result. Finally, all the information extraction results are combined to generate the query result. In this process, a comprehensive understanding and analysis of multimodal input information is achieved, and targeted information extraction is performed on the modal sample according to the sub-demand text, avoiding the broadness and lack of targetedness of the traditional solution in which text description information is generated for the database information in advance. The display device provided in this embodiment fully mines the information about user needs in the multimodal data, which is of great significance for improving the utilization effect of data in the database, improving the accuracy of the reply and the user experience. More specifically, it can extract key information that better meets the user's needs, so as to simplify the search results and improve the accuracy of the search results.

[0120] In some embodiments, the modal sample database stores the correspondence between the modal samples and the sample feature vectors corresponding to the modal samples. The processor 112 executes in the modal sample database to query at least one modal sample matching the at least one modal information, and is configured to: for each modal information, extract features of the modal information to obtain a corresponding feature vector; merge the feature vectors corresponding to each modal information together to obtain an integrated vector; and determine at least one modal sample matching the at least one modal information in the modal sample database according to the similarity between the integrated vector and the sample feature vectors corresponding to the modal samples in the modal sample database.

[0121] Among them, the feature vector is information in numerical form, which summarizes the key information so that the machine learning model can understand and process the data. For example, if the modal information is an image, the feature vector is an array containing numerical values.

[0122] The purpose of feature extraction for each modal information is to convert each modal information into a feature vector with a unified format.

[0123] Exemplarily, for each type of modal information, the modal information may be input into a multi-modal large model to perform feature extraction on the modal information to obtain a corresponding feature vector.

[0124] The merging of feature vectors is used to combine multiple feature vectors into a unified feature representation, and finally obtain a vector, which is the integrated vector.

[0125] The merging of feature vectors is to unify the format and content of multiple modal information inputs.

[0126] Exemplarily, a special vectorization model may be used to map multiple modal information into a unified vector space to obtain vector representations for the multiple modal information, that is, to obtain the integrated vector.

[0127] Among them, the modal sample database is a vector database. The modal sample database stores the correspondence between the modal samples and the sample feature vectors corresponding to the modal samples. When searching in the modal sample database, first find the sample feature vector whose similarity with the integrated vector is greater than or equal to the preset similarity. Then, based on the correspondence between the found sample feature vector, the modal sample and the sample feature vector corresponding to the modal sample, determine the modal sample corresponding to the integrated vector. It should be noted that since the integrated vector contains the feature vector corresponding to each modal information, the modal sample corresponding to the integrated vector contains the modal sample corresponding to each modal information. For example, the integrated vector contains feature vectors corresponding to 3 types of modal information, and the modal sample corresponding to the integrated vector contains modal samples corresponding to 3 types of modal information.

[0128] The display device provided in this embodiment realizes the comprehensive understanding and analysis of multimodal information, solves the problem that the format and content of different modal information input cannot be unified, and makes it easier to retrieve corresponding modal samples.

[0129] In some embodiments, the requirement text is obtained by inputting the prompt word and the at least one modal information into the first large model and then outputting it; the processor 112 performs feature extraction on the modal information to obtain a corresponding feature vector, and is configured as: in the process of the first large model processing the modal information, the feature vector output by the fully connected layer of the first large model is used as the feature vector corresponding to the modal information.

[0130] The first large model can be the multimodal large model described above, or it can be other vectorized models, such as the CLIP (Contrastive Language-lmage Pre-training) model. The CLIP model consists of two parts: an image encoder and a text encoder. The image encoder is responsible for converting the image into a feature vector, which can be a convolutional neural network or a Transformer model. The text encoder is responsible for converting the text into a feature vector, which is usually a Transformer model. The two encoders realize cross-modal information interaction and fusion by sharing a vector space.

[0131] In order to save resources, the output data of the fully connected layer of the multimodal large model can be used as the feature vector corresponding to the modal information.

[0132] Optionally, a CLIP model may be trained separately, and the CLIP model may be used to perform vectorization processing of multimodal information.

[0133] Exemplarily, before adopting the first large model, it is necessary to organize relevant training data according to the task to be processed by the first large model, and then fine-tune the first large model using LoRA (Low-Rank Adaptation of Large Language Models) training or full-parameter training. In this embodiment, the task to be processed by the first large model refers to outputting the user's demand text based on at least one modal information input by the user, and the training data may include multiple modal information, demand text corresponding to each modal information, etc.

[0134] It should be noted that the requirement text is obtained by inputting the prompt word and the at least one modal information into the first large model and then outputting it. After receiving the at least one modal information, the display device uses the first large model to understand the information of different modalities, and by setting a special prompt word (Prompt), the first large model can convert the at least one modal information into the requirement text. The prompt word (Prompt) can be understood as an injection instruction, which is used to instruct the first large model to think about the problem and output the result according to the preset ideas. The prompt word can be manually written according to task requirements and business experience.

[0135] In some embodiments, the processor 112 performs requirement decomposition on the requirement text based on the modal sample to obtain a sub-requirement text for the modal sample, and is configured to: input the modal sample and the requirement text into a second large model, and output the sub-requirement text for the modal sample.

[0136] Among them, the second largest model can be a multimodal large model or other models, as long as it can output the sub-requirement text for the modal sample, and this embodiment does not make too many restrictions. The second largest model is used to perform query rewriting on the requirement text based on the modal sample. Query rewriting is a technical means to optimize search results. It obtains the sub-requirement text after a series of modifications and conversions to the requirement text, so as to achieve the purpose of better understanding user needs and providing more relevant search results. Compared with the requirement text, the sub-requirement text can better extract key information that meets user needs for the modal sample.

[0137] For example, one modal sample queried is a video, such as a portion of the film and television footage of actor A in TV series B. The requirement text is "Please list several classic film and television footage of actor A, analyze the film and television ratings of all TV series he has acted in, and output the names of the top five TV series." The modal sample (video) and the requirement text are input into the second largest model, and the output sub-requirement text for the modal sample (video) is "Is there a role played by actor A in this video, and is it a classic shot of this role?"

[0138] For example, in the text database, the text record information of the TV series played by actor A is retrieved: "TV series B (2008), director a, screenwriter b, starring c, film and television rating 9.3, 112916 people commented, etc.". Combined with this text record information, the demand text is queried and rewritten through the second largest model to "Is this a TV series played by actor A? If so, please extract the TV series name and film and television rating score from this record." The rewritten sub-demand text and this text record information are re-input into the first largest model (for example, the multimodal large model), and a reply is obtained, "This is a TV series played by actor A. The name of the TV series is "TV Series B" and the film and television rating is 9.3." This reply is the accurate information extraction result for the text record information.

[0139] By rewriting the query and extracting information for each modal sample (or each search result), accurate information extraction results for each modal sample are obtained.

[0140] Exemplarily, before adopting the second largest model, it is necessary to organize relevant training data according to the tasks to be processed by the second largest model, and then fine-tune the second largest model using LoRA training or full parameter training. In this embodiment, the task to be processed by the second largest model refers to query rewriting the requirement text based on the modal sample, that is, outputting a sub-requirement text for the modal sample based on the modal sample and the requirement text. The training data may include multiple modal samples, requirement texts corresponding to each modal sample, and sub-requirement texts corresponding to each modal sample.

[0141] The display device provided in this embodiment can decompose the demand text to obtain a sub-demand text that is more in line with user needs and targeted at modal samples. The sub-demand text is more targeted and accurate. Information extraction based on the sub-demand text can extract key information that is more in line with user needs, thereby achieving the purpose of simplifying the search results and improving the accuracy of the search results.

[0142] In some embodiments, the processor executes information extraction results corresponding to the at least one modal sample to generate query results corresponding to the at least one modal information, and is configured to: input the information extraction results corresponding to the at least one modal sample into the third largest model, and output the query results corresponding to the at least one modal information.

[0143] The third large model may be a multimodal large model or other large models, which is not limited in this embodiment.

[0144] According to the corresponding information extraction results of the modal samples, the third model is used to generate query results according to the demand text.

[0145] Exemplarily, a prompt word of the third largest model is established, and the information extraction result corresponding to the at least one modal sample and the demand text are input into the third largest model, and the query result corresponding to the at least one modal information output by the third largest model is obtained. That is to say, by establishing a special prompt word (Prompt), the third largest model is used to generate a complete query result by using the demand text and the precise information extraction result corresponding to each modal sample, as well as the multimodal information related to the reply. For example, after obtaining the classic film and television shots of actor A and the film and television ratings of the TV series he has participated in, the third largest model is used to sort the TV series and generate colloquial replies according to the film and television ratings, and the query result (or reply text) "Actor A's classic film and television shots are as follows, and the five TV series with the highest film and television ratings he has played are "TV Series B", "TV Series C", "TV Series D", "TV Series E", and "TV Series F"" and the video results of the film and television shots are output.

[0146] Exemplarily, before adopting the third largest model, it is necessary to organize relevant training data according to the tasks to be processed by the third largest model, and then fine-tune the third largest model using LoRA training or full parameter training. In this embodiment, the task to be processed by the third largest model refers to outputting a query result corresponding to at least one modal information based on the information extraction result corresponding to at least one modal sample and the requirement text. The training data may include information extraction results corresponding to at least one modal sample, requirement text, and query results corresponding to at least one modal information.

[0147] In one embodiment, the first large model, the second large model and the third large model described above are all multimodal large models. By making full use of the content understanding and question rewriting capabilities of the multimodal large model, sufficient and targeted information extraction and content integration are performed for the modal samples corresponding to the multimodal information, which greatly improves the utilization of multimodal information, improves the response effect during information retrieval and improves user experience.

[0148] For ease of understanding, some of the steps performed by the processor 112 are described below with more specific embodiments. Figure 8 The process diagram is shown.

[0149] like Figure 8 As shown, the first step is to obtain the at least one modal information. The at least one modal information may include text, voice, picture, video, other device synchronization information, etc. Other device synchronization information, for example, modal information sent by a smart air conditioner. The second step is to understand and analyze the at least one modal information based on a multimodal large model to obtain the demand text, and to vectorize the at least one modal information to obtain the integrated vector described above. The third step is to perform information retrieval on the integrated vector to obtain at least one modal sample corresponding to the at least one modal information. The fourth step is to perform Query rewriting and information extraction on the demand text in combination with each modal sample. The fifth step is to generate query results. The sixth step is to display the query results. The query results can be displayed or played.

[0150] For ease of understanding, some of the steps performed by the processor 112 are described below with more specific embodiments. Fig. 9 The process diagram is shown.

[0151] like Fig. 9 As shown, the first step is to obtain at least one modal information. The at least one modal information may include text, voice, picture, video, other device synchronization information, etc. Other device synchronization information, for example, modal information sent by smart air conditioners. The second step is to understand and analyze the at least one modal information based on the multimodal large model to obtain the demand text, and input the at least one modal information into the vectorization model to obtain the integrated vector described above. The third step is to use the integrated vector to query in the modal sample database to obtain at least one modal sample. At least one modal sample is a video sample, a picture sample, and a text sample as shown in the figure. The fourth step is to use the large model to rewrite the query for each modal sample in combination with the demand text. The fifth step is to use the rewritten sub-demand text to search again for each modal sample to obtain the information extraction result of the modal sample. The sixth step is to generate a query result by integrating the information extraction result of each modal sample.

[0152] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.

[0153] The following uses a processor of a display device as an example to illustrate how a display device implements an information query method.

[0154] The technical solution of the present application is described in detail below in conjunction with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0155] Figure 5 The present invention provides a flowchart of an information query method. The information query method can be applied to the display device provided in any of the above embodiments. The display device is not limited to a mobile phone, a tablet computer, a smart home device, a smart TV, etc.

[0156] like Figure 5 As shown, the method comprises the following steps:

[0157] S510, based on at least one modal information used for query, obtaining a demand text for representing the user's query demand.

[0158] Among them, modality refers to different information carriers, such as text, pictures, audio and video. These different information carriers can be regarded as different modalities for transmitting and expressing information. At least one modality information can include text, pictures, videos, audio, knowledge graphs, etc.

[0159] Exemplarily, after receiving at least one modal information, a multimodal large language model can be used to understand different modal information, and then generate a demand text based on the understood modal information. For example, a user inputs multimodal information, and the multimodal information includes a picture and text, in which actor A is shown, and the text is such as "What classic film and television shots does this actor have, and what are the top five TV series with the highest film and television ratings?" After receiving the multimodal information, the multimodal large model understands and analyzes the picture and text. For example, the multimodal large model understands and analyzes the picture, and the result of the understanding analysis is that the actor in the picture is A. Combined with the text and the result of the understanding analysis (the actor is A), the final output demand text can be "Please list the classic film and television shots of actor A, and analyze the film and television ratings of all TV series he has played, and output the names of the top five TV series". Among them, the multimodal large model is an advanced artificial intelligence model that can process and understand multiple types of information. Specifically, the multimodal large model is a model that combines multimodal information such as text, images, videos, and audio for training. This large multimodal model can be replaced by other artificial intelligence models as long as they can understand different modal information.

[0160] S520: Search a modality sample database for at least one modality sample that matches the at least one modality information.

[0161] The modal sample database stores multimodal samples and is a multimodal database. Modal samples can be understood as modal data. The modal data in the multimodal database may have many formats, such as text, pictures, videos, voice, knowledge graphs, etc. The queried modal sample has the same modality as the at least one modal information. For example, if the at least one modal information is a picture, the queried at least one modal sample matching the at least one modal information is also a picture.

[0162] For example, when building and retrieving a database, it may include data in various modalities or formats, such as text, pictures, videos, voice, knowledge graphs, user device logs, etc.

[0163] Before using the at least one modal information to query in the modal sample database, it is necessary to convert the at least one modal information into a format that can be recognized by the modal sample database. For example, if the modal sample database is a vector database, it is necessary to convert the at least one modal information into a feature vector (or vector data), and then query in the modal sample database based on the feature vector. Among them, a vector database is a database system specifically used to store, manage and query high-dimensional vector data. Vector databases perform particularly well in efficient query and similarity search of unstructured data such as images, text, and audio.

[0164] S530: For each modal sample, based on the modal sample, the requirement text is decomposed to obtain a sub-requirement text for the modal sample.

[0165] In order to make the final query result more in line with user needs, it is necessary to extract key information from the modality sample according to the demand text obtained in step 510. That is, after modality samples of different modalities are retrieved, it is necessary to extract special information from the modality samples of each modality according to user needs.

[0166] Based on the modal sample, the requirement text is decomposed, which can be understood as combining the requirement text with the modal sample to obtain a more specific sub-requirement text. The sub-requirement text is for the modal sample and is used to extract key information in the modal sample that meets the user's needs.

[0167] For example, one of the modal samples queried is video, such as some video footage of actor A in TV series B. The requirement text is "Please list several classic video footages of actor A, analyze the video ratings of all TV series he has acted in, and output the names of the top five TV series." Combining the queried video footage of actor A in TV series B and the requirement text, the sub-requirement text can be obtained as "Is there a role played by actor A in this video, and is it a classic shot of this role?"

[0168] Exemplarily, based on the modal sample, the requirement text is decomposed to obtain the sub-requirement text for the modal sample by rewriting the requirement text in combination with the modal sample into a query statement for the modal sample. Then, based on the rewritten query statement, key information that meets the user's needs is extracted from the modal sample.

[0169] S540: Based on the sub-requirement text for the modal sample, extract information from the modal sample to obtain an information extraction result corresponding to the modal sample.

[0170] The sub-requirement text is used to extract information that meets the user's needs from the modal sample. For example, the sub-requirement text is "Is there a role played by actor A in this video? Is it a classic shot of this role?" After extracting information from the modal sample based on the sub-requirement text, the information extraction result obtained can be "This video contains the role of actor A, which is a classic shot in this TV series."

[0171] Exemplarily, the modal sample can be understood as modal data, and the way to extract information from the modal sample can be to determine through the sub-requirement text whether the modal data contains the information that the sub-requirement text wants to know. For example, the information that the sub-requirement text wants to know includes whether there is a role for actor A, whether the role is a classic shot, etc., then it is necessary to query the information such as actor A, actor A's role, and the role's shot in the modal data.

[0172] Exemplarily, the modal sample and the sub-requirement text may also be input into the multimodal large model, and the output of the multimodal large model may be the information extraction result corresponding to the modal sample.

[0173] S550: Generate a query result corresponding to the at least one modal information according to the information extraction result corresponding to the at least one modal sample.

[0174] The query result is generated by combining the corresponding information extraction results of each modality sample. The query result can be understood as a reply message to the user. The reply message is, for example, "The classic film and television shots of actor A are as follows, and the five TV series with the highest ratings he has starred in are "TV Series B", "TV Series C", "TV Series D", "TV Series E", and "TV Series F"".

[0175] The query result may be displayed to the user by controlling a display screen to display the query result, announcing the query result by voice, or directly sending the query result to a user terminal, etc., which is not limited in this embodiment.

[0176] Exemplarily, when generating the query result, it can be generated according to the data format support of the specific user interaction device. For example, the query result corresponding to the TV device can contain data in different formats such as text, video, and picture. For devices such as smart speakers, the generated query result only contains text data.

[0177] In summary, the method provided in this embodiment obtains at least one modal information for query, and then generates a demand text that meets the user's needs after a comprehensive understanding and analysis of the input at least one modal information, and retrieves the corresponding retrieved modal sample in the database according to the at least one modal information. Based on the modal sample, the demand text is decomposed, and information is extracted from the modal sample according to the sub-demand text obtained by the decomposition to obtain the corresponding information extraction result. Finally, all information extraction results are combined to generate the query result. In this process, a comprehensive understanding and analysis of multimodal input information is achieved, and targeted information extraction is performed on the modal sample according to the sub-demand text, avoiding the broadness and lack of targetedness of the traditional solution in which text description information is generated for the database information in advance. The method provided in this embodiment fully mines the information about user needs in multimodal data, which is of great significance for improving the utilization effect of data in the database, improving the accuracy of replies and user experience. More specifically, it can extract key information that better meets user needs, so as to simplify the search results and improve the accuracy of the search results.

[0178] For some examples, see Figure 6 The modal sample database stores the corresponding relationship between the modal samples and the sample feature vectors corresponding to the modal samples. Step S520 includes:

[0179] S610: For each type of modal information, extract features of the modal information to obtain a corresponding feature vector.

[0180] Among them, the feature vector is information in numerical form, which summarizes the key information so that the machine learning model can understand and process the data. For example, if the modal information is an image, the feature vector is an array containing numerical values.

[0181] The purpose of feature extraction for each modal information is to convert each modal information into a feature vector with a unified format.

[0182] Exemplarily, for each type of modal information, the modal information may be input into a multi-modal large model to perform feature extraction on the modal information to obtain a corresponding feature vector.

[0183] S620, merging the feature vectors corresponding to each type of modal information together to obtain an integrated vector.

[0184] The merging of feature vectors is used to combine multiple feature vectors into a unified feature representation, and finally obtain a vector, which is the integrated vector.

[0185] The merging of feature vectors is to unify the format and content of multiple modal information inputs.

[0186] Exemplarily, a special vectorization model may be used to map multiple modal information into a unified vector space to obtain vector representations for the multiple modal information, that is, to obtain the integrated vector.

[0187] S630: Determine at least one modal sample in the modal sample database that matches the at least one modal information based on the similarity between the integrated vector and the sample feature vectors corresponding to the modal samples in the modal sample database.

[0188] Among them, the modal sample database is a vector database. The modal sample database stores the correspondence between the modal samples and the sample feature vectors corresponding to the modal samples. When searching in the modal sample database, first find the sample feature vector whose similarity with the integrated vector is greater than or equal to the preset similarity. Then, based on the correspondence between the found sample feature vector, the modal sample and the sample feature vector corresponding to the modal sample, determine the modal sample corresponding to the integrated vector. It should be noted that since the integrated vector contains the feature vector corresponding to each modal information, the modal sample corresponding to the integrated vector contains the modal sample corresponding to each modal information. For example, the integrated vector contains feature vectors corresponding to 3 types of modal information, and the modal sample corresponding to the integrated vector contains modal samples corresponding to 3 types of modal information.

[0189] The method provided in this embodiment realizes the comprehensive understanding and analysis of multimodal information, solves the problem that the format and content of different modal information input cannot be unified, and makes it easier to retrieve corresponding modal samples.

[0190] For some examples, see Figure 7 , step S610 includes:

[0191] S710, in the process of the first large model processing the modal information, the feature vector output by the fully connected layer of the first large model is used as the feature vector corresponding to the modal information.

[0192] The first large model can be the multimodal large model described above, or it can be other vectorized models, such as the CLIP (Contrastive Language-lmage Pre-training) model. The CLIP model consists of two parts: an image encoder and a text encoder. The image encoder is responsible for converting the image into a feature vector, which can be a convolutional neural network or a Transformer model. The text encoder is responsible for converting the text into a feature vector, which is usually a Transformer model. The two encoders realize cross-modal information interaction and fusion by sharing a vector space.

[0193] In order to save resources, the output data of the fully connected layer of the multimodal large model can be used as the feature vector corresponding to the modal information.

[0194] Optionally, a CLIP model may be trained separately, and the CLIP model may be used to perform vectorization processing of multimodal information.

[0195] Exemplarily, before adopting the first large model, it is necessary to organize relevant training data according to the task to be processed by the first large model, and then fine-tune the first large model using LoRA training or full parameter training. In this embodiment, the task to be processed by the first large model refers to outputting the user's demand text based on at least one modal information input by the user, and the training data may include multiple modal information, demand text corresponding to each modal information, etc.

[0196] It should be noted that the requirement text is obtained by inputting the prompt word and the at least one modal information into the first large model and then outputting it. After receiving the at least one modal information, the display device uses the first large model to understand the information of different modalities, and by setting a special prompt word (Prompt), the first large model can convert the at least one modal information into the requirement text. The prompt word (Prompt) can be understood as an injection instruction, which is used to instruct the first large model to think about the problem and output the result according to the preset ideas. The prompt word can be manually written according to task requirements and business experience.

[0197] In some embodiments, step S530 performs requirement decomposition on the requirement text based on the modal sample to obtain sub-requirement texts for the modal sample, including:

[0198] Step 1: input the modal sample and the requirement text into the second largest model, and output the sub-requirement text for the modal sample.

[0199] Among them, the second largest model can be a multimodal large model or other models, as long as it can output the sub-requirement text for the modal sample, and this embodiment does not make too many restrictions. The second largest model is used to perform query rewriting on the requirement text based on the modal sample. Query rewriting is a technical means to optimize search results. It obtains the sub-requirement text after a series of modifications and conversions to the requirement text, so as to achieve the purpose of better understanding user needs and providing more relevant search results. Compared with the requirement text, the sub-requirement text can better extract key information that meets user needs for the modal sample.

[0200] For example, one modal sample queried is a video, such as a portion of the film and television footage of actor A in TV series B. The requirement text is "Please list several classic film and television footage of actor A, analyze the film and television ratings of all TV series he has acted in, and output the names of the top five TV series." The modal sample (video) and the requirement text are input into the second largest model, and the output sub-requirement text for the modal sample (video) is "Is there a role played by actor A in this video, and is it a classic shot of this role?"

[0201] For example, in the text database, the text record information of the TV series played by actor A is retrieved: "TV series B (2008), director a, screenwriter b, starring c, film and television rating 9.3, 112916 people commented, etc.". Combined with this text record information, the demand text is queried and rewritten through the second largest model to "Is this a TV series played by actor A? If so, please extract the TV series name and film and television rating score from this record." The rewritten sub-demand text and this text record information are re-input into the first largest model (for example, the multimodal large model), and a reply is obtained, "This is a TV series played by actor A. The name of the TV series is "TV Series B" and the film and television rating is 9.3." This reply is the accurate information extraction result for the text record information.

[0202] By rewriting the query and extracting information for each modal sample (or each search result), accurate information extraction results for each modal sample are obtained.

[0203] Exemplarily, before adopting the second largest model, it is necessary to organize relevant training data according to the tasks to be processed by the second largest model, and then fine-tune the second largest model using LoRA training or full parameter training. In this embodiment, the task to be processed by the second largest model refers to query rewriting the requirement text based on the modal sample, that is, outputting a sub-requirement text for the modal sample based on the modal sample and the requirement text. The training data may include multiple modal samples, requirement texts corresponding to each modal sample, and sub-requirement texts corresponding to each modal sample.

[0204] The method provided in this embodiment can decompose the demand text to obtain a sub-demand text that is more in line with user needs and targeted at modal samples. The sub-demand text is more targeted and accurate. Further information extraction based on the sub-demand text can extract key information that is more in line with user needs, thereby achieving the purpose of simplifying the search results and improving the accuracy of the search results.

[0205] In some embodiments, step S550 includes:

[0206] Step 1: Input the information extraction result corresponding to the at least one modal sample into the third model, and output the query result corresponding to the at least one modal information.

[0207] The third large model may be a multimodal large model or other large models, which is not limited in this embodiment.

[0208] According to the corresponding information extraction results of the modal samples, the third model is used to generate query results according to the demand text.

[0209] Exemplarily, a prompt word of the third largest model is established, and the information extraction result corresponding to the at least one modal sample and the demand text are input into the third largest model, and the query result corresponding to the at least one modal information output by the third largest model is obtained. That is to say, by establishing a special prompt word (Prompt), the third largest model is used to generate a complete query result by using the demand text and the accurate information extraction result corresponding to each modal sample, as well as the multimodal information related to the reply. For example, after obtaining the classic film and television shots of actor A and the film and television ratings of the TV series he has participated in, the third largest model is used to sort the TV series and generate colloquial replies according to the film and television ratings, and the query result (or reply text) "Actor A's classic film and television shots are as follows. The five TV series with the highest film and television ratings he has played are "TV Series B", "TV Series C", "TV Series D", "TV Series E", and "TV Series F"" and the video results of the film and television shots are output.

[0210] Exemplarily, before adopting the third largest model, it is necessary to organize relevant training data according to the tasks to be processed by the third largest model, and then fine-tune the third largest model using LoRA training or full parameter training. In this embodiment, the task to be processed by the third largest model refers to outputting a query result corresponding to at least one modal information based on the information extraction result corresponding to at least one modal sample and the requirement text. The training data may include information extraction results corresponding to at least one modal sample, requirement text, and query results corresponding to at least one modal information.

[0211] In one embodiment, the first large model, the second large model and the third large model described above are all multimodal large models. By making full use of the content understanding and question rewriting capabilities of the multimodal large model, sufficient and targeted information extraction and content integration are performed for the modal samples corresponding to the multimodal information, which greatly improves the utilization of multimodal information, improves the response effect during information retrieval and improves user experience.

[0212] For ease of understanding, the following describes part of the process of the information query method with a more specific embodiment. Figure 8 The process diagram is shown.

[0213] like Figure 8 As shown, the first step is to obtain the at least one modal information. The at least one modal information may include text, voice, picture, video, other device synchronization information, etc. Other device synchronization information, for example, modal information sent by a smart air conditioner. The second step is to understand and analyze the at least one modal information based on a multimodal large model to obtain the demand text, and to vectorize the at least one modal information to obtain the integrated vector described above. The third step is to perform information retrieval on the integrated vector to obtain at least one modal sample corresponding to the at least one modal information. The fourth step is to perform Query rewriting and information extraction on the demand text in combination with each modal sample. The fifth step is to generate query results. The sixth step is to display the query results. The query results can be displayed or played.

[0214] For ease of understanding, the following describes part of the process of the information query method with a more specific embodiment. Fig. 9 The process diagram is shown.

[0215] like Fig. 9As shown, the first step is to obtain at least one modal information. The at least one modal information may include text, voice, picture, video, other device synchronization information, etc. Other device synchronization information, for example, modal information sent by smart air conditioners. The second step is to understand and analyze the at least one modal information based on the multimodal large model to obtain the demand text, and input the at least one modal information into the vectorization model to obtain the integrated vector described above. The third step is to use the integrated vector to query in the modal sample database to obtain at least one modal sample. At least one modal sample is a video sample, a picture sample, and a text sample as shown in the figure. The fourth step is to use the large model to rewrite the query for each modal sample in combination with the demand text. The fifth step is to use the rewritten sub-demand text to search again for each modal sample to obtain the information extraction result of the modal sample. The sixth step is to generate a query result by integrating the information extraction result of each modal sample.

[0216] Fig.10 Schematic diagram of the structure of an information query device 10 provided in this application. Fig.10 As shown, the device 10 is applied to a display device provided in any of the above embodiments. The device 10 includes:

[0217] The acquisition module 11 is used to acquire a demand text for representing the user's query demand based on at least one modal information used for query.

[0218] The query module 12 is used to query at least one modality sample matching the at least one modality information in the modality sample database.

[0219] The acquisition module 11 is also used to perform requirement decomposition on the requirement text for each modal sample based on the modal sample to obtain sub-requirement text for the modal sample.

[0220] The information extraction module 13 is used to extract information from the modal sample based on the sub-requirement text of the modal sample to obtain an information extraction result corresponding to the modal sample.

[0221] The result generating module 14 is used to generate a query result corresponding to the at least one modality information according to the information extraction result corresponding to the at least one modality sample.

[0222] In one embodiment, the modal sample database stores the correspondence between the modal samples and the sample feature vectors corresponding to the modal samples. The query module 12 is used to: for each modal information, extract the features of the modal information to obtain the corresponding feature vector; merge the feature vectors corresponding to each modal information to obtain an integrated vector; and determine at least one modal sample in the modal sample database that matches the at least one modal information according to the similarity between the integrated vector and the sample feature vectors corresponding to the modal samples in the modal sample database.

[0223] In one embodiment, the requirement text is obtained by inputting the prompt word and the at least one modal information into the first large model and then outputting it; the query module 12 is used to: in the process of the first large model processing the modal information, the feature vector output by the fully connected layer of the first large model is used as the feature vector corresponding to the modal information.

[0224] In one embodiment, the information extraction module 13 is used to: input the modal sample and the requirement text into the second large model, and output the sub-requirement text for the modal sample.

[0225] In one embodiment, the result generation module 14 is used to: input the information extraction result corresponding to the at least one modal sample into the third largest model, and output the query result corresponding to the at least one modal information.

[0226] In one embodiment, the result generation module 14 is used to: establish prompt words for the third largest model, and input the information extraction results corresponding to the at least one modal sample and the requirement text into the third largest model to obtain the query results corresponding to the at least one modal information output by the third largest model.

[0227] The information query device provided in the embodiment of the present application can execute the information query method in the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here. Fig.10 The division of the modules shown is only a schematic diagram, and the present application does not limit the division of the modules and the naming of the modules.

[0228] The present application also provides a computer-readable storage medium, which may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes. Specifically, the computer-readable storage medium stores program instructions, and the program instructions are used for the methods in the above embodiments.

[0229] The present application also provides a program product, which includes an execution instruction, which is stored in a readable storage medium. At least one control module of the display device can read the execution instruction from the readable storage medium, and at least one control module executes the execution instruction so that the display device implements the handwriting erasing method provided by the various embodiments described above.

[0230] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0231] For the convenience of explanation, the above description has been made in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations can be obtained. The selection and description of the above embodiments are to better explain the principles and practical applications, so that those skilled in the art can better use the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. A display device, characterized in that: include: At least one processor, the processor being configured to: Based on at least one modality information used for querying, obtaining a requirement text for representing the user's query requirement; In a modality sample database, searching for at least one modality sample matching the at least one modality information; For each modality sample, based on the modality sample, the requirement text is decomposed to obtain a sub-requirement text for the modality sample; Based on the sub-requirement text for the modal sample, extract information from the modal sample to obtain an information extraction result corresponding to the modal sample; According to the information extraction result corresponding to the at least one modality sample, a query result corresponding to the at least one modality information is generated.

2. The display device according to claim 1, characterized in that The modality sample database stores a correspondence between modality samples and sample feature vectors corresponding to the modality samples; the processor executes in the modality sample database to query at least one modality sample matching the at least one modality information, and is configured to: For each type of modal information, feature extraction is performed on the modal information to obtain a corresponding feature vector; Merge the eigenvectors corresponding to each modal information together to obtain an integrated vector; At least one modality sample in the modality sample database that matches the at least one modality information is determined according to the similarity between the integrated vector and the sample feature vectors corresponding to the modality samples in the modality sample database.

3. The display device according to claim 2, characterized in that The demand text is obtained by inputting the prompt word and the at least one modal information into the first large model and then outputting it; the processor performs feature extraction on the modal information to obtain a corresponding feature vector, and is configured as follows: In the process of processing the modal information by the first large model, the feature vector output by the fully connected layer of the first large model is used as the feature vector corresponding to the modal information.

4. The display device according to claim 1, characterized in that The processor performs requirement decomposition on the requirement text based on the modal sample to obtain a sub-requirement text for the modal sample, and is configured to: The modal sample and the requirement text are input into the second large model, and the sub-requirement text for the modal sample is output.

5. The display device according to claim 1, characterized in that The processor executes the information extraction result corresponding to the at least one modality sample to generate a query result corresponding to the at least one modality information, and is configured to: The information extraction result corresponding to the at least one modal sample is input into the third largest model, and the query result corresponding to the at least one modal information is output.

6. The display device according to claim 5, characterized in that The processor executes inputting the information extraction result corresponding to the at least one modality sample into the third model, and outputting the query result corresponding to the at least one modality information, and is configured to: The prompt words of the third largest model are established, and the information extraction result corresponding to the at least one modal sample and the requirement text are input into the third largest model to obtain the query result corresponding to the at least one modal information output by the third largest model.

7. An information query method, characterized in that: Applied to the display device as claimed in any one of claims 1 to 6; the method comprising: Based on at least one modality information used for querying, obtaining a requirement text for representing the user's query requirement; In a modality sample database, searching for at least one modality sample matching the at least one modality information; For each modality sample, based on the modality sample, the requirement text is decomposed to obtain a sub-requirement text for the modality sample; Based on the sub-requirement text for the modal sample, extract information from the modal sample to obtain an information extraction result corresponding to the modal sample; According to the information extraction result corresponding to the at least one modality sample, a query result corresponding to the at least one modality information is generated.

8. The method according to claim 7, characterized in that The modality sample database stores a correspondence between modality samples and sample feature vectors corresponding to the modality samples; the step of searching the modality sample database for at least one modality sample matching the at least one modality information comprises: For each type of modal information, feature extraction is performed on the modal information to obtain a corresponding feature vector; Merge the eigenvectors corresponding to each modal information together to obtain an integrated vector; At least one modality sample in the modality sample database that matches the at least one modality information is determined according to the similarity between the integrated vector and the sample feature vectors corresponding to the modality samples in the modality sample database.

9. The method according to claim 8, characterized in that The demand text is obtained by inputting the prompt word and the at least one modal information into the first large model and then outputting it; the processor performs feature extraction on the modal information to obtain a corresponding feature vector, and is configured as follows: In the process of processing the modal information by the first large model, the feature vector output by the fully connected layer of the first large model is used as the feature vector corresponding to the modal information.

10. The method according to claim 7, characterized in that The step of performing requirement decomposition on the requirement text based on the modal sample to obtain a sub-requirement text for the modal sample includes: The modal sample and the requirement text are input into the second large model, and the sub-requirement text for the modal sample is output.