Information processing method and device, equipment, medium and program product

By supporting the input of multimodal query information and cross-modal feature fusion technology, the accuracy and quality problems caused by single-modal queries in traditional search engines are solved, achieving more efficient information search and more accurate query results.

CN120632027APending Publication Date: 2025-09-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510714325.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional search engines only allow query input in a single information modality, which leads to a decrease in the accuracy and quality of query results and fails to fully express the user's search intent.

Method used

Provided are an information processing method and device that support the input and processing of multimodal query information. Through cross-modal feature fusion technology, they can identify and understand the user's multimodal query intentions and generate more accurate query results.

Benefits of technology

By combining multimodal query information, the accuracy of query results and search quality are significantly improved, enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632027A_ABST
    Figure CN120632027A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an information processing method and device, equipment, a medium and a program product. The method comprises the following steps: displaying an information search entry; in response to a trigger operation on the information search entry, receiving input multi-modal query information; and displaying a query result obtained by performing information search based on the multi-modal query information. By adopting the embodiment of the invention, the search accuracy can be remarkably improved by utilizing the complementarity between the sub-query information of different information modalities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, in particular to the field of information search, and specifically to an information processing method, an information processing apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Search engine technology is widely used in various industries, such as finance, education, publishing, and pharmaceuticals, to solve the problem that users have difficulty finding accurate information in massive amounts of information.

[0003] Search engines primarily analyze user-entered queries (i.e., questions used to find information) to find matching results. However, traditional search engines often allow users to enter simple queries, which significantly reduces the accuracy of search results and, consequently, the quality of search results. Summary of the Invention

[0004] Embodiments of the present application provide an information processing method, apparatus, device, medium, and program product that can significantly improve search accuracy by utilizing the complementarity between sub-query information of different information modalities.

[0005] In one aspect, an embodiment of the present application provides an information processing method, the method comprising:

[0006] Displaying an information search entry, which is used to trigger an information search;

[0007] In response to a triggering operation on an information search portal, receiving input multimodal query information; the multimodal query information includes sub-query information of at least two information modalities;

[0008] Displays query results obtained by searching for information based on multimodal query information.

[0009] On the other hand, an embodiment of the present application provides an information processing device, the device comprising:

[0010] A display unit, used to display an information search entry, which is used to trigger the execution of an information search;

[0011] A processing unit, configured to receive input multimodal query information in response to a triggering operation on an information search portal; the multimodal query information includes sub-query information of at least two information modalities;

[0012] The processing unit is further configured to display query results obtained by searching for information based on the multimodal query information.

[0013] In one implementation, the processing unit is configured to, in response to a triggering operation on an information search portal, receive input multimodal query information and specifically:

[0014] In response to a triggering operation on an information search portal, an information search interface is displayed; the information search interface includes an information input area; wherein the information search portal includes at least one of the following: a search bar in a content search interface, a dialog box in an interactive dialogue interface, and a search pop-up window in a service interface;

[0015] According to the information input operation performed in the information input area, the input multimodal query information is displayed in the information input area; the information modality includes at least one of the following: text, image, voice, video, animation and file.

[0016] In one implementation, the information search interface further includes a multimodal input prompt, which is used to prompt that the information search interface allows the input of multimodal query information; the information input area includes multiple sub-input areas, each sub-input area corresponding to an information modality; and the processing unit is configured to display the input multimodal query information in the information input area based on the information input operation performed in the information input area, specifically for:

[0017] In response to an information input operation performed in at least two of the plurality of sub-input areas, displaying sub-query information belonging to a corresponding information modality in a corresponding sub-input area;

[0018] According to a confirmation operation on the sub-query information displayed in the at least two sub-input areas, the sub-query information in the at least two sub-input areas is combined into multimodal query information.

[0019] In one implementation, the information input area further includes an area adding option; and the processing unit is further configured to:

[0020] Displaying a modal selection window according to a triggering operation for adding an option to a region; the modal selection window includes a region creation option corresponding to at least one information modality;

[0021] In response to a selection operation on a target area creation option in the modal selection window, a target sub-input area corresponding to a target information modality is displayed in the information input area; the target information modality is an information modality corresponding to the target area creation option, and the target area creation option is any area creation option in the modal selection window.

[0022] In one implementation, the multimodal input prompt is presented as a modal combination option; the processing unit is further configured to:

[0023] In response to a selection operation on the modal combination option, displaying a plurality of sub-input areas in the information input area;

[0024] The processing unit is also used to:

[0025] In response to a deselection operation on the modal combination option, a plurality of sub-input areas are deleted in the information input area.

[0026] In one implementation, the processing unit, when displaying the input multimodal query information in the information input area according to the information input operation performed in the information input area, is specifically configured to:

[0027] In response to an information input operation performed in the information input area, receiving input of first sub-query information;

[0028] If the initial query intent obtained based on the analysis of the first sub-query information does not meet the intent recognition condition, prompt input information is displayed; the prompt input information is used to prompt: input a second sub-query information representing the same query intent as the first sub-query information;

[0029] In response to an information input operation for the second sub-query information, the input second sub-query information is received; the multimodal query information includes the first sub-query information and the second sub-query information.

[0030] In one implementation, the processing unit is further configured to:

[0031] According to the closing operation performed on the input prompt information, a notification message is displayed; the notification message is used to prompt an information search based on the first sub-query information.

[0032] In one implementation, the prompt input information is used to prompt that the second sub-query information to be input belongs to the first information modality; the processing unit is further configured to:

[0033] According to the mode switching operation on the first information mode, the switched second information mode is displayed; the second information mode is different from the first information mode.

[0034] In one implementation, the multimodal query information includes an image belonging to an image modality; and the processing unit is further configured to:

[0035] In response to an image input operation, displaying the input image;

[0036] According to the information annotation operation performed on the image, the annotated annotation information is displayed at the annotated position;

[0037] In response to a confirmation operation on the annotated image, the image and the annotated image are respectively used as two sub-query information included in the multimodal query information.

[0038] In one implementation, the processing unit is further configured to:

[0039] Displays the modality identifier corresponding to each sub-query information contained in the multimodal query information; any modality identifier is used to indicate the information modality to which the corresponding sub-query information belongs;

[0040] In response to a deletion operation performed on a target modality identifier, performing an information search to generate a new query result based on the remaining sub-query information in the multimodal query information except the sub-query information corresponding to the deleted target modality identifier; the target modality identifier is a modality identifier corresponding to any sub-query information;

[0041] Display new query results.

[0042] In one implementation, the processing unit, when used to display query results obtained by performing an information search based on multimodal query information, is specifically configured to:

[0043] Displaying a result display interface, the result display interface including query results obtained from the information search based on the multimodal query information; the query results including recommended items that meet the query intent represented by the multimodal query information, and the query results including an item purchase portal for the recommended items;

[0044] In response to a trigger operation on the item purchase portal, the purchase details corresponding to the recommended item are displayed on the result display interface; the purchase details include payment information and item information of the recommended item;

[0045] Output the payment result based on the payment operation performed in the result display interface.

[0046] In one implementation, the query result includes query interpretation information that conforms to the query intent represented by the multimodal query information and recommended items that conform to the query intent; the processing unit is further configured to:

[0047] Perform feature extraction and fusion processing on each sub-query information in the multimodal query information to generate cross-modal fusion features;

[0048] Perform contextual semantic prediction on cross-modal fusion features to obtain query intent represented by multimodal query information;

[0049] Search the knowledge base for information based on the query intent and generate query interpretation information that matches the query intent; and

[0050] Search the item library for recommended items that meet the query intent.

[0051] In one implementation, the processing unit is configured to perform feature extraction and fusion processing on each sub-query information in the multimodal query information to generate a cross-modal fusion feature, specifically for:

[0052] Using the feature extraction rules corresponding to the information modality, feature encoding is performed on the sub-query information belonging to the corresponding information modality to obtain the initial feature information corresponding to each sub-query information;

[0053] Mapping the initial feature information corresponding to each sub-query information to the cross-modal feature shared space to obtain the target feature information of each sub-query information in the cross-modal feature shared space;

[0054] The information correlation between the target feature information corresponding to at least two sub-query information is calculated, and feature fusion processing is performed according to the information correlation between the target feature information to generate cross-modal fusion features.

[0055] In another aspect, an embodiment of the present application provides a computer device, comprising:

[0056] a processor for loading and executing computer programs;

[0057] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned information processing method is implemented.

[0058] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor and executing the above-mentioned information processing method.

[0059] On the other hand, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed by a processor, the above-mentioned information processing method is implemented.

[0060] In an embodiment of the present application, when the information search entry for triggering the execution of an information search is triggered, it indicates that the user has a need to conduct an information search, and at this time, the multimodal query information input by the user is received; wherein, the multimodal query information contains sub-query information of at least two information modes, that is, the multimodal query information contains at least two sub-query information belonging to different information modes, such as text belonging to a text mode and an image belonging to an image mode. All sub-query information belonging to different information modes contained in the multimodal query information is used to describe the same search intention. Compared with information search based on sub-query information of a single information mode, the semantic incompatibility between multiple sub-query information of different information modes is used to help improve and supplement the search intention that the user wants to express. In this way, information search based on multimodal query information with clearer search intention can obtain query results that are more consistent with the search intention represented by the multimodal query information, thereby improving the search and generation accuracy of the query results, greatly improving the quality of information search, and improving the user search experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0062] Figure 1 This is a schematic diagram of the architecture of an information processing system provided by an exemplary embodiment of the present application;

[0063] Figure 2 This is a flowchart of an information processing method provided by an exemplary embodiment of the present application;

[0064] Figure 3a is a schematic diagram of an exemplary embodiment of the present application providing a search engine that is a search browser;

[0065] Figure 3b This is a schematic diagram of a search engine deployed in a social application provided by an exemplary embodiment of the present application;

[0066] Figure 4 This is a schematic diagram of an information search portal integrated into a dialog box provided by an exemplary embodiment of the present application;

[0067] Figure 5 This is a schematic diagram of an information search portal integrated into a search pop-up window provided by an exemplary embodiment of the present application;

[0068] Figure 6a This is a schematic diagram of inputting multimodal query information in an information input area provided by an exemplary embodiment of the present application;

[0069] Figure 6b is a schematic diagram of an information input area including multiple sub-input areas provided by an exemplary embodiment of the present application;

[0070] Figure 7 This is a schematic diagram of displaying modal combination options in an information search interface provided by an exemplary embodiment of the present application;

[0071] Figure 8 is a schematic diagram of a user-created sub-input area provided by an exemplary embodiment of the present application;

[0072] Figure 9a This is a schematic diagram of a user inputting second sub-query information when inputting prompt information, provided by an exemplary embodiment of the present application;

[0073] Figure 9b This is a schematic diagram of a user performing a mode switching operation provided by an exemplary embodiment of the present application;

[0074] Figure 10 is a schematic diagram of performing a post-processing operation on an image provided by an exemplary embodiment of the present application;

[0075] Figure 11 is a flowchart of another information processing method provided by an exemplary embodiment of the present application;

[0076] Figure 12 is a schematic diagram of dynamically adjusting multimodal query information provided by an exemplary embodiment of the present application;

[0077] Figure 13 This is a schematic diagram of a one-key smart payment provided by an exemplary embodiment of the present application;

[0078] Figure 14 This is a technical architecture diagram of an information processing method provided by an exemplary embodiment of the present application;

[0079] Figure 15 This is a schematic diagram of components of each layer included in a technical architecture provided by an exemplary embodiment of the present application;

[0080] Figure 16 This is a schematic diagram of the background technology flow of an information processing method provided by an exemplary embodiment of the present application;

[0081] Figure 17 This is a flowchart of payment processing in an information processing method provided by an exemplary embodiment of the present application;

[0082] Figure 18 is a structural diagram of an information processing device provided by an exemplary embodiment of the present application;

[0083] Figure 19 It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0084] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0085] In the embodiments of the present application, an information processing solution is proposed, specifically an information search solution based on artificial intelligence (AI). The following is a brief introduction to the technical terms and related concepts involved in the information processing solution provided in the embodiments of the present application, including:

[0086] 1. Information search.

[0087] Information search, also known as information query or information retrieval, refers to the search engine's search for information based on the query information entered by the user, aiming to find query results that match the query information.

[0088] (1) A search engine is a retrieval system that uses search strategies to retrieve query results from the Internet and feeds them back to users based on user query requirements and specific algorithms. This retrieval system mainly collects information from the Internet, organizes and processes the information, and provides retrieval services to users, helping them quickly find the information they need.

[0089] (2) The query information input by the user refers to the information that can represent the user's search intention for this information search; for example, the query information is what kind of cold causes a dry and sore throat, and the search intention represented by this query information is to query the cause of the dry and sore throat. It is worth noting that different query information may have different information modalities; information modality can be simply referred to as modality, which refers to the information representation form of the query information or the user's perception form of the query information; information modality can include at least one of the following: text (or called text modality), image (or called image modality), voice (or called voice modality), video (or called video modality), animation (or called animation modality), audio (or called audio modality) and file (or called file modality).

[0090] (3) Query results that match the query information refer to search results that are generated or searched by a search engine and are relevant to the search intent represented by the query information; such search results can provide a response to the search intent. For example, if the search intent represented by the query information is to inquire about the causes of dry and sore throats, the search results that are relevant to the search intent may be that viral colds are the cause of dry and sore throats.

[0091] 2. Artificial intelligence.

[0092] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and create new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning.

[0093] The search engine provided in the embodiments of the present application relies on a variety of basic technologies and software technologies in the field of artificial intelligence to achieve information search, including but not limited to natural language processing (NLP) and machine learning (ML). Among them:

[0094] (1) Natural Language Processing (NLP) is an important field in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing involves natural language, which is the language people use in daily life, and is closely related to linguistic research; it also involves computer science and mathematics, and is an important technology for model training in the field of artificial intelligence. Among them, the pre-trained model is developed from the Large Language Model in the field of NLP. After fine-tuning, the Large Language Model can be widely used in downstream tasks. Natural language processing technology generally includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0095] (2) Machine learning is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. The emerging technology - pre-training model, is the latest development of deep learning and integrates the above technologies.

[0096] The embodiments of the present application mainly relate to deep neural network models (DNN), convolutional neural networks (CNN), Transformer architectures and attention mechanisms in the field of machine learning; among them: ① The deep neural network model is a neural network with multiple hidden layers that makes predictions or classifications by learning the representation of data; its powerful representation learning ability has enabled it to achieve remarkable results in many fields, such as image recognition, speech recognition and natural language processing; ② Convolutional neural network is a neural network specifically used to process grid data (such as images); it is mainly composed of input layer, convolution layer, pooling layer, fully connected layer and output layer; the convolution layer extracts local features of the input data through convolution operation; the pooling layer downsamples the output of the convolution layer to reduce the dimension of the data; the fully connected layer maps the extracted features to the sample label space; convolutional neural networks have outstanding performance in image recognition, image classification, target detection and other fields. ③The Transformer architecture is a neural network model based on the self-attention mechanism. Its architecture mainly consists of two parts: an encoder and a decoder. Both the encoder and decoder are composed of multiple identical layers stacked together, each containing a self-attention sublayer and a feedforward neural network sublayer. The Transformer uses the self-attention mechanism to capture dependencies in sequences and implements parallel processing, greatly improving the model's training speed and performance. ④The attention mechanism is a neural network mechanism that simulates human attention behavior. It allows the model to focus on different parts of the input data when processing information. To adapt to different attention needs and scenarios, attention mechanisms include self-attention, cross-modal attention, and multi-head attention.

[0097] In actual applications, when searching for information, search engines often encounter problems such as a low degree of matching between query information and query results, which leads to poor retrieval performance and query results that do not meet the user's search needs, i.e., low query result accuracy. A major reason for the low degree of matching between query information and query results is that the query information input to the search engine cannot fully explain or represent the user's search intent. This severely reduces the accuracy and quality of the query results when searching for information based on vague or incomplete search intent. For example, traditional search engines only allow users to input query information in one information modality (such as voice or text) during a search round. Even if multiple information modalities are allowed to be input during multiple search rounds, the query information in multiple information modalities is used independently in each search round. However, single-modal query information cannot fully express the user's search intent to a large extent, so information search based on simulated search intent greatly reduces the accuracy of the query results.

[0098] To improve the comprehensiveness and completeness of search intent representation and the accuracy of information search, the information processing solution provided in the embodiments of this application primarily combines multimodal information fusion and other methods from the query information dimension to form a new information search solution. This information search solution supports the input of multimodal query information composed of multiple modal information combinations into the search engine, thereby generating more accurate query results based on multimodal search.

[0099] Among them, multimodality (or cross-modality) is relative to unimodality. Unimodality refers to the use of query information of only one information modality for information search. For example, a search engine that relies only on text for information search is a unimodal retrieval system. Multimodality refers to the simultaneous use of sub-query information of multiple information modalities for information processing or interaction to achieve information search; for example, a search engine that uses multimodal query information that combines text, images, and videos as input to perform information search is a multimodal retrieval system. This multimodal retrieval system can integrate sub-query information from different information modalities, thereby utilizing the diverse information channels provided by sub-query information of different information modalities to achieve complementarity in the recognition process of search intent, which helps to understand and perceive information more comprehensively, enhance the recognition accuracy of search intent, and thus improve the search accuracy of query results.

[0100] In a specific implementation, the general process flow of the information processing solution provided by the embodiment of the present application may include: an information search portal is displayed on a computer device, and the information search portal is used to trigger the execution of an information search, that is, the information search portal is the portal for starting or entering a search engine. If a user has a need for information search, the user can perform a trigger operation on the information search portal. At this time, the computer device responds to the trigger operation on the information search portal and receives multimodal query information input by the user to the search engine; the multimodal query information is combined modal information input by the user containing sub-query information of at least two information modalities, such as the multimodal query information including sub-query information 1 belonging to the text modality and sub-query information 2 belonging to the image modality. In this way, the computer device will perform cross-modal feature fusion based on the multiple sub-query information of the multiple information modalities contained in the multimodal query information, so as to perform information search based on the intention features after the cross-modal feature fusion, and obtain and output query results that match the multimodal query information.

[0101] It has been found in practice that the information processing scheme provided by the embodiment of the present application has obvious advantages when implementing information search. The following is an example of comparing the scheme of the present application with the existing single-modal search engine to implement information search, to illustrate the advantages of the embodiment of the present application: Traditional search engines have shortcomings such as information modality fragmentation when performing information search, specifically only supporting the input of query information of a single information modality (such as only text, only voice or only image), or although supporting query information of multiple information modalities, each works independently (that is, query information of one information modality is used to implement one round of information search), lacking an effective multimodal fusion mechanism, and it is difficult to fully utilize the complementarity of sub-query information of different information modalities. However, the embodiment of the present application supports users to combine sub-query information of multiple information modalities to construct multimodal query information for a round of information search process, and realizes the use of the complementarity of different information modalities to significantly improve the correctness of search intent recognition, so that information search based on the correct search intent can ensure the search accuracy of the query results and improve the quality of information search.

[0102] The information processing solution provided in the embodiments of the present application is a multimodal information-based automatic information search solution. It allows for receiving multimodal query information in a multimodal combination, fully understanding the search intent to search for highly accurate query information, and solving the user's questions. This makes the information processing solution provided in the embodiments of the present application applicable to a variety of information search scenarios that require the use of search engines (specifically, the multimodal search engines provided in the embodiments of the present application) to perform information search. Information search scenarios may include, but are not limited to: ① Customer support: Multimodal search engines can be used as customer support tools to answer common user queries, reduce the workload of customer service staff, and improve customer satisfaction. ② Internal enterprise knowledge base: Enterprises can use multimodal search engines to build internal knowledge bases to help employees quickly find the information they need and improve work efficiency. ③ AI assistant: Multimodal search engines can serve as AI assistants for individuals or enterprises, assisting users with multi-round interactive conversations between users and AI assistants, providing functions such as multi-round conversations. ④ Online education: Multimodal search engines can be applied in the field of online education to provide students with personalized learning resources and real-time question-answering services. For example, providing an information search portal in the learning service interface facilitates students to quickly retrieve knowledge through the information search portal during the learning process. ⑤ E-commerce: Multimodal search engines can help users answer questions during the shopping process, provide shopping recommendations, and enhance the shopping experience. ⑥ Financial Services: Multimodal search engines can provide real-time consulting services to customers of financial institutions such as banks and insurance companies, answering questions about accounts, transactions, and products. ⑦ Medical Consultation: Multimodal search engines can provide patients with basic medical consultation services, answering questions about diseases, treatments, and medications. ⑧ Travel Consultation: Multimodal search engines can provide tourists with real-time travel information, answering questions about attractions, hotels, transportation, and more. ⑨ News and Information Retrieval: Multimodal search engines can help users quickly find the news and information they need, improving information retrieval efficiency. ⑩ Medical Search: Multimodal search engines can help users accurately identify diseases using multimodal search dimensions, helping users quickly locate the disease. For example, when a user has a medical need, they can use voice to ask questions such as "My blood pressure is a little high, what medicine do you recommend?" and upload an image of a blood pressure monitor. Based on semantic signals and the image of the blood pressure monitor, the multimodal search engine can provide professional advice and appropriate medication recommendations, allowing users to place an order with one click without complex operations. For example, visually impaired users can describe their needs through voice, and users with physical disabilities can control the entire process through voice, reducing operational barriers.

[0103] Furthermore, the information processing solution provided in the embodiments of the present application can be applied to any search engine that can implement information search (specifically, a multimodal search engine). The product form of a search engine may include but is not limited to applications and plug-ins. Among them, a plug-in is a component that can extend the functionality of software or add new features; a plug-in can be in the form of a computer program, or a component (such as a binary file, configuration file, script, or other type of data package). The plug-in is designed to be compatible with the application (i.e., application) or operating system, and can be loaded and executed at runtime to implement information search.

[0104] An application refers to a computer program that is used to complete one or more specific tasks. Applications are classified according to their application functions. Applications that can integrate the multimodal combined search provided by the embodiment of the present application may include but are not limited to: ① Video application refers to an application that can search and view videos or film and television dramas; through the input of multimodal query information, such as the combined input of stills and text lines of film and television dramas, the search accuracy of the video application for film and television dramas is improved. ② Content application refers to an application that can retrieve massive content on the Internet, and through the input of multimodal query information, it helps users retrieve query results that match the search intent from massive content. ③ Social application refers to an application that realizes instant messaging and social interaction based on the Internet; providing the multimodal search engine provided by the embodiment of the present application in the social application can not only realize the retrieval of massive information on the Internet during the social process without the need for cross-application information search, but also can perform multimodal rapid retrieval of social conversation content to help users quickly locate historical conversations. Applications are classified according to their operating modes. Applications that can integrate the multimodal combined search provided in the embodiments of the present application may include but are not limited to: ① a client that needs to download an installation package and install and run the installation package before it can be run in the terminal; or ② a mini-program that can be used without downloading and installing, which is usually a sub-program of the client; ③ a web (World Wide Web) application opened through a browser, etc.

[0105] The following description will be made by taking the integration of the search engine provided in the embodiment of the present application into a social application (such as a social client) as an example, which will be specifically explained here.

[0106] It should be understood that the above description is merely an exemplary product performance and information search scenario provided by the embodiments of this application, and does not limit the product performance and information search scenario of the information processing solution provided by the embodiments of this application. The multimodal search engine provided by the embodiments of this application can provide efficient, accurate, and convenient information search services in various information search scenarios, demonstrating high value and practicality in various information search scenarios, and helping to improve user experience and satisfaction.

[0107] To facilitate understanding of the information processing solution provided in the embodiment of this application, Figure 1 The scene diagram shown in FIG. 1 briefly introduces the information search scene involved in the embodiment of the present application; Figure 1 As shown, the system includes an object 101 , a terminal 102 and a server 103 . The embodiment of the present application does not limit the number and naming of the object 101 , the terminal 102 and the server 103 .

[0108] Wherein, object 101 refers to a user with information search needs. Terminal 102 is a terminal device held by object 101 for realizing information retrieval; a search engine integrating the information processing solution provided by the embodiment of the present application is deployed in the terminal device, and object 101 starts the search engine in terminal 102 to search for information. Wherein, terminal 102 may include but is not limited to: smart phones (such as smart phones deploying Android system, or smart phones deploying Internetworking Operating System (IOS)), tablet computers, portable personal computers, mobile Internet devices (Mobile Internet Devices, MID), vehicle-mounted devices, head-mounted devices, intelligent chat robots and aircraft and other terminal devices. The embodiment of the present application does not limit the type of terminal devices, which is explained here.

[0109] Server 103 is the server corresponding to terminal 102, and is used to interact with terminal 102 for data exchange to provide computing and application service support for terminal 102. Server 103 is specifically a backend service corresponding to the search engine deployed in terminal 102, providing technical support and search services for the search engine through terminal 102. Server 103 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0110] Among them, the various devices included in the information processing system can be directly or indirectly connected to each other through limited or wireless means. The embodiments of the present application do not limit the communication method between devices; for example, the communication method between devices may include but is not limited to: using HTTP requests, remote procedure calls (Remote Procedure Call), sockets, and shared content.

[0111] The information processing solution provided in the embodiment of the present application can be executed by a computer device, which may include Figure 1 The terminal 102 and server 103 in the information processing system shown, or including one of the terminal 102 and server 103; the embodiment of the present application does not limit the execution subject of the information processing solution. Figure 1 The information processing system shown briefly introduces the program process.

[0112] In a specific implementation, when the object 101 has a need to perform an information search, the object 101 opens or starts a search engine in the terminal 102, such as entering the search engine by triggering the information search entrance of the search engine displayed on the display screen of the terminal 102. Then, the object 101 enters multimodal query information in the search engine, and the multimodal query information contains sub-query information of at least two information modalities, such as an image (such as at least one picture) of an image modality and a text (such as a sentence) of a text modality. The terminal 102 generates an information query request based on the multimodal query information, and the information query request carries the multimodal query information, and the information query request is used to request the server 103 to perform an information search based on the multimodal query information and return a query result. After receiving the information query request sent by the terminal 102, the server 103 generates a query result that matches the multimodal query information through a series of data analysis in response to the information query request. Among them, data analysis includes but is not limited to: feature analysis of sub-query information of different information modalities and cross-modal feature fusion; intent analysis of cross-modal fusion features after cross-modal fusion, and knowledge retrieval and item search based on the analyzed query intent (i.e., search intent) to generate query results; the specific implementation content of data analysis is not elaborated in detail in this embodiment of the present application, and will be described in detail in subsequent specific embodiments. Server 103 returns the query results that match the multimodal query information to terminal 102, so that terminal 102 outputs the query results on the display screen, making it convenient for subject 101 to view the answer in a timely manner.

[0113] Based on the above brief introduction to the information processing solution and information processing system provided in the embodiments of the present application, the following points should be explained:

[0114] ① The above-mentioned embodiments of this application Figure 1 The system architecture shown is for the purpose of more clearly illustrating the technical solutions of the embodiments of the present application and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. It is known to those skilled in the art that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. In other words, Figure 1The architectural diagram of the information processing system shown is an exemplary architectural diagram; in actual applications, the number and types of devices included in the information processing system may change, and the embodiments of the present application do not limit the architectural diagram of the information processing system.

[0115] For example, the embodiments of the present application support the construction of vertical professional knowledge graphs for at least one field, such as constructing professional knowledge graphs (Knowledge Graph) covering fields such as medical care, childcare, and health. The knowledge graph can cover multi-dimensional relationships such as disease-symptoms-treatment-drugs; in this way, the search engine can achieve more complex reasoning and professional knowledge retrieval when combining the knowledge graph for information search. Among them, the knowledge graph is a structured knowledge base that represents various knowledge in the form of entities and relationships, supports intelligent reasoning and answer generation; or, the knowledge graph can be simply understood as a semantic network used to describe the relationship between entities, mainly used to describe entities, attributes, and the relationship between entities; its core idea is to convert a large amount of information into a graph, where the nodes in the graph represent entities and the edges represent the relationship between entities. In this case, Figure 1 The information processing system shown also includes a database for storing the knowledge graph, which may be another server independent of the server 103, or a storage space in the deployment server 103.

[0116] ② The collection and processing of relevant data in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations. The acquisition of personal information must be subject to the knowledge or consent of the individual subject (or the legal basis for obtaining the information), and subsequent data use and processing must be carried out within the scope of authorization of laws and regulations and the subject of personal information. For example, when the embodiments of this application are applied to specific products or technologies, such as when the terminal 102 sends the user's multimodal query information to the server 103, the user's permission or consent is required, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant region.

[0117] Based on the information processing scheme described above, the embodiment of the present application proposes a more detailed information processing method. The information processing method proposed in the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0118] See Figure 2 , Figure 2 A flowchart of an information processing method provided by an exemplary embodiment of the present application is shown. Figure 2 The information processing method shown can be executed by a computer device (such as a terminal and / or server that deploys the multimodal search engine provided by this application), and the information processing method includes but is not limited to steps S201-S203:

[0119] S201: Displaying an information search entry.

[0120] Among them, the information search entrance is the entrance used to trigger the execution / conduct of information search, specifically the entrance used to trigger the start of the search engine to conduct information search. Depending on the different information search scenarios or the different product forms of the search engine, the form of the information search entrance used to trigger the entry into the search engine varies. Exemplarily, the information search entrance includes at least one of the following: a search bar in the content search interface; a dialog box in the interactive dialogue interface; and a search pop-up window provided by the search engine and displayed in the service interface. For ease of understanding, schematic diagrams of the information search entrances of the above three exemplary search engines are given below in conjunction with the accompanying drawings; wherein:

[0121] (1) The information search entry is the search bar in the content search interface.

[0122] In other words, the search bar in the content search interface integrates an information search portal, or the information search portal is expressed as a search bar. The content search interface where the search bar resides is provided by a search engine, or by an application that has a search engine embedded within it. This embodiment of the application does not limit the provider of the content search interface, its interface style, or its content.

[0123] For example, assuming that the content search interface is an interface provided by a search engine to implement information search, if the search engine is a search browser with information search capabilities, the content search interface provided by the search engine may be the main page of the search browser. Figure 3a As shown, if the user has an information search need, the user can open a search browser in the terminal, which triggers the display of the content search interface 301 of the search browser - the main page. The content search interface 301 includes a search bar 302, which is used to receive multimodal query information input by the user and start information search.

[0124] For another example, assume that the content search interface is an application interface provided by a social application, such as a conversation list interface provided by the social application, and the conversation list interface includes a search bar for entering a search engine embedded in the social application. Figure 3b As shown, if a user needs to search for information during a social conversation in a social application, the user can quickly call up or enter a search engine and search for information through the search bar 302 in the conversation list interface 303 provided by the social application, without having to jump across applications or interfaces. This improves the convenience of information search during the social interaction process. It should be noted that the content search interface can be other interfaces besides the contact interface provided by the social application, such as the address book interface 304 (or called the contact interface), and this embodiment of the application is not limited to this.

[0125] (2) The information search entry is a dialog box in the interactive dialogue interface.

[0126] That is to say, the dialog box in the interactive dialogue interface is integrated with an information search portal, or it can be expressed as an information search portal in the form of a dialog box. Among them, the interactive dialogue interface is provided by a search engine with interactive dialogue capabilities, and is an interface for sending and receiving information between users and intelligent assistants; the dialog box in the interactive dialogue interface is an input box for users to input multimodal query information. Interactive dialogue, also known as intelligent question answering (QA), open-ended question answering or intelligent dialogue, belongs to the field of human-computer interaction and is an advanced form of information retrieval. The embodiment of the present application introduces multimodal search in the interactive dialogue, which greatly enriches the richness of the information modality input by the user during the interactive dialogue process and improves the dialogue quality of the interactive dialogue.

[0127] For example, the information search portal is a dialog box in an interactive dialogue interface provided by a search engine. Figure 4 .like Figure 4 As shown, the interactive dialogue interface 401 includes a dialogue display area 402 and a dialog box 403; the dialogue display area 402 is used to display the query information input by the user (such as query information 1) and the query result (such as query result 1) output by the search engine through the intelligent assistant for the query information. The query information can be multimodal query information or unimodal query information (as opposed to multimodal query information, which only contains query information of one information modality), and the dialog box 403 is used for the user to input new query information (such as query information 2 is multimodal query information).

[0128] (3) The information search entry is the search pop-up window in the service interface.

[0129] That is, the search pop-up window in the service interface integrates the information search portal, or the information search portal is expressed as a search pop-up window; the display layer of the search pop-up window is higher than the display layer of the service interface and is displayed above the service interface. The service interface can be any service interface in the search engine (such as Figure 3aThe service interface may also be any application interface provided by other applications independent of the search engine, and the other applications and the search engine are deployed simultaneously in the computer device; for example, the other application is a document application, and the service interface is a document interface of any document provided by the document application; for another example, the other application is a video application, and the service interface is any video playback interface provided by the video application, etc. This method of providing users with an information search entry through a search pop-up window on the service interface greatly improves the flexibility and convenience of users in searching for information when using computer devices. Information search is no longer limited to being performed within the search engine or relying on a search engine embedded in an application, but the search engine can be called up in any service interface to search for information.

[0130] Taking the document interface of the target document provided by the document application that is independent of the search engine as an example, the schematic diagram of displaying the search pop-up window on the document interface can be seen in Figure 5 .like Figure 5 As shown, assume that a document application and a search engine are both deployed on a smartphone, and a user opens a document interface 501 of a target document through the document application; the target document's content is displayed in the document interface 501. Optionally, considering that the user may need to search for information on all or part of the document content while browsing the target document, a search pop-up window 502 is supported to remain displayed on the document interface during the display of the document interface. Optionally, when a user selects specific content 503 in the document interface 501 and predicts that the user may need to search for information on the specific content 503, a search window 502 is automatically displayed on the document interface 501. In this way, when the user needs to search for information, they can enter multimodal query information for the specific content 503 in the pop-up search window 502, or the specific content 503 is displayed as part of the multimodal query information by default, and the search window 502 is displayed, allowing the user to continue editing the multimodal query information.

[0131] It should be noted that the above Figure 3a 、 Figure 3b 、 Figure 4 and Figure 5 These are several exemplary forms of information search portals provided in the embodiments of this application and do not limit the embodiments of this application. For ease of explanation, the following description uses the information search portal as a search bar in a content search interface, specifically a search bar included in a conversation list interface provided by a social application, as an example to introduce the information processing method based on multimodal search provided in the embodiments of this application.

[0132] S202: In response to a triggering operation on an information search portal, receiving input multimodal query information.

[0133] The multimodal query information entered by the user through the triggering information search portal includes sub-query information of at least two information modalities; in other words, the multimodal query information includes at least two sub-queries, and at least two of the at least two sub-queries belong to different information modalities. For example, the multimodal query information includes sub-query information 1, sub-query information 2, and sub-query information 3; wherein sub-query information 1 belongs to the information modality of image, sub-query information 2 belongs to the information modality of text, and sub-query information 3 belongs to the information modality of text; alternatively, sub-query information 2 and sub-query information 3 can be combined into sub-query information 4, in which case the multimodal query information includes sub-query information 1 belonging to the image modality and sub-query information 4 belonging to the text modality.

[0134] In a specific implementation, when a user performs a trigger operation on an information search portal, it indicates that the user has a need to perform an information search. At this time, the computer device responds to the user's trigger operation on the information search portal and displays an information search interface. The information search interface is an interface provided by the search engine for receiving multimodal query information input by the user; specifically, the information search interface includes an information input area, and the user performs an information input operation in the information input area. The computer device displays the input multimodal query information in the information input area according to the information input operation performed by the user in the information input area.

[0135] Among them, the embodiment of the present application provides users with two ways of inputting multimodal query information through the information search interface, including: the user actively inputs at least two sub-query information belonging to different information modalities in the information search interface to form multimodal query information; or the user inputs information according to the prompt output by the search engine, and inputs at least two sub-query information belonging to different information modalities multiple times to form multimodal query information. The embodiment of the present application provides active and passive (i.e., inputting information only after inputting information according to the prompt output by the search engine) multimodal query information input methods. When the user knows that the search engine supports multimodal search, the user can actively perform multimodal information search. When the user is unaware that the search engine has multimodal search capabilities, the user can be prompted to perform multimodal search. This greatly expands the ways in which users can provide sub-query information of multiple information modalities during the information search process, thereby improving the quality of information search.

[0136] The following is a detailed description of the specific implementation process of the two multimodal query information input methods proposed above;

[0137] 1. In the information search interface, the user actively inputs at least two sub-query information belonging to different information modalities to form multimodal query information.

[0138] Specifically, if the user can perceive that the information search interface displayed by triggering the information search entry supports multimodal search, then when the user wants to conduct an information search based on multimodality, the user can actively enter multimodal query information in the information search interface to conduct an information search. In an embodiment of the present application, a multimodal input prompt is displayed in the information search area to let the user know that the information search interface supports multimodal search; the multimodal input prompt is used to prompt the user to allow the user to enter multimodal query information in the information search interface to conduct an information search.

[0139] The embodiments of the present application do not limit the presentation form of the multimodal input prompt in the information search interface; under different presentation forms of the multimodal input prompt, the user may input multimodal query information in different ways in the information search interface. The following are several exemplary processes for users to input multimodal query information provided by the embodiments of the present application; among them:

[0140] (1) The multimodal input prompt is displayed in the form of a character string in the information search interface. In this case, after the user sees the multimodal input prompt in the information input area, he or she can perceive that the information search interface allows the input of multimodal query information, and the user can directly perform the information input operation in the information input area; and the multiple sub-query information entered by the information input operation is displayed in the information input area, so that the user can confirm the sub-query information included in the multimodal query information before publishing the multimodal query information, thereby improving the input accuracy of the multimodal query information.

[0141] For example, when the multimodal input prompt is displayed in the information search area in the form of a string, the user enters the multimodal query information in the information input area. Figure 6a .like Figure 6aAs shown, when information search entry 601 is triggered, information search interface 602 is displayed. This information search interface 602 includes an information input area 603. A multimodal input prompt 604 is displayed in the form of a string. For example, multimodal input prompt 604 reads "Enter text, voice, and image to search." If a user desires a multimodal search, they can perform an information input operation in information input area 603. In this case, the information input operation includes entering at least two sub-queries belonging to different information modalities in information input area 603. In detail, the information input area 603 includes input identifiers corresponding to different information modalities, such as an input identifier 6031 corresponding to an image modality and an input identifier 6032 corresponding to a voice modality; when the user directly edits a character string through the virtual keyboard in the information input area 603, the sub-query information of the text modality can be input; then, the user performs a trigger operation on the input identifier 6032, and the user can input an image (such as real-time shooting, pulling from a storage space, or downloading from the Internet, etc.) as the sub-query information; then, the user performs a trigger operation on the input identifier 6032, and the user can use the input voice as the sub-query information by inputting a voice signal. To facilitate the user to confirm the voice content, the embodiment of the present application supports directly converting the voice into text and displaying it in the information input area; based on the above-mentioned multiple information input operations, in response to the user's confirmation operation on the multiple sub-query information entered in the information input area, the multiple sub-query information is constituted into multi-modal query information, and the search engine starts to search for information based on the multi-modal query information. It should be understood that in addition to inputting sub-query information under the corresponding information mode by triggering the input identifier corresponding to the above-mentioned information mode, the user can also directly input sub-query information of different information modes by pasting in the information input area.

[0142] Furthermore, to facilitate users entering sub-query information belonging to different information modalities in the information input area to form multimodal query information, embodiments of the present application also support setting up multiple sub-input areas in the information input area, each sub-input area corresponding to a different information modality. In this way, when a user needs to enter multimodal query information, they can perform information input operations in the sub-input area corresponding to the corresponding information modality. By setting up sub-input areas corresponding to different information modalities, it can, to a certain extent, encourage users to choose to enter multimodal query information, thereby improving the quality of information search.

[0143] In a specific implementation, a computer device displays sub-query information belonging to a corresponding information modality in a corresponding sub-input area in response to an information input operation performed by a user in at least two of the multiple sub-input areas; when the user confirms the input of the multiple sub-query information as a source of information search questions, the computer device combines the sub-query information in the at least two sub-input areas into a multi-modal query information based on the user's confirmation operation on the sub-query information displayed in the at least two sub-input areas. For example, a schematic diagram of an information input area including multiple sub-input areas can be seen in Figure 6b .like Figure 6b As shown, the information input area includes a sub-input area 604 corresponding to the text modality, a sub-input area 605 corresponding to the voice modality, and a sub-input area 606 corresponding to the image modality. The user can enter sub-query information of the corresponding information modality in all or some of the sub-input areas corresponding to the information modality according to their input needs. For example, the user first performs an information input operation in the sub-input area 604 corresponding to the text modality, at which point sub-query information 1 belonging to the text modality is displayed in sub-input area 604. The user then enters sub-query information 2 belonging to the image modality in sub-input area 606 corresponding to the image modality. If the user decides to conduct an information search based on sub-query information 1 and sub-query information 2 in the multimodal manner, the user can select a confirmation option 607 in the information search interface, indicating that sub-query information 1 and sub-query information 2 are combined into multimodal query information. The computer device then conducts an information search based on this multimodal query information.

[0144] (2) The multimodal input prompt is presented as a modal combination option. That is, the information search interface includes a modal combination option that is allowed to be triggered, and the purpose of the modal combination option is to inform the user that a multimodal information search is allowed in the information search interface. In this case, if the user wants to conduct a multimodal information search, the user can select the modal combination option in the information search interface; in response to the user's selection of the modal combination option, the computer device will display multiple sub-input areas in the information input area, and the user can select the sub-input areas in the information input area according to the selection operation. Figure 6b In the process shown, an information input operation is performed in at least two of multiple sub-input areas; the computer device displays sub-query information belonging to the corresponding information modality in the corresponding sub-input area based on the information input operation performed by the user in at least two sub-input areas.

[0145] For example, a schematic diagram showing modal combination options in the information search interface can be found in Figure 7 ;like Figure 7As shown, a modal combination option 701 is displayed in the information search interface (specifically, the information input area in the information search interface), and an input identifier corresponding to a single information mode can also be displayed adjacent to the modal combination option 701 (such as Figure 6a When the user selects the modality combination option 701, indicating that the user wants to input multiple sub-query information, the computer device displays multiple sub-input areas in the information input area; in this way, the user selects the sub-input area according to the input method. Figure 6b The described information input operation is performed in at least two sub-input areas, and when a selection operation is performed on a confirmation option in the information search interface, it is determined that the sub-query information displayed by each sub-query information on which the information input operation is performed constitutes multimodal query information.

[0146] Furthermore, if the user wants to exit the multimodal information search (e.g., want to use the traditional single-modal information search or exit the search engine directly) after selecting the modal combination option, the user can cancel the selection of the modal combination option (e.g., Figure 7 As shown, the user performs a selection operation on the modal combination option that is in a selected state, switching the modal combination option from a selected state to an unselected state), and at this time the computer device responds to the unselect operation on the modal combination option, and deletes multiple sub-input areas in the information input area, indicating exit from the multimodal information search.

[0147] It is worth noting that, considering the limited display area of ​​the display screen of a computer device, it is difficult for the computer device to display the sub-input areas corresponding to all information modes in the information input area; to save interface space, the embodiment of the present application supports the computer device to display the sub-input areas corresponding to commonly used information modes in the information input area by default, such as the sub-input area corresponding to the text mode, the sub-input area corresponding to the voice mode, and the sub-input area corresponding to the image mode by default. In this case, in order to enrich the richness of the information mode to which the sub-query information input by the user belongs, the embodiment of the present application supports the user to create and customize new sub-input areas for inputting sub-query information under other information modes, helping the user to input personalized sub-query information of the information mode and improving the user's information input experience.

[0148] An exemplary diagram of a user-created sub-input area can be found in Figure 8 ;like Figure 8As shown, the information input area includes, in addition to the sub-input areas provided by the computer device (specifically, a search engine), an add area option 801. When a user triggers this add area option 801, the computer device displays a modality selection window 802 based on the user's triggering operation. Modality selection window 802 includes a region creation option corresponding to at least one information modality, such as region creation option 8021 for video modality, region creation option 8022 for audio modality, and so on. The user creates a new sub-input area by selecting a region creation option in modality selection window 802 according to their region creation needs. Specifically, in response to selecting the target region creation option in modality selection window 802, the computer device displays a target sub-input area 803 corresponding to the target information modality in the information input area. The target information modality is the information modality corresponding to the target region creation option, and the target region creation option is any region creation option in modality selection window 802, such as region creation option 8021. Based on this, the user can input sub-query information belonging to the video modality in the target sub-input area 803 displayed in the information input area to construct multi-modal query information.

[0149] 2. According to the prompt input output by the search engine, the user inputs at least two sub-query information belonging to different information modalities multiple times to form multimodal query information.

[0150] In a specific implementation, when a user allows multimodal information search in an unknown information search interface, or when a user allows multimodal information search in a known information search interface but wants to gradually enrich the query information, the user can first perform an information input operation in the information input area of ​​the information search interface. At this time, the computer device responds to the information input operation performed by the user in the information input area and receives the first sub-query information input. The computer device will first perform intent recognition based on the first sub-query information to identify the user's initial query intent for this information search; if the computer device detects that the initial query intent obtained based on the analysis of the first sub-query information does not meet the intent recognition conditions (such as the number of slots extracted during intent recognition is less than the number threshold, or the intent classification result is less than the similarity threshold, etc.), that is, the initial query intent is relatively vague, then the query results generated based on the initial query intent are often of poor quality. Therefore, the computer device can display prompt input information, which is used to prompt the user to enter a second sub-query information that represents the same query intent as the first sub-query information. After the user views the prompt input information, he or she may perform the information input operation again. At this time, the computer device receives the second sub-query information input in response to the information input operation performed again by the user. In this way, the computer device may simultaneously perform operations such as intent recognition based on the second sub-query information and the first sub-query information (the two constitute multi-modal query information) to realize information search. Compared with information search based only on single-modal information (i.e., the first sub-query information), the accuracy of identifying the query intent is improved, thereby improving the quality of information search.

[0151] Furthermore, if the user wishes to search for information using only the first sub-query information, embodiments of the present application also support the user's ability to independently close the input prompt information, allowing the search engine to continue searching for information based on the first sub-query information. Specifically, when the input prompt information is displayed on the computer device, the user can close the input prompt information; at this time, the computer device displays a notification message based on the user's closing operation on the input prompt information; the notification message is used to inform the user that the search engine will continue searching for information based on the first sub-query information.

[0152] A schematic diagram of a user inputting the second sub-query information when the input prompt information is displayed can be seen in Figure 9a ;like Figure 9aAs shown, the information input area includes an input identifier corresponding to the voice modality. When the input identifier of the voice modality is triggered, the user can input a voice signal belonging to the voice modality, and the computer device (specifically, a search engine) converts the voice signal into text. If the computer device performs an initial intent analysis based on the voice signal, it obtains the initial query intent represented by the voice signal. If the initial query intent does not meet the intent recognition conditions, a prompt input message 901 is output. The input prompt message 901 is used to prompt that the second sub-query information to be input belongs to the first information modality. The prompt input message 901 can be expressed as the input method corresponding to the first information modality recommended by the search engine for the user to input; for example, if the first information modality is the image modality, the prompt input message 901 is expressed as an image input interface for image input. On the one hand, if the user wants to continue to input the second sub-query information, the user can perform an information input operation for the second sub-query information, such as inputting the second sub-query information - image - in real time in the image input interface, obtaining it from local storage space, or downloading it from the Internet. On the other hand, if the user does not want to input the second sub-query information, the user can perform a close operation on the prompt input information 901; specifically, the prompt input information 901 is associated with a close option 902, such as a close component provided in the image input interface; the user can close the input prompt information 901 and exit the multimodal information search by triggering the close option 902, and the search engine automatically performs information search based on the first sub-query information.

[0153] Furthermore, in order to enhance the user's autonomous selectivity of the information modality to which the second sub-query information to be re-entered belongs, the embodiment of the present application supports user customization of the information modality to which the second sub-query information belongs, so that the user can personalize the supplementary input of the second sub-query information according to his or her own modal input requirements, thereby enhancing the flexibility of the setting of the information modality to which the second sub-query information belongs and improving the user's search experience. In a specific implementation, assuming that the prompt input information output by the search engine is used to prompt that the second sub-query information to be entered belongs to the first information modality, if the user does not want to enter the first information modality, but wants to enter a second information modality different from the first information modality, the user performs a modal switching operation on the first information modality; at this time, the computer device displays the switched second information modality according to the user's modal switching operation on the first information modality, specifically displays the information input interface corresponding to the second information modality, so that the user can perceive the second sub-query information input belonging to the second information modality according to the information input interface.

[0154] For example, a schematic diagram of a user performing a mode switching operation can be found in Figure 9b ;like Figure 9bAs shown, when input prompt information 901 is displayed and input prompt information 901 indicates that the second sub-query information to be input is in image mode, if the user wishes to switch the second sub-query information to be input to video mode, the user can perform a mode switching operation; the mode switching operation includes, but is not limited to, a sliding operation in a specified direction (e.g., a sliding operation to the left or right), a triggering operation on a switching control 903, or a fixed gesture operation (e.g., a sliding operation drawing an S-shaped trajectory). At this time, an information input interface 904 for receiving the second sub-query information in video mode is output; the user enters a video as the second sub-query information in information input interface 904.

[0155] In summary, the above describes the specific implementation process of the embodiment of the present application for the user to actively or passively input multimodal query information. Based on the above description, the following points need to be explained:

[0156] ① The above description is based on an example in which the multimodal query information input by the user contains sub-query information of two information modalities. However, in actual applications, the number of sub-query information contained in the multimodal query information input by the user may be greater than two, and the types of information modalities to which at least two sub-query information belong may be at least two. The embodiments of the present application do not limit the number of sub-query information contained in the multimodal query information and the types of information modalities to which they belong.

[0157] ② The information input operation performed by the user is related to the information mode to which the input sub-query information belongs; that is, the specific operation process of the information input operation is adapted to the information mode to which the sub-query information belongs. For example, if the sub-query information belongs to the text mode, the information input operation includes an editing operation of inputting a character string through a virtual keyboard or a physical keyboard; for another example, if the sub-query information belongs to the image mode, the information input operation includes a reading operation of reading an image from a local storage space, or a shooting operation of calling a camera to shoot an image, etc. The embodiments of the present application do not limit the specific operation process of the information input operation.

[0158] ③ The embodiment of the present application supports post-operation on the sub-query information after the sub-query information is input, so as to enhance the semantic expression ability of each sub-query information contained in the multimodal query information, thereby promoting the accuracy of intent recognition and improving the quality of information search. Among them, when performing a post-operation on a sub-query information, the sub-query information can be post-operated according to the semantics represented by other sub-query information to enhance the semantic complementarity between the sub-query information, such as enhancing the text semantics expressed by sub-query information 2 - text in sub-query information 1 - image. The following takes the case where the sub-query information belongs to the image modality as an example to exemplify the process of post-operation on the image; Figure 10As shown, in response to a user input operation, specifically an image input operation, the computer device displays the input image. The user can then perform information annotation operations on the image, including but not limited to editing annotation information on the image using editing tools. Based on the user's information annotation operation, the computer device displays the annotated information at the annotated location on the image. For example, if the user edits annotation information 1002 at the interface element candle 1001 on the image, this annotation information 1002 can assist the search engine in focusing on annotation information 1002 and the associated interface elements when processing the image. After the user completes the information annotation operation on the image, they need to confirm the annotated image. In response to this confirmation operation, the computer device adds the image and the annotated image as two sub-query information included in the multimodal query information. In other words, the annotated image is added to the multimodal query information as sub-query information, distinguishing it from unannotated images. This allows the search engine to conduct information searches based on the richer sub-query information, improving the accuracy of information searches.

[0159] based on Figure 10 As described above, it is worth noting that the specific operation mode of the post-operation is adapted to the information mode to which the sub-query information belongs. The embodiment of the present application does not limit the specific operation mode of the post-operation. In addition, the timing for the user to perform the post-operation on the sub-query information can be to perform the post-operation on the sub-query information after entering the sub-query information (such as Figure 10 ), or perform post-operation on a single sub-query information after entering all sub-query information.

[0160] S203: Displaying query results obtained by searching for information based on the multimodal query information.

[0161] Based on the aforementioned steps S201-S202, after the user inputs the multimodal query information into the search engine, the search engine generates query results that match the multimodal query information based on the multimodal query information. The computer device then outputs the query results obtained from the information search based on the multimodal query information to the user, so that the user can obtain the search content in a timely manner.

[0162] In summary, the embodiments of the present application support multimodal query information input by users, and the multimodal query information contains at least two sub-query information belonging to different information modalities, such as text belonging to text modality and image belonging to image modality. All sub-query information belonging to different information modalities contained in the multimodal query information is used to describe the same search intention. Through multimodal fusion technology, the collaborative work of voice, image, and text is realized, and the search accuracy and the ability to understand user intentions are improved. It is particularly suitable for scenarios that express complex needs. Compared with information search based on sub-query information of a single information modality, the semantic incompatibility between multiple sub-query information of different information modalities is utilized to help improve and supplement the search intention that the user wants to express; thereby, information search is performed based on multimodal query information with clearer search intentions, and the search and generation accuracy of query results is improved, which greatly improves the quality of information search and improves the user search experience.

[0163] See Figure 11 , Figure 11 A flowchart of another information processing method provided by an exemplary embodiment of the present application is shown. Figure 11 The information processing method shown can be executed by a computer device (such as a terminal and / or server that deploys the multimodal search engine provided by this application), and the information processing method includes but is not limited to steps S1101-S1105:

[0164] S1101: Display information search entry.

[0165] S1102: In response to a triggering operation on an information search portal, receiving input multimodal query information.

[0166] It should be noted that the specific implementation process shown in steps S1101-S1102 is the same as the above Figure 2 The specific implementation process shown in steps S201-S202 in the illustrated embodiment is the same. Please refer to the relevant description of the specific implementation process shown in steps S201-S202, and no further details will be given here.

[0167] In addition, in order to help users dynamically change multimodal query information and quickly conduct a new round of information search, the embodiment of the present application supports users to rewrite the multimodal query information entered in the previous round of information search, so that the search engine can quickly conduct a new information search based on the rewritten multimodal query information, significantly improving the efficiency of information search. In a specific implementation, after the multimodal query information is entered in this round of information search, it supports displaying the modal identifier corresponding to each sub-query information contained in the multimodal query information in the information search interface (specifically the information search area in the information search interface), and any modal identifier is used to indicate the information mode to which the corresponding sub-query information belongs. Then, the user can perform an identifier editing operation on the modal identifier, including but not limited to deleting the target modal identifier (the target modal identifier is the modal identifier corresponding to any sub-query information contained in the multimodal query information), adding a new modal identifier, etc.; in this way, the computer device responds to the identifier editing operation on the modal identifier, displays the new multimodal query information, and calls the search engine to generate a new query result based on the new multimodal query information.

[0168] Take the delete operation with the edit operation as an example, Figure 12 As shown, the information search interface displays the modality identifier corresponding to each sub-query information included in the multimodal query information, such as modality identifier 1201 indicating text modality, modality identifier 1202 indicating voice modality, and modality identifier 1203 indicating image modality. If the user then wishes to search based solely on sub-query information in voice modality and sub-query information in image modality, the user can delete modality identifier 1201 corresponding to the text modality. In this case, modality identifier 1201 becomes the target modality identifier, and the remaining modality identifiers 1202 and 1203 are displayed in the information search interface. The search engine then searches for information based on the remaining sub-query information in the multimodal query information (i.e., the sub-query information indicated by modality identifier 1202 and the sub-query information indicated by modality identifier 1203), generates a new query result, and outputs the new query result.

[0169] S1103: Displaying a result display interface, where the result display interface includes query results obtained by performing information search based on the multimodal query information.

[0170] S1104: In response to a triggering operation on an item purchase entry in the result display interface, purchase detail information corresponding to the recommended item is displayed on the result display interface.

[0171] S1105: Output the payment result according to the payment operation performed in the result display interface.

[0172] In steps S1103-S1105, the search engine performs an information search based on the multimodal query information input by the user. After generating query results, the search results may be displayed in a result display interface. The result display interface and the information search interface may be the same interface, meaning that the query results are displayed directly in the information search interface; alternatively, the result display interface and the information search interface may be separate interfaces. The embodiments of this application do not limit the result display interface.

[0173] In an embodiment of the present application, the search engine can identify the query intent of the user based on multimodal query information and a combination of voice, pictures and other modal information. It can not only provide professional answers to questions based on the identified query intent, but also recommend related items based on the query content, accurately identifying the user's real search needs and answering the user's questions while exploring potential purchase intentions. Based on this, the query results obtained by the search engine for information search based on multimodal query information include not only query explanation information that meets the query intent represented by the multimodal query information, but also recommended items that meet the query intent. Among them, the query explanation information serves as an answer to the user's search intent. For example, if the user's search intent is the cause of a dry and sore throat, then the query explanation information may be viral infection; the recommended items that meet the query intent are commodities that can help users solve problems. For example, in the case where the search intent is the cause of a dry and sore throat, the recommended items that meet the query intent are medicines for treating viral infections.

[0174] It should be noted that the results display interface may include more than just query explanations and recommended items. The interface content within the results display interface is tailored to the type of information search scenario. For example, in a video search scenario, the results display interface may include recommended videos, allowing users to directly view, redirect, download, and cache these videos.

[0175] Furthermore, to help users quickly purchase recommended items, the embodiments of this application support one-click smart payment technology; the so-called one-click smart payment technology refers to a technology that uses artificial intelligence technology to intelligently identify search intent based on object behavior, historical data, etc., automatically outputs recommended items to users, and supports users to conveniently enter smart payment from the result display interface to achieve quick payment. This one-click smart payment technology connects the information search and payment links, achieving a seamless connection from information search to item purchase, avoiding the separation of information search and payment scenarios. Compared with the user needing to go through multiple steps to complete the entire process from information search to item purchase, the one-click smart payment technology greatly simplifies the process of users purchasing items after information search, enriches the information search task, and improves payment convenience.

[0176] The following combination Figure 13 Introduce the process of purchasing recommended items with one-click smart payment in the result display interface; Figure 13 As shown, in response to a user's confirmation of an input multimodal query, the search engine performs an information search based on the multimodal query and displays the generated query results in a results display interface 1301. Result display interface 1301 includes query interpretation information 1302 used to answer the multimodal query and recommended items 1303 based on the user's potential purchase intent. The display area for recommended items 1303 includes an item purchase portal 1304, which allows for quick purchase of recommended items 1303 from the query results without requiring the user to close the search engine and perform multiple steps to purchase the recommended items in a dedicated purchase application. If a user wishes to purchase recommended item 1303, they can trigger item purchase portal 1304. In response to the triggering of item purchase portal 1304, the computer device displays purchase details corresponding to the recommended items in result display interface 1301. This purchase details includes payment information and item information for the recommended items. The user performs a payment operation on the purchase details information corresponding to the recommended item in the result display interface. Specifically, the result display interface is connected to a payment interface corresponding to at least one payment method, such as payment interface 1305 and payment interface 1306. In this way, the user can quickly perform a payment operation on the result display interface through any payment interface, and the computer device processes the payment according to the payment operation and outputs the payment result.

[0177] The aforementioned steps S1101-S1105 mainly introduce the information search and intelligent payment processes included in the information processing method provided in the embodiment of the present application from the interface dimension. The following introduces the process of implementing information search and intelligent payment in the background from the background technology dimension. In the specific implementation, after the search engine obtains the sub-query information contained in the multimodal query information, it can perform feature extraction and fusion processing on each sub-query information in the multimodal query information to generate cross-modal fusion features; then perform contextual semantic prediction processing on the cross-modal fusion features to obtain the query intent represented by the multimodal query information. On the one hand, knowledge retrieval is performed in the knowledge base according to the query intent to generate query explanation information that meets the query intent; on the other hand, recommended items that meet the query intent are searched from the item library according to the query intent.

[0178] The following describes the specific process of the above-mentioned information search and item recommendation in conjunction with the technical architecture provided in the embodiment of the present application. The technical architecture of the information processing method provided in the embodiment of the present application includes a search engine layer, a multimodal fusion layer (or called a multimodal fusion engine), an intent recognition layer, a knowledge retrieval layer, an item recommendation layer, and a payment processing layer. The data flow between each layer can be found in Figure 14 ;like Figure 14 As shown:

[0179] (1) Multimodal input fusion technology: The user inputs multimodal query information into the search engine through the information search interface provided by the search engine layer; the multimodal query information includes sub-query information belonging to at least two different information modalities. The computer device uses the feature extraction rules corresponding to the information modality (i.e., the encoding rules used by the modal input module corresponding to each information modality for feature encoding) to feature encode the sub-query information belonging to the corresponding information modality, and obtain the initial feature information corresponding to each sub-query information; specifically, each sub-query information contained in the multimodal query information is input into the modal input module corresponding to the corresponding information modality for feature encoding, and obtains the initial feature information corresponding to the sub-query information. For example: the sub-query information belonging to the text modality is input into the text input module, and the text input module is used to feature encode the sub-query information belonging to the text modality, aiming to extract the key text feature information of the sub-query information belonging to the text modality; the text input modality has the ability to use AI models such as bag-of-words models and word embedding models to implement feature encoding of the sub-query information belonging to the text modality. For example, subquery information belonging to the speech modality is input into the speech input module, which is used to perform feature encoding on the subquery information belonging to the speech modality, aiming to extract key speech feature information of the subquery information belonging to the speech modality. The speech input module, also known as the speech processing module, primarily uses deep neural network models for speech recognition, supports speech recognition of multiple dialects and accents, and can optimize the speech characteristics of special groups such as children and the elderly to extract more accurate speech features. For another example, subquery information belonging to the image modality is input into the image input module, which is used to perform feature encoding on the subquery information belonging to the image modality, aiming to extract key image feature information of the subquery information belonging to the image modality. The image input module, also known as the image processing module, deploys a visual model that integrates CNN (convolutional neural network) and Transformer architecture, capable of extracting key visual features from various types of medical and daily life images, supporting professional image analysis in medical, educational, and other fields.

[0180] After each modal input module processes the sub-query information under the corresponding information modality, the computer device will input the processing results into the multimodal fusion engine. The multimodal fusion engine adopts a cross-modal attention mechanism to deeply fuse the initial feature information corresponding to each sub-query information, and generates a unified semantic representation through a cross-encoder to achieve information complementarity and enhancement between modalities. Specifically, it includes: ① The multimodal fusion engine maps the initial feature information corresponding to each sub-query information to a cross-modal feature sharing space to obtain the target feature information of each sub-query information in the cross-modal feature sharing space. Among them, the cross-modal feature sharing space is a feature space that can accommodate feature information of multiple modalities, that is, the cross-modal feature sharing space includes feature information of at least two information modalities; by mapping data of different perceptual modalities (such as images, text, voice, video, etc.) into a shared feature space, the feature information of these different modalities can be compared, associated and analyzed in the shared cross-modal feature space. ② Calculate the information correlation between the target feature information corresponding to at least two sub-query information, and perform feature fusion processing based on the information correlation between the target feature information to generate cross-modal fusion features; that is, calculate the similarity of the target feature information of each sub-query information in the cross-modal feature sharing space, aiming to merge semantically similar target feature information and realize the rapid and accurate fusion of feature information of multiple information modalities.

[0181] (2) Intent recognition and knowledge retrieval technology: The computer device fuses the target feature information of each information modality in the cross-modal feature sharing space to form a cross-modal fusion feature. The cross-modal fusion feature is then processed by the intent recognition layer to identify the user's query intent, which can be subdivided into question intent and shopping intent. The intent recognition layer can be understood as an intent classification system. It supports pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers) and combines it with domain-specific fine-tuning to identify the real needs behind the query, such as distinguishing different query intents such as consultation, purchase, and first aid.

[0182] After identifying the query intent, the entity recognition module (not in Figure 14(drawn in) relies on named entity recognition technology to extract key information (such as information related to the query information, such as symptoms, drug names and age, etc.) from cross-modal fusion features, and entity link the extracted key information with the knowledge graph to achieve accurate indexing of relevant knowledge from the knowledge graph based on the key information. Among them, the knowledge graph is a pre-built professional knowledge graph covering fields such as medical care, parenting and health, covering multi-dimensional relationships such as disease-symptoms-treatment-drugs, and supporting complex reasoning and professional knowledge retrieval. Furthermore, the information generation layer performs answer generation processing based on the knowledge retrieved from the knowledge graph, specifically generating query explanation information.

[0183] (3) Intelligent recommendation and payment technology: The item recommendation layer is based on the identified user's query intent, specifically the shopping intent, and is based on a context-aware recommendation algorithm. It combines the user's query intent, historical behavior, and current information search scenario to intelligently recommend the most relevant items, such as recommending professional medicines in a medical consultation scenario. The search engine outputs the searched recommended items and the generated query explanation information through the result display interface; and the recommended items in the result display interface carry an item purchase entrance. This item purchase entrance is connected to the payment interface layer, which deploys multiple payment interface APIs to achieve seamless connection from the display of recommended items to order payment, realize intelligent payment, and reduce user operation steps.

[0184] Further, Figure 14 The detailed components or modules included in each layer of the technical architecture can be found in Figure 15 .like Figure 15As shown, the information search interface includes input components corresponding to each information modality, including but not limited to voice input components, image input components, and text input components, as well as functional components such as result display components (for displaying query explanation information and recommended items) and payment interaction components (for realizing smart payment). The multimodal processing engine includes modal input modules corresponding to each information modality and a multimodal fusion engine, etc., which are mainly used to realize feature encoding and feature fusion of sub-query information under each information modality and generate cross-modal fusion features. The cross-modal fusion features output by the multimodal processing engine are input to the intent recognition layer. The intent recognition layer includes intent classification models (or algorithms), entity recognition models, and semantic understanding models (through natural language processing technology, to understand the true meaning of the user's question, rather than just staying at the keyword matching level), etc., which are mainly used to perform intent recognition on the cross-modal fusion features output by the multimodal processing engine to obtain the user's query intent. The knowledge retrieval layer includes domain-specific knowledge graphs and knowledge retrieval models. It is primarily used to perform knowledge retrieval based on the query intent output by the intent recognition layer to generate query interpretation information. Simultaneously, the result display engine recommends items based on the query intent output by the intent recognition layer and displays the searched recommended items. The knowledge retrieval layer outputs the query interpretation information to the information search interface, while the result display engine also outputs the searched recommended items to the information search interface to display the query results. If a user makes a payment for a recommended item in the query results, the payment processing layer executes the payment process based on its integrated payment processing module and security verification module, achieving a coherent implementation process for information search and intelligent payment.

[0185] Furthermore, the above Figure 14 and Figure 15 The data flow between the various layers or modules shown in the figure can be used to implement the complete interactive process of the information processing method provided in the embodiment of the present application. Figure 16 .like Figure 16 As shown, the background implementation process of the information processing method may include but is not limited to the following steps s1-s19:

[0186] s1: The user enters multimodal query information in the information search interface provided by the search engine layer.

[0187] s2: The search engine layer generates an information search request based on the multimodal query information and sends the information search request to the multimodal processing engine.

[0188] s3: The multimodal processing engine performs feature encoding and cross-modal feature fusion on the sub-queries contained in the multimodal query information to generate cross-modal fusion features. Specifically, different modal input modules are used to perform feature encoding on the sub-queries in the corresponding information modality to obtain the initial feature information for each sub-query. This initial feature information for each sub-query is then mapped to a cross-modal feature sharing space to obtain the target feature information for each sub-query in the cross-modal feature sharing space. Finally, the target feature information for each sub-query in the cross-modal feature sharing space is fused to generate cross-modal fusion features.

[0189] s4: The multimodal processing engine sends the cross-modal fusion features to the intent recognition module.

[0190] s5: The intent recognition module performs intent recognition processing on the cross-modal fusion features to generate query intent; the query intent includes at least question intent and shopping intent.

[0191] s6: The intent recognition module sends the query intent to the knowledge retrieval layer, requesting knowledge retrieval and generation.

[0192] s7: The knowledge retrieval layer queries the relevant knowledge content from the knowledge graph according to the query intent, generates query interpretation information based on the knowledge content, and sends the query interpretation information to the intent recognition module.

[0193] s8: The intent recognition module sends the query intent to the item recommendation layer, requesting item search and recommendation.

[0194] s9: The item recommendation layer searches for recommended items that match the user's shopping intent based on the query intent and returns the recommended items to the intent recognition module.

[0195] s10: The intent recognition module sends the query results to the search engine layer; the query results include query explanation information and recommended items.

[0196] s11: The book search engine layer displays the query results in the result display interface.

[0197] s12: The user searches for the recommended items in the results on the result display interface and executes a trigger operation by clicking on the item purchase entry.

[0198] s13: The search engine layer generates a payment request in response to the user's triggering operation on the item purchase portal, and sends the payment request to the payment processing layer.

[0199] s14: The payment processing layer receives the payment request and returns the shopping details of the recommended items to the search engine layer.

[0200] s15: The search engine layer displays shopping details in the result display interface.

[0201] s16: The search engine layer receives the payment operation performed by the user based on the shopping details information.

[0202] s17: The search engine layer generates payment information based on the payment operation and submits the payment information to the payment processing layer.

[0203] s18: The payment processing layer processes the payment for the recommended items based on the payment information, generates the payment result, and returns the payment result to the search engine layer.

[0204] s19: The search engine layer outputs the payment result.

[0205] The more detailed payment processing flow shown in steps s12-s19 above can be found in Figure 17 .like Figure 17 As shown, the payment processing flow may include but is not limited to the following steps s21-s214:

[0206] s21: The user triggers the purchase entry of the recommended item in the result display interface provided by the search engine layer.

[0207] s22: The search engine layer generates a payment request based on the trigger operation for the item purchase entrance and sends the payment request to the smart payment service; the payment request is used to request the generation of a payment order for the recommended item.

[0208] s23: The smart payment service generates a payment order for the recommended item in response to the payment request.

[0209] s24: The smart payment service returns the order information of the payment order to the search engine layer; the order information is the shopping details information.

[0210] s25: The search engine layer displays shopping details information through the result display interface.

[0211] s26: The user performs payment operations in the result display interface provided by the search engine layer.

[0212] s27: The search engine layer sends the payment information generated by the user's payment operation to the smart payment service.

[0213] s28: The smart payment service sends a payment processing request to the payment gateway, where the payment processing request is used to request payment processing for the recommended item based on the payment information.

[0214] s29: The payment gateway and the payment institution interact to facilitate the payment institution to perform security verification and payment processing based on the payment information and generate payment results.

[0215] s210: The payment institution returns the payment result to the payment gateway.

[0216] s211: The payment gateway sends a payment status notification to the smart payment service based on the payment result. The payment status notification is used to notify the smart payment service to update the order status of the payment order.

[0217] s212: The smart payment service sends a status notification to the search engine layer; if the payment result is successful, the status notification is used to notify that the payment order has been completed.

[0218] s213: The search engine layer outputs the payment result to the user.

[0219] s214: Smart payment service updates the order status of the payment order.

[0220] In summary, the information processing method provided by the embodiment of the present application has at least the following advantages: ① Improved search accuracy: Compared with the single-modal search mode, the embodiment of the present application supports improved accuracy based on multi-modal fusion search in complex information search scenarios, especially in professional fields such as medical care and childcare, and can more accurately understand user needs through the combination of pictures and texts. ② Improved user operation efficiency: Compared with the traditional cumbersome steps of application jump and page jump, the embodiment of the present application can realize one-click smart purchase after the recommended items are obtained through information search, which simplifies user operations while shortening the time of item purchase tasks, especially in emergency shopping scenarios. The effect is more significant. ③ Improved item purchase conversion rate: Supports seamless connection between information search and brandless payment, which greatly improves the purchase conversion rate of recommended items.

[0221] The method of the embodiment of the present application is described in detail above. In order to facilitate the above-mentioned scheme of the embodiment of the present application to be better implemented, accordingly, the device of the embodiment of the present application is provided below. In the embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuit or memory) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the module or unit function.

[0222] Figure 18 A schematic diagram of the structure of an information processing device provided by an exemplary embodiment of the present application is shown; the information processing device can be used to execute Figure 2 and Figure 11 Some or all of the steps in the method embodiment shown. Figure 18 , the device includes the following units:

[0223] Display unit 1801, used to display an information search entry, which is used to trigger an information search;

[0224] Processing unit 1802 is configured to receive input multimodal query information in response to a triggering operation on an information search portal; the multimodal query information includes sub-query information of at least two information modalities;

[0225] The processing unit 1802 is further configured to display query results obtained by performing information search based on the multimodal query information.

[0226] In one implementation, the processing unit 1802 is configured to, in response to a triggering operation on an information search portal, receive input multimodal query information and specifically:

[0227] In response to a triggering operation on an information search portal, an information search interface is displayed; the information search interface includes an information input area; wherein the information search portal includes at least one of the following: a search bar in a content search interface, a dialog box in an interactive dialogue interface, and a search pop-up window in a service interface;

[0228] According to the information input operation performed in the information input area, the input multimodal query information is displayed in the information input area; the information modality includes at least one of the following: text, image, voice, video, animation and file.

[0229] In one implementation, the information search interface further includes a multimodal input prompt, which is used to prompt that the information search interface allows the input of multimodal query information; the information input area includes multiple sub-input areas, each sub-input area corresponding to an information modality; the processing unit 1802 is configured to display the input multimodal query information in the information input area based on the information input operation performed in the information input area, specifically for:

[0230] In response to an information input operation performed in at least two of the plurality of sub-input areas, displaying sub-query information belonging to a corresponding information modality in a corresponding sub-input area;

[0231] According to a confirmation operation on the sub-query information displayed in the at least two sub-input areas, the sub-query information in the at least two sub-input areas is combined into multimodal query information.

[0232] In one implementation, the information input area further includes an area addition option; the processing unit 1802 is further configured to:

[0233] Displaying a modal selection window according to a triggering operation for adding an option to a region; the modal selection window includes a region creation option corresponding to at least one information modality;

[0234] In response to a selection operation on a target area creation option in the modal selection window, a target sub-input area corresponding to a target information modality is displayed in the information input area; the target information modality is an information modality corresponding to the target area creation option, and the target area creation option is any area creation option in the modal selection window.

[0235] In one implementation, the multimodal input prompt is presented as a modal combination option; the processing unit 1802 is further configured to:

[0236] In response to a selection operation on the modal combination option, displaying a plurality of sub-input areas in the information input area;

[0237] The processing unit 1802 is further configured to:

[0238] In response to a deselection operation on the modal combination option, a plurality of sub-input areas are deleted in the information input area.

[0239] In one implementation, the processing unit 1802 is configured to, when displaying the input multimodal query information in the information input area according to the information input operation performed in the information input area, specifically:

[0240] In response to an information input operation performed in the information input area, receiving input of first sub-query information;

[0241] If the initial query intent obtained based on the analysis of the first sub-query information does not meet the intent recognition condition, prompt input information is displayed; the prompt input information is used to prompt: input a second sub-query information representing the same query intent as the first sub-query information;

[0242] In response to an information input operation for the second sub-query information, the input second sub-query information is received; the multimodal query information includes the first sub-query information and the second sub-query information.

[0243] In one implementation, the processing unit 1802 is further configured to:

[0244] According to the closing operation performed on the input prompt information, a notification message is displayed; the notification message is used to prompt an information search based on the first sub-query information.

[0245] In one implementation, the prompt input information is used to prompt that the second sub-query information to be input belongs to the first information modality; the processing unit 1802 is further configured to:

[0246] According to the mode switching operation on the first information mode, the switched second information mode is displayed; the second information mode is different from the first information mode.

[0247] In one implementation, the multimodal query information includes an image belonging to an image modality; the processing unit 1802 is further configured to:

[0248] In response to an image input operation, displaying the input image;

[0249] According to the information annotation operation performed on the image, the annotated annotation information is displayed at the annotated position;

[0250] In response to a confirmation operation on the annotated image, the image and the annotated image are respectively used as two sub-query information included in the multimodal query information.

[0251] In one implementation, the processing unit 1802 is further configured to:

[0252] Displays the modality identifier corresponding to each sub-query information contained in the multimodal query information; any modality identifier is used to indicate the information modality to which the corresponding sub-query information belongs;

[0253] In response to a deletion operation performed on a target modality identifier, performing an information search to generate a new query result based on the remaining sub-query information in the multimodal query information except the sub-query information corresponding to the deleted target modality identifier; the target modality identifier is a modality identifier corresponding to any sub-query information;

[0254] Display new query results.

[0255] In one implementation, the processing unit 1802, when configured to display query results obtained by performing an information search based on multimodal query information, is specifically configured to:

[0256] Displaying a result display interface, the result display interface including query results obtained from the information search based on the multimodal query information; the query results including recommended items that meet the query intent represented by the multimodal query information, and the query results including an item purchase portal for the recommended items;

[0257] In response to a trigger operation on the item purchase portal, the purchase details corresponding to the recommended item are displayed on the result display interface; the purchase details include payment information and item information of the recommended item;

[0258] Output the payment result based on the payment operation performed in the result display interface.

[0259] In one implementation, the query result includes query interpretation information that matches the query intent represented by the multimodal query information and recommended items that match the query intent; the processing unit 1802 is further configured to:

[0260] Perform feature extraction and fusion processing on each sub-query information in the multimodal query information to generate cross-modal fusion features;

[0261] Perform contextual semantic prediction on cross-modal fusion features to obtain query intent represented by multimodal query information;

[0262] Search the knowledge base for information based on the query intent and generate query interpretation information that matches the query intent; and

[0263] Search the item library for recommended items that meet the query intent.

[0264] In one implementation, the processing unit 1802 is configured to perform feature extraction and fusion processing on each sub-query information in the multimodal query information to generate a cross-modal fusion feature, specifically for:

[0265] Using the feature extraction rules corresponding to the information modality, feature encoding is performed on the sub-query information belonging to the corresponding information modality to obtain the initial feature information corresponding to each sub-query information;

[0266] Mapping the initial feature information corresponding to each sub-query information to the cross-modal feature shared space to obtain the target feature information of each sub-query information in the cross-modal feature shared space;

[0267] The information correlation between the target feature information corresponding to at least two sub-query information is calculated, and feature fusion processing is performed according to the information correlation between the target feature information to generate cross-modal fusion features.

[0268] According to one embodiment of the present application, Figure 18 The various units in the information processing device shown can be individually or entirely combined into one or several other units to form a whole, or one (or some) of the units can be further divided into multiple functionally smaller units to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the information processing device may also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units. According to another embodiment of the present application, the following can be executed by running on a general-purpose computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM). Figure 2 and Figure 11 A computer program (including program code) for each step of the corresponding method shown in FIG. Figure 18The information processing apparatus shown in and the information processing method of the embodiment of the present application are implemented. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.

[0269] Based on the same inventive concept, the principles and beneficial effects of the problems solved by the information processing device provided in the embodiment of the present application are similar to the principles and beneficial effects of the problems solved by the information processing method in the method embodiment of the present application. Please refer to the principles and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.

[0270] Figure 19 FIG2 shows a schematic diagram of a computer device provided by an exemplary embodiment of the present application. Figure 19 , the computer device includes a processor 1901, a communication interface 1902 and a computer-readable storage medium 1903. The processor 1901, the communication interface 1902 and the computer-readable storage medium 1903 can be connected via a bus or other means. The communication interface 1902 is used to receive and send data. The computer-readable storage medium 1903 can be stored in the memory of the computer device, the computer-readable storage medium 1903 is used to store computer programs, and the processor 1901 is used to execute the computer programs stored in the computer-readable storage medium 1903. The processor 1901 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more computer programs, and is specifically suitable for loading and executing one or more computer programs to realize the corresponding method flow or corresponding function.

[0271] The embodiment of the present application also provides a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space that stores the processing system of the computer device. In addition, one or more computer programs suitable for being loaded and executed by the processor 1901 are also stored in the storage space. It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor.

[0272] In one embodiment, one or more computer programs are stored in a computer-readable storage medium; the processor 1901 loads and executes the one or more computer programs stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned information processing method embodiment; in a specific implementation, the one or more computer programs in the computer-readable storage medium are loaded and executed by the processor 1901 to execute the steps of each embodiment of the present application; wherein, the steps of each embodiment of the present application can be referred to the relevant descriptions of the aforementioned embodiments, and will not be repeated here.

[0273] Based on the same inventive concept, the principles and beneficial effects of solving problems by the computer device provided in the embodiment of the present application are similar to the principles and beneficial effects of solving problems by the information processing method in the method embodiment of the present application. Please refer to the principles and beneficial effects of the implementation of the method. For the sake of concise description, they will not be repeated here.

[0274] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned information processing method is implemented.

[0275] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0276] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes a computer program (one or more). When the computer program is loaded and executed on a computer device, the computer program executes the above-mentioned process or function of the embodiment of the present application. The computer device can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer program can be stored in a computer-readable storage medium or transmitted by a computer-readable storage medium. The computer program can be transmitted from a website site, computer device, server or data center to another website site, computer device, server or data center by wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer device can access or a data storage device such as a server or data center that includes one or more available media integrations. Available media can be magnetic media (for example, floppy disk, hard disk, tape), optical media (for example, DVD) or semiconductor media (for example, solid-state drive (SSD)) etc.

[0277] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An information processing method, characterized in that: include: Displaying an information search entry, wherein the information search entry is used to trigger execution of an information search; In response to a triggering operation on the information search portal, receiving input multimodal query information; The multimodal query information includes sub-query information of at least two information modalities; Display query results obtained by performing information search based on the multimodal query information.

2. The method according to claim 1, wherein The receiving of input multimodal query information in response to a triggering operation on the information search portal includes: In response to a triggering operation on the information search portal, an information search interface is displayed; the information search interface includes an information input area; wherein the information search portal includes at least one of the following: a search bar in a content search interface, a dialog box in an interactive dialogue interface, and a search pop-up window in a service interface; According to the information input operation performed in the information input area, the input multimodal query information is displayed in the information input area; the information modality includes at least one of the following: text, image, voice, video, animation and file.

3. The method according to claim 2, wherein The information input area includes a plurality of sub-input areas, each of the sub-input areas corresponding to one of the information modalities; and displaying the input multimodal query information in the information input area according to the information input operation performed in the information input area includes: In response to an information input operation performed in at least two of the plurality of sub-input areas, displaying sub-query information belonging to the corresponding information modality in the corresponding sub-input area; According to a confirmation operation on the sub-query information displayed in at least two of the sub-input areas, the sub-query information in at least two of the sub-input areas is combined into multimodal query information.

4. The method according to claim 3, wherein The information input area also includes an area adding option; the method further includes: Displaying a modal selection window according to a triggering operation for adding an option to the area; the modal selection window includes at least one area creation option corresponding to the information modality; In response to a selection operation of a target area creation option in the modal selection window, a target sub-input area corresponding to a target information modality is displayed in the information input area; the target information modality is an information modality corresponding to the target area creation option, and the target area creation option is any one of the area creation options in the modal selection window.

5. The method according to claim 3, wherein The information search interface further includes a multimodal input prompt, wherein the multimodal input prompt is used to prompt that the information search interface allows input of multimodal query information; the multimodal input prompt is presented as a modal combination option; and in response to an information input operation performed in at least two of the plurality of sub-input areas, before displaying sub-query information belonging to the corresponding information modality in the corresponding sub-input area, the method further includes: In response to a selection operation on the modal combination option, displaying a plurality of sub-input areas in the information input area; The method further comprises: In response to a deselection operation on the modal combination option, a plurality of the sub-input areas are deleted in the information input area.

6. The method according to claim 2, wherein The step of displaying the input multimodal query information in the information input area according to the information input operation performed in the information input area includes: receiving input of first sub-query information in response to an information input operation performed in the information input area; If the initial query intent obtained based on the analysis of the first sub-query information does not meet the intent recognition condition, prompting input information is displayed; the prompting input information is used to prompt: input a second sub-query information representing the same query intent as the first sub-query information; In response to an information input operation for the second sub-query information, the input second sub-query information is received; the multimodal query information includes the first sub-query information and the second sub-query information.

7. The method according to claim 6, wherein The method further comprises: A notification message is displayed according to a closing operation performed on the input prompt information; the notification message is used to prompt the user to perform the information search based on the first sub-query information.

8. The method according to claim 6 or 7, wherein: The prompt input information is used to prompt that the second sub-query information to be input belongs to the first information modality; before receiving the second sub-query information input in the information input area, the method further includes: According to the mode switching operation on the first information mode, a second information mode after switching is displayed; the second information mode is different from the first information mode.

9. The method according to claim 1, 2, 3 or 6, wherein: The multimodal query information includes an image belonging to an image modality; and the method further includes: In response to an image input operation, displaying the input image; Displaying the marked information at the marked position according to the information marking operation performed on the image; In response to a confirmation operation on the annotated image, the image and the annotated image are respectively used as two sub-query information included in the multimodal query information.

10. The method according to claim 1, wherein The method further comprises: Displaying the modality identifier corresponding to each sub-query information contained in the multimodal query information; any modality identifier is used to indicate the information modality to which the corresponding sub-query information belongs; In response to a deletion operation performed on a target modality identifier, performing an information search to generate a new query result based on the remaining sub-query information in the multimodal query information except the sub-query information corresponding to the deleted target modality identifier; the target modality identifier is a modality identifier corresponding to any of the sub-query information; The new query result is displayed.

11. The method according to claim 1, wherein The displaying of query results obtained by performing information search based on the multimodal query information includes: Displaying a result display interface, the result display interface including query results obtained from an information search based on the multimodal query information; the query results including recommended items that meet the query intent represented by the multimodal query information, and the query results including an item purchase portal for the recommended items; In response to a triggering operation on the item purchase portal, displaying purchase details corresponding to the recommended item on the result display interface; the purchase details information includes payment information and item information of the recommended item; Output the payment result according to the payment operation performed in the result display interface.

12. The method according to claim 1 or 11, wherein: The query result includes query interpretation information that meets the query intent represented by the multimodal query information, and recommended items that meet the query intent; Before displaying the query results obtained by searching for information based on the multimodal query information, the method further includes: Performing feature extraction and fusion processing on each sub-query information in the multimodal query information to generate a cross-modal fusion feature; Performing contextual semantic prediction processing on the cross-modal fusion features to obtain the query intent represented by the multimodal query information; Performing knowledge retrieval in the knowledge base according to the query intent to generate query interpretation information that meets the query intent; and According to the query intent, the recommended items that meet the query intent are searched from the item library.

13. The method according to claim 12, wherein: The performing feature extraction and fusion processing on each sub-query information in the multimodal query information to generate a cross-modal fusion feature includes: Using a feature extraction rule corresponding to the information modality, feature encoding is performed on the sub-query information belonging to the corresponding information modality to obtain initial feature information corresponding to each sub-query information; Mapping the initial feature information corresponding to each sub-query information to a cross-modal feature sharing space to obtain target feature information of each sub-query information in the cross-modal feature sharing space; The information correlation between the target feature information corresponding to at least two of the sub-query information is calculated, and feature fusion processing is performed according to the information correlation between the target feature information to generate a cross-modal fusion feature.

14. An information processing device, characterized in that: include: A display unit, configured to display an information search entry, wherein the information search entry is configured to trigger execution of an information search; a processing unit, configured to receive input multimodal query information in response to a triggering operation on the information search portal; the multimodal query information includes sub-query information of at least two information modalities; The processing unit is further configured to display query results obtained by performing information search based on the multimodal query information.

15. A computer device, characterized in that: include: a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the information processing method according to any one of claims 1 to 13 is implemented.

16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer application program, and when the computer application program is executed, the information processing method according to any one of claims 1 to 13 is implemented.

17. A computer program product, characterized in that The computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the information processing method according to any one of claims 1 to 13 is implemented.