Image-based information search method and device, electronic equipment and medium
The server-side image search system addresses inefficiencies in existing image-based search technologies by identifying primary objects, generating recommendations, and using large language models to provide sequential answers, thereby enhancing user engagement and reducing operational complexity.
Patent Information
- Application Number
- CN202510467896.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-15
AI Technical Summary
The existing image search technology has shortcomings in meeting the complex needs of users. Users need to take photos again or adjust the location of the image frame to obtain ideal answers. The entry and recommendation strategies lead to high input costs, making it difficult to fully meet the user's extended search needs.
By receiving the image to be identified, the first subject is determined and the image recognition results and recommendation information are generated, the reply content is generated using a large language model, and the relevant information is displayed sequentially on the client, reducing the complexity of the user's further search.
Stimulate users to expand their search needs, improve the richness of search interaction functions and user experience, reduce operation complexity, and improve information acquisition efficiency and accuracy.
Smart Images

Figure CN120316293A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and in particular to the fields of large models, computer vision, deep learning, natural language processing, and search technology. Specifically, it relates to an image-based information search method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art
[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has technologies at both the hardware level and the software level. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing: Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] Searching through image recognition is a technology that uses a computer to analyze the content of an image and find similar or relevant images in a large amount of image data. However, when the search results do not meet the user's needs, the user needs to perform additional operations, such as taking a new photo or repeatedly switching different pages. The user's operation is not convenient, the retrieval efficiency is low, and the user experience is poor. Summary of the Invention
[0004] The present disclosure provides an image-based information search method, apparatus, electronic device, computer-readable storage medium, and computer program product.
[0005] According to one aspect of the present disclosure, there is provided an image-based information search method for a server, including: in response to receiving a to-be-recognized image sent by a client, determining a first subject in the to-be-recognized image; determining a first image recognition result corresponding to the first subject and a plurality of recommended information corresponding to the first image recognition result, and sending the first image recognition result and the plurality of recommended information to the client for display, where the plurality of recommended information is used to guide the user to ask follow-up questions; based on the plurality of recommended information, generating, by a first large language model, multiple reply contents corresponding one-to-one to the plurality of recommended information; in response to receiving a follow-up question request for a target recommended information among the plurality of recommended information, determining, among the multiple reply contents, a first reply content corresponding to the target recommended information and at least one second reply content other than the first reply content; and sending the first reply content and the at least one second reply content to the client, so that the first reply content is displayed on the client, and in response to a continuous browsing operation after browsing the first reply content, the corresponding reply contents in the at least one second reply content are sequentially displayed on the client.
[0006] According to another aspect of the present disclosure, there is provided an image-based information search method, including: in response to receiving an image to be recognized, sending the image to be recognized to a server; obtaining a first image recognition result and a plurality of recommended information sent by the server to display the first image recognition result and the plurality of recommended information on a recognition result page, where the first image recognition result is determined based on a first subject in the image to be recognized, the plurality of recommended information is determined based on the first image recognition result, and the plurality of recommended information is used to guide a user to ask follow-up questions; in response to receiving a search request for a target recommended information among the plurality of recommended information, obtaining a first reply content and at least one second reply content sent by the server, and displaying the first reply content on a follow-up question result page, where the first reply content is a reply content corresponding to the target recommended information among multiple reply contents generated by a first large language model based on the plurality of recommended information, and the at least one second reply content is other reply contents among the multiple reply contents except the first reply content; and in response to receiving a continuous browsing operation after browsing the first reply content, sequentially displaying corresponding reply contents among the at least one second reply content on the follow-up question result page.
[0007] According to another aspect of the present disclosure, there is provided an image-based information search device, including: a receiving unit configured to determine a first subject in the image to be recognized in response to receiving the image to be recognized sent by a client; an identification result determination unit configured to determine a first image recognition result corresponding to the first subject and a plurality of recommended information corresponding to the first image recognition result, and send the first image recognition result and the plurality of recommended information to the client for display, where the plurality of recommended information is used to guide a user to ask follow-up questions; a reply content generation unit configured to generate multiple reply contents corresponding one-to-one to the plurality of recommended information based on the plurality of recommended information through a first large language model; a reply content determination unit configured to, in response to receiving a follow-up question request for a target recommended information among the plurality of recommended information, determine a first reply content corresponding to the target recommended information and at least one second reply content other than the first reply content among the multiple reply contents; and a reply content sending unit configured to send the first reply content and the at least one second reply content to the client, so that the first reply content is displayed on the client, and in response to a continuous browsing operation after browsing the first reply content, corresponding reply contents among the at least one second reply content are sequentially displayed on the client.
[0008] According to another aspect of the present disclosure, there is provided an image-based information search device, including: a sending unit configured to send the image to be recognized to a server in response to receiving the image to be recognized; a recognition result display unit configured to obtain a first image recognition result and a plurality of recommended information sent by the server to display the first image recognition result and the plurality of recommended information on a recognition result page, where the first image recognition result is determined based on a first subject in the image to be recognized, the plurality of recommended information is determined based on the first image recognition result, and the plurality of recommended information is used to guide the user to ask follow-up questions; a reply content display unit configured to obtain a first reply content and at least one second reply content sent by the server in response to receiving a search request for a target recommended information among the plurality of recommended information, and display the first reply content on a follow-up question result page, where the first reply content is the reply content corresponding to the target recommended information among multiple reply contents generated by a first large language model based on the plurality of recommended information, and the at least one second reply content is other reply contents among the multiple reply contents except the first reply content; and a reply content switching unit configured to sequentially display the corresponding reply contents in the at least one second reply content on the follow-up question result page in response to receiving a continuous browsing operation after browsing the first reply content.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the present disclosure.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the present disclosure.
[0011] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program that implements the method described in the present disclosure when executed by a processor.
[0012] According to one or more embodiments of the present disclosure, while obtaining an image recognition result based on a picture, it is also possible to actively stimulate the user's search demand for related content by providing recommended information related to the image recognition result; and based on the continuous browsing operation, the reply contents corresponding to multiple recommended information can be sequentially displayed, reducing the complexity of the user's further search, stimulating the user to extend and express their search demand, and enhancing the richness of the search interaction function, thereby improving the user experience.
[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings exemplarily illustrate embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0015] Figure 1 A schematic diagram showing an exemplary system in which various methods described herein can be implemented according to an embodiment of the present disclosure;
[0016] Figure 2 A flowchart showing a method for image-based information search for a server according to an embodiment of the present disclosure;
[0017] Figure 3 A flowchart showing a method for image-based information search for a client according to an embodiment of the present disclosure;
[0018] Figure 4 A schematic diagram showing a client image receiving page according to an embodiment of the present disclosure;
[0019] Figure 5 A schematic diagram showing a client recognition result page according to an embodiment of the present disclosure;
[0020] Figure 6A and 6B A schematic diagram showing a follow-up question result page according to an embodiment of the present disclosure;
[0021] Figure 7 A block diagram showing the structure of an image-based information search device for a server according to an embodiment of the present disclosure;
[0022] Figure 8 A block diagram showing the structure of an image-based information search device for a client according to an embodiment of the present disclosure; and
[0023] Figure 9 A block diagram showing an exemplary electronic device capable of implementing the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0025] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0026] The terms used in the description of various examples in the present disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0027] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0028] Figure 1 A schematic diagram of an exemplary system 100 in which various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure is shown. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.
[0029] In an embodiment of the present disclosure, the server 120 can run one or more services or software applications that enable a method for performing image-based information search.
[0030] In certain embodiments, the server 120 can also provide other services or software applications, which can include non-virtual environments and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software-as-a-service (SaaS) model.
[0031] In Figure 1 the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from system 100. Thus, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0032] Users may use client devices 101, 102, 103, 104, 105, and / or 106 to input images to be recognized, receive recognition results, or search results. The client device may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.
[0033] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. The client device is capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0034] Network 110 can be any type of network well-known to those skilled in the art, and it can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IP10, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, virtual network, virtual private network (VPN), intranet, extranet, blockchain network, public switched telephone network (PSTN), infrared network, wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0035] Server 120 can include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNI10 servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices of the server). In various embodiments, server 120 can run one or more services or software applications that provide the functions described below.
[0036] The computing units in server 120 can run one or more operating systems including any of the above operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0037] In some implementations, server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0038] In some embodiments, the server 120 may be a server of a distributed system or a server incorporating a blockchain. The server 120 may also be a cloud server or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in a cloud computing service system, which addresses the deficiencies of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0039] The system 100 may further include one or more databases 130. In certain embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information related to images to be recognized and search history. The databases 130 may reside in various locations. For example, the databases used by the server 120 may be local to the server 120 or may be remote from the server 120 and may communicate with the server 120 via a network-based or dedicated connection. The databases 130 may be of different types. In certain embodiments, the databases used by the server 120 may be relational databases, for example. One or more of these databases may store, update, and retrieve data to and from the databases in response to commands.
[0040] In certain embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0041] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatuses described according to the present disclosure.
[0042] Currently, there are many deficiencies in the related technologies of photo search in meeting the complex needs of users. For example, after a user uploads an image with unclear intent, the recognition result may not meet the user's needs, but the user can only initiate a retrieval again by taking a new photo or changing the image selection position in the hope of obtaining an ideal answer. Or, some users who recognize images have further search needs, but due to issues with the entry and recommendation strategies, the user input cost is high. Overall, it is difficult to fully meet the complex and extended search needs of users, and there is room for optimization.
[0043] Therefore, embodiments according to the present disclosure provide an image-based information search method for a server side. Figure 2 The flowchart of the image-based information search method according to an embodiment of the present disclosure is shown, as Figure 2As shown, method 200 includes: in response to receiving an image to be recognized sent by a client, determining a first subject in the image to be recognized (step 210); determining a first image recognition result corresponding to the first subject and a plurality of recommended information corresponding to the first image recognition result, and sending the first image recognition result and the plurality of recommended information to the client for display, where the plurality of recommended information is used to guide the user to ask follow-up questions (step 220); based on the plurality of recommended information, generating, through a first large language model, multiple reply contents corresponding one by one to the plurality of recommended information (step 230); in response to receiving a follow-up question request for a target recommended information among the plurality of recommended information, determining, among the multiple reply contents, a first reply content corresponding to the target recommended information and at least one second reply content other than the first reply content (step 240); and sending the first reply content and the at least one second reply content to the client, so that the first reply content is displayed on the client, and in response to a continuous browsing operation after browsing the first reply content, the corresponding reply contents in the at least one second reply content are sequentially displayed on the client (step 250).
[0044] According to an embodiment of the present disclosure, while obtaining an image recognition result based on a picture, it is also possible to actively stimulate the user's search demand for expanding relevant content by providing recommended information related to the image recognition result; and based on the continuous browsing operation, the reply contents corresponding to multiple recommended information can be sequentially displayed, reducing the complexity of the user's further search, stimulating the user to extend and express their search demand, and enhancing the richness of the search interaction function, improving the user experience.
[0045] In step 210, in response to receiving an image to be recognized sent by a client, a first subject in the image to be recognized is determined.
[0046] In some embodiments, when receiving an image to be recognized sent by a client, a first subject in the image to be recognized is determined. The first subject is a part of the image to be recognized, such as a corresponding object (such as a commodity, an animal, a vehicle, etc.) in the image to be recognized. Exemplarily, the first subject can be determined by extracting key information from the image, recognizing the user's search intention, and other means.
[0047] In some embodiments, the first subject may be a single object (such as a commodity, an animal, a building, etc.), or a combination of multiple objects or a specific scene.
[0048] According to some embodiments, determining the first subject in the image to be recognized includes: in response to receiving the image to be recognized, determining one or more subjects in the image to be recognized; and in response to determining that the image to be recognized includes multiple subjects, determining the first subject among the multiple subjects and at least one second subject among the multiple subjects other than the first subject.
[0049] In some embodiments, after receiving the image to be recognized, the system needs to determine one or more subjects in the image. This step aims to comprehensively extract each object contained in the image, making subsequent searches more targeted. The implementation method can be, for example, using image recognition technology to analyze and extract the object contours, features, etc. in the image. For example, when receiving an image to be recognized of a group photo of flowers, the system can recognize multiple subjects such as Kalanchoe blossfeldiana, roses, and lilies. In this case, by further determining the first subject among the multiple subjects and at least one second subject other than the first subject, it can provide convenience for the user to select different search objects in subsequent searches. Among them, the first subject can refer to the subject that is default selected by the system as the target image to be recognized.
[0050] In specific implementation, the first subject can be determined according to preset rules, such as according to the area size of the subject in the image, the degree of position centrality, through intention recognition methods, etc. For example, in the above-mentioned group photo of flowers, since Kalanchoe blossfeldiana occupies the center of the image and has a large area, the system can determine it as the first subject, and roses and lilies as the second subjects. Or, analyze each subject in the image to obtain the intention recognition result corresponding to the image, and this intention recognition result includes the original user search intention related to the subjects in the image. When the first subject may be a combination of multiple objects or a specific scenario, the intention recognition result may also involve the analysis of the relationships between objects, the attributes of objects (such as color, size, shape, etc.), and the arrangement of objects, etc.
[0051] In some embodiments, for the intention recognition of the image to be recognized, it can be implemented based on an intention recognition model, and this intention recognition model can be a trained generative model. This generative model generates an intention recognition result by identifying and locating each object in the image. For example, it can also be based on the user characteristics of relevant users, and this result can be used as the original search intention of the user.
[0052] In this way, by performing differential processing on different subjects, it can meet the diverse needs of users for different subjects and improve the accuracy and practicality of image processing.
[0053] According to some embodiments, determining the first subject in the image to be recognized further includes: in response to receiving a subject switching instruction, using the second subject corresponding to the subject switching instruction as the new first subject, and using the first subject before receiving the subject switching instruction as the new second subject.
[0054] In some embodiments, if the image contains two or more subjects, when the system receives a subject switching instruction, it can perform a subject switching operation to meet the user's requirement of flexibly switching the subject of concern according to their own needs and provide a more personalized image recognition service. Specifically, the system will redefine the second subject corresponding to the subject switching instruction as the new first subject, and at the same time change the first subject before receiving the subject switching instruction into the new second subject.
[0055] In this way, the system can flexibly change the selection of the image subject according to the user's instruction, so that the image recognition and related recommended information can closely focus on the key points currently concerned by the user, significantly improving the user experience, enhancing the interactivity and practicality of the image recognition system, and improving the accuracy of the image recognition result.
[0056] In step 220, determine the first image recognition result corresponding to the first subject and a plurality of recommended information corresponding to the first image recognition result, and send the first image recognition result and the plurality of recommended information to the client for display, where the plurality of recommended information is used to guide the user to ask follow-up questions.
[0057] In the present disclosure, the first image recognition result can provide the user with a basic understanding of the image subject, while the recommended information can be used to guide the user to ask follow-up questions and expand the search requirements.
[0058] Still taking the image to be recognized including Kalanchoe blossfeldiana as an example, the first image recognition result can be the image and / or text information about Kalanchoe blossfeldiana recognized based on the first subject of "Kalanchoe blossfeldiana". This text information can be, for example: "Kalanchoe blossfeldiana is an evergreen herbaceous succulent plant and an ideal indoor potted plant for winter and spring". Further, in this example, the recommended information corresponding to the first image recognition result can be "What plants look good when paired with Kalanchoe blossfeldiana", "Where is the suitable place to put Kalanchoe blossfeldiana at home", etc. Based on these recommended information, it can stimulate the user to further explore relevant knowledge.
[0059] According to some embodiments, the first image recognition result includes at least one of a title and an abstract, and determining the first image recognition result corresponding to the first subject includes: in response to obtaining at least one first search result that matches the first subject, based on the at least one first search result and the first subject, generating at least one of the title and the abstract through a second large language model, wherein each of the at least one first search results includes an image and text for describing the image, and the similarity between the image and the image of the first subject is greater than a preset threshold.
[0060] In some embodiments, each obtained search result may include an image and text for describing the image. Moreover, for each obtained search result, the similarity between the included image and the image of the first subject is greater than a preset threshold. By filtering based on the dimension of the relevance between the search result and the first subject in the image to be recognized, high-quality and highly relevant search results that match the first subject can be obtained.
[0061] In the above embodiments, after obtaining the matching search results, a second large language model can be used to generate at least one of the title and the abstract. For example, the second large language model can be a multimodal large language model, which can comprehensively analyze the image and text information in the search results, combine the characteristics of the first subject, and generate a general title and abstract.
[0062] Therefore, the above method of combining search results and a multimodal large language model for image recognition can provide more targeted and attractive first image recognition results for users, help users quickly understand the key information of the first subject, and improve the information acquisition efficiency and experience of users during the image recognition process.
[0063] According to some embodiments, the first image recognition result includes at least one of a title and an abstract, and determining the first image recognition result corresponding to the first subject includes: in response to not obtaining at least one first search result that matches the first subject, generating at least one of the title and the abstract based on the first subject through a second large language model.
[0064] In some embodiments, correspondingly, it may be due to reasons such as the first subject in the image being too unique and novel, having less relevant public data, or the limited coverage of the search algorithm, resulting in the inability to obtain at least one search result that matches the first subject. In this case, at least one of the title and the abstract can be directly generated based on the first subject with the help of a multimodal large language model.
[0065] For example, if the first subject is a rare plant, the large model can generate relevant titles and abstracts based on the appearance of the plant in the image (such as leaf shape, flower color, etc.) and the identification text that may exist in the image. For example, the title can be "Mysterious Rare Plant Appears", and the abstract can be "This plant has unique XXX-shaped leaves and bright XXX flowers, similar to the XX plant under the XX genus of the XX family, but its special form may have been affected by a certain environment".
[0066] In this way, even if there are no ready-made matching search results, the system can still provide users with a first image recognition result with certain reference value through the intelligent analysis of the multi-modal large language model, helping users initially understand the key information of the first subject and improving the user experience and information acquisition efficiency during the image recognition process.
[0067] According to some embodiments, determining a plurality of recommended information corresponding to the first image recognition result includes: obtaining at least one first search result that matches the first subject, where each search result in the at least one first search result includes an image and text for describing the image, and the similarity between the image and the image of the first subject is greater than a preset threshold; generating the plurality of recommended information based on the at least one first search result and the first subject through a third large language model.
[0068] In some embodiments, the system first obtains at least one search result that matches the first subject. For example, the system can search and filter in a massive image database by means of an image matching algorithm and big data resources. Each search result includes an image and text for describing the image, and the similarity between these images and the image of the first subject is greater than a preset threshold. The preset threshold can be set in advance according to the actual application scenario and requirements to ensure the relevance and similarity between the search results and the first subject.
[0069] In the above embodiments, based on the at least one search result and the first subject obtained, the system can use a third large language model to generate at least one recommended information. For example, the third large language model can be a multi-modal large language model. Since the multi-modal large language model has powerful information integration and analysis capabilities, it can comprehensively consider the image and text information in the search results and the characteristics of the first subject, thereby generating targeted and practical recommended information. These recommended information can guide users to ask further questions and conduct searches, helping users more comprehensively understand the knowledge and information related to the first subject.
[0070] According to some embodiments, generating the plurality of recommendation information based on the at least one first search result and the first subject through a third large language model includes: in response to determining that the to-be-identified image part corresponding to the first subject includes text data, obtaining the text data; and generating the plurality of recommendation information based on the at least one first search result, the first subject, and the text data through a third large language model.
[0071] In some embodiments, when the to-be-identified image part corresponding to the first subject further includes text data, the text data can also be extracted first, because the text in the image often contains key information and can provide important clues for generating recommendation information. For example, when the to-be-identified image is a picture of a drug instruction manual and the first subject is the drug, the text on the instruction manual includes the ingredients, efficacy, usage and dosage of the drug, etc. Obtaining this text data can enable the system to more comprehensively understand the image information. Then, at least one search result, the first subject, and the text data can be combined, and at least one recommendation information can be generated through a third large language model.
[0072] Generating recommendation information by combining the information in three aspects can better meet the user's need for in-depth exploration of the to-be-identified image and guide the user to ask more valuable questions and conduct more valuable searches.
[0073] According to some embodiments, obtaining at least one first search result matching the first subject includes: determining the respective historical search information corresponding to at least one of the at least one first search results, wherein during the historical search process, the corresponding user searched for the corresponding search result among the at least one first search results based on the historical search information and the corresponding user browsed the corresponding search result during the search process; and generating the plurality of recommendation information based on the at least one first search result and the first subject through a third large language model includes: generating the plurality of recommendation information based on the at least one first search result, the first subject, and the respective historical search information corresponding to the at least one first search result through a third large language model.
[0074] In some embodiments, when obtaining at least one search result matching the first subject, the system can further determine the respective historical search information corresponding to these search results. Specifically, in some examples, the system can trace back which search information the user based on during the past search process (such as the search process based on a search engine) to find these specific search results, and confirm that the corresponding user browsed these results during the search process.
[0075] In some examples, the historical search information may be the historical search information determined based on a search engine. Exemplarily, the historical search process implemented based on the search engine can be obtained, and it can be determined which users retrieved one or more of the at least one search result based on which search terms or search statements during one or more historical search processes and browsed the one or more search results. In some examples, the search terms or search statements, and / or the keywords determined based on the search terms or search statements can be used as the historical search information.
[0076] It can be understood that in the present disclosure, other search records corresponding to the historical search process can also be used as the historical search information, which is not limited herein.
[0077] Therefore, in the stage of generating recommendation information, the system can comprehensively analyze at least one search result, the first subject, and the historical search information corresponding to each of these search results by means of, for example, a multimodal large language model to generate recommendation information that better meets the actual needs of users.
[0078] The recommendation information generated by tracing the historical search information can more accurately match the potential needs of users, significantly improve the efficiency and accuracy of users in obtaining image-related information, and optimize the user experience in the process of image recognition and information acquisition.
[0079] In step 230, based on the multiple recommendation information, a plurality of response contents corresponding one-to-one to the multiple recommendation information are generated by a first large language model.
[0080] According to some embodiments, generating a plurality of response contents corresponding one-to-one to the multiple recommendation information by a first large language model based on the multiple recommendation information includes: for each of the multiple recommendation information, performing the following operations: searching for at least one second search result based on the recommendation information; and obtaining a response content corresponding to the recommendation information by the first large language model based on the at least one second search result.
[0081] In some embodiments, for each of the multiple recommendation information, the system can first search based on the recommendation information to obtain at least one second search result. The purpose of this step is to collect content closely related to the recommendation information from a wide range of information sources. These information sources may cover professional databases, Internet web pages, industry reports, etc. After obtaining at least one second search result, the large language model can be used to analyze and integrate these results to obtain a response content corresponding to the recommendation information.
[0082] By collecting and integrating information based on recommendation information and generating high-quality response content with the help of large language models, accurate and detailed information services can be provided for users, effectively improving the quality and efficiency of intelligent interaction.
[0083] In step 240, in response to receiving a follow-up request for a target recommendation information among the multiple recommendation information, among the multiple response contents, determine a first response content corresponding to the target recommendation information and at least one second response content other than the first response content. In step 250, send the first response content and the at least one second response content to the client, so that the first response content is displayed on the client, and in response to a continuous browsing operation after browsing the first response content, the corresponding response content in the at least one second response content is sequentially displayed on the client.
[0084] In some embodiments, when receiving a follow-up request for a target recommendation information among multiple recommendation information, the first response content corresponding to the target recommendation information can be determined first, and the response contents corresponding to other recommendation information are used as the second response content. After the user browses the first response content and performs a continuous browsing operation, the corresponding response content in at least one second response content is then displayed.
[0085] In some examples, each of the multiple response contents may include corresponding recommendation information, such as as a title. In this way, users can intuitively know which recommendation information the response content is generated based on, thereby improving the user experience.
[0086] In the present disclosure, the continuous browsing operation may be, for example, a sliding browsing operation (including up and down sliding, left and right sliding, etc.) on the display page, inputting a preset continuous browsing instruction, etc., which is not limited herein.
[0087] In this way, a complete process from information generation received from an image to accurate feedback according to user follow-up is realized, improving the intelligence and personalization of the image recognition service and optimizing the user experience in the information interaction process.
[0088] According to some embodiments, determine one or more expected image recognition tasks that match the to-be-recognized image; and in response to receiving a trigger request for a target image recognition task among the one or more expected image recognition tasks, determine a second image recognition result that matches the target image recognition task.
[0089] In some embodiments, after receiving an image to be recognized, the system can expand the tasks that the user may wish to accomplish through image recognition, namely the search intent, based on information such as the features of the image, past data, and the user's behavior pattern. Thus, one or more desired image recognition tasks that match the image are determined to provide more targeted services for the user.
[0090] In the example, when the system receives a trigger request for a target image recognition task among these desired image recognition tasks, it further determines a second image recognition result that matches the target image recognition task. Thus, it can accurately provide detailed information that meets the user's expectations for the user's specific needs.
[0091] In this way, the system can actively discover the user's potential needs, accurately provide the required information when the user selects a specific task, improve the intelligence and accuracy of the image recognition service, expand the user's needs, reduce the complexity of the user's operations, and optimize the user's experience of using the image recognition function.
[0092] According to some embodiments, determining one or more desired image recognition tasks that match the image to be recognized includes: predicting an image search intent that matches the image to be recognized; and when the image search intent meets the recommendation trigger condition, obtaining one or more desired image recognition tasks corresponding to the image search intent.
[0093] In some embodiments, the recommendation trigger condition is a trigger condition set in advance by the system for generating recommendation content based on an image. The recommendation trigger condition can be set or adjusted according to different image contents or the user's original search intent to ensure the accuracy, relevance, and necessity of the recommendation content.
[0094] Exemplarily, the recommendation trigger condition can be based on a series of preset rules or algorithms to determine whether to generate recommendation content for the user. For example, the recommendation trigger condition can include: the image does not include code information (such as QR codes, barcodes, etc.). For instance, when the image search intent is a product intent, the desired image recognition task can be to find the same model; when the image search intent is a text intent, the desired image recognition tasks can be to extract text, translate, etc.; when the image search intent is a question intent, the desired image recognition tasks can be to solve the problem, extract text, etc.; when the image search intent is a face intent, the desired image recognition task can be to measure facial features; when the image search intent is a material / expression intent, the desired image recognition tasks can be to find AI similar images, change expressions, etc.
[0095] By presetting the recommendation trigger conditions, the system can more accurately explore the potential needs of users, provide relevant image recognition tasks according to a reasonable trigger mechanism, lay a foundation for users to obtain the required information subsequently, improve the intelligence and personalization of the image recognition service, enable users to obtain valuable content more conveniently, and enhance the experience and satisfaction of users when using the image recognition function.
[0096] According to some embodiments, determining one or more expected image recognition tasks that match the image to be recognized includes: in response to recognizing target type data in the image to be recognized, obtaining one or more expected image recognition tasks preset corresponding to the target type data.
[0097] Exemplarily, algorithms including big data analysis and image feature recognition can be used to recognize the target type data in the image to be recognized. In some examples, the association relationship between the target type data and the corresponding expected image recognition tasks can be preset in advance to achieve the rapid determination of the expected image recognition tasks.
[0098] In some embodiments, the system can pre - establish a task association database for storing the correspondence between various target type data and expected image recognition tasks, so as to accurately provide users with the possible required image recognition tasks for different image contents, and improve the pertinence and practicality of the service.
[0099] For example, when it is recognized that the image to be processed includes text data, the expected image recognition task can include text extraction; further, when the text is in a foreign language, the expected image recognition task can further include translation. In particular, the expected image recognition task can include, for example, a "find similarity" task. For any image to be recognized, a "find similarity" task can be preset as the matching image recognition task.
[0100] According to some embodiments, the target type data includes at least one of the following items: text type data, code type data, and wherein, the expected image recognition tasks corresponding to the text type data include at least one of the following items: text extraction task, text translation task; and the expected image recognition task corresponding to the code type data includes: code scanning task.
[0101] In some embodiments, when the system processes the received image to be recognized, it will first judge the type data. Here, the type data can include at least one of text type data and code type data. For example, for a commodity image containing text type data, the expected image recognition tasks can include "text extraction" tasks, "text translation" tasks; for an image containing code type data such as two - dimensional codes, barcodes, etc., the expected image recognition tasks can include "code scanning" tasks.
[0102] In this way, the system can provide accurate and practical image recognition tasks for users based on different types of data in the image, effectively improving the efficiency of image recognition and the user experience.
[0103] According to an embodiment of the present disclosure, there is also provided an image-based information search method for a client. Figure 3 The flowchart of the image-based information search method according to an embodiment of the present disclosure is shown as Figure 3 shown. The method 300 includes: in response to receiving an image to be recognized, sending the image to be recognized to the server (step 310); obtaining a first image recognition result and a plurality of recommended information sent by the server to display the first image recognition result and the plurality of recommended information on the recognition result page, where the first image recognition result is determined based on a first subject in the image to be recognized, and the plurality of recommended information is determined based on the first image recognition result, and the plurality of recommended information is used to guide the user to ask follow-up questions (step 320); in response to receiving a search request for a target recommended information among the plurality of recommended information, obtaining a first reply content and at least one second reply content sent by the server, and displaying the first reply content on the follow-up question result page, where the first reply content is the reply content corresponding to the target recommended information among the multiple reply contents generated by a first large language model based on the plurality of recommended information, and the at least one second reply content is the other reply contents except the first reply content among the multiple reply contents (step 330); and in response to receiving a continuous browsing operation after browsing the first reply content, sequentially displaying the corresponding reply contents in the at least one second reply content on the follow-up question result page (step 340).
[0104] In the present disclosure, the terms and features in the image-based information search method for the client have the same or similar meanings as the related terms and features in the image-based information search method for the server. And in some embodiments, they have the same or similar implementation manners as those described in the above embodiments, and will not be elaborated here.
[0105] According to some embodiments, one or more subjects in the image to be recognized are determined, and the one or more subjects are respectively identified by subject anchors. The first subject is the corresponding subject among the one or more subjects.
[0106] In the present disclosure, the operation of determining the subject in the image to be recognized has been described in the above embodiments, and will not be elaborated here.
[0107] In this embodiment, to display the different identified subjects, they can be identified through subject anchors. A subject anchor is a technical means capable of precisely marking the position and features of the image subject, which can be image coordinates, specific identifiers, or a set of data describing the subject features. Through this identification method, different subjects can be clearly distinguished, facilitating subsequent image processing.
[0108] Figure 4 FIG. shows a schematic diagram of a client image receiving page according to an embodiment of the present disclosure. As Figure 4 shown, during the process of receiving the image to be recognized or after receiving the image to be recognized, the area of the image to be recognized obtained through the viewfinder 401 or the area including each subject in the image to be recognized can be used. And, one or more subjects in the image to be recognized can be determined, such as Figure 4 the corresponding subject 402 identified by the anchor point 403 as shown in
[0109] According to some embodiments, one or more subjects in the image to be recognized and a first subject among the one or more subjects are determined, where the first subject is a part of the image in the image to be recognized; while displaying the first recognition result and the at least one recommended information on the recognition result page, the first subject is displayed in a first display area of the recognition result page, and the one or more subjects and the image to be recognized are respectively displayed in a plurality of second display areas of the recognition result page.
[0110] In the present disclosure, the operations of determining the first subject and the second subject have been described in the above-mentioned embodiments and will not be elaborated here.
[0111] In some embodiments, while displaying the recognition result and the recommended information on the recognition result page, the relevant subject image can be further displayed, so that the user can intuitively learn which image and which part of the image the current recognition result is based on.
[0112] Figure 5 FIG. shows a schematic diagram of a client recognition result page according to an embodiment of the present disclosure. As Figure 5 shown, on the recognition result page, the obtained first recognition result 501 and at least one recommended information 502 can be displayed. As described above, the first recognition result can include a title and an abstract. Further, the first display area 503 of the recognition result page displays a first subject (a part of the image to be recognized), and the one or more subjects and the image to be recognized are displayed in its second display area 504.
[0113] Therefore, through the above embodiments, the user can have a more comprehensive understanding of the overall image and various subjects in the image, improving the richness of the image interaction function and further enhancing the user experience.
[0114] According to some embodiments, the second display area corresponding to the first subject among the multiple second display areas is presented in a first display mode, and the second display areas corresponding to other subjects except the first subject and the image to be recognized are presented in a second display mode.
[0115] In some embodiments, during the display process of the image recognition result, different display modes can be adopted for different display objects in the multiple second display areas of the recognition result page. Among them, the second display area corresponding to the first subject is presented in a first display mode. For example, the first display mode may be more prominent in visual effects, such as using highlights, more vivid colors, unique border styles, etc. Thus, the user can quickly determine the main part of the image being displayed.
[0116] Furthermore, the second display areas corresponding to other subjects except the first subject and the image to be recognized are presented in a second display mode. For example, the second display mode may be less prominent in visual effects compared to the first display mode, such as not using highlights, having relatively soft colors, or using a more concise border style, etc. Thus, the user can understand other subjects of the image and the overall situation of the image.
[0117] Continue to refer to Figure 5 , the second display area 505 corresponding to the first subject in the multiple second display areas 504 is presented in a first display mode (i.e., a bold border), and the second display areas 506 corresponding to other subjects except the first subject and the image to be recognized are presented in a second display mode (i.e., a thinner border). It can be seen that the second display area corresponding to the first subject is more prominent in visual effects than the second display areas corresponding to other subjects.
[0118] By this way of distinguishing the display mode, when the user views the recognition result page, they can not only quickly focus on the first subject and obtain key information, but also simultaneously master other elements and the overall situation of the image, effectively improving the efficiency and experience of the user in obtaining image information.
[0119] According to some embodiments, in response to receiving a subject switching instruction, the other subject or the image to be recognized in the second display area corresponding to the subject switching instruction is used as the new first subject and is displayed in the first display area.
[0120] In some embodiments, if the first subject determined by the system is not the user's target concern or the user has a need to change the first subject, the subject displayed in the first display area can be switched to the subject that the user needs to focus on. Accordingly, the display mode of the corresponding subject in the second display area can also be adjusted accordingly. This interaction design effectively enhances the interactivity between the user and the image recognition system, allows the user to explore various types of information in the image more flexibly, and improves the image recognition effect.
[0121] According to some embodiments, one or more expected image recognition tasks determined to match the image to be recognized are obtained, so that while the first image recognition result and the at least one recommended information are displayed on the recognition result page of the search application, the one or more expected image recognition tasks are also displayed on the recognition result page; in response to receiving a trigger request for a target image recognition task among the one or more expected image recognition tasks, a second image recognition result matching the target image recognition task is obtained to display the second image recognition result on the recognition result page.
[0122] In the present disclosure, the operation of determining the expected image recognition task has been described in the embodiments above and will not be elaborated here.
[0123] Continue to refer to Figure 5 , in addition to the first image recognition result 501 and the at least one recommended information 502, one or more expected image recognition tasks 507 can also be displayed on the recognition result page to further explore the user's search needs, so that when the user views the recognition result, various image recognition task options matching the image to be recognized can be intuitively seen, broadening the user's exploration direction of image-related information. After receiving a trigger request for a target image recognition task among the one or more expected image recognition tasks, a second image recognition result matching the target image recognition task can be obtained.
[0124] As described above, for a product image containing text-type data, the expected image recognition tasks can include a "text extraction" task and a "text translation" task. After triggering the "text translation" task, the translation text corresponding to the text in the image to be recognized can be obtained.
[0125] In some examples, the obtained second image recognition result can be displayed on the recognition result page. For example, a new recognition result page generated in the form of a card or a pop-up window. Of course, a new result page can also be generated to display the second image recognition result on the new result page.
[0126] In this way, the system can actively explore the potential needs of users, provide diverse image recognition task options, and accurately provide the required information according to the user's selection, significantly improving the intelligence level and user experience of the image recognition service. When users use the search application for image recognition, they can obtain valuable information more efficiently and comprehensively.
[0127] According to some embodiments, in response to receiving a continuous browsing operation after browsing the first reply content, successively displaying the corresponding reply content in the at least one second reply content on the follow-up result page includes: in response to receiving a continuous browsing operation after browsing the first reply content, arranging the first reply content as a new second reply content after the at least one second reply content, and based on the browsing direction corresponding to the continuous browsing operation, displaying the corresponding second reply content in the at least one second reply content as a new first reply content.
[0128] In some embodiments, in order to provide users with a more coherent and more browsing-habit-compliant information display method and meet the users' need for in-depth and comprehensive understanding of relevant information, multiple reply contents can be displayed to users simultaneously. The multiple reply contents can be displayed in sequence and can be cyclically displayed in response to the users' continuous browsing operation. That is, when a continuous browsing operation by the user is detected, the original first reply content can be used as a new second reply content and arranged after the original at least one second reply content.
[0129] According to some embodiments, the first reply content and the at least one second reply content are displayed in the form of cards on the follow-up result page.
[0130] In this embodiment, when the reply content is displayed in the form of cards, each reply content can be independently presented on a card, making the information hierarchy clear and enabling users to easily distinguish different replies. At the same time, the card-based display method facilitates operations such as viewing, swiping, and clicking on individual cards by users. For example, users can swipe through the cards to browse each reply content in sequence, and clicking on a card may expand more detailed information or perform related interactive operations. In addition, the layout of the cards can be flexibly adjusted according to the interface design of the client and the screen size of the device to adapt to different display environments and ensure a good user experience on various devices.
[0131] Therefore, displaying the reply content in the form of cards not only improves the clarity and readability of information display but also enhances the interactivity between the user and the interface and the convenience of operation, further optimizing the user experience in the process of image recognition and information acquisition.
[0132] Figure 6A and 6BShows a schematic diagram of a follow-up question result page according to an embodiment of the present disclosure. As Figure 6A shown, in the follow-up question result page 601, the first reply content 602 is currently displayed, and it can be seen that the second reply content 604 is arranged below the first reply content 602. When the user finishes browsing the first reply content 602 and triggers a continuous upward sliding operation to continue browsing, the follow-up question result page changes from Figure 6A to Figure 6B shown, where the second reply content 604 can be displayed as the new first reply content 606, and the original first reply content 602 can be arranged as the new second reply content 605 after the current new first reply content 606. As Figure 6A and Figure 6B shown, each reply content includes a title 603 generated from the corresponding recommended information.
[0133] This processing method in the continuous browsing scenario of multiple reply contents enables users to obtain a more fluent experience when browsing information, without the need to manually search for other relevant reply contents, improving the efficiency of users to obtain information.
[0134] According to some embodiments, the follow-up question page further includes at least one of the following items: the image to be recognized, the follow-up question search box.
[0135] In some embodiments, the appearance of the image to be recognized on the follow-up question page allows users to always closely associate the original image with various obtained reply contents, thereby more intuitively understanding the reply contents. At the same time, displaying the image to be recognized on the follow-up question page can also help users confirm whether the object of the current follow-up question is accurate, avoiding information acquisition errors caused by memory deviation and other reasons.
[0136] In some embodiments, the follow-up question search box can further enhance the user's independent exploration ability. Users can enter more specific and personalized follow-up question content in the search box to perform more accurate queries. The setting of the follow-up question search box makes the follow-up question process no longer limited to the system-predefined recommended information, and can meet the diverse information needs of users.
[0137] According to an embodiment of the present disclosure, as Figure 7As shown, an image-based information search device 700 for a server is also provided, including: a receiving unit 710 configured to determine a first subject in the to-be-recognized image in response to receiving the to-be-recognized image sent by a client; a recognition result determination unit 720 configured to determine a first recognition result corresponding to the first subject and a plurality of recommended information corresponding to the first recognition result, and send the first recognition result and the plurality of recommended information to the client for display, where the plurality of recommended information is used to guide the user to ask follow-up questions; a reply content generation unit 730 configured to generate, based on the plurality of recommended information, multiple reply contents corresponding one by one to the plurality of recommended information through a first large language model; a reply content determination unit 740 configured to, in response to receiving a follow-up request for a target recommended information among the plurality of recommended information, determine a first reply content corresponding to the target recommended information and at least one second reply content other than the first reply content from among the multiple reply contents; and a reply content sending unit 750 configured to send the first reply content and the at least one second reply content to the client, so that the first reply content is displayed on the client, and in response to a continuous browsing operation after browsing the first reply content, the corresponding reply contents in the at least one second reply content are sequentially displayed on the client.
[0138] Here, the operations of the above units 710-750 of the image-based information search device 700 for the server are respectively similar to the operations of steps 210-250 described above, and will not be elaborated here.
[0139] According to an embodiment of the present disclosure, as Figure 8As shown, an image-based information search device 800 for a client is also provided, including: a sending unit 810 configured to send the image to be recognized to the server in response to receiving the image to be recognized; a recognition result display unit 820 configured to obtain a first image recognition result and a plurality of recommended information sent by the server to display the first image recognition result and the plurality of recommended information on a recognition result page, where the first image recognition result is determined based on a first subject in the image to be recognized, the plurality of recommended information is determined based on the first image recognition result, and the plurality of recommended information is used to guide the user to ask follow-up questions; a reply content display unit 830 configured to obtain a first reply content and at least one second reply content sent by the server and display the first reply content on a follow-up question result page in response to receiving a search request for a target recommended information among the plurality of recommended information, where the first reply content is the reply content corresponding to the target recommended information among the multiple reply contents generated by a first large language model based on the plurality of recommended information, and the at least one second reply content is the other reply content among the multiple reply contents except the first reply content; and a reply content switching unit 840 configured to sequentially display the corresponding reply content in the at least one second reply content on the follow-up question result page in response to receiving a continuous browsing operation after browsing the first reply content.
[0140] Here, the operations of the above-mentioned units 810-840 of the target video generation device 800 for the client are respectively similar to the operations of the steps 310-340 described above, and will not be elaborated here.
[0141] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0142] According to an embodiment of the present disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0143] Reference Figure 9, a block diagram of an electronic device 900 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0144] As Figure 9 shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0145] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, an output unit 907, a storage unit 908, and a communication unit 909. The input unit 906 can be any type of device that can input information into the electronic device 900. The input unit 906 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 907 can be any type of device that can present information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 908 can include but are not limited to magnetic disks, optical disks. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMa10 device, a cellular communication device, and / or the like.
[0146] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as method 200 or 300. For example, in some embodiments, method 200 or 300 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the method 200 or 300 described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute method 200 or 300 in any other suitable manner (e.g., by means of firmware).
[0147] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0148] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0149] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0150] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0151] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0152] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0153] It should be understood that the various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0154] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after this disclosure.
Claims
1. An image-based information search method, comprising: Responding to receiving an image to be recognized sent by a client, determining a first subject in the image to be recognized; Determining a first image recognition result corresponding to the first subject and a plurality of recommended information corresponding to the first image recognition result, and sending the first image recognition result and the plurality of recommended information to the client for display, wherein the plurality of recommended information is used to guide the user to ask follow-up questions; Based on the plurality of recommended information, generating, by a first large language model, multiple reply contents corresponding one by one to the plurality of recommended information; Responding to receiving a follow-up question request for a target recommended information among the plurality of recommended information, determining, among the multiple reply contents, a first reply content corresponding to the target recommended information and at least one second reply content other than the first reply content; And Sending the first reply content and the at least one second reply content to the client, so that the first reply content is displayed on the client, and in response to a continuous browsing operation after browsing the first reply content, the corresponding reply contents among the at least one second reply content are sequentially displayed on the client.
2. The method according to claim 1, wherein Determining the first subject in the image to be recognized includes: Responding to receiving an image to be recognized, determining one or more subjects in the image to be recognized; and Responding to determining that the image to be recognized includes multiple subjects, determining a first subject among the multiple subjects and at least one second subject among the multiple subjects other than the first subject.
3. The method according to claim 2, wherein, Determining the first subject in the image to be recognized further includes: Responding to receiving a subject switching instruction, using the second subject corresponding to the subject switching instruction as the new first subject, and using the first subject before receiving the subject switching instruction as the new second subject.
4. The method according to claim 1, further comprising: Determining one or more expected image recognition tasks matching the image to be recognized; And Responding to receiving a trigger request for a target image recognition task among the one or more expected image recognition tasks, determining a second image recognition result matching the target image recognition task.
5. The method according to claim 4, wherein Determining one or more expected image recognition tasks matching the image to be recognized includes: Predicting an image search intention matching the image to be recognized; and When the image search intention meets a recommendation trigger condition, obtaining one or more expected image recognition tasks corresponding to the image search intention.
6. The method according to claim 4, wherein Determining one or more expected image recognition tasks matching the image to be recognized includes: Responding to recognizing target type data in the image to be recognized, obtaining one or more expected image recognition tasks preset corresponding to the target type data.
7. The method according to claim 6, wherein, The target type data includes at least one of the following items: text type data, code type data, and wherein, The expected image recognition tasks corresponding to the text type data include at least one of the following items: text extraction task, text translation task; and The expected image recognition task corresponding to the code type data includes: a code scanning task.
8. The method according to claim 1, wherein The first image recognition result includes at least one of a title and an abstract, and wherein determining the first image recognition result corresponding to the first subject includes: In response to obtaining at least one first search result that matches the first subject, based on the at least one first search result and the first subject, generating at least one of the title and the abstract through a second large language model, wherein each of the at least one first search results includes an image and text for describing the image, and wherein the image similarity of the image to the image of the first subject is greater than a preset threshold.
9. The method according to claim 1 or 8, wherein The first image recognition result includes at least one of a title and an abstract, and wherein determining the first image recognition result corresponding to the first subject includes: In response to not obtaining at least one first search result that matches the first subject, generating at least one of the title and the abstract through the second large language model based on the first subject.
10. The method according to claim 1, wherein Determining a plurality of recommended information corresponding to the first image recognition result includes: Obtaining at least one first search result that matches the first subject, wherein each of the at least one first search results includes an image and text for describing the image, and wherein the image similarity of the image to the image of the first subject is greater than a preset threshold; Generating the plurality of recommended information through a third large language model based on the at least one first search result and the first subject.
11. The method according to claim 10, wherein, Generating the plurality of recommended information through a third large language model based on the at least one first search result and the first subject includes: In response to determining that the part of the image to be recognized corresponding to the first subject includes text data, obtaining the text data; Generating the plurality of recommended information through a third large language model based on the at least one first search result, the first subject, and the text data.
12. The method according to claim 10, wherein, Obtaining at least one first search result that matches the first subject includes: Determining the historical search information corresponding to each of one or more first search results in the at least one first search result, wherein during the historical search process, the corresponding user searched for the corresponding first search result in the one or more first search results based on the historical search information and the corresponding user browsed the corresponding first search result during the search process; and Generating the plurality of recommended information through a third large language model based on the at least one first search result and the first subject includes: Generating the plurality of recommended information through a third large language model based on the at least one first search result, the first subject, and the historical search information corresponding to each of the one or more first search results.
13. The method according to claim 1, wherein, Generating a plurality of reply contents corresponding one-to-one to the plurality of recommended information through a first large language model includes: For each of the plurality of recommended information, perform the following operations: Searching for at least one second search result based on the recommended information; and Based on the at least one second search result, obtain response content corresponding to the recommendation information through a first large language model.
14. An image-based information search method, comprising: In response to receiving an image to be recognized, send the image to be recognized to a server; Obtain a first image recognition result and a plurality of recommendation information sent by the server to display the first image recognition result and the plurality of recommendation information on a recognition result page, where the first image recognition result is determined based on a first subject in the image to be recognized, and the plurality of recommendation information is determined based on the first image recognition result, and the plurality of recommendation information is used to guide the user to ask follow-up questions; In response to receiving a search request for a target recommendation information among the plurality of recommendation information, obtain a first response content and at least one second response content sent by the server, and display the first response content on a follow-up question result page, where the first response content is the response content corresponding to the target recommendation information among the multiple response contents generated by the first large language model based on the plurality of recommendation information, and the at least one second response content is the other response contents except the first response content among the multiple response contents; and In response to receiving a continuous browsing operation after browsing the first response content, sequentially display the corresponding response contents in the at least one second response content on the follow-up question result page.
15. The method according to claim 14, wherein, In response to receiving a continuous browsing operation after browsing the first response content, sequentially displaying the corresponding response contents in the at least one second response content on the follow-up question result page includes: In response to receiving a continuous browsing operation after browsing the first response content, arrange the first response content as a new second response content after the at least one second response content, and based on the browsing direction corresponding to the continuous browsing operation, display the corresponding second response content in the at least one second response content as a new first response content.
16. The method according to claim 14 or 15, wherein Display the first response content and the at least one second response content in the form of cards on the follow-up question result page.
17. The method according to claim 14 or 15, wherein, The follow-up question page further includes at least one of the following items: the image to be recognized, a follow-up question search box.
18. The method according to claim 14, further comprising: Determine one or more subjects in the image to be recognized, and respectively identify the one or more subjects through subject anchors, where the first subject is the corresponding subject among the one or more subjects.
19. The method according to claim 14 or 18, further comprising: Determine one or more subjects in the image to be recognized and a first subject among the one or more subjects, where the first subject is a part of the image to be recognized; While displaying the first image recognition result and the plurality of recommendation information on the recognition result page, display the first subject in a first display area of the recognition result page, and display the one or more subjects and the image to be recognized in a plurality of second display areas of the recognition result page respectively.
20. The method according to claim 19, wherein, The second display area corresponding to the first subject among the multiple second display areas is displayed in a first display mode, and the second display areas corresponding to other subjects and the image to be recognized except the first subject are displayed in a second display mode.
21. The method according to claim 20, further comprising: In response to receiving a subject switching instruction, using the other subject or the image to be recognized in the second display area corresponding to the subject switching instruction as a new first subject, and displaying it in the first display area.
22. The method according to any one of claims 14-21, further comprising: Obtaining one or more expected image recognition tasks determined to match the image to be recognized, and displaying the one or more expected image recognition tasks on the recognition result page of the search application while displaying the first recognition result and the at least one recommended information on the recognition result page; In response to receiving a trigger request for a target image recognition task among the one or more expected image recognition tasks, obtaining a second recognition result matching the target image recognition task, and displaying the second recognition result on the recognition result page.
23. An image-based information search device, comprising: A receiving unit configured to determine a first subject in the image to be recognized in response to receiving the image to be recognized sent by a client; A recognition result determination unit configured to determine a first recognition result corresponding to the first subject and a plurality of recommended information corresponding to the first recognition result, and send the first recognition result and the plurality of recommended information to the client for display, wherein the plurality of recommended information is used to guide the user to ask follow-up questions; A reply content generation unit configured to generate a plurality of reply contents corresponding one-to-one to the plurality of recommended information through a first large language model based on the plurality of recommended information; A reply content determination unit configured to, in response to receiving a follow-up question request for a target recommended information among the plurality of recommended information, determine a first reply content corresponding to the target recommended information and at least one second reply content other than the first reply content in the plurality of reply contents; and A reply content sending unit configured to send the first reply content and the at least one second reply content to the client, so that the first reply content is displayed on the client, and in response to a continuous browsing operation after browsing the first reply content, the corresponding reply content in the at least one second reply content is sequentially displayed on the client.
24. An image-based information search device, comprising: A sending unit configured to send the image to be recognized to a server in response to receiving the image to be recognized; An identification result display unit configured to obtain a first image recognition result and a plurality of recommendation messages sent by the server, and display the first image recognition result and the plurality of recommendation messages on an identification result page, where the first image recognition result is determined based on a first subject in the image to be recognized, and the plurality of recommendation messages are determined based on the first image recognition result, and the plurality of recommendation messages are used to guide the user to ask follow-up questions; A reply content display unit configured to, in response to receiving a search request for a target recommendation message among the plurality of recommendation messages, obtain a first reply content and at least one second reply content sent by the server, and display the first reply content on a follow-up question result page, where the first reply content is the reply content corresponding to the target recommendation message among the multiple reply contents generated by a first large language model based on the plurality of recommendation messages, and the at least one second reply content is the other reply contents among the multiple reply contents except the first reply content; and A reply content switching unit configured to, in response to receiving a continuous browsing operation after browsing the first reply content, sequentially display the corresponding reply contents among the at least one second reply content on the follow-up question result page.
25. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-22.
26. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-22.
27. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-22.
Citation Information
Patent Citations
Information display method and device, electronic equipment and storage medium
CN117572999A
Image-based search processing method and device, equipment and storage medium
CN118690033A
Image-based search processing method and device, equipment and storage medium
CN118708740A
Recommendation question generation method and device based on large model, electronic equipment and medium
CN119537537A