Image-based search processing method and apparatus, electronic device, storage medium, and program

The image-based search processing method addresses inefficiencies in traditional image-based search by generating recommendation content alongside results, enabling users to clarify their intent without re-uploading images, thus improving search efficiency and user experience.

JP2025181604AActive Publication Date: 2025-12-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024194250
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-30
Filing Date
2024-11-06
Publication Date
2025-12-11
Estimated Expiration
2044-11-06

Smart Images

  • Figure 2025181604000001_ABST
    Figure 2025181604000001_ABST
Patent Text Reader

Abstract

To provide an image-based search processing method and apparatus, an electronic device, a storage medium, and a program, relating to the technical field of artificial intelligence, particularly to the technical fields of computer vision, deep learning, natural language processing, and searching, and applicable to scenarios such as image recognition, image searching, intelligent recommendation, and the like.SOLUTION: A method includes: determining, in response to receiving an image sent by a terminal, a first requirement corresponding to the image, the first requirement including an original search intent of a user; generating a search result and recommended content corresponding to the image according to the first requirement in a case where the first requirement meets a recommendation triggering condition; and returning the search result and the recommended content to the terminal to cause the terminal to output the search result and the recommended content on a search result page. Accordingly, search efficiency may be improved, and the intelligence and convenience of the search may be improved.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of artificial intelligence, in particular to the field of computer vision, deep learning, natural language processing and search technologies, which can be applied to scenarios such as image recognition, image search and intelligent recommendation, and particularly to an image-based search processing method, apparatus, device and storage medium. [Background technology]

[0002] In a traditional image-based search process, a user inputs an image and is presented with search results. Summary of the Invention [Problem to be solved by the invention]

[0003] However, if the search results do not meet the user's requirements, the user needs to clarify or refine the requirements or take new photos, which increases the complexity of the search and reduces search efficiency. [Means for solving the problem]

[0004] The present disclosure provides an image-based search processing method, apparatus, electronic device, storage medium, and program.

[0005] According to a first aspect of the present disclosure, there is provided an image-based search processing method, the method comprising: In response to receiving an image transmitted from a terminal, determining a first request corresponding to the image, the first request including an original search intent of a user; generating search results and recommendation content corresponding to the image based on the first request when the first request satisfies a recommendation trigger condition; and returning the search results and the recommendation content to the terminal so that the terminal outputs the search results and the recommendation content on a search result page.

[0006] According to a second aspect of the present disclosure, there is provided an image-based search processing method, the method comprising: In response to receiving an image entered by a user via a search application, transmitting the image to a server; receiving search results and recommendations returned from the server, the search results being determined when the image satisfies a recommendation trigger condition, and the search results being determined based on a first request corresponding to the image, the first request including a user's original search intent; and displaying the search results and the recommendation content on a search result page.

[0007] According to a third aspect of the present disclosure, there is provided an image-based search processing device, comprising: a first determination module for determining, in response to receiving an image transmitted from a terminal, a first request corresponding to the image, the first request including a user's original search intent; a generating module for generating search results and recommendation content corresponding to the image based on the first request when the first request satisfies a recommendation trigger condition; and a first communication module for returning the search results and the recommendation content to the terminal so that the terminal outputs the search results and the recommendation content on a search result page.

[0008] According to a fourth aspect of the present disclosure, there is provided an image-based search processing device, comprising: a second communication module for, in response to receiving an image input by a user through a search application, transmitting the image to a server and receiving search results and recommendation content returned from the server, the search results being determined when the image satisfies a recommendation trigger condition, and the search results being determined based on a first request corresponding to the image, the first request including an original search intent of the user; and and an output control module for displaying the search results and the recommendation content on a search result page.

[0009] According to a fifth aspect of the present disclosure, there is provided an electronic device, the device comprising: at least one processor; a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the at least one processor to perform any one of the image-based search processing methods in the embodiments of the present disclosure.

[0010] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute any one of the image-based search processing methods in the embodiments of the present disclosure.

[0011] According to a seventh aspect of the present disclosure, there is provided a program for, when executed by a processor, realizing any one of the image-based search processing methods in the embodiments of the present disclosure.

[0012] According to an aspect of the present disclosure, when searching based on an image, a method is adopted to output search results and recommendation content in the first round of replies. If the search results cannot meet the user's requirements, the user not only does not need to retake and upload the image, but also does not need to spend more time and energy expressing the requirements. The user can quickly clarify the search requirements through the recommendation content, thereby improving search efficiency and enhancing the intelligence and convenience of the search.

[0013] It should be understood that the contents described herein are not intended to describe key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be better understood through the following specification.

[0014] The accompanying drawings are for a better understanding of the solutions of the present disclosure, but are not intended to limit the present disclosure. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a flowchart of an image-based search processing method according to an embodiment of the present disclosure. [Figure 2] 10 is a flowchart of another image-based search processing method according to an embodiment of the present disclosure. [Figure 3] 1 is a flowchart based on image retrieval according to an embodiment of the present disclosure. [Figure 4] FIG. 10 is a schematic diagram illustrating an interface after triggering a camera entry in a search application according to an embodiment of the present disclosure. [Figure 5] 1 is a first schematic diagram illustrating search results and recommendation contents based on an image search according to an embodiment of the present disclosure. [Figure 6] FIG. 2 is a second schematic diagram illustrating search results and recommendation contents based on an image search according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a first schematic diagram illustrating search results generated based on a selection operation of recommendation content according to an embodiment of the present disclosure. [Figure 8] FIG. 10 is a second schematic diagram illustrating search results generated based on a selection operation of recommendation content according to an embodiment of the present disclosure. [Figure 9] FIG. 1 is a schematic diagram illustrating a technical framework for implementing a search according to an embodiment of the present disclosure. [Figure 10] 1 is a first schematic diagram illustrating a configuration of an image-based search processing device according to an embodiment of the present disclosure. [Figure 11] FIG. 2 is a second schematic diagram illustrating the configuration of an image-based search processing device according to an embodiment of the present disclosure. [Figure 12] 1 is a schematic diagram illustrating a scenario of an image-based search processing method according to an embodiment of the present disclosure; [Figure 13] FIG. 1 is a schematic diagram illustrating the configuration of an electronic device for implementing an image-based search processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0016] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. Various details of the embodiments of the present disclosure are included herein for ease of understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, in the following description, descriptions of known functions and structures will be omitted for clarity and conciseness.

[0017] The terms "first," "second," "third," etc. in the examples and claims of the specification of this disclosure and in the above-described drawings are intended to distinguish between similar objects without necessarily being used to describe a particular order or priority. Furthermore, the terms "comprise" and "have" and variations thereof are intended to be non-exclusive inclusions, e.g., the inclusion of a series of things or units. Methods, systems, products, or devices need not be limited to the things or units explicitly listed, but may include other things or units not explicitly listed or inherent to these processes, methods, products, or devices.

[0018] Traditional image recognition products rely on image recognition, image search (topic search, same image search, plant / animal / product identification, etc.), and image processing (real-time translation, document scanning, etc.) technologies to meet user needs by taking photos or uploading images from photo albums for searching. However, this method only retrieves existing online search results, and if users have follow-up questions, they must return to the image recognition product's text search box to continue asking. For example, if a user notices a warning light on the dashboard while driving a car, they take a photo and search, which only satisfies a basic cognitive need. They must then ask the text search box, "How do I fix the tire pressure light?" to get a satisfactory answer. If the recognition results are not as expected, the user must take a photo or upload the image from the album again to perform image recognition. Even if an image recognition product has an image search function, the user can upload the image and then fill in the text questions themselves to search. However, this method provides few satisfactory answers, makes it difficult to answer the user's questions, and fails to meet the user's potential needs such as "what to do" and "why." If the user is not satisfied with the answer, they have no choice but to re-upload the image and replace the written questions with another search. This makes the user's search operation route complicated, and the desired results cannot be obtained, so the user's needs cannot be met.

[0019] Artificial intelligence (AI) dialogue products primarily use text rather than image-based search. AI dialogue products can understand and generate natural language and engage in deep, multi-round dialogue, but they lack the ability to guide, reverse-prompt, or recommend. Conversation generation is based on user input and preset rules within the model. If the user's request or question is unclear, questions or prompts can be generated to prompt the user to provide more information, but they do not provide reverse-prompts or recommendations based on search results or query terms. For example, if a user uploads multiple body images without a specific query, the model's generation may be ineffective and not meet the user's expectations. The user may not necessarily know the cause or the optimal direction for prompts, resulting in an unsatisfactory answer.

[0020] Therefore, image recognition products such as traditional image recognition products and AI dialogue products have at least the following problems in the search process.

[0021] First, there is a lack of detailed guidance for unclear requests.

[0022] In traditional image recognition products, if a user uploads an image with an unclear purpose, the image recognition results may not necessarily meet the user's needs. The user must retake the photo or reposition the image frame to restart the search to find the ideal answer. In AI interactive products, when a user uploads an image with an unclear purpose, the AI ​​interactive product can provide possible answers based on existing information or obtain more information through generalization guidance. For example, "I'm not very sure about your specific request. Can you provide more information or clarify the problem? That way, we can help you better." This guidance method is vague and impersonal. If the user is unable to express their request precisely, this guidance lacks direction, leading to the user spending more time and energy expressing their request, relying on the user's expressive ability. This reduces search efficiency, fails to guarantee satisfaction, and may even reduce the user's trust in the AI ​​product.

[0023] Second, there is a lack of understanding scenarios for the new features.

[0024] With the development of technology, image recognition products have realized more and more new functions. However, new users are often instructed on these new functions using animations and sample diagrams. This method consumes time and understanding costs for users. While it can achieve good results for tool functions (such as translation, text extraction, code scanning, etc.), it cannot clearly help users understand the usage scenarios of non-tool functions, especially AI-related functions (such as drawing from diagrams and transcribing from diagrams).

[0025] Third, there is a lack of response to expansion requests.

[0026] When searching, there may be a need for augmentation other than recognition, such as when taking a photo of a product. However, in image-based searches, the search results of conventional image recognition products could not meet the augmentation needs.

[0027] To at least partially solve one or more of the above-mentioned problems and other potential problems, the present disclosure proposes an image-based search processing method, in which, during an image search, a server first determines a first request corresponding to an image, the first request including a user's original search intent, and if the first request satisfies a recommendation trigger condition, generates search results and recommendation content corresponding to the image based on the first request, and then a first round of replies is made by a terminal based on the image search, and the search results and recommendation content are output on a search result page. By adopting a method of outputting search results and recommendation content in the first round of replies based on image search, if the search results cannot meet the user's requirements, the user does not need to take new photos and upload images, and the user does not need to spend more time and energy expressing their requirements. The user can quickly clarify their search requirements through the recommendation content, thereby improving search efficiency, improving search intelligence and convenience, and further enhancing the search experience.

[0028] The embodiments of the present disclosure provide an image-based search processing method, which can be applied to a server, which has an image search function and supports image search, image recognition, request guidance, function recommendation, etc. In practical applications, the server includes, but is not limited to, a normal server, a cloud server, etc. As shown in Figure 1, the image-based search processing method:

[0029] In S101, in response to receiving an image sent from a terminal, a first request corresponding to the image is determined, where the first request includes a user's original search intent.

[0030] In S102, if the first request satisfies a recommendation trigger condition, a search result and a recommendation content corresponding to the image are generated based on the first request.

[0031] In S103, the search results and the recommendation contents are returned to the terminal so that the terminal can output the search results and the recommendation contents on a search result page.

[0032] In an embodiment of the present disclosure, a search application client that supports an image search function is provided on a terminal. A camera entry icon is provided in the search box on the top page of the search application. By triggering this camera entry icon, a user can directly take an image and then search for that image. Alternatively, by triggering this camera entry icon, a user can also select an image from an album as an image for search.

[0033] In the embodiments of the present disclosure, original search intent refers to the initial, direct purpose or request a user has when performing a search operation. When a user uploads a single image for search, they usually have a specific goal or request, and these goals or requests constitute the original search intent. Original search intents include, but are not limited to, the following types: First, information acquisition. A user may want to learn more about a concept, event, person, or product. Second, problem-solving. A user may encounter a problem and hope to find a solution or suggestion through search. Third, purchasing decision. A user may be considering purchasing a product or service and may want to compare different options through search or find product reviews. Fourth, entertainment or recreation. A user may want to find interesting content to kill time, such as watching videos, reading articles, or looking at images. Fifth, navigation or location. A user may want to find the exact location or contact information of a place, business, or event. Sixth, comparison or evaluation. A user may want to compare the pros and cons of different products or services to make a decision. The above is merely an exemplary illustration and is not intended as a limitation on all possible content of the original search intent, and a full list is not provided here.

[0034] In an embodiment of the present disclosure, the recommendation trigger condition is a trigger condition preset by the system for generating recommendation content based on an image. The recommendation trigger condition can be set or adjusted based on different image content or original search intent to ensure the accuracy, relevance, and necessity of the recommendation content. In some embodiments, the recommendation trigger condition may be to determine whether to generate recommendation content for a user based on a set of preset rules or algorithms.

[0035] In an embodiment of the present disclosure, the recommendation trigger conditions may include at least one of the following: the image does not contain code-type information (e.g., a two-dimensional code, a barcode, etc.); the image contains multiple different types of objects (e.g., a beach and a surfboard, a table and a cup, a keyboard and a flower, etc.); the image contains multiple similar objects (e.g., pedestrians in a crowd, spectators in the stands, multiple porcelain bowls, multiple books, multiple shoes, etc.). The above is merely an exemplary description and is not intended to limit all possible ways of recommendation trigger conditions, so a list will be omitted here. In actual application, the system can adjust the scope of recommendation trigger conditions based on current resource utilization, which helps the system maintain efficient and stable operation while providing timely and relevant recommendations to users.

[0036] In an embodiment of the present disclosure, the search results are generated based on the first request corresponding to an image, and the search results may include at least one of the following image identification results: text in the identified image, text content in the image, text described by the image, description text generated based on the image content, image information, other images similar to the uploaded image, image source, original source or link of the provided image, identified objects in the image, objects, faces, or scenarios in the identified image, web page or product information, such as a purchase link or user review for a product returned if a product is identified, video, a related video clip or full video returned if the image is a portion of a video, audio, an audio file or audio description related to the image content, a web page link returned if the image is associated with a web page or online content, knowledge information, such as encyclopedia, interpretation, or background information based on the image content, cultural, historical, or scientific information related to objects or scenarios in the image, location information, such as location information, maps, or travel advice for a landmark or specific location returned if the image includes the location, and social media content, such as related posts, user comments, or labels returned if the image is shared on social media. The above is merely an exemplary description and is not intended as a limitation on all possible image identification results, and a list will not be provided here.

[0037] In embodiments of the present disclosure, recommendations are used to help users more clearly express or understand their search intent and provide them with content for potential further manipulation. For example, recommendations are a series of additional content, suggestions, alternative search options, or prompts for further clarification that are automatically recommended to the user that may be related to or of interest to the original search intent. When the images uploaded by the user are not clear or the search intent is not specific, recommendations can help the user further clarify what they want to search for and help the user clarify their search request.

[0038] In real applications, recommendations may include other images similar to the original image, keywords or labels related to the original search intent, popular or trending content related to the original search intent, prompts to clarify possible search intent (e.g., "Are you looking for XX brand products?"), suggestions for related classifications or subcategories, relevant user reviews and comments, etc.

[0039] In an embodiment of the present disclosure, the recommendation content may include recommendation content generated based on the first request corresponding to the image. Specific recommendation content may vary depending on the content of the image and the user's original search intent. Recommendation content that may be generated based on the image content includes, but is not limited to, the following types: First, product recommendation. If the image shows a single product (e.g., clothing, electronic products, household goods, etc.), similar products, related products, the same product in different colors or sizes, similar products with high user ratings, etc. may be recommended. Second, sightseeing recommendation. If the image is a photo of a tourist attraction, sightseeing courses related to the tourist attraction, nearby hotels, local specialty foods, transportation options, travel tips, etc. may be recommended. Third, style or design recommendation. If the image shows a specific design style (e.g., interior design, apparel design, artwork, etc.), other design works, designers, design courses, or design software with a similar style may be recommended. Fourth, book or movie recommendation. If the image resembles a book or movie cover, books, movies, and dramas on related themes may be recommended. Fifth, gourmet food recommendation. If the image shows a dish or ingredients, recommendations can include recipes for that dish, cooking techniques, links to purchase related ingredients, and nearby restaurants that serve that dish. 6. Celebrity or artist recommendations. If the image is a photo or work of a celebrity or artist, recommendations can include other works by that celebrity or artist, related news, social media accounts, concert or exhibition information, etc. 7. Tutorial or educational material recommendations. If the image shows a skill or activity such as arts and crafts, cooking, or fitness, recommendations can include tutorials, instructional videos, online lessons, or books related to that skill. 8. Parts recommendations. If the image is of a product (e.g., a cell phone, camera, car, etc.), recommendations can include parts, accessories, protective covers, and expansion equipment that are compatible with the product. 9. Health or fitness recommendations.If the image shows fitness activities or healthy foods, related fitness plans, healthy recipes, fitness equipment, or nutritional supplements can be recommended. Tenth, service or application recommendations. If the image is related to some service or application (such as a web page accessed by scanning a QR code, or a screenshot within a game), related services, applications, special events, etc. can be recommended. The above is merely an exemplary description and is not intended as a limitation on all possible content of recommendations, so a list will be omitted here.

[0040] In the embodiments of the present disclosure, the search result page refers to the result page presented by the search application client after receiving the images taken by the user. The present disclosure does not limit the page layout style or specifications of the search result page.

[0041] In practical application, after receiving an image taken by a user, the search application client on the terminal sends the image to a server, the server determines a first request corresponding to the image, and if the first request satisfies a recommendation trigger condition, generates search results and recommendation contents corresponding to the image based on the first request, and the server returns the search results and recommendation contents to the terminal, so that the search results and recommendation contents are output on a search result page by the terminal. That is, the user only needs to upload an image, that is, after the user uploads the image, the user does not need to input any other auxiliary information before the search application client outputs the first round of search results, and the search application client can directly output the search results and recommendation contents corresponding to the image.

[0042] In the solution of the embodiment of the present disclosure, the server determines a first request corresponding to the image in response to receiving an image sent from a receiving terminal. If the first request satisfies a recommendation trigger condition, it generates search results and recommendation content corresponding to the image based on the first request, and outputs the search results and recommendation content on the search result page when the terminal returns the first round of search results. While the server returns search results based on the image and search terms sent from the terminal, the present disclosure eliminates the need for the user to input search terms along with the image, simplifying the image search process and improving search efficiency and search experience. Compared to a processing method that only outputs search results, the present disclosure also outputs recommendation content at the same time as outputting search results, eliminating the need for the user to spend more time and effort expressing their request. Users can quickly clarify their search intent based on the recommendation content, reducing the number of searches and improving search efficiency, while also improving search accuracy and search satisfaction.

[0043] In an embodiment of the present disclosure, the recommendation content includes at least one of a guidance content, a counter-question content, and a function entry, wherein:

[0044] The guidance content includes a first type of information for guiding the user to clarify search intent;

[0045] The reverse query includes a second type of information for the user to search for to clarify the user's search intent;

[0046] The function entry indicates an entry into available functions that are recommended to the user.

[0047] In some embodiments, navigational content is generally used to prompt the user to further clarify search intent or navigate to more specific search directions.

[0048] As an example, if a user uploads an image containing multiple flowers, recommendations may include, "In light of the fact that you uploaded images of flowers, would you like to know the names of these flowers, how to care for them, or are you looking for a link to buy them?"

[0049] As an example, if a user uploads an image containing shoes, the recommendation might be, "Are you looking for this shoe in a different color? Or would you like to know your size and availability?"

[0050] In some embodiments, the reverse query explores the user's true intentions through questioning methods, helping the system to more accurately understand the user's requirements.

[0051] For example, after a user uploads a landscape image, the recommendation could be, "Maybe you'd like to travel here? Here are some related travel routes and strategies for you to use as a reference."

[0052] For example, for a gourmet image, you could ask, "Want to learn how to make this dish? We have detailed recipes and video tutorials."

[0053] In some embodiments, the function entry provides the user with direct entry to available functions or services, allowing the user quick access to the next action.

[0054] For example, while displaying image search results, a function entry "more similar images" can be provided to allow the user to view more related images.

[0055] For example, when recommending a product, a functional entry such as "Add to Shopping Cart" or "Buy Now" can be provided so that the user can purchase it directly.

[0056] For example, if a user searches for services such as travel, repairs, etc., the recommendations may include function entries such as "book service" or "contact customer service."

[0057] The above is merely an illustrative example and is not intended to be a limitation on all possible content recommendations, so a list will not be provided here. In practical applications, these recommendations may be generated in conjunction with information that the user has agreed to disclose, such as the user's search history, browsing behavior, geographic location, and personal preferences, to ensure accuracy and individualization of the recommendations.

[0058] In this way, the recommendations are more humane and intelligent, and can more accurately meet the user's personalized needs.By analyzing the user's feedback and behavioral data on the recommendations, it is possible to quickly gain insight into the user's search intent and preferences, thereby providing more accurate search results that meet the user's expectations and improving the accuracy and effectiveness of the recommendations.

[0059] In an embodiment of the present disclosure, determining a first request corresponding to the image includes determining a description model suitable for a content type of the image, generating image content description information for the image based on the description model, and determining the first request based on the image content description information.

[0060] In some embodiments, an image classification model or object detection algorithm is used to identify objects in an image, such as people, animals, natural scenes, products, etc., and to determine the content type of the image. Here, the image classification model can use computer vision techniques to perform image classification and object detection. Note that the image classification model is a pre-trained model, and the present disclosure does not limit the training method or structure of the image classification model. Any algorithm that can detect objects in an image can be the object detection algorithm of the present disclosure, and the present disclosure does not impose any restrictions on the object detection algorithm.

[0061] In embodiments of the present disclosure, the descriptive model may be a predefined template, a rule set, or a machine learning model.

[0062] Different content types may have different description requirements and foci: for example, a product image may need to describe attributes such as the product's brand, model, and color, while a natural landscape image may need to describe its overall atmosphere, major tourist attractions, etc.

[0063] In some embodiments, a predefined template or rule set may be used to generate the basic description, for example, for a product image, there may be a template that includes fields such as brand, model, color, etc.

[0064] In some embodiments, for more complex scenarios, machine learning-based text generation models can be used to generate more natural and detailed explanations.

[0065] In some embodiments, an image is used as input to generate a description of the image content using a description model. This typically involves converting image features into natural language text. Specifically, image elements such as recognized objects, colors, and textures are input as features into the description model, and the description model generates a corresponding natural language text description based on these features, which may be a simple sentence, phrase, or paragraph. Furthermore, language generation techniques such as text summarization and text generation can be used to optimize the generated description to make it more fluid and easy to understand.

[0066] In some embodiments, generating image content description information for the image based on the description model comprises:

[0067] When the content type is a two-dimensional code or a barcode class, identifying the link of the two-dimensional code or the information of the barcode through code identification technology;

[0068] Identifying text information in an image using Optical Character Recognition (OCR) technology when the content type is text or title class;

[0069] When the content type is a person or a plant class, generate a word guessing result (e.g., "this is xxx") using a word guessing model, and generate image content description information using a multimodal large language model when the word guessing model does not generate a result;

[0070] This may include using a multimodal large language model to generate image content description information (e.g., This is an image of a tire pressure alarm on a car dashboard) where the content type is a class of material, face, animal, facial expression, product, or other class.

[0071] In the embodiment of the present disclosure, the image content description information may be description information generated based on features such as objects, scenarios, and colors in the image.

[0072] In some embodiments, generating a first request based on the image content description information includes performing semantic analysis on the image content description information, extracting keywords, phrases, or topics from the image content description information, which may be nouns, verbs, adjectives, or words with specific meanings, and weighting or filtering the extracted keywords. For example, weighting the extracted keywords can be based on factors such as the importance of the keywords in the sentence, their frequency of appearance, or their relevance to the user's search history, and filtering out keywords that are unimportant, redundant, or do not match the user's search habits. Through weighting and filtering, the user's main interests and requests can be determined. Natural language processing (NLP) techniques, such as text classification and sentiment analysis, can be used to further analyze the user's intent and request, such as classifying the description information into predefined categories, to more accurately understand the user's intent. Sentiment analysis techniques can be used to determine the user's emotional attitude toward the image content, which can help estimate the user's potential requests or preferences. Entity identification can further identify entities in descriptive information, such as brands, products, and locations, which may be closely related to the user's search intent. Based on the keywords, categories, sentiment trends, and entity information obtained above, the user's intent and requirements are comprehensively analyzed to further determine the user's original search intent, i.e., the type of information the user wants to obtain or the requirements they want to fulfill through their search. Finally, the user's original search intent is converted into specific and actionable search and recommendation requirements, i.e., the first request. The first request can include conditions such as specific search keywords, categories, geographic restrictions, time ranges, and target ranges, allowing the search and recommendation system to find content that meets the user's requirements based on the first request.

[0073] In some embodiments, determining the first request based on the image content description information includes: comprehensively analyzing the user's intention and request in accordance with the user's context, such as historical search records, user images, etc., to further determine the user's original search intention; and converting the user's original search intention into a specific and actionable search and recommendation request, i.e., the first request.

[0074] In this way, by generating image content description information based on a description model suitable for the content type of the image and determining the user's first request based on this, the accuracy of the determined first request can be improved, thereby improving the accuracy of search results and recommended content and improving search efficiency.

[0075] In some embodiments, determining a first request corresponding to the image includes analyzing each object in the image to obtain an intention estimation result corresponding to the image and including the object in the image; generating image content description information for the image based on a description model appropriate for a content type of the image; and determining the first request based on the intention estimation result and the image content description information.

[0076] In some embodiments, the system may use computer vision techniques such as convolutional neural networks, target detection algorithms, etc. to identify and locate each object in the image. The system first performs a detailed analysis of each object in the image, including but not limited to the object's color, texture, layout, spatial relationships, etc.

[0077] In some embodiments, the system further estimates the intent of the image based on the identification of each object in the image. The intent estimation typically includes target objects in the image that are important to understanding the intent of the image. The target object may be a single object, such as a product, an animal, or a building, or a combination of multiple objects or a specific scenario. Intent estimation may also involve analyzing the relationships between objects, the attributes of the objects (e.g., color, size, shape), and how the objects are arranged.

[0078] In some embodiments, intent estimation is performed using a machine learning or deep learning model to obtain one or more possible intent estimation results corresponding to the image.

[0079] In some embodiments, determining the first request based on the intent estimation result and the image content description information includes determining the original search intent in accordance with the intent estimation result and the image content description information, and determining the first request based on the original search intent.

[0080] In some embodiments, the method includes processing the image content description information according to the intent estimation result to convert the image content description information into a search query sentence; and processing the search query sentence based on the intent estimation result to obtain a first request, where the first request can more accurately express the user's original search intent.

[0081] Here, processing the search query statement includes, but is not limited to, adding appropriate modifiers, qualifying terms, or search operators.

[0082] For example, after obtaining the image content description information, the system processes the image content description information in accordance with the previously obtained intention estimation result. A target object in the intention estimation result may be particularly emphasized or focused on in the image content description information to more accurately determine the user's first request. For example, if the intention estimation result indicates that the user may be looking for a specific product, the system may particularly focus on a part related to the product and set it as the first request.

[0083] In this way, by analyzing each object in an image to infer the intent behind the image and then matching the intent estimation results with the image content description information to determine the user's primary request, the accuracy of search and recommendations can be effectively improved. Combining the intent estimation with the image content description information can more accurately reflect the user's actual request, improving the accuracy of search and recommendations and helping to reduce irrelevant or low-quality search results and recommendations. Furthermore, the introduction of intent estimation allows the system to better understand the user's search intent, thereby providing more intelligent and personalized services.

[0084] In an embodiment of the present disclosure, determining the first request based on the intent estimation result and the image content description information includes: determining the original search intent according to the intent estimation result and the image content description information; and determining the first request based on the original search intent.

[0085] In some embodiments, the intent estimation result and the image content description information are subjected to a fusion analysis. For example, if the intent estimation result indicates that the user is likely searching for "travel destinations" and the image description information identifies "beach" and "sunset" scenarios, it can be more accurately estimated that the user's original search intent is searching for "beach resorts." Based on the results of the fusion analysis, the system can determine the user's original search intent and further determine a first request that satisfies that intent. For example, for the search intent of "beach resorts," the first request may include recommendations for specific beach resort hotels, travel plans, travel tips, etc.

[0086] In this way, by determining the original search intent based on the intent estimation result and the image content description information, the system can more comprehensively understand the user's search request, thereby improving the accuracy of the original search intent estimation and ultimately improving the accuracy of the first request. With the accurate original search intent map and first request, the system can be guided to provide more accurate and more closely related search results and recommendation content, thereby improving the accuracy of the search results and recommendation content.

[0087] In an embodiment of the present disclosure, generating search results and recommendation content corresponding to the image based on the first request includes generating search results corresponding to the image based on the first request, and generating recommendation content corresponding to the image based on the first request and the search results, where the search results include search results obtained based on the original search intent.

[0088] In some embodiments, generating search results corresponding to the image based on the first request includes generating search results related to the image based on the original search intent and results of the image content analysis. The search results typically directly respond to the user's original search intent and include content related to the image content, where possible. For example, if a user uploads an image containing a plant, the system can recognize the type of plant in the image and return information about the plant as a search result.

[0089] In some embodiments, concurrently with or after generating the search results, the system may also generate recommendations based on the first request and the search results, where the recommendations may be other resources similar to the search results, related topics or products that may be of interest to the user, personalized recommendations based on the user's behavior or preferences, etc., to provide the user with more relevant information or options to help the user more fully understand or fulfill their request.

[0090] In this way, by generating search results and recommended content corresponding to the image in response to the first request, it is possible to more accurately understand the user's search intent, provide search results and useful recommended content that are closely related to the image content, and improve search efficiency.

[0091] In an embodiment of the present disclosure, generating recommendation content corresponding to an image based on a first request includes determining a request category of the first request, and invoking a recommendation policy corresponding to the request category to generate recommendation content.

[0092] In some embodiments, the first request may include:

[0093] The first class of requirements involves the inability to identify the target body from the image.

[0094] A second class of requirements involves relationships between any at least two bodies in an image, where at least two bodies in the image are target bodies.

[0095] The third class of requirements includes requirements other than acknowledged requirements, such as extensible requirements and consumption requirements.

[0096] In some embodiments, determining the request category of the first request comprises:

[0097] The method includes analyzing a first request from a user and classifying the request as a first-class request when a clear target body cannot be directly recognized from the image, which may be due to the blurred content of the image and the unclear body.

[0098] In some embodiments, determining the request category of the first request comprises:

[0099] Analyzing the user's first request and classifying the request as a second class request if the user's first request concerns a relationship between any at least two bodies in the image and these bodies belong to the same category (i.e., bodies of the same class). For example, a user may want to know the relationship between two people in an image or a comparison between two similar objects.

[0100] In some embodiments, determining the request category of the first request comprises:

[0101] The method includes analyzing the user's first request, and if the first request includes all requests other than those of the above two classes, classifying the request into a third class request, which may be related to information related to but not directly identifiable with the image content, such as background information such as the source, time, and location of the image, or other requests not directly related to the image content.

[0102] In some embodiments, the recommendation policy for the first class of requests includes:

[0103] For the first class of requests, this involves recommending topics or categories that are related to the image content but broader in scope to help users further clarify their search intent.

[0104] For example, general related topics, related search results, or similar images that may be of interest to the user may be recommended.

[0105] In some embodiments, the recommendation policy for the second class of requests includes:

[0106] For the second class of requests, we focus on analyzing the relationships between bodies in images and generating relevant recommendations, which may include interpretations, comparisons, analyses, or other relevant information about the relationships between bodies, as well as recommending other images, articles, videos, or other resources related to these body relationships.

[0107] In some embodiments, the recommendation policy for the third class of requests includes:

[0108] For the third class of requests, the system needs to make personalized recommendations based on the specific request.

[0109] This may involve invoking other modules or services to obtain relevant information such as image recognition, knowledge maps, user images, etc. Recommendations may include background information related to the image content, related knowledge, other content that may be of interest to the user, etc.

[0110] In some embodiments, outputting the recommendation comprises:

[0111] For any request, the generated recommendation content should be displayed to the user in an appropriate manner, such as through a search result page, a pop-up, a push notification, etc. The display method should be concise and clear, and easy for the user to understand and operate.

[0112] In this way, when generating recommendation content corresponding to an image based on a user's first request, by determining the request category and applying the corresponding recommendation policy, the system can provide more accurate and personalized recommendation content based on the user's specific request, thereby improving the user's search experience and satisfaction.

[0113] In an embodiment of the present disclosure, generating search results corresponding to images based on the first request includes, if the first request includes at least two target objects, generating search results based on the at least two target objects.

[0114] When the number of objects in an image is large or the similarity is high, it may be necessary to use more complex algorithms and techniques to accurately identify the target object. For example, a deep learning model can be used to analyze the image in more detail, or to help determine the target object based on the user's past historical search records, contextual information, etc.

[0115] In some embodiments, generating search results based on the at least two target objects comprises:

[0116] determining an appropriate search policy based on the characteristics of the target object and the user's first request;

[0117] performing a search operation using the search policy to obtain search results related to the target object.

[0118] In some embodiments, generating search results based on the at least two target objects further comprises:

[0119] Further processing and filtering the search results obtained to remove duplicate, irrelevant or low quality content;

[0120] If necessary, the method may include at least one of sorting, categorizing, or organizing the search results to facilitate viewing or selection by the user.

[0121] In this way, when the first request involves at least two target objects, more relevant and valuable search results can be generated based on these target objects, improving the precision of the search.

[0122] In some embodiments, the image-based search processing method further includes: in response to receiving a follow-up request sent from the terminal, determining a new first request based on the follow-up request, where the follow-up request is generated by the terminal based on a selection operation on the recommendation content; generating new search results and / or new recommendation content corresponding to the image based on the new first request; and returning the new search results and / or the new recommendation content to the terminal, so that the terminal outputs the new search results and / or the new recommendation content on the search result page.

[0123] In some implementations, the user's response to the recommendations can serve as further feedback of the user's intent, helping the system to more accurately tailor search results or recommendations.

[0124] In some embodiments, receiving a response action includes the system receiving these response actions when the user interacts with the recommendation (e.g., clicks on a link, selects an option, enters text, etc.) These response actions typically contain the user's feedback on the current recommendation and reflect the user's further needs and preferences for the search results.

[0125] In some embodiments, determining the new first request includes the system analyzing the user's response actions and inferring the user's new request or intention based thereon.

[0126] Here, the new first request may be different from the original search intent because it more specifically reflects the immediate needs and interests of the user who viewed the recommended content.

[0127] In some embodiments, the system re-generates search results or recommendations based on the determined new first request, which may involve re-analyzing image content, updating the user's context, adjusting recommendation policies, etc.

[0128] In some embodiments, outputting new search results or recommendations includes the system replacing or supplementing existing search results or recommendations to display the new search results or recommendations on the search results page, thereby enabling users to find information they are truly interested in more quickly and increasing search efficiency and satisfaction.

[0129] In this way, by analyzing user responses and adjusting search results or recommendations accordingly, the system can better understand the user's intent and requirements, providing more intelligent and personalized services. By reducing unnecessary user operations and browsing time, the system can more effectively utilize computing and network resources and improve overall efficiency. Furthermore, using user responses as a direct feedback mechanism can help the system continuously optimize search and recommendation policies, improving the accuracy and adaptability of the algorithms.

[0130] In order to further improve the accuracy of the search, based on any of the above embodiments, the image-based search processing method may further include determining a second request corresponding to the image, where the second request includes the user's potential search intent.

[0131] In some embodiments, generating search results and recommendations corresponding to the image based on the first request includes, if a second request is present, generating search results corresponding to the image based on the first request, including search results obtained based on the original search intent, and generating recommendations corresponding to the image based on the second request, including search results obtained based on the implicit search intent.

[0132] In some embodiments, generating search results and recommendations corresponding to the image based on the first request includes, if a second request is present, generating search results corresponding to the image based on the first request or generating recommendations corresponding to the image in accordance with the first request and the second request, wherein the search results include at least search results obtained based on the original search intent.

[0133] In some embodiments, generating search results and recommendations corresponding to the image based on the first request includes, if a second request is present, generating recommendations corresponding to the image based on the second request, or generating recommendations corresponding to the image in accordance with the first request, the search results, and the second request, wherein the recommendations include at least search results obtained based on the underlying search intent.

[0134] In an embodiment of the present disclosure, the second requirement to reflect the user's potential interests or needs may be different from the first requirement, but is equally important. The recommendation content may include other information, resources, or suggestions related to the user's potential interests or potential search intent, aiming to provide the user with more options and deeper exploration opportunities.

[0135] In the embodiments of the present disclosure, implicit search intent refers to search requests or points of interest that a user may have during a search but have not explicitly expressed. These requests or points of interest may be related to the original search intent and may be broader or deeper. The system can infer implicit search intent by analyzing factors such as user behavior, historical search records, and contextual information. For example, implicit search intent includes, but is not limited to, the following: First, similar products or substitutes. When a user uploads an image of a product, in addition to searching for detailed information about the product, the user may also be interested in products with similar designs, functions, or prices. For example, if a user uploads an image of a smartwatch, the implicit intent may be to search for other brands or models of smartwatches. Second, related parts or accessories. When a user searches for an item, the user may also be interested in parts or accessories related to that item. For example, if a user uploads an image of a camera, the implicit intent may be to search for parts such as camera bags, lenses, and filters. Third, tutorials or techniques. The user may be interested in how to use or maintain the search item. For example, if a user uploads an image of a kitchen appliance, the potential intent may be to find tutorials or repair guides for that appliance. Fourth, style or design inspiration. A user may search for other related design inspiration or styles based on the style or design of the searched item. For example, if a user uploads an image of a modern interior, the potential intent may be to find suggestions for more modern interiors. Fifth, price comparison or special offers. A user may want to know the market price of a searched item or search for special offers. For example, if a user uploads an image of an electronic product, the potential intent may be to find the lowest price or discount information for that product. Sixth, brand or manufacturer information. A user may want to know details about the brand or manufacturer of the searched item.For example, if a user uploads an image of a designer bag, the potential intent may be to find the brand's official website or to find out more about the brand. Seventh, social sharing or discussion. A user may want to share the search item on social media or participate in related discussions. For example, if a user uploads an image of fashion clothing, the potential intent may be to find social media topics or community discussions related to the clothing. The above is merely an illustrative explanation and is not intended as a limitation on all possible content of potential search intents, and a list will not be repeated here.

[0136] In this way, by comprehensively considering the user's first and second requirements, it is possible to not only meet the user's direct requirements, but also take into account the user's latent requirements and interests, thereby providing more comprehensive and personalized search results and recommendations, thereby improving the user's search experience and satisfaction.

[0137] In an embodiment of the present disclosure, generating recommendation content corresponding to an image according to the first request, the search results, and the second request includes: integrating the first request, the search results, and the second request to obtain comprehensive user request information; and generating recommendation content corresponding to the image based on the comprehensive user request information.

[0138] In some embodiments, generating a recommendation corresponding to the image based on the aggregate user request information includes determining a recommendation policy based on the aggregate user request information, and invoking the recommendation policy to generate a recommendation corresponding to the image.

[0139] In some embodiments, generating recommendation content corresponding to the image based on the comprehensive user request information includes adjusting weights of recommendation parameters based on the comprehensive user request information and generating recommendation content corresponding to the image based on the weights of the recommendation parameters. First, the system needs to collect and integrate the user's first request (including the original search intent), second request (including the potential search intent), and possible other related information (e.g., the user's historical search records, behavioral patterns, context information, etc.). These information jointly constitute the comprehensive user request information. Next, the system analyzes the comprehensive user request information to identify the user's interests, preferences, and possible search gaps or omissions. Based on the analysis of the user request information, the system adjusts weights of the recommendation parameters. Recommendation parameters may include content relevance, user interest, timeliness, variety, etc. Different parameters may have different importance for different users or different search scenarios. The recommendation content may include other information related to the image theme, related images, videos, articles, user reviews, etc. The system can also continue to optimize the recommendation content based on user feedback and real-time data.

[0140] In this way, the recommendation content can be generated according to the first requirement, search results, and second requirement, so that the recommendation content can be generated to better meet the user's expectations, provide more comprehensive and personalized services, meet the diversified requirements of users, and improve the user's search experience and satisfaction.

[0141] In an embodiment of the present disclosure, determining a second request corresponding to the image includes: determining a search stage of the original search intent; predicting an implicit search intent based on the search stage of the original search intent; and determining a second request based on the implicit search intent.

[0142] In some embodiments, search stages generally refer to different stages in a user's search process, reflecting different levels of user needs and information requirements. Search stages include, but are not limited to, the awareness stage, information search stage, awareness stage, evaluation stage, and purchase / action stage. During the awareness stage, a user recognizes that they have a need or problem, but may not be clear on what it is specifically. During the information search stage, a user begins to actively search for information related to their need. During the evaluation stage, a user collects and evaluates information from different sources to make a decision. During the purchase / action stage, a user takes an actual action, such as purchasing, downloading, or contacting, based on the information provided to the user by the system. By analyzing the image entered by the user and possible contextual information (e.g., search history, browsing behavior, etc.), it is possible to determine which search stage a user is currently in.

[0143] In some embodiments, predicting potential search intent based on the search phase of the original search intent includes predicting the user's potential search intent after determining the search phase the user is in. This typically involves inferring information the user may be interested in or need in the current phase. For example, if the user is in the information search phase, they may be looking for more information about a product, comparing different products, researching prices, etc. Based on these potential requests, the user's potential search intent can be predicted.

[0144] In some embodiments, determining a second request based on the implicit search intent includes determining a second request of the user based on the predicted implicit search intent. These second requests are typically related to the original search intent but are more specific or detailed. For example, if the original search intent is looking for a camera and the user is in an information search phase, the second request may be to understand the detailed specifications of the camera, view user reviews, or compare different camera models.

[0145] In this way, by determining the search stage of the original search intent, predicting the latent search intent, and finally determining the second request, more accurate and personalized search results and recommendation content can be provided, and more accurate prediction and recommendation can be achieved.

[0146] In an embodiment of the present disclosure, determining a second request corresponding to the image includes determining a focus corresponding to the image based on the first request, obtaining relevant content corresponding to the first request based on the focus, and determining a second request based on the relevant content.

[0147] In some embodiments, determining points of interest corresponding to the image based on the first request includes identifying key elements in the image, such as objects, scenarios, or activities, based on the original search intent, and determining primary points of interest based on the identified key elements. For example, if a user uploads a landscape image, the system can identify elements in the image, such as mountains, lakes, and trees, and determine that the user's points of interest are likely to be natural landscapes or tourist attractions.

[0148] In some embodiments, obtaining relevant content responsive to the first request based on the points of interest includes searching for text, images, videos, or other media content that matches the points of interest using a search engine, database, or online resource. For example, if the points of interest are natural scenery, the system may search for content such as travel guides, photography, historical information, etc. related to mountain ranges, lakes, trees, etc.

[0149] In some embodiments, determining the second request based on the relevant content includes, after obtaining content related to the original search intent, the system can begin analyzing the content to infer the user's potential search intent. This typically involves semantic content analysis, thematic modeling, or machine learning techniques to identify potential points of interest or requests in the content. For example, after analyzing content related to natural scenery, the system may discover that the user is interested in other similar tourist spots, or requests for sightseeing itineraries, photography techniques, etc. These can be considered potential search intents.

[0150] In some embodiments, the system can also utilize data analysis and machine learning techniques to automatically optimize the secondary request determination process. By analyzing large amounts of sample data, the system can learn how to more accurately identify users' potential requests and continuously improve the quality of searches and recommendations.

[0151] In this way, determining the secondary requirement based on the primary requirement is a reasoning process from concrete to abstract, from explicit to implicit. By determining the secondary requirement based on relevant content, the quality of the secondary requirement can be improved, thereby providing more accurate and personalized search results and recommendation content, helping to meet the diverse needs of users, increase the accuracy and individualization of searches, and improve the search experience. Furthermore, the system can allocate resources more intelligently, such as by prioritizing the processing of search requests or recommendation content related to the secondary requirement. At the same time, the system can also adjust recommendation policies according to the secondary requirement to improve the accuracy and efficiency of recommendations.

[0152] In an embodiment of the present disclosure, an image-based search processing method includes: in response to receiving a follow-up request sent from a terminal, determining a new second request based on the follow-up request generated by the terminal based on a selection operation for the recommended content; generating new search results and / or new recommended content corresponding to the image based on the new second request; and returning the new search results and / or new recommended content to the terminal, so that the terminal outputs the new search results and / or new recommended content on a search result page.

[0153] In some embodiments, selection actions on the recommended content include when a user interacts with the recommended content, such as clicking, viewing, commenting on, or sharing a recommended item, and the system receives these response actions as user feedback.

[0154] In some embodiments, determining a new secondary request based on the follow-up request includes analyzing the user's response actions to understand what types of recommended content the user is more interested in, which content is being ignored, and whether the user's behavioral patterns are changing. Based on these analyses, the system attempts to determine the user's new potential search intent or focus, i.e., the new secondary request.

[0155] In some embodiments, generating new search results corresponding to the image based on the new second request includes deeply analyzing the new second request to understand the intent and expectations behind it, converting the new second request into actionable search or analysis parameters, and performing a search based on the actionable search or analysis parameters, and applying a recommendation algorithm to filter or generate content related to the image as the new search results. In practical applications, collaborative filtering, content filtering, or a mixed search policy may be considered. Also, the search results may be filtered based on the new second request to eliminate irrelevant or low-quality content.

[0156] In some embodiments, generating new recommendations corresponding to the image based on the new second requirement includes reevaluating the recommendation policy after determining the new second requirement and adjusting parameters of the recommendation policy based on the new requirement, and generating new recommendations related to the image based on the new recommendation policy, which may more accurately reflect the user's current interests and requirements.

[0157] In this way, the recommendation policy is dynamically adjusted and new recommendations are generated based on the user's response to the recommendations, enhancing the intelligence and personalization of the search and recommendation system. By dynamically adjusting the recommendation policy in response to the user's real-time feedback, the system can adapt more quickly to user changes and provide more accurate and timely recommendations. Users realize that the system "understands" their interests and requirements and can provide valuable recommendations during their searches, significantly enhancing their search experience. Personalized recommendations can attract users' attention, increase their interaction with the system, and improve user engagement and stickiness. The system can adjust resource allocation based on the user's real-time feedback, using more resources to generate recommendations that interest the user and increasing resource utilization efficiency. By constantly learning and adapting to user behavior, the system can continuously improve its intelligence and accuracy, providing users with better service.

[0158] In some embodiments, generating a recommendation corresponding to the image based on the second request includes matching a target scenario to the second request and providing function entries that match the target scenario, wherein the recommendation includes the function entries.

[0159] In some embodiments, matching the target scenario to the second request comprises:

[0160] After determining the user's second request, the system searches for a goal scenario that best matches the second request. These application scenarios may cover a variety of potential user requests and usage scenarios, such as shopping, travel, learning, entertainment, etc. The system uses an algorithm or machine learning model to match the most relevant goal scenario based on factors such as the nature of the second request, keywords, historical data, and the user's behavioral patterns.

[0161] In some embodiments, after determining the target scenario, the system provides function entries that match the target scenario. These function entries are interface elements that the user can click or manipulate directly to further explore or fulfill their requirements. The design of the function entries should suit the characteristics of the target scenario and the user's requirements. For example, if the target scenario is shopping, the function entries may include a product list, a search box, a shopping cart icon, etc. If the target scenario is sightseeing, the function entries may include tourist attraction recommendations, schedule planning, hotel reservations, etc.

[0162] In this way, by providing function entries that match the target scenario, the system can more directly meet the user's requirements, reduce the user's search and browsing time, and improve search efficiency.When the system can accurately respond to the user's requirements and provide corresponding function entries, the user can experience the intelligence and convenience of the system, and increase their satisfaction with the system.

[0163] In an embodiment of the present disclosure, an image-based search processing method includes: when the first request does not satisfy a recommendation trigger condition, generating search results corresponding to the image based on the first request; and returning the search results to the terminal, so that the terminal outputs the search results on a search result page.

[0164] In this way, if the first request does not satisfy the recommendation trigger condition, search results corresponding to the image are generated based on the first request and these search results are output on the search result page. Even if personalized recommendations cannot be provided, the user can still obtain basic search results related to the query, ensuring a basic search experience.

[0165] In an embodiment of the present disclosure, before determining the first request corresponding to the image, the image-based search processing method may further include: performing a risk control check on the image; if the image meets the risk control criteria, entering a flow of determining the first request corresponding to the image; and if the image does not meet the risk control criteria, returning prompt information to the terminal to indicate that the image is invalid.

[0166] In some embodiments, risk control checks can include multiple aspects such as image content recognition, classification, keyword filtering, copyright checks, and the like.

[0167] In some embodiments, the risk control criteria may be preset by the system or may be dynamically adjusted based on laws and regulations, user agreements, platform policies, etc. These criteria are typically used to determine whether an image contains objectionable content, infringes the rights and interests of others, conforms to platform regulations, etc.

[0168] In some embodiments, if the image meets the risk control criteria, the system enters a flow to determine a first request corresponding to the image and generates appropriate search results or recommendations for the user based on the user's query intent and the image content. If the image does not meet the risk control criteria, the system will not enter into the subsequent search process and will notify the user that the image is invalid. This notification can be displayed as text, a pop-up, a sound, or the like to inform the user that the image cannot be processed or that it poses a security risk.

[0169] In this way, by conducting risk control inspections on images, potential bad content and risks can be identified and filtered out early, ensuring that users do not come into contact with undesirable information or suffer losses when using search services, thereby improving search safety.

[0170] In an embodiment of the present disclosure, before determining the first request corresponding to the image, the image-based search processing method further includes: performing code detection on the image; if the image meets the code detection criteria, the flow of determining the first request corresponding to the image is not entered, but the code detection result of the image is obtained through the code detection model, and the code detection result is returned to the terminal.

[0171] In some embodiments, before determining the first request for the image, a code detection step is introduced to identify and process images that may contain special codes such as two-dimensional codes, barcodes, etc. If the image contains such codes, the system may directly analyze these codes to obtain relevant information, rather than further determining the first request corresponding to the image.

[0172] In some embodiments, if the image meets the code detection criteria, i.e., if the image contains a clear, complete, and analyzable code, the system does not continue the process of determining the first request corresponding to the image. The system analyzes the code in the image through a code detection model (e.g., a two-dimensional code identification algorithm, a barcode identification algorithm, etc.) to obtain a code detection result. This result may be a web address, a string of characters, an identifier, etc., depending on the information stored in the code. Based on the code detection result, the system can take different subsequent actions. For example, if the decoding result is a web address, the system can directly guide the user to access this web address; if the decoding result is an identifier, the system can search for information related to the identifier in an internal database and present it to the user.

[0173] In this way, for images containing analyzable code, the system can directly parse the code to obtain relevant information without the need for further processing of the image content, significantly improving processing efficiency and reducing unnecessary consumption of computing resources.

[0174] An embodiment of the present disclosure provides an image-based search processing method, which is applicable to a terminal equipped with a search application client having an image search function. In practical application, the terminal may include, but is not limited to, devices such as a mobile phone, a tablet, a wearable device, or a personal computer. As shown in Figure 2, the image-based search processing method includes the following steps:

[0175] In S201, in response to receiving an image input by a user via a search application, the image is transmitted to a server.

[0176] At S202, the search results and recommendation content returned from the server are received, where the search results are determined when the image satisfies the recommendation trigger condition, and the search results are determined based on a first request corresponding to the image, and the first request includes the user's original search intent.

[0177] In S203, the search results and the recommended contents are displayed on the search result page.

[0178] In an embodiment of the present disclosure, to increase the variety of recommendation content, the recommendation content may include content related to but different from the image.

[0179] In this way, after the device receives the image entered by the user through the search application, the search application outputs search results and recommendations on its search result page during the first round of replies. By combining the image search and recommendation functions, the user can be provided with more abundant search results and a more comprehensive presentation of information, greatly improving search efficiency. By viewing the recommendations, the user can discover more content related to their interests and enjoy a better search experience. Furthermore, the presentation of recommendations can also increase the user's click rate and conversion rate, further enhancing the value of the search application.

[0180] In some embodiments, recommendations are generated based on the first request and the search results.

[0181] In some embodiments, recommendations are generated based on the user's initial request (i.e., original search intent) and actual search results, increasing the likelihood of satisfying the user's actual request. Recommendations that are relevant to a user's interests, preferences, or high search goals can improve user satisfaction and overall search experience, and increase stickiness between the user and the platform.

[0182] In this way, by generating recommendation content based on the first request and search results, the accuracy and relevance of recommendations can be improved, the intelligence and adaptability of the system can be enhanced, and the ever-changing needs of users can be better met.

[0183] In some implementations, the recommendations are generated based on the second request, and the recommendations include search results obtained based on the implicit search intent.

[0184] In some embodiments, by identifying and analyzing users' potential search intent, the system can predict users' potential and extended requests in advance and provide relevant recommendations, thereby significantly saving users time from re-searching and improving search efficiency and users' search experience.

[0185] In this way, recommended content is generated based on the user's latent search intent, and the recommended content can be highly relevant to the user's latent requirements, ensuring that the user can directly access content that may be of interest to the user but has not been explicitly specified or searched for, thereby increasing the relevance and accuracy of the recommended content.

[0186] In some embodiments, the recommendations are generated in accordance with a first request, search results, and a second request, the first request including the user's implicit search intent, and the second request including the user's original search intent.

[0187] In some embodiments, the system can accurately capture the user's immediate needs and recommend content that is highly relevant to the needs by combining the user's original search intent and the search results. At the same time, the system can predict the user's future needs by considering the user's potential search intent, further improving the accuracy of recommendations.

[0188] In this way, by generating recommendation content according to the user's first request, search results, and second request, more accurate and relevant recommendation content can be provided, and a personalized recommendation service can be provided, thereby further improving the user's search experience and satisfaction.

[0189] In some embodiments, the image-based search processing method includes, in response to receiving a selection operation for the recommendation content, sending a follow-up request to a server corresponding to the selection operation; receiving new search results and / or new recommendation content obtained by the server based on the follow-up request; and displaying the new search results and / or new recommendation content on a search result page.

[0190] In some embodiments, the front-end interface of the search application listens for user selections on recommendations. The selections may be clicks, touches, slides, or other forms of user input. When a user selection is detected, necessary information related to the selection (e.g., the identity document (ID) of the selected recommendation, the user ID, a timestamp, etc.) is collected, and the server performs appropriate logical operations to obtain new search results, recommendations, etc., in response to the user selection.

[0191] In this way, search results and recommendations can be dynamically adjusted based on the user's real-time selections and actions regarding recommendations, providing the user with more personalized and relevant content.

[0192] In some embodiments, the image-based search processing method includes: in response to receiving a selection operation on the recommended content, displaying text information corresponding to the selection operation in a corresponding input box on the search result page; in response to receiving an editing operation on the text information, displaying the edited text information in the input box; using the edited text information as a response result to the recommended content; and generating a follow-up question request based on the response result.

[0193] In some embodiments, after viewing the recommendations on the search results page, the user can select the recommendations of interest by clicking or using other interactive methods. When the user selects a recommendation, the client automatically enters text information about the recommendation into a search box, which the user can view or further edit. The user can directly edit the text information displayed in the search box, for example, by changing, adding, or deleting some of the text, to meet the user's actual needs. After completing the editing, the user can click a submit button or use other interactive methods to send the edited text information to the system as a response to the recommendation. After receiving the response sent by the user, the system can perform further processing, such as improving the recommendation algorithm, recording user preferences, and performing operations requested by the user.

[0194] In this way, allowing the user to select and edit the text information of the recommendation content and send it as a response result not only improves the user's operational efficiency, but also helps the system understand the user's more specific search intent, thereby providing more accurate search results or recommendation content.

[0195] In some embodiments, before outputting the search results and the recommendation content, the image-based search processing method may further include not outputting the recommendation content in response to receiving an operation to turn off the recommendation function.

[0196] In some embodiments, the system provides a user with an option or settings entry to turn off the recommendations feature before the search results page is displayed or while the user is performing a search operation. This can be done by a switch, checkbox, or similar widget. If the user chooses to turn off the recommendations feature, this setting information is stored in the user's session data, browser data, or user preferences so that it can be identified in subsequent requests. When the user initiates a search request, the system checks the user's recommendations setting. If the recommendation setting is activated, recommendations are not generated or included during the processing of the search request, and only search results are generated based on the user's search request and search algorithms. If recommendations are not turned off, additional recommendations are generated. If recommendations are turned off, only search results are generated, and the generated search results (without recommendations and if the user has turned off the recommendations feature) are returned to the user and displayed on the front row page.

[0197] In this way, before the search results and recommendation content are output, the user can turn off the recommendation function, and this requirement can be met by not outputting the recommendation content, and the user can turn off the recommendation function, so that the personalized requirements of different users can be met. For users who do not need or do not like the recommendation content, the search result page will be more concise and direct, which can improve search efficiency and satisfaction.

[0198] In some embodiments, after outputting the search results and the recommendation content, the image-based search processing method may further include deleting or hiding the recommendation content on the search result page in response to receiving an operation to turn off the recommendation function.

[0199] In some embodiments, a switch or button may be provided at an appropriate position on the search result page (such as the top or bottom of the page or in a sidebar) to control the display or hiding of the recommendations. When a user clicks the switch or button, an event processor is activated. This processor is responsible for receiving the user's operation instruction and determining whether the user has chosen to turn off the recommendations function. If the user determines that they want to turn off the recommendations function, the event processor sends a request to a back-end server or a front-end data processing module to remove or hide the recommendations on the search result page. Based on the received command, the front-end page updates the layout and content of the search result page in real time to ensure that the recommendations are no longer displayed to the user.

[0200] In this way, after outputting the search results and recommendation content, the system supports the user to turn off the recommendation function, and after receiving the user's operation to turn off the recommendation function, the system can accordingly delete or hide the recommendation content on the search result page. In this way, the system can dynamically adjust its functions and interface design, adapt to the needs of different user groups, and reflect the flexibility of the system.

[0201] As shown in Figure 3, the flow chart based on image retrieval is divided into three parts: image judgment and processing, first-round response, and multiple-round response.

[0202] In S301, the user uploads an image, and then S302 is executed.

[0203] In S302, it is determined whether the image satisfies the risk control criteria, and if so, S303 is executed, while if not, S304 is executed.

[0204] In S303, the user is prompted to retake or replace the image.

[0205] In S304, intent determination and policy distribution are performed. Specifically, in S304a, for a cognitive request, a policy for searching short entities is adopted, where a short entity usually refers to a short and specific information segment that can meet the user's cognitive request. In S304b, for an image search request, a policy for searching similar images is adopted. In S304c, for a text request, translation request, or problem-solving request, an OCR identification policy is adopted. In S304d, for an unknown request, a generative policy is adopted.

[0206] In S305, the model obtains the contents of the intention determination field, OCR field, word guess field, generated description field, etc. generated based on the image, and then executes S306.

[0207] In S306, it is determined whether the image belongs to the range of post-search recommendation, and if it belongs to the range of post-search recommendation and the method of post-search recommendation is reverse question, S307 is executed, if it belongs to the range of post-search recommendation and the method of post-search recommendation is recommendation, S308 is executed, and if it does not belong to the range of post-search recommendation, S309 is executed.

[0208] In S307, the image recognition result and the inverse question are output, and then S310 is executed.

[0209] In S308, the image recognition result and existing function recommendations are output, and then S310 is executed.

[0210] In S309, the image recognition result is output, and S311 is executed.

[0211] In S310, it is determined whether or not a reverse question or recommendation function has been clicked. If not, the flow ends, and if so, S312 is executed.

[0212] In S311, it is determined whether the user has started the second round of follow-up questions. If not, the flow ends; if yes, S314 is executed.

[0213] In S312, the result of the follow-up question is determined based on the recommendation function or the content of the reverse question, and then S313 is executed.

[0214] In S313, the result of the follow-up question is returned.

[0215] In S314, multiple rounds of follow-up questions are entered, and then S315 is executed.

[0216] In S315, the result of the follow-up question is returned.

[0217] The interaction process based on image search can include four steps.

[0218] In step 1, take / select a photo and send it.

[0219] FIG. 4 is a schematic diagram showing the interface after triggering a camera entry in the search application. As shown in FIG. 4, the bottom of the interface displays, from left to right, a history search button, a photo button, and an album button. If the user does not select the history search button or the album button within a preset time, the interface automatically takes a photo and calls a model to identify the automatically taken image. If the user clicks the photo button to take a photo, the interface calls a model to identify the image taken by the user. If the user selects the album button, the album images are loaded and the model is called to identify the image selected by the user. If the user clicks the history search button, the interface calls a history search record, allowing the user to view the search history.

[0220] In step 2, describe the image, display the recommendation content, complete the recognition of the main requirements, and provide extensions and guided entries.

[0221] The model understands image information based on image recognition and image search capabilities, helps users complete the main question of "what is the image?", and at the same time makes judgments based on post-search recommendation policies and displays recommended content.

[0222] In step 3, based on the user's click selection operation on the recommended content, multiple rounds of dialogue are entered to complete the extended recognition or requirement clarification.

[0223] When a user clicks on a recommended feature, a search is initiated using the current image as a query for the selected feature, and search results are obtained.

[0224] When a user clicks on a reverse question, the model will identify the content of the image intent, OCR field, word guessing field, etc., and provide different levels of recommended reverse questions for different intents. After clicking, the user can ask a reverse question in the input box at the bottom to further refine their requirements, or start a question directly, for example, "I want to know more about [famous tourist spot]" or "I want to know more about [famous tourist spot; indoors]."

[0225] Figure 5 is a schematic diagram showing the search results and recommendations based on an image search. As shown in Figure 5, after a user uploads an image, the search result displayed on the search results page is, "This looks like an image of Cocotohai. Is there anything I can help you with?" The recommendations displayed on the search results page include a recommendation function called "Picture to Text," and include guidance information such as, "Please provide more information so we can provide a better answer. Want to know more? 'Specific locations,' 'Famous tourist spots,' 'History and culture,' 'Local specialties.'" The search results page also displays a collection of related images.

[0226] Figure 6 is a second schematic diagram showing the search results and recommendations based on image search. As shown in Figure 6, after a user uploads an image, the search result displayed on the search result page is "This looks like an image of Cocotohai. Is there anything I can help you with?" The recommendation displayed on the search result page includes guidance information, such as "Please tell us more information so we can provide a better answer. Want to know more: 'Specific locations,' 'Famous tourist spots,' 'History and culture,' 'Local specialties.'" The search result page also displays a collection of related images.

[0227] As shown in Figure 6, after the user clicks on "famous tourist spots," a prompt message appears at the bottom of the search results page: "If you would like to know more about [famous tourist spots], you can continue with your request." The input box displays the text message "I would like to know more about [famous tourist spots]." When the user clicks the submit button, the search results page displays the tourist spot photos recommended for the user and the encyclopedia information of Cocotohai, as shown in Figure 7.

[0228] As shown in Figure 6, after the user clicks on "famous tourist spots", the prompt information "If you want to know more about [famous tourist spots], you can continue with your request" will be displayed at the bottom of the search result page. The input box displays the text information "I want to know more about [famous tourist spots]", the user edits the information in the input box, adds the two characters "indoor" to the input box, and finally the input box displays the result "I want to know more about [famous tourist spots; indoor]", and the user clicks the submit button. As shown in Figure 8, the search result page displays the search results for indoor tourist spot information in Cocotohai.

[0229] The schematic diagram of the technical framework for implementing search can be divided into four parts as shown in Figure 9.

[0230] The first part performs filtering and intention determination after the user sends an image.

[0231] After a user uploads an image, for the sake of accuracy in policy recognition and safety in risk control, the image is first filtered using fuzzy graphs. To better generate image descriptions, code detection technology and intent determination technology are used to distinguish image categories (including types such as 2D codes, barcodes, text, themes, people, plants, materials, human faces, animals, facial expressions, products, and more), which helps generate image content description information for different types of images in the next step.

[0232] In the second part, image content description information is generated for each intended image.

[0233] Depending on the image type, different paths are taken to generate the image content description information.

[0234] For 2D code and barcode type images, code identification technology is used to identify the 2D code link or barcode information (e.g., the 2D code link is xxxx), and for text and theme images, OCR technology is used to identify the text information in the image (e.g., the text in the image is xxxx).

[0235] For human and plant type images, the word prediction model is used preferentially to generate word prediction results (e.g., this is xxx), and if no word prediction results are generated, the multimodal large language model is used to generate image description results.

[0236] For images of materials, human faces, animals, facial expressions, products, and other types, which are the dominant scenarios for multimodal large language models, we use this model to generate image description results (e.g., this is a photo of a tire pressure alarm on the dashboard of a xx car).

[0237] In order to ensure the safety of the risk control of the output content, the image content description information must be checked for the presence of sensitive word content through a sensitive word filtering model, and after passing this test, the image description field will be generated and the first round of description results will be presented to the user.

[0238] The third part determines the post-search recommendation policy.

[0239] Based on the contents of the intent determination field, OCR field, word guessing field, generated description field, etc. generated by the model based on the image, it determines whether the image belongs to the scope of post-search recommendations and whether the policy is valid as a reverse query or a recommendation. If the policy is valid, the first round of search identification results and post-search recommendations are displayed on the search result page.

[0240] In the fourth part, after the user asks a question, the context is transferred to a multimodal large language model to generate an answer and ask follow-up questions in multiple rounds.

[0241] After the user sends an image and the multimodal large language model generates an image description, a multi-round dialogue process is entered, in which the dialogue context is sent to the multimodal large language model, the multimodal large language model generates and returns results, and the user continuously asks follow-up questions based on the above content, until the request is met and the dialogue is terminated.

[0242] As described above, the present invention has at least the following effects.

[0243] It provides sophisticated guidance for unclear requests. Users upload images of unclear types through photography or photo albums, and after the description information of the image content is generated, the system guides users to clarify the entity they want to inquire about through the method of reverse query for multi-body images. After users click and select, they can enter multi-round based on the image, and they do not need to submit the image again.

[0244] The idea is to add new functions and actively match them with application scenarios. An adaptive function is embedded in the post-search recommendation module. For example, after a user uploads a product image and provides a description based on product identification, the post-search recommendation module is embedded with a function to transcribe the image (such as advertising copy, recommendation copy, online shopping review copy, etc.), actively providing application scenarios for the new functions while the user is using the product.

[0245] Proactively respond to scalable and continuously consumable search. Satisfy users' extended needs, for example, by providing a product price comparison function for product intent.

[0246] An embodiment of the present disclosure provides an image-based search processing device applied to a server, as shown in FIG. 10 , which includes: a first determination module 1001 for determining a first request corresponding to the image in response to receiving an image sent from a terminal, where the first request includes a user's original search intent; a generation module 1002 for generating search results and recommendation content corresponding to the image based on the first request when the first request satisfies a recommendation trigger condition; and a first communication module 1003 for returning the search results and recommendation content to the terminal, so that the terminal outputs the search results and recommendation content on a search result page.

[0247] In some embodiments, the recommendation content includes at least one of a prompt content, a counter-question content, and a function entry, wherein:

[0248] The guidance content includes a first type of information for guiding the user to clarify search intent;

[0249] The reverse question includes a second type of information for the user to explore to clarify the search intent,

[0250] The function entry indicates an entry to available functions that are recommended to the user.

[0251] In some embodiments, the first determination module 1001:

[0252] analyzing each object in the image to obtain an intention estimation result corresponding to the image, the intention estimation result including a target object in the image;

[0253] generating image content description information for the image based on a description model appropriate for the content type of the image;

[0254] and determining a first requirement based on the intention estimation result and the image content description information.

[0255] In some embodiments, the first determination module 1001 is used to determine an original search intent according to the intent estimation result and the image content description information, and to determine a first request based on the original search intent.

[0256] In some embodiments, the generating module 1002 includes:

[0257] a first generating sub-module for generating search results corresponding to the image based on the first request, the search results including search results obtained based on the original search intent;

[0258] and a second generating sub-module for generating recommendation content corresponding to the image based on the first request and the search results.

[0259] In some embodiments, the first determination module 1001 is further used to, in response to receiving a follow-up request sent from the terminal, determine a new first request based on the follow-up request, where the follow-up request is generated by the terminal based on a selection operation for the recommended content; the generation module 1002 is further used to generate new search results and / or new recommended content corresponding to the image based on the new first request; and the first communication module 1003 is further used to return the new search results and / or new recommended content to the terminal, so that the terminal outputs the new search results and / or new recommended content on the search result page.

[0260] In some embodiments, the image-based search processor comprises:

[0261] The system further comprises a second determination module 1004 (not shown in FIG. 10) for determining a second request corresponding to the image and including the user's potential search intent.

[0262] In some embodiments, the first generation submodule is used for generating search results corresponding to images based on a first request, or generating recommendation content corresponding to images in accordance with the first request and the second request, where the search results include at least search results obtained based on the original search intent; and the second generation submodule is used for generating recommendation content corresponding to images based on the second request, or generating recommendation content corresponding to images in accordance with the first request, the search results, and the second request, where the recommendation content includes at least search results obtained based on the underlying search intent.

[0263] In some embodiments, the second determination module 1004 (not shown in FIG. 10)

[0264] Once you have determined the search stage of the original search intent,

[0265] Predicting potential search intent based on the search stage of the original search intent;

[0266] and determining a second request based on the underlying search intent.

[0267] In some embodiments, the second determination module 1004 (not shown in FIG. 10)

[0268] determining a point of interest corresponding to the image based on the first request;

[0269] Obtaining relevant content corresponding to the first request based on the focus; and

[0270] and determining a second request based on the relevant content.

[0271] In some embodiments, the second determination module 1004 (not shown in FIG. 10 ) is further used to, in response to receiving a follow-up request sent from the terminal, determine a new second request based on the follow-up request, where the follow-up request is generated by the terminal based on a selection operation for the recommended content; the generation module 1002 is further used to generate new search results and / or new recommended content corresponding to the image based on the new second request; and accordingly, the first communication module 1003 is further used to return the new search results and / or new recommended content to the terminal, so that the terminal outputs the new search results and / or new recommended content on the search result page.

[0272] In some embodiments, the second generation submodule is further used to match a target scenario to the second request and provide function entries that match the target scenario, wherein the recommendation content includes the function entries.

[0273] In some embodiments, the generation module 1002 is further used for generating search results corresponding to images based on the first request when the first request does not satisfy the recommendation trigger condition, and the first communication module 1003 is further used for returning the search results to the terminal, so that the terminal outputs the search results on a search result page.

[0274] In some embodiments, the image-based search processor further comprises:

[0275] The system is provided with a first detection module 1005 (not shown in FIG. 10) for performing a risk control check on the image before determining the first request corresponding to the image, and entering a flow for determining the first request corresponding to the image if the image meets the risk control criteria, while returning prompt information to the terminal to indicate that the image is invalid if the image does not meet the risk control criteria.

[0276] In some embodiments, the image-based search processor further comprises:

[0277] The device further includes a second detection module 1006 (not shown in FIG. 10) for performing code detection on the image before determining the first request corresponding to the image, and if the image meets the code detection criteria, obtaining the code detection result for the image through the code detection model without entering the flow for determining the first request corresponding to the image, and returning the code detection result to the terminal.

[0278] The functions of each processing module in the image-based search processing device applied to the server according to the embodiments of the present disclosure can be understood by referring to the relevant description of the image-based search processing method applied to the server described above, and it should be understood by those skilled in the art that each processing module in the image-based search processing device according to the embodiments of the present disclosure can be realized by a generation circuit that realizes the function according to the embodiments of the present disclosure, and can be realized by the operation of software that executes the function according to the embodiments of the present disclosure on an electronic device.

[0279] According to the image-based search processing device according to the embodiment of the present disclosure, it is possible to improve search efficiency, and enhance the intelligence and convenience of the search.

[0280] An embodiment of the present disclosure provides an image-based search processing device applied to a terminal, as shown in FIG. 11 , in which a second communication module 1101 is used for sending an image to a server in response to receiving an image input by a user through a search application, and receiving search results and recommendation contents returned from the server, where the search results are determined when the image satisfies a recommendation trigger condition, and the search results are determined based on a first request corresponding to the image, and the first request includes the user's original search intent; and an output control module 1102 is used for displaying the search results and recommendation contents on a search result page.

[0281] In some embodiments, recommendations are generated based on the first request and the search results.

[0282] In some implementations, the recommendations are generated based on a second request that includes the user's implicit search intent and include search results obtained based on the implicit search intent.

[0283] In some embodiments, a search result is generated in accordance with the first request, the search results, and a second request, the second request including the user's implicit search intent.

[0284] In some embodiments, the second communication module 1101 is further used for sending a follow-up request to the server in response to receiving a selection operation for the recommended content, and for receiving new search results and / or new recommended content obtained by the server based on the follow-up request, and the output control module 1102 is for displaying the new search results and / or new recommended content on the search result page.

[0285] In some embodiments, the second communication module 1101 is further used for, in response to receiving a selection operation on the recommended content, displaying text information corresponding to the selection operation in a corresponding input box on the search result page; and, in response to receiving an editing operation on the text information, the output control module 1102 is used for, in response to receiving an editing operation on the text information, displaying the edited text information in the input box, using the edited text information as a response result to the recommended content, and generating a follow-up question request based on the response result.

[0286] In some embodiments, the output control module 1102 is used to not output the recommendation content in response to receiving an action to turn off the recommendation function before outputting the search results and recommendation content.

[0287] In some embodiments, the output control module 1102 is further used to delete or hide the recommendation content on the search result page after outputting the search results and recommendation content in response to receiving an operation to turn off the recommendation function.

[0288] The functions of each processing module in the image-based search processing device applied to the terminal according to the embodiments of the present disclosure can be understood by referring to the relevant description of the image-based search processing method applied to the server described above, and it should be understood by those skilled in the art that each processing module in the image-based search processing device according to the embodiments of the present disclosure can be realized by a generation circuit that realizes the function according to the embodiments of the present disclosure, and can be realized by the operation of software that executes the function according to the embodiments of the present disclosure on an electronic device.

[0289] According to an embodiment of the image-based search processing device of the present invention, when performing an image-based search, a method of outputting search results and recommendation content in the first round of reply is adopted. If the search results do not meet the user's requirements, the user not only does not need to re-take and upload the image, but also does not need to spend more time and energy expressing the requirements. The user can quickly clarify the search requirements through the recommendation content, thereby improving search efficiency and improving the intelligence and convenience of the search.

[0290] For a description of specific functions and examples of the modules and sub-modules of the apparatus according to the embodiments of the present disclosure, please refer to the corresponding related descriptions in the above-mentioned method embodiments, and they will not be repeated here.

[0291] An embodiment of the present disclosure provides a schematic diagram illustrating a scenario of an image-based search processing method, as shown in Figure 12. As described above, the image-based search processing method provided by an embodiment of the present disclosure is applied to an electronic device.

[0292] Specifically, the electronic device

[0293] In response to receiving the image transmitted from the terminal, determine a first request and a second request corresponding to the image, where the first request includes the user's original search intent and the second request includes the user's latent search intent;

[0294] generating search results corresponding to the image based on the first request, and generating recommendation content corresponding to the image based on the second request;

[0295] The operation may specifically include returning the search results and the recommendation content to the terminal, so that the terminal outputs the search results and the recommendation content on a search result page.

[0296] The scenario diagram shown in FIG. 12 is merely illustrative and not limiting, and those skilled in the art can make various obvious changes and / or substitutions based on the example of FIG. 12, and the obtained solution still falls within the disclosure scope of the embodiments of the present disclosure.

[0297] In the technical solution of the present disclosure, the acquisition, storage, and application of users' personal information comply with the provisions of relevant laws and regulations and do not violate public order and morals.

[0298] According to embodiments of the present disclosure, the present disclosure further provides an electronic device, a non-transitory computer-readable storage medium, and a program product.

[0299] 13 is a block diagram of an electronic device 1300 for implementing an embodiment of the present disclosure. The electronic device refers to various types of digital computers, including, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device also refers to various types of mobile devices, including, for example, personal digital assistants, cellular phones, intelligent phones, wearable devices, and other similar computing devices. The components, their connections, and functions described in this disclosure are merely exemplary and do not limit the implementation of what is described and specified in this disclosure.

[0300] 13, device 1300 includes a computing unit 1301 that can perform various appropriate operations and processes based on computer program instructions stored in a read-only memory (ROM) 1302 or loaded from a storage unit 1308 into a random access memory (RAM) 1303. The RAM 1303 can further store various programs and data required for the operation of device 1300. The computing unit 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0301] The components of device 1300 are connected to an I / O interface 1305, which includes an input unit 1306 such as a keyboard or mouse, an output unit 1307 such as various displays and speakers, a storage unit 1308 such as a magnetic disk or optical disk, and a communication unit 1309 such as a network card, modem, wireless communication transceiver, etc. The communication unit 1309 allows device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various carrier networks.

[0302] The computing unit 1301 may be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units that execute various machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1301 performs each of the methods and processes described above, such as the thread stripping method. For example, in some embodiments, the thread stripping method may be implemented as a computer software program tangibly embodied in a machine-readable medium such as the storage unit 1308. In some examples, some or all of the computer program may be loaded and / or installed into the device 1300 via the ROM 1302 and / or the communication unit 1309. When the computer program is loaded into the RAM 1303 and executed by the computing unit 1301, it may perform one or more of the thread stripping methods described above. Additionally, in other embodiments, the computing unit 1301 may be configured to perform the thread stripping method in any other suitable manner (eg, firmware).

[0303] Various embodiments of the systems or techniques described in this disclosure may be implemented using digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. Each of these embodiments may involve execution by one or more computer programs executed and / or interpreted by a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, capable of receiving data and instructions from, and transferring data and instructions to, a storage system, at least one input device, and at least one output device.

[0304] Program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programming data processing apparatus, such that when the program code is executed by the processor or controller, it can perform the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on-site, partially on-site, as a separate software package partially on-site and partially on a remote computer, or entirely on a remote computer or server.

[0305] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or stores a program for use by or in conjunction with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Further examples of machine-readable storage media include one or more wired electrical connections, a portable computer disk cartridge, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any combination of the foregoing.

[0306] To provide for user interaction, the systems and techniques described herein can be implemented on a computer that includes a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor, etc.) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball, etc.) for the user to provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, haptic feedback, etc.), and input from the user can be received in any form (e.g., acoustic input, voice input, tactile input, etc.).

[0307] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., as a data server), middleware components (e.g., an application server), front-end components (e.g., a user computer having a graphical user interface or network browser through which a user can interact with embodiments of the systems and techniques described herein), or any combination of such background, middleware, or front-end components. Components of the system can be connected to each other via any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0308] The computer system may include a client and a server. Typically, the client and server are remote from each other and generally interact via a communication network. The client-server relationship is created by a computer program running on a corresponding computer. The server may be a cloud server, a server in a distributed system, or a server incorporating a blockchain.

[0309] It should be understood that the various aspects of the flow charts shown above can be reordered, added, or deleted. For example, the operations described in this disclosure can be performed in parallel, sequentially, or in a different order. This disclosure is not limited to this, as long as the technical solutions disclosed in this disclosure can achieve the desired results.

[0310] The above specific examples do not constitute limitations on the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions are possible depending on design considerations and other factors. Any modifications, equivalent replacements, improvements, etc. within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. 1. An image-based search processing method, comprising: In response to receiving an image transmitted from a terminal, determining a first request corresponding to the image, the first request including an original search intent of a user; generating search results and recommendation content corresponding to the image based on the first request when the first request satisfies a recommendation trigger condition; returning the search results and the recommendation content to the terminal so that the terminal outputs the search results and the recommendation content on a search result page. Image-based search processing method.

2. The recommendation content includes at least one of a prompt content, a counter-question content, and a function entry, wherein: The prompting content includes a first type of information for prompting the user to clarify search intent; the reverse query includes a second type of information for searching to clarify the user's search intent; The function entry indicates an entry to available functions recommended to the user. The image-based search processing method of claim 1 .

3. Determining a first request corresponding to the image includes: analyzing each object in the image to obtain an intention estimation result corresponding to the image, the intention estimation result including a target object in the image; generating image content description information for said image based on a description model appropriate for a content type of said image; determining the first request based on the intention estimation result and the image content description information; The image-based search processing method of claim 1 .

4. determining the first request based on the intention estimation result and the image content description information, determining the original search intent according to the intent estimation result and the image content description information; determining the first request based on the original search intent; The image-based search processing method of claim 3.

5. generating search results and recommendation content corresponding to the image based on the first request, generating search results corresponding to the image based on the first request, the search results including search results obtained based on the original search intent; generating the recommendation content corresponding to the image based on the first request and the search results; The image-based search processing method of claim 1 .

6. The image-based search processing method includes: determining a new first request based on a follow-up question request transmitted from the terminal in response to receiving the follow-up question request, the follow-up question request being generated by the terminal based on a selection operation on the recommendation content; generating new search results and / or new recommendations corresponding to the image based on the new first request; and returning the new search results and / or the new recommendation content to the terminal so that the terminal outputs the new search results and / or the new recommendation content on the search result page. The image-based search processing method of claim 1 .

7. The image-based search processing method includes: determining a second request corresponding to the image, the second request including a potential search intent of the user; generating search results and recommendation content corresponding to the image based on the first request, generating search results corresponding to the image based on the first request, or generating the recommendation content corresponding to the image in accordance with the first request and the second request, wherein the search results include at least search results obtained based on the original search intent; generating the recommendation content corresponding to the image based on the second request, or generating the recommendation content corresponding to the image in accordance with the first request, the search results, and the second request, wherein the recommendation content includes at least the search results obtained based on the implicit search intent; The image-based search processing method of claim 1 .

8. Determining a second request corresponding to the image includes: Once the search stage of the original search intent has been determined, predicting the potential search intent based on the search stage of the original search intent; determining the second request based on the implicit search intent; The image-based search processing method of claim 7.

9. Determining a second request corresponding to the image includes: determining a point of interest corresponding to the image based on the first request; obtaining relevant content corresponding to the first request based on the attention points; determining the second request based on the relevant content. The image-based search processing method of claim 7.

10. The image-based search processing method includes: determining a new second request based on a follow-up question request transmitted from the terminal in response to receiving the follow-up question request, the follow-up question request being generated by the terminal based on a selection operation on the recommendation content; and generating new search results and / or new recommendations corresponding to the image based on the new second request; and returning the new search results and / or the new recommendation content to the terminal so that the terminal outputs the new search results and / or the new recommendation content on the search result page. The image-based search processing method of claim 7.

11. generating the recommendation content corresponding to the image based on the second request, Matching a target scenario to the second requirement; providing function entries that match the target scenario, wherein the recommendation content includes function entries. The image-based search processing method of claim 7.

12. The image-based search processing method includes: generating search results corresponding to the image based on the first request if the first request does not satisfy the recommendation trigger condition; returning the search results to the terminal such that the terminal outputs the search results on a search result page. The image-based search processing method of claim 1 .

13. Before determining a first request corresponding to the image, the image-based search processing method includes: performing a risk control check on the image, and if the image meets the risk control criteria, entering a flow for determining a first request corresponding to the image; and if the image does not meet the risk control criteria, returning prompt information to the terminal to indicate that the image is invalid. The image-based search processing method of claim 1 .

14. Before determining a first request corresponding to the image, the image-based search processing method includes: performing code detection on the image; and if the image meets a code detection criterion, not entering a flow for determining a first request corresponding to the image, obtaining a code detection result for the image through a code detection model, and returning the code detection result to the terminal. The image-based search processing method of claim 1 .

15. 1. An image-based search processing method, comprising: In response to receiving an image entered by a user via a search application, transmitting the image to a server; receiving search results and recommendations returned from the server, the search results being determined when the image satisfies a recommendation trigger condition, and the search results being determined based on a first request corresponding to the image, the first request including a user's original search intent; displaying the search results and the recommendation content on a search result page. Image-based search processing method.

16. the recommendation content is generated based on the first request and the search results.

16. The image-based search processing method of claim 15.

17. the recommendation content is generated based on a second request including an implicit search intent of the user, and includes search results obtained based on the implicit search intent; 16. The image-based search processing method of claim 15.

18. The recommendation content is generated in accordance with the first request, the search results, and a second request, the second request including a potential search intent of the user.

16. The image-based search processing method of claim 15.

19. The image-based search processing method includes: transmitting a follow-up question request to the server in response to receiving a selection operation for the recommendation content; the server receives new search results and / or new recommendations based on the follow-up request, and displays the new search results and / or the new recommendations on the search result page.

16. The image-based search processing method of claim 15.

20. The image-based search processing method includes: In response to receiving a selection operation for the recommendation content, displaying text information corresponding to the selection operation in a corresponding input box on the search result page; displaying the edited character information in the input box in response to receiving an edit operation on the character information; and generating the follow-up question request based on the edited text information as a response to the recommendation content.

20. The image-based search processing method of claim 19.

21. Before outputting the search results and the recommendation content, the image-based search processing method includes: The method further includes not outputting the recommendation content in response to receiving an operation to turn off the recommendation function.

16. The image-based search processing method of claim 15.

22. After outputting the search results and the recommendation content, the image-based search processing method includes: and further comprising: removing or hiding the recommendation content on the search result page in response to receiving an operation to turn off the recommendation function.

16. The image-based search processing method of claim 15.

23. An image-based search processing device, comprising: a first determination module for determining, in response to receiving an image transmitted from a terminal, a first request corresponding to the image, the first request including a user's original search intent; a generating module for generating search results and recommendation content corresponding to the image based on the first request when the first request satisfies a recommendation trigger condition; a first communication module for returning the search results and the recommendation content to the terminal so that the terminal outputs the search results and the recommendation content on a search result page; Image-based search processor.

24. The image-based search processing device includes: a second determination module for determining a second request corresponding to the image, the second request including a potential search intent of a user; The request generation module: a first generating sub-module for generating search results corresponding to the image based on the first request, or generating the recommendation content corresponding to the image in accordance with the first request and the second request, wherein the search results include at least search results obtained based on the original search intent; a second generation sub-module for generating the recommendation content corresponding to the image based on the second request, or for generating the recommendation content corresponding to the image in accordance with the first request, the search results, and the second request, wherein the recommendation content includes at least the search results obtained based on the implicit search intention; 24. The image-based retrieval processing device of claim 23.

25. An image-based search processing device, comprising: a second communication module for receiving an image input by a user through a search application, transmitting the image to a server, and receiving search results and recommendation content returned from the server, the search results being determined when the image satisfies a recommendation trigger condition, and the search results being determined based on a first request corresponding to the image, the first request including an original search intent of the user; and an output control module for displaying the search results and the recommendation content on a search result page; Image-based search processor.

26. at least one processor; a memory communicatively coupled to the at least one processor; 23. An electronic device, wherein the memory stores instructions executable by the at least one processor, the instructions, when executed by the at least one processor, causing the at least one processor to perform the method of any one of claims 1 to 22.

27. A non-transitory computer readable storage medium having stored thereon computer instructions that cause a computer to perform the method of any one of claims 1 to 22.

28. A program for implementing the control method according to any one of claims 1 to 12 when executed by a processor in a computer.

Citation Information

Patent Citations

  • Image-based search method and device, electronic equipment and storage medium

    CN114647756A

  • Content retrieval method and device, equipment and medium

    CN116501960A

  • Interactive searching method and apparatus

    JP2015225657A

  • Information retrieval method and device

    JP2016115294A

  • Intelligent system and method for visual search query

    JP2022051559A