An image shooting interaction method and a shooting terminal

Through the combination of multi-entry shooting interface and vertical response model, the problem of photographers switching between different functional entrances is solved, and efficient and accurate image recognition and response services are achieved in a single interface.

CN119854628BActive Publication Date: 2025-07-18HANGZHOU ROBAM APPLIANCES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510320952.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

In the prior art, photographers need to frequently switch the interactive interfaces of different functional portals, resulting in a fragmentation of the user experience. The call of background model programs lacks a diversion mechanism, which is difficult to match the actual needs of photographers.

Method used

It provides an image shooting interaction method, which sets vertical response model and shunt decision model through a multi-entry shooting interface, automatically matches the photographer's functional needs, and ensures the accuracy of identification results and the reliability of response results.

Benefits of technology

It realizes the completion of diversified functional requirements in a single interface, reduces the sense of operation switching, and improves the accuracy and reliability of identification and response results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854628B_ABST
    Figure CN119854628B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of information processing, and particularly to an image shooting interaction method and a shooting terminal. The method is applied to the shooting terminal; the shooting terminal provides a multi-entry shooting interface; a first entry mode is set in the multi-entry shooting interface, and the first entry mode has a corresponding vertical response model, and the vertical response model is matched with the corresponding image category. The image shooting interaction method provided by this application is applied to a shooting terminal that provides a multi-entry shooting interface, aggregating scattered multi-functional entries into a unified multi-entry shooting interface, so that the shooter can meet diverse functional requirements without exiting the interface, eliminating the fragmentation of shooting operations; in addition, different functional entry modes in the multi-entry shooting interface respectively correspond to multiple background model programs, which can distinguish the actual functional requirements of the shooter, and divert the obtained images to the appropriate background model programs for processing to improve the accuracy and reliability of the response results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technologies, and in particular, to an image shooting interaction method and a shooting terminal. Background Art

[0002] With the rapid development of Internet technologies, the interaction entrances between shooters and service systems have become increasingly diverse and complex.

[0003] However, in related technologies, an interaction interface and a background model program are usually developed separately for each function entrance. When a shooter calls different background model programs based on required function requirements, they often need to switch to different interaction interfaces or even application programs, resulting in fragmentation of each function entrance and seriously affecting the shooter's usage experience. In addition, the call of the background model program completely depends on the shooter's understanding of the function requirements, lacking a diversion mechanism, resulting in the called model not necessarily fully matching the function requirements actually needed by the shooter. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an image shooting interaction method and a shooting terminal.

[0005] In a first aspect, an embodiment of the present invention provides an image shooting interaction method, which is applied to a shooting terminal; the shooting terminal provides a multi-entrance shooting interface; a first entrance mode is set in the multi-entrance shooting interface, and the first entrance mode has a corresponding vertical response model, and the vertical response model matches the corresponding image category;

[0006] The method includes:

[0007] In response to a shooting instruction triggered through the first entrance mode, obtain a first target image captured, and input the first target image into a general recognition model to determine a recognition result;

[0008] In response to the recognition result matching the vertical recognition model corresponding to the first entrance mode, input the recognition result into the corresponding vertical recognition model to obtain a corresponding response result;

[0009] In response to the recognition result not matching the vertical recognition model corresponding to the first entrance mode, input the recognition result into a corresponding diversion decision model to determine a corresponding vertical recognition model;

[0010] Input the recognition result into the determined vertical recognition model to obtain a corresponding response result.

[0011] In combination with the first aspect, a second entrance mode is also set in the multi-entrance shooting interface, and the method further includes:

[0012] In response to a shooting instruction triggered through the second entry mode, obtain the captured second target image, and input the second target image into a general recognition model to determine the recognition result;

[0013] Input the recognition result into the corresponding shunt decision model to determine the corresponding vertical recognition model;

[0014] Input the recognition result into the determined vertical recognition model to obtain the corresponding response result.

[0015] Combined with the first aspect, a third entry mode is further set in the multi-entry shooting interface, and the method further includes:

[0016] In response to an upload instruction triggered through the third entry mode, obtain the uploaded third target image, and input the third target image into a general recognition model to determine the recognition result;

[0017] Input the recognition result into the corresponding shunt decision model to determine the corresponding vertical recognition model;

[0018] Input the recognition result into the determined vertical recognition model to obtain the corresponding response result.

[0019] Combined with the first aspect, the recognition result includes the recognized image features and the description text corresponding to the image features, and the response result is generated by a large language model.

[0020] Combined with the first aspect, a fourth entry mode is further set in the multi-entry shooting interface, and the method further includes:

[0021] In response to a session instruction triggered through the fourth entry mode, obtain the corresponding session content, which is used to input into a general recognition model to determine the recognition result and / or used to input into a vertical recognition model to obtain the corresponding response result.

[0022] Combined with the first aspect, the first entry mode is the shooting and recognizing face mode, the vertical recognition model corresponding to the first entry mode is the face recognition model, and the corresponding image category is the face image; the recognition result includes facial index information and skin quality index information;

[0023] The step of inputting the recognition result into the corresponding vertical recognition model to obtain the corresponding response result includes:

[0024] Input the facial index information and the skin quality index information into the face recognition model to output the first response result; the first response result includes at least one of the health status, body index, and improvement suggestions of the shooter; wherein, the face recognition model is trained based on a fatigue degree algorithm.

[0025] In combination with the first aspect, the first entry mode is the tongue coating shooting and recognition mode. The vertical recognition model corresponding to the first entry mode is the tongue coating recognition model, and the corresponding image category is the tongue coating image. The recognition result includes tongue coating index information;

[0026] The step of inputting the recognition result into the corresponding vertical recognition model to obtain the corresponding response result includes:

[0027] Input the tongue coating index information into the tongue coating recognition model to output a second response result. The second response result includes at least one of the constitutive index, physical index, physical health status, and improvement suggestions of the shooter. Among them, the tongue coating recognition model is trained based on a health degree algorithm.

[0028] In combination with the first aspect, the first entry mode is the food ingredient shooting and recognition mode. The vertical recognition model corresponding to the first entry mode is the food ingredient recognition model, and the corresponding image category is the food ingredient image. The recognition result includes at least food ingredient information;

[0029] The step of inputting the recognition result into the corresponding vertical recognition model to obtain the corresponding response result includes:

[0030] In response to the category information of the food ingredient information matching the sub-category in the preset food ingredient name list, input the food ingredient information into the food ingredient recognition model to output a third response result for the food ingredient. Among them, the food ingredient recognition model is trained based on the g-Dino detection method and the ROI region extraction method.

[0031] In combination with the first aspect, the second entry mode is the shooting and recognition diversion mode. The second entry mode corresponds to a general recognition model and a diversion decision model. The diversion decision model includes:

[0032] A multi-modal feature extraction module for extracting features from the image to be processed and outputting a feature vector;

[0033] A diversion decision fully connected network for processing the input feature vector to determine the target vertical response model. Among them, the diversion decision fully connected network includes an input layer, a fully connected layer, and an output layer. A Softmax layer or a Sigmoid layer is also provided between the fully connected layer and the output layer.

[0034] In a second aspect, the present application also provides a shooting terminal. The shooting terminal includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the shooting terminal to execute the above method.

[0035] The embodiments of the present invention bring the following beneficial effects: An image capture interaction method and a capture terminal provided by the present application, the method is applied to the capture terminal; the capture terminal provides a multi-entry capture interface; a first entry mode is set in the multi-entry capture interface, and the first entry mode has a corresponding vertical response model, and the vertical response model matches the corresponding image category; the method includes: in response to a capture instruction triggered through the first entry mode, obtaining a first target image captured, and inputting the first target image into a general recognition model to determine a recognition result; in response to the recognition result matching the vertical recognition model corresponding to the first entry mode, inputting the recognition result into the corresponding vertical recognition model to obtain a corresponding response result; in response to the recognition result not matching the vertical recognition model corresponding to the first entry mode, inputting the recognition result into a corresponding diversion decision model to determine the corresponding vertical recognition model; inputting the recognition result into the determined vertical recognition model to obtain a corresponding response result.

[0036] The image capture interaction method provided by the present application is applied to a capture terminal that provides a multi-entry capture interface, aggregating scattered function entries into a unified multi-entry capture interface, so that the photographer can meet diverse functional requirements without exiting the interface, eliminating the fragmentation of capture operations; in addition, different function entry modes in the multi-entry capture interface respectively correspond to multiple background model programs, which can distinguish the actual functional requirements of the photographer, divert the obtained images to the appropriate background model programs for processing, so as to improve the accuracy and reliability of the response results.

[0037] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by practicing the present invention. The objectives and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the specification, the claims, and the drawings.

[0038] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, is described in detail as follows. Description of the Drawings

[0039] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a flowchart of an implementation example of the image capture interaction method provided by the embodiments of the present invention;

[0041] Figure 2 Another flowchart of the image capture interaction method provided by the embodiment of the present invention;

[0042] Figure 3 Schematic diagram of the shunt decision fully connected network in the image capture interaction method provided by the embodiment of the present invention;

[0043] Figure 4 Schematic diagram of another shunt decision fully connected network in the image capture interaction method provided by the embodiment of the present invention;

[0044] Figure 5 Schematic diagram of the structure of the shooting terminal provided by the embodiment of the present invention;

[0045] Figure 6 A schematic diagram of a multi - entrance shooting interface of the shooting terminal provided by the embodiment of the present invention.

[0046] Reference numerals:

[0047] 130 - Processor, 131 - Memory, 132 - Bus, 133 - Communication interface. Detailed implementation manners

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] To facilitate the understanding of this embodiment, the application scenarios and design concepts of the embodiments of the present application will be briefly introduced below.

[0050] When the existing interaction interface interacts with different entrances, it is necessary to frequently switch the interface, resulting in a strong sense of fragmentation for the shooter. Separately configuring shunt algorithms for multiple entrances is inconvenient for development and maintenance, and it is difficult to dynamically adjust the best image parsing results and provide high - quality responses according to the scenario.

[0051] Based on this, the embodiments of the present application provide an image capture interaction method and a shooting terminal.

[0052] Embodiment 1

[0053] The present application provides an image capture interaction method, which is applied to a shooting terminal; the shooting terminal provides a multi - entrance shooting interface; a first entrance mode is set in the multi - entrance shooting interface, and the first entrance mode has a corresponding vertical response model, and the vertical response model matches the corresponding image category.

[0054] Combined withFigure 1 As shown in the figure, the method includes:

[0055] S110: In response to a shooting instruction triggered by a first entry mode, obtain a captured target image and input the target image into a general recognition model to determine a recognition result.

[0056] S120: If the recognition result matches the vertical recognition model corresponding to the first entry mode, input the recognition result into the corresponding vertical recognition model to obtain a corresponding response result.

[0057] S130: If the recognition result does not match the vertical recognition model corresponding to the first entry mode, input the recognition result into a corresponding diversion decision model to determine the corresponding vertical recognition model.

[0058] S140: Input the recognition result into the determined vertical recognition model to obtain a corresponding response result.

[0059] The image shooting interaction method provided by this application is applied to a shooting terminal. A multi-entry shooting interface is provided on the shooting terminal, and scattered multi-source entry requests are uniformly mapped to a single aggregated page. By interacting on a single page, the number of jump steps is reduced, the operation coherence is improved, and the sense of fragmentation in the operator's operation switching is eliminated. It can not only efficiently process different types of images, but also provide personalized responses and suggestions based on the recognition result, improve the service accuracy, and improve the service accuracy.

[0060] It can be understood that the number of first entry modes is not limited. When there are multiple first entry modes, there is a one-to-one correspondence between the multiple first entry modes and multiple vertical response models. Each vertical response model recognizes and responds to a specified category of images. Therefore, after the recognition result is obtained by recognizing the first target image based on the general recognition model in step S110, if the recognition result matches the current vertical recognition model, then in step S120, the first target image is further recognized and responded to by this vertical recognition model; if the recognition result does not match the current vertical recognition model, it cannot be accurately further recognized and responded to. At this time, the first target image is diverted by the diversion decision model in step S130 to match a suitable vertical recognition model for it. Thus, in step S140, further analysis and response are performed based on the determined vertical recognition model to obtain a response result. In addition, when the shooter forgets to select an entry mode or selects the wrong first entry mode, after the initial recognition is obtained through the general model, the recognition result is matched with the vertical recognition model corresponding to this first entry mode. If they do not match, the recognition result is input into the diversion decision model to determine the corresponding vertical recognition model for further recognition and response, so as to achieve automatic matching correction, determine a suitable vertical recognition model for recognition and response, and improve the accuracy and user experience.

[0061] Combined with the second aspect, a second entry mode is also set in the multi-entry shooting interface. Combined Figure 2 As shown, the method further includes:

[0062] S210, in response to a shooting instruction triggered through the second entry mode, obtain a second target image captured, and input the second target image into a general recognition model to determine a recognition result.

[0063] S220, input the recognition result into a corresponding diversion decision model to determine a corresponding vertical recognition model.

[0064] S230, input the recognition result into the determined vertical recognition model to obtain a corresponding response result.

[0065] In this embodiment, the multi-entry aggregation interface is provided with at least two types of entry modes, one or more first entry modes and a second entry mode. Among them, each first entry mode corresponds to a vertical recognition model; and the first entry mode also corresponds to a diversion decision model, which is used to divert the recognition result determined by the general recognition model to the vertical recognition model that matches the image category of the second target image for recognition and response to obtain a corresponding response result. Among them, as an implementable way, the vertical recognition model is a large language model. At this time, both recognition and response are completed by the vertical recognition model; as another implementable way, the vertical recognition model can only complete recognition, and the response is handed over to the subsequent large language model to complete. This is only an example here and is not limited.

[0066] It can be understood that the number of the first type of entry modes is not limited. Each first type of entry mode corresponds to a vertical recognition model, and can be adjusted and configured according to actual application requirements, which is not limited here.

[0067] In this embodiment, there are three first type of entry modes. Please refer to Figure 6 , the three first type of entry modes are respectively named: "Shoot ingredients", "AI face diagnosis", "AI tongue diagnosis"; the second type of entry mode is named "Recognize all things". The image captured in the "Recognize all things" mode is the "second target image". The general recognition model is used for preliminary image recognition and analysis to provide a preliminary recognition result, and the corresponding recognition result is dynamically diverted through the diversion decision model to the corresponding vertical recognition models in "Shoot ingredients", "AI face diagnosis", and "AI tongue diagnosis" for fine-grained recognition, analysis and response to provide in-depth analysis in a specific field, so as to obtain a response result. Among them, the general recognition model has good versatility and can recognize various types of images, but its recognition accuracy and granularity are not as good as those of the vertical recognition model specialized in a certain image category.

[0068] Exemplarily, a general recognition model is used for preliminary and extensive classification and feature extraction of input images. It can identify the main object categories and basic features in the image, can process various categories of images to provide preliminary classification results, but the accuracy may be limited. For example, there is a carrot on the table, and the recognition result obtained by the general recognition model may be: the image contains "food" or "vegetable". The vertical recognition model focuses on the recognition of specific categories of images and provides more refined and professional analysis. It can conduct in-depth analysis by combining domain knowledge and has high accuracy and high specialization. Combining the above example, the recognition result after being recognized by the vertical recognition model may be: This is a fresh carrot, rich in vitamin A, suitable for cold dressing or stir-frying.

[0069] It can be understood that different image recognition tasks require different processing methods and professional knowledge. Through the preliminary classification of the general recognition module and the diversion of the diversion decision model, the most suitable vertical recognition model can be quickly determined, reducing unnecessary calculations and time waste.

[0070] In the case of triggering the first type of entry mode ("taking pictures of ingredients", "AI face diagnosis", "AI tongue diagnosis"), after the first target image obtained by shooting is subjected to preliminary image recognition and analysis by the general recognition model, the recognition result is matched with the image categories that the vertical recognition model can recognize. If they match, the current vertical recognition model can be used for recognition and response; if they do not match, it is necessary to re-divert to determine the matching vertical recognition model, so as to conduct targeted recognition and response for different categories of the first target image taken, which helps to obtain more fine-grained recognition results and provide strong data support for subsequent responses.

[0071] Combined with the first aspect, the second entry mode is a shooting recognition diversion mode, and the diversion decision model corresponding to the second entry mode includes: a multi-modal feature extraction module and a diversion decision fully connected network connected in sequence.

[0072] Among them, the multi-modal feature extraction module is used for feature extraction of the image to be processed and the input text, and outputs an image feature vector (marked as "v" in this application).

[0073] The diversion decision fully connected network is used for data processing of the input feature vector to determine the target second type of entry mode.

[0074] Among them, the second entry mode is a shooting recognition diversion mode, and the second entry mode corresponds to the general recognition model and the diversion decision model; the diversion decision fully connected network includes: an input layer, a fully connected layer, and an output layer, and a Softmax layer or a Sigmoid layer is also provided between the fully connected layer and the output layer.

[0075] Please refer to Figure 3 Exemplarily proposed is a structure of a shunt decision fully connected network. As an implementable manner, the shunt decision fully connected network includes two cascaded fully connected layers, and finally outputs probability values in the range of (0, 1) through the Softmax layer to map the feature H (the feature H can be an image feature vector v, or can be obtained by fusing the image feature v with other feature vectors such as text feature vectors and speech feature vectors) to the shunt weights P of multiple different algorithm channels, so as to determine the optimal algorithm channel. The "algorithm channel" refers to a vertical recognition model. Among them, a ReLU activation function is set between the two fully connected layers to introduce non-linear features.

[0076] Among them, multiple vertical recognition models correspond one-to-one with multiple algorithm channels. Combining the above example, P = [P face , P tongue , P food , where P face corresponds to the algorithm channel of "AI face diagnosis", P tongue corresponds to the algorithm channel of "AI tongue diagnosis", and P food corresponds to the algorithm channel of "taking pictures of ingredients". For example, if P = [0.7, 0.2, 0.1], it means that there is a 70% probability of entering "AI face diagnosis", a 20% probability of entering AI tongue diagnosis, and a 10% probability of entering ingredient recognition. At this time, image recognition and response are performed using the vertical recognition model corresponding to the first entry mode of "AI face diagnosis".

[0077] It can be understood that Figure 4 only an exemplary shunt decision fully connected network structure for simultaneously performing two tasks is provided. The input layer is connected to a shared hidden layer, and the shared hidden layer is connected to the task 1 fully connected layer and the task 2 fully connected layer. The task 1 fully connected layer uses parameters W2⁽¹⁾ and b2⁽¹⁾, and the output is processed through the Softmax layer to obtain the task 1 probability P1; the fully connected layer of task 2 uses parameters W2⁽²⁾ and b2⁽²⁾, and the output is processed through the Sigmoid layer to obtain the task 2 score S. In the actual application process, the number of tasks can be adjusted according to actual needs. This is only an example here and is not limited. In this way, it is possible to perform recognition and response on multiple images taken and / or uploaded.

[0078] Combined with the first aspect, a third entry mode is also set in the multi-entry shooting interface. Combining Figure 5 as shown, the method further includes:

[0079] S310, in response to the upload instruction triggered through the third entry mode, obtain the uploaded target image, and input the target image into the general recognition model to determine the recognition result.

[0080] S320. Input the recognition result into the corresponding shunt decision model to determine the corresponding vertical recognition model.

[0081] S330. Input the recognition result into the determined vertical recognition model to obtain the corresponding response result.

[0082] In this embodiment, a third entry mode is further set in the multi-entry shooting interface. As shown in combination with Figure 6 a third entry mode of "Upload Image" is set in the upper left corner of the multi-entry shooting interface. Through this, image upload is performed to obtain a third target image.

[0083] At this time, there are several situations:

[0084] Situation 1: The photographer does not select to trigger the first entry mode and the second entry mode, and performs an image upload operation using the previously used entry mode or the initialized entry mode. If the previously used entry mode or the initialized entry mode is the first entry mode (combining the above example, such as "AI face consultation"), after obtaining the uploaded third target image, the recognition result obtained by preliminarily recognizing the third target image through the general recognition model is matched with the vertical recognition model of the first entry mode. If the match fails, step S320 needs to be executed to re-shunt, determine a suitable vertical recognition model, and then perform recognition and response.

[0085] Situation 2: The photographer does not select to trigger the first entry mode and the second entry mode, and performs an image upload operation using the previously used entry mode or the initialized entry mode. If the previously used entry mode or the initialized entry mode is the first entry mode (combining the above example, such as "AI face consultation"), after obtaining the uploaded third target image, the recognition result obtained by preliminarily recognizing the third target image through the general recognition model is matched with the vertical recognition model of the first entry mode. If the match is successful, step S330 is executed for further recognition and response by this vertical recognition model.

[0086] Situation 3: The photographer does not select to trigger the first entry mode and the second entry mode, and performs an image upload operation using the previously used entry mode or the initialized entry mode. If the previously used entry mode or the initialized entry mode is the second entry mode (combining the above example, such as "Recognize Everything"), after obtaining the recognition result by preliminarily recognizing the uploaded third target image through the general recognition model, step S320 is executed to perform shunt processing on it to determine a suitable vertical recognition model for recognition and response.

[0087] In Case 4, the photographer has selected to trigger a certain first entry mode, and at this time, the image is uploaded in the entry mode selected by the photographer. After obtaining the uploaded third target image, the recognition result obtained by preliminarily recognizing the third target image through the general recognition model is matched with the vertical recognition model of this first entry mode. If the match fails, step S320 needs to be executed to re-split and determine a suitable vertical recognition model for recognition and response. In this way, automatic error correction can be achieved, and the phenomenon that the uploaded image does not match the selected entry mode can be avoided.

[0088] In Case 5, the photographer has selected to trigger a certain first entry mode, and at this time, the image is uploaded in the entry mode selected by the photographer. After obtaining the uploaded third target image, the recognition result obtained by preliminarily recognizing the third target image through the general recognition model is matched with the vertical recognition model of this first entry mode. If the match is successful, step S330 is executed for further recognition and response by this vertical recognition model.

[0089] In Case 6, if the entry mode is the second entry mode (combining the above example, such as "Recognize Everything"), after obtaining the recognition result from the preliminary recognition of the uploaded third target image through the general recognition model, step S320 is executed for its shunt processing to determine a suitable vertical recognition model for recognition and response.

[0090] Combined with the first aspect, the recognition result includes the recognized image features and the description text corresponding to the image features, and the response result is generated by the large language model.

[0091] It can be understood that the general recognition model can recognize the basic features in the image, such as objects, texts, scenes (such as environmental types, activities, behaviors, time, weather, atmosphere, etc. like nature, city, indoor, etc.), as well as image categories (such as faces, food ingredients, buildings, landscapes, etc.), and the descriptions corresponding to the image features (i.e., the description text). For example, the image feature may be "a red circular object", and the corresponding description text may be "A red circular object appears in the center of the picture".

[0092] Among them, the response result generated by the large language model refers to using a large-scale pre-trained language model (such as GPT, BERT, etc.) to generate a natural language description or answer based on the input text or image recognition result. These models can generate smooth, coherent and logically consistent text content by learning a large amount of text data.

[0093] Suppose the input image is a photo of a kitchen, which contains some food ingredients and cooking utensils. The general recognition model recognizes the following features:

[0094] Food ingredients: tomatoes, eggs, onions;

[0095] Cooking utensils: wok, spatula;

[0096] Scenario: Kitchen countertop.

[0097] The response result generated by the large language model may be: "This photo shows a kitchen worktop with some common cooking ingredients and tools. Fresh tomatoes, eggs, and onions can be seen, and there is also a wok and a spatula beside. It seems to be preparing to cook a delicious home-cooked meal. If you need specific recipe suggestions or cooking skills, please let me know."

[0098] This generated response not only describes the specific content in the image but also forms multi-modal information together with the original image. This multi-modal information combines visual data (image) and language data (descriptive text), and can more comprehensively reflect the image content. For example, the image itself may only show a device, while the generated descriptive text can supplement information such as the operating state and fault possibility of the device, further providing useful information or suggestions. In subsequent processing, the multi-modal information can be used for vertical recognition (i.e., in-depth recognition for a specific field) to provide more comprehensive and in-depth services, enhancing the shooter's experience.

[0099] Combined with the first aspect, a fourth entry mode is also set in the multi-entry shooting interface, and the method further includes:

[0100] S410, in response to a session instruction triggered by the fourth entry mode, obtain corresponding session content, where the session content is used to be input into the general recognition model to determine the recognition result, and / or, used to be input into the vertical recognition model to obtain a corresponding response result.

[0101] A fourth entry mode is also set in the multi-entry shooting interface. When the shooter inputs a session instruction through the fourth entry mode in the multi-entry shooting interface and clicks "Send" in the fourth entry mode (as shown in Figure 6 ), the system will obtain corresponding session content. Here, the "session content" can be the shooter's voice input (such as "Please identify the operating state of the electrical appliance in the image"), text input (such as "Analyze the problems in this picture"), or other forms of interaction information (such as gestures or button clicks).

[0102] The obtained session content will first be input into the general recognition model to determine a preliminary recognition result. The general recognition model can process various types of inputs and extract key features or intentions.

[0103] If necessary, the conversation content can also be directly or indirectly input into a specific vertical recognition model to obtain more professional response results. If the recognition result of the general recognition model matches that of a certain vertical recognition model, the result can be directly passed to that vertical recognition model; if not, the shunt decision model can be used to determine which vertical model to use.

[0104] As an example, the photographer took a photo of a kitchen through the first entry mode, which included ingredients and cooking utensils. The system recognized the objects in the image and generated a description: "This is a photo of a kitchen, containing fresh tomatoes, eggs and onions, with a wok and a spatula beside."

[0105] The photographer input a voice through the fourth entry mode: "I need to cook a simple home-cooked dish. Please recommend a recipe." The system converted this voice into text and input it into the general recognition model, and recognized that the photographer's intention was to look for a recipe. Then, the system passed this request to a dedicated vertical recognition model, and finally generated a response result: "Based on the ingredients you provided, it is recommended that you can try to cook a scrambled egg with tomatoes. This is a simple and delicious home-cooked dish."

[0106] In this way, the system can not only handle image recognition tasks, but also handle conversation instructions based on text or voice, providing more rich and flexible services.

[0107] Combined with the first aspect, the first entry mode is the face recognition mode for shooting, the vertical recognition model corresponding to the first entry mode is the face recognition model, and the corresponding image category is the face image; the recognition result includes at least one of facial index information and skin quality index information, and the facial index information includes at least one of nasolabial folds, dark circles under the eyes, and skin color grading; the skin quality index information includes at least one of spots, acne marks, and blackheads.

[0108] In step S120, the recognition result is input into the corresponding vertical recognition model to obtain the corresponding response result, which specifically includes:

[0109] S121, input the facial index information and the skin quality index information into the face recognition model, and output the first response result; the first response result includes at least one of the photographer's physical health status, physical indicators, and improvement suggestions; among them, the face recognition model is trained based on the fatigue degree algorithm.

[0110] In this embodiment, after the photographer takes or uploads a face image, it is preliminarily recognized by the general recognition model, and the relevant features of the face and skin quality are extracted. Subsequently, under the condition that the face image matches the currently selected or shunted face recognition model, recognition and response are performed based on the fatigue degree algorithm to generate the first response result.

[0111] Among them, the facial index information includes, but is not limited to, the positions and shapes of facial feature points such as nasolabial folds, dark circles under the eyes, eyes, eyebrows, mouth, nose, etc.; the skin quality index information includes, but is not limited to, the color, texture, glossiness of the skin, etc., and possible defects (such as spots, pimples, dark circles under the eyes, etc.).

[0112] In this embodiment, the face recognition model is trained based on a fatigue degree algorithm. This means that the model pays special attention to facial features related to fatigue during the training process, such as: the degree of eye opening and closing, the degree of dark circles under the eyes, the relaxation of facial muscles, the change in skin glossiness, etc.

[0113] In this embodiment, the first response result includes at least one of the health status, body indicators, and improvement suggestions of the photographer. Among them, based on the changes in facial features, the face recognition model can evaluate the overall health status of the photographer. For example, frequent dark circles under the eyes may indicate insufficient sleep or excessive stress; the face recognition model can also estimate some basic body indicators, such as fatigue index, stress level, etc. This is usually achieved by analyzing facial features and skin quality changes. Further, according to the evaluation results, the face recognition model can also provide personalized improvement suggestions. For example: if obvious dark circles under the eyes are detected, it is recommended that the photographer ensure sufficient sleep time; if dry skin is found, it is recommended to drink more water and use moisturizing skin care products; if signs of fatigue are detected, it is recommended to rest and relax appropriately.

[0114] As an example, the first response result may be: "Based on the analysis of your facial features, you may have insufficient sleep recently. It is recommended that you ensure 7-8 hours of sufficient sleep every day, and at the same time pay attention to a balanced diet and appropriate exercise."

[0115] Further, combining the natural language processing ability of the large language model, extract the feature scenarios from the facial images, and push function cards based on the feature scenarios; the function cards can be recommended recipes, health suggestions, shopping recommendations, etc. For example, when the feature scenario is the kitchen, the system can push recipes that match the current health status, methods of rest and relaxation based on facial fatigue signs, etc.

[0116] Combined with the first aspect, the first entry mode is the mode of photographing and recognizing the tongue coating, the vertical recognition model corresponding to the first entry mode is the tongue coating recognition model, and the corresponding image category is the tongue coating image; the recognition result includes the tongue coating index information.

[0117] Inputting the recognition result into the corresponding vertical recognition model in step S120 to obtain the corresponding response result specifically includes:

[0118] S122. Input the tongue coating index information into the tongue coating recognition model to output a second response result, where the second response result includes at least one of the constitutive index, physical index, physical health status, and improvement suggestions of the photographer. The tongue coating recognition model is trained based on a health degree algorithm.

[0119] In this embodiment, after the photographer takes or uploads a tongue coating image, it is preliminarily recognized by a general recognition model to extract the main features in the tongue coating image, that is, the tongue coating index information, including the shape, color, texture, blood vessel status, etc. of the tongue coating. Subsequently, under the condition that the tongue coating image matches the currently selected or shunted tongue coating recognition model, these features are input into a dedicated tongue coating recognition model, which is trained based on a health degree algorithm. The generated second response result may be: "Based on the analysis of the color and blood vessel status of your tongue coating, you may have damp-heat in your body. It is recommended that you drink more water, pay attention to a light diet, and avoid spicy foods. In addition, appropriate exercise also helps to regulate your physical state."

[0120] The tongue coating recognition model is trained based on a health degree algorithm. This means that the model pays special attention to tongue coating features related to health during training. For example, different tongue coating colors may reflect different physical conditions (for example, a yellow tongue coating may indicate heat in the body, and a white tongue coating may indicate cold-dampness); the color and shape of the sublingual veins can reflect the blood circulation situation (for example, blood vessel dilation may indicate inflammation in the body or poor blood circulation). Based on the above features, after data processing by the health degree algorithm, a second response result is output. In this embodiment, the second response result includes at least one of the constitutive index, physical index, physical health status, and improvement suggestions of the photographer. Among them, based on the changes in tongue coating features, the overall health status of the photographer can be evaluated. For example, frequent appearance of a yellow tongue coating may indicate heat in the body; according to the color, shape, etc. of the tongue coating, it can be used to evaluate the constitutive type of the photographer, such as qi deficiency, blood stasis, etc., and can also be used to evaluate physical indicators, such as body temperature, blood pressure, etc. Further, according to the evaluation results, the model can provide personalized improvement suggestions. For example: If a yellow tongue coating is detected, it is recommended that the photographer drink more water, pay attention to a light diet, and avoid spicy foods; if dilation of the sublingual veins is found, it is recommended that the photographer exercise appropriately to promote blood circulation; if a thick and greasy tongue coating is detected, it is recommended that the photographer regulate the stomach and avoid greasy foods.

[0121] Combined with the first aspect, the first entry mode is the mode of photographing and recognizing ingredients. The vertical recognition model corresponding to the first entry mode is the ingredient recognition model, and the corresponding image category is the ingredient image. The recognition result includes at least ingredient information.

[0122] In step S120, inputting the recognition result into the corresponding vertical recognition model to obtain the corresponding response result specifically includes:

[0123] S123, in response to the category information of the ingredient information matching the sub-category in the preset ingredient name list, input the ingredient information into the ingredient recognition model and output the third response result of the ingredient; wherein, the ingredient recognition model is trained based on the g-Dino detection method and the ROI region extraction method.

[0124] In this embodiment, after the photographer takes or uploads an ingredient image, it is preliminarily recognized by the general recognition model to extract the main features in the ingredient image, that is, the ingredient information, including the shape, color, texture, etc. of the ingredient. Subsequently, under the condition that the ingredient image matches the currently selected or shunted ingredient recognition model and the category information of the ingredient image matches the sub-category in the ingredient name list, these features are input into a dedicated ingredient recognition model, which is trained based on the g-Dino detection method and the ROI region extraction method. The generated second response result may be: "The picture you took contains the following ingredients: tomatoes, beef, and onions. It is recommended that you can try making a stewed beef with tomatoes, which is a simple and delicious home-cooked dish. In addition, note that onions should be stored in a ventilated and dry place, and the remaining beef should be frozen to extend its shelf life."

[0125] Among them, the g-Dino detection method is an advanced object detection algorithm that can accurately locate and identify ingredient objects in images; the ROI region extraction method is used to extract the region of interest (Region of Interest), that is, the specific position and range of the ingredient, to improve the recognition accuracy. According to the analysis result of the ingredient recognition model, the system outputs the third response result, which at least includes: ingredient name, ingredient type (further subdivided ingredient types, such as vegetables, fruits, meats, etc.), quantity or weight, recommended use (providing cooking suggestions or recipe recommendations based on the ingredient information), and storage suggestions (providing the best storage method and shelf life suggestions based on the ingredient characteristics).

[0126] As an example, the photographer took an ingredient image containing tomatoes, eggs, and onions. The general recognition model pre-processes the image and extracts the main features, such as the color of the tomatoes, the shape of the eggs, the texture of the onions, etc. At this time, the ingredient image matches the ingredient vertical recognition model, and the above features (i.e., the recognition results) recognized by the general recognition model are input into a dedicated ingredient recognition model, which is trained based on the g-Dino detection method and the ROI region extraction method. The generated third response result may be: "The picture you took contains the following ingredients: tomatoes, eggs, and onions. It is recommended that you can try making a scrambled egg with tomatoes, which is a simple and delicious home-cooked dish. In addition, note that onions should be stored in a ventilated and dry place to extend its shelf life.

[0127] In a second aspect, the present application provides a photographing terminal, which, in combination with Figure 5 as shown, the photographing terminal includes a memory 131 and a processor 130. The memory 131 is used to store a computer program, and the processor 130 runs the computer program to enable the photographing terminal to execute the above method.

[0128] Further, in combination with Figure 5 as shown, the photographing terminal further includes a bus 132 and a communication interface 133. The processor 130, the communication interface 133, and the memory 131 are connected through the bus 132.

[0129] Among them, the memory 131 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 133 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 132 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only a bidirectional arrow is used in

[0130] The processor 130 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method may be completed by the integrated logic circuit of the hardware in the processor 130 or the instructions in the form of software. The above-mentioned processor 130 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention may be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 131, and the processor 130 reads the information in the memory 131 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0131] In this embodiment, the shooting terminal provides a multi-entry shooting interface; a first entry mode is set in the multi-entry shooting interface, and the first entry mode has a corresponding vertical response model, and the vertical response model matches the corresponding image category.

[0132] In this way, the scattered function entrances are aggregated into a unified multi-entry shooting interface, so that the shooter can meet diverse functional requirements without exiting this interface, eliminating the fragmentation feeling of shooting operations; in addition, different functional entry modes in the multi-entry shooting interface respectively correspond to multiple background model programs, which can distinguish the actual functional requirements of the shooter and divert the obtained images to the appropriate background model programs for processing to improve the accuracy and reliability of the response results.

[0133] Fourthly, an embodiment of the present application provides a readable storage medium, in which computer program instructions are stored. When the computer program instructions are read and run by a processor, the above-mentioned method is executed.

[0134] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0135] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "installed", "connected", and "coupled" shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0136] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0137] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0138] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. An image capture interaction method, characterized in that, The method is applied to a photographing terminal; the photographing terminal provides a multi-entry photographing interface; in the multi-entry photographing interface, a first entry mode and a second entry mode are set, the first entry mode has a corresponding vertical response model, and the vertical response model matches the corresponding image category; The method includes: In response to a photographing instruction triggered through the first entry mode, obtaining a first target image captured, and inputting the first target image into a general recognition model to determine a recognition result; the recognition result includes the recognized image features and a description text corresponding to the image features; In response to the recognition result matching the vertical recognition model corresponding to the first entry mode, inputting the recognition result into the corresponding vertical recognition model to obtain a corresponding response result; In response to the recognition result not matching the vertical recognition model corresponding to the first entry mode, inputting the recognition result into a corresponding shunt decision model to determine a corresponding vertical recognition model; Inputting the recognition result into the determined vertical recognition model to obtain a corresponding response result; In response to a photographing instruction triggered through the second entry mode, obtaining a second target image captured, and inputting the second target image into a general recognition model to determine a recognition result; Inputting the recognition result into a corresponding shunt decision model to determine a corresponding vertical recognition model; Inputting the recognition result into the determined vertical recognition model to obtain a corresponding response result; wherein, the response result is generated by a large language model.

2. The method according to claim 1, wherein A third entry mode is further set in the multi-entry photographing interface, and the method further includes: In response to an upload instruction triggered through the third entry mode, obtaining a third target image uploaded, and inputting the third target image into a general recognition model to determine a recognition result; Inputting the recognition result into a corresponding shunt decision model to determine a corresponding vertical recognition model; Inputting the recognition result into the determined vertical recognition model to obtain a corresponding response result.

3. The method according to claim 1, wherein A fourth entry mode is further set in the multi-entry photographing interface, and the method further includes: In response to a session instruction triggered through the fourth entry mode, obtaining corresponding session content, where the session content is used to be input into the general recognition model to determine the recognition result, and / or, used to be input into the vertical recognition model to obtain a corresponding response result.

4. The method according to claim 1 or 3, characterized in that, The first entry mode is a mode for photographing and recognizing a face, the vertical recognition model corresponding to the first entry mode is a face recognition model, and the corresponding image category is a face image; the recognition result includes facial index information and skin quality index information; The step of inputting the recognition result into the corresponding vertical recognition model to obtain a corresponding response result includes: Inputting the facial index information and the skin quality index information into the face recognition model, and outputting a first response result; the first response result includes at least one of the physical health status, physical indexes, and improvement suggestions of the photographer; wherein, the face recognition model is trained based on a fatigue degree algorithm.

5. The method according to claim 1 or 2, characterized in that, The first entry mode is the tongue coating photographing and recognition mode. The vertical recognition model corresponding to the first entry mode is the tongue coating recognition model, and the corresponding image category is the tongue coating image. The recognition result includes tongue coating index information; The step of inputting the recognition result into the corresponding vertical recognition model to obtain the corresponding response result includes: Inputting the tongue coating index information into the tongue coating recognition model to output a second response result. The second response result includes at least one of the constitutive index, physical index, physical health status, and improvement suggestions of the photographer. Among them, the tongue coating recognition model is trained based on a health degree algorithm.

6. The method according to claim 1 or 2, characterized in that, The first entry mode is the food ingredient photographing and recognition mode. The vertical recognition model corresponding to the first entry mode is the food ingredient recognition model, and the corresponding image category is the food ingredient image. The recognition result at least includes food ingredient information; The step of inputting the recognition result into the corresponding vertical recognition model to obtain the corresponding response result includes: In response to the category information of the food ingredient information matching the sub-category in the preset food ingredient name list, inputting the food ingredient information into the food ingredient recognition model to output a third response result for the food ingredient. Among them, the food ingredient recognition model is trained based on the g-Dino detection method and the ROI region extraction method.

7. The method according to claim 1, wherein The second entry mode is the photographing and recognition shunting mode. The second entry mode corresponds to the general recognition model and the shunting decision model. The shunting decision model includes: A multi-modal feature extraction module for extracting features from the image to be processed and outputting a feature vector; A shunting decision fully connected network for processing the input feature vector to determine the target vertical response model. Among them, the shunting decision fully connected network includes an input layer, a fully connected layer, and an output layer. A Softmax layer or a Sigmoid layer is also provided between the fully connected layer and the output layer.

8. A photographing terminal, characterized in that, The photographing terminal includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the photographing terminal to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image-based search processing method and device, equipment and storage medium

    CN118690033A