Food description information display method, device, equipment and computer readable medium
By dynamically encode and fusion of food images and text description information, and combining features to supplement information, more comprehensive food description information is generated, the problem of inaccurate food description information in the prior art is solved, and more accurate and comprehensive food feature description is achieved.
Patent Information
- Application Number
- CN202510207485.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When generating food description information, it is difficult to fully describe the information in the food image, and the characteristics of the food image are highly similar, resulting in the generated food description information having a single feature, which may be biased, and it is difficult to meet user needs.
By receiving food images and initial food description information, image and text feature information are generated, dynamic position encoding, fusing image and text position feature information, retrieving feature supplementary information corresponding to multimodal fusion feature information, and generating more comprehensive target food description information.
It realizes more accurate and comprehensive food description information, which can better meet the needs of users, and generates clear food taste and nutritional content information by combining image and text feature information.
Smart Images

Figure CN119693941B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and in particular to a method, device, electronic device, and computer-readable medium for displaying food description information. Background Art
[0002] In social networks, target users often learn about food through food images shared by other users. If the target users want to learn more about the food shown in the food images shared by others, they can use food recognition technology to recognize the food images. Food recognition technology is a technology that generates food description information based on food images. At present, when generating food description information, the method usually adopted is: generating food description information only by recognizing food images.
[0003] However, when using the above method to generate detailed description information about food, the following technical problems often occur:
[0004] The food description information generated only by recognizing food images often cannot fully describe the information of the food shown in the food images, and the image feature information between foods is often quite similar. As a result, when food description information is generated only from the perspective of food image recognition, the obtained food features are relatively single, and the generated food description information may be biased, which makes it difficult to meet user needs.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the invention
[0006] The content of this disclosure is used to introduce concepts in a brief form, which will be described in detail in the detailed implementation section below. The content of this disclosure is not intended to identify the key features or essential features of the technical solution claimed for protection, nor is it intended to limit the scope of the technical solution claimed for protection.
[0007] Some embodiments of the present disclosure propose a method, device, electronic device, and computer-readable medium for displaying food description information to solve one or more of the technical problems mentioned in the above background technology section.
[0008] In a first aspect, some embodiments of the present disclosure provide a method for displaying food description information, comprising: in response to receiving food identification request information sent by a user terminal, receiving a food image and initial food description information, wherein the initial food description information is text description information of the food corresponding to the food image; generating image feature information corresponding to the food image; performing a first dynamic position encoding on the image feature information to obtain image position feature information; generating text feature information corresponding to the initial food description information; performing a second dynamic position encoding on the text feature information to obtain text position feature information; fusing the image position feature information and the text position feature information, Obtain multimodal fusion feature information; retrieve feature supplementary information corresponding to the multimodal fusion feature information that characterizes and describes food features; generate target food description information based on the multimodal fusion feature information and the feature supplementary information; in response to determining that the food identification request information includes a target playback identifier, obtain the target food audio corresponding to the target food description information from a target audio database, wherein the target playback identifier indicates that the target food description information is played in the form of audio; send the target food description information and the target food audio to the user terminal, so that the user terminal can display the target food description information and play the target food audio by audio.
[0009] In a second aspect, some embodiments of the present disclosure provide a food description information display device, comprising: a receiving unit, configured to receive a food image and initial food description information in response to receiving food identification request information sent by a user terminal, wherein the initial food description information is text description information of the food corresponding to the food image; a first generating unit, configured to generate image feature information corresponding to the food image; a first executing unit, configured to perform a first dynamic position encoding on the image feature information to obtain image position feature information; a second generating unit, configured to generate text feature information corresponding to the initial food description information; a second executing unit, configured to perform a second dynamic position encoding on the text feature information to obtain text position feature information; and a fusion unit, configured to perform a second dynamic position encoding on the image position feature information and the text position feature information. The target food description information is fused with the feature information to obtain the multimodal fusion feature information; the third execution unit is configured to retrieve the feature supplementary information corresponding to the multimodal fusion feature information and representing the food feature; the third generation unit is configured to generate the target food description information based on the multimodal fusion feature information and the feature supplementary information; the acquisition unit is configured to obtain the target food audio corresponding to the target food description information from the target audio database in response to determining that the food recognition request information includes a target playback identifier, wherein the target playback identifier represents that the target food description information is played in the form of audio; the fourth execution unit is configured to send the target food description information and the target food audio to the user terminal, so that the user terminal displays the target food description information and plays the target food audio as audio.
[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0012] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the target food description information obtained by the food description information display method of some embodiments of the present disclosure is more accurate and comprehensive, and can better meet the needs of users. Specifically, the reason why the obtained food description information is often not accurate enough and difficult to meet the needs of users is that the food description information generated only by recognizing the food image often cannot fully describe the information of the food displayed in the food image, and the image feature information between foods often has a relatively similar situation, so that the food description information is generated only from the perspective of food image recognition, the obtained food features are relatively single, and the generated food description information may have deviations, which makes it difficult to meet the needs of users. Based on this, the food description information display method of some embodiments of the present disclosure, first, in response to receiving the food recognition request information sent by the user terminal, receives the food image and the initial food description information, wherein the initial food description information is the text description information of the food corresponding to the food image. In this way, the image of the food that the user wants to know and the existing food description information can be obtained, which provides basic data for the subsequent generation of image feature information and text feature information. Then, the image feature information corresponding to the above food image is generated, and the above image feature information is firstly dynamically encoded to obtain the image position feature information. Thus, by performing the first dynamic position encoding on the image feature information, the obtained image position feature information can represent the overall structural information of the image. Secondly, the text feature information corresponding to the above-mentioned initial food description information is generated, and the above-mentioned text feature information is subjected to the second dynamic position encoding to obtain the text position feature information. Thus, by performing the second dynamic position encoding on the text feature information, the obtained text position feature information can represent the contextual relationship of the initial food description information. Next, the above-mentioned image position feature information and the above-mentioned text position feature information are fused to obtain multimodal fusion feature information. Thus, the feature information of the image and the text are combined to form a multimodal feature representation, so that the characteristics of the food can be better represented. Further, the feature supplementary information corresponding to the above-mentioned multimodal fusion feature information that represents and describes the food characteristics is retrieved. Thus, according to the fused feature information, the relevant feature supplementary information is retrieved in the database storing the full amount of food feature information, so that the feature supplementary information and the multimodal fusion feature information can be fused to obtain more comprehensive features. Furthermore, based on the above-mentioned multimodal fusion feature information and the above-mentioned feature supplementary information, the target food description information is generated. Therefore, the target food information can clearly represent the taste, nutritional components and other information of the food, thereby obtaining a comprehensive and rich description of the target food that meets the needs of users as much as possible.Finally, in response to determining that the food identification request information includes a target playback identifier, the target food audio corresponding to the target food description information is obtained from the target audio database, wherein the target playback identifier represents the target food description information to be played in the form of audio, and the target food description information and the target food audio are sent to the user terminal so that the user terminal can display the target food description information and play the target food audio. Thus, the target food description information is displayed, and the user can intuitively see the detailed description information of the food. The target food description information is played in audio, which expands the audience range and improves the availability of the target food description information and user satisfaction. Based on this, by dynamically position encoding the image feature information and the text feature information respectively, and fusing the image position feature information and the text position feature information, and retrieving the relevant feature supplementary information from the pre-constructed database, the obtained target food description information is more accurate and comprehensive, and can better meet the needs of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flow chart of some embodiments of the method for displaying food description information according to the present disclosure;
[0015] Figure 2 is a schematic diagram of the structure of some embodiments of the food description information display device according to the present disclosure;
[0016] Figure 3 is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure;
[0017] Figure 4 This is a schematic screenshot of an internal test operation page for displaying food description information according to the food description information display method disclosed in the present invention. DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0024] refer to Figure 1 , shows a process 100 of some embodiments of the method for displaying food description information according to the present disclosure. The method for displaying food description information comprises the following steps:
[0025] Step 101: In response to receiving food identification request information sent by a user terminal, a food image and initial food description information are received.
[0026] In some embodiments, the execution subject (e.g., computing device) of the above-mentioned food description information display method may receive a food image and initial food description information in response to receiving a food identification request information sent by a user terminal. The above-mentioned user terminal may be a terminal device used by a target user interacting with the above-mentioned execution subject. For example, the above-mentioned user terminal may include but is not limited to a mobile phone and a tablet computer. A food identification button may be provided on the above-mentioned user terminal. In practice, the target user may click the above-mentioned food identification button to enable the above-mentioned user terminal to generate food identification request information and send the food identification request information to the above-mentioned execution subject. The above-mentioned food identification request information may be information requesting identification of food. The above-mentioned food image may be an image showing the food that the target user wants to know. The above-mentioned initial food description information may be text description information of the food corresponding to the above-mentioned food image. As an example, the above-mentioned food may be fried fish with Hami melon. The above-mentioned food image may be an image showing a dish of fried fish with Hami melon. The above-mentioned initial food description information may be: a food composed of Hami melon and fish as raw materials, the Hami melon is yellow in color, and the fish is crucian carp.
[0027] Step 102: Generate image feature information corresponding to the food image.
[0028] In some embodiments, the execution subject may generate image feature information corresponding to the food image. The image feature information may characterize the feature semantic content of the food image. In practice, the execution subject may generate image feature information corresponding to the food image through conventional image feature extraction technology. The image feature information includes but is not limited to the color features of the food corresponding to the food image. Specifically, first, the execution subject may preprocess the food image to obtain a preprocessed food image. As an example, the execution subject may perform noise reduction on the food image, and use the noise-reduced food image as the preprocessed food image. The execution subject may also crop the food image, and use the cropped food image as the preprocessed food image. Then, the execution subject may perform feature extraction on the preprocessed food image through image feature extraction technology to generate image feature information corresponding to the food image.
[0029] In some optional implementations of some embodiments, the execution subject may generate the image feature information corresponding to the food image through the following steps:
[0030] In the first step, the food image is divided according to a preset division method to obtain a food image block set. The preset division method may be a grid division. As an example, the execution subject may divide the food image according to a 5×5 grid division method to obtain a food image block set including 25 food image blocks. The food image blocks in the food image block set may be evenly distributed and have the same shape and size.
[0031] In the second step, for each food image block in the food image block set, the following first generation step is performed:
[0032] The first sub-step is to generate image block feature information corresponding to the food image block. The image block feature information may represent the feature semantic content of the food image block. In practice, the execution subject may generate the image block feature information corresponding to the food image block by image feature extraction technology.
[0033] The second sub-step is to generate the image block position information corresponding to the above-mentioned food image block. The above-mentioned image block position information may be the serial number of the above-mentioned food image block being marked. In practice, firstly, the above-mentioned execution subject may mark each food image block in the above-mentioned food image block set according to a preset marking sequence, and obtain the serial number corresponding to each food image block in the above-mentioned food image block set. Then, the above-mentioned execution subject may determine the serial number corresponding to the above-mentioned food image block as the above-mentioned image block position information.
[0034] The third sub-step is to combine the above-mentioned image block feature information and the above-mentioned image block position information into target image block feature information.
[0035] The third step is to determine the obtained feature information of each target image block as image feature information. The above image feature information may be feature information in vector form. As an example, the above image feature information may be Among them, the above Indicates the above-mentioned image feature information. Represents the target image block feature information. Among them, Represents the feature information of an image block. Indicates the image block location information, The value range of . The above The size of can represent the number of target image block feature information contained in the above image feature information. Indicates the target image block feature information whose image block position information is 2.
[0036] Step 103: Perform first dynamic position coding on the image feature information to obtain image position feature information.
[0037] In some embodiments, the execution entity may perform a first dynamic position encoding on the image feature information to obtain image position feature information.
[0038] In some optional implementations of some embodiments, the execution subject may perform first dynamic position encoding on the image feature information through the following steps to obtain the image position feature information:
[0039] The first step is to obtain a preset even dimension weight coefficient and a preset odd dimension weight coefficient. Among them, the even dimension weight coefficient can be a weight coefficient under an even dimension. The even dimension can represent that the index corresponding to the element in the feature information is an even number. The odd dimension weight coefficient can be a weight coefficient under an odd dimension. The odd dimension can represent that the index corresponding to the element in the feature information is an odd number. The even dimension weight coefficient and the odd dimension weight coefficient can be pre-trained weight coefficients.
[0040] In the second step, for each target image block feature information in the above image feature information, the following first processing step is performed:
[0041] The first sub-step is to generate local feature information corresponding to the target image block feature information based on the image feature information. In practice, first, the execution subject can obtain at least one target image block feature information corresponding to the target image block feature information from the image feature information to obtain a first image block feature information set. Specifically, the first image block feature information set can be determined by the following steps: First, the execution subject can use the surrounding image blocks of the image block corresponding to the target image block feature information as the first image block set. Among them, the surrounding image blocks of the image block can include four image blocks that are immediately adjacent to the image block in the upper, lower, left, and right directions. Secondly, the execution subject can determine the target image block feature information corresponding to each first image block in the first image block set as the first image block feature information set. Then, for each first image block feature information in the first image block feature information set, in response to determining that the similarity value between the first image block feature information and the target image block feature information is greater than or equal to a preset threshold, the execution subject can determine the first image block feature information as the second image block feature information. Specifically, the execution subject may determine the similarity value between the first image block feature information and the target image block feature information by the cosine similarity method. Here, there is no limitation on the specific setting of the preset threshold. Next, for each second image block feature information in the determined second image block feature information set, the execution subject may determine the product of the preset weight value and the second image block feature information as the local sub-feature information. Among them, the preset weight value increases as the similarity value corresponding to the second image block feature information increases. Finally, the execution subject may determine the sum of each determined local sub-feature information as the local feature information.
[0042] The second sub-step is to perform the following second processing step for each element in the above target image block feature information:
[0043] Sub-step one, based on the image block position information included in the target image block feature information, the feature dimension of the target image block feature information and the element position information corresponding to the element, generate a first frequency adjustment factor. Wherein, the target image block feature information may be feature information in the form of a vector. The feature dimension of the target image block feature information may be the number of elements contained in the vector corresponding to the target image block feature information. The element position information may be the index of the element in the target image block feature information. The first frequency adjustment factor may be a numerical value generated for the image block position information and the element position information corresponding to the element. The first frequency adjustment factor helps to capture the relationship between different element positions by using the periodicity of trigonometric functions in the subsequent process, so that the target image block feature information has different frequencies at different element positions. The first frequency adjustment factor may be determined according to the following formula: .
[0044] Among them, the above represents the first frequency adjustment factor. Indicates the image block position information included in the above target image block feature information. Indicates the element position information corresponding to the element in the above target image block feature information. The feature dimension representing the feature information of the target image block.
[0045] Sub-step 2: in response to determining that the above element position information is an even number, executing the following second generating step:
[0046] First, the product of the element corresponding to the element position information in the local feature information and the even-dimension weight coefficient is determined as the first dynamic feature information. The first dynamic feature information can be determined according to the following formula: .
[0047] Among them, the above represents the above even dimension weight coefficient. Indicates the image block location information as The local feature information corresponding to the feature information of the target image block. Indicates that the element position information in the above local feature information is The elements of is an even number. Represents the first dynamic feature information.
[0048] Secondly, based on the first dynamic feature information and the first frequency adjustment factor, first dynamic adjustment feature information is generated. The first dynamic feature adjustment information can be determined according to the following formula: .
[0049] Among them, the above represents the first dynamic adjustment feature information, wherein, is an even number. Express and The sum of is subjected to sine operation.
[0050] Sub-step three, in response to determining that the above element position information is an odd number, executing the following third generation step:
[0051] First, the product of the element corresponding to the element position information in the local feature information and the odd dimension weight coefficient is determined as the second dynamic feature information. The second dynamic feature information can be determined according to the following formula: .
[0052] Among them, the above represents the odd-dimensional weight coefficient. Indicates the image block location information as The local feature information corresponding to the feature information of the target image block. Indicates that the element position information in the above local feature information is The elements of is an odd number. Represents the second dynamic feature information.
[0053] Secondly, based on the second dynamic characteristic information and the first frequency adjustment factor, second dynamic adjustment characteristic information is generated. The second dynamic characteristic adjustment information can be determined according to the following formula: .
[0054] Among them, the above represents the second dynamic adjustment feature information, wherein, is an odd number. Express and The cosine operation is performed on the sum of .
[0055] In the third sub-step, the generated first dynamic adjustment feature information and the generated second dynamic adjustment feature information are integrated to obtain the image block dynamic adjustment feature information. In practice, the execution subject may integrate the first dynamic adjustment feature information and the second dynamic adjustment feature information according to the element position information corresponding to the first dynamic adjustment feature information and the element position information corresponding to the second dynamic adjustment feature information to obtain the image block dynamic adjustment feature information. As an example, the first dynamic adjustment feature information may be , , the above-mentioned second dynamic adjustment feature information can be , and . Then the above image block dynamically adjusts the feature information to be .
[0056] The fourth sub-step is to fuse the target image block feature information and the image block dynamic adjustment feature information to obtain the image block position feature information. In practice, the execution subject may add the target image block feature information and the image block dynamic adjustment feature information to obtain the image block position feature information.
[0057] The third step is to determine the obtained position feature information of each image block as the above-mentioned image position feature information.
[0058] Step 104: Generate text feature information corresponding to the initial food description information.
[0059] In some embodiments, the execution subject may generate text feature information corresponding to the initial food description information. The text feature information may characterize the features of the initial food description information. In practice, the execution subject may generate text feature information corresponding to the initial food description information by using text feature extraction technology.
[0060] In some optional implementations of some embodiments, the execution subject may generate text feature information corresponding to the initial food description information through the following steps:
[0061] The first step is to perform word segmentation processing on the initial food description information to generate a word segmentation sequence. In practice, the execution subject can perform word segmentation on the initial description information by a rule-based word segmentation method to generate a word segmentation sequence.
[0062] The second step is to perform the following fourth generation step for each word in the above word segmentation sequence:
[0063] The first sub-step is to generate the word segmentation feature information corresponding to the above word segmentation. In practice, the above execution subject can generate the word segmentation feature information corresponding to the above word segmentation by one-hot encoding.
[0064] The second sub-step is to generate the segmentation position information corresponding to the segmentation. The segmentation position information may be the sequence number of the segmentation. In practice, first, the execution subject may mark each segmentation in the segmentation sequence according to a preset marking order to obtain the sequence number corresponding to each segmentation in the segmentation sequence. Then, the execution subject may determine the sequence number corresponding to the segmentation as the segmentation position information.
[0065] The third sub-step is to combine the above-mentioned word segmentation feature information and the above-mentioned word segmentation position information into target word segmentation feature information.
[0066] The third step is to determine the obtained feature information of each target word segmentation as text feature information. The above text feature information can be feature information in vector form. As an example, the above text feature information can be Among them, the above Indicates the above text feature information. Represents the target word segmentation feature information. Among them, Represents word segmentation feature information, Indicates word segmentation position information. The value range of . The above The size of can represent the number of target word segmentation feature information contained in the above text feature information. For example, Indicates the target word segmentation feature information with word segmentation position information of 2.
[0067] Step 105: Perform second dynamic position coding on the text feature information to obtain text position feature information.
[0068] In some embodiments, the execution entity may perform a second dynamic position encoding on the text feature information to obtain text position feature information.
[0069] In some optional implementations of some embodiments, the execution subject may perform second dynamic position encoding on the text feature information through the following steps to obtain text position feature information:
[0070] The first step is to perform the following third processing step for each target word segmentation feature information in the above text feature information:
[0071] The first sub-step is to generate the first adjustment feature information and the second adjustment feature information corresponding to the target word segmentation feature information based on the above-mentioned word segmentation sequence. In practice, the above-mentioned execution subject can input the above-mentioned word segmentation sequence into a pre-trained bidirectional language model to obtain the first adjustment feature information and the second adjustment feature information corresponding to the above-mentioned target word segmentation feature information. Among them, the above-mentioned bidirectional language model can be a BERT model. In practice, firstly, the above-mentioned execution subject can format the above-mentioned word segmentation sequence and input it into the BERT model. Secondly, the BERT model can input the above-mentioned format-converted word segmentation sequence into the hidden layer. The above-mentioned hidden layer can perform forward processing and backward processing on the above-mentioned format-converted word segmentation sequence to obtain the first feature information and the second feature information corresponding to the above-mentioned target word segmentation feature information. Among them, the above-mentioned forward processing can be a semantic understanding processing of the above-mentioned format-converted word segmentation sequence from left to right. The above-mentioned backward processing can be a semantic understanding processing of the above-mentioned format-converted word segmentation sequence from right to left. Finally, the above-mentioned execution subject can determine the above-mentioned first feature information as the first adjustment feature information, and determine the above-mentioned second feature information as the second adjustment feature information.
[0072] The second sub-step is to execute the following fifth generation step for each element in the target word segmentation feature information:
[0073] Sub-step one, based on the word segmentation position information included in the above-mentioned target word segmentation feature information, the feature dimension of the above-mentioned target word segmentation feature information and the element position information corresponding to the above-mentioned element, generate a second frequency adjustment factor. Among them, the above-mentioned target word segmentation feature information can be feature information in the form of a vector. The feature dimension of the above-mentioned target word segmentation feature information can be the number of elements contained in the vector corresponding to the above-mentioned target word segmentation feature information. The above-mentioned element position information can be the index of the above-mentioned element in the above-mentioned target word segmentation feature information. The above-mentioned second frequency adjustment factor can be a numerical value generated for the above-mentioned word segmentation position information and the element position information corresponding to the above-mentioned element. The second frequency adjustment factor helps to capture the relationship between different element positions by using the periodicity of trigonometric functions in the subsequent process, so that the target word segmentation feature information has different frequencies at different element positions. The above-mentioned second frequency adjustment factor can be determined according to the following formula: .
[0074] Among them, the above represents the second frequency adjustment factor. Indicates the word segmentation position information included in the target word segmentation feature information. Indicates the element position information corresponding to the element in the above target word segmentation feature information, The feature dimension representing the above target word segmentation feature information.
[0075] Sub-step 2: in response to determining that the element position information is an even number, generating third dynamic adjustment characteristic information based on the element corresponding to the element position information in the first adjustment characteristic information and the second frequency adjustment factor. The third dynamic adjustment characteristic information can be determined according to the following formula: .
[0076] Among them, the above represents the third dynamic adjustment feature information, wherein, is an even number, the above Indicates the word segmentation position information included in the target word segmentation feature information.
[0077] Above Express and The sum of sine operations is performed. Indicates the first adjustment feature information. Indicates that the element position information in the first adjustment feature information is The elements of Is an even number.
[0078] Sub-step three, in response to determining that the element position information is an odd number, generating fourth dynamic adjustment characteristic information based on the element corresponding to the element position information in the second adjustment characteristic information and the second frequency adjustment factor. The fourth dynamic adjustment characteristic information can be determined according to the following formula: .
[0079] Among them, the above represents the fourth dynamic adjustment characteristic information, wherein, is an odd number, the above Indicates the word segmentation position information included in the target word segmentation feature information. Express and The cosine operation is performed on the sum of Indicates the second adjustment feature information. Indicates that the element position information in the second adjustment feature information is The elements of Is an odd number.
[0080] In the third sub-step, the generated third dynamic adjustment feature information and the generated fourth dynamic adjustment feature information are integrated to obtain the word segmentation dynamic adjustment feature information. In practice, the execution subject can integrate the third dynamic adjustment feature information and the fourth dynamic adjustment feature information according to the element position information corresponding to the third dynamic adjustment feature information and the element position information corresponding to the fourth dynamic adjustment feature information to obtain the word segmentation dynamic adjustment feature information. As an example, the third dynamic adjustment feature information can be , , the above-mentioned fourth dynamic adjustment characteristic information can be , and . Then the above word segmentation dynamic adjustment feature information can be .
[0081] The fourth sub-step is to merge the target word segmentation feature information and the word segmentation dynamic adjustment feature information to obtain the word segmentation position feature information. In practice, the execution subject may add the target word segmentation feature information and the word segmentation dynamic adjustment feature information to obtain the word segmentation position feature information.
[0082] In the second step, the obtained position feature information of each word segmentation is determined as the above-mentioned text position feature information.
[0083] Step 106: fuse the image position feature information and the text position feature information to obtain multimodal fusion feature information.
[0084] In some embodiments, the execution subject may fuse the image position feature information and the text position feature information to obtain multimodal fusion feature information. In practice, the execution subject may splice the image position feature information and the text position feature information to obtain multimodal fusion feature information.
[0085] Step 107, retrieving feature supplementary information corresponding to the multimodal fusion feature information that represents and describes the food features.
[0086] In some embodiments, the execution entity may retrieve the feature supplementary information corresponding to the multimodal fusion feature information that represents and describes the food features.
[0087] In some optional implementations of some embodiments, the execution subject may retrieve the feature supplementary information representing and describing the food features corresponding to the multimodal fusion feature information through the following steps:
[0088] The first step is to identify the food name of the food image to obtain a first food name set. In practice, first, the execution subject can identify the food displayed in the food image by image recognition technology to obtain at least one food name corresponding to the food. Then, the execution subject can determine the at least one food name as the first food name set.
[0089] The second step is to extract the food names in the initial food description information to obtain a second food name set. In practice, first, the execution subject can segment the initial food description information using a rule-based segmentation method. Then, the execution subject can extract at least one segmentation that is a food name from each segmentation obtained to obtain a second food name set.
[0090] In the third step, the union of the first food name set and the second food name set is determined as the target food name set.
[0091] Step 4: For each target food name in the target food name set, obtain the food characteristic information group corresponding to the target food name from the pre-built food information database to obtain the candidate food characteristic information group. The food information database stores the full amount of food names and the food characteristic information group corresponding to each food name. The food characteristic information group contains at least one food characteristic information.
[0092] Step 5: for each candidate food feature information in the obtained candidate food feature information set, perform the following determination steps:
[0093] The first sub-step is to generate a similarity value corresponding to the multimodal fusion feature information and the candidate food feature information. In practice, the execution subject may generate a similarity value corresponding to the multimodal fusion feature information and the candidate food feature information by using a cosine similarity method.
[0094] In the second sub-step, in response to the similarity value satisfying the first preset condition, the candidate food feature information is determined as the target food feature information. The first preset condition may be that the similarity value is less than or equal to the first preset threshold and the similarity value is greater than or equal to the second preset threshold. Here, the specific settings of the first preset threshold and the second preset threshold are not limited.
[0095] The sixth step is to determine the determined characteristic information of each target food as the characteristic supplement information. In practice, the execution subject may fuse the characteristic information of each target food to obtain the characteristic supplement information.
[0096] In the process of adopting technical solutions to solve the problems mentioned in the background technology, the following problems are often accompanied:
[0097] When retrieving the feature supplementary information corresponding to the multimodal fusion feature information that represents and describes the food characteristics, the retrieval accuracy of the feature supplementary information describing the food characteristics cannot be effectively guaranteed, resulting in the subsequent determined food description information not matching the food image and the initial food description information.
[0098] In the face of the above problems, the inventor decided to adopt the following solutions:
[0099] In some optional implementations of some embodiments, the execution subject may retrieve the feature supplementary information representing and describing the food features corresponding to the multimodal fusion feature information through the following steps:
[0100] In the first step, the multimodal fusion feature information is input into a pre-trained food identification generation model to generate at least one food identification and at least one food identification probability. The food identification generation model is a classification model for generating food identification. The food identification generation model and the target food description information database are uniformly set. The food identification generation model is a classification model trained based on a multimodal sample set and a gradient descent method. For example, the food identification generation model can be a multi-layer series convolution layer + a 2-branch fully connected layer. The 2-branch fully connected layer includes: a fully connected layer for outputting food identification and a fully connected layer for outputting food identification probability. The target food description information database is a database that stores the full food description information corresponding to the full food set. The full food description information can be a full description of the corresponding food. The full food description information has 3 levels of food description information. The user's understanding of the first level of food description information is higher than the second level of food description information. The user's understanding of the second level of food description information is higher than the third level of food description information.
[0101] The second step is to generate at least one food description information query statement corresponding to the at least one food identifier. There is a one-to-one correspondence between the food identifier in the at least one food identifier and the food description information query statement in the at least one food description information query statement. The food description information query statement may be an SQL query statement.
[0102] The third step is to execute the at least one food description information query statement to obtain the corresponding at least one full food description information from the target food description information database, wherein the full food description information in the at least one full food description information and the food description information query statement in the at least one food description information query statement have a one-to-one correspondence.
[0103] The fourth step is to sort the at least one food identifier in descending order according to the at least one food identifier probability to obtain a food identifier sequence.
[0104] The fifth step is to select food identifiers whose corresponding food identifier probabilities are higher than the target probability value from the above food identifier sequences to obtain the remaining food identifier sequences.
[0105] Step 6: For each food identifier in the remaining food identifier sequence, perform the following sixth generation step:
[0106] The first sub-step is to determine the full food description information corresponding to the above food label as the target full food description information.
[0107] The second sub-step is to determine the first-level food description vector, the second-level food description vector and the third-level food description vector corresponding to the above-mentioned target full food description information. Among them, the first-level food description vector is the food description feature vector of the first-level food description information. The second-level food description vector is the food description feature vector of the second-level food description information. The first-level food description vector and the second-level food description vector can be pre-generated offline and stored in the food description vector repository.
[0108] The third sub-step is to determine the feature similarity between the first-level food description vector and the multimodal fusion feature information as the first feature similarity. The feature similarity may be cosine similarity. The first feature similarity may be a value between 0 and 1, and the higher the corresponding value, the more similar the features of the first-level food description vector and the multimodal fusion feature information are.
[0109] The fourth sub-step is to determine the feature similarity between the second-level food description vector and the multimodal fusion feature information as the second feature similarity. The second feature similarity can be a value between 0 and 1, and the higher the corresponding value, the more similar the features of the second-level food description vector and the multimodal fusion feature information are.
[0110] The fifth sub-step is to determine the feature similarity between the third-level food description vector and the multimodal fusion feature information as the third feature similarity. The third feature similarity can be a value between 0 and 1, and the higher the corresponding value, the more similar the features of the third-level food description vector and the multimodal fusion feature information are.
[0111] The sixth sub-step is to determine the sequence position of the above-mentioned food identifier in the above-mentioned remaining food identifier sequence.
[0112] A seventh sub-step is to combine the column position, the first feature similarity, the second feature similarity and the third feature similarity to generate an information group.
[0113] In the seventh step, the obtained information group set is input into a pre-set food description information determination script to filter out the target information group. The food description information determination script may be a source file for determining food description information set by relevant technical personnel based on food-focused knowledge.
[0114] Step 8: Determine the target food image corresponding to the target information group.
[0115] In the ninth step, the target food image and the food image are annotated with positive points and negative points to obtain a first annotated image and a second annotated image. The positive points are points on food in the image, and the negative points are points on non-food.
[0116] In step 10, the first annotated image and the second annotated image are input into an interactive image similarity information generation model to generate image similarity information. The image similarity information may be a value between 0 and 1. The image similarity information generation model may be a SAM model.
[0117] In the eleventh step, in response to determining that the image similarity information is higher than the target value, the full amount of food description information corresponding to the target information group is determined as feature supplementary information.
[0118] The above-mentioned "in some optional implementation methods of some embodiments", as one of the inventive points, solves another technical problem of the present disclosure, "the retrieval accuracy of the feature supplementary information describing the food characteristics cannot be effectively guaranteed, resulting in the subsequent determined food description information not matching the food image and the initial food description information". Based on this, the present disclosure, firstly, through the food identification generation model, can accurately and preliminarily determine at least one food and at least one food probability that matches the multimodal fusion feature information. Then, by dividing the food into multiple levels of knowledge, the food knowledge at each level is compared in turn to determine the similarity of the food knowledge at multiple levels, and the similarity between each food and the multimodal fusion feature information can be obtained (reflected by each information in the information group). Finally, through the pre-set food description information determination script and the interactive image similarity information generation model, the food is screened in many aspects to finally accurately determine the feature supplementary information corresponding to the multimodal fusion feature information.
[0119] Step 108: Generate target food description information based on the multimodal fusion feature information and feature supplementary information.
[0120] In some embodiments, the execution subject may generate target food description information based on the multimodal fusion feature information and the feature supplementary information. In practice, first, the execution subject may use natural language generation technology to generate multimodal description information and supplementary description information corresponding to the multimodal feature information and the feature supplementary information, respectively. Secondly, the execution subject may determine the multimodal description information and the supplementary description information as the target food description information.
[0121] In some optional implementations of some embodiments, the execution subject may generate the target food description information based on the multimodal fusion feature information and the feature supplementary information through the following steps:
[0122] The first step is to integrate the feature supplement information and the multimodal fusion feature information to generate target fusion feature information. In practice, first, the execution subject can remove the feature information in the multimodal fusion feature information that is repeated with the feature supplement information to obtain the multimodal fusion feature information without repeated feature information. Secondly, the execution subject can splice the feature supplement information and the multimodal fusion feature information without repeated feature information to generate target fusion feature information.
[0123] In the second step, the target fusion feature information is input into a text generation model to obtain the target food description information. The text generation model may be a neural network model used to generate text content corresponding to the target food description information. For example, the text generation model may be a Transformer language model.
[0124] Step 109: in response to determining that the food identification request information includes a target playback identifier, obtaining the target food audio corresponding to the target food description information from a target audio database.
[0125] In some embodiments, the execution entity may obtain the target food audio corresponding to the target food description information from the target audio database in response to determining that the food identification request information includes an indication that the target food description information is played in the form of audio. In practice, the food identification request information may also include a target playback identifier. Among them, the target playback identifier may indicate that the target food description information is played in the form of audio. The target audio database may store a database of audio resources corresponding to all foods. In practice, an audio playback button may be provided on the user terminal. The target user may click on the audio playback button to cause the user terminal to generate a target playback identifier indicating that the target food description information is played in the form of audio.
[0126] Step 1010, sending the target food description information and the target food audio to the user terminal, so that the user terminal can display the target food description information and play the target food audio.
[0127] In some embodiments, the execution entity may send the target food description information and the target food audio to the user terminal, so that the user terminal can display the target food description information and play the target food audio.
[0128] For example, Figure 4 This is a schematic screenshot of an internal test operation page for displaying food description information according to the food description information display method disclosed in the present invention.
[0129] Figure 4 An interface for interaction between a user terminal and a system corresponding to the food description information display method of the present disclosure is shown. "User terminal" can be various electronic devices with a display screen and supporting information browsing, including but not limited to smartphones and tablets. "AI nutritionist" is the name of the interactive object that interacts with the user in the user interaction interface. The target user can input the image and text information of the food he wants to know through the "text and image input box". "Food image" can be the image of the food that the target user wants to know. "Initial food description information" can be the text description information of the food corresponding to the food image entered by the target user. For example, the above-mentioned food image can be an image of a stinky mandarin fish, and the above-mentioned initial food description information can be "Why does this fish smell stinky? Will it be delicious?". "Target food description information" is generated based on the food description information display method. The target food description information can more comprehensively describe the taste, nutritional ingredients and other information of the food. For example, the target food description information may be “This dish is stinky mandarin fish, and it looks delicious! Stinky mandarin fish is a traditional Chinese dish known for its unique fermented flavor and rich taste. It has a total weight of about 275 grams and contains 346.5 kilocalories, as well as 4.54 grams of carbohydrates, 13.8 grams of fat, and 51.92 grams of protein. In addition, it also contains some cellulose (0.85 grams)”.
[0130] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the target food description information obtained by the food description information display method of some embodiments of the present disclosure is more accurate and comprehensive, and can better meet the needs of users. Specifically, the reason why the obtained food description information is often not accurate enough and difficult to meet the needs of users is that the food description information generated only by recognizing the food image often cannot fully describe the information of the food displayed in the food image, and the image feature information between foods often has a relatively similar situation, so that the food description information is generated only from the perspective of food image recognition, the obtained food features are relatively single, and the generated food description information may have deviations, which makes it difficult to meet the needs of users. Based on this, the food description information display method of some embodiments of the present disclosure, first, in response to receiving the food recognition request information sent by the user terminal, receives the food image and the initial food description information, wherein the initial food description information is the text description information of the food corresponding to the food image. In this way, the image of the food that the user wants to know and the existing food description information can be obtained, which provides basic data for the subsequent generation of image feature information and text feature information. Then, the image feature information corresponding to the above food image is generated, and the above image feature information is firstly dynamically encoded to obtain the image position feature information. Thus, by performing the first dynamic position encoding on the image feature information, the obtained image position feature information can represent the overall structural information of the image. Secondly, the text feature information corresponding to the above-mentioned initial food description information is generated, and the above-mentioned text feature information is subjected to the second dynamic position encoding to obtain the text position feature information. Thus, by performing the second dynamic position encoding on the text feature information, the obtained text position feature information can represent the contextual relationship of the initial food description information. Next, the above-mentioned image position feature information and the above-mentioned text position feature information are fused to obtain multimodal fusion feature information. Thus, the feature information of the image and the text are combined to form a multimodal feature representation, so that the characteristics of the food can be better represented. Further, the feature supplementary information corresponding to the above-mentioned multimodal fusion feature information that represents and describes the food characteristics is retrieved. Thus, according to the fused feature information, the relevant feature supplementary information is retrieved in the database storing the full amount of food feature information, so that the feature supplementary information and the multimodal fusion feature information can be fused to obtain more comprehensive features. Furthermore, based on the above-mentioned multimodal fusion feature information and the above-mentioned feature supplementary information, the target food description information is generated. Therefore, the target food information can clearly represent the taste, nutritional components and other information of the food, thereby obtaining a comprehensive and rich description of the target food that meets the needs of users as much as possible.Finally, in response to determining that the food identification request information includes a target playback identifier, the target food audio corresponding to the target food description information is obtained from the target audio database, wherein the target playback identifier represents the target food description information to be played in the form of audio, and the target food description information and the target food audio are sent to the user terminal so that the user terminal can display the target food description information and play the target food audio. Thus, the target food description information is displayed, and the user can intuitively see the detailed description information of the food. The target food description information is played in audio, which expands the audience range and improves the availability of the target food description information and user satisfaction. Based on this, by dynamically position encoding the image feature information and the text feature information respectively, and fusing the image position feature information and the text position feature information, and retrieving the relevant feature supplementary information from the pre-constructed database, the obtained target food description information is more accurate and comprehensive, and can better meet the needs of users.
[0131] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a food description information display device, and these device embodiments are Figure 1 Corresponding to the method embodiments shown, the food description information display device can be specifically applied to various electronic devices.
[0132] like Figure 2As shown, a food description information display device 200 includes: a receiving unit 201, a first generating unit 202, a first executing unit 203, a second generating unit 204, a second executing unit 205, a fusion unit 206, a third executing unit 207, a third generating unit 208, an acquiring unit 209 and a fourth executing unit 2010. The receiving unit 201 is configured to: receive a food image and initial food description information in response to receiving a food identification request information sent by a user terminal, wherein the initial food description information is text description information of the food corresponding to the food image. The first generating unit 202 is configured to: generate image feature information corresponding to the food image. The first executing unit 203 is configured to: perform a first dynamic position encoding on the image feature information to obtain image position feature information. The second generating unit 204 is configured to: generate text feature information corresponding to the initial food description information. The second executing unit 205 is configured to: perform a second dynamic position encoding on the text feature information to obtain text position feature information. The fusion unit 206 is configured to: fuse the above-mentioned image position feature information and the above-mentioned text position feature information to obtain multimodal fusion feature information. The third execution unit 207 is configured to: retrieve the feature supplementary information corresponding to the above-mentioned multimodal fusion feature information that represents the description of food features. The third generation unit 208 is configured to: generate target food description information based on the above-mentioned multimodal fusion feature information and the above-mentioned feature supplementary information. The acquisition unit 209 is configured to: in response to determining that the above-mentioned food recognition request information includes a target playback identifier, obtain the target food audio corresponding to the above-mentioned target food description information from the target audio database, wherein the above-mentioned target playback identifier represents that the above-mentioned target food description information is played in the form of audio. The fourth execution unit 2010 is configured to: send the above-mentioned target food description information and the above-mentioned target food audio to the above-mentioned user terminal, so that the above-mentioned user terminal displays the above-mentioned target food description information and plays the above-mentioned target food audio by audio.
[0133] It is understandable that the units described in the food description information display device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the food description information display device 200 and the units included therein, and will not be described in detail here.
[0134] Reference below Figure 3 , which shows a structural schematic diagram of an electronic device (eg, an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0135] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 to a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0136] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as required.
[0137] In particular, according to some embodiments of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are executed.
[0138] It should be noted that the computer-readable medium mentioned above in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0139] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0140] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: receives a food image and initial food description information in response to receiving a food identification request information sent by a user terminal, wherein the above-mentioned initial food description information is text description information of the food corresponding to the above-mentioned food image; generates image feature information corresponding to the above-mentioned food image; performs a first dynamic position encoding on the above-mentioned image feature information to obtain image position feature information; generates text feature information corresponding to the above-mentioned initial food description information; performs a second dynamic position encoding on the above-mentioned text feature information to obtain text position feature information; performs a second dynamic position encoding on the above-mentioned image position feature information and obtains text position feature information; The feature information is fused to obtain multimodal fusion feature information; the feature supplementary information representing and describing the food features corresponding to the multimodal fusion feature information is retrieved; based on the multimodal fusion feature information and the feature supplementary information, the target food description information is generated; in response to determining that the food identification request information includes a target playback identifier, the target food audio corresponding to the target food description information is obtained from a target audio database, wherein the target playback identifier indicates that the target food description information is played in the form of audio; the target food description information and the target food audio are sent to the user terminal, so that the user terminal displays the target food description information and plays the target food audio by audio.
[0141] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0142] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0143] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including a receiving unit, a first generating unit, a first executing unit, a second generating unit, a second executing unit, a fusion unit, a third executing unit, a third generating unit, an acquiring unit, and a fourth executing unit. Among them, the names of these units do not constitute a limitation on the units themselves in certain cases, for example, the receiving unit may also be described as "a unit that receives food images and initial food description information in response to receiving food identification request information sent by a user terminal".
[0144] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0145] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for displaying food description information, comprising: In response to receiving food identification request information sent by a user terminal, receiving a food image and initial food description information, wherein the initial food description information is text description information of the food corresponding to the food image; generating image feature information corresponding to the food image; Performing a first dynamic position encoding on the image feature information to obtain image position feature information; Generating text feature information corresponding to the initial food description information; Performing second dynamic position coding on the text feature information to obtain text position feature information; Fusing the image position feature information and the text position feature information to obtain multimodal fusion feature information; Retrieving feature supplementary information corresponding to the multimodal fusion feature information that characterizes and describes food features; Generate target food description information based on the multimodal fusion feature information and the feature supplementary information; In response to determining that the food identification request information includes a target playback identifier, acquiring a target food audio corresponding to the target food description information from a target audio database, wherein the target playback identifier indicates that the target food description information is played in the form of audio; The target food description information and the target food audio are sent to the user terminal, so that the user terminal displays the target food description information and plays the target food audio.
2. The method according to claim 1, wherein: The generating of image feature information corresponding to the food image includes: Dividing the food image according to a preset division method to obtain a food image block set; For each food image block in the food image block set, the following first generation step is performed: Generating image block feature information corresponding to the food image block; Generating image block position information corresponding to the food image block; Combining the image block feature information and the image block position information into target image block feature information; The obtained feature information of each target image block is determined as image feature information.
3. The method according to claim 2, wherein: The performing first dynamic position encoding on the image feature information to obtain the image position feature information includes: Obtaining a preset even dimension weight coefficient and a preset odd dimension weight coefficient; For each target image block feature information in the image feature information, the following first processing step is performed: Based on the image feature information, generating local feature information corresponding to the target image block feature information; For each element in the target image block feature information, the following second processing step is performed: generating a first frequency adjustment factor based on the image block position information included in the target image block feature information, the feature dimension of the target image block feature information, and the element position information corresponding to the element; In response to determining that the element position information is an even number, the following second generating step is performed: Determine the product of the element corresponding to the element position information in the local feature information and the even-dimensional weight coefficient as the first dynamic feature information; generating first dynamic adjustment characteristic information based on the first dynamic characteristic information and the first frequency adjustment factor; In response to determining that the element position information is an odd number, the following third generating step is performed: Determine the product of the element corresponding to the element position information in the local feature information and the odd-dimension weight coefficient as the second dynamic feature information; generating second dynamic adjustment characteristic information based on the second dynamic characteristic information and the first frequency adjustment factor; Integrate the generated first dynamic adjustment feature information and the generated second dynamic adjustment feature information to obtain the image block dynamic adjustment feature information; The target image block feature information and the image block dynamic adjustment feature information are merged to obtain the image block position feature information; The obtained position feature information of each image block is determined as image position feature information.
4. The method according to claim 1, wherein: The generating of text feature information corresponding to the initial food description information includes: Performing word segmentation processing on the initial food description information to generate a word segmentation sequence; For each word in the word segmentation sequence, the following fourth generation step is performed: Generate word segmentation feature information corresponding to the word segmentation; Generate word segmentation position information corresponding to the word segmentation; Combining the word segmentation feature information and the word segmentation position information into target word segmentation feature information; The obtained feature information of each target word segmentation is determined as text feature information.
5. The method according to claim 4, wherein: The step of performing a second dynamic position encoding on the text feature information to obtain the text position feature information includes: For each target word segmentation feature information in the text feature information, the following third processing step is performed: Based on the word segmentation sequence, generating first adjustment feature information and second adjustment feature information corresponding to the target word segmentation feature information; For each element in the target word segmentation feature information, the following fifth generation step is performed: Generate a second frequency adjustment factor based on the word segmentation position information included in the target word segmentation feature information, the feature dimension of the target word segmentation feature information, and the element position information corresponding to the element; In response to determining that the element position information is an even number, generating third dynamic adjustment characteristic information based on the element corresponding to the element position information in the first adjustment characteristic information and the second frequency adjustment factor; In response to determining that the element position information is an odd number, generating fourth dynamic adjustment characteristic information based on the element corresponding to the element position information in the second adjustment characteristic information and the second frequency adjustment factor; Integrate the generated third dynamic adjustment feature information and the generated fourth dynamic adjustment feature information to obtain word segmentation dynamic adjustment feature information; The target word segmentation feature information and the word segmentation dynamic adjustment feature information are merged to obtain word segmentation position feature information; The obtained position feature information of each word segmentation is determined as text position feature information.
6. The method according to claim 1, wherein: The retrieving the feature supplementary information corresponding to the multimodal fusion feature information that characterizes and describes the food features includes: Performing food name recognition on the food image to obtain a first food name set; Extracting food names from the initial food description information to obtain a second food name set; determining a union of the first food name set and the second food name set as a target food name set; For each target food name in the target food name set, obtaining a food feature information group corresponding to the target food name from a pre-constructed food information database to obtain a candidate food feature information group, wherein the food information database stores a full set of food names and a food feature information group corresponding to each food name; For each candidate food feature information in the obtained candidate food feature information set, the following determination steps are performed: Generating a similarity value corresponding to the multimodal fusion feature information and the candidate food feature information; In response to the similarity value satisfying a first preset condition, determining the candidate food feature information as target food feature information; The determined characteristic information of each target food is determined as characteristic supplementary information.
7. The method according to claim 1, wherein: The generating target food description information based on the multimodal fusion feature information and the feature supplementary information includes: Integrate the feature supplement information and the multimodal fusion feature information to generate target fusion feature information; The target fusion feature information is input into a text generation model to obtain target food description information.
8. A food description information display device, comprising: a receiving unit, configured to receive a food image and initial food description information in response to receiving food identification request information sent by a user terminal, wherein the initial food description information is text description information of the food corresponding to the food image; A first generating unit is configured to generate image feature information corresponding to the food image; A first execution unit is configured to perform a first dynamic position encoding on the image feature information to obtain image position feature information; A second generating unit is configured to generate text feature information corresponding to the initial food description information; A second execution unit is configured to perform a second dynamic position encoding on the text feature information to obtain text position feature information; a fusion unit configured to fuse the image position feature information and the text position feature information to obtain multimodal fusion feature information; A third execution unit is configured to retrieve feature supplementary information corresponding to the multimodal fusion feature information that characterizes and describes food features; a third generating unit, configured to generate target food description information based on the multimodal fusion feature information and the feature supplementary information; an acquisition unit, configured to acquire, in response to determining that the food identification request information includes a target playback identifier, a target food audio corresponding to the target food description information from a target audio database, wherein the target playback identifier indicates that the target food description information is played in the form of audio; The fourth execution unit is configured to send the target food description information and the target food audio to the user terminal, so that the user terminal displays the target food description information and plays the target food audio.
9. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Controllable image description method, system and equipment based on multi-modal fusion and medium
CN115631394A
Multi-stage fusion multi-modal social user position inference method
CN118171735A