Data generation method, system and equipment
By identifying the pixel proportion and secondary splitting conditions of the visual elements of the image, determining the description weights and attributes of the visual elements and sub-elements, and generating structured directed text, it solves the problem that text descriptions in the prior art are difficult to highlight key points and content randomness, and realizes accurate, detailed and directional text output.
Patent Information
- Application Number
- CN202510694787.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The text description generated by the prior art after image recognition is difficult to highlight the key content of the image, and the generated text content is relatively random and cannot meet the directed output requirements for text content in specific scenarios.
By identifying the visual elements of the input image, the description weight is determined based on the pixel proportion of each visual element, and when the secondary splitting conditions are met, the element proportion and description attributes of each child element in the visual element are determined to generate structured directional text.
It realizes the accuracy and meticulousness of text descriptions, and can allocate space according to the importance of the image content, avoid random content generation, meet the directional output requirements in different scenarios, and improve the logic and readability of the text.
Smart Images

Figure CN120219770A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technologies, and in particular to a method, system, and device for data generation. Background Art
[0002] In the field of artificial intelligence, text generation technology based on image recognition has been widely used, and is commonly used in scenarios such as image content interpretation and image retrieval assistance.
[0003] In the prior art, when AI generates descriptive text after recognizing a picture, it usually adopts relatively general algorithms and models, lacking refined processing of image content. Specifically, first, it cannot effectively distinguish the importance levels of various elements in the image, resulting in the generated text description being difficult to highlight the key content of the image. For example, in an image containing a person and a landscape, it cannot reasonably allocate the space for description according to the difference in the proportion of the two in the image. Second, the generated text content has a large degree of randomness, and it may be difficult to perform targeted output according to user needs, unable to meet the requirements for targeted output of text content in specific scenarios.
[0004] Therefore, how to achieve refined processing of picture content and improve the accuracy of text description has become an urgent problem to be solved. Summary of the Invention
[0005] The present invention provides a method, system, and device for data generation, which can achieve refined processing of picture content and improve the accuracy of text description.
[0006] In a first aspect of the present invention, a method for data generation is provided, including: Identifying visual elements of an input image, and determining description weights according to the pixel ratio of each visual element; When the pixel ratio meets the quadratic splitting condition, determining the element ratio of each sub-element in the visual element, and the description attribute corresponding to the sub-element, where the description attribute includes a positive attribute and a negative attribute; Processing the description weights according to the description attribute and the element ratio to determine the description sub-weights of each sub-element; Performing correlation processing on the required space based on the description weights and / or description sub-weights, and the levels of the visual elements and / or sub-elements to generate a structured and targeted text.
[0007] Optionally, in a possible implementation manner of the first aspect, the identifying visual elements of the input image and determining description weights according to the pixel ratio of each visual element includes: Receiving the input image and identifying multiple visual elements of the input image; Obtain the number of pixels of each of the visual elements, and generate a pixel ratio based on the ratio of the number of pixels to the total number of pixels of the input image; Generate a proportion ratio based on the pixel ratio of each of the visual elements, and generate a corresponding description weight according to the proportion ratio.
[0008] Optionally, in a possible implementation manner of the first aspect, when the pixel ratio meets the secondary splitting condition, determine the element ratio of each sub-element in the visual element and the description attribute corresponding to the sub-element, including: If the pixel ratio corresponding to the visual element is greater than or equal to the ratio threshold, determine that the visual element meets the secondary splitting condition; Obtain the element area corresponding to each sub-element in the corresponding visual element and the element area corresponding to the visual element; According to the ratio between each element area and the element area, obtain the element ratio corresponding to each sub-element, and determine the element ratio of each sub-element based on the element ratio of each sub-element; Receive the click information of the sub-element from the client, retrieve the attribute window corresponding to the sub-element and send it to the client, and determine the description attribute corresponding to the sub-element according to the click information of any description attribute in the attribute window by the client.
[0009] Optionally, in a possible implementation manner of the first aspect, the correlation processing of the required space based on the description weight and / or description sub-weight, and the level of the visual element and / or sub-element to generate a structured directional text includes: Process the preset space based on the description weight to obtain the first space corresponding to each visual element, and split the corresponding first space according to the description sub-weight to obtain the second space and description attribute corresponding to each description sub-weight; Parse the picture content corresponding to the visual element according to the first space to generate a first text, and parse the picture content corresponding to the sub-element according to the second space and description attribute to generate a second text; Construct a parent node with the input image, sub-nodes with the visual elements, and grandchild nodes with the sub-elements, and generate an architecture tree according to the parent node, sub-nodes, and grandchild nodes; Associate the first text and the second text according to the architecture tree to generate a structured directional text.
[0010] Optionally, in a possible implementation manner of the first aspect, the parsing the picture content corresponding to the sub-element according to the second space and description attribute to generate a second text includes: Parse the picture content corresponding to each sub-element, and perform directional adjustment on the picture content according to the description attribute to obtain the adjusted content; Process the adjusted content according to the second piece to generate a second text.
[0011] Optionally, in a possible implementation of the first aspect, after receiving the input image and identifying multiple visual elements of the input image, it further includes: Obtain the element types of the visual elements, classify the visual elements according to the element types to obtain an element set; Obtain the pixel proportion of each visual element in the element set, and perform a descending order sorting according to the pixel proportion to obtain an element sequence; Determine the judgment ratio of the pixel proportion of the last visual element in the element sequence to the pixel proportion of the first visual element, and remove the last visual element in the element sequence when the judgment ratio is less than a preset value; Repeat the above steps. When the judgment ratio is greater than or equal to the preset value, use the remaining visual elements in the element sequence as new visual elements.
[0012] Optionally, in a possible implementation of the first aspect, the parsing the picture content corresponding to the visual element according to the first piece to generate a first text includes: Divide the first piece into a position piece and a content piece, obtain the relative position information of the visual element, and generate a position text corresponding to the position piece according to the relative position information; Parse the picture content of the visual element, and generate a content text corresponding to the content piece according to the picture content; Obtain a second text according to the position text and the content text.
[0013] Optionally, in a possible implementation of the first aspect, the obtaining the relative position information of the visual element and generating a position text corresponding to the position piece according to the relative position information includes: Generate an orientation template according to the pixel proportion corresponding to the visual element, divide the input image based on the orientation template to generate multiple positioning grids, wherein the division spacing of the orientation template is proportional to the pixel proportion; Determine the positioning grid intersected by the visual element as the initial grid, and obtain the target grid after removing the initial grid with an intersection ratio less than a preset intersection ratio; Determine the relative position information of the visual element based on the target grid, and generate a position text corresponding to the position piece according to the relative position information.
[0014] Optionally, in a possible implementation of the first aspect, the determining the relative position information of the visual element based on the target grid and generating a position text corresponding to the position piece according to the relative position information includes: Determine the uppermost target cell as the first target cell, and determine the first proportional position according to the vertical quantity relationship of the first target cell; Determine the lowermost target cell as the second target cell, and determine the second proportional position according to the vertical quantity relationship of the second target cell; Determine the leftmost target cell as the third target cell, and determine the third proportional position according to the horizontal quantity relationship of the third target cell; Determine the rightmost target cell as the fourth target cell, and determine the fourth proportional position according to the horizontal quantity relationship of the fourth target cell; Determine the relative position information of the visual elements according to the first proportional position, the second proportional position, the third proportional position and the fourth proportional position, and generate position text corresponding to the position space according to the relative position information.
[0015] In a second aspect of the present invention, there is provided a generation system for data, including: An identification module, configured to identify visual elements of an input image, and determine a description weight according to the pixel ratio of each visual element; A determination module, configured to determine the element ratio of each sub-element in the visual element and the description attribute corresponding to the sub-element when the pixel ratio meets the quadratic splitting condition, where the description attribute includes a positive attribute and a negative attribute; A processing module, configured to process the description weight according to the description attribute and the element ratio to determine the description sub-weight of each sub-element; A generation module, configured to perform an association process on a required space based on the description weight and / or the description sub-weight, and the level of the visual element and / or the sub-element, and generate a structured directional text.
[0016] In a third aspect of an embodiment of the present invention, there is provided an electronic device, including: a memory, a processor, and a computer program, where the computer program is stored in the memory, and the processor runs the computer program to execute the method according to the first aspect and various possible aspects of the first aspect of the present invention.
[0017] The beneficial effects of the present invention are as follows: 1. The present invention can determine the corresponding description weight by identifying the pixel ratio of the visual elements of the input image, and generate the text description content corresponding to each visual element according to the description weight, so that the text description can allocate space according to the importance degree of each visual element in the image, effectively solving the problem of unprominent description focus, and thus enabling users to quickly obtain the core information of the image.
[0018] 2. When the pixel ratio of the visual elements meets the quadratic splitting condition, the present invention can further refine and split the visual elements and describe their features, determine the element ratios and description attributes of each sub - element in the visual elements, and generate more detailed text descriptions according to their corresponding element ratios and description attributes, thereby improving the detail level and accuracy of the text descriptions.
[0019] 3. The present invention can, based on the forward and reverse description attributes, select appropriate attributes to describe the elements according to user requirements, avoid random content generation, achieve directional output of the text, and meet the requirements for the text content tendency in different scenarios.
[0020] 4. The present invention can achieve structured directional text output through operations such as dividing the space of visual elements, generating position and content text, constructing an architecture tree, and text association, so that the text content is arranged in an orderly manner according to the hierarchical relationship of the image content. This not only ensures that the text structure is clear and well - defined, but also realizes directional description, improves the logic and readability of the text, and facilitates users to understand the image content. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flowchart of a data generation method provided by the present invention; Figure 2 is a schematic structural diagram of a data generation system provided by the present invention; Figure 3 is a schematic hardware structure diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0024] See Figure 1 , which is a schematic diagram of a data generation method provided by an embodiment of the present invention. Figure 1The execution entity of the method shown can be a software and / or hardware device. The execution entity of this application can include but is not limited to at least one of the following: user equipment, network equipment, etc. Among them, the user equipment can include but is not limited to computers, smartphones, personal digital assistants (PDAs for short), and the electronic devices mentioned above. The network equipment can include but is not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing. Among them, cloud computing is a type of distributed computing, which consists of a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. It includes steps S1 to S4, specifically as follows: S1. Identify the visual elements of the input image, and determine the description weight according to the pixel ratio of each of the visual elements.
[0025] Among them, the input image refers to the image that the user needs to perform intelligent text output on. The visual element refers to each element in the input image. For example, in a landscape photo, the visual elements can be elements such as trees, flowers, grasslands, sky, etc. The pixel ratio refers to the proportion of each visual element in the image. The description weight refers to the proportion of the space occupied in the text description assigned to each visual element according to the pixel ratio. The higher the pixel ratio, the greater the description weight, and the higher the description space obtained in the subsequent text generation.
[0026] This solution can determine the description weight by identifying the pixel ratio of the visual elements of the input image, so that the text description can allocate space according to the importance of each element in the image, which can effectively solve the problem of unremarkable description focus in the prior art, enable the user to quickly obtain the core information of the image, and when the pixel ratio meets the secondary splitting condition, this solution can further determine the element ratio and description attributes of the sub-elements in the visual element, perform refined splitting and feature description on the key elements, so as to improve the detail and accuracy of the text description. At the same time, this solution can select appropriate attributes to describe the elements based on the forward and reverse description attributes according to the user's needs, avoid random generation of content, realize the directional output of the text, and meet the requirements for the text content tendency in different scenarios.
[0027] Specifically, through image recognition technology, the input image can be analyzed to extract the features in the image, thereby identifying various visual elements in the image, such as people, trees, sky, etc. These visual elements are the basic units for constructing the image content. After identifying the visual elements, the number of pixels occupied by each visual element can be counted. For example, if an image has a total of 10,000 pixels and people occupy 3,000 pixels, then the pixel ratio of people in this image is 30%. In this way, the pixel ratio corresponding to each visual element can be calculated. The pixel ratio can reflect the importance and conspicuousness of the visual element in the image. Taking the pixel ratio as the basis for describing the weight, the higher the pixel ratio, the greater the description weight corresponding to the visual element. For example, in the above example, the pixel ratio of people is 30% and the pixel ratio of trees is 10%. Then, when generating the description text subsequently, the description length and detail level of people will theoretically be higher than those of trees.
[0028] In some embodiments, the specific implementation manner of step S1 may be: S11, receive the input image and identify multiple visual elements of the input image.
[0029] Specifically, the image uploaded by the user, that is, the input image, can be received. Through image recognition technology, feature extraction is performed on the input image, and then multiple specific objects in the image, that is, visual elements, can be identified.
[0030] S12, obtain the number of pixels of each of the visual elements, and generate a pixel ratio according to the ratio of the number of pixels to the total number of pixels of the input image.
[0031] After identifying the visual elements, the pixels included in each visual element can be accurately counted. This process relies on the pixel-level representation of the image. Each pixel in the image has its specific coordinate position and color value. By traversing the pixel points belonging to each visual element in the image and counting point by point, the number of pixels of each visual element can be obtained. After obtaining the number of pixels of each visual element, dividing the number of pixels of each visual element by the total number of pixels of the input image can obtain the pixel ratio of the visual element. Among them, the number of pixels refers to the total number of all pixel points of the visual element, and the total number of pixels refers to the sum of all pixel points in the input image.
[0032] S13, generate a ratio based on the pixel ratio of each of the visual elements, and generate a corresponding description weight according to the ratio.
[0033] Specifically, arrange the pixel proportions of each visual element in descending order, calculate the ratio of the pixel proportions of adjacent visual elements to form a proportion relationship. For example, if there are three visual elements A, B, and C in an image with pixel proportions of 30%, 20%, and 10% respectively, then the proportion among A, B, and C is 3:2:1. The proportion can clearly show the differences in the relative importance of each visual element. According to the proportion and in combination with the pre-set weight assignment rules, generate a description weight for each visual element. A common rule is that the higher the proportion of a visual element, the higher the description weight assigned to it. The description weights determined in this way can ensure that when generating text descriptions later, the visual elements with high importance and large proportions in the image will receive more description space and attention, thus highlighting the core content of the image.
[0034] Through the above implementation methods, the generated text can be made more in line with the importance distribution of the image itself in terms of structure and content, thus highlighting the core content of the image.
[0035] S2. When the pixel proportion meets the secondary splitting condition, determine the element proportion of each sub-element in the visual element and the description attribute corresponding to the sub-element. The description attribute includes a positive attribute and a negative attribute.
[0036] Among them, the secondary splitting condition means that the pixel proportion is greater than a preset proportion threshold. A sub-element refers to a component obtained by further disassembling a visual element that meets the secondary splitting condition. For example, taking the human visual element as an example, the sub-elements include eyes, nose, mouth, ears, etc. The element proportion refers to the proportion of each sub-element in the visual element to which it belongs. The description attribute refers to the attribute used to describe the characteristics of the sub-element. The positive attribute refers to the attribute that can reflect the advantages and positive characteristics of the sub-element, and the negative attribute refers to the attribute that reflects the disadvantages and negative characteristics of the sub-element.
[0037] Specifically, a pixel proportion threshold can be preset as the secondary splitting condition. For example, when the pixel proportion of a certain visual element exceeds 30%, it is considered that this visual element occupies an important position in the image and needs more detailed analysis to meet the secondary splitting condition. For the visual elements that meet the secondary splitting condition, further disassemble them. Taking the human visual element as an example, the sub-elements may include eyes, nose, mouth, ears, etc. Analyze the proportion of each sub-element in this visual element. For example, in the human head, the eyes may account for 30% of the head area. Set positive and negative attributes for each sub-element. The positive attribute refers to the attribute that can reflect the advantages and positive characteristics of the sub-element. For example, the positive attributes of the eyes can be "bright and energetic" and "big and lively", and the negative attribute is the attribute that reflects the disadvantages and negative characteristics, such as "dim and listless" and "small and inattentive". The setting of these attributes provides a basis for subsequent targeted descriptions.
[0038] Based on the above embodiments, the specific implementation manner of step S2 can be as follows: S21. If the pixel ratio corresponding to the visual element is greater than or equal to the ratio threshold, it is determined that the visual element meets the condition for secondary splitting.
[0039] Specifically, a pixel ratio threshold can be set in advance. This threshold can be used as a standard for determining whether a visual element needs further in-depth analysis. For example, the ratio threshold can be set to 30%. The setting of this value is usually based on the statistical analysis of a large amount of image data and the requirements of the actual application scenario, aiming to screen out visual elements that occupy an important position in the image. After calculating the pixel ratio of each visual element, the pixel ratio of each visual element can be compared with the pre-set ratio threshold one by one. If the pixel ratio of a certain visual element is greater than or equal to the threshold, for example, the pixel ratio of the "person" visual element in the input image reaches 40%, exceeding the 30% threshold, then it can be determined that the "person" visual element meets the condition for secondary splitting, and subsequent more detailed analysis and processing will be carried out on it. On the contrary, if the pixel ratio is less than the ratio threshold, the visual element will not be split secondarily. Among them, the ratio threshold refers to the pre-set pixel ratio critical value used to measure the importance of the visual element in the image.
[0040] S22. Obtain the element area corresponding to each sub-element in the corresponding visual element and the element area corresponding to the visual element.
[0041] For a visual element that meets the condition for secondary splitting, more refined recognition can be performed on it with the help of computer vision algorithms. Taking the "person" visual element as an example, through technologies such as human key point detection and semantic segmentation, the person can be further refined and can also be split into smaller sub-elements such as eyes, nose, mouth, ears, etc. After identifying the sub-elements, the image area occupied by each sub-element can be calculated. This process is based on the pixel information of the image. By counting the number of pixels included in the sub-element and converting the number of pixels into an actual area value, at the same time, the total element area of the visual element can be calculated. For example, the total area of the "person" visual element is 1000 square pixels, where the area of the "eye" sub-element is 200 square pixels, the area of the "nose" sub-element is 100 square pixels, etc.
[0042] Among them, the element area refers to the actual area size occupied by a certain sub-element in the visual element in the image, and the element area refers to the actual area size occupied by the entire visual element in the image.
[0043] S23. According to the ratio between each element area and the element area, obtain the element ratio corresponding to each sub-element, and determine the element ratio of each sub-element based on the element ratio of each sub-element.
[0044] Specifically, divide the element area of each sub - element by the element area of the visual element to which it belongs to obtain the element proportion of the sub - element. For example, when the area of the "eye" sub - element is 200 square pixels, the area of the "nose" sub - element is 100 square pixels, and the total area of the "person" visual element is 1000 square pixels, then the element proportion of the "eye" is 20%, and the element proportion of the "nose" is 10%. In this way, the element proportions of all sub - elements corresponding to the visual element can be calculated. After obtaining the element proportions of each sub - element, organize and compare these proportion values according to certain rules to form the element proportion relationship between each sub - element. For example, if the element proportions of "eye", "nose", and "mouth" are 20%, 10%, and 15% respectively, then their element proportion can be expressed as 4:2:3. This element proportion relationship can clearly show the relative importance and distribution of each sub - element in the visual element to which it belongs, providing a basis for the allocation of subsequent description content.
[0045] Among them, the element proportion refers to the proportion of the element area of the sub - element in the element area of the visual element to which it belongs, and the element proportion relationship refers to the proportion relationship between the element proportions of each sub - element in the visual element.
[0046] S24, receive the click information of the user - end on the sub - element, retrieve the corresponding property window of the sub - element and send it to the user - end, and determine the description property corresponding to the sub - element according to the click information of the user - end on any description property in the property window.
[0047] Specifically, the visual element that meets the condition of secondary splitting and its sub - elements can be displayed on the user interface of the user - end. The user can select the sub - element of interest through operations such as mouse clicking and touching. For example, the user can click on the "eye" sub - element in the "person" visual element in the image preview interface. When receiving the click information of the user on the sub - element, the corresponding property window of the sub - element can be retrieved from the pre - set property database. The property window contains the positive and negative property options of the sub - element. Send the property window to the user - end for display. The user selects the corresponding description property in the property window according to their own needs and understanding of the image. When the user clicks on a property option, the click information can be received, and the property selected by the user is determined as the description property corresponding to the sub - element. For example, when the user clicks on the option of the positive property, the description property of the "eye" sub - element is determined to be positive. Subsequently, when generating the text description, the "eye" will be described around the positive property to achieve the orientation of the text content.
[0048] Among them, the client refers to the terminal held by the user who applies to generate the text description corresponding to the input picture. For example, it can be a mobile terminal. The click information refers to the relevant information about the click behavior received when the user clicks on an element on the interface of the client. The property window refers to an interface window retrieved from a pre-set property database according to the sub-element clicked by the user and displayed to the user.
[0049] Through the above implementation manners, the directional output of the text can be realized, meeting the requirements of users for text content in different scenarios.
[0050] In some other embodiments, the description attributes corresponding to the sub-elements can also be determined in the following manner: A1. Determine the element type corresponding to the sub-element. The element type includes the proportion type and the pixel type.
[0051] Specifically, according to the analysis of the characteristics of image elements and the description requirements, the sub-elements can be divided into the proportion type and the pixel type. The proportion type focuses on the analysis from the perspective of the area proportion of the sub-element in the visual element to which it belongs, and is applicable to sub-elements that need to measure importance by relative size. The pixel type focuses on the characteristics of the pixel value of the sub-element itself and is used to analyze features directly related to the pixel value such as color and brightness. After completing the recognition and preliminary analysis of the sub-elements, the element type can be automatically determined for each sub-element according to the characteristics of the sub-element and the preset classification rules. For example, for sub-elements such as "eyes", "nose", and "mouth" in the "person" visual element, since their relative size in the visual element is important for description, they can be classified into the proportion type. For the "leaf" sub-element, if the focus is on features such as its color and texture reflected by the pixel value, it can be classified into the pixel type.
[0052] Among them, the element type refers to classifying the sub-element into different categories according to the characteristics of the sub-element and the analysis requirements. The proportion type refers to the type of sub-element that pays more attention to its area proportion in the visual element to which it belongs in the analysis. The pixel type refers to the type corresponding to the sub-element that pays more attention to the characteristics of its own pixel value in the analysis.
[0053] A2. Obtain the element proportion corresponding to the sub-element of the proportion type. If the element proportion is greater than or equal to the preset threshold, determine its corresponding description attribute as a positive attribute based on the proportion determination strategy.
[0054] For a sub - element determined to be of the proportion type, the corresponding calculated element proportion data can be obtained. For example, in the visual element of "person", the element proportion of the sub - element "eyes" is 30%. And a preset element proportion threshold, that is, a preset threshold, can be set in advance for the sub - elements of the proportion type. For example, it can be 20%. This preset threshold can be used to distinguish the pros and cons tendency of the sub - elements. Compare the element proportion of the sub - element with the preset threshold. If the element proportion is greater than or equal to the preset threshold, for example, 30% of "eyes" ≥ 20%, then its description attribute can be determined to be a positive attribute according to the proportion determination strategy. Among them, the preset threshold refers to a preset value used to compare with the element proportion of the sub - element of the proportion type to judge the description attribute corresponding to the sub - element. The proportion determination strategy refers to the strategy of determining the description attribute corresponding to the sub - element according to the element proportion.
[0055] A3. If the element proportion is less than the preset threshold, determine its corresponding description attribute to be a negative attribute based on the proportion determination strategy.
[0056] If the element proportion of the sub - element is less than the preset threshold, for example, when the element proportion of "eyes" is 10%, since it is less than the preset threshold, its corresponding description attribute can be determined to be a negative attribute according to the proportion determination strategy.
[0057] A4. Obtain the pixel mean value corresponding to the sub - element of the pixel type. Based on the pixel value determination strategy, determine that the description attribute corresponding to the sub - element whose pixel mean value is within the normal pixel range is a positive attribute, and determine that the description attribute corresponding to the sub - element whose pixel mean value is outside the normal pixel range is a negative attribute.
[0058] Specifically, for the sub - element of the pixel type, all pixel points included in the sub - element can be traversed to obtain the color value of each pixel point, and then the average value of these pixel values is calculated to obtain the pixel mean value of the sub - element. And according to the common pixel value range of this type of element in the image, a normal pixel range can be set for the sub - element of the pixel type. Compare the pixel mean value corresponding to the sub - element of the pixel type with the normal pixel range. If the pixel mean value is within the normal pixel range, it indicates that the pixel feature of the sub - element conforms to the conventional good state, and its description attribute can be determined to be a positive attribute. For example, positive descriptions such as "bright color" and "clear texture" can be given. If the pixel mean value exceeds the normal pixel range, its description attribute is determined to be a negative attribute. For example, negative descriptions such as "abnormal color" and "unbalanced brightness" can be corresponding.
[0059] Among them, the pixel mean refers to the average value of the color values of all pixel points corresponding to the sub-elements of the pixel type. The pixel value determination strategy refers to the strategy of determining its description attribute according to the comparison result between the pixel mean of the sub-element and the normal pixel interval. The normal pixel interval refers to an interval preset according to the common pixel value range of the sub-elements of the pixel type in the image. If the pixel mean of the sub-element is within this interval, it is considered that its pixel feature conforms to the normal good state.
[0060] S3. Process the description weight according to the description attribute and the element ratio to determine the description sub-weight of each sub-element.
[0061] Among them, the description sub-weight is a quantitative index used to measure the importance of each sub-element in the image in the finally generated text description. The higher the value of the description sub-weight, the more attention the sub-element receives in the text description, and more text content can be allocated to it.
[0062] It can be understood that the description sub-weight is a quantitative index used to measure the importance of each sub-element in the image in the finally generated text description. It comprehensively considers the characteristics of the sub-element itself (element ratio and description attribute), and the importance of the visual element to which the sub-element belongs in the entire image (description weight). The size of the description sub-weight value directly determines the description space and detail level that the corresponding sub-element will obtain when generating text. The higher the value, the more attention the sub-element receives in the text description, and more text content will be allocated to it.
[0063] Specifically, the element ratios of each sub - element that has been calculated in its respective visual element can be obtained. For example, in the visual element of "person", the element ratio of the "eye" sub - element is 30%, the element ratio of the "nose" sub - element is 15%, and the element ratio of the "mouth" sub - element is 20%. These ratios can reflect the relative importance and area proportion of each sub - element within its respective visual element. And the description weights of each visual element that have been determined in the entire image can be obtained. For example, for the visual element of "person", since its pixel proportion in the image is relatively high, the determined description weight is 0.6 (assuming the total weight is 1 and the weights of each visual element are distributed according to pixel proportion). Multiply the element ratio corresponding to each sub - element by the description weight of the visual element to which the sub - element belongs, and the description sub - weights corresponding to each sub - element can be obtained. For example, for the "eye" sub - element, its element ratio is 30% (i.e., 0.3), and the description weight of the "person" visual element to which it belongs is 0.6. Then the description sub - weight of the "eye" can be 0.3×0.6 = 0.18. For the "nose" sub - element with an element ratio of 15% (i.e., 0.15), its description sub - weight is 0.15×0.6 = 0.09. For the "mouth" sub - element with an element ratio of 20% (i.e., 0.2), its description sub - weight is: 0.2×0.6 = 0.12.
[0064] S4. Based on the described weights and / or description sub - weights, and the levels of the visual elements and / or sub - elements, perform an association process on the required text length to generate a structured and targeted text.
[0065] Among them, the required text length refers to the total amount of text expected to be achieved for generating a structured and targeted text that describes a specific image or scene. The structured and targeted text refers to a text with a specific structure and content tendency formed after integrating and associating the first text that describes the visual elements in the image with the second text that describes the sub - elements under the visual elements based on the architecture tree generated by image analysis.
[0066] Establish an association rule with the required length according to the description weight, descriptor weight, and the preset levels of visual elements and sub-elements. For example, if the description weight corresponding to a visual element is high, more length can be allocated; if the descriptor weight of the sub-element corresponding to a visual element is high, more text volume can be given to the corresponding sub-element in the description of that visual element. According to the association rule, allocate the required length to the descriptions of each visual element and sub-element. For example, when the required length is 1000 words, according to the weights and levels of each visual element and sub-element, 400 words may be allocated to the character visual element, and the description corresponding to the eye sub-element with a high descriptor weight may be 100 words. Based on the allocated length, combine the characteristics, description attributes, etc. of each visual element and sub-element, and generate text according to a certain text structure, that is, structured and targeted text. During the generation process, if the user has a positive description requirement, positive attributes are preferentially used; if there is a negative description requirement, negative attributes are emphasized, ensuring that the finally output text not only meets the structured requirements, that is, the descriptions of each element and sub-element are clear and hierarchical, but also achieves orientation, that is, meets the user's requirements for content tendency.
[0067] Based on the above embodiments, the specific implementation manner of step S4 may be: S41, process the preset length based on the description weight to obtain the first length corresponding to each visual element, and split the corresponding first length according to the descriptor weight to obtain the second length and description attributes corresponding to each descriptor weight.
[0068] Specifically, the preset length can be obtained, that is, the total number of words of the text that the user expects to generate, such as 1000 words. According to the determined description weights of each visual element, the preset length is proportionally allocated to each visual element to obtain the first length corresponding to each visual element. For example, assuming that there are visual elements such as people, trees, and grass in the input image, where the description weight of people is 0.6 and the description weight of trees is 0.4, then the first length corresponding to people is 600 words, and the first length corresponding to trees is 400 words. These first lengths can reflect the approximate importance and length allocation of each visual element in the overall description. For the first length obtained for each visual element, it can be further split according to the determined description sub-weights of each sub-element under that visual element. For example, taking the visual element of people as an example, when the description sub-weights of eyes, nose, and mouth are 0.18, 0.09, and 0.12 respectively, and the first length of "people" is 600 words, then the second length corresponding to "eyes" is 600×(0.18÷(0.18 + 0.09 + 0.12))≈240 words, the second length corresponding to "nose" is approximately 120 words, and the second length corresponding to "mouth" is approximately 160 words. During the splitting process, the description attributes corresponding to each sub-element are also obtained. For example, when the eyes have a positive attribute, "bright and energetic" can be combined to provide a content direction for subsequent text generation.
[0069] Among them, the preset length refers to the total number of words of the text that the user expects to generate. The first length refers to the number of words corresponding to each visual element after the preset length is proportionally allocated to each visual element according to the description weights of the visual elements. The second length refers to the number of words corresponding to each sub-element after the first length obtained for each visual element is further split according to the description sub-weights of each sub-element under that visual element.
[0070] S42, Parse the picture content corresponding to the visual element according to the first length to generate the first text, and parse the picture content corresponding to the sub-element according to the second length and the description attribute to generate the second text.
[0071] Specifically, for the first passage corresponding to each visual element, conduct an in-depth analysis of the content of the visual element in the image. Taking the visual element of "trees" as an example, its first passage is 400 words. The characteristics such as the shape, color, and surrounding environment of the trees in the image can be analyzed. Descriptions can be made in a certain logical order, such as from the whole to the part, from far to near, etc., using appropriate language to generate a first text of about 400 words regarding "trees". For the second passage and description attributes corresponding to each sub-element, the details of the sub-element in the image can be focused on. For example, taking the "eyes" sub-element under the visual element of "person" as an example, its second passage is 240 words and the description attribute is a positive attribute. By carefully observing details such as the shape, color, and expression in the eyes, and combining with the positive description attribute, a second text of about 240 words can be generated.
[0072] Among them, the picture content refers to the specific image information presented by each visual element and its sub-elements in the input image. The first text refers to the text generated after an in-depth analysis of the content of each visual element in the image according to the first passage allocated to it. The second text refers to the text generated after a focused analysis of the details of the sub-elements in the image in combination with the second passage allocated to them and the determined description attributes.
[0073] In some embodiments, the step of "analyzing the picture content corresponding to the sub-elements according to the second passage and description attributes to generate the second text" in step S42 includes the following steps: S421, analyze the picture content corresponding to each of the sub-elements, and perform directional adjustment on the picture content according to the description attributes to obtain the adjusted content.
[0074] Specifically, a detailed analysis can be conducted on the specific content of each sub-element in the image. For example, taking the "eyes" sub-element as an example, all visible image detail information such as the shape contour of the eyes, the cleanliness of the sclera, and the growth condition of the eyelashes can be identified. Through image recognition algorithms and feature extraction techniques, these details are converted into a data form that can be processed by a computer to obtain the corresponding picture content. After obtaining the picture content, directional adjustment can be performed on the picture content according to the description attributes. That is to say, if the description attribute is a positive attribute, then only the features that can reflect the positive attribute can be selected. For example, for "eyes", positive features such as bright iris color, clear expression, and long and slender eyelashes can be selected, and neutral or negative features such as possible tiny blood streaks and slight eye bags can be ignored. Then, these positive features are directionally strengthened and highlighted to make the positive features more distinct, so as to obtain the adjusted content that meets the requirements of the positive attribute. If the description attribute is a negative attribute, then vice versa, negative features are selected and emphasized for adjustment.
[0075] Among them, the directional adjustment refers to the process of making targeted and directional modifications and optimizations to the content of the sub-element images parsed based on the descriptive attributes of the sub-elements. The adjusted content refers to the result obtained by screening and processing the original image content according to the descriptive attributes, and includes the characteristic information that meets the requirements of the descriptive attributes.
[0076] S422, process the adjusted content according to the second passage to generate a second text.
[0077] Specifically, according to the word limit of the second passage, the adjusted content can be planned. For example, if the second passage of the "eye" sub-element is 240 words, then the approximate word distribution for describing each feature can be determined. For example, 80 words can be allocated to describe the eye shape and overall appearance, 100 words to describe the look in the eyes and dynamics, and 60 words to describe other details such as eyelashes. According to the plan, the adjusted content can be transformed into smooth text, so that the descriptions of each feature can be naturally linked together. At the same time, it can ensure that the number of words in the text meets the requirements of the second passage, and generate the second text corresponding to the sub-element.
[0078] Based on the above steps, this solution also includes the following embodiments: B1, obtain the element type of the visual element, and classify the visual elements according to the element type to obtain an element set.
[0079] Specifically, the element type is a classification identifier for visual elements. For example, visual elements can be divided into "landscape types" (such as trees, rivers, mountains), "human types", etc. By using image recognition technology and pre-set classification rules, the characteristics of each visual element are analyzed to determine its element type. For example, for an image of a park, visual elements such as "trees", "flowers", and "people" are recognized. According to their characteristics, "trees" and "flowers" can be classified as "landscape types", and "people" can be classified as "human types". After determining the element type corresponding to each visual element, the visual elements can be classified according to the element type, and the visual elements with the same element type are grouped together to finally form multiple element sets. Through this classification method, the visual elements in the image can be systematically organized, providing convenience for subsequent analysis and processing.
[0080] Among them, the element type refers to the category identifier set when classifying visual elements, used to distinguish visual elements of different natures. The element set refers to the set formed by classifying and organizing visual elements according to the element type after determining the element type corresponding to each visual element, and grouping the visual elements with the same element type together.
[0081] B2, obtain the pixel proportion of each visual element in the element set, and sort them in descending order according to the pixel proportion to obtain an element sequence.
[0082] Specifically, after obtaining the feature set, the pixel ratio of each visual element in each feature set can be obtained. This step is similar to the method of calculating the pixel ratio when determining the description weight before, that is, the number of pixels of each visual element is counted and divided by the total number of pixels of the image to obtain the pixel ratio. For example, in the landscape feature set, the number of pixels of "trees" is 3000, and the total number of pixels of the image is 10000, then the pixel ratio of "trees" is 30%, and the number of pixels of "flowers" is 1000, and its pixel ratio is 10%. The visual elements in each feature set are sorted from large to small according to the pixel ratio. For example, after sorting the landscape feature set, "trees (30%) > flowers (10%)" is obtained. The sorting results of the visual elements in each feature set are integrated to obtain the feature sequence corresponding to each feature set. By sorting in descending order, the importance distribution of each visual element in the image can be clearly presented. The visual elements with a high pixel ratio are placed at the front of the sequence, which is convenient for subsequent key analysis and processing.
[0083] Among them, the feature sequence refers to the ordered sequence obtained by arranging the visual elements in each feature set in descending order according to their pixel proportions from large to small.
[0084] B3, determining a judgment ratio of a pixel ratio of the last visual element in the element sequence to a pixel ratio of the first visual element, and removing the last visual element in the element sequence when the judgment ratio is less than a preset value.
[0085] Specifically, the pixel ratio of the last visual element and the pixel ratio of the first visual element in the element sequence can be taken out, and the ratio of the two can be calculated as the judgment ratio. For example, if the element sequence is "people (40%)>trees (30%)>flowers (10%)", then the judgment ratio is 10%÷40%=0.25. The calculated judgment ratio is compared with the preset threshold. If the judgment ratio is less than the preset value, it can be considered that the importance of the last visual element in the element sequence in the image is too different from that of the first visual element, and it is a relatively minor element. At this time, the visual element can be removed from the element sequence. For example, when the preset value is 0.3, the judgment ratio 0.25 calculated above is less than 0.2, then "flowers" can be removed from the element sequence.
[0086] The judgment ratio refers to the ratio of the pixel ratio of the last visual element in the element sequence to the pixel ratio of the first visual element, and the preset value refers to a pre-set threshold for comparison with the judgment ratio.
[0087] B4, repeating the above steps, when the judgment ratio is greater than or equal to a preset value, taking the remaining visual elements in the element sequence as new visual elements.
[0088] Specifically, after completing one elimination operation, the above steps can be executed again to recalculate the pixel proportion of the remaining visual elements and sort them in descending order. Then, calculate the new judgment proportion and compare it with the preset value. If the judgment proportion is still less than the preset value, continue to eliminate the last visual element in the element sequence. If the judgment proportion is greater than or equal to the preset value, it can be considered that the remaining visual elements in the element sequence are relatively close in terms of importance and there are no particularly minor elements. When the judgment proportion is greater than or equal to the preset value, the remaining visual elements in the element sequence can be determined as the new visual elements, and these new visual elements will be used as the objects for subsequent analysis and processing, such as determining the description weight, generating text descriptions, etc. Through this continuous screening and optimization process, relatively minor visual elements in the image can be removed, and more critical and representative visual elements can be retained, thereby improving the accuracy and effectiveness of the subsequent processing results.
[0089] In some embodiments, the first text can also be generated in the following manner: C1. Divide the first length into a position length and a content length, obtain the relative position information of the visual elements, and generate a position text corresponding to the position length according to the relative position information.
[0090] Specifically, before generating the first text, the first length assigned to each visual element can be divided into a position length and a content length according to a certain ratio, and this ratio can be preset according to actual needs. For example, it can be set that the position length accounts for 30% of the first length and the content length accounts for 70%. Assuming that the first length of the visual element "tree" is 400 words, then the position length is 400×30% = 120 words, and the content length is 400×70% = 280 words. Moreover, the relative position relationship between the visual elements in the image can be analyzed, and based on the obtained relative position information, combined with the assigned position length, appropriate language can be used to describe the position relationship of each visual element to generate the position text corresponding to each visual element.
[0091] Among them, the position length refers to the text length set in advance to describe the position information of a certain visual element in the image when generating the first text about the visual element. The content length refers to the text length used to describe the specific content characteristics of the visual element. The relative position information refers to the mutual position relationship between the visual elements in the image. The position text refers to the text content generated by using appropriate language to describe the position relationship of each visual element based on the obtained relative position information corresponding to the visual element and combined with the assigned position length.
[0092] In some embodiments, the steps of "obtain the relative position information of the visual elements and generate a position text corresponding to the position length according to the relative position information" in step C1 include the following steps: C11. Generate an orientation template based on the pixel ratio corresponding to the visual element, and divide the input image based on the orientation template to generate multiple positioning grids, where the division interval of the orientation template is proportional to the pixel ratio.
[0093] Specifically, the pixel ratio of the visual element in the image is an important basis. The pixel ratio reflects the importance of the visual element in the image and the size of the occupied space. The larger the pixel ratio, the more important the visual element is in the image or the larger the occupied space. Then, when generating the orientation template, the orientation template can be generated according to the determined pixel ratios of the visual elements. The generated orientation template can be used to divide the input image. After division, multiple positioning grids are generated, and the division interval of the orientation template is proportional to the pixel ratio, that is, the larger the pixel ratio, the larger the division interval. In this way, the image can be divided more reasonably to adapt to the spatial distribution of different visual elements. For example, if the pixel ratio of "trees" is 30% and the pixel ratio of "people" is 50%, since the pixel ratio of "people" is higher, the division interval of the orientation template for "people" generated will be larger than that of "trees".
[0094] Among them, the orientation template refers to a regular layout template for dividing the input image generated based on the pixel ratio corresponding to the visual element. The positioning grid refers to the grid with clear boundaries and positions obtained after dividing the input image based on the orientation template. The division interval refers to the interval distance between adjacent positioning grids when dividing the input image based on the orientation template.
[0095] C12. Determine the positioning grids where the visual elements intersect as the initial grids, and obtain the target grids after removing the initial grids with an intersection ratio less than the preset intersection ratio.
[0096] Specifically, by judging the intersection situation between the visual element and each positioning grid, the positioning grids where the visual element intersects can be determined as the initial grids. This step is to find out the range involved by the visual element in the positioning grids after image division. Since some intersections may be only very small parts and have little effect on describing the positional relationship of the visual element, a standard intersection ratio, that is, the preset intersection ratio, can be set. The initial grids with an intersection ratio less than the preset intersection ratio are removed, and the remaining ones are the target grids. In this way, the positioning grid area that is more meaningful for describing the position of the visual element can be screened out, improving the accuracy of subsequent position information determination.
[0097] Among them, the initial grid refers to the grid that has an intersection relationship with the visual element determined by judging the intersection situation between the visual element and each positioning grid generated after dividing the input image based on the orientation template. The intersection ratio refers to the ratio of the area of the intersection part between the visual element and a certain positioning grid to the total area of the positioning grid. The preset intersection ratio refers to a preset ratio threshold used to judge whether the initial grid is of great significance for describing the position of the visual element. The target grid refers to the remaining positioning grid after removing the initial grids with an intersection ratio less than the preset intersection ratio after determining the initial grids that intersect with the visual element.
[0098] C13. Determine the relative position information of the visual element based on the target grid, and generate a position text corresponding to the position space according to the relative position information.
[0099] Specifically, after obtaining the target grids corresponding to each visual element, the position information of the target grids corresponding to each visual element in the orientation template can be analyzed to determine the relative position information of the visual element. According to the determined relative position information, combined with the pre-allocated position space, appropriate language is used to describe the position relationship of each visual element, and finally a position text is generated. For example, if the position space is 120 words, then within the range of these 120 words, the position information corresponding to the relevant visual elements can be clearly and accurately described to obtain the corresponding position text.
[0100] In some embodiments, the specific implementation manner of step C13 may be: C131. Determine the topmost target grid as the first target grid, and determine the first proportional position according to the vertical quantity relationship of the first target grid.
[0101] Specifically, since a visual element may correspond to multiple target grids, the target grid in the topmost position can be found among the target grids corresponding to all visual elements and defined as the first target grid. This target grid represents the highest position of the visual element in the vertical direction. Analyze the vertical quantity relationship of the first target grid, that is, count which one the first target grid is from the top and which one it is from the bottom. Through these two quantity information, the relative position ratio of the first target grid in the distribution of all vertical target grids can be calculated. For example, if there are a total of 10 target grids, and the first target grid is the second from the top and the ninth from the bottom, then it can be clearly determined that the first target is at 2 / 10 of the upper part and is relatively close to the top in the vertical direction.
[0102] Among them, the first target grid refers to the target grid in the highest position in the vertical direction among the multiple target grid sets corresponding to the visual element, and the first proportional position refers to the relative position reflecting the first target grid in the distribution of all vertical target grids calculated according to the vertical quantity relationship of the first target grid.
[0103] C132. Determine the bottommost target cell as the second target cell, and determine the second proportional position according to the vertical quantity relationship of the second target cell.
[0104] Specifically, among the target cells corresponding to all visual elements, find the target cell at the bottommost position and define it as the second target cell. This target cell can represent the lowest position of the visual element in the vertical direction. Similar to the analysis method of the first target cell, count the position of the second target cell starting from above and the position starting from below. Based on the counted vertical quantities, calculate the relative position ratio of the second target cell in the vertical target cell distribution to obtain the second proportional position. For example, if it is the 8th from above and the 3rd from below, then it is at the 3 / 10th position from the bottom, indicating that this target cell is relatively close to the bottom in the vertical direction.
[0105] Among them, the second target cell refers to the target cell at the bottommost position in the vertical direction in the set of all target cells corresponding to the visual element, and the second proportional position refers to the relative position calculated based on the vertical quantity relationship of the second target cell and used to represent the second target cell in the entire vertical target cell distribution.
[0106] C133. Determine the leftmost target cell as the third target cell, and determine the third proportional position according to the horizontal quantity relationship of the third target cell.
[0107] Specifically, among the target cells corresponding to all visual elements, find the target cell at the leftmost position and determine it as the third target cell. This target cell can represent the leftmost position of the visual element in the horizontal direction. Count the position of the third target cell starting from the left and the position starting from the right. According to the horizontal quantity statistical results, calculate the relative position ratio of the third target cell in the horizontal target cell distribution. For example, if it is the 1st from the left and the 10th from the right, then it is at the 1 / 10th position from the left, and it can be considered that this target cell is very close to the left in the horizontal direction.
[0108] Among them, the third target cell refers to the target cell at the leftmost position in the horizontal direction in the set of target cells corresponding to the visual element, and the third proportional position refers to the relative position calculated based on the horizontal quantity relationship of the third target cell and used to represent the third target cell in the entire horizontal target cell distribution.
[0109] C134. Determine the rightmost target cell as the fourth target cell, and determine the fourth proportional position according to the horizontal quantity relationship of the fourth target cell.
[0110] Among the target cells corresponding to all visual elements, find the target cell in the rightmost position and define it as the fourth target cell. This target cell can represent the rightmost position of the visual element in the horizontal direction. Count the positions of the fourth target cell starting from the left and from the right. Based on the counted left and right quantities, calculate the relative position ratio of the fourth target cell in the horizontal distribution of target cells. For example, if it is the 9th from the left and the 2nd from the right, then it is at the 2 / 10 position on the right, indicating that the target cell is relatively close to the right side in the horizontal direction.
[0111] Among them, the fourth target cell refers to the target cell in the rightmost position in the horizontal direction among a set of target cells corresponding to the visual element, and the fourth proportional position refers to the relative position calculated based on the left and right quantity relationship of the fourth target cell and used to reflect the relative position of the fourth target cell in the entire horizontal distribution of target cells.
[0112] C135, determine the relative position information of the visual element according to the first proportional position, the second proportional position, the third proportional position, and the fourth proportional position, and generate a position text corresponding to the position space according to the relative position information.
[0113] Specifically, comprehensively analyze the first proportional position, the second proportional position, the third proportional position, and the fourth proportional position calculated previously. These proportional positions comprehensively describe the position characteristics of the visual element in the image from both vertical and horizontal dimensions. By integrating this information, the relative position of the visual element can be accurately determined. According to the determined relative position information, combined with the pre-allocated position space, use appropriate language to describe the position relationship of each visual element. For example, if the position space is 120 words, it is necessary to clearly and accurately express the position of the visual element in the image within these 120 words and generate the corresponding position text.
[0114] C2, analyze the picture content of the visual element, and generate a content text corresponding to the content space according to the picture content.
[0115] Specifically, deeply analyze the specific content of each visual element in the image. According to the analyzed picture content and combined with the allocated content space, the language can be organized to describe the characteristics of the visual element and generate a content text corresponding to the content space. Among them, the content text refers to the text describing the picture content of the visual element.
[0116] C3, obtain the first text according to the position text and the content text.
[0117] Specifically, by integrating the position text and the content text, the first text corresponding to each visual element can be obtained.
[0118] S43. Construct a parent node with the input image, child nodes with the visual elements, and grandchild nodes with the subelements, and generate an architecture tree based on the parent node, child nodes, and grandchild nodes.
[0119] Specifically, the input image can be used as the parent node, which represents all the information of the entire image and is the foundation of the entire architecture tree. The identified visual elements are used as child nodes and associated under the parent node. For example, in an image containing "person", "tree", and "grassland", the child nodes corresponding to the visual elements "person", "tree", and "grassland" are all connected to the parent node corresponding to the input image. And for visual elements that meet the condition of secondary splitting, their subelements are used as grandchild nodes and connected to the corresponding visual element child nodes. For example, the subelement nodes such as "eyes", "nose", and "mouth" under the "person" visual element are all connected under the "person" child node, forming a hierarchical node relationship. Through the construction of the above nodes, a complete tree-like architecture tree can be formed. This architecture tree clearly shows the hierarchical relationship and inclusion relationship among the input image, visual elements, and subelements, providing an intuitive framework for the subsequent structured organization of the text. It can be clearly seen from the architecture tree which subelement belongs to which visual element, and the visual element belongs to the entire image, facilitating the orderly organization and management of the text content.
[0120] Among them, the parent node refers to the node constructed with the input image, which can represent all the information of the entire image. The child node refers to the node constructed with each identified visual element and associated under the parent node. The grandchild node refers to the node corresponding to each subelement and connected to the corresponding visual element child node. The architecture tree refers to the complete tree-like structure formed by constructing a parent node with the input image, child nodes with the visual elements, and grandchild nodes with the subelements and connecting them according to the hierarchical relationship.
[0121] S44. Associate the first text and the second text according to the architecture tree to generate a structured oriented text.
[0122] Specifically, based on the generated architecture tree, the generated first text and second text can be integrated. According to the hierarchical order of the architecture tree, first arrange the first texts corresponding to each visual element in sequence, and then insert the second texts corresponding to the subelements under each visual element below the first text of that visual element. For example, first present the first text of the "tree" visual element, then present the first text of the "person" visual element, and after the first text of "person", insert the second texts of the subelements such as "eyes", "nose", and "mouth" in sequence, so that the text content is arranged in an orderly manner according to the hierarchical relationship of the image content, ensuring that the finally generated structured oriented text not only meets the structured requirements, with clear descriptions and distinct hierarchies of each visual element and subelement, but also can achieve orientation and meet the user's requirements for content tendency.
[0123] Through the above embodiments, structured and directional text output can be achieved, enabling the text content to be arranged in an orderly manner according to the hierarchical relationship of the image content. This not only ensures clear text structure and distinct levels but also realizes directional description, enhancing the logic and readability of the text and facilitating users' understanding of the image content.
[0124] See Figure 2 , which is a schematic structural diagram of a data generation system provided by an embodiment of the present invention. The data generation system includes: An identification module for identifying visual elements of an input image and determining description weights according to the pixel ratios of the visual elements. A determination module for determining the element ratios of each sub-element in the visual elements and the description attributes corresponding to the sub-elements when the pixel ratio meets the secondary splitting condition. The description attributes include positive attributes and negative attributes. A processing module for processing the description weights according to the description attributes and element ratios to determine the description sub-weights of each sub-element. A generation module for performing an association process on the required length based on the description weights and / or description sub-weights, as well as the levels of the visual elements and / or sub-elements, to generate structured and directional text.
[0125] See Figure 3 , which is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention. The electronic device 30 includes: a processor 31, a memory 32, and a computer program. Among them The memory 32 is used to store the computer program, and the memory can also be a flash memory. The computer program is, for example, an application program or a functional module that implements the above method.
[0126] The processor 31 is used to execute the computer program stored in the memory to implement each step executed by the device in the above method. For specific details, reference can be made to the relevant descriptions in the previous method embodiments.
[0127] Optionally, the memory 32 can be either independent or integrated with the processor 31.
[0128] When the memory 32 is a device independent of the processor 31, the device may further include: A bus 33 for connecting the memory 32 and the processor 31.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating data, characterized in that, Including: Identifying visual elements of the input image and determining description weights according to the pixel ratio of each visual element; When the pixel ratio meets the condition of secondary splitting, determining the element ratio of each sub-element in the visual element and the description attribute corresponding to the sub-element, where the description attribute includes a positive attribute and a negative attribute; Processing the description weights according to the description attribute and the element ratio to determine the description sub-weights of each sub-element; Based on the description weights and / or description sub-weights, and the levels of the visual elements and / or sub-elements, performing an association process on the required length to generate a structured directional text.
2. The method according to claim 1, characterized in that The identifying visual elements of the input image and determining description weights according to the pixel ratio of each visual element includes: Receiving the input image and identifying multiple visual elements of the input image; Obtaining the pixel quantity of each visual element and generating a pixel ratio according to the ratio of the pixel quantity to the total pixel quantity of the input image; Generating a ratio proportion based on the pixel ratio of each visual element and generating a corresponding description weight according to the ratio proportion.
3. The method according to claim 1, wherein When the pixel ratio meets the condition of secondary splitting, determining the element ratio of each sub-element in the visual element and the description attribute corresponding to the sub-element includes: If the pixel ratio corresponding to the visual element is greater than or equal to the ratio threshold, determining that the visual element meets the condition of secondary splitting; Obtaining the element area corresponding to each sub-element in the corresponding visual element and the element area corresponding to the visual element; According to the ratio between each element area and the element area, obtaining the element ratio corresponding to each sub-element, and determining the element ratio of each sub-element based on the element ratio of each sub-element; Receiving the click information of the sub-element from the client, retrieving the attribute window corresponding to the sub-element and sending it to the client, and determining the description attribute corresponding to the sub-element according to the click information of any description attribute in the attribute window by the client.
4. The method according to claim 1, wherein The performing an association process on the required length based on the description weights and / or description sub-weights, and the levels of the visual elements and / or sub-elements to generate a structured directional text includes: Processing a preset length according to the description weights to obtain the first length corresponding to each visual element, and splitting the corresponding first length according to the description sub-weights to obtain the second length corresponding to each description sub-weight and the description attribute; Analyzing the picture content corresponding to the visual element according to the first length to generate a first text, and analyzing the picture content corresponding to the sub-element according to the second length and the description attribute to generate a second text; Constructing a parent node with the input image, sub-nodes with the visual elements, and grandchild nodes with the sub-elements, and generating an architecture tree according to the parent node, sub-nodes, and grandchild nodes; Associating the first text and the second text according to the architecture tree to generate a structured directional text.
5. The method according to claim 4, wherein The analyzing the picture content corresponding to the sub-element according to the second length and the description attribute to generate a second text includes: Analyzing the picture content corresponding to each sub-element, and performing directional adjustment on the picture content according to the description attribute to obtain the adjusted content; The adjusted content is processed according to the second length to generate a second text.
6. The method according to claim 4, wherein After receiving the input image and identifying the multiple visual elements of the input image, the method further includes: Acquiring the element type of the visual element, classifying the visual elements according to the element type, and obtaining an element set; Obtaining the pixel ratio of each visual element in the element set, and sorting the elements in descending order according to the pixel ratio to obtain an element sequence; Determine a judgment ratio of a pixel ratio of the last visual element in the element sequence to a pixel ratio of the first visual element, and remove the last visual element in the element sequence when the judgment ratio is less than a preset value; Repeat the above steps, and when the judgment ratio is greater than or equal to a preset value, use the remaining visual elements in the element sequence as new visual elements.
7. The method according to claim 6, characterized in that, The step of parsing the image content corresponding to the visual element according to the first length to generate the first text includes: dividing the first length into a position length and a content length, acquiring relative position information of the visual element, and generating a position text corresponding to the position length according to the relative position information; Analyzing the image content of the visual element, and generating content text corresponding to the content length according to the image content; A first text is obtained according to the position text and the content text.
8. The method according to claim 7, characterized in that, The step of acquiring the relative position information of the visual element and generating a position text corresponding to the position length according to the relative position information includes: Generate an orientation template according to the pixel ratio corresponding to the visual element, divide the input image based on the orientation template to generate a plurality of positioning grids, wherein the division spacing of the orientation template is proportional to the pixel ratio; Determine the positioning grid where the visual elements intersect as the initial grid, and remove the initial grids whose intersection ratio is less than a preset intersection ratio to obtain the target grid; The relative position information of the visual element is determined based on the target grid, and a position text corresponding to the position length is generated according to the relative position information.
9. The method according to claim 8, characterized in that The determining the relative position information of the visual element based on the target grid, and generating the position text corresponding to the position length according to the relative position information, comprises: Determine the uppermost target grid as the first target grid, and determine a first proportional position according to the upper and lower quantity relationship of the first target grid; Determine the target grid at the bottom as the second target grid, and determine the second proportional position according to the upper and lower quantity relationship of the second target grid; Determine the leftmost target grid as the third target grid, and determine a third proportional position according to the left and right quantity relationship of the third target grid; Determine the rightmost target grid as the fourth target grid, and determine the fourth proportional position according to the left and right quantity relationship of the fourth target grid; The relative position information of the visual element is determined according to the first proportional position, the second proportional position, the third proportional position and the fourth proportional position, and the position text corresponding to the position length is generated according to the relative position information.
10. A generation system for data, characterized in that, include: A recognition module, used to recognize visual elements of an input image and determine a description weight according to a pixel ratio of each visual element; A determination module, configured to determine, when the pixel ratio meets the secondary splitting condition, the element ratio of each sub-element in the visual element and the description attribute corresponding to the sub-element, where the description attribute includes a positive attribute and a negative attribute; A processing module, configured to process the description weights according to the description attribute and the element ratio to determine the description sub-weights of each sub-element; A generation module, configured to perform an association process on the required length based on the description weight and / or the description sub-weight, and the level of the visual element and / or the sub-element, and generate a structured directional text.
Citation Information
Patent Citations
Manuscript generation method and related device, electronic equipment and storage medium
CN117033567A
Drawing content effective space ratio acquisition method and image feature vector extraction method
CN117095040A
Low-illumination image description method based on branch prediction
CN117112828A
Camera shooting picture quality optimization method and camera
CN118784974A
Vision-text collaborative abstract generation method and system based on multi-modal learning
CN119862861A