Method, system and device for generating data
By identifying the pixel proportion and secondary split of the visual elements of the image, determining the description weights and attributes of the visual elements and sub-elements, and generating structured directed text, solving the problem of intricate image description and randomness in the prior art, and realizing the accurate and directed output of the text.
Patent Information
- Application Number
- CN202510694787.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-05-28
AI Technical Summary
When generating text after image recognition, the prior art cannot effectively distinguish the importance of each element in the image, making it difficult for text description to highlight the key content of the image, and the generated text content is relatively random, making it difficult to meet the directional output requirements in specific scenarios.
By identifying the visual elements of the input image, determining the description weight based on the pixel proportion, and when the secondary splitting conditions are met, the element proportion and description attributes of each child element in the visual element are determined, structured directed text is generated, and directional output is performed in accordance with user needs.
It realizes the refined processing of text description, highlights the core information of the image, improves the accuracy and logic of text description, and meets the directional output needs in different scenarios.
Smart Images

Figure CN120219770B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data processing technology, and in particular to a method, system and device for generating data. Background Art
[0002] In the field of artificial intelligence, text generation technology based on image recognition has been widely used and is often used in scenarios such as image content interpretation and image retrieval assistance.
[0003] In the existing technology, when AI generates descriptive text after recognizing an image, it usually adopts relatively general algorithms and models, and lacks refined processing of the image content. Specifically, it is manifested as follows: First, it is unable to effectively distinguish the importance of each element in the image, resulting in the generated text description being difficult to highlight the key content of the image. For example, in an image containing people and scenery, the length of the description cannot be reasonably allocated based on the difference in the proportion of the two in the image. Second, the generated text content is highly random, which may make it difficult to output it in a targeted manner according to user needs, and cannot meet the requirements for targeted output of text content in specific scenarios.
[0004] Therefore, how to achieve refined processing of image content and improve the accuracy of text description has become an urgent problem that needs to be solved. Summary of the Invention
[0005] The present invention provides a method, system and device for generating data, which can realize the refined processing of image content and improve the accuracy of text description.
[0006] A first aspect of the present invention provides a method for generating data, comprising:
[0007] Identify visual elements of the input image and determine description weights based on the pixel ratio of each visual element;
[0008] When the pixel ratio satisfies the secondary splitting condition, determining the element ratio of each sub-element in the visual element and the descriptive attributes corresponding to the sub-element, wherein the descriptive attributes include forward attributes and reverse attributes;
[0009] Processing the description weight according to the description attribute and the element ratio to determine the descriptor weight of each sub-element;
[0010] Based on the description weight and / or descriptor weight, and the level of the visual elements and / or sub-elements, the required length is associated and processed to generate a structured directional text.
[0011] Optionally, in a possible implementation manner of the first aspect, identifying visual elements of the input image and determining description weights according to pixel proportions of each visual element includes:
[0012] receiving an input image and identifying a plurality of visual elements of the input image;
[0013] Obtaining the number of pixels of each visual element, and generating a pixel ratio according to a ratio of the number of pixels to the total number of pixels of the input image;
[0014] A proportion is generated based on the pixel proportion of each visual element, and a corresponding description weight is generated according to the proportion.
[0015] Optionally, in a possible implementation of the first aspect, when the pixel ratio satisfies the secondary splitting condition, determining the element ratio of each sub-element in the visual element and the descriptive attributes corresponding to the sub-element includes:
[0016] If the pixel ratio corresponding to the visual element is greater than or equal to the ratio threshold, it is determined that the visual element meets the secondary splitting condition;
[0017] Obtaining the element area corresponding to each sub-element in the corresponding visual element and the element area corresponding to the visual element;
[0018] Obtaining an element ratio corresponding to each of the sub-elements according to a ratio between the area of each element and the area of the element, and determining an element ratio of each sub-element based on the element ratio of each sub-element;
[0019] Receive the user terminal's click information on the sub-element, retrieve the property window corresponding to the sub-element and send it to the user terminal, and determine the description property corresponding to the sub-element according to the user terminal's click information on any description property in the property window.
[0020] Optionally, in a possible implementation of the first aspect, the generating structured directional text by associating the required length based on the description weight and / or descriptor weight and the level of the visual element and / or sub-element includes:
[0021] Processing the preset length based on the description weight to obtain a first length corresponding to each visual element, splitting the corresponding first length according to the descriptor weight to obtain a second length and description attribute corresponding to each descriptor weight;
[0022] Parsing the image content corresponding to the visual element according to the first length to generate a first text, and parsing the image content corresponding to the sub-element according to the second length and description attributes to generate a second text;
[0023] Constructing a parent node with the input image, constructing a child node with the visual element, and constructing a grandchild node with the child element, and generating a structure tree according to the parent node, the child node, and the grandchild node;
[0024] The first text and the second text are associated according to the architecture tree to generate a structured directional text.
[0025] Optionally, in a possible implementation of the first aspect, parsing the image content corresponding to the sub-element according to the second length and description attribute to generate the second text includes:
[0026] Parsing the image content corresponding to each of the sub-elements, and performing directionally adjusted adjustments to the image content according to the description attributes to obtain adjusted content;
[0027] The adjusted content is processed according to the second length to generate a second text.
[0028] Optionally, in a possible implementation manner of the first aspect, after receiving the input image and identifying multiple visual elements of the input image, the method further includes:
[0029] Acquiring element types of the visual elements, classifying the visual elements according to the element types, and obtaining an element set;
[0030] Obtaining a pixel ratio of each visual element in the element set, and sorting the elements in descending order according to the pixel ratio to obtain an element sequence;
[0031] Determining a judgment ratio of a pixel ratio of the last visual element in the element sequence to a pixel ratio of the first visual element in the element sequence, and eliminating the last visual element in the element sequence when the judgment ratio is less than a preset value;
[0032] Repeat the above steps, and when the judgment ratio is greater than or equal to a preset value, use the remaining visual elements in the element sequence as new visual elements.
[0033] Optionally, in a possible implementation of the first aspect, parsing the image content corresponding to the visual element according to the first length to generate the first text includes:
[0034] dividing the first length into a position length and a content length, obtaining relative position information of the visual elements, and generating a position text corresponding to the position length according to the relative position information;
[0035] Analyzing the image content of the visual element and generating content text corresponding to the content length according to the image content;
[0036] A second text is obtained according to the position text and the content text.
[0037] Optionally, in a possible implementation of the first aspect, obtaining the relative position information of the visual element and generating the position text corresponding to the position length according to the relative position information includes:
[0038] Generating an orientation template according to the pixel ratio corresponding to the visual element, dividing the input image based on the orientation template to generate a plurality of positioning grids, wherein the division spacing of the orientation template is proportional to the pixel ratio;
[0039] Determine the positioning grid where the visual elements intersect as the initial grid, and remove the initial grids whose intersection ratio is less than a preset intersection ratio to obtain the target grid;
[0040] The relative position information of the visual elements is determined based on the target grid, and a position text corresponding to the position length is generated according to the relative position information.
[0041] Optionally, in a possible implementation of the first aspect, determining relative position information of visual elements based on the target grid, and generating position text corresponding to the position length according to the relative position information includes:
[0042] Determine the uppermost target grid as a first target grid, and determine a first proportional position according to a quantitative relationship between the upper and lower target grids;
[0043] Determine the lowermost target grid as the second target grid, and determine a second proportional position according to the upper and lower quantity relationship of the second target grid;
[0044] Determine the leftmost target grid as the third target grid, and determine a third proportional position based on the left and right quantity relationship of the third target grid;
[0045] Determine the rightmost target grid as the fourth target grid, and determine a fourth proportional position based on the left and right quantity relationship of the fourth target grid;
[0046] The relative position information of the visual element is determined according to the first proportional position, the second proportional position, the third proportional position and the fourth proportional position, and the position text corresponding to the position length is generated according to the relative position information.
[0047] A second aspect of the present invention provides a system for generating data, comprising:
[0048] A recognition module is used to identify visual elements of the input image and determine a description weight based on the pixel ratio of each visual element;
[0049] a determination module, configured to determine, when the pixel ratio satisfies a secondary splitting condition, an element ratio of each sub-element in the visual element and a descriptive attribute corresponding to the sub-element, the descriptive attribute including a forward attribute and a reverse attribute;
[0050] a processing module, configured to process the description weight according to the description attribute and the element ratio, and determine the descriptor weight of each sub-element;
[0051] A generation module is used to perform association processing on the required length based on the description weight and / or descriptor weight, and the level of the visual element and / or sub-element, to generate structured directional text.
[0052] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising: a memory, a processor, and a computer program, wherein the computer program is stored in the memory, and the processor runs the computer program to execute the first aspect of the present invention and various methods that may be involved in the first aspect.
[0053] The beneficial effects of the present invention are as follows:
[0054] 1. The present invention can determine the corresponding description weight by identifying the pixel ratio of the visual elements of the input image, and generate the text description content corresponding to each visual element according to the description weight, so that the text description can be allocated according to the importance of each visual element in the image, which can effectively solve the problem of unclear description focus, thereby allowing users to quickly obtain the core information of the image.
[0055] 2. When the pixel ratio of a visual element meets the secondary splitting conditions, the present invention can further refine the splitting and feature description of the visual element, determine the element ratio and description attributes of each sub-element in the visual element, and generate a more detailed text description based on the corresponding element ratio and description attributes, thereby improving the detail and accuracy of the text description.
[0056] 3. The present invention can select appropriate attributes to describe elements based on forward and reverse description attributes according to user needs, avoid random content generation, achieve directional output of text, and meet the requirements for text content tendency in different scenarios.
[0057] 4. The present invention can realize structured directional text output through operations such as dividing the length of visual elements, generating position and content text, building an architecture tree and associating text, so that the text content is arranged in order according to the hierarchical relationship of the image content, ensuring that the text structure is clear and the levels are distinct, and realizing directional description, thereby improving the logic and readability of the text and facilitating users' understanding of the image content. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 A flow chart of a method for generating data provided by the present invention;
[0059] Figure 2 A schematic structural diagram of a data generation system provided by the present invention;
[0060] Figure 3 A schematic diagram of the hardware structure of an electronic device provided by the present invention. DETAILED DESCRIPTION
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0062] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0063] See also Figure 1 , is a schematic diagram of a method for generating data provided by an embodiment of the present invention, Figure 1 The execution subject of the method shown may be a software and / or hardware device. The execution subject of the present application may include but is not limited to at least one of the following: user equipment, network equipment, etc. Among them, user equipment may include but is not limited to computers, smart phones, personal digital assistants (PDAs) and the electronic devices mentioned above. Network equipment may include but is not limited to a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of computers or network servers, wherein cloud computing is a type of distributed computing, a super virtual computer composed of a group of loosely coupled computers. This embodiment does not limit this. It includes steps S1 to S4, as follows:
[0064] S1, identifying visual elements of an input image, and determining a description weight according to a pixel ratio of each visual element.
[0065] Among them, the input image refers to the image that the user needs to output intelligent text, the visual elements refer to the various elements in the input image. For example, in a landscape photo, the visual elements can be trees, flowers, grass, sky and other elements. The pixel ratio refers to the proportion of each visual element in the image, and the description weight refers to the length ratio of the text description assigned to each visual element based on the pixel ratio. The higher the pixel ratio, the greater the description weight, and the longer the description length obtained in the subsequent text generation.
[0066] This solution can determine the description weight by identifying the pixel ratio of the visual elements of the input image, so that the text description can be allocated in length according to the importance of each element in the image. It can effectively solve the problem of unclear description focus in the existing technology, allowing users to quickly obtain the core information of the image. When the pixel ratio meets the secondary splitting conditions, this solution can further determine the element ratio and description attributes of the sub-elements in the visual elements, and perform refined splitting and feature description of the key elements, thereby improving the detail and accuracy of the text description. At the same time, this solution can select appropriate attributes to describe the elements according to user needs based on the forward and reverse description attributes, avoid random content generation, realize targeted text output, and meet the requirements for text content tendency in different scenarios.
[0067] Specifically, image recognition technology can be used to analyze the input image and extract features from the image, thereby identifying various visual elements in the image, such as people, trees, sky, etc. These visual elements are the basic units for constructing image content. After identifying the visual elements, the number of pixels occupied by each visual element can be counted. For example, an image has a total of 10,000 pixels, of which people occupy 3,000 pixels, then the pixel proportion of people in the image is 30%. In this way, the pixel proportion corresponding to each visual element can be calculated. The pixel proportion can reflect the importance and conspicuousness of the visual element in the image. The pixel proportion is used as the basis for the description weight. The higher the pixel proportion, the greater the description weight corresponding to the visual element. For example, in the above example, the pixel proportion of the person is 30%, and the pixel proportion of the tree is 10%. Then, when generating the descriptive text later, the length and detail of the description of the person will theoretically be higher than the description of the tree.
[0068] In some embodiments, the specific implementation of step S1 may be:
[0069] S11, receiving an input image, and identifying multiple visual elements of the input image.
[0070] Specifically, an image uploaded by a user, namely an input image, may be received, and features of the input image may be extracted through image recognition technology to further identify multiple specific objects in the image, namely visual elements.
[0071] S12, obtaining the number of pixels of each visual element, and generating a pixel ratio according to the ratio of the number of pixels to the total number of pixels of the input image.
[0072] After identifying the visual elements, the pixels contained in each visual element can be accurately counted. This process relies on the pixel-level representation of the image. Each pixel in the image has its own specific coordinate position and color value. By traversing the pixels belonging to each visual element in the image and counting them point by point, the number of pixels of each visual element is obtained. After obtaining the number of pixels of each visual element, the number of pixels of each visual element is divided by the total number of pixels in the input image to obtain the pixel ratio of the visual element. The number of pixels refers to the total number of all pixels of the visual element, and the total number of pixels refers to the sum of all pixels in the input image.
[0073] S13, generating a proportion based on the pixel proportion of each visual element, and generating a corresponding description weight according to the proportion.
[0074] Specifically, the pixel proportions of each visual element are arranged in order from large to small, and the ratios of the pixel proportions of adjacent visual elements are calculated to form a proportion relationship. For example, there are three visual elements A, B, and C in the image, and their pixel proportions are 30%, 20%, and 10%, respectively. Then the proportions of A, B, and C are 3:2:1. The proportions can clearly show the relative importance differences between the visual elements. According to the proportions, combined with the pre-set weight allocation rules, a description weight is generated for each visual element. The common rule is that the higher the proportion of the visual element, the higher the description weight is assigned. The description weight determined in this way can ensure that when the text description is generated later, the visual elements with high importance and large proportion in the image will get more description space and attention, thereby highlighting the core content of the image.
[0075] Through the above implementation, the generated text can be made more consistent with the importance distribution of the image itself in terms of structure and content, thereby highlighting the core content of the image.
[0076] S2: When the pixel ratio satisfies a secondary splitting condition, determine the element ratio of each sub-element in the visual element and the descriptive attributes corresponding to the sub-element, where the descriptive attributes include forward attributes and reverse attributes.
[0077] Among them, the secondary splitting condition means that the pixel ratio is greater than the preset ratio threshold, and the sub-element refers to the component after further decomposition of the visual element that meets the secondary splitting condition. For example, the visual element of a person is taken as an example, the sub-elements include eyes, nose, mouth, ears, etc. The element ratio refers to the proportion of each sub-element in the corresponding visual element. The descriptive attribute refers to the attribute used to describe the characteristics of the sub-element. The positive attribute refers to the attribute that can reflect the advantages and positive characteristics of the sub-element. The negative attribute refers to the attribute that reflects the disadvantages and negative characteristics of the sub-element.
[0078] Specifically, a pixel ratio threshold can be pre-set as a secondary splitting condition. For example, when the pixel ratio of a certain visual element exceeds 30%, it is considered that the visual element occupies an important position in the image and needs to be analyzed more carefully to meet the secondary splitting condition. For visual elements that meet the secondary splitting condition, they are further decomposed. Taking the visual element of a person as an example, sub-elements may include eyes, nose, mouth, ears, etc. The proportion of each sub-element in the visual element is analyzed. For example, in the head of a person, the eyes may account for 30% of the head area. Positive attributes and reverse attributes are set for each sub-element. Positive attributes refer to attributes that can reflect the advantages and positive characteristics of the sub-element. For example, the positive attributes of the eyes can be "bright and lively" and "big and smart". The reverse attributes are attributes that reflect shortcomings and negative characteristics, such as "dull" and "small and lifeless". The setting of these attributes provides a basis for subsequent directional description.
[0079] Based on the above embodiment, the specific implementation of step S2 may be:
[0080] S21: If the pixel ratio corresponding to the visual element is greater than or equal to the ratio threshold, it is determined that the visual element meets the secondary splitting condition.
[0081] Specifically, a pixel ratio threshold can be pre-set, which can be used as a criterion for determining whether a visual element requires further in-depth analysis. For example, the ratio threshold can be set to 30%. This value is usually set based on statistical analysis of a large amount of image data and the needs of actual application scenarios, aiming to screen out visual elements that occupy an important position in the image. After completing the calculation of the pixel ratio of each visual element, the pixel ratio of each visual element can be compared one by one with the pre-set ratio threshold. If the pixel ratio of a certain visual element is greater than or equal to the threshold, for example, the pixel ratio of the "person" visual element in the input image reaches 40%, exceeding the threshold of 30%, then it can be determined that the "person" visual element meets the secondary splitting conditions and will be analyzed and processed more carefully later. Conversely, if the pixel ratio is less than the ratio threshold, the visual element will not be split twice. Among them, the ratio threshold refers to a pre-set pixel ratio critical value used to measure the importance of a visual element in an image.
[0082] S22, obtaining the element area corresponding to each sub-element in the corresponding visual element and the element area corresponding to the visual element.
[0083] For visual elements that meet the conditions for secondary splitting, they can be identified more finely with the help of computer vision algorithms. Taking the "person" visual element as an example, the person can be further refined through human key point detection, semantic segmentation and other technologies, and can also be split into smaller sub-elements such as eyes, nose, mouth, and ears. After identifying the sub-elements, the image area occupied by each sub-element can be calculated. This process is based on the pixel information of the image. By counting the number of pixels contained in the sub-element, the number of pixels is converted into the actual area value. At the same time, the overall element area of the visual element can be calculated. For example, the total area of the "person" visual element is 1,000 square pixels, of which the area of the "eye" sub-element is 200 square pixels, and the area of the "nose" sub-element is 100 square pixels.
[0084] Among them, the element area refers to the actual area occupied by a sub-element of a visual element in the image, and the feature area refers to the actual area occupied by the entire visual element in the image.
[0085] S23, obtaining the element proportion corresponding to each of the sub-elements according to the ratio between the area of each element and the element area, and determining the element ratio of each sub-element based on the element proportion of each sub-element.
[0086] Specifically, the element area of each sub-element is divided by the element area of the visual element to which it belongs to obtain the element proportion of the sub-element. For example, when the area of the "eye" sub-element is 200 square pixels, the area of the "nose" sub-element is 100 square pixels, and the total area of the "character" visual element is 1000 square pixels, the element proportion of the "eye" is 20%, and the element proportion of the "nose" is 10%. In this way, the element proportion of all sub-elements corresponding to the visual element can be calculated. After obtaining the element proportion of each sub-element, these proportion values are sorted and compared according to certain rules to form an element proportion relationship between the sub-elements. For example, the element proportions of "eyes", "nose" and "mouth" are 20%, 10% and 15% respectively, then the element proportion between them can be expressed as 4:2:3. This element proportion relationship can clearly show the relative importance and distribution of each sub-element in the corresponding visual element, and provide a basis for the subsequent allocation of descriptive content.
[0087] Among them, the element proportion refers to the ratio of the element area of the sub-element to the element area of the visual element to which it belongs, and the element proportion refers to the proportional relationship between the element proportions of each sub-element in the visual element.
[0088] S24, receiving a click message from the user terminal on the sub-element, retrieving a property window corresponding to the sub-element and sending it to the user terminal, and determining a description property corresponding to the sub-element according to the click message from the user terminal on any description property in the property window.
[0089] Specifically, visual elements and their sub-elements that meet the secondary splitting conditions can be displayed on the user interface of the user end, and the user can select the sub-element of interest by clicking the mouse, touching, etc. For example, the user can click the "eye" sub-element in the "person" visual element in the image preview interface. After receiving the user's click information on the sub-element, the attribute window corresponding to the sub-element can be retrieved from the pre-set attribute database. The attribute window contains the forward attribute and reverse attribute options of the sub-element, and the attribute window is sent to the user end for display. The user selects the corresponding descriptive attribute in the attribute window according to his or her own needs and understanding of the image. When the user clicks on an attribute option, the click information can be received, and the attribute selected by the user can be determined as the descriptive attribute corresponding to the sub-element. For example, when the user clicks on the forward attribute option, the descriptive attribute of the "eye" sub-element is determined to be positive. When generating the text description later, the "eye" will be described around the forward attribute to achieve the orientation of the text content.
[0090] Among them, the user terminal refers to the terminal held by the user who applies to generate the text description corresponding to the input image, for example, it can be a mobile phone terminal. The click information refers to the relevant information about the click behavior received when the user clicks on the element on the interface at the user terminal. The attribute window refers to an interface window that is retrieved from a pre-set attribute database and displayed to the user based on the sub-element clicked by the user.
[0091] Through the above implementation, the directional output of text can be achieved to meet the user's requirements for text content in different scenarios.
[0092] In some other embodiments, the description attribute corresponding to the sub-element may also be determined in the following manner:
[0093] A1. Determine the element type corresponding to the sub-element, where the element type includes a proportion type and a pixel type.
[0094] Specifically, according to the needs of analyzing and describing the characteristics of image elements, sub-elements can be divided into proportion type and pixel type. The proportion type focuses on analyzing from the perspective of the area proportion of the sub-element in the visual element to which it belongs, and is suitable for sub-elements whose importance needs to be measured by relative size. The pixel type focuses on the characteristics of the sub-element's own pixel value, and is used to analyze color, brightness and other features directly related to the pixel value. After completing the identification and preliminary analysis of the sub-elements, the element type can be automatically determined for each sub-element based on the characteristics of the sub-element and the preset classification rules. For example, for sub-elements such as "eyes", "nose" and "mouth" in the "person" visual element, since their relative size in the visual element is important for description, they can be divided into proportion type. For "leaf" sub-elements, if the focus is on its color, texture and other features reflected by pixel values, they can be divided into pixel type.
[0095] Among them, element type refers to the classification of sub-elements into different categories based on their characteristics and analysis requirements. Proportion type refers to the sub-element type that pays more attention to the area ratio of the sub-element in the corresponding visual element in the analysis. Pixel type refers to the type corresponding to the sub-element that pays more attention to the pixel value characteristics of the sub-element in the analysis.
[0096] A2, obtaining the element proportion corresponding to the sub-element of the proportion type, and if the element proportion is greater than or equal to a preset threshold, determining its corresponding description attribute as a positive attribute based on the proportion determination strategy.
[0097] For sub-elements that have been determined to be of the proportion type, the corresponding calculated element proportion data can be obtained. For example, in the visual element "character", the element proportion of the "eyes" sub-element is 30%, and an element proportion threshold can be pre-set for the sub-element of the proportion type, that is, a preset threshold, for example, it can be 20%. The preset threshold can be used to distinguish the pros and cons of the sub-elements, and the element proportion of the sub-element is compared with the preset threshold. If the element proportion is greater than or equal to the preset threshold, for example, 30% ≥ 20% of "eyes", its descriptive attribute can be determined as a positive attribute according to the proportion determination strategy. Among them, the preset threshold refers to a pre-set numerical value, which is used to compare with the element proportion of the sub-element of the proportion type to determine the descriptive attribute corresponding to the sub-element, and the proportion determination strategy refers to a strategy for determining the descriptive attribute corresponding to the sub-element according to the element proportion.
[0098] A3: If the element ratio is less than a preset threshold, the corresponding description attribute is determined as a reverse attribute based on the ratio determination strategy.
[0099] If the element proportion of the sub-element is less than the preset threshold, for example, when the element proportion of "eye" is 10%, since it is less than the preset threshold, the corresponding description attribute can be determined as a reverse attribute according to the proportion determination strategy.
[0100] A4, obtain the pixel mean corresponding to the sub-element of the pixel type, and determine based on the pixel value determination strategy that the descriptive attribute corresponding to the sub-element whose pixel mean is within the normal pixel interval is a positive attribute, and determine that the descriptive attribute corresponding to the sub-element whose pixel mean is outside the normal pixel interval is a negative attribute.
[0101] Specifically, for a pixel-type sub-element, all pixels contained in the sub-element can be traversed to obtain the color value of each pixel, and then the average of these pixel values can be calculated to obtain the pixel mean of the sub-element. In addition, a normal pixel interval can be set for the pixel-type sub-element based on the common pixel value range of this type of element in the image, and the pixel mean corresponding to the pixel-type sub-element can be compared with the normal pixel interval. If the pixel mean is within the normal pixel interval, it means that the pixel feature of the sub-element meets the conventional good state, and its descriptive attribute can be determined to be a positive attribute, such as "bright color" and "clear texture". If the pixel mean exceeds the normal pixel interval, its descriptive attribute is determined to be a negative attribute, such as "abnormal color" and "brightness imbalance".
[0102] Among them, the pixel mean refers to the average value of the color values of all pixels corresponding to the sub-elements of the pixel type. The pixel value determination strategy refers to the strategy of determining its descriptive attributes based on the comparison result of the pixel mean of the sub-element and the normal pixel interval. The normal pixel interval refers to a pre-set interval based on the common pixel value range of the sub-elements of the pixel type in the image. If the pixel mean of the sub-element is within this interval, it is considered that its pixel characteristics are in a conventional good state.
[0103] S3, processing the description weight according to the description attribute and the element ratio to determine the descriptor weight of each sub-element.
[0104] Among them, the descriptor weight refers to a quantitative indicator used to measure the importance of each sub-element in the image in the final generated text description. The higher the value of the descriptor weight, the more attention the sub-element receives in the text description and the more text content can be allocated to it.
[0105] The descriptor weight is a quantitative measure of the importance of each sub-element in the resulting text description. It takes into account both the sub-element's characteristics (element proportions and descriptive attributes) and the importance of the visual elements to which it belongs within the overall image (descriptor weight). The descriptor weight directly determines the length and level of detail in the generated text description for the corresponding sub-element. A higher descriptor weight indicates greater attention to the sub-element in the text description, resulting in more textual content allocated to it.
[0106] Specifically, the calculated element ratio of each sub-element in the visual element to which it belongs can be obtained. For example, in the visual element "person", the element ratio of the "eye" sub-element is 30%, the element ratio of the "nose" sub-element is 15%, and the element ratio of the "mouth" sub-element is 20%. These ratios can reflect the relative importance and area proportion of each sub-element in the visual element to which it belongs, and the description weight of each visual element in the entire image can be obtained. For example, for the visual element "person", because it has a high pixel ratio in the image, the determined description weight is 0.6 (assuming the total weight is 1, and each visual element is weighted according to the pixel ratio). Weight), multiply the element ratio corresponding to each sub-element by the description weight of the visual element to which the sub-element belongs, and the descriptor weight corresponding to each sub-element can be obtained. For example, for the "eye" sub-element, its element ratio is 30% (that is, 0.3), and the description weight of the "person" visual element to which it belongs is 0.6, then the descriptor weight of "eye" can be 0.3×0.6=0.18, for the "nose" sub-element, its element ratio is 15% (that is, 0.15), then its descriptor weight is 0.15×0.6=0.09, for the "mouth" sub-element, its element ratio is 20% (that is, 0.2), and its descriptor weight is: 0.2×0.6=0.12.
[0107] S4, performing association processing on the required length based on the description weight and / or descriptor weight, and the level of the visual element and / or sub-element, to generate a structured directional text.
[0108] Among them, the required length refers to the expected total amount of text set to generate structured targeted text describing a specific image or scene. Structured targeted text refers to a text with a specific structure and content tendency formed by integrating and associating the first text describing the visual elements in the image with the second text describing the sub-elements under the visual elements based on the architecture tree generated by image analysis.
[0109] According to the description weight, descriptor weight and the pre-set level of visual elements and sub-elements, establish the association rules with the required length. For example, if the description weight corresponding to the visual element is high, then more length can be allocated. If the descriptor weight of the sub-element corresponding to the visual element is high, then the corresponding sub-element can be given more text in the description of the visual element. According to the association rules, the required length is allocated to the description of each visual element and sub-element. For example, when the required length is 1,000 words, it may be allocated to the character according to the weight and level of each visual element and sub-element. The visual elements are 400 words, of which the description of the eye sub-element with a high descriptor weight may be 100 words. According to the allocated length, combined with the characteristics, descriptive attributes and other information of each visual element and sub-element, the text is generated according to a certain text structure, that is, structured directional text. During the generation process, if the user has a positive description requirement, the positive attribute will be used first. If there is a negative description requirement, the negative attribute will be emphasized to ensure that the final output text meets the structural requirements, that is, the description of each element and sub-element is clear and well-organized, and is directional, that is, it meets the user's demand for content tendency.
[0110] Based on the above embodiment, the specific implementation of step S4 may be:
[0111] S41, processing the preset length based on the description weight to obtain a first length corresponding to each visual element, splitting the corresponding first length according to the descriptor weight to obtain a second length and description attribute corresponding to each descriptor weight.
[0112] Specifically, the preset length set by the user can be obtained, that is, the total number of words in the text that the user expects to generate, for example, 1000 words. According to the determined description weight of each visual element, the preset length is proportionally distributed to each visual element to obtain the first length corresponding to each visual element. For example, assuming that there are visual elements such as people, trees, and grass in the input image, where the description weight of the person is 0.6 and the description weight of the tree is 0.4, then the first length corresponding to the person is 600 words, and the first length corresponding to the tree is 400 words. These first lengths can reflect the approximate importance and length distribution of each visual element in the overall description. For the first length obtained for each visual element, the determined length can be used. The descriptor weights of each sub-element under the visual element further split the first length. For example, taking the character visual element as an example, when the descriptor weights of the eyes, nose, and mouth are 0.18, 0.09, and 0.12 respectively, and the first length of "character" is 600 words, then the second length corresponding to "eyes" is 600×(0.18÷(0.18+0.09+0.12))≈240 words, the second length corresponding to "nose" is about 120 words, and the second length corresponding to "mouth" is about 160 words. During the splitting process, the descriptive attributes corresponding to each sub-element are obtained at the same time. For example, when the eyes are a positive attribute, they can be combined with "bright and energetic" to provide content direction for subsequent text generation.
[0113] Among them, the preset length refers to the total number of words of text expected to be generated set by the user. The first length refers to the number of text words corresponding to each visual element after the preset length is proportionally distributed to each visual element based on the description weight of each visual element. The second length refers to the first length obtained for each visual element. After the first length is further split according to the descriptor weights of each sub-element under the visual element, the number of text words corresponding to each sub-element is obtained.
[0114] S42: Parsing the image content corresponding to the visual element according to the first length to generate a first text, and parsing the image content corresponding to the sub-element according to the second length and description attributes to generate a second text.
[0115] Specifically, for the first length corresponding to each visual element, an in-depth analysis is conducted on the content of the visual element in the image. Taking the "tree" visual element as an example, its first length is 400 words. The shape, color, environment and other characteristics of the trees in the image can be analyzed, and then described in a certain logical order, such as from the whole to the part, from far to near, etc., using appropriate language to generate a first text of about 400 words about "trees". For the second length and description attributes corresponding to each sub-element, the details of the sub-element in the image can be focused on. For example, taking the "eyes" sub-element under the "person" visual element as an example, its second length is 240 words, and the description attribute is a positive attribute. By carefully observing the details such as the shape, color, and eyes of the eyes, combined with the positive description attribute, a second text of about 240 words can be generated.
[0116] Among them, image content refers to the specific image information presented by each visual element and its sub-elements in the input image. The first text refers to the text generated after an in-depth analysis of the content of each visual element in the image based on the first length allocated to it. The second text refers to the text generated after a focused analysis of the details of the sub-element in the image under each visual element, combined with the second length allocated to it and the determined descriptive attributes.
[0117] In some embodiments, the step S42 of "parsing the image content corresponding to the sub-element according to the second length and description attribute to generate the second text" includes the following steps:
[0118] S421, parsing the image content corresponding to each of the sub-elements, and performing directionally adjusted image content according to the description attributes to obtain adjusted content.
[0119] Specifically, the specific content of each sub-element in the image can be analyzed in detail. For example, taking the "eye" sub-element as an example, all visible image detail information such as the shape and outline of the eyes, the cleanliness of the whites of the eyes, and the growth of the eyelashes can be identified. Through image recognition algorithms and feature extraction techniques, these details are converted into computer-processable data to obtain the corresponding image content. After obtaining the image content, the image content can be adjusted in a targeted manner according to the descriptive attributes. That is to say, if the descriptive attribute is a positive attribute, then only features that can reflect the positive attribute can be selected. For example, for "eyes", positive features such as bright iris color, clear eyes, and long eyelashes can be selected, and neutral or negative features such as possible tiny bloodshot eyes and slight eye bags can be ignored. Then, these positive features can be strengthened and highlighted in a targeted manner to make the positive features more distinct, thereby obtaining adjustment content that meets the requirements of the positive attributes. If the descriptive attribute is a negative attribute, then conversely, negative features can be selected and emphasized for adjustment.
[0120] Among them, targeted adjustment refers to the process of targeted and directional modification and optimization of the parsed sub-element image content based on the descriptive attributes of the sub-element. The adjusted content refers to the result obtained after screening and processing the original image content according to the descriptive attributes, which contains feature information that meets the requirements of the descriptive attributes.
[0121] S422: Process the adjusted content according to the second length to generate a second text.
[0122] Specifically, according to the word limit of the second length, the adjusted content can be planned. For example, if the second length of the "eyes" sub-element is 240 words, then the approximate word count for describing each feature can be determined. For example, 80 words can be allocated to describe the shape of the eyes and the overall appearance, 100 words to describe the eyes and dynamics, and 60 words to describe other details such as eyelashes. According to the plan, the adjusted content can be converted into fluent text, so that the descriptions of each feature can be naturally connected together. At the same time, the word count of the text can be ensured to meet the second length requirement, and the second text corresponding to the sub-element can be generated.
[0123] On the basis of the above steps, this solution also includes the following embodiments:
[0124] B1, obtaining the element type of the visual element, classifying the visual elements according to the element type, and obtaining an element set.
[0125] Specifically, the feature type is a classification identification of visual features. For example, visual features can be divided into "landscape" (such as trees, rivers, mountains), "people", etc. Through image recognition technology and pre-set classification rules, the characteristics of each visual feature are analyzed to determine its feature type. For example, for an image of a park, the visual features such as "trees", "flowers", and "people" are identified. According to their features, "trees" and "flowers" can be classified as "landscape", and "people" can be classified as "people". After determining the feature type corresponding to each visual feature, the visual features can be classified according to the feature type, and visual features of the same feature type can be grouped together to form multiple feature sets. This classification method can systematically organize the visual elements in the image, providing convenience for subsequent analysis and processing.
[0126] Among them, the feature type refers to the category identifier set when classifying visual elements, which is used to distinguish visual elements of different natures. The feature set refers to the set formed by classifying and organizing visual elements according to the feature type after determining the feature type corresponding to each visual element, and grouping visual elements with the same feature type into one group.
[0127] B2, obtaining the pixel ratio of each visual element in the element set, and sorting the elements in descending order according to the pixel ratio to obtain an element sequence.
[0128] Specifically, after obtaining the feature set, the pixel ratio of each visual element in each feature set can be obtained. This step is similar to the method of calculating the pixel ratio when determining the description weight previously, that is, counting the number of pixels of each visual element and dividing it by the total number of pixels of the image to obtain the pixel ratio. For example, in the landscape feature set, the number of pixels of "trees" is 3000, and the total number of pixels of the image is 10,000, then the pixel ratio of "trees" is 30%, and the number of pixels of "flowers" is 1000, and its pixel ratio is 10%. The visual elements in each feature set are sorted from large to small according to the pixel ratio. For example, after sorting the landscape feature set, "trees (30%) > flowers (10%)" is obtained. By integrating the sorting results of the visual elements in each feature set, the feature sequence corresponding to each feature set can be obtained. By sorting in descending order, the importance distribution of each visual element in the image can be clearly presented. The visual elements with a high pixel ratio are placed at the front of the sequence, which is convenient for subsequent key analysis and processing.
[0129] Among them, the feature sequence refers to the ordered sequence obtained by arranging the visual elements in each feature set in descending order according to their pixel proportions from large to small.
[0130] B3, determining a judgment ratio of a pixel ratio of the last visual element in the element sequence to a pixel ratio of the first visual element, and eliminating the last visual element in the element sequence when the judgment ratio is less than a preset value.
[0131] Specifically, the pixel ratio of the last visual element and the pixel ratio of the first visual element in the element sequence can be taken out, and the ratio of the two can be calculated as the judgment ratio. For example, if the element sequence is "people (40%) > trees (30%) > flowers (10%)", the judgment ratio is 10% ÷ 40% = 0.25. The calculated judgment ratio is compared with the preset threshold. If the judgment ratio is less than the preset value, it can be considered that the importance of the last visual element in the element sequence in the image is too large compared with the first visual element, and it is a relatively minor element. At this time, the visual element can be removed from the element sequence. For example, when the preset value is 0.3, the judgment ratio 0.25 calculated above is less than 0.2, then "flowers" can be removed from the element sequence.
[0132] The judgment ratio refers to the ratio of the pixel ratio of the last visual element to the pixel ratio of the first visual element in the element sequence, and the preset value refers to a pre-set threshold for comparison with the judgment ratio.
[0133] B4, repeating the above steps, and when the judgment ratio is greater than or equal to a preset value, taking the remaining visual elements in the element sequence as new visual elements.
[0134] Specifically, after completing a elimination operation, the above steps can be performed again to recalculate the pixel ratio of the remaining visual elements and sort them in descending order, and then calculate the new judgment ratio and compare it with the preset value. If the judgment ratio is still less than the preset value, continue to eliminate the last visual element in the element sequence. If the judgment ratio is greater than or equal to the preset value, it can be considered that the remaining visual elements in the element sequence at this time are relatively close in importance, and there are no particularly minor elements. When the judgment ratio is greater than or equal to the preset value, the remaining visual elements in the element sequence can be determined as new visual elements. These new visual elements will be used as objects for subsequent analysis and processing, such as for determining description weights, generating text descriptions, etc. Through this continuous screening and optimization process, relatively minor visual elements in the image can be removed, and more critical and representative visual elements can be retained, thereby improving the accuracy and effectiveness of subsequent processing results.
[0135] In some embodiments, the first text may also be generated in the following manner:
[0136] C1. Divide the first length into a position length and a content length, obtain relative position information of the visual elements, and generate a position text corresponding to the position length according to the relative position information.
[0137] Specifically, before generating the first text, the first length allocated to each visual element can be divided into position length and content length according to a certain ratio. This ratio can be pre-set according to actual needs. For example, it can be set to 30% of the first length for the position length and 70% for the content length. Assuming that the first length of the "tree" visual element is 400 words, then the position length is 400×30%=120 words, and the content length is 400×70%=280 words. In addition, the relative position relationship between the visual elements in the image can be analyzed. Based on the relative position information obtained and combined with the allocated position length, appropriate language can be used to describe the position relationship of each visual element to generate the position text corresponding to each visual element.
[0138] Among them, position length refers to the text length that is pre-allocated to describe the position information of a visual element in the image when generating the first text about a visual element; content length refers to the text length used to describe the specific content characteristics of the visual element; relative position information refers to the positional relationship between the visual elements in the image; position text refers to the text content generated by describing the positional relationship of each visual element based on the relative position information corresponding to the acquired visual elements, combined with the allocated position length, using appropriate language.
[0139] In some embodiments, the step C1 of "obtaining relative position information of the visual element, and generating position text corresponding to the position length according to the relative position information" includes the following steps:
[0140] C11, generating an orientation template according to the pixel ratio corresponding to the visual element, dividing the input image based on the orientation template to generate a plurality of positioning grids, wherein the division spacing of the orientation template is proportional to the pixel ratio.
[0141] Specifically, the pixel ratio of the visual element in the image is an important basis. The pixel ratio reflects the importance of the visual element in the image and the size of the space it occupies. The larger the pixel ratio, the more important the visual element is in the image or the larger the space it occupies. Therefore, when generating an orientation template, the orientation template can be generated according to the determined pixel ratio of each visual element. The generated orientation template can be used to divide the input image to generate multiple positioning grids after division. Moreover, the division spacing of the orientation template is proportional to the pixel ratio, that is, the larger the pixel ratio, the larger the division spacing. In this way, the image can be divided more reasonably to adapt to the spatial distribution of different visual elements. For example, if the pixel ratio of "trees" is 30% and the pixel ratio of "people" is 50%, since the pixel ratio of "people" is higher, the division spacing of the generated orientation template for "people" will be larger than that for "trees".
[0142] Among them, the orientation template refers to a regular layout template used to divide the input image, which is generated by the pixel ratio corresponding to the visual elements. The positioning grid refers to the grid with clear boundaries and positions obtained after dividing the input image based on the orientation template. The division spacing refers to the interval distance between adjacent positioning grids when dividing the input image based on the orientation template.
[0143] C12, determining the positioning grid where the visual elements intersect as the initial grid, and eliminating the initial grids whose intersection ratio is less than a preset intersection ratio to obtain the target grid.
[0144] Specifically, by judging the intersection between the visual element and each positioning grid, the positioning grid that intersects with the visual element can be determined as the initial grid. This step is to find out the range of the visual element in the positioning grid after the image is divided. Since some intersections may be only very small parts, they are not very useful for describing the position relationship of the visual elements. Therefore, a standard intersection ratio can be set, that is, a preset intersection ratio, and the initial grids with an intersection ratio less than the preset intersection ratio are eliminated. What remains are the target grids. In this way, the positioning grid areas that are more meaningful for describing the position of the visual element can be screened out, thereby improving the accuracy of subsequent position information determination.
[0145] Among them, the initial grid refers to the locating grid that has an intersection relationship with the visual element determined by judging the intersection between the visual element and the various locating grids generated after dividing the input image based on the orientation template. The intersection ratio refers to the ratio of the area of the intersection between the visual element and a certain locating grid to the total area of the locating grid. The preset intersection ratio refers to a pre-set ratio threshold used to determine whether the initial grid is of great significance for describing the position of the visual element. The target grid refers to the locating grids remaining after the initial grids that intersect with the visual element are determined and the initial grids with an intersection ratio less than the preset intersection ratio are eliminated.
[0146] C13, determining relative position information of the visual elements based on the target grid, and generating position text corresponding to the position length according to the relative position information.
[0147] Specifically, after obtaining the target grid corresponding to each visual element, the position information of the target grid corresponding to each visual element in the orientation template can be analyzed to determine the relative position information of the visual element. According to the determined relative position information, combined with the pre-allocated position length, appropriate language is used to describe the positional relationship of each visual element, and finally a position text is generated. For example, if the position length is 120 words, then the position information corresponding to the relevant visual element can be clearly and accurately described within the range of 120 words to obtain its corresponding position text.
[0148] In some embodiments, the specific implementation of step C13 may be:
[0149] C131 , determining the uppermost target grid as the first target grid, and determining a first proportional position based on the upper and lower quantity relationship of the first target grid.
[0150] Specifically, since a visual element may correspond to multiple target grids, the target grid at the top position can be found among the target grids corresponding to all visual elements and defined as the first target grid. This target grid represents the highest position of the visual element in the vertical direction. The upper and lower quantitative relationship of the first target grid is analyzed, that is, the first target grid is counted from the top and the bottom. Through these two quantitative information, the relative position ratio of the first target grid in the entire vertical target grid distribution can be calculated. For example, if there are 10 target grids in total, the first target grid is the second from the top and the ninth from the bottom, then it can be clearly determined that the first target is located at 2 / 10 above, and it is relatively close to the top in the vertical direction.
[0151] Among them, the first target grid refers to the target grid at the top in the vertical direction in the set of multiple target grids corresponding to the visual element, and the first proportional position refers to the relative position of the first target grid in the entire vertical target grid distribution, which is calculated based on the upper and lower quantitative relationship of the first target grid.
[0152] C132: Determine the lowermost target grid as the second target grid, and determine a second proportional position based on the upper and lower quantity relationship of the second target grid.
[0153] Specifically, among the target grids corresponding to all visual elements, find the target grid at the bottom position and define it as the second target grid. This target grid can represent the lowest position of the visual element in the vertical direction. Similar to the analysis method of the first target grid, the position of the second target grid counted from the top and the position counted from the bottom are counted. Based on the upper and lower numbers obtained by statistics, the relative position ratio of the second target grid in the vertical target grid distribution is calculated to obtain the second proportional position. For example, if it is the 8th from the top and the 3rd from the bottom, then it is at 3 / 10 at the bottom, indicating that the target grid is relatively close to the bottom in the vertical direction.
[0154] Among them, the second target grid refers to the target grid at the bottom of the vertical direction in the set of all target grids corresponding to the visual element, and the second proportional position refers to the relative position of the second target grid in the entire vertical target grid distribution calculated based on the upper and lower quantitative relationship of the second target grid.
[0155] C133: Determine the leftmost target grid as the third target grid, and determine a third proportional position based on the left and right quantity relationship of the third target grid.
[0156] Specifically, among the target grids corresponding to all visual elements, find the target grid at the leftmost position and determine it as the third target grid. This target grid can represent the leftmost position of the visual element in the horizontal direction. The position of the third target grid counted from the left and the position counted from the right are counted. According to the left and right quantity statistics, the relative position ratio of the third target grid in the horizontal target grid distribution is calculated. For example, if it is the 1st from the left and the 10th from the right, then it is at 1 / 10 on the left, and it can be considered that the target grid is very close to the left in the horizontal direction.
[0157] Among them, the third target grid refers to the target grid at the leftmost position in the horizontal direction in a set of target grids corresponding to the visual element, and the third proportional position refers to the relative position of the third target grid in the entire horizontal target grid distribution, which is calculated based on the left and right quantitative relationship of the third target grid.
[0158] C134, determining the rightmost target grid as the fourth target grid, and determining a fourth proportional position based on the left and right quantity relationship of the fourth target grid.
[0159] Among all the target grids corresponding to the visual elements, find the target grid at the rightmost position and define it as the fourth target grid. This target grid can represent the rightmost position of the visual element in the horizontal direction. Count the positions of the fourth target grid from the left and from the right. Based on the left and right numbers obtained by statistics, calculate the relative position ratio of the fourth target grid in the horizontal target grid distribution. For example, if it is the 9th from the left and the 2nd from the right, then it is at 2 / 10 on the right, indicating that the target grid is relatively close to the right in the horizontal direction.
[0160] Among them, the fourth target grid refers to the target grid at the rightmost position in the horizontal direction in a set of target grids corresponding to the visual element, and the fourth proportional position refers to the relative position of the fourth target grid in the entire horizontal target grid distribution calculated based on the left and right quantitative relationship of the fourth target grid.
[0161] C135 , determining relative position information of the visual element according to the first proportional position, the second proportional position, the third proportional position, and the fourth proportional position, and generating position text corresponding to the position length according to the relative position information.
[0162] Specifically, the first proportional position, second proportional position, third proportional position and fourth proportional position calculated previously are comprehensively analyzed. These proportional positions comprehensively describe the position characteristics of the visual elements in the image from both vertical and horizontal dimensions. By integrating this information, the relative positions of the visual elements can be accurately determined. According to the determined relative position information and combined with the pre-allocated position length, appropriate language is used to describe the positional relationship of each visual element. For example, if the position length is 120 words, it is necessary to clearly and accurately express the position of the visual element in the image within these 120 words to generate the corresponding position text.
[0163] C2, analyzing the image content of the visual element, and generating content text corresponding to the content length according to the image content.
[0164] Specifically, the specific content of each visual element in the image is deeply analyzed. Based on the analyzed image content and the allocated content length, language can be organized to describe the characteristics of the visual element, generating content text corresponding to the content length. The content text refers to the text that describes the image content of the visual element.
[0165] C3. Obtain a first text according to the position text and the content text.
[0166] Specifically, by integrating the position text and the content text, the first text corresponding to each visual element can be obtained.
[0167] S43, constructing a parent node with the input image, constructing a child node with the visual element, and constructing a grandchild node with the child element, and generating a structure tree according to the parent node, the child node and the grandchild node.
[0168] Specifically, the input image can be used as the parent node, which represents all the information of the entire image and is the root of the entire architecture tree. The identified visual elements are used as child nodes and associated under the parent node. For example, in an image containing "person", "tree", and "grass", the child nodes corresponding to the visual elements of "person", "tree", and "grass" are all connected to the parent node corresponding to the input image, and for the visual elements that meet the secondary splitting conditions, the child elements under them are used as grandchild nodes and connected to the corresponding visual element child nodes. For example, the child element nodes such as "eyes", "nose", and "mouth" under the "person" visual element are all connected under the "person" child node, forming a hierarchical node relationship. Through the construction of the above nodes, a complete tree-like architecture tree can be formed. This architecture tree clearly shows the hierarchical relationship and inclusion relationship between the input image, visual elements, and child elements, providing an intuitive framework for the subsequent structured organization of text. From the architecture tree, it can be clearly seen which visual element each child element belongs to, and the visual elements belong to the entire image, which is convenient for orderly organization and management of text content.
[0169] Among them, the parent node refers to the node constructed with the input image, which can represent all the information of the entire image. The child node refers to the node constructed by the identified visual elements and associated with the parent node. The grandchild node refers to the node corresponding to each child element, which is connected to the corresponding visual element child node. The architecture tree refers to the complete tree structure formed by constructing the parent node with the input image, the child node with the visual element, and the grandchild node with the child element, and connecting them according to the hierarchical relationship.
[0170] S44: Associating the first text and the second text according to the architecture tree to generate a structured directional text.
[0171] Specifically, the generated first text and second text can be integrated according to the generated architecture tree. The first text corresponding to each visual element can be arranged in sequence according to the hierarchical order of the architecture tree. Then, under the first text of each visual element, the second text corresponding to each sub-element under the visual element can be inserted. For example, the first text of the "tree" visual element is presented first, followed by the first text of the "person" visual element. After the first text of "person", the second text of sub-elements such as "eyes", "nose" and "mouth" are inserted in sequence, so that the text content is arranged in order according to the hierarchical relationship of the image content, ensuring that the structured directional text finally generated not only meets the structural requirements, but also has clear and hierarchical descriptions of each visual element and sub-element, and can be directional, in line with the user's demand for content preference.
[0172] Through the above implementation, structured directional text output can be achieved, so that the text content is arranged in order according to the hierarchical relationship of the image content, which not only ensures that the text structure is clear and the levels are distinct, but also realizes directional description, improves the logic and readability of the text, and facilitates users to understand the image content.
[0173] See also Figure 2 , is a schematic structural diagram of a data generation system provided by an embodiment of the present invention, the data generation system comprising:
[0174] A recognition module is used to identify visual elements of the input image and determine a description weight based on the pixel ratio of each visual element;
[0175] a determination module, configured to determine, when the pixel ratio satisfies a secondary splitting condition, an element ratio of each sub-element in the visual element and a descriptive attribute corresponding to the sub-element, the descriptive attribute including a forward attribute and a reverse attribute;
[0176] a processing module, configured to process the description weight according to the description attribute and the element ratio, and determine the descriptor weight of each sub-element;
[0177] A generation module is used to perform association processing on the required length based on the description weight and / or descriptor weight, and the level of the visual element and / or sub-element, to generate structured directional text.
[0178] See also Figure 3 , is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention, wherein the electronic device 30 includes: a processor 31, a memory 32 and a computer program;
[0179] The memory 32 is used to store the computer program, which may also be a flash memory. The computer program is, for example, an application program or a functional module for implementing the above method.
[0180] The processor 31 is configured to execute the computer program stored in the memory to implement the various steps performed by the device in the above method. For details, please refer to the relevant description in the above method embodiment.
[0181] Optionally, the memory 32 may be independent or integrated with the processor 31 .
[0182] When the memory 32 is a device independent of the processor 31, the device may further include:
[0183] The bus 33 is used to connect the memory 32 and the processor 31 .
[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating data, characterized in that: include: Identify visual elements of the input image and determine description weights based on the pixel ratio of each visual element; When the pixel ratio satisfies the secondary splitting condition, determining the element ratio of each sub-element in the visual element and the descriptive attributes corresponding to the sub-element, wherein the descriptive attributes include forward attributes and reverse attributes; Processing the description weight according to the description attribute and the element ratio to determine the descriptor weight of each sub-element includes: Processing the preset length based on the description weight to obtain a first length corresponding to each visual element, splitting the corresponding first length according to the descriptor weight to obtain a second length and description attribute corresponding to each descriptor weight; Parsing the image content corresponding to the visual element according to the first length to generate a first text, and parsing the image content corresponding to the sub-element according to the second length and the description attribute to generate a second text, including: dividing the first length into a position length and a content length, obtaining relative position information of the visual elements, and generating a position text corresponding to the position length according to the relative position information; Analyzing the image content of the visual element and generating content text corresponding to the content length according to the image content; Obtaining a first text according to the position text and the content text; Based on the description weight and / or descriptor weight, and the level of the visual elements and / or sub-elements, the required length is associated and processed to generate a structured directional text.
2. The method according to claim 1, characterized in that The identifying visual elements of the input image and determining description weights according to pixel proportions of the visual elements include: receiving an input image and identifying a plurality of visual elements of the input image; Obtaining the number of pixels of each visual element, and generating a pixel ratio according to a ratio of the number of pixels to the total number of pixels of the input image; A proportion is generated based on the pixel proportion of each visual element, and a corresponding description weight is generated according to the proportion.
3. The method according to claim 1, characterized in that When the pixel ratio satisfies the secondary splitting condition, determining the element ratio of each sub-element in the visual element and the descriptive attributes corresponding to the sub-element includes: If the pixel ratio corresponding to the visual element is greater than or equal to the ratio threshold, it is determined that the visual element meets the secondary splitting condition; Obtaining the element area corresponding to each sub-element in the corresponding visual element and the element area corresponding to the visual element; Obtaining an element ratio corresponding to each of the sub-elements according to a ratio between the area of each element and the area of the element, and determining an element ratio of each sub-element based on the element ratio of each sub-element; Receive the user terminal's click information on the sub-element, retrieve the property window corresponding to the sub-element and send it to the user terminal, and determine the description property corresponding to the sub-element according to the user terminal's click information on any description property in the property window.
4. The method according to claim 1, wherein The step of performing correlation processing on the required length based on the description weight and / or descriptor weight and the level of the visual element and / or sub-element to generate structured directional text includes: Constructing a parent node with the input image, constructing a child node with the visual element, and constructing a grandchild node with the child element, and generating a structure tree according to the parent node, the child node, and the grandchild node; The first text and the second text are associated according to the architecture tree to generate a structured directional text.
5. The method according to claim 4, characterized in that The step of parsing the image content corresponding to the sub-element according to the second length and the description attribute to generate the second text includes: Parsing the image content corresponding to each of the sub-elements, and performing directionally adjusted adjustments to the image content according to the description attributes to obtain adjusted content; The adjusted content is processed according to the second length to generate a second text.
6. The method according to claim 4, characterized in that After receiving an input image and identifying a plurality of visual elements of the input image, the method further includes: Acquiring element types of the visual elements, classifying the visual elements according to the element types, and obtaining an element set; Obtaining a pixel ratio of each visual element in the element set, and sorting the elements in descending order according to the pixel ratio to obtain an element sequence; Determining a judgment ratio of a pixel ratio of the last visual element in the element sequence to a pixel ratio of the first visual element in the element sequence, and eliminating the last visual element in the element sequence when the judgment ratio is less than a preset value; Repeat the above steps, and when the judgment ratio is greater than or equal to a preset value, use the remaining visual elements in the element sequence as new visual elements.
7. The method according to claim 1, characterized in that The acquiring the relative position information of the visual element and generating the position text corresponding to the position length according to the relative position information includes: Generating an orientation template according to the pixel ratio corresponding to the visual element, dividing the input image based on the orientation template to generate a plurality of positioning grids, wherein the division spacing of the orientation template is proportional to the pixel ratio; Determine the positioning grid where the visual elements intersect as the initial grid, and remove the initial grids whose intersection ratio is less than a preset intersection ratio to obtain the target grid; The relative position information of the visual elements is determined based on the target grid, and a position text corresponding to the position length is generated according to the relative position information.
8. The method according to claim 7, characterized in that The determining of relative position information of visual elements based on the target grid, and generating position text corresponding to the position length according to the relative position information, includes: Determine the uppermost target grid as a first target grid, and determine a first proportional position according to a quantitative relationship between the upper and lower target grids; Determine the lowermost target grid as the second target grid, and determine a second proportional position according to the upper and lower quantity relationship of the second target grid; Determine the leftmost target grid as the third target grid, and determine a third proportional position based on the left and right quantity relationship of the third target grid; Determine the rightmost target grid as the fourth target grid, and determine a fourth proportional position based on the left and right quantity relationship of the fourth target grid; The relative position information of the visual element is determined according to the first proportional position, the second proportional position, the third proportional position and the fourth proportional position, and the position text corresponding to the position length is generated according to the relative position information.
9. A data generation system corresponding to the data generation method according to claim 1, characterized in that: include: A recognition module is used to identify visual elements of the input image and determine a description weight based on the pixel ratio of each visual element; a determination module, configured to determine, when the pixel ratio satisfies a secondary splitting condition, an element ratio of each sub-element in the visual element and a descriptive attribute corresponding to the sub-element, the descriptive attribute including a forward attribute and a reverse attribute; a processing module, configured to process the description weight according to the description attribute and the element ratio, and determine the descriptor weight of each sub-element; A generation module is used to perform association processing on the required length based on the description weight and / or descriptor weight, and the level of the visual element and / or sub-element, to generate structured directional text.
Citation Information
Patent Citations
Manuscript generation method and related device, electronic equipment and storage medium
CN117033567A
Camera shooting picture quality optimization method and camera
CN118784974A