Photograph composition model composition sample set construction method and device, equipment and medium
By using a quantitative labeling system and structured prompt word templates, combined with a large visual language model and a regression model, a high-quality photographic composition dataset is automatically constructed. This solves the problems of high cost of manual annotation and lack of quantitative standards in existing technologies, and enables efficient training and application of intelligent composition models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MALANSHAN AUDIO & VIDEO LABORATORY
- Filing Date
- 2026-03-25
- Publication Date
- 2026-06-16
AI Technical Summary
The construction of existing photographic composition datasets relies on manual annotation, which is costly, inefficient, and lacks quantitative standards and deep semantic information, thus limiting the development of intelligent composition technology.
By acquiring a set of photographic composition rules, deconstructing them into a quantitative labeling system, designing structured prompt word templates, automatically labeling them using a large visual language model, evaluating errors and filtering samples using a target composition regression model, correcting abnormal samples on the user end, and constructing a high-quality, multi-dimensional composition dataset.
It achieves efficient and automated multi-dimensional graphing and annotation, constructs an information-rich and clearly structured dataset, provides a reliable foundation for intelligent graphing models, and promotes the application of the technology in real-world scenarios.
Smart Images

Figure CN122223474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a method, apparatus, device, and medium for constructing a composition sample set for photographic composition models. Background Technology
[0002] Composition, as a crucial element in photographic art that determines visual expression, has long relied on the photographer's subjective aesthetic sense and accumulated experience, involving complex functions such as the arrangement of elements, emotional communication, and visual guidance. Classic composition rules such as the rule of thirds, symmetry, leading lines, and the golden ratio form the cornerstone of photographic aesthetics. However, the composition process itself is highly subjective and ambiguous; different photographers may have drastically different judgments about the same scene, resulting in a lack of clear quantitative standards and high-quality supervision signals for automated modeling. Traditional composition datasets rely heavily on manual annotation, which is costly, inefficient, and prone to inconsistent high-quality data due to individual aesthetic differences. Existing datasets often only contain single-dimensional labels such as composition coordinates, lacking descriptions of deeper semantic information such as composition logic and scene type. This makes it difficult to support the model's comprehensive understanding of the composition process and improve its generalization ability, becoming a core bottleneck restricting the development of intelligent composition technology.
[0003] In summary, establishing quantitative standards and constructing high-quality, multi-dimensional graph datasets are pressing issues that need to be addressed. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for constructing a composition sample set for photographic composition models, capable of establishing quantitative standards and constructing a high-quality, multi-dimensional composition dataset. The specific solution is as follows: Firstly, this application provides a method for constructing a composition sample set for a photographic composition model, including: Obtain a set of composition rules for photography, deconstruct the set of composition rules, and define a quantitative labeling system for the composition process from a preset dimension; Based on the aforementioned quantitative labeling system, a structured target prompt word template is designed. The initial set of photographic images is input into the visual language large model. The composition annotation data corresponding to each image in the initial set of photographic images is output according to the target prompt word template. The initial set of photographic images and the corresponding composition annotation data are integrated to generate an initial composition sample set. A target composition regression model is determined, and the target composition annotation error of each sample in the initial composition sample set is evaluated using the target composition regression model. Based on the target composition annotation error, samples that meet the preset anomaly conditions are selected from the initial composition sample set as the filtered sample set. The abnormal samples in the filtered sample set are corrected by the user terminal to obtain a corrected sample set. The corresponding samples in the preliminary dataset are replaced by the corrected sample set to generate a target sample set. The target sample set is used to train a photographic composition model to obtain a trained composition model. The trained composition model is then used to respond to composition requests and output the corresponding photographic composition results.
[0005] Optionally, the preset dimensions include any one or more of the following: composition coordinate dimension, composition rating dimension, scene type dimension, frame type dimension, composition type dimension, subject description dimension, and composition rule description dimension.
[0006] Optionally, the step of inputting the initial photographic image set into the visual language large model, outputting the compositional annotation data corresponding to each image in the initial photographic image set according to the target prompt word template, and integrating the initial photographic image set and the corresponding compositional annotation data to generate an initial compositional sample set includes: Obtain an initial set of photographic images; the initial set of photographic images is a set of images that have not been labeled or mapped. The initial set of photographic images is input into the visual language model, and the compositional annotation data of each image in the initial set of photographic images is output according to the target prompt word template. The icon annotation data is paired with each image in the initial photographic image set to generate a pairing result; Based on the pairing results, integrate each image in the initial photographic image set and the corresponding composition annotation data to generate an initial composition sample set.
[0007] Optionally, the determination of the target mapping regression model includes: Obtain the initial composition regression model and the corresponding training samples; wherein, the training samples include images and corresponding composition annotation data; The initial composition regression model is trained using training samples to obtain a target composition regression model for labeling predicted cropping boxes in an image.
[0008] Optionally, evaluating the target composition annotation error of each sample in the initial composition sample set using the target composition regression model includes: The first composition coordinates of each sample in the initial composition sample set are determined using the target composition regression model, and the corresponding first composition box is determined based on the first composition coordinates; Obtain the second composition coordinates of each sample in the initial composition sample set, and determine the corresponding second composition frame based on the second composition coordinates; the second composition coordinates are the composition coordinates of each sample in the initial composition sample set as determined by experts; The corresponding coordinate regression loss is determined based on the first and second plot coordinates using the smoothed L1 loss algorithm. The overlap between the first and second frame frames is determined using the intersection-union loss algorithm based on the first and second frame frames coordinates. The coordinate regression loss and the overlap are weighted and fused to obtain the first mapping annotation error of each sample in the initial mapping sample set; Construct a target consistency discriminant, and use the target consistency discriminant to determine the second construction annotation error between the visual features of the images in the initial construction sample set and the corresponding construction annotation data; Obtain a preset composition rule library, and perform a preset logic verification operation on the composition annotation data according to the preset composition rule library to generate a third composition annotation error; The first, second, and third composition annotation errors are weighted and fused to obtain the target composition annotation error of each sample in the initial composition sample set.
[0009] Optionally, the step of selecting samples that meet preset anomaly conditions from the initial composition sample set based on the target composition annotation error, as the selected sample set, includes: The initial composition sample set is used to select samples whose target composition annotation error is greater than a preset error threshold, and these samples are used as the filtered sample set. Alternatively, the samples in the initial composition sample set are arranged in a preset order based on the target composition annotation error to generate a corrected sample set sequence, and samples with a preset proportion are selected from the corrected sample set sequence according to the preset order as the filtered sample set.
[0010] Optionally, after correcting the abnormal samples in the filtered sample set through the user terminal to obtain the corrected sample set, the method further includes: The corrected sample set is structured and stored according to a preset standard format to generate a standard dataset; wherein, the standard dataset includes image files, label files corresponding to the image files, data partitioning information for model training, and dataset version description, and the data partitioning information includes training set information, validation set information, and test set information; The standard dataset is maintained based on a preset data version management mechanism so that it can be used to train photographic composition models.
[0011] Secondly, this application provides a device for constructing a composition sample set for a photographic composition model, comprising: The system definition module is used to acquire a set of composition rules during photography, deconstruct the set of composition rules, and define a quantitative label system for the composition process from a preset dimension. The sample set generation module is used to design a structured target prompt word template based on the quantized label system, input the initial photographic image set into the visual language large model, output the composition annotation data corresponding to each image in the initial photographic image set according to the target prompt word template, and integrate the initial photographic image set and the corresponding composition annotation data to generate an initial composition sample set. The sample set acquisition module is used to determine the target composition regression model, use the target composition regression model to evaluate the target composition annotation error of each sample in the initial composition sample set, and select samples that meet the preset anomaly conditions from the initial composition sample set according to the target composition annotation error, as the filtered sample set. The sample set correction module is used to correct abnormal samples in the filtered sample set through the user terminal to obtain a corrected sample set. Based on the corrected sample set, the corresponding samples in the initial dataset are replaced to generate a target sample set. The target sample set is used to train a photographic composition model to obtain a trained composition model. The trained composition model is then used to respond to composition requests and output the corresponding photographic composition results.
[0012] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the method for constructing a composition sample set of photographic composition models as described above.
[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned method for constructing a composition sample set of a photographic composition model.
[0014] In summary, this application acquires a set of composition rules for photography, deconstructs the set of composition rules to define a quantitative labeling system for the composition process from a preset dimension; designs a structured target prompt word template based on the quantitative labeling system, inputs an initial set of photographic images into a visual language model, outputs composition annotation data corresponding to each image in the initial set of photographic images according to the target prompt word template, integrates the initial set of photographic images and the corresponding composition annotation data to generate an initial composition sample set; determines a target composition regression model, uses the target composition regression model to evaluate the target composition annotation error of each sample in the initial composition sample set, and filters samples that meet preset abnormal conditions from the initial composition sample set according to the target composition annotation error as a filtered sample set; corrects the abnormal samples in the filtered sample set through the user terminal to obtain a corrected sample set, replaces the corresponding samples in the initial set with the corrected sample set to generate a target sample set, so as to train a photographic composition model using the target sample set to obtain a trained composition model, and uses the trained composition model to respond to composition requests and output corresponding photographic composition results. As described above, this application first acquires and deconstructs photographic composition rules, defines a quantified labeling system from preset dimensions, and designs a structured prompt word template based on this. Then, the initial set of photographic images is input into a large visual language model, which automatically outputs the composition annotation data for each image using this template, thereby integrating and generating an initial composition sample set. Next, the annotation error of each sample in the initial set is evaluated using a defined target composition regression model, and samples that meet preset anomaly conditions are selected as the filtered sample set based on this error. Finally, these anomaly samples are manually corrected by the user, resulting in a high-quality corrected sample set for subsequent training of the photographic composition model. In this way, the large visual language model is used to automate and quantify the composition process in multiple dimensions, decomposing composition into multiple dimensions such as coordinates, rating, scene, frame, type, subject description, and rule description, constructing a high-quality composition dataset that is rich in information and has a clear structure. By introducing a lightweight composition model for data filtering and manual verification, a closed loop is established, further ensuring data quality, thus providing a reliable foundation for training the composition model and promoting the widespread application of intelligent composition technology in real-world scenarios. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0016] Figure 1This is a flowchart of a method for constructing a composition sample set for a photographic composition model disclosed in this application; Figure 2 This application discloses a flowchart of a method for constructing a composition sample set for a specific photographic composition model. Figure 3 This is a schematic diagram of a composition sample set construction device for a photographic composition model disclosed in this application; Figure 4 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Currently, composition, as a crucial element in photographic art that determines visual expression, has long relied on the photographer's subjective aesthetics and accumulated experience, involving complex functions such as the arrangement of elements, emotional communication, and visual guidance. Classic composition rules such as the rule of thirds, symmetry, leading lines, and the golden ratio form the cornerstone of photographic aesthetics. However, the composition process itself is highly subjective and ambiguous; different photographers may have drastically different judgments about the same scene, resulting in a lack of clear quantitative standards and high-quality supervision signals for automated modeling. Traditional composition dataset construction relies heavily on manual annotation, which is costly, inefficient, and difficult to generate consistent high-quality data due to individual aesthetic differences. Existing datasets often only contain single-dimensional labels such as composition coordinates, lacking descriptions of deeper semantic information such as composition logic and scene type, making it difficult to support the model's comprehensive understanding of the composition process and improve its generalization ability, becoming a core bottleneck restricting the development of intelligent composition technology. To solve the above technical problems, this application discloses a method, apparatus, device, and medium for constructing a composition sample set for a photographic composition model, which can establish quantitative standards and construct a high-quality, multi-dimensional composition dataset.
[0019] See Figure 1 As shown, this embodiment of the invention discloses a method for constructing a composition sample set for a photographic composition model, including: Step S11: Obtain the set of composition rules for photography, and deconstruct the set of composition rules to define a quantitative label system for the composition process from a preset dimension.
[0020] In this embodiment, the composition process is first systematically deconstructed, and preset dimensions are clearly defined. These preset dimensions include any one or more of the following: composition coordinate dimension, composition rating dimension, scene type dimension, frame type dimension, composition type dimension, subject description dimension, and composition rule description dimension. It is important to know that composition coordinates refer to the position and size of the cropping box; composition rating is an aesthetic score, such as 1-5 points; scene type includes images such as portraits, landscapes, and architecture; frame type includes close-up, foreground, medium shot, full shot, and background; composition type includes the rule of thirds, symmetrical composition, and leading line composition; subject description is a natural language description of the core object in the image; and composition rule description is a linguistic explanation of the aesthetic rules followed in the composition. In this way, the multi-dimensional quantitative labeling system provides a clear structured framework for subsequent annotation, ensuring the comprehensiveness and consistency of data annotation.
[0021] Step S12: Design a structured target prompt word template based on the quantized label system, input the initial photographic image set into the visual language large model, output the composition annotation data corresponding to each image in the initial photographic image set according to the target prompt word template, and integrate the initial photographic image set and the corresponding composition annotation data to generate an initial composition sample set.
[0022] In this embodiment, based on the aforementioned multi-dimensional labeling system, a structured prompt word template is designed. First, the core analytical dimensions are anchored using the established quantitative labeling system, integrating scattered image evaluation metrics into a logical instruction framework. The initial prompts provide contextual guidance relevant to the image scene, avoiding direct, rigid instructions. Instead, the model is guided to immerse itself in perceiving the image information, allowing it to accurately capture the visual core and atmosphere of the scene, thus solidifying its understanding. This is followed by a concrete task description module, abandoning vague expressions and breaking down core tasks such as composition, scoring, and classification into precise instructions that the model can execute, clearly defining each analytical task. The template refines boundary rules, such as detailed numerical rules for composition coordinates, evaluation criteria for scoring intervals, and enumeration categories for various labels, to avoid misunderstandings in the model. Finally, the output format requirements module standardizes all analytical results into a unified structured paradigm, specifying data types and expression formats while also considering readability and reusability. This transforms fragmented information into standardized structured data, ensuring that every statement in the template closely adheres to the label system and aligns with the task objectives. This guarantees the rigor of instructions while mitigating a mechanical feel through smooth sentence transitions, allowing the model to follow the template's logical flow and accurately output complete analytical results that meet expectations. The prompt word template includes modules such as image context description, task instructions, and output format requirements, ensuring that the model outputs structured data that meets expectations. For example: "Based on the image content, please output the composition coordinates (x, y, w, h, where x and y are the coordinates of the top left corner of the composition frame, and w and h are the width and height of the composition frame, respectively), composition score (1-5), scene type (portrait, nature, architecture, etc.), frame type (close-up, foreground, medium shot, full shot, background, etc.), composition type (rules of thirds, center, symmetry, leading lines, etc.), subject description, and composition rule description."
[0023] Furthermore, after designing the prompt word template, an initial set of photographic images is obtained. This initial set of images consists of images without any compositional annotations. The initial set of images is input into a large-scale visual language model, and the compositional annotation data for each image in the initial set is output based on the target prompt word template. The compositional annotation data is paired with each image in the initial set to generate a pairing result. Based on the pairing result, the images in the initial set and their corresponding compositional annotation data are integrated to generate an initial compositional sample set. Specifically, after completing the customized design of the prompt word template, the initial set of photographic images is prepared and obtained. This set consists entirely of original photographic materials without any compositional annotation processing, including original images of various shooting scenes and aspect ratios, without any pre-defined annotation information. Subsequently, these initial photographic images are imported into the large-scale visual language model in batches. Relying on the determined target prompt word template as the parsing basis, the model is driven to visually decompose each image and output the compositional annotation data corresponding to each image. Next, data pairing was performed, binding each individual image in the initial photographic image set to the specific compositional annotation data generated by the model, ensuring a one-to-one match between images and annotation information without any errors. Finally, based on all pairing results, all images and their corresponding compositional annotation data in the initial photographic image set were organized and integrated to form a complete initial compositional sample set: ; in, Original image; For the composition coordinates; Rate the composition; Scene type; Image format; Composition type; The main description; This describes the composition rules.
[0024] It is important to know that the icon annotation process is broken down into four logically progressive reasoning stages. Each stage outputs a structured intermediate result, which serves as the input for the next stage, forming a complete reasoning chain and significantly improving the accuracy and interpretability of large model annotation.
[0025] Phase 1: Scene and Subject Analysis. Input the original image; the model outputs the scene type. Main description and the main center position .
[0026] Phase Two: Matching Compositional Intent with Composition Type. Based on the scene and subject information identified in Phase One, the model infers the photographer's compositional intent. Examples of techniques include highlighting the main subject, creating depth, and emphasizing symmetry, and selecting a matching composition type. .
[0027] Phase Three: Composition Rule Reasoning Phase. Based on the scene type determined in the first two phases. Compositional Intent and composition type The specific graphing rules that model inference should follow And describe it in natural language.
[0028] Phase Four: Geometric Coordinate Reasoning. Based on the semantic reasoning results from the first three phases, the model outputs the final graph coordinates. To transform compositional rules into geometric constraints, a rule-based coordinate generation function is defined. This function calculates the ideal position and size of the cropping frame based on the subject's location, composition type, and rule description. The final coordinate calculation formula is as follows: ; The above four-stage reasoning process can be uniformly represented as a recursive form: , , , ,in, For the first The inference function of the stage, Given an input image, the final output is the composition coordinates. Meanwhile, in the above multi-stage reasoning process, each stage can dynamically retrieve examples similar to the intermediate state of the current stage from a pre-built graph example library, further improving the accuracy and consistency of reasoning at each stage through few-shot context learning.
[0029] Step S13: Determine the target composition regression model, use the target composition regression model to evaluate the target composition annotation error of each sample in the initial composition sample set, and select samples that meet the preset anomaly conditions from the initial composition sample set based on the target composition annotation error, as the filtered sample set.
[0030] In this embodiment, an initial composition regression model and corresponding training samples are obtained; wherein, the training samples include images and corresponding composition annotation data; the initial composition regression model is trained using the training samples to obtain a target composition regression model for annotating predicted cropping boxes in images. Specifically, in constructing a lightweight composition regression model... In the process, the initial model architecture was first designed. Based on a lightweight network framework, the model structure was simplified, the number of convolutional layers was reduced, and depthwise separable convolutions were used instead of traditional convolutions, reducing the number of model parameters and computational complexity while maintaining core regression capabilities. Subsequently, the initial dataset was compiled. As training samples, this dataset contains a large amount of image data from diverse scenes, along with detailed compositional annotation data for each image. The annotation information specifically covers key compositional parameters such as the coordinates and aspect ratio of the optimal cropping box in the image. During the training phase, the dataset is proportionally divided into training and validation sets. The image data is input into the initial compositional regression model, which outputs predicted cropping box parameters through forward propagation. The regression loss between the predicted and actual values is then calculated using the labeled data as a benchmark. The network weight parameters of the model are iteratively optimized using a gradient descent algorithm. During training, the loss changes on the validation set are continuously monitored, and hyperparameters such as the learning rate and batch size are adjusted until the regression error on the validation set stabilizes. Finally, a target compositional regression model capable of accurately annotating and predicting cropping boxes in images is obtained.
[0031] Furthermore, the first composition coordinates of each sample in the initial composition sample set are determined using the target composition regression model, and the corresponding first composition box is determined based on the first composition coordinates; the second composition coordinates of each sample in the initial composition sample set are obtained, and the corresponding second composition box is determined based on the second composition coordinates; the second composition coordinates are the composition coordinates of each sample in the initial composition sample set as determined by experts; the corresponding coordinate regression loss is determined using the smoothed L1 loss algorithm based on the first and second composition coordinates; the overlap between the first and second composition boxes is determined using the intersection-union loss algorithm based on the first and second composition coordinates. The coordinate regression loss and the overlap are weighted and fused to obtain the first composition annotation error of each sample in the initial composition sample set; a target consistency discriminator is constructed, and the second composition annotation error between the visual features of the images in the initial composition sample set and the corresponding composition annotation data is determined by the target consistency discriminator; a preset composition rule library is obtained, and a preset logic verification operation is performed on the composition annotation data according to the preset composition rule library to generate a third composition annotation error; the first composition annotation error, the second composition annotation error and the third composition annotation error are weighted and fused to obtain the target composition annotation error of each sample in the initial composition sample set. Specifically, after obtaining the target composition regression model, all sample images in the initial composition sample set are first input into the model one by one, and the model outputs the first composition coordinates corresponding to each sample. Based on this coordinate information, the corresponding first composition box is drawn on the image. Simultaneously, second composition coordinates, pre-judged and labeled by experts, are retrieved from the initial composition sample set. These coordinates represent the optimal composition position determined by combining professional composition rules and visual aesthetics. Based on these coordinates, a second composition box is also drawn on the image. Next, a smoothed L1 loss algorithm is used, substituting the first and second composition coordinates of each sample into the calculation. This algorithm reduces the impact of outliers on the loss calculation, resulting in an accurate coordinate regression loss. ; in, = (x, y, w, h) are the actual composition coordinates, i.e., the second composition coordinates; These are the predicted cropping frame coordinates, i.e., the first composition coordinates; .
[0032] Subsequently, the intersection-union ratio loss algorithm is used to calculate the ratio of the intersection area to the union area of the first and second frame frames, thereby quantifying the degree of overlap between the two frame frames: ; Among them, | | indicates the area of the intersection and union of the clipping frames.
[0033] Finally, based on the model training requirements, reasonable weight coefficients are set, and the coordinate regression loss and overlap value are weighted and fused. By weighted summation, the first construction annotation error corresponding to each sample in the initial construction sample set is finally obtained: ; Here, λ represents the balancing weight. This fully reflects the degree of deviation between the model's predicted graph and the expert-annotated graph.
[0034] Simultaneously, a multimodal consistency discriminator is constructed. The input image's visual features and text labels across various dimensions are compared, and the output is a consistency score, i.e., the second-order image annotation error. .
[0035] In addition, a rule base based on common sense about composition is established to logically validate labels. For example, under the rule of thirds, the subject should be located in the direction of the leading line, and the subject should not be completely centered. This generates a rule violation score, i.e., the third composition labeling error. The final comprehensive abnormality score is... .
[0036] Then, after calculating the error of all samples, a selection sample is set based on the error distribution. Samples with target mapping annotation errors greater than a preset error threshold can be selected from the initial mapping sample set as the selected sample set; or, the samples in the initial mapping sample set can be arranged according to a preset order based on the target mapping annotation error to generate a corrected sample set sequence, and samples with a preset proportion in the corrected sample set sequence can be selected as the selected sample set. Specifically, samples with error values greater than the overall mean plus twice the standard deviation are taken as underfitting samples, or the top p% of samples with the largest errors are directly selected as the selected sample set. These samples usually correspond to complex mappings that the model struggles to learn or anomalous data with inconsistent annotations, requiring further manual verification. This mechanism effectively identifies and locates low-quality annotations in the initial dataset, providing a clear target for subsequent corrections.
[0037] Step S14: Correct the abnormal samples in the filtered sample set through the user terminal to obtain a corrected sample set. Replace the corresponding samples in the preliminary dataset with the corrected sample set to generate a target sample set. Use the target sample set to train the photographic composition model to obtain a trained composition model. Use the trained composition model to respond to composition requests and output the corresponding photographic composition results.
[0038] In this embodiment, the selected abnormal samples undergo manual review. Professional photographers or annotators correct the original annotation information to ensure the accuracy and consistency of labels such as composition coordinates, scores, and rule descriptions. The corrected samples are then reintegrated into the initial dataset, replacing the original abnormal samples, to form the final high-quality multimodal composition target sample set. Simultaneously, the changes in labels before and after the correction are recorded to build a correction knowledge base, which can be used in subsequent iterations to optimize suggestion templates or fine-tune lightweight models, enabling the reuse of correction knowledge. The target sample set can be used for iterative training of high-performance composition models, effectively enhancing the model's generalization and adaptation capabilities and practical performance. This allows the trained composition model to respond quickly to user composition requests and accurately output photographic composition results that meet visual aesthetic requirements.
[0039] Furthermore, the corrected sample set is structured and stored according to a preset standard format to generate a standard dataset. This standard dataset includes image files, corresponding label files, data partitioning information for model training, and a dataset version description. The data partitioning information includes training set information, validation set information, and test set information. The standard dataset is maintained based on a preset data version management mechanism to facilitate the training of the photographic composition model. Specifically, the final dataset is structured and stored according to a standard format, including image files, label files, data partitioning information (training set / validation set / test set), and version description. A data version management mechanism is also established to facilitate subsequent model training, evaluation, and iterative optimization, ensuring the long-term availability and reproducibility of the dataset.
[0040] As described above, this embodiment first obtains and deconstructs photographic composition rules, defines a quantified labeling system from preset dimensions, and designs a structured prompt word template based on this. Then, the initial set of photographic images is input into a large visual language model, which automatically outputs the composition annotation data for each image using the template, thereby integrating and generating an initial composition sample set. Next, the annotation error of each sample in the initial set is evaluated using a determined target composition regression model, and samples that meet preset anomaly conditions are selected as the filtered sample set based on this error. Finally, these anomaly samples are manually corrected by the user, resulting in a high-quality corrected sample set for subsequent training of the photographic composition model. In this way, the large visual language model is used to automate and quantify the composition process in multiple dimensions, decomposing composition into multiple dimensions such as coordinates, rating, scene, frame, type, subject description, and rule description, constructing a high-quality composition dataset that is rich in information and has a clear structure. By introducing a lightweight composition model for data filtering and manual verification, the data quality is further guaranteed, thus providing a reliable foundation for training the composition model and promoting the widespread application of intelligent composition technology in real-world scenarios.
[0041] As can be seen from the previous embodiment, this application discloses a method for constructing a composition sample set for a photographic composition model, which can establish quantitative standards and construct a high-quality, multi-dimensional composition dataset, such as for the implementation of intelligent photographic composition assistants on mobile devices. Next, we will discuss methods for constructing such... Figure 2 The method for constructing the composition sample set of the photographic composition model shown is explained in detail.
[0042] First, this application comprehensively collects a set of general photographic composition rules for various mobile phones, such as the rule of thirds, symmetrical composition, leading lines composition, golden ratio, and subject centering. This set of rules is then deeply deconstructed, and a set of quantifiable and labelable composition tags is defined in detail from preset dimensions such as composition frame coordinates, subject proportion, visual center of gravity, and line direction, so that abstract composition aesthetics can be transformed into standardized numerical indicators.
[0043] Then, based on this quantitative labeling system, a structured, multi-round interactive target prompt word template with a unified format and clear direction was designed. An initial set of real-world mobile phone photos covering various themes such as portraits, landscapes, food, and street photography was input into a general visual language model. Combining thought chain reasoning and dynamic example retrieval, the model was guided to complete high-precision composition annotations step by step, driving the model to strictly follow the prompt word template requirements to generate composition annotation data for each image. Finally, the initial photos and the matched annotation data were integrated to generate an initial composition sample set for subsequent filtering.
[0044] Then, the pre-trained target mapping regression model is invoked. For each sample in the initial mapping sample set, a multi-dimensional anomaly screening mechanism is constructed by combining regression error, multimodal consistency score, and rule violation degree. The corresponding target mapping annotation error is calculated. Then, based on the preset error threshold, annotation misalignment and other anomaly judgment conditions, unqualified samples are screened out from the initial sample set to obtain a filtered sample set that retains only anomaly samples.
[0045] Subsequently, relying on the user-side backend of the mobile photography assistant, professional photographers meticulously corrected and calibrated the abnormal annotations in the filtered sample set, forming a precisely labeled and corrected sample set. These high-quality corrected samples then replaced the flawed samples in the original preliminary dataset, ultimately generating the target sample set suitable for model training. Simultaneously, an iterative closed-loop "annotation-training-filtering-correction" strategy and an active learning strategy were introduced to achieve progressive quality optimization and self-enhancing capabilities of the dataset.
[0046] Finally, the model is trained using the target sample set to obtain a convergent and stable post-trained composition model. When mobile phone users initiate composition requests such as composition optimization and automatic cropping, the model can quickly respond and output photographic composition results that conform to professional aesthetics and are suitable for mobile phone shooting scenarios.
[0047] See Figure 3 As shown, this embodiment of the invention discloses a device for constructing a composition sample set for a photographic composition model, comprising: The system definition module 11 is used to acquire a set of composition rules during photography, deconstruct the set of composition rules, and define a quantitative label system for the composition process from a preset dimension. The sample set generation module 12 is used to design a structured target prompt word template based on the quantized label system, input the initial photographic image set into the visual language large model, output the composition annotation data corresponding to each image in the initial photographic image set according to the target prompt word template, and integrate the initial photographic image set and the corresponding composition annotation data to generate an initial composition sample set. The sample set acquisition module 13 is used to determine the target composition regression model, use the target composition regression model to evaluate the target composition annotation error of each sample in the initial composition sample set, and select samples that meet the preset anomaly conditions from the initial composition sample set according to the target composition annotation error, as the filtered sample set. The sample set correction module 14 is used to correct the abnormal samples in the filtered sample set through the user terminal to obtain a corrected sample set, and replace the corresponding samples in the preliminary dataset with the corrected sample set to generate a target sample set, so as to use the target sample set to train the photographic composition model, obtain the trained composition model, and use the trained composition model to respond to the composition request and output the corresponding photographic composition result.
[0048] As described above, this application first acquires and deconstructs photographic composition rules, defines a quantified labeling system from preset dimensions, and designs a structured prompt word template based on this. Then, the initial set of photographic images is input into a large visual language model, which automatically outputs the composition annotation data for each image using this template, thereby integrating and generating an initial composition sample set. Next, the annotation error of each sample in the initial set is evaluated using a defined target composition regression model, and samples that meet preset anomaly conditions are selected as the filtered sample set based on this error. Finally, these anomaly samples are manually corrected by the user, resulting in a high-quality corrected sample set for subsequent training of the photographic composition model. In this way, the large visual language model is used to automate and quantify the composition process in multiple dimensions, decomposing composition into multiple dimensions such as coordinates, rating, scene, frame, type, subject description, and rule description, constructing a high-quality composition dataset that is rich in information and has a clear structure. By introducing a lightweight composition model for data filtering and manual verification, a closed loop is established, further ensuring data quality, thus providing a reliable foundation for training the composition model and promoting the widespread application of intelligent composition technology in real-world scenarios.
[0049] In some specific implementations, the preset dimensions include any one or more of the following: composition coordinate dimension, composition rating dimension, scene type dimension, frame type dimension, composition type dimension, subject description dimension, and composition rule description dimension.
[0050] In some specific implementations, the sample set generation module 12 may specifically include: An image set acquisition unit is used to acquire an initial photographic image set; the initial photographic image set is an image set that has not been labeled. The data output unit is used to input the initial set of photographic images into the visual language large model and output the compositional annotation data of each image in the initial set of photographic images according to the target prompt word template. The result generation unit is used to pair the icon annotation data with each image in the initial photographic image set to generate a pairing result; The sample set generation unit is used to integrate each image in the initial photographic image set and the corresponding composition annotation data according to the pairing result to generate an initial composition sample set.
[0051] In some specific implementations, the sample set acquisition module 13 may specifically include: The sample acquisition unit is used to acquire the initial mapping regression model and the corresponding training samples; wherein, the training samples include images and corresponding mapping annotation data; The model training unit is used to train the initial composition regression model using training samples to obtain a target composition regression model for labeling predicted cropping boxes in an image.
[0052] In some specific implementations, the sample set acquisition module 13 may specifically include: The first composition frame determination unit is used to determine the first composition coordinates of each sample in the initial composition sample set using the target composition regression model, and to determine the corresponding first composition frame based on the first composition coordinates. The second composition frame determination unit is used to obtain the second composition coordinates of each sample in the initial composition sample set, and determine the corresponding second composition frame based on the second composition coordinates; the second composition coordinates are the composition coordinates of each sample in the initial composition sample set as determined by experts; The loss determination unit is used to determine the corresponding coordinate regression loss based on the first and second plot coordinates using a smoothed L1 loss algorithm. The overlap determination unit is used to determine the overlap between the first frame and the second frame based on the first frame coordinates and the second frame coordinates using the cross-union ratio loss algorithm. The first error acquisition unit is used to weight and fuse the coordinate regression loss and the overlap to obtain the first mapping annotation error of each sample in the initial mapping sample set; An error determination unit is used to construct a target consistency discriminator, and to determine a second construction annotation error between the visual features of the images in the initial construction sample set and the corresponding construction annotation data through the target consistency discriminator; An error generation unit is used to acquire a preset composition rule library and perform a preset logical verification operation on the composition annotation data according to the preset composition rule library to generate a third composition annotation error. The second error acquisition unit is used to perform weighted fusion of the first construction annotation error, the second construction annotation error and the third construction annotation error to obtain the target construction annotation error of each sample in the initial construction sample set.
[0053] In some specific implementations, the sample set acquisition module 13 may specifically include: The sample filtering unit is used to filter samples from the initial composition sample set whose target composition annotation error is greater than a preset error threshold, and use them as the filtered sample set; or, to arrange each sample in the initial composition sample set according to the target composition annotation error in a preset order to generate a corrected sample set sequence, and to select samples with a preset proportion in the corrected sample set sequence according to the preset order, and use them as the filtered sample set.
[0054] In some specific embodiments, the apparatus for constructing the composition sample set of the photographic composition model may further include: The dataset generation module is used to structure and store the corrected sample set according to a preset standard format to generate a standard dataset; wherein, the standard dataset includes image files, label files corresponding to the image files, data partitioning information for model training, and dataset version description, and the data partitioning information includes training set information, validation set information, and test set information; The dataset maintenance module is used to maintain the standard dataset based on a preset data version management mechanism, so as to use the standard dataset to train the photographic composition model.
[0055] Furthermore, embodiments of this application also disclose an electronic device, Figure 4This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the method for constructing a composition sample set of photographic composition models disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0056] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0057] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk, or optical disk, etc. The resources stored thereon can include an operating system 221, computer programs 222, etc., and the storage method can be temporary storage or permanent storage.
[0058] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the method for constructing a composition sample set of a photographic composition model executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.
[0059] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned method for constructing a composition sample set of the photographic composition model. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0060] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0061] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0062] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0063] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0064] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for constructing a composition sample set for a photographic composition model, characterized in that, include: Obtain a set of composition rules for photography, deconstruct the set of composition rules, and define a quantitative labeling system for the composition process from a preset dimension; Based on the aforementioned quantitative labeling system, a structured target prompt word template is designed. The initial set of photographic images is input into the visual language large model. The composition annotation data corresponding to each image in the initial set of photographic images is output according to the target prompt word template. The initial set of photographic images and the corresponding composition annotation data are integrated to generate an initial composition sample set. A target composition regression model is determined, and the target composition annotation error of each sample in the initial composition sample set is evaluated using the target composition regression model. Based on the target composition annotation error, samples that meet the preset anomaly conditions are selected from the initial composition sample set as the filtered sample set. The abnormal samples in the filtered sample set are corrected by the user terminal to obtain a corrected sample set. The corresponding samples in the preliminary dataset are replaced by the corrected sample set to generate a target sample set. The target sample set is used to train a photographic composition model to obtain a trained composition model. The trained composition model is then used to respond to composition requests and output the corresponding photographic composition results.
2. The method for constructing a composition sample set for a photographic composition model according to claim 1, characterized in that, The preset dimensions include any one or more of the following: composition coordinate dimension, composition rating dimension, scene type dimension, frame type dimension, composition type dimension, subject description dimension, and composition rule description dimension.
3. The method for constructing a composition sample set for a photographic composition model according to claim 1, characterized in that, The process of inputting an initial set of photographic images into a large visual language model, outputting compositional annotation data corresponding to each image in the initial set of photographic images based on the target prompt word template, and integrating the initial set of photographic images and the corresponding compositional annotation data to generate an initial compositional sample set includes: Obtain an initial set of photographic images; the initial set of photographic images is a set of images that have not been labeled or mapped. The initial set of photographic images is input into the visual language model, and the compositional annotation data of each image in the initial set of photographic images is output according to the target prompt word template. The icon annotation data is paired with each image in the initial photographic image set to generate a pairing result; Based on the pairing results, integrate each image in the initial photographic image set and the corresponding composition annotation data to generate an initial composition sample set.
4. The method for constructing a composition sample set for a photographic composition model according to claim 3, characterized in that, The determined target mapping regression model includes: Obtain the initial composition regression model and the corresponding training samples; wherein, the training samples include images and corresponding composition annotation data; The initial composition regression model is trained using training samples to obtain a target composition regression model for labeling predicted cropping boxes in an image.
5. The method for constructing a composition sample set for a photographic composition model according to claim 4, characterized in that, The evaluation of the target composition annotation error of each sample in the initial composition sample set using the target composition regression model includes: The first composition coordinates of each sample in the initial composition sample set are determined using the target composition regression model, and the corresponding first composition box is determined based on the first composition coordinates; Obtain the second composition coordinates of each sample in the initial composition sample set, and determine the corresponding second composition frame based on the second composition coordinates; the second composition coordinates are the composition coordinates of each sample in the initial composition sample set as determined by experts; The corresponding coordinate regression loss is determined based on the first and second plot coordinates using the smoothed L1 loss algorithm. The overlap between the first and second frame frames is determined using the intersection-union loss algorithm based on the first and second frame frames coordinates. The coordinate regression loss and the overlap are weighted and fused to obtain the first mapping annotation error of each sample in the initial mapping sample set; Construct a target consistency discriminant, and use the target consistency discriminant to determine the second construction annotation error between the visual features of the images in the initial construction sample set and the corresponding construction annotation data; Obtain a preset composition rule library, and perform a preset logic verification operation on the composition annotation data according to the preset composition rule library to generate a third composition annotation error; The first, second, and third composition annotation errors are weighted and fused to obtain the target composition annotation error of each sample in the initial composition sample set.
6. The method for constructing a composition sample set for a photographic composition model according to claim 1, characterized in that, The step of selecting samples from the initial composition sample set that meet preset anomaly conditions based on the target composition annotation error, as the selected sample set, includes: The initial composition sample set is used to select samples whose target composition annotation error is greater than a preset error threshold, and these samples are used as the filtered sample set. Alternatively, the samples in the initial composition sample set are arranged in a preset order based on the target composition annotation error to generate a corrected sample set sequence, and samples with a preset proportion are selected from the corrected sample set sequence according to the preset order as the filtered sample set.
7. The method for constructing a composition sample set for a photographic composition model according to claim 1, characterized in that, After correcting the abnormal samples in the filtered sample set through the user terminal to obtain the corrected sample set, the method further includes: The corrected sample set is structured and stored according to a preset standard format to generate a standard dataset; wherein, the standard dataset includes image files, label files corresponding to the image files, data partitioning information for model training, and dataset version description, and the data partitioning information includes training set information, validation set information, and test set information; The standard dataset is maintained based on a preset data version management mechanism so that it can be used to train photographic composition models.
8. A device for constructing a composition sample set for a photographic composition model, characterized in that, include: The system definition module is used to acquire a set of composition rules during photography, deconstruct the set of composition rules, and define a quantitative label system for the composition process from a preset dimension. The sample set generation module is used to design a structured target prompt word template based on the quantized label system, input the initial photographic image set into the visual language large model, output the composition annotation data corresponding to each image in the initial photographic image set according to the target prompt word template, and integrate the initial photographic image set and the corresponding composition annotation data to generate an initial composition sample set. The sample set acquisition module is used to determine the target composition regression model, use the target composition regression model to evaluate the target composition annotation error of each sample in the initial composition sample set, and select samples that meet the preset anomaly conditions from the initial composition sample set according to the target composition annotation error, as the filtered sample set. The sample set correction module is used to correct abnormal samples in the filtered sample set through the user terminal to obtain a corrected sample set. Based on the corrected sample set, the corresponding samples in the initial dataset are replaced to generate a target sample set. The target sample set is used to train a photographic composition model to obtain a trained composition model. The trained composition model is then used to respond to composition requests and output the corresponding photographic composition results.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the method for constructing a composition sample set of a photographic composition model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store computer programs; wherein, when the computer programs are executed by a processor, they implement the method for constructing a composition sample set of a photographic composition model as described in any one of claims 1 to 7.