Bird image vertical large model script generation method for multiple scenes
By periodically validating the validation set and optimizing the parameters, the problems of low reliability and efficiency in generating text in existing technologies have been solved, and efficient and accurate text generation for bird images has been achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIHOME TECHNOLOGY CO LTD
- Filing Date
- 2026-04-24
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies do not consider periodically correcting generation parameters based on a validation set, resulting in low reliability and efficiency in generating bird image captions.
By periodically verifying the quality of the generated text based on the validation set, correcting the generation parameters using the deviation and document viscosity characterization values, and optimizing the text generation by combining bird feature identifiability and overclocking probability, the accuracy and structural rationality of the generated text are ensured.
It improves the reliability and efficiency of generated text, reduces performance drift during long-term operation, and ensures that the generated text conforms to the reading path and accuracy of popular science communication.
Smart Images

Figure CN122113869A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bird science popularization technology, and in particular to a method for generating vertical large-scale model text for bird images in multiple scenes. Background Technology
[0002] The production of bird science popularization content is characterized by strong context, strong facts, and strong timeliness. Especially in scenarios such as nature education, birdwatching event organization, museum and nature reserve interpretation, urban biodiversity promotion, popular science books for teenagers, and short-form content on new media, there is often a need to quickly generate clear, factually accurate, and reader-friendly popular science text based on a single bird photograph or a set of bird photos. Such texts not only need to provide the bird species name but also need to supplement it with a concise introduction suitable for dissemination.
[0003] Targeted bird-related text needs to be generated for multiple scenarios. When users take bird photos, the goal is to generate integrated educational text for use in birdwatching notes, social media posts, and community outreach. Accompanying explanatory text should be generated for bird images taken by tourists or captured by monitoring equipment. Card-style text should be automatically generated based on educational photos or specimen images. Batch generation of educational explanations or supplementary text for reviewing large numbers of images is also required.
[0004] Chinese Patent Application Publication No. CN118673136B discloses a method, apparatus, electronic device, and storage medium for generating copy. The method includes: acquiring multiple images input by a user and / or an initial requirement description for the copy to be generated; generating copy based on a copy generation model, applying the multiple images and / or the initial requirement description to obtain a draft copy; acquiring a modification requirement description input by the user for the draft copy; modifying the draft copy based on the copy generation model, applying the modification requirement description, or applying the multiple images and the modification requirement description, to generate the target copy. It is evident that the above technical solution has the following problem: it does not consider periodically correcting the generation parameters based on a validation set, affecting the reliability of the generated copy and thus impacting the copy generation efficiency. Summary of the Invention
[0005] To address this issue, the present invention provides a method for generating text for vertical large-scale bird images in multiple scenes, thereby overcoming the problem in existing technologies that do not consider periodically correcting generation parameters based on a validation set, which affects the reliability of the generated text and consequently the efficiency of text generation.
[0006] To achieve the above objectives, this invention provides a method for generating vertical large-scale model text for bird images in multiple scenes, comprising: The acquired image to be analyzed is preprocessed to obtain a corrected image; Extract bird features and background features from the corrected image; The actual species name is determined based on the identified bird features and background characteristics; Based on the actual species names, several references were retrieved from the bird feature database included in the vertical large model. Identify valid sentences in each reference based on core bird knowledge phrases and sentence lengths in the references; Extract content from each valid statement to obtain several text groups; The vertical large model sorts the text groups to obtain the output copy; Periodically verify the validity of the generated output text based on the validation set, including: When an anomaly is identified in the output text generation, the output text generation parameters are corrected based on the deviation amount, including correcting the vertical large model’s identification criteria for effective sentences based on the document viscosity characterization value. Alternatively, if the generated output text is deemed satisfactory, continue generating output text for each item using the current generation parameters.
[0007] Furthermore, the process of periodically verifying the quality of the output text of the large vertical model based on the validation set includes: When the similarity representation value is less than or equal to the preset similarity representation value, the generation of the output copy is determined to be abnormal, and the generation parameters of the output copy are corrected based on the deviation. The validation set consists of several validation groups, and each validation group includes the image to be analyzed and the corresponding validation text. The output text is obtained from the image to be analyzed in a single validation group, and the similarity representation value is determined based on the output text and the validation text.
[0008] Furthermore, the process of determining similarity representation values based on the output copy and the validation copy includes: The output text is divided into a set of sentences to be checked, and the verification text is divided into a set of sentences to be checked. Each sentence set contains several statements. Based on the vector similarity between each statement in the set of sentences to be tested and each statement in the set of sentences to be verified, the test similarity value for each statement in the set of sentences to be tested and the verification similarity value for a single statement in the set of sentences to be verified are determined respectively. Semantic similarity is determined based on each test similarity value and each verification similarity value; Based on the similarity values to be tested, a set of matching pairs is determined with the order of the sentences to be tested as the benchmark. The word order consistency is determined based on the order of each sentence in the set of sentences to be tested and the order of each sentence in the set of matching pairs. A single matching pair consists of a single sentence to be validated and a corresponding validated sentence; The semantic similarity and word order consistency are assigned corresponding weight coefficients and summed to obtain the similarity representation value.
[0009] Furthermore, the parameters for generating the output copy are adjusted based on the deviation, including: The difference between the preset similarity characterization value and the similarity characterization value is calculated to obtain the deviation. When the deviation is less than or equal to the preset deviation, the vertical large model is corrected for the identification of valid sentences based on the literature viscosity characterization value; When the deviation exceeds the preset deviation, the process of dynamically correcting the determination of the actual species name based on the identifiability of bird features in each image to be analyzed is employed.
[0010] Furthermore, the process of determining the actual species name based on the identified bird features and background characteristics includes: Determining the season of an image based on background features; Bird feature vectors are determined based on bird characteristics in order to retrieve several candidate species from the bird feature database; The matching value for each candidate category is determined based on the season of the image; The comprehensive evaluation value is determined based on the feature vector similarity score and matching value corresponding to each candidate species, so as to determine the actual species name; The process of dynamically correcting the actual species name based on the bird feature identifiability of each image to be analyzed includes adjusting the weight coefficient corresponding to the matching value to the corresponding value based on the bird feature identifiability.
[0011] Furthermore, the process of determining the identifiability of bird features includes: The analytical quantification value is determined based on the total area of bird features and the sharpness of the image to be analyzed; The pose quantization value is determined based on the area of local head features and the total area of bird features. The bird feature recognition degree is obtained by summing the corresponding weight coefficients assigned to the analysis quantization value and the posture quantization value.
[0012] Furthermore, after adjusting the weight coefficients corresponding to the matching values, the determination of whether to strengthen the correction of the actual species name based on the overclocking probability of bird characteristics includes: When the overclocking probability is greater than the preset overclocking probability, the determination of the actual type name of the correction is enhanced.
[0013] Furthermore, the process of determining the overclocking probability includes: Perform a two-dimensional fast Fourier transform on each patch of bird features in the image to be analyzed; The overclocking probability is determined based on the energy in the frequency domain spectrum.
[0014] Furthermore, the determination of the actual type name for overclocking probability enhancement correction includes: Adjust the high-frequency gain during the sharpening process to the corresponding value.
[0015] Furthermore, the identification of effective sentences based on the vertical large model corrected by the document viscosity characterization value includes, Calculate the average repeatability of each reference to obtain the literature viscosity characterization value; The increase in the criteria for identifying effective sentences is positively correlated with the viscosity characterization value of the document.
[0016] Compared with existing technologies, the beneficial effects of this invention lie in its ability to break down the text into sentences by splitting the sentence set, transforming the comparison of the entire text into a comparison of the smallest semantic units. It determines the bidirectional maximum matching vector similarity and average value, finding the closest similarity among the verification sentences for the sentence to be tested and taking the maximum value: ensuring that each output sentence can find a semantic anchor in the standard text; otherwise, the similarity value of that sentence will be very low, indicating that the output sentence lacks basis or is off-topic. It reversely determines the verification sentence to the output sentence, preventing the output from only covering a part of the verification. The average similarity value of the sentence to be tested measures whether the output sentence as a whole matches the verification content. The average similarity value of the verification measures whether the information that should be verified is covered by the output. The semantic similarity value is a quantitative value that balances accuracy and coverage after bidirectional averaging. It determines the matching pair set and the word order deviation inversion pairs; popular science texts not only need sentence matching but also structural matching. If the word order is disordered, it will increase the reader's understanding cost and may even lead to a reversal of the explanatory logic. The number of inversion pairs measures how many pairs in the corresponding verification sentence index sequence of the output sentence are out of order, equivalent to how many sentences are misplaced. Semantic similarity quantifies whether something is correct, while word order deviation quantifies whether it flows smoothly. Weighted summation upgrades the evaluation from merely semantically correct to semantically correct and structurally sound, with weighting coefficients reflecting application-side preferences. This reduces the proportion of low-usability texts that are correct in content but disorganized in read, making the generated results more aligned with the reading path of science popularization. Periodic validation based on a validation set determines pass / fail status; the validation group binds the image to be analyzed with the validation text, which serves as a reference standard output for that image. When the similarity representation value is less than or equal to a preset similarity representation value, the deviation between the output and the standard answer has reached an unacceptable level, triggering correction. This forms a closed loop of automatic quality inspection and regression validation, reducing performance drift over long-term operation. Regularly correcting the generation parameters based on the validation set improves the reliability of the generated text, thereby increasing generation efficiency.
[0017] Furthermore, parameters are generated based on the deviation amount. The deviation amount quantifies the degree of deviation. When the deviation amount is less than or equal to the preset deviation amount, the output is close to the verification text, and the bird species identification and general direction are correct. This is identified as an anomaly in the text processing, and the effective sentence recognition standard is adjusted accordingly. When the deviation amount is greater than the preset deviation amount, the difference is large, indicating an image recognition anomaly. When the similarity representation value is far below the threshold, it means that there is a systematic error in the core facts. At this time, even if the text extraction is excellent, it will still be written around the wrong species, resulting in a serious overall deviation. In this case, the weight coefficient corresponding to the matching value used in determining the actual species name is adjusted to the corresponding value based on the bird feature identifiability. This achieves layered error correction for minor deviations in text and serious deviations in recognition. The bird feature identifiability is determined by two main factors: the number of pixels of the subject in the image (the larger the proportion, the more details); and image clarity (the clearer the image, the more reliable the texture and edges). The quantified value is multiplied by the two to estimate the effective pixel information. The proportion of bird features reflects whether the resolution resources are sufficient. Normalized sharpness reflects the usability of microstructures such as edges and feather patterns. Determining pose quantification values is crucial, as many key recognition points depend on pose deployment. If a bird is huddled, occluded, or positioned to the side or back, although its proportion may be significant, key features are not visible. Pose quantification values measure whether the current pose is close to a readable state. The area of local head features is used to determine the preset extended pose area, making the head more stable and easier to detect. The overall extended area is extrapolated from the head area, establishing an adaptive reference at different distance scales to avoid incomparability between distant and near views due to using fixed pixel thresholds. Bird feature identifiability measures the degree to which bird features are clearly identifiable. Based on bird feature identifiability, the weight coefficients corresponding to the matching values are adjusted to the corresponding values. When bird feature identifiability is low, the subject is small, blurry, occluded, or backlit, resulting in unstable visual similarity and easy misclassification of similar species. In such cases, it is necessary to supplement information using background season and behavioral priors, thus requiring a higher weight increase for the matching value. The lower the identifiability, the weaker the visual evidence, the stronger the prior constraints should be, and the greater the increase in matching weights. Improving the stability of species identification under low-quality image conditions reduces the overall misjudgment rate, increases the reliability of generated text, and thus improves the efficiency of text generation.
[0018] Furthermore, the decision to adjust denoising and sharpening is based on the over-frequency probability. Even with increased matching weights due to anomaly detection, artifacts generated during preprocessing may still affect feature extraction. Brightening and sharpening are often necessary in low-light conditions like dawn and dusk. Excessive denoising or sharpening can produce high-frequency artifacts such as false edges, smeared textures, and ringing. These artifacts microscopically alter feather patterns and edge morphology, shifting feature vectors and making similarity recognition less stable. After adjusting the recognition strategy, frequency domain artifacts are checked to determine whether to adjust the high-frequency sharpening gain. The over-frequency probability represents the proportion of abnormal high-frequency components in the image, used to quantify artifact risk. A two-dimensional fast Fourier transform maps texture and edge information to the frequency domain; both real details and artifacts are reflected in high-frequency energy changes. The bird's body is processed in blocks to capture local artifacts microscopically. The frequency domain suppression high-frequency filtering frequency is determined based on the family and genus, as different families and genera have different feather scales and edge sharpness. The overclocking probability is defined as the proportion of components in the frequency domain spectrum that are higher than the suppression frequency. A higher proportion of these components indicates more pixels exhibiting abnormal high-frequency energy, corresponding to over-sharpening, ringing, and residual noise. This makes artifact detection more closely match the differences in bird textures, reduces mistuning, and improves robustness in low-light scenes at dawn and dusk. The reduction in high-frequency gain is positively correlated with the overclocking probability; a higher overclocking probability indicates more prevalent or severe artifacts. In such cases, a stronger reduction in high-frequency sharpening gain is needed to bring abnormal high frequencies back to the normal range. Suppressing artifacts while preserving as much detail as possible improves the stability of subsequent feature extraction and similarity calculation, reducing the probability of false features created by sharpening. This improves the reliability of generated text, thereby increasing text generation efficiency. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the steps of a vertical large-model text generation method for bird images in multiple scenes according to an embodiment of the present invention. Figure 2 This is a logic diagram for verifying the validity of the output text generated by a vertical large model based on a validation set, according to an embodiment of the present invention. Figure 3 This is a logic diagram of the generation parameters for the output text based on deviation correction in an embodiment of the present invention; Figure 4 This is a logic diagram for determining whether to enhance the correction of the actual species name based on the overclocking probability of bird characteristics in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0021] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0022] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0023] Furthermore, it should be noted that, in the description of this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0024] Please see Figure 1 The diagram shows the steps of a method for generating text for a vertical large model of bird images in multiple scenes according to an embodiment of the present invention. The method includes: S1, preprocess the acquired image to be analyzed to obtain a corrected image; S2, extract bird features and background features from the corrected image; S3, determine the actual species name based on the identified bird features and background features; S4, based on the actual species names, retrieve several references from the bird feature database included in the vertical large model; S5 identifies valid sentences in each reference based on core bird knowledge phrases and sentence lengths in the references. S6: Extract content from each valid statement to obtain several text groups; S7, the vertical large model sorts each text group to obtain the output copy; S8, periodically checks the validity of the generated output text based on the validation set, including: When an anomaly is identified in the output text generation, the output text generation parameters are corrected based on the deviation amount, including correcting the vertical large model’s identification criteria for effective sentences based on the document viscosity characterization value. Alternatively, if the generated output text is deemed satisfactory, continue generating output text for each item using the current generation parameters.
[0025] After acquiring several text sets, each text set is input into a self-developed vertical model for bird science popularization for copywriting generation and optimization. This vertical model is based on a general pre-trained language model and is constructed using professional corpus in the field of ornithology for incremental pre-training and instruction fine-tuning. The domain corpus includes "The Birds of China", journals of the Chinese Ornithological Society, monitoring reports of nature reserves, and authoritative popular science publications, accounting for more than 65% of the total training data.
[0026] The vertical large-scale model employs LoRA low-rank adaptation technology for efficient parameter fine-tuning, enhancing the understanding and generation capabilities of professional knowledge such as bird morphological description, habitat analysis, and behavioral interpretation while maintaining the general capabilities of the base model. A knowledge-enhanced attention mechanism is introduced into the model architecture, embedding structured knowledge from the bird feature database—including species classification levels, geographical distribution information, and seasonal migration patterns—into the attention calculation process in the form of a knowledge graph to ensure the factual accuracy of the generated content.
[0027] During the generation process, the vertical large model performs semantic fusion and logical reorganization on each text group, eliminating information conflicts between multi-source documents, supplementing implicit professional related knowledge, and organizing the content structure according to ornithological narrative norms, morphology, distribution, habits, and conservation, outputting coherent, accurate, and context-appropriate candidate texts. After the candidate texts undergo terminology consistency verification and readability determination by the post-processing module, they enter the sorting and optimization process in step S7.
[0028] Specifically, the process of preprocessing the acquired image to be analyzed to obtain a corrected image may include applying Gaussian filtering for noise reduction and unifying the image size to 224x224 pixels by scaling and padding while maintaining the original aspect ratio, in order to eliminate the impact of differences in shooting environment on recognition.
[0029] Specifically, in S2, the process of extracting bird features and background features from the corrected image may include inputting the corrected image into a pre-trained EfficientNetV2B0 convolutional neural network and extracting the high-dimensional feature map output by the last convolutional layer of the network. Using a pre-trained semantic segmentation head, feature regions belonging to the foreground birds and feature regions belonging to the background environment are decoded from this feature map, thereby obtaining bird feature vectors and background feature vectors.
[0030] In S3, the process of determining the actual species name based on the identified bird features and background features includes: The season of an image is determined by the vegetation color saturation and hue based on background features. If the proportion of green in the background features is >60%, it is determined to be summer; if the total proportion of yellow and red in the background features is >40%, it is determined to be autumn; and the rest are determined to be winter or spring. The extracted bird feature vectors are compared with the cosine similarity of the global average feature vectors corresponding to each bird species in the pre-stored bird feature database to determine the similarity score, and several candidate species are determined based on the similarity score. The matching value of each candidate species is calculated based on the season of the image; the bird feature database pre-stores the occurrence probability of each bird species in each season, and the occurrence probability of a single candidate species in the corresponding image season is the matching value between the single candidate species and the image season. The similarity score and matching value are assigned weight coefficients and summed to determine the comprehensive evaluation value. The candidate category corresponding to the maximum value is determined as the actual category name. The similarity score weight is 0.6, the matching value weight is 0.4, and the sum of the similarity score weight and the matching value weight is 1.
[0031] In a single embodiment, the process of retrieving several references from the bird feature database included in the vertical large model based on actual species names may include: searching the literature database within the vertical large model using the actual species name as a keyword; and quantifying the relevance of each document to the query topic using a weighted summation method based on the matching frequency of the keyword in the document title, abstract, and keyword fields. For a specific actual species name, such as the Yellow-bellied Tit, the following steps are performed for each candidate document during the literature retrieval: The number of times a keyword appears in the document title is recorded as C_title; The number of times a keyword appears in the document's keyword field is recorded, denoted as C_keywords; The number of times the keyword appears in the document abstract is counted and denoted as C_abstract.
[0032] Considering the varying importance of different fields, the title is weighted at w_title=3, keywords at w_keywords=2, and the abstract at w_abstract=1. The corresponding document's similarity score R_doc is calculated as follows: R_doc = w_title C_title + w_keywords C_keywords + w_abstract C_abstract.
[0033] If the word "yellow-bellied tit" appears once in the title of a single document, twice in the keyword field, and five times in the abstract, then its repetition rate R_doc = 3×1 + 2×2 + 1×5 = 12.
[0034] After selecting several references with the highest repetition rate, the arithmetic mean of the repetition rates of these references is calculated, which is the literature viscosity characterization value. The literature viscosity characterization value reflects the richness and concentration of information about the bird species in the current knowledge base. The larger the value, the more abundant the available high-quality literature resources, so that the selection criteria for effective sentences can be improved accordingly in subsequent steps.
[0035] Specifically, based on the core bird knowledge phrases and sentence lengths in the references, valid sentences are extracted from each reference. The content of each valid sentence is extracted to obtain several text groups. A PDF parser and a substitution resolution algorithm can be used to extract candidate sentences with a length of more than 10 words that are logically complete and include core bird knowledge phrases from the documents. Valid sentences are then determined from the candidate sentences. Using the automatic extraction technology of key text information, the valid sentences are broken down and summarized into text groups with different themes such as morphological characteristics, habitat, and living habits.
[0036] Specifically, the vertical large model pre-stores core ornithological knowledge phrases, which can include morphology, habits, and distribution. The Shannon information entropy of each candidate statement is calculated. If the Shannon information entropy is greater than the recognition standard, indicating a large amount of information, it is determined to be a valid statement and retained; otherwise, it is filtered out as invalid information.
[0037] Please see Figure 2 As shown, this is a logic diagram illustrating the validation of the generated output text of a vertical large-scale model based on a validation set, according to an embodiment of the present invention. The present invention includes a process for periodically validating the generated output text of a vertical large-scale model based on a validation set, comprising: The verification set includes several verification groups, and each verification group includes the image to be analyzed and the verification text corresponding to the image to be analyzed. The output text is obtained from the image to be analyzed in a single validation group, and the similarity representation value is determined based on the output text and the validation text. When the similarity representation value is less than or equal to the preset similarity representation value, the generation of the output copy is determined to be abnormal, and the generation parameters of the output copy are corrected based on the deviation. When the similarity value is greater than the preset similarity value, the output copy is deemed qualified, and the generation of each output copy continues to be completed using the current generation parameters.
[0038] The preset similarity representation value is selected within the range [0.75, 0.85]. Those skilled in the art can select and determine it according to the actual application requirements. In this embodiment, 0.85 is preferred.
[0039] Specifically, a validation set can be constructed by selecting 1,000 clearly labeled bird images and their corresponding expert-written popular science texts from authoritative publications such as the Field Guide to the Birds of China and the official popular science platform of the Ornithological Society of China.
[0040] Specifically, the process of determining similarity values based on the output copy and the validation copy includes: The output text is divided into a set of sentences to be checked and the verification text is divided into a set of sentences to be checked. Each set of sentences contains several statements. Calculate the vector similarity between a single statement in the set of sentences to be tested and each statement in the set of sentences to be verified, and take the maximum value to obtain the test similarity value for a single statement in the set of sentences to be tested; Calculate the average of each similarity value to be tested to obtain the average similarity value to be tested. Calculate the vector similarity between a single statement in the verification sentence set and each statement in the sentence set to be verified, and take the maximum value to obtain the verification similarity value for a single statement in the verification sentence set; Calculate the average of each verification similarity value to obtain the verification average similarity value; The semantic similarity value is obtained by averaging the average similarity value of the verification and the average similarity value of the test. The semantic similarity score is obtained by calculating the ratio of the semantic similarity score to the preset semantic similarity score. Based on the similarity values to be tested, a set of matching pairs is determined to obtain the verification sentences corresponding to each sentence in the set of sentences to be tested; The statements in the check statement set are numbered according to their order. The check statement indices that best match each statement in the sentence set to be checked are arranged according to their order, resulting in an index sequence. The number of inversion pairs in this index sequence is determined, and the theoretical maximum number of inversion pairs (n) of this index sequence is calculated. The ratio of (n-1) / 2, where n is the total number of sentences, yields the normalized word order bias.
[0041] Determine word order consistency as 1 and reduce normalized word order bias; The semantic similarity and word order consistency are assigned corresponding weight coefficients and summed to obtain the similarity representation value.
[0042] Specifically, the weighting coefficients for semantic similarity and word order deviation can be determined based on the actual usage scenario. For example, for science popularization for children, structural order is more important, while for scientific research records, semantic coverage is more important. The sum of the weighting coefficients for semantic similarity and word order is 1. In this embodiment, it is preferable that the weighting coefficients for both semantic similarity and word order are 0.5.
[0043] Specifically, to calculate the vector similarity between a single statement in the set of sentences to be tested and each statement in the set of sentences to be verified, the paraphrase-multilingual-MiniLM-L12-v2 model can be used to convert the statements into 384-dimensional semantic vectors, and the cosine similarity between the two vectors can be calculated as the vector similarity.
[0044] Specifically, based on the similarity values to be tested, a greedy matching algorithm is used to determine the set of matching pairs: the sentence pairs with the highest similarity are matched first, and the matched items are removed until all sentences to be tested find a unique corresponding verification sentence, thereby ensuring a one-to-one correspondence.
[0045] In a single embodiment, the index sequence of the set of sentences to be checked is [3, 1, 2, 4], which is arranged according to the order of the sentences in the set of sentences to be checked. The index sequence [3, 1, 2, 4] means that the output of sentence 1 corresponds to the check of sentence 3, the output of sentence 2 corresponds to the check of sentence 1, and the more disordered the sequence is, the more inversion pairs there are.
[0046] Specifically, by splitting the sentence set, the copy is broken down into sentences, transforming the entire comparison into a comparison of the smallest semantic units. The maximum bidirectional matching vector similarity and average value are determined. The sentence to be tested is found to be the closest to the verification sentence, and the maximum value is taken: ensuring that each output sentence can find a semantic anchor in the standard copy; otherwise, the similarity value of that sentence will be very low, indicating that the output sentence lacks basis or is off-topic. The verification sentence is reversed to determine the output sentence, preventing the output from only covering a part of the verification. The average similarity value of the sentence to be tested measures whether the output sentence as a whole matches the verification content. The average similarity value of the verification measures whether the information that should be verified is covered by the output. The semantic similarity value is a quantitative value that balances accuracy and coverage after bidirectional averaging. The matching pair set and word order deviation inversion pairs are determined. Popular science copy not only needs sentence correctness but also structural correctness. If the word order is disordered, it will increase the reader's understanding cost and may even lead to a reversal of the explanatory logic. The number of inversion pairs measures how many pairs in the corresponding verification sentence index sequence of the output sentence are out of order, equivalent to how many sentences are misplaced. Semantic similarity quantifies whether something is correct, while word order deviation quantifies whether it flows smoothly. Weighted summation upgrades the evaluation from merely semantically correct to semantically correct and structurally sound, with weighting coefficients reflecting application-side preferences. This reduces the proportion of low-usability texts that are correct in content but disorganized in read, making the generated results more aligned with the reading path of science popularization. Periodic validation based on a validation set determines pass / fail status; the validation group binds the image to be analyzed with the validation text, which serves as a reference standard output for that image. When the similarity representation value is less than or equal to a preset similarity representation value, the deviation between the output and the standard answer has reached an unacceptable level, triggering correction. This forms a closed loop of automatic quality inspection and regression validation, reducing performance drift over long-term operation. Regularly correcting the generation parameters based on the validation set improves the reliability of the generated text, thereby increasing generation efficiency.
[0047] Please see Figure 3 As shown, this is a logic decision diagram of the generation parameters for the output text based on deviation correction in an embodiment of the present invention. The generation parameters for the output text based on deviation correction in the present invention include: The difference between the preset similarity characterization value and the similarity characterization value is calculated to obtain the deviation. When the deviation is less than or equal to the preset deviation, the vertical large model is corrected for the identification of valid sentences based on the literature viscosity characterization value; When the deviation exceeds the preset deviation, the weight coefficient corresponding to the matching value used in determining the actual species name will be adjusted to the corresponding value based on the bird feature identifiability.
[0048] Specifically, the preset deviation amount can be determined by analyzing the deviation amount distribution of the two types of errors in historical error cases. The preset deviation amount serves as the threshold for distinguishing between small-deviation textual logic errors and large-deviation image recognition errors.
[0049] Specifically, the process of adjusting the weight coefficients corresponding to the matching values based on the identifiability of bird features includes: The increase in the weight coefficient corresponding to the matching value is negatively correlated with the identifiability of bird features.
[0050] In this embodiment, optionally, Compare the bird feature identifiability with the preset identification comparison value; determine the weight coefficient corresponding to the matching value as the matching weight coefficient; When the identifiability of bird features is less than or equal to the preset identification comparison value, the matching weight coefficient is adjusted to 1.2 times the initial matching weight coefficient; When the identifiability of bird features is greater than the preset identification comparison value, the matching weight coefficient is adjusted to 1.1 times the initial matching weight coefficient; The preset recognition comparison value is 0.6.
[0051] Specifically, the process of determining the identifiability of bird features includes: The ratio of the area of bird features to the total area of the image to be analyzed is used to obtain the proportion of bird features. Calculate the ratio of the sharpness of the image to be analyzed to the preset sharpness to obtain the normalized sharpness; The product of the proportion of bird features and the normalized clarity is calculated to obtain the analytical quantification value; The area of the pre-defined extended posture is determined based on the area of local head features in bird characteristics. The ratio of the area of bird features to the area of a preset extended posture is calculated to obtain the posture quantification value; The bird feature recognition degree is obtained by summing the corresponding weight coefficients assigned to the analysis quantization value and the posture quantization value.
[0052] In a single embodiment, the preset sharpness is set to 0.5, which is calculated based on the Brenner gradient function, i.e., the average value of the gray-level difference between adjacent pixels in the diagonal direction of the image that exceeds a preset critical threshold.
[0053] The weighting coefficient for the analytical quantization value was set to 0.6, and the weighting coefficient for the pose quantization value was set to 0.4, in order to reflect the prior knowledge that subject sharpness has a greater impact on the recognition result than pose integrity.
[0054] Determining the preset stretched posture area based on the area of local head features in bird characteristics can include multiplying the detected head area A_head by an empirical coefficient of 12.5 to estimate the preset stretched posture area A_stretched = 12.5. A_head. The empirical coefficient can be determined based on the statistical mean of the body-to-head ratio of common birds such as grey herons and white wagtails.
[0055] Specifically, parameters are generated based on deviation. The deviation is used to quantify the degree of deviation. When the deviation is less than or equal to a preset deviation, the output is close to the verification text, indicating that the bird species identification and general direction are correct. This is identified as an anomaly in the text processing, and the effective sentence recognition standard is adjusted accordingly. When the deviation is greater than the preset deviation, the difference is large, indicating an image recognition anomaly. When the similarity representation value is far below the threshold, it means that there is a systematic error in the core facts. In this case, even if the text extraction is excellent, it will still be written around the wrong species, leading to a serious overall deviation. At this time, the weight coefficient corresponding to the matching value used in determining the actual species name is adjusted to the corresponding value based on the bird feature identifiability. This achieves layered error correction for minor deviations in text and severe deviations in recognition. The bird feature identifiability is determined by two main factors: the number of pixels of the subject in the image (the larger the proportion, the more detail); and image clarity (the clearer the image, the more reliable the texture and edges). The quantified value is multiplied by these two factors to estimate the effective pixel information. The proportion of bird features reflects whether the resolution resources are sufficient. Normalized sharpness reflects the usability of microstructures such as edges and feather patterns. Determining pose quantification values is crucial, as many key recognition points depend on pose deployment. If a bird is huddled, occluded, or positioned to the side or back, although its proportion may be significant, key features are not visible. Pose quantification values measure whether the current pose is close to a readable state. The area of local head features is used to determine the preset extended pose area, making the head more stable and easier to detect. The overall extended area is extrapolated from the head area, establishing an adaptive reference at different distance scales to avoid incomparability between distant and near views due to using fixed pixel thresholds. Bird feature identifiability measures the degree to which bird features are clearly identifiable. Based on bird feature identifiability, the weight coefficients corresponding to the matching values are adjusted to the corresponding values. When bird feature identifiability is low, the subject is small, blurry, occluded, or backlit, resulting in unstable visual similarity and easy misclassification of similar species. In such cases, it is necessary to supplement information using background season and behavioral priors, thus requiring a higher weight increase for the matching value. The lower the identifiability, the weaker the visual evidence, the stronger the prior constraints should be, and the greater the increase in matching weights. Improving the stability of species identification under low-quality image conditions reduces the overall misjudgment rate, increases the reliability of generated text, and thus improves the efficiency of text generation.
[0056] Please see Figure 4 As shown, this is a logic diagram for determining whether to strengthen the correction of the actual species name based on the overclocking probability of bird characteristics in an embodiment of the present invention. After adjusting the weight coefficients corresponding to the matching value, the determination of whether to strengthen the correction of the actual species name based on the overclocking probability of bird characteristics includes: When the overclocking probability is less than or equal to the preset overclocking probability, the current preprocessing parameters are continuously used to process each image to be analyzed. When the overclocking probability is greater than the preset overclocking probability, the determination of the actual type name of the correction is enhanced.
[0057] The preset overclocking probability is selected within the range [0.1, 0.15]. A preferred value of 0.15 allows for the selection of a representative set of bird images taken under different lighting conditions. Different levels of sharpening algorithms are applied to these images, generating a series of images ranging from insufficiently sharpened to moderately sharpened to overly sharpened. These images are then subjectively scored by image processing experts or bird photographers. Scoring criteria include 1- no visible artifacts, 2- slight artifacts that do not affect interpretation, and 3- obvious artifacts interfering with recognition. When more than 70% of the scorers consider an image to have obvious artifacts interfering with recognition, the overclocking probability value of that image is recorded. The overclocking probability values of all images marked as having obvious artifacts are statistically analyzed, and the 25th percentile is used as the preset overclocking probability threshold.
[0058] Specifically, the process of determining the over-sharpening probability includes: dividing the bird features into several 32x32 pixel patches, performing a two-dimensional fast Fourier transform on each patch to convert the spatial domain image into a frequency domain spectrum; counting the number of patches in the frequency domain spectrum with energy greater than the frequency domain suppression high-frequency filter frequency, calculating the ratio of this ratio to the total number of patches, and determining this ratio as the over-sharpening probability to quantify the risk of artifacts caused by over-sharpening.
[0059] The process of determining the reference bird family and genus based on the image to be analyzed, and determining the frequency domain suppression high-frequency filter frequency based on the reference bird family and genus, includes classifying the image to the family level using a lightweight MobileNetV3 model. In a single embodiment, the radial power spectral density is statistically analyzed based on artifact-free sample images of the reference bird family and genus. The frequency is normalized to the Nyquist frequency of 1 (dimensionless), and the average value of the frequency f95 corresponding to 95% cumulative energy is taken as the benchmark. A safety margin of 0.03 is added to obtain the cutoff frequency fc, which is limited to 0.15 to 0.45. For passerines with finer textures, fc is set to 0.30; for hawksbill animals with smoother textures, fc is set to 0.20.
[0060] The determination of the actual types of overclocking probability enhancement correction names includes... Adjust the high-frequency gain during the sharpening process to the corresponding value.
[0061] Specifically, based on the overclocking probability, the high-frequency gain during the sharpening process is adjusted to a corresponding value, where... The reduction in high-frequency gain is positively correlated with the probability of overclocking.
[0062] In this embodiment, optionally, Compare the overclocking probability with the preset overclocking probability comparison value; When the overclocking probability is less than or equal to the preset overclocking probability comparison value, the high-frequency gain during the sharpening process is adjusted to 0.94 times the initial high-frequency gain. When the overclocking probability is greater than the preset overclocking probability comparison value, the high-frequency gain during the sharpening process will be adjusted to 0.86 times the initial high-frequency gain. The preset overclocking probability comparison value can be determined through subjective user experiments. When the overclocking probability exceeds the preset overclocking probability comparison value, 90% of testers will consider that the image contains visible artifacts. In this embodiment, the preset overclocking probability comparison value is preferably set to 0.2.
[0063] Specifically, the decision to adjust denoising and sharpening is based on the over-frequency probability. Even with increased matching weights due to anomaly detection, artifacts generated during preprocessing may still affect feature extraction. Brightening and sharpening are often necessary in low-light conditions like dawn and dusk. Excessive denoising or sharpening can produce high-frequency artifacts such as false edges, smeared textures, and ringing. These artifacts microscopically alter feather patterns and edge morphology, shifting feature vectors and making similarity recognition less stable. After adjusting the recognition strategy, frequency domain artifacts are checked to determine whether to adjust the high-frequency sharpening gain. The over-frequency probability is the proportion of abnormal high-frequency components in the image, used to quantify artifact risk. A two-dimensional fast Fourier transform maps texture and edge information to the frequency domain; both real details and artifacts are reflected in high-frequency energy changes. The bird's body is processed in blocks to capture local artifacts microscopically. The frequency domain suppression high-frequency filtering frequency is determined based on the family and genus, as different families and genera have different feather scales and edge sharpness. The overclocking probability is defined as the proportion of components in the frequency domain spectrum that are higher than the suppression frequency. A higher proportion of these components indicates more pixels exhibiting abnormal high-frequency energy, corresponding to over-sharpening, ringing, and residual noise. This makes artifact detection more closely match the differences in bird textures, reduces mistuning, and improves robustness in low-light scenes at dawn and dusk. The reduction in high-frequency gain is positively correlated with the overclocking probability; a higher overclocking probability indicates more prevalent or severe artifacts. In such cases, a stronger reduction in high-frequency sharpening gain is needed to bring abnormal high frequencies back to the normal range. Suppressing artifacts while preserving as much detail as possible improves the stability of subsequent feature extraction and similarity calculation, reducing the probability of false features created by sharpening. This improves the reliability of generated text, thereby increasing text generation efficiency.
[0064] Specifically, the identification of effective sentences based on the vertical large model corrected by the viscosity characterization value of the literature includes, Adjust the criteria for recognizing valid statements to the corresponding values; Calculate the average repeatability of each reference to obtain the literature viscosity characterization value; The increase in the criteria for identifying effective sentences is positively correlated with the viscosity characterization value of the document.
[0065] In this embodiment, optionally, T_new = T_base (1 + 0.2 The formula (V) represents the document viscosity characterization value, where T_new is the adjusted identification standard, T_base is the initial identification standard, and V is the document viscosity characterization value. This formula indicates that the increase in the identification standard for effective sentences is positively correlated with the document viscosity characterization value; that is, the richer the document resources, the higher the requirements for sentence quality.
[0066] Specifically, after adjusting the recognition criteria for valid sentences, the system re-determines whether the generated output text is qualified based on the validation set. If the generated output text is still determined to be abnormal, the weight coefficient corresponding to the matching value is adjusted to the corresponding value based on the identifiability of bird features.
[0067] Based on the document viscosity characterization value, the identification standard of effective sentences is adjusted to the corresponding value. By dynamically matching the status of knowledge base resources and the strictness of screening, the standard is increased when document resources are abundant to reduce the processing cost of low-quality documents, and the standard is decreased when document resources are scarce to ensure the success rate of information acquisition, thereby optimizing the document retrieval efficiency.
[0068] This invention replaces manual review with an automated verification-feedback-correction closed loop, achieving continuous optimization of generated quality without service interruption. Dynamic parameter adjustment avoids resource waste or information omissions caused by fixed standards. Computational resources are dynamically allocated based on image quality, avoiding over-reliance on unreliable visual features on low-quality images and reducing duplicate searches due to misidentification. Preventative suppression of artifacts caused by over-sharpening avoids cascading efficiency losses due to artifact-induced feature extraction errors, incorrect species identification, and invalid document retrieval.
[0069] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0070] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating text for large vertical models of bird images across multiple scenes, characterized in that, include: The acquired image to be analyzed is preprocessed to obtain a corrected image; Extract bird features and background features from the corrected image; The actual species name is determined based on the identified bird features and background characteristics; Based on the actual species names, several references were retrieved from the bird feature database included in the vertical large model. Identify valid sentences in each reference based on core bird knowledge phrases and sentence lengths in the references; Extract content from each valid statement to obtain several text groups; The vertical large model sorts the text groups to obtain the output copy; Periodically verify the validity of the generated output text based on the validation set, including: When an anomaly is identified in the output text generation, the output text generation parameters are corrected based on the deviation amount, including correcting the vertical large model’s identification criteria for effective sentences based on the document viscosity characterization value. Alternatively, if the generated output text is deemed satisfactory, continue generating output text for each item using the current generation parameters.
2. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 1, characterized in that, The process of periodically verifying the quality of the output text of a large vertical model based on a validation set. include, When the similarity representation value is less than or equal to the preset similarity representation value, the generation of the output copy is determined to be abnormal, and the generation parameters of the output copy are corrected based on the deviation. The validation set consists of several validation groups, and each validation group includes the image to be analyzed and the corresponding validation text. The output text is obtained from the image to be analyzed in a single validation group, and the similarity representation value is determined based on the output text and the validation text.
3. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 2, characterized in that, The process of determining similarity metrics based on the output copy and the validation copy includes: The output text is divided into a set of sentences to be checked, and the verification text is divided into a set of sentences to be checked. Each sentence set contains several statements. Based on the vector similarity between each statement in the set of sentences to be tested and each statement in the set of sentences to be verified, the test similarity value for each statement in the set of sentences to be tested and the verification similarity value for a single statement in the set of sentences to be verified are determined respectively. Semantic similarity is determined based on each test similarity value and each verification similarity value; Based on the similarity values to be tested, a set of matching pairs is determined with the order of the sentences to be tested as the benchmark. The word order consistency is determined based on the order of each sentence in the set of sentences to be tested and the order of each sentence in the set of matching pairs. A single matching pair consists of a single sentence to be validated and a corresponding validated sentence; The semantic similarity and word order consistency are assigned corresponding weight coefficients and summed to obtain the similarity representation value.
4. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 3, characterized in that, The parameters for generating the output copy based on the deviation correction include: The difference between the preset similarity characterization value and the similarity characterization value is calculated to obtain the deviation. When the deviation is less than or equal to the preset deviation, the vertical large model is corrected for the identification of valid sentences based on the literature viscosity characterization value; When the deviation exceeds the preset deviation, the process of dynamically correcting the determination of the actual species name based on the identifiability of bird features in each image to be analyzed is employed.
5. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 4, characterized in that, The process of determining the actual species name based on identified bird features and background characteristics includes: Determining the season of an image based on background features; Bird feature vectors are determined based on bird characteristics in order to retrieve several candidate species from the bird feature database; The matching value for each candidate category is determined based on the season of the image; The comprehensive evaluation value is determined based on the feature vector similarity score and matching value corresponding to each candidate species, so as to determine the actual species name; The process of dynamically correcting the actual species name based on the bird feature identifiability of each image to be analyzed includes adjusting the weight coefficient corresponding to the matching value to the corresponding value based on the bird feature identifiability.
6. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 5, characterized in that, The process of determining the identifiability of bird features includes: The analytical quantification value is determined based on the total area of bird features and the sharpness of the image to be analyzed; The pose quantization value is determined based on the area of local head features and the total area of bird features. The bird feature recognition degree is obtained by summing the corresponding weight coefficients assigned to the analysis quantization value and the posture quantization value.
7. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 6, characterized in that, After adjusting the weight coefficients corresponding to the matching values, the determination of whether to strengthen the correction based on the overclocking probability of bird characteristics and the determination of the actual species name include: When the overclocking probability is greater than the preset overclocking probability, the determination of the actual type name of the correction is enhanced.
8. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 7, characterized in that, The process of determining the probability of overclocking includes: Perform a two-dimensional fast Fourier transform on each patch of bird features in the image to be analyzed; The overclocking probability is determined based on the energy in the frequency domain spectrum.
9. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 8, characterized in that, The determination of the actual types of overclocking probability enhancement correction includes: Adjust the high-frequency gain during the sharpening process to the corresponding value.
10. The method for generating vertical large-scale model text for bird images in multiple scenes according to claim 9, characterized in that, The identification of effective statements based on the vertical large model corrected by the viscosity characterization value of the literature includes, Calculate the average repeatability of each reference to obtain the literature viscosity characterization value; The increase in the criteria for identifying effective sentences is positively correlated with the viscosity characterization value of the document.