Simulation defect identification method and system based on dynamically generated material and storage medium

By performing edge detection and feature region segmentation on the generated materials, calculating pixel contrast and feature position errors, and automatically identifying and correcting simulation defects in the generated materials, the problem of visual differences in AI-generated video content is solved, improving the visual realism and practical value of the materials.

CN122223345APending Publication Date: 2026-06-16QILINGSHI CULTURE MEDIA (SHANGHAI) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QILINGSHI CULTURE MEDIA (SHANGHAI) CO LTD
Filing Date
2026-04-09
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Existing video content generated based on artificial intelligence algorithms shows significant visual differences from real video content in non-semantic representations, lacking a deep understanding of the physical laws and detailed structures of real scenes, resulting in simulation defects in the generated videos.

Method used

A simulation defect identification method based on dynamically generated materials is adopted. By acquiring generated content data and prompt data, target semantics and display features, edge features are extracted. Edge detection and adaptive methods are used to perform edge detection and feature region segmentation, calculate pixel contrast degree and feature position error, and realize automated identification and correction.

Benefits of technology

It achieves efficient and accurate identification of simulation defects such as blurred edges and disordered structural positions in generated materials, improving the visual realism and practical value of the materials and meeting the application needs of high-realism scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122223345A_ABST
    Figure CN122223345A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence and discloses a simulation defect identification method and system based on dynamically generated materials and a storage medium. The method comprises the following steps: acquiring generated content data and corresponding prompt data, extracting target semantic features and display features, dividing display feature regions and extracting edge regions, calculating edge contrast degree values and adaptively performing fuzzy processing, matching reference positions in a feature position library, calculating error matching degrees of target actual positions and the reference positions, and generating a defect early warning prompt if the error matching degrees do not meet the requirements. The application realizes automatic and accurate identification of edge blur, position disorder and other defects, does not require manual frame-by-frame checking, greatly improves detection efficiency, effectively improves the visual authenticity of generated materials, and meets the needs of high-fidelity scene requirements such as short drama production and e-commerce advertising.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a simulation defect identification method, system and storage medium based on dynamically generated materials. Background Technology

[0002] With the iterative evolution of artificial intelligence technology, generative AI has achieved leapfrog development in the field of dynamic material production and has been gradually applied to multiple scenarios such as entertainment, education, and commerce. Technological breakthroughs centered on diffusion models and the DiffusionTransformer architecture have propelled dynamic material generation from proof of concept to large-scale commercial use, significantly lowering the barrier to video production, shortening the production cycle, and enabling efficient generation of dynamic materials.

[0003] Currently, video content generated based on artificial intelligence algorithms is showing a diversified development trend, with generation methods covering various types such as text-to-video, image-to-video, and video-to-video. Application scenarios include short drama production, science popularization courseware, e-commerce advertising, and personalized greeting videos. Various generation models can support video generation at different resolutions and durations, and can flexibly adjust content style, scene settings, and dynamic effects according to user needs.

[0004] Existing video content is generated based on artificial intelligence algorithms. Its core advantage lies in its ability to accurately match the semantic needs of user input, and its overall semantic performance can meet the basic requirements of daily application scenarios. However, since the generation model is essentially based on statistical fitting to generate content, it lacks a deep understanding of the physical laws and detailed structures of real scenes, resulting in significant visual differences between the non-semantic aspects of the generated video and the real video content. Summary of the Invention

[0005] To reduce the visual difference between the non-semantic representation of generated data and real video content, this application provides a simulation defect identification method, system, and storage medium based on dynamically generated materials.

[0006] Firstly, this application provides a simulation defect identification method based on dynamically generated materials, employing the following technical solution: A simulation defect identification method based on dynamically generated materials includes the following steps: Acquire generated content data and corresponding generated prompt data; extract target semantic features from generated prompt data based on a preset word meaning extraction algorithm; and extract target display features from generated content data based on the target semantic features and a preset content extraction algorithm. Based on the target display features, the generated content data is divided into multiple regions to obtain the display feature regions. The edges of the target display features are extracted based on the edge extraction algorithm, and the edge regions are selected within a preset length range from the edges. Calculate the contrast value of the pixel content in the edge region. The clearer the pixel content, the higher the contrast value. The more blurred the pixel content, the lower the contrast value. If the contrast value is greater than the preset contrast reference value, the pixel content in the edge region is blurred according to the preset blur parameter. The larger the blur parameter, the higher the degree of blur processing. The smaller the blur parameter, the lower the degree of blur processing. Obtain the feature location library corresponding to the target feature, and match the corresponding target reference position from the feature location library according to the target feature. The target reference position is the relative distance between at least two target features. The actual position of the target is generated by calculating the relative positions between target features in the generated content data, and the error matching degree between the target reference position and the actual position of the target is calculated. If the error matching degree is less than the preset error reference value, a defect warning prompt corresponding to the target feature will be generated.

[0007] By adopting the above technical solution, through automatic extraction of semantic and display features, partition detection of edge clarity and quantitative comparison of feature relative position deviation, simulation defects such as blurred edges and disordered structural positions in generated dynamic materials can be automatically and accurately identified without manual frame-by-frame verification, which greatly improves defect detection efficiency, effectively improves the visual realism and simulation effect of generated materials, and meets the material application needs of high realism scenes.

[0008] Optionally, the step of dividing the generated content data into multiple regions based on the target display features to obtain the display feature regions further includes the following sub-steps: If there are multiple target display features, the interference region between adjacent target display features is identified, and the area of ​​the interference region is calculated as the interference area; the area of ​​the target display feature with the smallest area is taken as the base area; the ratio of the interference area to the base area is calculated as the area ratio. If the area ratio is less than the preset area reference ratio, the target display feature and the surrounding area within the preset setting range will be used as the corresponding display feature area of ​​the target display feature; otherwise, the area of ​​the target display feature that is not obscured and the surrounding area within the preset setting range that is not interfered with will be used as the corresponding display feature area of ​​the target display feature.

[0009] By adopting the above technical solution, feature regions can be divided by adaptively judging the degree of interference between multiple features. This can effectively avoid the region division error caused by the mutual interference of target features, ensure that each displayed feature region is independent and clear with reasonable boundaries, and reduce feature confusion.

[0010] Optionally, the step of dividing the generated content data into multiple regions based on the target display features to obtain the display feature regions further includes the following sub-steps: The ratio of the calculated area ratio to the area reference ratio is used as the area adjustment value; If the area ratio is less than the preset area reference ratio, the setting range is negatively adjusted according to the area adjustment value with a preset first adjustment step size; otherwise, the setting range is positively adjusted according to the area adjustment value with a preset second adjustment step size, wherein the second adjustment step size is less than the first adjustment step size.

[0011] By adopting the above technical solution, the setting range of the feature region can be adaptively and dynamically adjusted according to the degree of feature interference, and fine control can be made by using differentiated step size, so that the feature region division can better fit the actual interference state and reduce feature omissions or misjudgments caused by unreasonable region range.

[0012] Optionally, the step of calculating the contrast value of pixel content in the edge region further includes the following sub-steps: The variance values ​​corresponding to the edge pixels of the target display feature in the edge region are calculated based on the preset pixel coordinate axis to obtain the pixel variance group corresponding to the edge of the target display feature; The variance of each element within a pixel variance group is calculated as the edge variance. A preset edge reference variance is obtained, and the ratio of the edge variance to the edge reference variance is calculated as the contrast value.

[0013] By adopting the above technical solution, the discrete differences of edge pixels can be accurately quantified through two-layer variance operation, which can objectively and finely characterize the edge clarity and blur degree, reduce subjective judgment bias, and achieve sensitive detection of subtle edge defects.

[0014] Optionally, the step of calculating the contrast value of pixel content in the edge region further includes the following sub-steps: If the comparison value is less than the preset comparison reference value, the length range is adjusted according to the negative correlation of the comparison value.

[0015] By adopting the above technical solution, the edge region length range can be adaptively adjusted according to the contrast degree value to ensure that the edge analysis range is highly matched with the actual clarity, thereby reducing the detection deviation caused by a fixed range.

[0016] Optionally, the steps of calculating the actual target position generated by the relative positions between target features in the generated content data, and calculating the error matching degree between the target reference position and the actual target position, further include the following sub-steps: Calculate the pixel coordinates of the target feature, calculate the average of the sum of the pixel coordinates as the center coordinate, use the center coordinate as the position coordinate of the target feature, calculate the pixel distance between the nearest adjacent target feature based on the position coordinate, and use the set of pixel distances as the actual position of the target. Establish a one-to-one correspondence between the related sub-items in the target's actual location and the target's reference location; The absolute value of the difference between the one-to-one corresponding sub-items is calculated as the distance difference, and the ratio between the distance difference and the corresponding sub-item in the target reference position is calculated as the error sub-item; The error matching degree is calculated by weighted average of the error sub-items.

[0017] By adopting the above technical solution, and by locating the feature center coordinates, quantifying the pixel distance, and calculating the error matching degree by weighting, the relative position deviation of the feature can be accurately and objectively quantified, reducing the one-sidedness of a single distance judgment.

[0018] Optionally, the steps of calculating the actual target position generated by the relative positions between target features in the generated content data, and calculating the error matching degree between the target reference position and the actual target position, further include the following sub-steps: Obtain the coordinates of the display center of the generated content data as the display midpoint position; The distance to the midpoint is calculated based on the center position of the target feature and the display midpoint position; In the generated content data, the sum of the distances to the midpoints of the target features is the sum of the midpoint distances. The percentage of midpoint distance is calculated based on the ratio between the midpoint distance and the value of the target feature; Adjust the weight of the corresponding error sub-item based on the negative correlation of the midpoint distance percentage.

[0019] By adopting the above technical solution, the visual importance is determined based on the distance between the feature and the display center, and the error weight is adaptively adjusted. This gives the positional deviation of the visual core area a higher detection weight, which is more in line with the visual perception law of the human eye, and greatly improves the rationality and pertinence of the error matching degree calculation.

[0020] Optionally, the method may further include the following sub-steps: Based on the defect warning prompts, the corresponding relevant areas with defects are marked to obtain the qualified areas; The marked area is visualized as an edge layer, and different colors are used to distinguish and mark different types of target display features.

[0021] By adopting the above technical solution, defect areas can be accurately marked and visualized using edge layers and different target features using multiple colors. This allows for intuitive and clear location of defects and differentiation of defect types, facilitating rapid identification and location of various generated material defects.

[0022] Secondly, this application provides a simulation defect identification system based on dynamically generated materials, employing the following technical solution: A simulation defect identification system based on dynamically generated materials includes a processor, wherein the processor executes the steps of the simulation defect identification method based on dynamically generated materials as described in any one of the preceding claims.

[0023] Thirdly, this application provides a storage medium, which adopts the following technical solution: A storage medium storing a program, which, when executed by a processor, implements the steps of the simulation defect identification method based on dynamically generated materials as described above.

[0024] In summary, this application includes at least one of the following beneficial technical effects: Through an integrated design that combines precise semantic-visual feature mapping, multi-feature interference adaptive region segmentation, edge clarity dual-layer variance quantization, visual weight adaptation for positional error calculation, and defect visualization marking, it achieves fully automated and high-precision identification of simulated defects such as blurred / overly sharp edges, misaligned positions, and disproportionate ratios in generated dynamic materials. This eliminates the need for manual frame-by-frame verification, significantly improving detection efficiency and accuracy. Furthermore, it adapts to the material requirements of different resolutions and scene types. Intuitive defect location and classification labels facilitate rapid correction of subsequent materials, effectively improving the visual realism and practical value of generated materials, and meeting the application needs of scenarios with high requirements for material realism, such as short drama production, e-commerce advertising, and professional visualization. Attached Figure Description

[0025] Figure 1 This is a flowchart illustrating the steps of a simulation defect identification method based on dynamically generated materials.

[0026] Figure 2 This is a diagram showing the steps of adaptive segmentation and display of feature regions using multi-target interferometry.

[0027] Figure 3 This is a step diagram showing how to adaptively adjust the feature region range based on the degree of multi-target interference. Detailed Implementation

[0028] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.

[0029] In the description of this specification, the references to "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples" refer to specific features, structures, materials, or characteristics described in connection with the described embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0030] The simulation defect identification method based on dynamically generated materials described in this application is applicable to dynamic videos, image sequences, and other materials output by generative artificial intelligence models such as diffusion models and DiffusionTransformer. It aims to address common non-semantic simulation defects in existing AI-generated content, such as blurred / overly sharp edges, texture distortion, multi-target interference, physical inconsistencies in target relative positions, and spatial structural errors. The method achieves fully automated, high-precision, and quantifiable defect detection and early warning, effectively improving the visual realism and practical value of dynamic materials. (Refer to...) Figure 1 The specific steps of this application method are as follows: S100 acquires generated content data and generated prompt data, and extracts target semantic features and target display features: S101 Data Acquisition and Preprocessing: Obtain the generated content data to be detected and the corresponding generated prompt data: The generated content data consists of visual materials output by the generative AI model, including single-frame images (in JPG, PNG, etc.) or continuous video frames (frame rate 15-60fps, resolution supports 720P, 1080P, 4K, etc.). The data contains RGB three-channel color information. To eliminate color interference, unify the calculation dimensions, and improve the stability of subsequent feature extraction and edge detection, the generated content data is grayscaled using a weighted average algorithm defined by the international standard ITU-R BT.601. The core formula is: Gray = 0.299 × R + 0.587 × G + 0.114 × B; where R, G, and B are the red, green, and blue channel values ​​of the corresponding pixels (range 0-255), and Gray is the single-channel pixel value after grayscale conversion (range 0-255). For hardware deployment scenarios, an optimized integer operation version can be used: Gray=(R×299+G×587+B×114+500) / 1000, which reduces floating-point operation overhead through rounding; or a shift operation version can be used: Gray=(R×19595+G×38469+B×7472)>>16, which further improves computational efficiency.

[0031] The generated prompt data consists of natural language descriptions input by the user, such as "The male lead is holding a gun and standing in front of the braised duck stall" or "The female lead is walking side by side with the male lead," which includes core semantic information such as target objects, characters, and scene associations.

[0032] S102 Target Semantic Feature Extraction: The generated prompt data is semantically parsed using a pre-defined word meaning extraction algorithm. The specific process is as follows: 1. Semantic word segmentation: The jieba word segmentation algorithm is used to segment the generated prompt data, splitting continuous text into independent lexical units; 2. Stop word filtering: Remove stop words such as "of", "in", "and" that have no actual semantic meaning, and retain the core vocabulary; 3. Entity recognition: Use the BERT pre-trained model for named entity recognition to extract the core semantic entities that the user expects to generate, that is, the target semantic features, such as noun entities like "gun", "spicy duck", "leading male actor", "leading female actor", etc., to clarify the semantic orientation of the detection target; 4. Keyword normalization: Unify synonymous expressions, such as normalizing "pistol" to "gun" to ensure the consistency of semantic features.

[0033] S103 Target display feature extraction: According to the extracted target semantic features, use a preset content extraction algorithm to locate and extract visual targets in the grayscale generated content data. The specific implementation is as follows: For static targets, such as spicy duck and gun, use the YOLOv8 object detection algorithm or the MaskR-CNN semantic segmentation algorithm. Based on the visual feature templates corresponding to the target semantic features, such as shape, texture, and contour features, accurately locate the pixel range of the target in the image, and output the bounding box coordinates and pixel set of the target; For dynamic targets, such as the leading male actor and leading female actor in the video, combine the inter-frame optical flow method to track the target motion trajectory to ensure the consistent recognition of the same target in consecutive video frames and reduce the loss of targets between frames; The extracted target display features need to establish a one-to-one mapping relationship with the target semantic features, that is, clarify the text semantic entity corresponding to each visual target.

[0034] S200 Divide the display feature area and extract the edge and edge detection area: S201 Basic display feature area division: Refer to Figure 2 , according to the bounding box coordinates and pixel distribution range of the target display features, use a region segmentation algorithm to divide the entire generated content data into multiple independent display feature areas, ensuring that a single area only contains one target display feature and reducing pixel interference between multiple targets.

[0035] S202 Multi-feature interference adaptive area division: When the number of extracted target display features ≥ 2, it is necessary to process the problem of multi-target overlap interference. The specific process is as follows: 1. Interference area recognition: Identify the pixel area where the bounding boxes of adjacent target display features overlap through pixel coordinate intersection operation, which is the interference area; 2. Area parameter calculation: Count the total number of pixels in the interference region and use it as the interference area S_inter; traverse all target display features, count the total number of pixels for each feature, and record the minimum value as the base area S_base; calculate the area ratio R_area = S_inter / S_base; 3. Area division rules: Preset area reference ratio R_ref (value range 0.05-0.1, the smallest value greater than 0, for example 0.08, representing the maximum allowable relative interference): If R_area < R_ref: it indicates that the degree of multi-target interference is low and does not affect the main features of the target. The complete pixel area of ​​the target display features and the pixels within the surrounding preset range (initial value is 5-10 pixels, which can be adjusted according to the image resolution) are uniformly divided into the display feature area corresponding to the target. If R_area≥R_ref: This indicates severe multi-target interference and significant occlusion. Only the effective pixel area (non-interference area) that is not occluded in the target display features and the surrounding non-interference pixels within a set range are classified as display feature areas to reduce feature confusion caused by interference areas.

[0036] S203 setting range adaptive adjustment: Reference Figure 3 To further optimize the accuracy of region division, adjustment parameters are calculated based on the above area ratios, and the size of the set range is dynamically adjusted: 1. Calculate the area adjustment value R_adj = R_area / R_ref (when R_area = 0, R_adj = 0); 2. Preset adjustment step size: First adjustment step size Step_1 (value range 0.1-0.3, e.g. 0.2), second adjustment step size Step_2 (value range 0.05-0.1, e.g. 0.08), and Step_2 < Step_1, to ensure fine-tuning characteristics when interference is severe; 3. Range adjustment rules: If R_area < R_ref: the range is negatively correlated with R_adj and the adjustment formula is Range_new = Range_init × (1 - R_adj × Step_1); when R_adj = 0 (no interference), Range_new is set to the maximum value (e.g., 15 pixels) to ensure the inclusion of the complete environment around the target. If R_area≥R_ref: Set the range to be positively correlated with R_adj and adjust it. The adjustment formula is Range_new=Range_init×(1+R_adj×Step_2). By slightly expanding the non-interference area, the feature loss caused by occlusion is compensated.

[0037] S204 Edge Extraction and Edge Region Determination: 1. Edge Extraction: The edges of targets within each display feature region are extracted using the Canny edge detection algorithm or the Sobel / Laplacian edge detection algorithm. Specific parameter settings are as follows (based on industrial vision optimization standards): Gaussian kernel size ksize: Select 3×3 (low noise image), 5×5 (medium noise image) or 7×7 (high noise image) according to the image noise level, for preprocessing denoising, to balance the denoising effect and edge preservation integrity; Dual threshold settings: low threshold threshold1 = 50-100, high threshold threshold2 = 150-250, satisfying threshold2 > threshold1, where gradient value > threshold2 is a strong edge, threshold1 < gradient value < threshold2 and connected to a strong edge is a weak edge, and gradient value < threshold1 is a non-edge, effectively filtering out false edges; Post-processing: Morphological closing operation (3×3 convolution kernel) is used to connect edge breakpoints, remove isolated short pseudo edges, and ensure edge continuity.

[0038] 2. Edge Region Selection: Using the extracted edge pixels as the central reference, a continuous pixel region is selected within a preset length range (initial value 3-8 pixels) along the inner and outer normal directions of the edge, serving as the edge region. This region primarily reflects the sharpness, texture integrity, and natural transition characteristics of the target edge, and is the core detection area for determining defects such as blurriness and distortion.

[0039] S300 calculates edge contrast values ​​and performs adaptive blurring: S301 Two-layer Variance Quantification Comparison Value: The grayscale differences of edge pixels are accurately quantified through two-layer variance calculation, reducing subjective judgment bias. The specific process is as follows: 1. Pixel variance group calculation: Based on the preset pixel coordinate system, with the top left corner of the image as the origin, the horizontal direction to the right as the x-axis, and the vertical direction downward as the y-axis, the edge pixels in the edge region are divided into segments of 5-10 pixels along the edge direction. The variance Var_i of the gray value of each segment is calculated, i=1,2,...,n, where n is the number of segments, forming a pixel variance group [Var_1, Var_2,..., Var_n]. The single variance Var_i represents the degree of gray value dispersion of the local edge; 2. Edge variance calculation: Calculate the variance of all elements within the pixel variance group to obtain the edge variance Var_edge. The formula is: ; in The mean of the pixel variance group is used, while the edge variance enhances the ability to represent the overall uniformity of the edges. 3. Contrast Value Definition: Obtain the preset edge reference variance Var_ref, based on the statistical mean of variance of over 100,000 real-world object edges, with a value range of 30-80. Adjust according to the target type, such as Var_ref=60 for hard objects and Var_ref=40 for soft objects. Calculate the contrast value C=Var_edge / Var_ref. C is positively correlated with edge sharpness: the larger C is, the more dramatic the edge grayscale difference and the sharper the outline; the smaller C is, the smoother the edge grayscale transition and the more blurred the outline.

[0040] S302 edge region length range adaptive adjustment: If the contrast value C is less than the preset contrast reference value C_ref, with a range of 1.0-1.5 (e.g., 1.2), which is the critical value representing whether the edge is clear or not, it indicates that the edge has a blurry defect. In this case, C is adjusted according to the negative correlation between C and the length range. The adjustment formula is L_new = L_init × C, where L_init is the initial length range and L_new is the adjusted length range; Example: If L_init = 6 pixels and C = 0.8 < 1.2, then L_new = 6 × 0.8 = 4.8, which is rounded down to 5 pixels. This reduces the inclusion of invalid blurry pixels outside the edge in the calculation and improves the accuracy of blurry defect detection.

[0041] S303 Adaptive Fuzzification Processing: A preset comparison reference value C_ref (consistent with S302) is used to determine and process the edge state based on the comparison level value: If C > C_ref: This indicates that the edge is too sharp, has abrupt transitions or jagged distortion, which does not match the natural transition characteristics of real object edges, and the pixels in the edge area need to be blurred. Blur method: Gaussian blur algorithm or mean blur, bilateral filtering is used. The blur intensity is controlled by the preset blur parameter K. The value of K ranges from 1 to 5. K is positively correlated with the size of the blur kernel: K=1 corresponds to a 3×3 kernel, K=2 corresponds to a 5×5 kernel, ..., K=5 corresponds to an 11×11 kernel. The larger K is, the higher the blur processing intensity and the smoother the edge transition. Parameter matching rule: C and K are positively correlated, that is, the larger C is, the sharper the edge and the larger the value of K. For example, when C=2.0 (C_ref=1.2), K=4; when C=1.6, K=2, to ensure the targeting and appropriateness of blurring processing and reduce the loss of details caused by excessive blurring.

[0042] S400 acquires a feature location database and matches it with the target reference location: S401 Feature Location Library Construction: A feature location database is pre-built and stored. This database is trained based on a large amount of real-world scene images and videos and has the following characteristics: Data sources: Real footage covering multiple scenarios such as entertainment, education, and commerce, including scenes of human interaction, product display, and natural landscapes, as well as public datasets such as COCO and ImageNet, ensuring the universality and authenticity of the data; Storage content: It includes the standard relative positional relationships of various target semantic features in real physical space and conventional visual logic. Specifically, it is a data set consisting of standard pixel distance, standard orientation angle, and standard spatial ratio between two or more related target features, i.e., target reference position; Example data: For example, the standard relative distance between the male lead and the female lead is 50-100 pixels (1080P resolution), and the azimuth angle is 0°-30° (side by side); the standard relative distance between the male lead and the gun is 10-20 pixels (handheld), and the spatial ratio is 1:0.3 (the ratio of the character's height to the gun's length). Update mechanism: Supports online incremental updates, continuously optimizing the accuracy and coverage of reference locations by adding real-world scenario data.

[0043] S402 target reference position matching: Based on the extracted target semantic features and semantic relationships, such as the handheld relationship between "male lead" and "gun," and the interactive relationship between "male lead" and "female lead," a feature matching algorithm is used to accurately retrieve the corresponding target reference positions in the feature location database, serving as the standard benchmark for subsequent position rationality determination. For example, when the target semantic features are "male lead," "female lead," and "gun," and the semantic relationship is "male lead holding a gun and standing side by side with female lead," the matched target reference positions are {(male lead, female lead): 50-100 pixels, (male lead, gun): 10-20 pixels, (female lead, gun): 60-110 pixels}.

[0044] S500 calculates the match between the target's actual position and the error: Calculation of the actual location of target S501: 1. Feature Center Coordinate Localization: For each target display feature, iterate through the coordinates (x_i, y_i) of all its pixels, i = 1, 2, ..., m, where m is the total number of pixels for that feature. Calculate the center coordinates (X_c, Y_c) using the following formula: Using the center coordinates as the standardized position coordinates of the target features reduces the positioning error caused by the deviation of a single pixel. 2. Neighbor Distance Calculation: Based on the center coordinates, the pixel distance D_j between each target feature and its nearest neighbor is calculated using the Euclidean distance formula, where j = 1, 2, ..., k, and k is the number of neighboring target pairs. The formula is: The set consisting of all D_j is defined as the actual location of the target, and the true spatial distribution among multiple targets is represented in the form of a numerical range.

[0045] S502 Error Matching Degree Calculation: 1. Data Correspondence: Pair the actual location of the target with the distance sub-items in the target reference location that are associated with the attributes and logically corresponding. For example, pair the distance between the male lead and the female lead in the actual location with the distance in the same group in the reference location to ensure data comparison in the same dimension. 2. Error Sub-item Calculation: For each pair of paired sub-items, calculate the absolute value of the distance difference ΔD_j=|D_{j,act}-D_{j,ref}|, where D_{j,act} is the actual distance and D_{j,ref} is the reference distance. Then, using the reference distance as the benchmark, calculate the error sub-item E_j=ΔD_j / D_{j,ref}, where E_j represents the relative deviation of a single set of distances. 3. Visual weight adjustment: Incorporating the principles of human visual perception, the error sub-items are optimized using weighted adjustments. Display center point positioning: Obtain the geometric center coordinates (X_m, Y_m) of the generated content data as the display center point position, such as (960, 540) of a 1920×1080 resolution image; Midpoint distance calculation: Calculate the Euclidean distance between the center coordinates of each target feature and the displayed midpoint position. ; Midpoint distance percentage: Calculates the sum of midpoint distances for all target features. Percentage of midpoint distance for a single target feature ; Weight adjustment: Adjust the weight W_j of the corresponding error sub-item according to the negative correlation of P_j. The adjustment formula is W_j=1-P_j or W_j=e^{-kP_j}, where k is the adjustment coefficient. The closer the target feature is to the center of the image, the smaller P_j is and the larger W_j is. That is, the positional deviation of the visual core area has a higher detection weight. 4. Error matching degree normalization: Perform a weighted average calculation on all weighted error sub-items, using the following formula: The Match value represents the error matching degree (in percentage form). The higher the Match value, the closer the actual position of the target is to the reference position and the more reasonable the spatial structure. The lower the Match value, the greater the positional deviation and the more chaotic the spatial logic.

[0046] S600 Defect Identification, Early Warning, and Visual Marking: S601 Defect Identification and Early Warning: A preset error reference value, Match_ref (ranging from 30% to 50%, for example, 40%), is used as the threshold for determining the reasonableness of the location. If Match < Match_ref: it is determined that the current target display features have simulation defects such as abnormal relative position, disordered spatial structure, disproportionate scale, and physical relationship that violates common sense; Generate defect warning prompts: The warning information includes the defect type (such as "abnormal relative position of the target" or "blurred edge"), the target associated with the defect (such as "male lead and gun"), the coordinates of the defect location (such as the center coordinates of the target and the range of the edge area), and the degree of deviation (such as the error matching degree of 35%). It supports output in formats such as text and JSON, which is convenient for subsequent processing workflows to call.

[0047] S602 Defect Area Visual Marking: To visually represent the location and type of defects, marking is performed based on defect warning prompts: 1. Marking area determination: Based on the defect location coordinates in the early warning prompt, locate the relevant area where the defect exists, and accurately mark it to obtain the marking area; 2. Edge layer display: The marked area is overlaid on the original generated content data as a semi-transparent edge layer. The layer opacity is set to 50%-70%, which highlights the defect area without obscuring the original content. 3. Multi-color differentiation: Different colors are used to mark the display features of different types of targets. For example, the defect area corresponding to "male lead" is red (RGB: 255, 0, 0), "gun" is blue (RGB: 0, 0, 255), and "braised duck" is green (RGB: 0, 255, 0). The line width is set to 2-3 pixels to facilitate quick differentiation of defect-related targets and types.

[0048] This application also discloses a simulation defect identification system based on dynamically generated materials, including a processor, wherein the processor executes the steps of the simulation defect identification method based on dynamically generated materials as described in any of the above embodiments.

[0049] This application also discloses a storage medium storing a program, which, when executed by a processor, implements the steps of the simulation defect identification method based on dynamically generated materials described above.

[0050] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A simulation defect identification method based on dynamically generated materials, characterized in that, Includes the following steps: Acquire generated content data and corresponding generated prompt data; extract target semantic features from generated prompt data based on a preset word meaning extraction algorithm; and extract target display features from generated content data based on the target semantic features and a preset content extraction algorithm. Based on the target display features, the generated content data is divided into multiple regions to obtain the display feature regions. The edges of the target display features are extracted based on the edge extraction algorithm, and the edge regions are selected within a preset length range from the edge. Calculate the contrast value of the pixel content in the edge region. The clearer the pixel content, the higher the contrast value. The more blurred the pixel content, the lower the contrast value. If the contrast value is greater than the preset contrast reference value, the pixel content in the edge region is blurred according to the preset blur parameter. The larger the blur parameter, the higher the degree of blur processing. The smaller the blur parameter, the lower the degree of blur processing. Obtain the feature location library corresponding to the target feature, and match the corresponding target reference position from the feature location library according to the target feature. The target reference position is the relative distance between at least two target features. The actual position of the target is generated by calculating the relative positions between target features in the generated content data, and the error matching degree between the target reference position and the actual position of the target is calculated. If the error matching degree is less than the preset error reference value, a defect warning prompt corresponding to the target feature will be generated.

2. The simulation defect identification method based on dynamically generated materials according to claim 1, characterized in that, The step of dividing the generated content data into multiple regions based on the target display features to obtain the display feature regions also includes the following sub-steps: If there are multiple target display features, the interference region between adjacent target display features is identified, and the area of ​​the interference region is calculated as the interference area; the area of ​​the target display feature with the smallest area is taken as the base area; the ratio of the interference area to the base area is calculated as the area ratio. If the area ratio is less than the preset area reference ratio, the target display feature and the surrounding area within the preset setting range will be used as the display feature area corresponding to the target display feature. Otherwise, the area where the target display feature is not obscured, as well as the area within the set range where there is no interference, will be taken as the display feature area corresponding to the target display feature.

3. The simulation defect identification method based on dynamically generated materials according to claim 2, characterized in that, The step of dividing the generated content data into multiple regions based on the target display features to obtain the display feature regions also includes the following sub-steps: The ratio of the calculated area ratio to the area reference ratio is used as the area adjustment value; If the area ratio is less than the preset area reference ratio, the setting range is adjusted negatively according to the area adjustment value with a preset first adjustment step size. Otherwise, the setting range is positively adjusted according to the area adjustment value with a preset second adjustment step size, wherein the second adjustment step size is smaller than the first adjustment step size.

4. The simulation defect identification method based on dynamically generated materials according to claim 1, characterized in that, The step of calculating the contrast value of pixel content in the edge region also includes the following sub-steps: The variance values ​​corresponding to the edge pixels of the target display feature in the edge region are calculated based on the preset pixel coordinate axis to obtain the pixel variance group corresponding to the edge of the target display feature; The variance of each element within a pixel variance group is calculated as the edge variance. A preset edge reference variance is obtained, and the ratio of the edge variance to the edge reference variance is calculated as the contrast value.

5. The simulation defect identification method based on dynamically generated materials according to claim 4, characterized in that, The step of calculating the contrast value of pixel content in the edge region also includes the following sub-steps: If the comparison value is less than the preset comparison reference value, the length range is adjusted according to the negative correlation of the comparison value.

6. The simulation defect identification method based on dynamically generated materials according to claim 1, characterized in that, The steps of calculating the actual target position generated by the relative positions between target features in the generated content data, and calculating the error matching degree between the target reference position and the actual target position, also include the following sub-steps: Calculate the pixel coordinates of the target feature, calculate the average of the sum of the pixel coordinates as the center coordinate, use the center coordinate as the position coordinate of the target feature, calculate the pixel distance between the nearest adjacent target feature based on the position coordinate, and use the set of pixel distances as the actual position of the target. Establish a one-to-one correspondence between the related sub-items in the target's actual location and the target's reference location; The absolute value of the difference between the one-to-one corresponding sub-items is calculated as the distance difference, and the ratio between the distance difference and the corresponding sub-item in the target reference position is calculated as the error sub-item; The error matching degree is calculated by weighted average of the error sub-items.

7. The simulation defect identification method based on dynamically generated materials according to claim 6, characterized in that, The steps of calculating the actual target position generated by the relative positions between target features in the generated content data, and calculating the error matching degree between the target reference position and the actual target position, also include the following sub-steps: Obtain the coordinates of the display center of the generated content data as the display midpoint position; The distance to the midpoint is calculated based on the center position of the target feature and the display midpoint position; In the generated content data, the sum of the distances to the midpoints of the target features is the sum of the midpoint distances. The percentage of midpoint distance is calculated based on the ratio between the midpoint distance and the value of the target feature; Adjust the weight of the corresponding error sub-item based on the negative correlation of the midpoint distance percentage.

8. The simulation defect identification method based on dynamically generated materials according to claim 1, characterized in that, The method also includes the following sub-steps: Based on the defect warning prompts, the corresponding relevant areas with defects are marked to obtain the qualified areas; The marked area is visualized as an edge layer, and different colors are used to distinguish and mark different types of target display features.

9. A simulation defect identification system based on dynamically generated materials, characterized in that, The device includes a processor that performs the steps of the simulation defect identification method based on dynamically generated materials as described in any one of claims 1-8.

10. A storage medium, characterized in that, The storage medium stores a program, which, when executed by a processor, implements the steps of the simulation defect identification method based on dynamically generated materials as described in any one of claims 1-8.