A visual AI accurate recognition method based on image recognition enhancement

By performing scene recognition and hierarchical enhancement processing on images, the misidentification problem of traditional visual AI recognition technology in complex scenarios is solved, and high-precision and efficient image recognition effect is achieved.

CN120107765BActive Publication Date: 2025-09-05BEIJING TAIHE GUANFU TECH CO LTD
2 Cites 0 Cited by

Patent Information

Application Number
CN202510181697.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-09-05
Estimated Expiration
2045-02-19

Smart Images

  • Figure CN120107765B_ABST
    Figure CN120107765B_ABST
Patent Text Reader

Abstract

The present application relates to a method for accurate visual AI recognition based on image recognition enhancement. The method comprises: performing scene recognition on the image data to be recognized to obtain scene recognition data corresponding to the image data to be recognized; identifying the scene initial primary data and the scene initial secondary data from the image data to be recognized based on the scene recognition data; performing recognition enhancement on the scene initial primary data based on the scene primary data enhancement model to obtain scene enhanced primary data; and performing recognition enhancement on the scene initial secondary data based on the scene secondary data enhancement model to obtain scene enhanced secondary data; and recognizing the image content of the image data to be recognized based on the scene enhanced primary data and the scene enhanced secondary data to obtain image content target recognition data. The use of this method can effectively improve the efficiency of visual AI recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a visual AI precise recognition method based on image recognition enhancement. Background Art

[0002] Traditional visual AI recognition technology primarily relies on deep learning and convolutional neural networks (CNNs). By training models with large amounts of labeled data, they can perform tasks such as object recognition, face recognition, and image segmentation. Data augmentation, transfer learning, and multi-scale analysis are often used to improve the model's adaptability and accuracy in diverse environments. Mainstream applications include autonomous driving, security monitoring, and medical image analysis. However, traditional visual AI recognition technology is highly dependent on data quality and quantity, and in complex and dynamic scenarios, it can lead to misidentification or inaccurate recognition, resulting in low visual AI recognition efficiency. Summary of the Invention

[0003] Based on this, it is necessary to provide a method, device and computer equipment for visual AI precise recognition based on image recognition enhancement, which can effectively improve the efficiency of visual AI recognition, in order to address the above technical problems.

[0004] In a first aspect, the present application provides a visual AI accurate recognition method based on image recognition enhancement, comprising:

[0005] Performing scene recognition on the image data to be recognized to obtain scene recognition data corresponding to the image data to be recognized;

[0006] identifying scene initial primary data and scene initial secondary data from the image data to be identified according to the scene identification data;

[0007] Identify and enhance the initial main data of the scene according to the scene main data enhancement model to obtain scene enhanced main data;

[0008] and, performing recognition enhancement on the initial scene secondary data according to the scene secondary data enhancement model to obtain scene enhanced secondary data;

[0009] The image content of the image data to be identified is identified according to the scene enhancement primary data and the scene enhancement secondary data to obtain image content target identification data.

[0010] In a second aspect, the present application also provides a visual AI accurate recognition device based on image recognition enhancement, comprising:

[0011] A scene recognition module, configured to perform scene recognition on the image data to be recognized, and obtain scene recognition data corresponding to the image data to be recognized;

[0012] a data classification module, configured to identify scene initial primary data and scene initial secondary data from the image data to be identified based on the scene identification data;

[0013] A data enhancement module is used to identify and enhance the initial main data of the scene according to the scene main data enhancement model to obtain scene enhanced main data;

[0014] and, a data enhancement module, further configured to perform recognition enhancement on the initial secondary data of the scene according to the scene secondary data enhancement model to obtain scene enhanced secondary data;

[0015] The content recognition module is used to recognize the image content of the image data to be recognized based on the scene enhancement primary data and the scene enhancement secondary data to obtain image content target recognition data.

[0016] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements any step of a visual AI precise recognition method based on image recognition enhancement:

[0017] The aforementioned visual AI precision recognition method, apparatus, and computer device based on image recognition enhancement perform scene recognition on the image to be identified, obtaining scene recognition data related to the image content, thereby laying the foundation for subsequent refined recognition. Based on the scene recognition results, the image data is divided into "primary scene data" and "secondary scene data," each of which undergoes independent recognition enhancement processing. The primary scene data enhancement model enhances the most prominent and important parts of the image (such as core objects and main structures), improving the recognition system's perception and judgment of these key parts and reducing recognition errors caused by factors such as image blur, poor lighting, or occlusion. Meanwhile, the secondary scene data enhancement model enhances minor elements or background information in the image, enabling the system to better recognize image details such as background objects and environmental features. This reduces interference from background noise or unimportant information and avoids misidentification of irrelevant information. Combining the enhanced primary and secondary data, the image to be identified undergoes final comprehensive recognition, obtaining more accurate and comprehensive image content and target recognition data. This hierarchical and phased enhancement and recognition strategy is implemented. This effectively improves the accuracy, robustness, and adaptability of image recognition systems, ensuring high-precision recognition results, especially in complex and dynamic scenes, effectively increasing the efficiency of visual AI recognition. This significantly optimizes the image recognition process and enhances the practicality and effectiveness of intelligent image recognition technology in diverse applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 This is a diagram of an application environment for a visual AI precise recognition method based on image recognition enhancement in one embodiment;

[0020] Figure 2 1 is a flow chart of a method for accurate visual AI recognition based on image recognition enhancement in one embodiment;

[0021] Figure 3 1 is a flow chart of a method for identifying initial primary data and initial secondary data of a first scenario in one embodiment;

[0022] Figure 4 Schematic diagram of a flow chart of a method for identifying primary scene recognition data and secondary scene recognition data in one embodiment;

[0023] Figure 5 1 is a flow chart of a second method for identifying primary scene recognition data and secondary scene recognition data in one embodiment;

[0024] Figure 6 1 is a flow chart of a method for identifying initial primary data and initial secondary data of a second scenario in one embodiment;

[0025] Figure 7 Schematic diagram of a flow chart of a method for obtaining scene enhancement main data in one embodiment;

[0026] Figure 8 1 is a flow chart of a method for obtaining scene enhancement secondary data in one embodiment;

[0027] Figure 9 Schematic diagram of a flow chart of a method for obtaining image content target recognition data in one embodiment;

[0028] Figure 10 1 is a structural block diagram of a visual AI precise recognition device based on image recognition enhancement in one embodiment;

[0029] Figure 11 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0031] The embodiment of the present application provides a visual AI accurate recognition method based on image recognition enhancement, which can be applied to Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store data that server 104 needs to process. The data storage system can be integrated with server 104 or placed on a cloud or other network server. Server 104 can be implemented as a standalone server or a server cluster consisting of multiple servers.

[0032] In an exemplary embodiment, Figure 2 As shown in the figure, a visual AI accurate recognition method based on image recognition enhancement is provided, which is applied to Figure 1 The server in FIG. 1 is taken as an example to illustrate the method, including the following steps 202 to 210. Among them:

[0033] Step 202: Perform scene recognition on the image data to be recognized to obtain scene recognition data corresponding to the image data to be recognized.

[0034] The image data to be recognized can be raw image data that has not yet been processed and analyzed, typically coming from cameras, image sensors, or other devices. This data contains rich image information, but requires processing and recognition algorithms to extract meaningful content, such as objects, scenes, and people.

[0035] Scene recognition involves performing preliminary scene classification and feature extraction on image data using image processing and analysis algorithms. This process typically utilizes deep learning models (such as convolutional neural networks) to identify scene types in images, such as street scenes, natural scenery, and indoor scenes.

[0036] Scene recognition data can be scene-related information obtained from the image recognition process. This data includes information such as the scene type, relevant features, and object location of the image, which is used for further analysis and processing.

[0037] Specifically, feature extraction and analysis are performed on the image data to be identified through deep learning models or computer vision algorithms (such as convolutional neural networks, CNNs) to identify the main scene features in the image. This process will eventually output a scene recognition data, which contains the scene type, features and other relevant information involved in the image.

[0038] Step 204 : Identify scene initial primary data and scene initial secondary data from the image data to be identified based on the scene identification data.

[0039] The initial scene primary data can be the most important and representative elements or information in the image, extracted from the scene recognition data. Primary data typically refers to prominent objects or structures in the image, such as people, vehicles, and buildings. This data plays a central role in image analysis, helping to determine the subject or focal point of the image.

[0040] Initial scene secondary data can be less prominent, but still important, background information or auxiliary objects in the image extracted from the scene recognition data. Secondary data includes the image's environment, background, and auxiliary objects. While these elements do not directly affect the core understanding of the image, they are still crucial for describing and understanding the overall scene.

[0041] Specifically, deep learning models (such as convolutional neural networks, YOLO or Faster R-CNN, etc.) are combined with scene recognition data to extract features and classify the image data to be identified, and identify the main objects or core elements in the scene recognition data, such as people, vehicles, buildings, etc., which constitute the initial main data of the scene. The identification of the initial main data of the scene depends on the target detection capability of the model. By accurately locating the position and category of the object, the most representative elements are extracted. For the identification of the initial secondary data of the scene, it is usually necessary to use semantic segmentation technology (such as DeepLab, U-Net, etc.) in combination with scene recognition data to perform pixel-level classification of the image data to be identified and distinguish secondary elements such as background, environment or auxiliary objects. These initial secondary data of the scene usually do not directly affect the understanding of the core content, but they are crucial to the integrity and detailed analysis of the scene and can provide richer contextual information.

[0042] Step 206 : Identify and enhance the initial scene main data according to the scene main data enhancement model to obtain scene enhanced main data.

[0043] Among them, the scene main data enhancement model can be a deep learning model or algorithm used to enhance the scene main data. The model uses data enhancement technology (such as rotation, scaling, cropping, color change, etc.) or specific optimization algorithms (such as attention mechanism, transfer learning, etc.) to further improve the recognition accuracy of the main data, so that these key data can be extracted and identified more clearly and accurately.

[0044] Recognition enhancement involves further processing and optimizing already recognized data to improve its accuracy and reliability. This typically involves techniques such as detail enhancement, noise removal, and boundary extraction. The goal is to enable the model to more accurately identify key data in complex or low-quality images, thereby enhancing the model's robustness and adaptability.

[0045] The scene-enhanced primary data can be more accurate and clearer than the primary data obtained after being enhanced using the scene primary data enhancement model. This enhanced data improves the accuracy of recognition results by enhancing image details, making the boundaries of key elements (such as people, objects, and buildings) more defined and their shapes clearer.

[0046] Specifically, the image processing techniques in the scene primary data enhancement model are applied, such as adjusting image contrast, brightness, and color saturation, to enhance the visibility of key areas in the initial scene primary data, making details more prominent. Furthermore, edge enhancement or denoising algorithms in the scene primary data enhancement model are used to further improve the clarity of the primary data, especially in low-quality or blurry images. Local region magnification techniques in the scene primary data enhancement model are then used to refine the primary data regions, amplifying their features and ensuring clearer details, thereby reducing the risk of misidentification or missed recognition. Simultaneously, the image restoration methods in the scene primary data enhancement model, combined with image super-resolution techniques, amplify low-resolution regions and add details, thereby improving the resolution of the primary data. Furthermore, contextual information is leveraged, taking into account the image background and surrounding environment, to optimize the position and size of the primary data. For example, the spatial relationships in the image or information about previously identified objects are combined to automatically adjust the positioning of the primary data to ensure greater accuracy, ultimately resulting in scene-enhanced primary data.

[0047] Step 208 : performing recognition enhancement on the initial scene secondary data according to the scene secondary data enhancement model to obtain scene enhanced secondary data.

[0048] The scene secondary data enhancement model can be a deep learning model or algorithm used to enhance scene secondary data. This model aims to improve the recognition of secondary information in the image (such as background, environment, and auxiliary objects). Commonly used techniques include semantic segmentation and image super-resolution, which aim to enhance the detail and accuracy of secondary data and ensure that details of background or auxiliary elements are not overlooked.

[0049] Scene-enhanced secondary data can be clearer and more accurate secondary data obtained by using the scene secondary data enhancement model. After optimization, the details of the background and other secondary elements are clearer and more recognizable, providing more complete scene information and making image analysis more comprehensive.

[0050] Specifically, semantic segmentation techniques in the scene secondary data enhancement model, such as U-Net or DeepLab, are used to perform more refined pixel-level classification on the initial secondary data of the scene, thereby improving the resolution of background and environmental elements and ensuring that details are not overlooked. Then, image enhancement techniques in the scene secondary data enhancement model are used, such as applying appropriate blurring and contrast adjustment to the background area, to make the secondary data more prominent and ensure that it is distinguished from the primary data. In addition, super-resolution reconstruction methods can be used to restore details of low-resolution secondary data to improve the overall clarity of the image. In order to prevent secondary data from interfering with the recognition of primary data, regional adaptive technology can be used to dynamically adjust the processing method according to the characteristics of different regions in the image to enhance the recognizability of secondary data. Combined with background modeling technology, such as using deep neural networks to predict unimportant elements or objects in the scene, the background part of the image is further refined, and finally scene-enhanced secondary data is obtained.

[0051] Step 210 : Identify the image content of the image data to be identified based on the scene enhancement primary data and the scene enhancement secondary data to obtain image content target identification data.

[0052] Image content target recognition data can be information containing target objects extracted after a comprehensive analysis of the image data to be recognized. Image content target recognition data includes the location information, category labels, and other relevant attributes of all objects (such as people, objects, buildings, etc.) recognized in the image.

[0053] Specifically, the scene enhancement primary data and scene enhancement secondary data are respectively integrated into the image data to be identified to obtain synthetic image data to be identified; further target detection algorithms (such as YOLOv5, RetinaNet, etc.) are used to further identify the synthetic image data to be identified. These algorithms can locate and classify targets based on the enhanced data to generate target position boxes and corresponding category labels in the image. In addition, image description generation technology can be combined to further describe the image content through the model and output detailed semantic information of the image, such as labels such as "pedestrians and vehicles on the street." Finally, image content target recognition data is generated. These data can be used in subsequent applications such as image search, autonomous driving, security monitoring and other fields.

[0054] In the aforementioned visual AI-based precision recognition method based on image recognition enhancement, scene recognition data related to the image content is obtained by performing scene recognition on the image to be identified, laying the foundation for subsequent refined recognition. Based on the scene recognition results, the image data is divided into "primary scene data" and "secondary scene data," each of which undergoes independent recognition enhancement processing. The primary scene data enhancement model enhances the most prominent and important parts of the image (such as core objects and main structures), improving the recognition system's perception and judgment of these key parts and reducing recognition errors caused by factors such as image blur, poor lighting, and occlusion. Meanwhile, the secondary scene data enhancement model enhances minor elements or background information in the image, enabling the system to better recognize image details such as background objects and environmental features. This reduces interference from background noise or unimportant information and avoids misidentification of irrelevant information. Combining the enhanced primary and secondary data, the image to be identified is finally comprehensively recognized, resulting in more accurate and comprehensive image content and target recognition data. This hierarchical and phased enhancement and recognition strategy is implemented. This effectively improves the accuracy, robustness, and adaptability of image recognition systems, ensuring high-precision recognition results, especially in complex and dynamic scenes, effectively increasing the efficiency of visual AI recognition. This significantly optimizes the image recognition process and enhances the practicality and effectiveness of intelligent image recognition technology in diverse applications.

[0055] In an exemplary embodiment, Figure 3 As shown, according to the scene recognition data, the scene initial primary data and the scene initial secondary data are identified from the image data to be identified, including steps 302 to 308.

[0056] Step 302 : performing a generalization analysis on the acquisition scenes of the image data to be recognized based on the scene recognition data to obtain generalization data of each scene.

[0057] Among them, the scene generalization data can be obtained by summarizing, abstracting and inferring the current scene recognition data, which is scene information that is similar to the original scene but has a certain degree of universality.

[0058] Specifically, key features are extracted from the scene recognition data, such as the type, location, color, lighting and other information of the object. Then, based on these features, pattern recognition or data-driven methods are used to map these features to other possible scene types that are similar to the current scene. For example, if the current scene recognition result shows an image of a city street scene, generalization analysis may infer other similar city street scenes, or scenes containing similar features (such as buildings, vehicles, pedestrians, etc.). This process not only expands and matches the features of the current scene, but also helps to infer potential other scenes that are similar in features to the current scene, thereby providing more reference data for subsequent recognition, and ultimately obtaining a set of generalized data that can represent multiple related scenes.

[0059] Step 304 : defining scene recognition rules for the scene recognition data and the generalized data of each scene respectively, and obtaining a set of scene data recognition rules.

[0060] Scene recognition rules can be the criteria and methods used to identify specific scene types when analyzing images. These rules are defined based on image features (such as object shape, color distribution, texture, etc.) to help the model identify the scene in the image.

[0061] The scene data recognition rule set can include a series of scene recognition rules that cover the recognition standards and strategies for different scene features in an image. This set can include multiple categories of rules, each corresponding to one or more scene features, such as buildings, roads, and natural landscapes.

[0062] Specifically, the scene features in the scene recognition data and the generalized data for each scene are analyzed, and a set of scene recognition rules corresponding to each scene data is developed. These rules include classification criteria based on image features, object boundary detection methods, and scene partitioning methods. For example, for architectural scenes, rules such as "buildings should have a rectangular appearance" and "high lighting intensity" can be defined; while for natural scenes, rules such as "green vegetation should appear in the scene" or "there should be a clear sky element in the background" can be defined. The rules corresponding to each scene are organized into a set to obtain the recognition rule set for each scene data.

[0063] Step 306 : Using each scene data recognition rule set, respectively identify scene recognition primary data and scene recognition secondary data from the image data to be recognized.

[0064] The primary data for scene recognition can be the most significant and critical scene information identified from an image, typically the core objects or elements in the image, such as people, vehicles, buildings, etc. This data often occupies a prominent position in the image and is crucial for understanding the subject or content of the image.

[0065] Among them, scene recognition secondary data can be elements identified from images that have an auxiliary role in scene understanding but are relatively less significant. This data usually contains background, environment, or secondary objects such as the sky, roads, plants, etc. Although they are not as important as primary data in the core understanding of the image, they provide important contextual information, help build a complete scene context, and help distinguish different scene types.

[0066] Specifically, the previously defined scene recognition rule sets are applied to the analysis of the images to be recognized. These rule sets are constructed based on the characteristics of the scene (such as object type, position, size, color, etc.), and can guide the model to perform detailed regional division and classification of the image. First, for the primary data, each rule set will guide the model to prioritize the identification of significant elements in the image, such as people, vehicles or buildings, which are usually the core part of the image and have strong spatial positioning characteristics, thereby obtaining the primary data for each scene recognition. Next, for the secondary data, each rule set will help the model identify background information or auxiliary objects in the image, such as the sky, roads, plants, etc. Although these data do not directly affect the main content of the image, they are crucial to the overall understanding of the scene and background reconstruction, thereby obtaining the secondary data for each scene recognition.

[0067] Step 308 : Fusing the primary data for each scene recognition and the secondary data for each scene recognition to obtain the initial primary data for the scene and the initial secondary data for the scene.

[0068] Specifically, since the primary data for each scene recognition represents the core information in the image, it is usually highly recognizable and dominant elements, such as people, vehicles, buildings, etc. During fusion, the primary data will be given a higher weight, and through refined processing (such as boundary adjustment, feature enhancement, etc.), its positioning will be ensured to be more accurate and clear, and the initial primary data of the scene will be obtained. Then, the secondary data for each scene recognition usually contains background or environmental elements in the image, such as roads, sky, vegetation, etc. Although they are not as important as the primary data for the core understanding of the scene, they provide rich contextual information and contribute to the comprehensive analysis of the scene. When fusing the secondary data, spatial relationship processing is used to ensure that these auxiliary information work together with the primary data and do not interfere with the core elements, so as to obtain the initial secondary data of the scene.

[0069] In this embodiment, generalized analysis is performed through scene recognition data, and the model is not limited to the current scene, but can also infer related and similar scenes, thereby improving the breadth and adaptability of recognition; further definition of scene recognition rules ensures the standardization and consistency of the recognition process; the use of rule sets for the recognition of primary and secondary data helps to distinguish key objects from background information, and improves the accuracy and detail of recognition; ultimately, by fusing primary and secondary data, the integrity and richness of scene information is guaranteed, providing high-quality input for subsequent image analysis, classification, and application, thereby improving the overall recognition effect and application value.

[0070] In an exemplary embodiment, Figure 4 As shown, each scene data recognition rule set is used to respectively identify the scene recognition primary data and the scene recognition secondary data from the image data to be recognized, including steps 402 to 406. In which:

[0071] Step 402 : According to each scene data recognition rule set, model parameters of the scene data recognition model to be set are respectively set to obtain each set scene data recognition model.

[0072] The scene data recognition model to be set may be a model that has not yet been configured according to the recognition rules and parameters of a specific scene.

[0073] The set scene data recognition model may be a model that has been configured according to specific scene recognition rules and data features.

[0074] Specifically, model parameters of the corresponding scene data recognition model to be set are set for different scene types according to each set of scene recognition rule sets. These rule sets contain scene recognition standards and features, such as object type, size, color distribution, positional relationship, etc. According to these rules, the parameters of the scene data recognition model to be set (such as weights, thresholds, activation functions, etc.) will be adjusted to ensure that each scene data recognition model to be set can better capture and match the characteristics of a specific scene, and ultimately obtain multiple set scene data recognition models, each of which is optimized for a specific scene type.

[0075] In step 404 , each of the set scene data recognition models is used to identify model recognition primary data and model recognition secondary data from the image data to be recognized.

[0076] Among them, the main data for model recognition can be the most significant and important scene information extracted by the model from the image data to be recognized.

[0077] Among them, the model recognizes secondary data, which can be auxiliary information extracted from the image by the model. This data is usually not as significant as the primary data, but is crucial for a complete understanding of the scene.

[0078] Specifically, each pre-set scene data recognition model will identify the model recognition primary data and model recognition secondary data in the image based on its parameters and rule set. Specifically, each pre-set scene data recognition model will scan different areas in the image, identify key elements that meet the scene recognition rules (such as people, vehicles, buildings, etc.), and mark these elements as model recognition primary data; at the same time, each pre-set scene data recognition model will also identify background information or secondary objects in the image (such as roads, sky, plants, etc.) and classify them as model recognition secondary data.

[0079] Step 406 : For any set scene data recognition model, perform information differentiation and optimization on the model recognition primary data and the model recognition secondary data to obtain the scene recognition primary data and the scene recognition secondary data.

[0080] Among them, information differentiation optimization can be a process of further processing and optimizing the primary data and secondary data identified by the model, with the aim of reducing misidentification and redundant information.

[0081] Specifically, for each set scene data recognition model, the model recognition primary data and model recognition secondary data identified by it are further optimized, wherein the goal of the optimization process is to reduce misidentification and redundant information through information differentiation (such as based on the similarity, spatial relationship, priority, etc. between data). For example, the model recognition primary data can be fine-tuned through boundary optimization technology to avoid misidentification as model recognition secondary data; or the weight of the model recognition primary data can be enhanced through regional importance analysis to ensure that it is more prominent. At the same time, the model recognition secondary data can also be adjusted through background optimization, denoising and other methods to improve the accuracy and clarity of each part of the information in the image. After these optimization processes, the final scene recognition primary data and scene recognition secondary data are obtained.

[0082] In this embodiment, by dynamically setting model parameters based on a set of scene data recognition rules, each recognition model can be optimized for different scene characteristics, enhancing the model's flexibility and adaptability. Using the defined model, primary and secondary data are accurately extracted from the image, ensuring complete recognition of key objects and auxiliary background. This information differentiation and optimization effectively reduces errors and redundancy in the recognition data, improving data clarity and accuracy, and thus enhancing the ability to understand image content. This provides a solid foundation for efficient image analysis and subsequent processing in complex scenarios.

[0083] In an exemplary embodiment, Figure 5 As shown, information differentiation and optimization are performed on the model recognition primary data and the model recognition secondary data to obtain the scene recognition primary data and the scene recognition secondary data, including steps 502 to 504. In which:

[0084] Step 502 : performing classification error correction and multi-view correction on the main model recognition data in sequence to obtain the main scene recognition data.

[0085] Among them, classification error correction can be an operation to correct the misclassification results identified by the model during the image recognition process.

[0086] Among them, multi-view correction can be to correct the recognition results of objects or scenes by combining information from different viewpoints or angles.

[0087] Specifically, classification error correction is first performed on the model's main recognition data. Since misclassification may occur in image recognition, the model will identify and correct these errors based on preset rules or post-processing algorithms. For example, if a building is mistakenly identified as a natural landscape, the correction process will adjust the model's classification criteria and weights to re-determine the correct category of the area and obtain classification error correction data. Next, the model will use image data from different angles or different views to further enhance the accuracy of the recognition results. Since the same scene may appear differently from different perspectives, by fusing features and information from multiple perspectives, the model can better determine the correct location and category of the main data, and ultimately obtain the main data for scene recognition.

[0088] In step 504 , depth of field boundary correction and environmental dynamic correction are sequentially performed on the model recognition secondary data to obtain scene recognition secondary data.

[0089] The depth of field boundary correction may be to analyze the depth information in the image and adjust and optimize the boundary between the object and the background.

[0090] Among them, dynamic environmental correction can take into account the impact of environmental changes (such as weather changes, lighting changes, or object movement, etc.) on the scene and dynamically adjust and correct the recognition results.

[0091] Specifically, by analyzing the depth information in the model-recognized secondary data and the model-recognized primary data, the boundary between the objects and the background in the scene is adjusted to adjust the hierarchy between the secondary data (such as distant views, background, etc.) and the primary data, making the hierarchy between the two clearer. For example, in an urban street scene image, distant buildings and nearby pedestrians should have a clear depth of field distinction, and the correction process will ensure that the positioning of these elements is more accurate. Then, dynamic elements in the processed image, such as weather changes, lighting changes, and moving objects, are dynamically corrected. The model can adaptively adjust the secondary data according to changes in the environment, ensuring that the background and environmental information are not distorted by external dynamic changes, thereby obtaining more stable and accurate scene recognition secondary data.

[0092] In this embodiment, by correcting classification errors in the primary data used for model recognition, misclassification in object recognition is effectively reduced, improving recognition accuracy. Multi-view correction further integrates feature information from different perspectives, making the primary data more stable and consistent in complex scenes. Simultaneously, depth-of-field boundary correction is performed on the secondary data, optimizing the depth relationship between objects and backgrounds and enhancing the sense of spatial hierarchy. Dynamic environmental correction enhances the stability of the background under varying environmental conditions, ensuring the integrity and authenticity of background details. This results in more accurate and natural overall scene recognition results, providing reliable support for image analysis and applications.

[0093] In an exemplary embodiment, Figure 6 As shown, the process of fusing the primary data of each scene recognition and the secondary data of each scene recognition to obtain the initial primary data of the scene and the initial secondary data of the scene includes steps 602 to 606.

[0094] In step 602 , confidence calculation is performed on each scene recognition primary data and each scene recognition secondary data respectively with the image data to be recognized, so as to obtain the confidence of each primary data and the confidence of each secondary data.

[0095] The confidence calculation may be a calculation method for evaluating the reliability and accuracy of each piece of identification data.

[0096] Among them, the main data confidence can be the reliability score of the core elements of the scene (such as people, vehicles, buildings, etc.) appearing in the image.

[0097] The secondary data confidence may be a confidence score of the background or auxiliary elements (such as roads, sky, plants, etc.) recognized by the model.

[0098] Specifically, each primary scene recognition data item and each secondary scene recognition data item are matched against the image data to be recognized, and the confidence level of each data item is calculated. The confidence level is calculated based on the consistency between the data features and the image, such as the position, shape, and color of the object. For each primary scene recognition data item and each secondary scene recognition data item, a corresponding recognition result is generated. The model generates a confidence score for each primary data item and each secondary data item, indicating the reliability of the recognition result.

[0099] In step 604 , the recognition data whose confidences of the primary data and the secondary data are greater than the confidence threshold are selected as the primary data for target recognition and the secondary data for target recognition.

[0100] The confidence threshold may be data used to distinguish the validity of the recognition data.

[0101] The target recognition main data may be main data with a confidence level greater than a confidence threshold.

[0102] The target recognition secondary data may be secondary data having a confidence level greater than a confidence threshold.

[0103] Specifically, the confidence level of each primary data point and each secondary data point is compared with a confidence threshold. The primary data point and secondary data point with a confidence level greater than the preset threshold are selected as valid target recognition primary data and secondary data. This threshold is usually set by the model based on previous training and actual needs to exclude recognition results with too low confidence or unreliable results.

[0104] In step 606 , differential aggregation is performed on each target identification primary data and each target identification secondary data to obtain the scene initial primary data and the scene initial secondary data.

[0105] Among them, differentiated aggregation can be to adopt different processing methods according to the importance or characteristics of the data when merging or fusing different data.

[0106] Specifically, a differentiated aggregation process is performed on the selected primary and secondary object recognition data, enabling targeted fusion based on their importance and characteristics. For primary object recognition data, this data is typically core elements of the scene, such as people, vehicles, and buildings, and is assigned a higher weight. During aggregation, the model fuses primary data of the same type using techniques such as weighted averaging or region merging, ensuring that these important elements are spatially and semantically coherent and accurate. For example, for multiple identified objects (such as multiple views of the same building), the aggregation process considers features such as their position, shape, and size in the image, selecting the most representative features for merging to produce the initial primary scene data. While secondary object recognition data is less critical than primary data for scene understanding, it provides important contextual information to help construct a complete scene. For example, secondary elements in the background, such as roads, sky, and plants, are aggregated using spatial fusion or region optimization to ensure good coherence and consistency with the primary data and avoid errors caused by background clutter, resulting in the initial secondary scene data. In this process, the differentiated aggregation of primary and secondary data can not only retain their respective characteristics but also improve the recognition accuracy of the overall scene.

[0107] In this embodiment, confidence calculations are performed on the primary and secondary scene recognition data to ensure the reliability of each recognition result. This allows the selection of target data with confidence levels above a threshold, while eliminating invalid or erroneous recognitions with low confidence levels, thereby improving the accuracy and reliability of the recognition results. By performing differentiated aggregation on the selected target primary and secondary data, the core features of the primary data and the auxiliary information of the secondary data are retained, achieving an optimized combination of the initial scene data. This results in a more accurate and complete final initial scene primary and secondary data, providing high-quality input for image analysis and understanding, and improving overall recognition performance.

[0108] In an exemplary embodiment, Figure 7 As shown, the initial scene main data is identified and enhanced according to the scene main data enhancement model to obtain the scene enhanced main data, including steps 702 to 712.

[0109] Step 702 : Perform object semantic analysis on each object in the initial main data of the scene to obtain semantic vector data of each object.

[0110] Among them, object semantic analysis can be to extract the semantic features of objects in the image through a deep learning model, including the category, function and attributes of the object, such as "car", "red" or "moving object", and generate a semantic label for each object.

[0111] Among them, object semantic vector data can be the representation of the semantic features of the object as a high-dimensional numerical vector, where each dimension in the vector corresponds to a certain semantic attribute of the object, such as category, color or function, so that semantically similar objects are closer in the semantic space, thereby facilitating the calculation and analysis of the semantic relationship between objects.

[0112] Specifically, semantic analysis is performed on each object in the initial scene data to extract semantic features, such as its category (e.g., "car," "building") and attributes (e.g., "red," "circle"). By analyzing the image using deep learning techniques (e.g., convolutional neural networks), these semantic features can be converted into high-dimensional semantic vector data for each object. Object semantic vector data numerically represents the object's position and features in semantic space, providing foundational data for subsequent reasoning and relationship mining.

[0113] Step 704 : constructing an object semantic relationship graph corresponding to each object based on the semantic vector data of each object.

[0114] The object semantic relationship graph may be a graph structure constructed based on object semantic vectors, wherein nodes represent objects in a scene and edges represent semantic associations between objects.

[0115] Specifically, based on the semantic vector data of each object, an object semantic relationship graph is constructed. During specific execution, the semantic relationships between objects are identified from the semantic vector data of each object, and the specific object semantic relationship graph is expressed through nodes (representing objects) and edges (representing semantic relationships between objects) in the graph. Each object is represented as a node in the graph, and the relationships between objects (such as similarity, dependency, or coexistence) are connected by edges. For example, if "table" and "chair" often appear together in multiple scenes, the edge between them will indicate this coexistence relationship, forming a close semantic connection.

[0116] Step 706 : inferring semantic association information of each object based on the object semantic relationship graph to obtain object semantic enhancement data.

[0117] The semantic association information may be semantic relationship data between objects inferred from the object semantic relationship graph, such as two objects appearing together in multiple scenes or having similar functions.

[0118] Among them, object semantic enhancement data can be an enhanced semantic representation obtained through semantic analysis and reasoning, which includes the semantic features of the object and its semantic association with other objects.

[0119] Specifically, through the edges between object nodes in the object semantic relationship graph, the model can identify deep connections between objects, such as co-occurrence relationships (for example, "sofa" and "coffee table" are often in the same room), or dependencies between objects (for example, "lamp" depends on "power" to work). This relationship information is used to enhance the semantic understanding of objects, thereby inferring the semantic associations between objects and obtaining object semantic enhancement data. For example, if the semantic vector of an object is similar to the semantic vector of another object, the model can infer that they may belong to the same category or have similar functions.

[0120] Step 708 : Calculate the spatial topological relationship between the objects based on the initial main data of the scene to obtain the object spatial layout data.

[0121] The spatial topological relationship may be data describing the relative positions and spatial relationships of objects in a scene, such as whether an object is located above, next to, or far away from another object.

[0122] Among them, the object space layout data can be the overall distribution information of objects in the scene derived based on the spatial topological relationship, which describes the position, arrangement and regional distribution of objects in two-dimensional or three-dimensional space, and provides a basis for regional enhancement.

[0123] Specifically, the scene main data enhancement model uses image processing techniques (such as target detection or depth estimation) to determine the bounding box or three-dimensional position coordinates of each object. The scene main data enhancement model then evaluates the spatial relationship between objects by calculating the relative distance, angle, overlap and other information between the bounding boxes of these objects. For example, the model will determine whether two objects are located above, below, left, right, front and back of each other, or whether they occlude each other. For complex scenes, the scene main data enhancement model can also calculate the relative position and arrangement of objects in three-dimensional space by estimating the depth information of the objects. Finally, based on these calculation results, a spatial topological relationship is obtained; further, based on the spatial topological relationship, the groups, regions or hierarchies of objects in the scene are identified, such as which objects form an area, which objects are closely arranged with each other, and which objects have obvious spatial isolation, to obtain the object spatial layout data.

[0124] Step 710 : performing regional enhancement on each object according to the object spatial layout data to obtain object topology enhancement data.

[0125] Among them, regional enhancement can be to optimize the spatial distribution of objects in the scene by adjusting the position, size or mutual distance of objects after calculating the spatial layout data of objects, so that the arrangement of objects is more natural and reasonable and conforms to the logic of the real scene.

[0126] Among them, the object topology enhancement data can be the optimized object spatial position information generated by regional enhancement, which reflects the reasonable layout and spatial relationship of the object in the scene.

[0127] Specifically, based on the object spatial layout data, the relative positions and arrangement of objects are adjusted to make the layout more consistent with the laws of the real world and the logic of the scene. For example, if the position of an object appears unnatural or inconsistent with other objects, regional enhancement will adjust its position to make the layout of the entire scene more reasonable and in line with physical laws. In this way, object topology enhancement data is obtained.

[0128] Step 712 : Perform weighted fusion on the object semantic enhancement data, the object topology enhancement data, and the initial main scene data to obtain the scene enhancement main data.

[0129] Specifically, object semantic enhancement data, object topology enhancement data, and the initial scene data are weighted and fused. Different types of data (such as semantic, topological, and initial data) are assigned different weights based on their importance in scene understanding. For example, semantic data may require a higher weight to strengthen the association between objects, while topological data provides spatial relationships between objects, helping to better understand how objects are organized in three-dimensional space. Through this weighted fusion, the scene enhancement data is ultimately obtained.

[0130] In this embodiment, by performing object semantic analysis on the initial main data of the scene, extracting rich semantic features, constructing an object semantic relationship graph, and inferring the semantic associations between objects, the semantic understanding ability of object recognition is improved. By calculating the spatial topological relationship between objects, accurate object spatial layout data is obtained, and the objects are further regionally enhanced to optimize the position and structure of the objects in the scene. Ultimately, by weighted fusion of semantic enhancement data, topological enhancement data, and initial data, the complementarity and optimization of multi-dimensional information are achieved. The obtained scene enhancement main data has higher semantic understanding, spatial rationality, and overall consistency, which greatly improves the recognition accuracy of image content and the ability to restore scenes, providing strong support for intelligent analysis and applications.

[0131] In an exemplary embodiment, Figure 8 As shown, the scene initial secondary data is identified and enhanced according to the scene secondary data enhancement model to obtain scene enhanced secondary data, including steps 802 to 808.

[0132] Step 802: Perform regional background analysis on the initial secondary scene data to obtain background feature data for each region.

[0133] Among them, regional background analysis can be to divide the secondary data in the scene (such as background) into multiple areas (such as foreground, distant view, edge area, etc.), and perform feature extraction and analysis on each area separately to identify its background features such as color, texture, brightness, etc.

[0134] The regional background feature data may be a feature description of each background region obtained through regional background analysis, including color distribution, texture pattern, lighting conditions, etc.

[0135] Specifically, the initial secondary data of the scene is divided into multiple regions, such as foreground, background and boundary regions, and the background features of each region, such as color distribution, texture pattern and lighting information, are extracted to generate corresponding regional background feature data.

[0136] Step 804 : performing local detail optimization processing on the background feature data of each region to obtain detail optimized data of each region.

[0137] Among them, local detail optimization processing can be the process of refining and improving the background features of each area, using image sharpening, denoising, super-resolution and other technologies to enhance the details of the background, such as improving edge clarity, reducing noise interference, and making the background clearer and richer.

[0138] Among them, the regional detail optimization data can be the output after local detail optimization processing, which includes the background data of each area after detail enhancement, with higher resolution, clearer texture and more accurate color.

[0139] Specifically, the background feature data of each area is refined, such as applying image sharpening algorithms to enhance edge details, denoising algorithms to reduce background noise, and super-resolution technology to improve the clarity of low-resolution areas, thereby obtaining richer and clearer regional detail optimization data.

[0140] Step 806 : Perform global consistency adjustment on the detail optimization data of each region to obtain background enhancement data of each scene.

[0141] Among them, global consistency adjustment can be achieved by adjusting the visual styles between different areas through methods such as color balance, lighting balance, and texture matching, so that the entire background remains unified in tone, brightness, and details, avoiding the differences between areas affecting the overall look and feel.

[0142] Among them, scene background enhancement data can be background data obtained after regional analysis, local detail optimization and global consistency adjustment, with higher detail clarity and overall consistency, providing stable, natural and high-quality background information for scene analysis.

[0143] Specifically, global consistency adjustments are made to the detail optimization data of all areas. Through methods such as color balance, brightness balance, and texture matching, consistency in hue, lighting, and style is ensured between areas, visual differences between different areas are eliminated, and ultimately background enhancement data for each scene is generated.

[0144] Step 808 : performing weighted fusion on the background enhancement data of each scene and the initial secondary data of the scene to obtain scene enhanced secondary data.

[0145] Specifically, the optimized details and consistency features of each scene background enhancement data are evaluated for their importance within the overall scene context and assigned a higher weight. Meanwhile, the initial secondary data, serving as the foundation for the original background information and providing complete environmental context, is appropriately weighted. Subsequently, a weighted fusion algorithm is used to perform feature matching and numerical merging of the optimized background details with the original background data, ensuring that the integrity of the initial background is preserved while also incorporating detail enhancement and consistency adjustments. The final output is the scene-enhanced secondary data.

[0146] In this embodiment, by performing regional background analysis on the initial secondary data of the scene, the background feature data of each area is accurately extracted, ensuring the meticulousness and integrity of the background information. Subsequently, local detail optimization processing improves the clarity and detail expression of the background area, reducing noise and blur. Through global consistency adjustment, the differences in color, brightness and texture of each area are unified, making the entire background more harmonious and natural. Finally, the optimized background enhancement data is weightedly fused with the initial background data, which not only retains the original information of the background, but also introduces the optimized details and consistency. The generated scene enhancement secondary data has higher detail richness and visual consistency, which improves the realism of the overall scene and the stability of image recognition.

[0147] In an exemplary embodiment, Figure 9 As shown, the image content of the image data to be identified is identified based on the scene enhancement primary data and the scene enhancement secondary data to obtain image content target identification data, including steps 902 to 910. In which:

[0148] Step 902 : performing feature extraction on the image data to be identified based on the scene enhancement primary data and the scene enhancement secondary data to obtain a feature vector of the image to be identified.

[0149] The feature vector of the image to be identified may be a high-dimensional numerical representation obtained after feature extraction of the image, which includes key information such as the shape, color, texture and background of the object in the image.

[0150] Specifically, the scene enhancement primary data and the scene enhancement secondary data are superimposed on the image data to be identified, and the detailed features of each area in the image, such as object contours, texture, color and background information, are extracted through a convolutional neural network (CNN) to generate a high-dimensional feature vector of the image to be identified.

[0151] Step 904 : performing object-scene separation on the image feature vector to be identified to obtain an object-separated image feature vector and a background-separated image feature vector.

[0152] Among them, object-scene separation can be the process of distinguishing objects from the background in an image. Usually, a semantic segmentation algorithm is used to separate the foreground objects and background areas in the image.

[0153] The object separation image feature vector may be a numerical representation separated from the feature vector of the image to be identified, specifically representing the object in the image, and includes information such as the category, boundary, and position of the object.

[0154] Among them, the background separation image feature vector can be a numerical representation of the background part extracted from the image features, including the color, texture, lighting and other features of the background area, which is used for background analysis and optimization.

[0155] Specifically, through semantic segmentation algorithms (such as U-Net or DeepLab), the object area and background area in the image are distinguished, and object separation image feature vectors (representing information such as the shape, category, and position of the object) and background separation image feature vectors (representing background texture, lighting, and environment features) are generated respectively.

[0156] Step 906 : Perform feature enhancement on the object separation image feature vector and the background separation image feature vector respectively to obtain an object separation image enhancement vector and a background separation image enhancement vector.

[0157] Among them, feature enhancement can be to optimize and enhance the extracted image features through algorithms, making the edges of objects clearer and the background details richer, thereby improving the recognizability and classification accuracy of image content.

[0158] Among them, the object separation image enhancement vector can be the object feature data optimized by the feature enhancement algorithm, which highlights the key details and features of the object and improves the recognition accuracy of the object in the image.

[0159] Among them, the background separation image enhancement vector can be a numerical representation of the background features optimized by feature enhancement technology, making the background information clearer and more consistent, and providing more stable environmental features for the overall image analysis.

[0160] Specifically, for the object separation image feature vector, the attention mechanism is used to highlight the edges, shapes and details of key objects; for the background separation image feature vector, the background information is enhanced through methods such as texture refinement and illumination equalization, and finally the object separation image enhancement vector and the background separation image enhancement vector are obtained, which improves the recognizability of each part of the image.

[0161] Step 908 : Identify the image content of the image data to be identified based on the object separation image enhancement vector and the background separation image enhancement vector to obtain initial image content identification data.

[0162] Among them, image content recognition can be performed by analyzing the image through a deep learning algorithm to identify the objects, categories, locations and background environment in the image and output structured recognition results.

[0163] Among them, the initial image content recognition data can be the preliminary output obtained after image content recognition, which includes object category labels, location information, confidence scores, etc., and is used to represent all targets detected in the image.

[0164] Specifically, based on the enhanced object and background feature vectors, deep learning classification and detection algorithms (such as Faster R-CNN or YOLO) are used to identify the image content, identify the object categories, location boxes and background environment in the image, and generate initial image content recognition data, which includes object labels, coordinates, confidence levels, and other content.

[0165] Step 910: Perform non-maximum suppression processing on the initial image content recognition data to obtain image content target recognition data.

[0166] Among them, non-maximum suppression processing can be used to filter detection frames. By removing frames with low confidence and overlapping with high-confidence detection frames, it ensures that each object in the final output has only one optimal detection frame to avoid repeated detection.

[0167] Specifically, the non-maximum suppression (NMS) algorithm is used to screen multiple recognition boxes with high overlap in the initial recognition data, remove boxes with low confidence that overlap with high-confidence boxes, and ensure that each object ultimately has only one optimal detection box, thereby outputting the final image content target recognition data, which accurately labels each object in the image and its corresponding category.

[0168] In this embodiment, feature extraction is performed using scene-enhanced primary and secondary data to generate high-quality feature vectors for the image to be identified, ensuring that the rich features of the objects and background in the image are fully captured. Object-scene separation effectively distinguishes object and background features, avoiding background interference in object recognition. Object and background features are further enhanced, respectively, to improve the clarity of object details and the stability of background information. Image content recognition is performed based on the enhanced feature vectors, ensuring high-precision initial recognition results. Finally, redundant detection frames are removed through non-maximum suppression. The output image content target recognition data has higher accuracy and reliability, providing strong support for precise target recognition in complex image scenes.

[0169] Based on the same inventive concept, the embodiment of the present application also provides a visual AI accurate recognition device based on image recognition enhancement for realizing the above-mentioned visual AI accurate recognition method based on image recognition enhancement, such as Figure 10 As shown, it includes: scene recognition module 1002, data classification module 1004, data enhancement module 1006 and content recognition module 1008. Each module in the above-mentioned visual AI precise recognition device based on image recognition enhancement can be fully or partially implemented by software, hardware or a combination thereof.

[0170] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 11 The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface.

[0171] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0172] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0173] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of each of the above-described method embodiments.

[0174] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.

[0175] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0176] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A visual AI accurate recognition method based on image recognition enhancement, characterized in that: The method comprises: Performing scene recognition on the image data to be recognized to obtain scene recognition data corresponding to the image data to be recognized; identifying scene initial primary data and scene initial secondary data from the image data to be identified according to the scene identification data; Performing object semantic analysis on each object in the initial main data of the scene to obtain semantic vector data of each object; Constructing an object semantic relationship graph corresponding to each of the objects according to the semantic vector data of each of the objects; Inferring semantic association information of each object according to the object semantic relationship graph to obtain object semantic enhancement data; Calculating the spatial topological relationship between the objects based on the initial main data of the scene to obtain the object spatial layout data; Performing regional enhancement on each of the objects according to the object spatial layout data to obtain object topology enhancement data; performing weighted fusion on the object semantic enhancement data, the object topology enhancement data, and the initial main scene data to obtain scene enhancement main data; and, performing recognition enhancement on the initial scene secondary data according to the scene secondary data enhancement model to obtain scene enhanced secondary data; The image content of the image data to be identified is identified according to the scene enhancement primary data and the scene enhancement secondary data to obtain image content target identification data.

2. The method according to claim 1, characterized in that The step of identifying scene initial primary data and scene initial secondary data from the image data to be identified based on the scene identification data includes: Performing a generalization analysis on the acquisition scenes of the image data to be identified based on the scene recognition data to obtain generalization data of each scene; respectively defining scene recognition rules for the scene recognition data and each of the scene generalization data to obtain a set of scene data recognition rules; Using each of the scene data recognition rule sets, respectively identifying scene recognition primary data and scene recognition secondary data from the image data to be recognized; The scene recognition primary data and the scene recognition secondary data are integrated to obtain the scene initial primary data and the scene initial secondary data.

3. The method according to claim 2, characterized in that The using each of the scene data recognition rule sets to respectively identify scene recognition primary data and scene recognition secondary data from the image data to be recognized includes: According to each of the scene data recognition rule sets, model parameters of the scene data recognition model to be set are respectively set to obtain each set scene data recognition model; Using each of the set scene data recognition models, identifying model recognition primary data and model recognition secondary data from the image data to be recognized; For any of the set scene data recognition models, information differentiation and optimization are performed on the model recognition primary data and the model recognition secondary data to obtain the scene recognition primary data and the scene recognition secondary data.

4. The method according to claim 3, characterized in that The performing information differentiation and optimization on the model recognition primary data and the model recognition secondary data to obtain the scene recognition primary data and the scene recognition secondary data includes: performing classification error correction and multi-view correction on the model recognition main data in sequence to obtain the scene recognition main data; The depth of field boundary correction and the environment dynamic correction are sequentially performed on the model recognition secondary data to obtain the scene recognition secondary data.

5. The method according to claim 2, characterized in that The fusing of the scene recognition primary data and the scene recognition secondary data to obtain the scene initial primary data and the scene initial secondary data includes: Calculating the confidence of each of the scene recognition primary data and each of the scene recognition secondary data with the image data to be recognized, to obtain the confidence of each primary data and the confidence of each secondary data; Selecting the recognition data whose confidence levels are greater than the confidence threshold among the primary data and the secondary data as the primary data and the secondary data for target recognition; Differentiated aggregation is performed on each of the target identification primary data and each of the target identification secondary data to obtain the scene initial primary data and the scene initial secondary data.

6. The method according to claim 1, characterized in that The identifying and enhancing the initial scene secondary data according to the scene secondary data enhancement model to obtain scene enhanced secondary data includes: Performing regional background analysis on the initial secondary data of the scene to obtain background feature data of each region; Performing local detail optimization processing on the background feature data of each region to obtain detail optimized data of each region; Performing global consistency adjustment on the detail optimization data of each region to obtain background enhancement data of each scene; The background enhancement data of each scene and the initial secondary data of the scene are weightedly fused to obtain the scene enhancement secondary data.

7. The method according to any one of claims 1 to 6, characterized in that The step of identifying the image content of the image data to be identified based on the scene enhancement primary data and the scene enhancement secondary data to obtain image content target identification data includes: performing feature extraction on the image data to be identified according to the scene enhancement primary data and the scene enhancement secondary data to obtain a feature vector of the image to be identified; Performing object-scene separation on the feature vector of the image to be identified to obtain an object-separated image feature vector and a background-separated image feature vector; Performing feature enhancement on the object separation image feature vector and the background separation image feature vector respectively to obtain an object separation image enhancement vector and a background separation image enhancement vector; Identifying the image content of the image data to be identified based on the object separation image enhancement vector and the background separation image enhancement vector to obtain initial image content identification data; Non-maximum suppression processing is performed on the image content initial recognition data to obtain the image content target recognition data.

8. A visual AI accurate recognition device based on image recognition enhancement, characterized in that: The device comprises: A scene recognition module, configured to perform scene recognition on the image data to be recognized, and obtain scene recognition data corresponding to the image data to be recognized; a data classification module, configured to identify scene initial primary data and scene initial secondary data from the image data to be identified based on the scene identification data; A data enhancement module, configured to perform object semantic analysis on each object in the initial main data of the scene to obtain semantic vector data of each object; Constructing an object semantic relationship graph corresponding to each of the objects according to the semantic vector data of each of the objects; Inferring semantic association information of each object according to the object semantic relationship graph to obtain object semantic enhancement data; Calculating the spatial topological relationship between the objects based on the initial main data of the scene to obtain the object spatial layout data; Performing regional enhancement on each of the objects according to the object spatial layout data to obtain object topology enhancement data; performing weighted fusion on the object semantic enhancement data, the object topology enhancement data, and the initial main scene data to obtain scene enhancement main data; and, a data enhancement module, further configured to perform recognition enhancement on the initial secondary data of the scene according to the scene secondary data enhancement model to obtain scene enhanced secondary data; The content recognition module is used to recognize the image content of the image data to be recognized based on the scene enhancement primary data and the scene enhancement secondary data to obtain image content target recognition data.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Scene understanding information generation method and device, equipment and medium

    CN119251657A

  • Image processing method, image processing apparatus and display device

    US20160293138A1