Visual AI accurate identification method based on image identification enhancement
By performing scene recognition and hierarchical enhancement processing on images, the problem of inefficiency of traditional visual AI recognition in complex scenarios is solved, and higher recognition accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510181697.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-19
AI Technical Summary
Traditional visual AI recognition technology has a high dependence on data quality and quantity, and may misidentify or fail to accurately identify in complex and dynamic scenarios, resulting in inefficiency.
Scene recognition is performed by identifying images, which are divided into scene main data and secondary data, and are respectively identified and enhanced, and the final comprehensive identification is performed by combining the enhanced main data and secondary data.
The accuracy, robustness and adaptability of the image recognition system are improved, especially in complex and dynamic scenarios, which can ensure high-precision recognition results and effectively improve the efficiency of visual AI recognition.
Smart Images

Figure CN120107765A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a visual AI precise recognition method based on image recognition enhancement. Background Art
[0002] Visual AI recognition technology in traditional technology mainly relies on deep learning and convolutional neural networks (CNN). By training models with a large amount of labeled data, it can achieve tasks such as object recognition, face recognition, and image segmentation. Usually through data enhancement, transfer learning, and multi-scale analysis, the model can improve its adaptability and accuracy in different environments. Mainstream applications include autonomous driving, security monitoring, and medical image analysis. However, traditional visual AI recognition technology is highly dependent on data quality and quantity, and in complex and dynamic scenarios, misidentification or inaccurate recognition may occur, resulting in low efficiency of visual AI recognition. Summary of the invention
[0003] Based on this, it is necessary to provide a visual AI precise recognition method, device and computer equipment based on image recognition enhancement that can effectively improve the efficiency of visual AI recognition in response to the above technical problems.
[0004] In the first aspect, the present application provides a visual AI accurate recognition method based on image recognition enhancement, comprising:
[0005] Performing scene recognition on the image data to be recognized to obtain scene recognition data corresponding to the image data to be recognized;
[0006] According to the scene recognition data, identifying scene initial primary data and scene initial secondary data from the image data to be recognized;
[0007] According to the scene main data enhancement model, the initial scene main data is identified and enhanced to obtain scene enhanced main data;
[0008] And, identifying and enhancing the initial secondary data of the scene according to the scene secondary data enhancement model to obtain scene enhanced secondary data;
[0009] The image content of the image data to be identified is identified according to the scene enhancement primary data and the scene enhancement secondary data to obtain image content target identification data.
[0010] In the second aspect, the present application also provides a visual AI accurate recognition device based on image recognition enhancement, comprising:
[0011] A scene recognition module, used to perform scene recognition on the image data to be recognized, and obtain scene recognition data corresponding to the image data to be recognized;
[0012] A data classification module, used for identifying scene initial primary data and scene initial secondary data from the image data to be identified according to the scene identification data;
[0013] A data enhancement module, used to identify and enhance the initial main data of the scene according to the scene main data enhancement model to obtain scene enhanced main data;
[0014] And, the data enhancement module is further used to identify and enhance the initial secondary data of the scene according to the scene secondary data enhancement model to obtain scene enhanced secondary data;
[0015] The content recognition module is used to recognize the image content of the image data to be recognized according to the scene enhancement primary data and the scene enhancement secondary data, so as to obtain image content target recognition data.
[0016] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, any step of a visual AI precise recognition method based on image recognition enhancement is implemented:
[0017] The above-mentioned visual AI accurate recognition method, device and computer equipment based on image recognition enhancement obtain scene recognition data related to the image content by performing scene recognition on the image to be recognized, thereby laying the foundation for subsequent refined recognition. According to the scene recognition results, the image data is divided into "scene main data" and "scene secondary data", and both are independently recognized and enhanced; the scene main data enhancement model enhances the recognition system's perception and judgment ability of these key parts by strengthening the most significant and important parts of the image (such as core objects, main structures, etc.), reducing the recognition errors caused by factors such as image blur, poor light or occlusion. At the same time, the scene secondary data enhancement model enhances the secondary elements or background information in the image, so that the system can better recognize the details in the image, such as background objects, environmental features, etc., thereby reducing the interference caused by background noise or unimportant information, and avoiding the misrecognition of irrelevant information. Combined with the enhanced primary data and secondary data, the image to be recognized is finally comprehensively recognized, and more accurate and comprehensive image content target recognition data is obtained. Through this hierarchical and staged enhancement and recognition strategy. Therefore, it can effectively improve the accuracy, robustness and adaptability of the image recognition system, especially in complex and dynamic scenes, it can ensure high-precision recognition results and effectively improve the efficiency of visual AI recognition. It can also significantly optimize the image recognition process and enhance the practicality and effectiveness of intelligent image recognition technology in a variety of applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 This is a diagram of an application environment of a visual AI precise recognition method based on image recognition enhancement in one embodiment;
[0020] Figure 2 It is a flowchart of a visual AI accurate recognition method based on image recognition enhancement in one embodiment;
[0021] Figure 3 It is a flowchart of a method for identifying the initial primary data and the initial secondary data of a first scene in one embodiment;
[0022] Figure 4 It is a flowchart of a method for identifying primary data and secondary data for scene recognition in a first embodiment;
[0023] Figure 5 It is a flowchart of a second method for identifying primary scene recognition data and secondary scene recognition data in one embodiment;
[0024] Figure 6 It is a flowchart of a method for identifying the second scene initial primary data and scene initial secondary data in one embodiment;
[0025] Figure 7 A schematic diagram of a flow chart of a method for obtaining scene enhancement main data in one embodiment;
[0026] Figure 8 It is a flowchart of a method for obtaining scene enhancement secondary data in one embodiment;
[0027] Fig. 9 A schematic diagram of a flow chart of a method for obtaining image content target recognition data in one embodiment;
[0028] Fig.10 It is a structural block diagram of a visual AI accurate recognition device based on image recognition enhancement in one embodiment;
[0029] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0031] The embodiment of the present application provides a visual AI accurate recognition method based on image recognition enhancement, which can be applied to Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. Among them, the server 104 can be implemented with an independent server or a server cluster composed of multiple servers.
[0032] In an exemplary embodiment, Figure 2 As shown in the figure, a visual AI accurate recognition method based on image recognition enhancement is provided, and the method is applied to Figure 1 The server in the example is used to illustrate, including the following steps 202 to 210. Among them:
[0033] Step 202: Perform scene recognition on the image data to be recognized to obtain scene recognition data corresponding to the image data to be recognized.
[0034] The image data to be identified can be raw image data that has not been processed and analyzed, usually from cameras, image sensors or other devices. This data contains rich image information, but requires processing and recognition algorithms to extract meaningful content, such as objects, scenes, people, etc.
[0035] Among them, scene recognition can be a preliminary scene classification and feature extraction process of the image data to be recognized through image processing and analysis algorithms. This process usually uses deep learning models (such as convolutional neural networks) to identify the scene types in the image, such as street scenes, natural scenery, indoor scenes, etc.
[0036] The scene recognition data may be scene-related information obtained from the image recognition process, which includes the scene type, related features, object location, etc. of the image, and is used for further analysis and processing.
[0037] Specifically, the image data to be identified is subjected to feature extraction and analysis through a deep learning model or a computer vision algorithm (such as a convolutional neural network, CNN) to identify the main scene features in the image. This process ultimately outputs a scene recognition data, which contains the scene type, features, and other relevant information involved in the image.
[0038] Step 204 , identifying scene initial primary data and scene initial secondary data from the image data to be identified according to the scene identification data.
[0039] The scene initial primary data can be the most important and representative elements or information in the image extracted from the scene recognition data. The primary data usually involves significant objects or structures in the image, such as people, vehicles, buildings, etc. These data play a core role in image analysis and can help determine the theme or focus of the image.
[0040] Among them, the scene initial secondary data can be less significant but still important background information or auxiliary objects in the image extracted from the scene recognition data. Secondary data includes the environment, background, auxiliary objects, etc. of the image. Although these elements do not directly affect the core understanding of the image, they are still crucial to the description and understanding of the overall scene.
[0041] Specifically, deep learning models (such as convolutional neural networks, YOLO or Faster R-CNN, etc.) are used in combination with scene recognition data to extract features and classify the image data to be recognized, and the main objects or core elements in the scene recognition data, such as people, vehicles, buildings, etc., are identified, which constitute the initial main data of the scene. The recognition of the initial main data of the scene depends on the target detection capability of the model, and the most representative elements are extracted by accurately locating the position and category of the object. For the recognition of the initial secondary data of the scene, it is usually necessary to use semantic segmentation technology (such as DeepLab, U-Net, etc.) in combination with scene recognition data to classify the image data to be recognized at the pixel level and distinguish secondary elements such as background, environment or auxiliary objects. These initial secondary data of the scene usually do not directly affect the understanding of the core content, but they are crucial to the integrity and detailed analysis of the scene and can provide richer contextual information.
[0042] Step 206, identifying and enhancing the initial main data of the scene according to the scene main data enhancement model to obtain enhanced main data of the scene.
[0043] Among them, the scene main data enhancement model can be a deep learning model or algorithm used to enhance the scene main data. The model further improves the recognition accuracy of the main data through data enhancement technology (such as rotation, scaling, cropping, color change, etc.) or specific optimization algorithms (such as attention mechanism, transfer learning, etc.), so that these key data can be extracted and identified more clearly and accurately.
[0044] Among them, recognition enhancement can be further processing and optimization of the already recognized data to improve its accuracy and reliability. Recognition enhancement usually includes technologies such as detail enhancement, noise elimination, and boundary extraction. The purpose is to enable the model to more accurately identify key data in complex or low-quality images and improve the robustness and adaptability of the model.
[0045] The scene enhanced main data may be more accurate and clear main data obtained after enhancement processing using the scene main data enhancement model. These enhanced data enhance the accuracy of the recognition results by strengthening the image details, making the boundaries of the main elements (such as people, objects, buildings, etc.) clearer and the shapes clearer.
[0046] Specifically, the image processing technology in the scene main data enhancement model is applied, such as adjusting the contrast, brightness, color saturation, etc. of the image, to enhance the visibility of the key areas in the scene's initial main data, making the details more prominent. In addition, the edge enhancement or denoising algorithm in the scene main data enhancement model is used to further improve the clarity of the main data, especially in low-quality or blurred images. Then, the local area magnification technology in the scene main data enhancement model is used to refine the main data area, amplify the features of these areas, and ensure that the details are clearer, thereby reducing the risk of misidentification or missed identification. At the same time, the image restoration method in the scene main data enhancement model is combined with the image super-resolution technology to enlarge and supplement the details of the low-resolution area to improve the resolution of the main data. In addition, the context information is used to optimize the position and size of the main data in combination with the background and surrounding environment in the image. For example, the positioning of the main data is automatically adjusted in combination with the spatial relationship of the image or the information of previously identified objects to ensure that it is more accurate, and finally the scene enhanced main data is obtained.
[0047] Step 208: perform recognition enhancement on the scene initial secondary data according to the scene secondary data enhancement model to obtain scene enhanced secondary data.
[0048] Among them, the scene secondary data enhancement model can be a deep learning model or algorithm used to enhance the scene secondary data. The goal of this model is to improve the recognition effect of secondary information in the image (such as background, environment, auxiliary objects). Commonly used technologies include semantic segmentation and image super-resolution, which aim to enhance the detail expression and accuracy of secondary data and ensure that the details of background or auxiliary elements are not ignored.
[0049] Among them, the scene enhanced secondary data can be the clearer and more accurate secondary data obtained after using the scene secondary data enhancement model. After the optimization processing, the details of the background and other secondary elements are clearer and more recognizable, providing more complete scene information, making the image analysis more comprehensive.
[0050] Specifically, the semantic segmentation technology in the scene secondary data enhancement model, such as U-Net or DeepLab, is used to perform more refined pixel-level classification on the initial secondary data of the scene, thereby improving the resolution of the background and environmental elements and ensuring that the details are not ignored. Then, the image enhancement technology in the scene secondary data enhancement model is used, such as applying appropriate blur processing and contrast adjustment to the background area, so that the secondary data is more prominent and ensures that it is distinguished from the main data. In addition, the low-resolution secondary data can be restored in detail through super-resolution reconstruction methods to improve the overall clarity of the image. In order to avoid the secondary data interfering with the recognition of the main data, the regional adaptive technology can be used to dynamically adjust the processing method according to the characteristics of different regions in the image to enhance the recognizability of the secondary data. Combined with background modeling technology, such as using deep neural networks to predict unimportant elements or objects in the scene, the background part of the image is further refined, and finally the scene enhanced secondary data is obtained.
[0051] Step 210 , identifying the image content of the image data to be identified based on the scene enhancement primary data and the scene enhancement secondary data, to obtain image content target identification data.
[0052] The image content target recognition data may be information containing the target object extracted after a comprehensive analysis of the image data to be recognized. The image content target recognition data includes the location information, category labels and other related attributes of all targets (such as people, objects, buildings, etc.) recognized in the image.
[0053] Specifically, the scene enhancement primary data and scene enhancement secondary data are respectively integrated into the image data to be identified to obtain synthetic image data to be identified; further use the target detection algorithm (such as YOLOv5, RetinaNet, etc.) to further identify the synthetic image data to be identified. These algorithms can locate and classify targets based on the enhanced data, and generate target position boxes and corresponding category labels in the image. In addition, it can also be combined with image description generation technology to further describe the image content through the model and output detailed semantic information of the image, such as labels such as "pedestrians and vehicles on the street". Finally, image content target recognition data is generated. These data can be used in subsequent applications such as image search, autonomous driving, security monitoring and other fields.
[0054] In the above-mentioned visual AI accurate recognition method based on image recognition enhancement, scene recognition data related to the image content is obtained by performing scene recognition on the image to be recognized, thereby laying the foundation for subsequent refined recognition. According to the scene recognition results, the image data is divided into "scene main data" and "scene secondary data", and both are independently recognized and enhanced; the scene main data enhancement model enhances the recognition system's perception and judgment ability of these key parts by strengthening the most significant and important parts of the image (such as core objects, main structures, etc.), reducing the recognition errors caused by factors such as image blur, poor light or occlusion. At the same time, the scene secondary data enhancement model enhances the secondary elements or background information in the image, so that the system can better recognize the details in the image, such as background objects, environmental features, etc., thereby reducing the interference caused by background noise or unimportant information, and avoiding the misrecognition of irrelevant information. Combined with the enhanced primary and secondary data, the image to be recognized is finally comprehensively recognized, and more accurate and comprehensive image content target recognition data is obtained. Through this hierarchical and staged enhancement and recognition strategy. Therefore, it can effectively improve the accuracy, robustness and adaptability of the image recognition system, especially in complex and dynamic scenes, it can ensure high-precision recognition results and effectively improve the efficiency of visual AI recognition. It can also significantly optimize the image recognition process and enhance the practicality and effectiveness of intelligent image recognition technology in a variety of applications.
[0055] In an exemplary embodiment, Figure 3 As shown, according to the scene recognition data, the scene initial primary data and the scene initial secondary data are recognized from the image data to be recognized, including steps 302 to 308. Among them:
[0056] Step 302 , based on the scene recognition data, generalization analysis is performed on the acquisition scenes of the image data to be recognized to obtain generalization data of each scene.
[0057] Among them, the scene generalization data can be obtained by summarizing, abstracting and inferring the current scene recognition data, which is scene information similar to the original scene but with certain universality.
[0058] Specifically, key features are extracted from the scene recognition data, such as the type, location, color, lighting and other information of the object. Then, based on these features, pattern recognition or data-driven methods are used to map these features to other possible scene types that are similar to the current scene. For example, if the current scene recognition result shows an image of a city street scene, generalization analysis may infer other similar city street scenes, or scenes containing similar features (such as buildings, vehicles, pedestrians, etc.). This process not only expands and matches the features of the current scene, but also helps to infer potential other scenes that are similar in features to the current scene, thereby providing more reference data for subsequent recognition, and ultimately obtaining a set of generalized data that can represent multiple related scenes.
[0059] Step 304 , respectively define the scene recognition rules for the scene recognition data and the generalized data for each scene, and obtain a set of recognition rules for each scene data.
[0060] Among them, scene recognition rules can be the standards and methods used to identify specific scene types when analyzing images. These rules are defined based on the characteristics of the image (such as object shape, color distribution, texture, etc.) to help the model recognize the scene in the image.
[0061] The scene data recognition rule set may be a set of scene recognition rules that cover recognition standards and strategies for different scene features in an image. This set may include multiple categories of rules, each of which corresponds to one or more scene features, such as buildings, roads, natural landscapes, etc.
[0062] Specifically, the scene features in the scene recognition data and the generalized data of each scene are analyzed, and a set of corresponding scene recognition rules are formulated for each scene data. These rules include classification standards based on image features, object boundary detection methods, scene partition processing methods, etc. For example, for architectural scenes, rules such as "buildings should have a rectangular appearance" and "high light intensity" can be defined; and for natural scenes, rules such as "green vegetation appears in the scene" or "there are obvious sky elements in the background" can be defined. The rules corresponding to each scene are sorted into a set to obtain a set of recognition rules for each scene data.
[0063] Step 306 : using each scene data recognition rule set, respectively identifying scene recognition primary data and scene recognition secondary data from the image data to be recognized.
[0064] Among them, the main data of scene recognition can be the most significant and key scene information identified from the image, usually the core objects or elements in the image, such as people, vehicles, buildings, etc. These data often occupy a prominent position in the image and are crucial to understanding the theme or content of the image.
[0065] Among them, scene recognition secondary data can be elements identified from images that are auxiliary but relatively less significant for scene understanding. These data usually contain background, environment or secondary objects, such as sky, road, plants, etc. Although they are not as important as primary data in the core understanding of the image, they provide important contextual information, help build a complete scene background and help distinguish different scene types.
[0066] Specifically, the previously defined scene recognition rule sets are applied to the analysis of the images to be recognized. These rule sets are constructed according to the characteristics of the scene (such as object type, position, size, color, etc.), and can guide the model to perform detailed regional division and classification of the image. First, for the main data, each rule set will guide the model to prioritize the recognition of significant elements in the image, such as people, vehicles or buildings, which are usually the core part of the image and have strong spatial positioning characteristics, and obtain the main data for each scene recognition. Then, for the secondary data, each rule set will help the model recognize the background information or auxiliary objects in the image, such as the sky, roads, plants, etc. Although these data do not directly affect the main content of the image, they are crucial to the overall understanding of the scene and the reconstruction of the background, and obtain the secondary data for each scene recognition.
[0067] Step 308: fuse the primary data for scene recognition and the secondary data for scene recognition to obtain the initial primary data for the scene and the initial secondary data for the scene.
[0068] Specifically, since each scene recognition primary data represents the core information in the image, it is usually a highly recognizable, dominant element, such as people, vehicles, buildings, etc. During fusion, the primary data will be given a higher weight, and through refined processing (such as boundary adjustment, feature enhancement, etc.) to ensure that its positioning is more accurate and clear, the initial primary data of the scene is obtained. Then, each scene recognition secondary data usually contains background or environmental elements in the image, such as roads, sky, vegetation, etc. Although they are not as important as the primary data for the core understanding of the scene, they provide rich contextual information and contribute to the comprehensive analysis of the scene. When fusing the secondary data, spatial relationship processing is used to ensure that these auxiliary information work together with the primary data and do not interfere with the core elements, so as to obtain the initial secondary data of the scene.
[0069] In this embodiment, generalized analysis is performed through scene recognition data, and the model is not limited to the current scene, but can also infer related and similar scenes, thereby improving the breadth and adaptability of recognition; scene recognition rules are further defined to ensure the standardization and consistency of the recognition process; the use of rule sets for the recognition of primary and secondary data helps to distinguish key objects and background information, and improve the accuracy and detail of recognition; ultimately, by fusing primary and secondary data, the integrity and richness of scene information is guaranteed, providing high-quality input for subsequent image analysis, classification, and application, thereby improving the overall recognition effect and application value.
[0070] In an exemplary embodiment, Figure 4 As shown, each scene data recognition rule set is used to respectively recognize scene recognition primary data and scene recognition secondary data from the image data to be recognized, including steps 402 to 406. Among them:
[0071] Step 402 : According to each scene data recognition rule set, model parameters of the scene data recognition model to be set are respectively set to obtain each set scene data recognition model.
[0072] The scene data recognition model to be set may be a model that has not yet been configured according to the recognition rules and parameters of a specific scene.
[0073] The set scene data recognition model may be a model that has been configured according to specific scene recognition rules and data features.
[0074] Specifically, model parameters of the corresponding scene data recognition models to be set are set for different scene types according to each set of scene recognition rule sets, and these rule sets contain scene recognition standards and features, such as object type, size, color distribution, position relationship, etc. According to these rules, the parameters of the scene data recognition models to be set (such as weights, thresholds, activation functions, etc.) will be adjusted to ensure that each scene data recognition model to be set can better capture and match the characteristics of a specific scene, and finally multiple set scene data recognition models are obtained, each of which is optimized for a specific scene type.
[0075] Step 404 , using each of the set scene data recognition models, recognizes model recognition primary data and model recognition secondary data from the image data to be recognized.
[0076] Among them, the main data for model recognition can be the most significant and important scene information extracted by the model from the image data to be recognized.
[0077] Among them, the model recognizes secondary data, which can be auxiliary information extracted from the image by the model. This data is usually not as significant as the primary data, but is crucial for a complete understanding of the scene.
[0078] Specifically, each set scene data recognition model will identify the model recognition primary data and model recognition secondary data in the image according to its parameters and rule sets. Specifically, each set scene data recognition model will scan different areas in the image, identify key elements that meet the scene recognition rules (such as people, vehicles, buildings, etc.), and mark these elements as model recognition primary data; at the same time, each set scene data recognition model will also identify background information or secondary objects in the image (such as roads, sky, plants, etc.) and classify them as model recognition secondary data.
[0079] Step 406, for any set scene data recognition model, perform information differentiation optimization on the model recognition primary data and the model recognition secondary data to obtain the scene recognition primary data and the scene recognition secondary data.
[0080] Among them, information differentiation optimization can be a process of further processing and optimizing the primary and secondary data identified by the model, with the aim of reducing misidentification and redundant information.
[0081] Specifically, for each set scene data recognition model, the model recognition primary data and model recognition secondary data identified by it are further optimized, wherein the goal of the optimization process is to reduce misidentification and redundant information through information differentiation (such as based on similarity, spatial relationship, priority, etc. between data). For example, the model recognition primary data can be fine-tuned through boundary optimization technology to avoid misidentification as model recognition secondary data; or the weight of the model recognition primary data can be enhanced through regional importance analysis to ensure that it is more prominent. At the same time, the model recognition secondary data can also be adjusted through background optimization, denoising and other methods to improve the accuracy and clarity of each part of the information in the image. After these optimization processes, the scene recognition primary data and scene recognition secondary data are finally obtained.
[0082] In this embodiment, by dynamically setting model parameters based on the scene data recognition rule set, each recognition model can be optimized for different scene features, thereby improving the flexibility and adaptability of the model. The set model is used to accurately extract primary and secondary data from the image, ensuring the complete recognition of key objects and auxiliary backgrounds. Through information differentiation optimization, errors and redundancies in recognition data are effectively reduced, the clarity and accuracy of data are improved, thereby enhancing the ability to understand image content, and providing a solid foundation for efficient image analysis and subsequent processing in complex scenes.
[0083] In an exemplary embodiment, Figure 5 As shown, information differentiation and optimization are performed on the model recognition primary data and the model recognition secondary data to obtain the scene recognition primary data and the scene recognition secondary data, including steps 502 to 504. Among them:
[0084] Step 502, classifying error correction and multi-view correction are performed on the main data of model recognition in sequence to obtain the main data of scene recognition.
[0085] Among them, classification error correction can be an operation to correct the misclassification results identified by the model during the image recognition process.
[0086] Among them, multi-view correction can be to correct the recognition result of the object or scene by combining information from different viewpoints or angles.
[0087] Specifically, the classification error correction is first performed on the model recognition main data. Since misclassification may occur in image recognition, the model will identify and correct these errors based on preset rules or post-processing algorithms. For example, if a building is mistakenly identified as a natural landscape, the correction process will re-determine the correct category of the area by adjusting the model's classification criteria and weights to obtain classification error correction data. Next, the model will use image data from different angles or different views to further enhance the accuracy of the recognition results. Since the same scene may appear differently from different perspectives, by integrating features and information from multiple perspectives, the model can better determine the correct location and category of the main data, and ultimately obtain the main data for scene recognition.
[0088] Step 504, sequentially performing depth of field boundary correction and environmental dynamic correction on the model recognition secondary data to obtain scene recognition secondary data.
[0089] The depth of field boundary correction may be to analyze the depth information in the image and adjust and optimize the boundary between the object and the background.
[0090] Among them, dynamic environmental correction can take into account the impact of environmental changes (such as weather changes, lighting changes, or object movement, etc.) on the scene, and dynamically adjust and correct the recognition results.
[0091] Specifically, by analyzing the depth information in the model-recognized secondary data and the model-recognized primary data, the boundary between the object and the background in the scene is adjusted to adjust the hierarchy between the secondary data (such as distant view, background, etc.) and the primary data, making the hierarchy between the two clearer. For example, in an urban street scene image, there should be a clear depth distinction between distant buildings and nearby pedestrians, and the correction process will ensure that the positioning of these elements is more accurate. Then, dynamic elements in the processed image, such as weather changes, lighting changes, moving objects, etc., are dynamically corrected. The model can adaptively adjust the secondary data according to changes in the environment to ensure that the background and environmental information will not be distorted due to external dynamic changes, thereby obtaining more stable and accurate scene recognition secondary data.
[0092] In this embodiment, by correcting the classification errors of the main data of model recognition, the misclassification phenomenon in object recognition is effectively reduced and the recognition accuracy is improved; the multi-view correction further integrates the feature information under different view angles, making the main data more stable and consistent in complex scenes. At the same time, the depth of field boundary correction is performed on the secondary data to optimize the depth relationship between the object and the background and enhance the sense of spatial hierarchy; the environmental dynamic correction enhances the stability of the background under different environmental changes, ensuring the integrity and authenticity of the background details, so that the overall scene recognition results are more accurate and natural, providing reliable support for image analysis and application.
[0093] In an exemplary embodiment, Figure 6 As shown, the primary data of each scene recognition and the secondary data of each scene recognition are integrated to obtain the initial primary data of the scene and the initial secondary data of the scene, including steps 602 to 606. Among them:
[0094] Step 602 , calculate the confidence of each scene recognition primary data and each scene recognition secondary data with the image data to be recognized, so as to obtain the confidence of each primary data and the confidence of each secondary data.
[0095] Among them, confidence calculation can be a calculation method for evaluating the reliability and accuracy of each identification data.
[0096] Among them, the main data confidence can be the reliability score of the core elements (such as people, vehicles, buildings, etc.) in the scene appearing in the image.
[0097] Among them, the secondary data confidence can be the credibility score of the background or auxiliary elements (such as roads, sky, plants, etc.) recognized by the model.
[0098] Specifically, each scene recognition primary data and each scene recognition secondary data are matched with the image data to be recognized, and the confidence of each data item is calculated. The confidence calculation is based on the consistency between the features of the data and the image, such as the position, shape, color, etc. of the object. For each scene recognition primary data and each scene recognition secondary data, there is a corresponding recognition result. The model will give a confidence score as the confidence of each primary data and each secondary data, indicating the reliability of the recognition result.
[0099] Step 604 , selecting the recognition data whose confidences of the primary data and the secondary data are greater than the confidence threshold as the primary data for target recognition and the secondary data for target recognition.
[0100] The confidence threshold may be data used to distinguish the validity of the identification data.
[0101] Among them, the target recognition main data can be the main data with a confidence level greater than a confidence level threshold.
[0102] Among them, the target recognition secondary data can be secondary data with a confidence level greater than a confidence level threshold.
[0103] Specifically, the confidence of each primary data and each secondary data is compared with the confidence threshold, and the primary data and secondary data with confidence greater than the preset threshold are selected as valid target recognition primary data and target recognition secondary data. This threshold is usually set by the model based on previous training and actual needs to exclude those recognition results with low confidence or unreliable.
[0104] Step 606, differentially aggregate each target recognition primary data and each target recognition secondary data to obtain scene initial primary data and scene initial secondary data.
[0105] Among them, differentiated aggregation can be to adopt different processing methods according to the importance or characteristics of the data when merging or fusing different data.
[0106] Specifically, for each selected target recognition primary data and each target recognition secondary data, the process of differential aggregation is used to carry out targeted fusion according to the importance and characteristics of the data. For target recognition primary data, it is usually the core elements in the scene, such as people, vehicles, buildings, etc., and these data will be given a higher weight. During aggregation, the model will fuse the same type of primary data through techniques such as weighted averaging or regional merging to ensure that these important elements are more coherent and accurate in space and semantics. For example, for multiple identified objects of the same object (such as multiple perspectives of the same building), the aggregation process will consider their position, shape, size and other features in the image, select the most representative features for merging, and obtain the initial primary data of the scene. For target recognition secondary data, although they are not as critical as primary data in scene understanding, they provide important contextual information to help build a complete scene. For example, secondary elements such as roads, skies, and plants in the background will be aggregated through spatial fusion or regional optimization to ensure that they have good connection and consistency with the primary data, and no errors are caused by background clutter, so as to obtain the initial secondary data of the scene. In this process, the differentiated aggregation of primary and secondary data can not only retain their respective characteristics but also improve the recognition accuracy of the overall scene.
[0107] In this embodiment, by calculating the confidence of the primary and secondary data of scene recognition, the reliability evaluation of each recognition result is ensured, thereby screening out target data with a confidence higher than the threshold, excluding invalid or erroneous recognitions with low confidence, and improving the accuracy and credibility of the recognition results. By differentially aggregating the screened target primary and secondary data, the core features of the primary data and the auxiliary information of the secondary data are retained, and the optimal combination of the scene initial data is achieved, making the final scene initial primary and secondary data more accurate and complete, providing high-quality input for image analysis and understanding, and improving the overall recognition performance.
[0108] In an exemplary embodiment, Figure 7 As shown, the initial main data of the scene is identified and enhanced according to the scene main data enhancement model to obtain the scene enhanced main data, including steps 702 to 712. Among them:
[0109] Step 702: perform object semantic analysis on each object in the initial main data of the scene to obtain semantic vector data of each object.
[0110] Among them, object semantic analysis can be to extract the semantic features of objects in the image through a deep learning model, including the category, function and attributes of the object, such as "car", "red" or "moving object", and generate semantic labels for each object.
[0111] Among them, the object semantic vector data can be the semantic features of the object represented as a high-dimensional numerical vector, in which each dimension corresponds to a certain semantic attribute of the object, such as category, color or function, so that semantically similar objects are closer in the semantic space, thereby facilitating the calculation and analysis of the semantic relationship between objects.
[0112] Specifically, semantic analysis is performed on each object in the initial main data of the scene to extract the semantic features of the object, such as the category of the object (such as "car", "building"), attributes (such as "red", "circle"), etc. These semantic features can be converted into high-dimensional semantic vector data of each object by analyzing the image through deep learning technology (such as convolutional neural network). The object semantic vector data represents the position and features of the object in the semantic space in the form of a numerical value, thereby providing basic data for subsequent reasoning and relationship mining.
[0113] Step 704: construct an object semantic relationship graph corresponding to each object based on the semantic vector data of each object.
[0114] The object semantic relationship graph may be a graph structure constructed based on the object semantic vector, wherein nodes represent objects in the scene and edges represent semantic associations between objects.
[0115] Specifically, based on the semantic vector data of each object, an object semantic relationship graph is constructed. During specific execution, the semantic relationship between objects is identified from the semantic vector data of each object, and the specific object semantic relationship graph is expressed through nodes (representing objects) and edges (representing semantic relationships between objects) in the graph. Each object is represented as a node in the graph, and the relationship between objects (such as similarity, dependency, or coexistence) connects these nodes through edges. For example, if "table" and "chair" often appear together in multiple scenes, the edge between them will indicate this coexistence relationship, forming a close semantic connection.
[0116] Step 706, inferring the semantic association information of each object according to the object semantic relationship graph to obtain object semantic enhancement data.
[0117] The semantic association information may be semantic relationship data between objects inferred from the object semantic relationship graph, such as two objects appearing together in multiple scenes or having similar functions.
[0118] Among them, the object semantic enhancement data can be an enhanced semantic representation obtained through semantic analysis and reasoning, including the semantic features of the object and its semantic association with other objects.
[0119] Specifically, through the edges between object nodes in the object semantic relationship graph, the model can identify deep connections between objects, such as co-occurrence relationships (for example, "sofa" and "coffee table" are often in the same room), or dependencies between objects (for example, "lamp" depends on "power" to work). This relationship information is used to enhance the semantic understanding of objects, thereby inferring the semantic associations between objects and obtaining object semantic enhancement data. For example, if the semantic vector of an object is similar to the semantic vector of another object, the model can infer that they may belong to the same category or have similar functions.
[0120] Step 708, calculating the spatial topological relationship between the objects based on the initial main data of the scene, and obtaining the object spatial layout data.
[0121] The spatial topological relationship may be data describing the relative positions and spatial relationships of objects in a scene, such as whether an object is located above, beside, or far from another object.
[0122] Among them, the object space layout data can be the overall distribution information of objects in the scene derived based on the spatial topological relationship, which describes the position, arrangement and regional distribution of objects in two-dimensional or three-dimensional space, and provides a basis for regional enhancement.
[0123] Specifically, the scene main data enhancement model determines the bounding box or three-dimensional position coordinates of each object through image processing techniques (such as target detection or depth estimation). Then the scene main data enhancement model evaluates the spatial relationship between objects by calculating the relative distance, angle, overlap and other information between the bounding boxes of these objects. For example, the model will determine whether two objects are located above, below, left, right, front and back of each other, or whether they occlude each other. For complex scenes, the scene main data enhancement model can also calculate the relative position and arrangement of objects in three-dimensional space by estimating the depth information of the objects. Finally, based on these calculation results, a spatial topological relationship is obtained; further, according to the spatial topological relationship, the groups, regions or hierarchies of objects in the scene are identified, such as which objects form an area, which objects are closely arranged to each other, and which objects have obvious spatial isolation, and the spatial layout data of the objects is obtained.
[0124] Step 710 , performing regional enhancement on each object according to the object spatial layout data to obtain object topology enhancement data.
[0125] Among them, regional enhancement can be to optimize the spatial distribution of objects in the scene by adjusting the position, size or mutual distance of objects after calculating the spatial layout data of objects, so that the arrangement of objects is more natural and reasonable and conforms to the logic of the real scene.
[0126] Among them, the object topology enhancement data can be the optimized object spatial position information generated by regional enhancement, which reflects the reasonable layout and spatial relationship of the object in the scene.
[0127] Specifically, according to the object spatial layout data, the relative positions and arrangements of objects are adjusted so that the layout of objects is more consistent with the laws of the real world and the logic of the scene. For example, if the position of an object seems unnatural or inconsistent with other objects, regional enhancement will adjust its position so that the layout of the entire scene is more reasonable and in line with physical rules. In this way, object topology enhancement data is obtained.
[0128] Step 712: weighted fusion is performed on the object semantic enhancement data, the object topology enhancement data and the initial main scene data to obtain the scene enhancement main data.
[0129] Specifically, the object semantic enhancement data, object topology enhancement data and the initial main data of the scene are weighted and fused. Different types of data (such as semantic, topological and initial data) are given different weights according to their importance in scene understanding. For example, semantic data may require a higher weight to strengthen the association between objects, while topological data provides spatial relationships between objects to help better understand how objects are organized in three-dimensional space. Through this weighted fusion, the main scene enhancement data is finally obtained.
[0130] In this embodiment, by performing object semantic analysis on the initial main data of the scene, extracting rich semantic features, constructing an object semantic relationship graph, and inferring the semantic associations between objects, the semantic understanding ability of object recognition is improved. By calculating the spatial topological relationship between objects, accurate object spatial layout data is obtained, and the objects are further regionally enhanced to optimize the position and structure of the objects in the scene. Finally, by weighted fusion of semantic enhancement data, topological enhancement data, and initial data, the complementarity and optimization of multi-dimensional information are achieved, and the obtained scene enhancement main data has higher semantic understanding, spatial rationality, and overall consistency, which greatly improves the recognition accuracy of image content and the ability to restore scenes, and provides strong support for intelligent analysis and applications.
[0131] In an exemplary embodiment, Figure 8 As shown, the scene initial secondary data is identified and enhanced according to the scene secondary data enhancement model to obtain scene enhanced secondary data, including steps 802 to 808. Among them:
[0132] Step 802, performing regional background analysis on the initial secondary scene data to obtain background feature data for each region.
[0133] Among them, regional background analysis can be to divide the secondary data in the scene (such as background) into multiple areas (such as foreground, distant view, edge area, etc.), and perform feature extraction and analysis on each area to identify its background features such as color, texture, brightness, etc.
[0134] The regional background feature data may be a feature description of each background region obtained through regional background analysis, including color distribution, texture pattern, lighting conditions, etc.
[0135] Specifically, the initial secondary data of the scene is divided into multiple regions, such as foreground, background and boundary regions, and the background features of each region, such as color distribution, texture pattern and lighting information, are extracted to generate corresponding regional background feature data.
[0136] Step 804: Perform local detail optimization processing on the background feature data of each region to obtain detail optimized data of each region.
[0137] Among them, local detail optimization processing can be the process of refining and improving the background features of each area, using image sharpening, denoising, super-resolution and other technologies to enhance the details of the background, such as improving edge clarity, reducing noise interference, and making the background clearer and richer.
[0138] Among them, the regional detail optimization data can be the output after local detail optimization processing, which includes the background data of each area after detail enhancement, with higher resolution, clearer texture and more accurate color.
[0139] Specifically, the background feature data of each area is refined, such as applying image sharpening algorithms to enhance edge details, denoising algorithms to reduce background noise, and super-resolution technology to improve the clarity of low-resolution areas, thereby obtaining richer and clearer regional detail optimization data.
[0140] Step 806: Perform global consistency adjustment on the detail optimization data of each region to obtain background enhancement data of each scene.
[0141] Among them, global consistency adjustment can be achieved by adjusting the visual styles between different areas through methods such as color balance, lighting balance, and texture matching, so that the entire background remains unified in tone, brightness, and details, avoiding the differences between areas affecting the overall look and feel.
[0142] Among them, the scene background enhancement data can be the background data obtained after regional analysis, local detail optimization and global consistency adjustment. It has higher detail clarity and overall consistency, providing stable, natural and high-quality background information for scene analysis.
[0143] Specifically, global consistency adjustments are made to the detail optimization data of all areas. Through methods such as color balance, brightness balance, and texture matching, consistency in tone, lighting, and style is ensured between areas, visual differences between different areas are eliminated, and ultimately background enhancement data for each scene is generated.
[0144] Step 808: weighted fusion is performed on the background enhancement data of each scene and the initial secondary data of the scene to obtain scene enhancement secondary data.
[0145] Specifically, according to the optimized details and consistency features in each scene background enhancement data, its importance in the overall scene background is evaluated and a higher weight is assigned to it; at the same time, the initial secondary data, as the basis of the original background information, provides a complete environmental context and is given an appropriate weight. Subsequently, through the weighted fusion algorithm, the optimized background details are feature matched and numerically merged with the original background data to ensure that the integrity of the initial background is retained, while detail enhancement and consistency adjustment are introduced, and the final output is the scene enhancement secondary data.
[0146] In this embodiment, by performing regional background analysis on the initial secondary data of the scene, the background feature data of each area is accurately extracted to ensure the meticulousness and integrity of the background information. Subsequently, the local detail optimization process improves the clarity and detail performance of the background area and reduces noise and blur. Through global consistency adjustment, the differences in color, brightness and texture of each area are unified to make the entire background more harmonious and natural. Finally, the optimized background enhancement data is weightedly fused with the initial background data, which not only retains the original information of the background, but also introduces the optimized details and consistency. The generated scene enhancement secondary data has higher detail richness and visual consistency, which improves the realism of the overall scene and the stability of image recognition.
[0147] In an exemplary embodiment, Fig. 9 As shown, according to the scene enhancement primary data and the scene enhancement secondary data, the image content of the image data to be identified is identified to obtain the image content target identification data, including steps 902 to 910. Among them:
[0148] Step 902: extract features of the image data to be identified based on the scene enhancement primary data and the scene enhancement secondary data to obtain a feature vector of the image to be identified.
[0149] The feature vector of the image to be identified may be a high-dimensional numerical representation obtained after feature extraction of the image, which includes key information such as the shape, color, texture and background of the object in the image.
[0150] Specifically, the scene enhancement primary data and the scene enhancement secondary data are superimposed on the image data to be identified, and the detailed features of each area in the image, such as object contour, texture, color and background information, are extracted through a convolutional neural network (CNN) to generate a high-dimensional feature vector of the image to be identified.
[0151] Step 904 , performing object-scene separation on the image feature vector to be identified, to obtain an object-separated image feature vector and a background-separated image feature vector.
[0152] Among them, object-scene separation can be the process of distinguishing objects from the background in an image. Usually, a semantic segmentation algorithm is used to separate the foreground objects and background areas in the image.
[0153] The object separation image feature vector may be a numerical representation separated from the image feature vector to be identified and specifically represents the object in the image, and includes information such as the category, boundary, and position of the object.
[0154] Among them, the background separation image feature vector can be a numerical representation of the background part extracted from the image features, including the color, texture, lighting and other characteristics of the background area, which is used for background analysis and optimization.
[0155] Specifically, through a semantic segmentation algorithm (such as U-Net or DeepLab), the object area and the background area in the image are distinguished, and an object separation image feature vector (representing information such as the shape, category, and position of the object) and a background separation image feature vector (representing background texture, lighting, and environment features) are generated respectively.
[0156] Step 906 , respectively perform feature enhancement on the object separation image feature vector and the background separation image feature vector to obtain an object separation image enhancement vector and a background separation image enhancement vector.
[0157] Among them, feature enhancement can be achieved by optimizing and enhancing the extracted image features through algorithms, making the edges of objects clearer and the background details richer, thereby improving the recognizability of image content and classification accuracy.
[0158] Among them, the object separation image enhancement vector can be the object feature data optimized by the feature enhancement algorithm, which highlights the key details and features of the object and improves the recognition accuracy of the object in the image.
[0159] Among them, the background separation image enhancement vector can be a numerical representation of the background features optimized by feature enhancement technology, making the background information clearer and more consistent, and providing more stable environmental features for the overall image analysis.
[0160] Specifically, for the object separation image feature vector, the attention mechanism is used to highlight the edges, shapes and details of key objects; for the background separation image feature vector, the background information is enhanced through methods such as texture refinement and illumination equalization, and finally the object separation image enhancement vector and the background separation image enhancement vector are obtained, which improves the recognizability of each part of the image.
[0161] Step 908: Identify the image content of the image data to be identified based on the object separation image enhancement vector and the background separation image enhancement vector to obtain initial image content identification data.
[0162] Among them, image content recognition can be performed by analyzing the image through a deep learning algorithm, identifying the objects, categories, locations and background environment in the image, and outputting structured recognition results.
[0163] Among them, the initial image content recognition data can be the preliminary output obtained after image content recognition, which includes object category labels, location information, confidence scores, etc., and is used to represent all targets detected in the image.
[0164] Specifically, based on the enhanced object and background feature vectors, deep learning classification and detection algorithms (such as Faster R-CNN or YOLO) are used to identify the image content, identify the object categories, location boxes and background environments in the image, and generate initial image content recognition data, which includes object labels, coordinates, confidence levels, and other content.
[0165] Step 910, performing non-maximum suppression processing on the image content initial recognition data to obtain image content target recognition data.
[0166] Among them, non-maximum suppression processing can be used to screen detection frames, by removing frames with low confidence and overlapping with high-confidence detection frames, to ensure that each object in the final output has only one optimal detection frame to avoid repeated detection.
[0167] Specifically, the non-maximum suppression (NMS) algorithm is used to screen multiple recognition boxes with high overlap in the initial recognition data, remove boxes with low confidence and overlapping with high confidence boxes, and ensure that each object ultimately has only one optimal detection box, thereby outputting the final image content target recognition data, which accurately labels each object in the image and its corresponding category.
[0168] In this embodiment, by using scene enhancement primary data and secondary data for feature extraction, a high-quality feature vector of the image to be identified is generated, ensuring that the rich features of the object and background in the image are fully captured. By separating the object and the scene, the object and background features are effectively distinguished, avoiding the interference of the background on object recognition. The object and background features are further enhanced respectively, improving the clarity of the object details and the stability of the background information. Image content recognition is performed based on the enhanced feature vector, ensuring high-precision initial recognition results, and finally removing redundant detection frames through non-maximum suppression. The output image content target recognition data has higher accuracy and reliability, providing strong support for accurate target recognition in complex image scenes.
[0169] Based on the same inventive concept, the embodiment of the present application also provides a visual AI accurate recognition device based on image recognition enhancement for implementing the above-mentioned visual AI accurate recognition method based on image recognition enhancement, such as Fig.10 As shown, it includes: scene recognition module 1002, data classification module 1004, data enhancement module 1006 and content recognition module 1008. Each module in the above-mentioned visual AI precise recognition device based on image recognition enhancement can be fully or partially implemented by software, hardware and their combination.
[0170] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig.11 The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface.
[0171] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.
[0172] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0173] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the above-mentioned method embodiments.
[0174] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.
[0175] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0176] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A visual AI accurate recognition method based on image recognition enhancement, characterized in that: The method comprises: Performing scene recognition on the image data to be recognized to obtain scene recognition data corresponding to the image data to be recognized; According to the scene recognition data, identifying scene initial primary data and scene initial secondary data from the image data to be recognized; According to the scene main data enhancement model, the initial scene main data is identified and enhanced to obtain scene enhanced main data; And, identifying and enhancing the initial secondary data of the scene according to the scene secondary data enhancement model to obtain scene enhanced secondary data; The image content of the image data to be identified is identified according to the scene enhancement primary data and the scene enhancement secondary data to obtain image content target identification data.
2. The method according to claim 1, characterized in that: The step of identifying scene initial primary data and scene initial secondary data from the image data to be identified according to the scene identification data includes: According to the scene recognition data, generalization analysis is performed on the acquisition scene of the image data to be recognized to obtain generalization data of each scene; Respectively defining the scene recognition rules for the scene recognition data and each of the scene generalization data to obtain a set of scene data recognition rules; Using each of the scene data recognition rule sets, respectively identifying scene recognition primary data and scene recognition secondary data from the image data to be recognized; The scene recognition primary data and the scene recognition secondary data are integrated to obtain the scene initial primary data and the scene initial secondary data.
3. The method according to claim 2, characterized in that The using each of the scene data recognition rule sets to respectively identify scene recognition primary data and scene recognition secondary data from the image data to be recognized includes: According to each of the scene data recognition rule sets, model parameters of the scene data recognition models to be set are respectively set to obtain each of the set scene data recognition models; Using each of the set scene data recognition models, identifying model recognition primary data and model recognition secondary data from the image data to be recognized; For any of the set scene data recognition models, information differentiation and optimization are performed on the model recognition primary data and the model recognition secondary data to obtain the scene recognition primary data and the scene recognition secondary data.
4. The method according to claim 3, characterized in that The performing information differentiation and optimization on the model recognition primary data and the model recognition secondary data to obtain the scene recognition primary data and the scene recognition secondary data includes: Performing classification error correction and multi-view correction on the model recognition main data in sequence to obtain the scene recognition main data; The depth of field boundary correction and the environment dynamic correction are performed on the model recognition secondary data in sequence to obtain the scene recognition secondary data.
5. The method according to claim 2, characterized in that: The fusing of each of the scene recognition primary data and each of the scene recognition secondary data to obtain the scene initial primary data and the scene initial secondary data includes: Calculating the confidence of each of the scene recognition primary data and each of the scene recognition secondary data with the image data to be recognized, to obtain the confidence of each primary data and the confidence of each secondary data; Selecting the identification data whose confidences of the primary data and the secondary data are greater than the confidence threshold as the primary data for target identification and the secondary data for target identification; Differentiated aggregation is performed on each of the target recognition primary data and each of the target recognition secondary data to obtain the scene initial primary data and the scene initial secondary data.
6. The method according to claim 1, characterized in that The identifying and enhancing the initial main data of the scene according to the scene main data enhancement model to obtain the scene enhanced main data includes: Performing object semantic analysis on each object in the initial main data of the scene to obtain semantic vector data of each object; Constructing an object semantic relationship graph corresponding to each of the objects according to the semantic vector data of each of the objects; Inferring semantic association information of each of the objects according to the object semantic relationship graph to obtain object semantic enhancement data; Calculating the spatial topological relationship between the objects based on the initial main data of the scene to obtain the spatial layout data of the objects; According to the object spatial layout data, each of the objects is subjected to regional enhancement to obtain object topology enhancement data; The object semantic enhancement data, the object topology enhancement data and the scene initial main data are weightedly fused to obtain the scene enhancement main data.
7. The method according to claim 1, characterized in that The identifying and enhancing the initial scene secondary data according to the scene secondary data enhancement model to obtain scene enhanced secondary data includes: Performing regional background analysis on the initial secondary data of the scene to obtain background feature data of each region; Performing local detail optimization processing on the background feature data of each region to obtain detail optimization data of each region; Performing global consistency adjustment on the detail optimization data of each region to obtain background enhancement data of each scene; The background enhancement data of each scene and the initial secondary data of the scene are weightedly fused to obtain the scene enhancement secondary data.
8. The method according to any one of claims 1 to 7, characterized in that: The step of identifying the image content of the image data to be identified according to the scene enhancement primary data and the scene enhancement secondary data to obtain image content target identification data includes: Extracting features of the image data to be identified according to the scene enhancement primary data and the scene enhancement secondary data to obtain a feature vector of the image to be identified; Performing object-scene separation on the feature vector of the image to be identified to obtain an object-separated image feature vector and a background-separated image feature vector; Respectively perform feature enhancement on the object separation image feature vector and the background separation image feature vector to obtain an object separation image enhancement vector and a background separation image enhancement vector; Identify the image content of the image data to be identified according to the object separation image enhancement vector and the background separation image enhancement vector to obtain initial image content identification data; The image content initial recognition data is subjected to non-maximum suppression processing to obtain the image content target recognition data.
9. A visual AI accurate recognition device based on image recognition enhancement, characterized in that: The device comprises: A scene recognition module, used to perform scene recognition on the image data to be recognized, and obtain scene recognition data corresponding to the image data to be recognized; A data classification module, used for identifying scene initial primary data and scene initial secondary data from the image data to be identified according to the scene identification data; A data enhancement module, used to identify and enhance the initial main data of the scene according to the scene main data enhancement model to obtain scene enhanced main data; And, the data enhancement module is further used to identify and enhance the initial secondary data of the scene according to the scene secondary data enhancement model to obtain scene enhanced secondary data; The content recognition module is used to recognize the image content of the image data to be recognized according to the scene enhancement primary data and the scene enhancement secondary data, so as to obtain image content target recognition data.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Image scene classification method and device based on local feature saliency
CN112001399A
Image recognition method and device, equipment and storage medium
CN112906819A
Image feature enhancement method based on scene and target decoupling
CN118570102A
Scene understanding information generation method and device, equipment and medium
CN119251657A
Target identification method
CN119314149A