Image analysis method, device and equipment based on large model and storage medium
Through the large-model-based image analysis method, combined with image cleaning and external knowledge base, the problem of low image recognition accuracy is solved, and high-precision analysis of dynamic images is achieved.
Patent Information
- Application Number
- CN202510483665.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-08
AI Technical Summary
现有技术中图像识别方法存在精度低,无法有效识别动态图像的问题。
The image analysis method based on the big model is adopted to obtain the image input by the user and analyze the requirements, image cleaning and feature information extraction are carried out, and analysis is combined with external knowledge base and big model to optimize the final results.
It improves the accuracy and generalization ability of image analysis, and can effectively identify target objects and behavior patterns in dynamic images.
Smart Images

Figure CN120279285A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an image analysis method, device, equipment and storage medium based on a large model. Background Art
[0002] Image recognition refers to the technology of using a computer to process, analyze and understand images in order to identify various different patterns of targets and objects, and is a practical application of applying deep learning algorithms. In today's digital age, the demand for the recognition and analysis of image data is increasing day by day, and it is widely used in many fields such as security monitoring, medical image diagnosis, and autonomous driving.
[0003] Currently, SVM (Support Vector Machine) or CNN (Convolutional Neural Networks) is usually used to process images, and common objects or scenes in the images can be recognized. However, this image recognition method has the problems of low image recognition accuracy and inability to recognize and analyze dynamic images. Therefore, how to propose an image analysis method with high accuracy and capable of recognizing dynamic images has become a technical problem to be solved at present. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide an image analysis method, device, equipment and storage medium based on a large model, which can use an external knowledge base to assist image analysis and improve the accuracy of image analysis. The specific solutions are as follows:
[0005] In a first aspect, the present application provides an image analysis method based on a large model, including:
[0006] Obtain a to-be-analyzed image input by a target user based on an interactive interface and an image analysis requirement corresponding to the to-be-analyzed image, and clean the to-be-analyzed image to obtain a cleaned image corresponding to the to-be-analyzed image; wherein, the data type of the image analysis requirement includes text data and voice data, and the to-be-analyzed image includes a dynamic image;
[0007] Analyze the cleaned image to obtain feature information corresponding to the cleaned image; wherein, the feature information includes the texture feature of the cleaned image, the edge detection result of the cleaned image, the corner points and key points in the cleaned image, and the target object and the behavior pattern of the target object in the cleaned image;
[0008] Obtain the target scene category corresponding to the cleaned image, obtain the corresponding target knowledge data from the target knowledge base associated with the preset image analysis large model according to the target scene category, and analyze the cleaned image based on the target knowledge data, the feature information, the image analysis requirement, and the preset image analysis large model to obtain the corresponding initial analysis result;
[0009] Optimize the initial analysis result to obtain the corresponding target analysis result, and use the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time.
[0010] Optionally, the cleaning the image to be analyzed to obtain the cleaned image corresponding to the image to be analyzed includes:
[0011] Identify the noise type in the image to be analyzed. If the noise type in the image to be analyzed is Gaussian noise, use the non-local means filtering algorithm to perform noise reduction processing on the image to be analyzed to obtain the corresponding denoised image;
[0012] If the noise type in the image to be analyzed is salt-and-pepper noise, use the adaptive median filtering algorithm to perform noise reduction processing on the image to be analyzed to obtain the corresponding denoised image;
[0013] If the noise type in the image to be analyzed is motion blur noise, use the blind deconvolution algorithm to perform noise reduction processing on the image to be analyzed to obtain the corresponding denoised image;
[0014] Adjust the image brightness and image contrast of the denoised image to obtain the cleaned image corresponding to the image to be analyzed.
[0015] Optionally, the analyzing the cleaned image based on the target knowledge data, the feature information, the image analysis requirement, and the preset image analysis large model includes:
[0016] Adjust the model parameters of the preset image analysis large model according to the target scene category corresponding to the cleaned image to obtain the corresponding adjusted image analysis large model;
[0017] Analyze the cleaned image by using the target knowledge data, the feature information, the image analysis requirement, and the adjusted image analysis large model.
[0018] Optionally, the optimizing the initial analysis result to obtain the corresponding target analysis result includes:
[0019] Determine the target information corresponding to the image analysis requirement from each of the initial analysis results, determine the redundant information in each of the initial analysis results based on the target information, and eliminate the redundant information in each of the initial analysis results to obtain the corresponding first analysis result;
[0020] Determine the confidence level corresponding to each of the first analysis results, and screen out the second analysis results with a confidence level greater than the preset confidence threshold from each of the first analysis results;
[0021] Perform semantic optimization on each of the second analysis results, and convert each of the semantically optimized second analysis results into the target output format corresponding to the image analysis requirement to obtain the corresponding target analysis result.
[0022] Optionally, the image analysis method based on the large model further includes:
[0023] Set a target time interval, and obtain new knowledge data corresponding to the target knowledge data from the target data source based on the target time interval;
[0024] Determine whether there is data error in the new knowledge data. If there is data error in the new knowledge data, discard the new knowledge data;
[0025] If there is no data error in the new knowledge data, add the new knowledge data to the target knowledge base to update the target knowledge base.
[0026] Optionally, after using the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time, it further includes:
[0027] Obtain the feedback information input by the target user based on the interactive interface, adjust the target analysis result in real time according to the feedback information, and use the interactive interface to display the adjusted analysis result in real time.
[0028] Optionally, after using the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time, it further includes:
[0029] Determine the target file export format according to the actual application requirements of the target user, and generate a target file corresponding to the adjusted analysis result according to the target file export format;
[0030] Use the interactive interface to display the download link corresponding to the target file to the target user, so that the target user can export the target file using the download link.
[0031] In a second aspect, the present application provides an image analysis device based on a large model, including:
[0032] An image cleaning module, configured to obtain an image to be analyzed input by a target user based on an interactive interface and an image analysis requirement corresponding to the image to be analyzed, and clean the image to be analyzed to obtain a cleaned image corresponding to the image to be analyzed; wherein, the data type of the image analysis requirement includes text data and voice data;
[0033] A feature information acquisition module, configured to analyze the cleaned image to obtain feature information corresponding to the cleaned image; wherein, the feature information includes the texture feature of the cleaned image, the edge detection result of the cleaned image, the corner points and key points in the cleaned image, and the target object in the cleaned image and the behavior pattern of the target object;
[0034] An image analysis module, configured to obtain a target scene category corresponding to the cleaned image, obtain corresponding target knowledge data from a target knowledge base associated with a preset image analysis large model according to the target scene category, and analyze the cleaned image based on the target knowledge data, the feature information, the image analysis requirement, and the preset image analysis large model to obtain a corresponding initial analysis result;
[0035] An analysis result display module, configured to optimize the initial analysis result to obtain a corresponding target analysis result, and use the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time.
[0036] In a third aspect, the present application provides an electronic device, including:
[0037] A memory, configured to store a computer program;
[0038] A processor, configured to execute the computer program to implement the foregoing image analysis method based on a large model.
[0039] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, and when the computer program is executed by a processor, the foregoing image analysis method based on a large model is implemented.
[0040] In this application, first, the image to be analyzed input by the target user based on the interactive interface and the image analysis requirements corresponding to the image to be analyzed are obtained, and the image to be analyzed is cleaned to obtain the cleaned image corresponding to the image to be analyzed; wherein, the data types of the image analysis requirements include text data and voice data, and the image to be analyzed includes dynamic images. Then, the cleaned image is analyzed to obtain the feature information corresponding to the cleaned image; wherein, the feature information includes the texture feature of the cleaned image, the edge detection result of the cleaned image, the corner points and key points in the cleaned image, the target object in the cleaned image, and the behavior pattern of the target object. After that, the target scene category corresponding to the cleaned image is obtained, and the corresponding target knowledge data is obtained from the target knowledge base associated with the preset image analysis large model according to the target scene category. Based on the target knowledge data, the feature information, the image analysis requirements, and the preset image analysis large model, the cleaned image is analyzed to obtain the corresponding initial analysis result. Finally, the initial analysis result is optimized to obtain the corresponding target analysis result, and the image analysis progress corresponding to the cleaned image and the target analysis result are displayed to the target user in real time through the interactive interface. It can be seen that in this application, by deeply analyzing the image, the scene category corresponding to the image to be analyzed is obtained, and the data in the corresponding data knowledge base is called according to the scene category to assist the image analysis large model in image analysis, improving the accuracy of the analysis result; by obtaining the behavior feature information of the target object in the image to be analyzed, the motion feature of the target object in the dynamic image can be effectively analyzed, thus realizing the analysis of the dynamic image; by associating the external knowledge base with the large language model and using the data in the external knowledge base to assist image recognition, the generalization ability of image recognition is improved. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0042] Figure 1 Flowchart of an image analysis method based on a large model disclosed in this application;
[0043] Figure 2 Structural schematic diagram of an image analysis system based on a large model disclosed in this application;
[0044] Figure 3Schematic flow diagram of a specific image analysis method based on a large model disclosed in this application;
[0045] Figure 4 Schematic structural diagram of an image analysis device based on a large model disclosed in this application;
[0046] Figure 5 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] The current image recognition and analysis methods have problems such as low image recognition accuracy and inability to recognize and analyze dynamic images. For this reason, this application provides an image analysis method based on a large model, which can analyze the behavior patterns of objects in images and improve the accuracy of image analysis by using an external knowledge base to assist image analysis.
[0049] See Figure 1 As shown, an image analysis method based on a large model disclosed in an embodiment of the present invention includes:
[0050] Step S11: Obtain the image to be analyzed input by the target user based on the interactive interface and the image analysis requirement corresponding to the image to be analyzed, and clean the image to be analyzed to obtain the cleaned image corresponding to the image to be analyzed; wherein, the data types of the image analysis requirement include text data and voice data, and the image to be analyzed includes dynamic images.
[0051] In this embodiment, the image analysis method based on a large model is applied to an intelligent image recognition and analysis system, such as Figure 2 As shown, this system realizes efficient recognition and in-depth analysis of images by integrating a user interaction module, an image preprocessing module, a large model core module, a result optimization module, and a knowledge base module. Users can upload images or input image feature requirements through the interactive interface, and the system automatically completes tasks such as image recognition, feature extraction, and content analysis, and provides an optimized result output. This system can be widely applied to fields such as security monitoring, medical image diagnosis, and autonomous driving, significantly improving the efficiency and quality of image processing.
[0052] Among them, the user interaction module: is the core interface for the system to communicate with users, provides an intuitive and easy-to-use operation environment, supports multiple input methods, and meets the needs of different users.
[0053] Image preprocessing module: It is a key part of the system to preliminarily process the input image, aiming to optimize the image quality, extract useful information, and provide a basis for subsequent recognition and analysis.
[0054] Large model core module: It is the core part of the system. Based on advanced artificial intelligence large models, it realizes efficient recognition and in-depth analysis of images and provides high-quality results.
[0055] Result optimization module: It is used to further optimize the recognition and analysis results generated by the large model, improve the accuracy and usability of the results, and meet the actual needs of users.
[0056] Knowledge base module: It provides rich domain knowledge and data support for the system, improves the adaptability and accuracy of the system, and at the same time maintains the timeliness and dynamic update of knowledge.
[0057] In this embodiment, the overall process of analyzing the image is as Figure 3 shown. First, the user interaction module is responsible for the interaction between the user and the system: The user can upload images by dragging or selecting files, supporting common image formats such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (BitMap, a standard image file format), etc.; The user can also input the feature requirements (i.e., image analysis requirements) or descriptions of the input image, such as "recognize all vehicles in the image" or "analyze abnormal areas in medical images", etc. In addition, when the user uses the system with both hands busy or inconvenient to operate the interface, voice command input is supported. Specifically, the user uploads an image or inputs image feature requirements through the interactive interface to clarify the goals of recognition and analysis. The user can click the "Upload" button on the interface to select an image file stored locally, supporting multiple formats (such as JPEG, PNG, BMP, etc.); In addition, the user can also input specific image feature requirements in the text box, such as "recognize all vehicles in the image" or "detect lesion areas in medical images". For users who are not convenient to input using the keyboard, the system also provides a voice input function. The user can quickly express requirements through voice commands. After the user completes the input, click the "Start Recognition" button, and the system will automatically enter the next processing flow.
[0058] The image cleaning module is responsible for cleaning the image to be analyzed: First, denoising is performed. The noise in the image is removed through filtering algorithms (such as Gaussian filtering, median filtering) to improve the image clarity. Then, the illumination is adjusted. The illumination condition of the image is automatically detected and adjusted to enhance the contrast and brightness of the image, making it more suitable for subsequent processing. Correspondingly, the process of cleaning the image to be analyzed to obtain the cleaned image corresponding to the image to be analyzed may specifically include: identifying the type of noise in the image to be analyzed. If the type of noise in the image to be analyzed is Gaussian noise, the non-local means filtering algorithm is used to perform denoising processing on the image to be analyzed to obtain the corresponding denoised image; if the type of noise in the image to be analyzed is salt-and-pepper noise, the adaptive median filtering algorithm is used to perform denoising processing on the image to be analyzed to obtain the corresponding denoised image; if the type of noise in the image to be analyzed is motion blur noise, the blind deconvolution algorithm is used to perform denoising processing on the image to be analyzed to obtain the corresponding denoised image; adjusting the image brightness and image contrast of the denoised image to obtain the cleaned image corresponding to the image to be analyzed. By performing denoising processing on the image according to the type of image noise, the pertinence of the image cleaning process is ensured. By adjusting the brightness and contrast of the denoised image, the image can be made more easily recognizable.
[0059] Step S12: Analyze the cleaned image to obtain the feature information corresponding to the cleaned image; wherein, the feature information includes the texture feature of the cleaned image, the edge detection result of the cleaned image, the corner points and key points in the cleaned image, and the target object in the cleaned image and the behavior pattern of the target object.
[0060] In this embodiment, the process of analyzing the cleaned image is also responsible for the image preprocessing module in the system. This module first performs edge detection, uses an edge detection algorithm (such as the Canny algorithm) to extract the edge information in the image for subsequent object recognition and segmentation; secondly, analyzes the texture, extracts the texture pattern by analyzing the texture feature of the image for distinguishing different materials or regions; finally, extracts the key points, identifies the key points (such as corner points, feature points) in the image to provide a basis for image matching and three-dimensional reconstruction.
[0061] In addition, this embodiment can perform pixel-level segmentation on the image, label the category to which each pixel belongs, and achieve a fine understanding of the image content. Then, sentiment analysis is performed to identify the sentiment tendency in the image (such as facial expression analysis of a person), and provide analysis results related to sentiment. Finally, anomaly detection may be performed according to the scene requirements. In a specific field (such as medical imaging), abnormal regions or features in the image are identified to assist in diagnosis. By deeply analyzing the image to be recognized, the texture features, edge detection results, corners, and key points corresponding to the image to be recognized are extracted, and the features of the image to be analyzed are obtained from multiple angles, thereby achieving high-precision image recognition and maintaining a high accuracy even in complex scenarios.
[0062] Step S13: Obtain the target scene category corresponding to the cleaned image, obtain corresponding target knowledge data from the target knowledge base associated with the preset image analysis large model according to the target scene category, and analyze the cleaned image based on the target knowledge data, the feature information, the image analysis requirements, and the preset image analysis large model to obtain corresponding initial analysis results.
[0063] In this embodiment, the image preprocessing module is used to initially determine the scene category to which the image belongs (such as indoor, outdoor, medical imaging, etc.), providing a basis for subsequent domain-specific processing. Then, object detection is performed to identify the main objects in the image (such as people, vehicles, organs, etc.), and their positions and categories are labeled, providing basic information for in-depth analysis. After that, the large model core module uses the preset image analysis large model to identify the objects in the image, supporting multi-category and multi-instance recognition. Then, scene recognition is performed to classify and describe the overall scene of the image, and identify the key elements in the scene and their mutual relationships. Finally, behavior recognition is performed: in a video or dynamic image, identify the behavior patterns of the objects (such as pedestrians walking, vehicles driving, etc.); then, relevant knowledge data is obtained from the knowledge base according to the scene category corresponding to the image to be analyzed, so as to use the relevant knowledge data to perform in-depth analysis on the image subsequently.
[0064] In this embodiment, analyzing the cleaned image based on the target knowledge data, the feature information, the image analysis requirements, and the preset image analysis large model includes: adjusting the model parameters of the preset image analysis large model according to the target scene category corresponding to the cleaned image to obtain a corresponding adjusted image analysis large model; using the target knowledge data, the feature information, the image analysis requirements, and the adjusted image analysis large model to analyze the cleaned image; that is, this embodiment can fine-tune the large model according to the domain requirements input by the user, load the data set and parameters of a specific domain, and improve the adaptability and accuracy of the model in this domain.
[0065] It should be noted that the knowledge base module has multi-domain knowledge bases, covering professional knowledge and data sets in multiple fields such as security monitoring, medical imaging, and autonomous driving, providing rich background information for the model. It also stores a professional term library, storing professional terms and standard descriptions in each field to ensure the professionalism and accuracy of the generated results. Moreover, the knowledge base module provides a variety of predefined image recognition and analysis templates. Users can quickly select suitable templates according to their needs. In addition, users can also customize templates, supporting users to create and save custom templates according to their own needs to improve the flexibility of the system.
[0066] In addition, the knowledge base module also has a real-time update function: it can be updated according to user feedback. According to the user's feedback and annotation information, the content of the knowledge base is updated in real time to optimize the system performance. It can also perform data synchronization updates, regularly synchronize the latest domain knowledge and data from external data sources to ensure the timeliness and accuracy of the knowledge base. Correspondingly, the image analysis method in this embodiment further includes: setting a target time interval, and obtaining new knowledge data corresponding to the target knowledge data from the target data source based on the target time interval; determining whether there are data errors in the new knowledge data. If there are data errors in the new knowledge data, the new knowledge data is discarded; if there are no data errors in the new knowledge data, the new knowledge data is added to the target knowledge base to update the target knowledge base.
[0067] In this embodiment, the model parameters of the preset image analysis large model can be adjusted according to the scene category corresponding to the image to be analyzed, so as to improve the accuracy of image analysis. By deeply analyzing the image to be analyzed, feature information from multiple angles is obtained, improving the accuracy of image analysis; by identifying the target objects and corresponding behavior patterns in the image, the analysis operation of dynamic images is realized.
[0068] Step S14: Optimize the initial analysis result to obtain the corresponding target analysis result, and use the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time.
[0069] In this embodiment, the process of optimizing the initial analysis results is completed by the result optimization module. The result optimization module includes functions such as content optimization, result screening, and format adjustment. Correspondingly, the process of optimizing the initial analysis results to obtain the corresponding target analysis results may specifically include: determining the target information corresponding to the image analysis requirements from each initial analysis result, determining the redundant information in each initial analysis result based on the target information, and removing the redundant information in each initial analysis result to obtain the corresponding first analysis result; determining the confidence level corresponding to each first analysis result, and screening out the second analysis results with a confidence level greater than the preset confidence threshold from each first analysis result; performing semantic optimization on each second analysis result, and converting each second analysis result after semantic optimization into the target output format corresponding to the image analysis requirements to obtain the corresponding target analysis result.
[0070] The above process is completed by the result optimization module: First, information screening is performed. According to the user's needs, the most relevant and important information is screened out, and redundant content is removed. Then, confidence screening is carried out. According to the confidence level output by the model, high-confidence results are screened out to ensure the reliability of the results; Content optimization: First, semantic coherence optimization is carried out. The generated text description or report is semantically optimized to ensure that the content is coherent and logically clear. Then, the annotation is optimized. The annotation information in the image is adjusted, redundant annotations are removed, and the annotation position and category are optimized; In addition, this embodiment can also typeset the results. According to the output format specified by the user, the results are typeset and optimized, supporting multiple formats (such as image annotation, document report, data table), and custom templates can also be used. The user can select or customize the result display template according to the needs. The system generates results that meet the user's needs according to the template. After that, the system calls the user interaction module to display the recognition and analysis results in an intuitive way, including annotated images, generated reports or statistical charts, etc. And it supports exporting the results in multiple formats, such as image files (annotated JPEG, PNG), document formats (PDF, Word), or data formats such as CSV (Comma-Separated Values, that is, the comma-separated file format), JSON (JavaScript Object Notation, that is, the JS object notation), which is convenient for the user to use later; Specifically, the system optimizes the recognition and analysis results, screens out important information, and adjusts the output format: The system further optimizes the preliminary results generated by the large model to improve the accuracy and usability of the results. The system will perform semantic coherence optimization on the recognition results, adjust the annotation information and text description to ensure that the results are logically clear and naturally expressed. Then, the system screens out the most important information according to the user's needs, removes redundant content, and highlights the key results. In addition, the system typesets and adjusts the format of the results according to the output format specified by the user (such as image annotation, document report, data table, etc.) to meet the diverse needs of the user.
[0071] In this embodiment, after using the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time, the following steps are further included: obtaining the feedback information input by the target user based on the interactive interface, adjusting the target analysis result in real time according to the feedback information, and using the interactive interface to display the adjusted analysis result in real time. That is, the user can view the result in real time and use the tools provided by the interface (such as annotation and comment functions) to adjust the result or supplement the requirements. At this time, the system can adjust the result in real time according to the user's annotation and display the adjusted result to the user in real time. That is, after receiving the user input, the system starts the recognition and analysis process in real time, dynamically displays the progress and preliminary results on the interface, the user can view the result in real time, and use the tools provided by the interface (such as annotation and comment functions) to adjust the result or supplement the requirements.
[0072] In this embodiment, after using the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time, the following steps are further included: determining the target file export format according to the actual application requirements of the target user, and generating a target file corresponding to the adjusted analysis result according to the target file export format; using the interactive interface to display the download link corresponding to the target file to the target user, so that the target user can export the target file using the download link; specifically, after the user completes the adjustment of the result, the user can choose to export the result in multiple common formats, such as JPEG, PNG, PDF, Word, or CSV, etc. The user can select a suitable export format according to the actual needs, and the system will generate a corresponding file according to the selected format and provide a download link or save path. After the user downloads or saves the file, the image recognition and analysis results generated by the system can be applied to actual work or research to complete the entire process. By using the interactive interface to receive the user's annotation information, adjusting the analysis result in real time according to the user's annotation information, and displaying the adjusted result to the user in real time, the interactivity between the user and the system is provided, thereby improving the user experience; by outputting the analysis result in different formats according to the user's needs, the applicability of the image analysis result is improved.
[0073] It can be seen that through the integration of the knowledge base module and the model fine-tuning technology, this application significantly improves the adaptability of the system, makes the process of analyzing the to-be-recognized images in various fields more targeted, thereby improving the accuracy of image analysis; provides an interactive user interface, through which the user can upload images, input requirements, and view the recognition and analysis results in real time, enhancing the user's sense of interaction; by deeply analyzing the to-be-recognized images and obtaining the feature information from multiple angles of the images, the understanding of the images by the model can be deepened, and thus the accuracy of image analysis is improved.
[0074] See Figure 4As shown in the figure, an embodiment of the present invention discloses an image analysis device based on a large model, including:
[0075] An image cleaning module 11, configured to obtain an image to be analyzed input by a target user based on an interactive interface and an image analysis requirement corresponding to the image to be analyzed, and clean the image to be analyzed to obtain a cleaned image corresponding to the image to be analyzed; wherein, the data type of the image analysis requirement includes text data and voice data;
[0076] A feature information acquisition module 12, configured to analyze the cleaned image to obtain feature information corresponding to the cleaned image; wherein, the feature information includes the texture feature of the cleaned image, the edge detection result of the cleaned image, the corner points and key points in the cleaned image, and the target object in the cleaned image and the behavior pattern of the target object;
[0077] An image analysis module 13, configured to obtain a target scene category corresponding to the cleaned image, obtain corresponding target knowledge data from a target knowledge base associated with a preset image analysis large model according to the target scene category, and analyze the cleaned image based on the target knowledge data, the feature information, the image analysis requirement, and the preset image analysis large model to obtain a corresponding initial analysis result;
[0078] An analysis result display module 14, configured to optimize the initial analysis result to obtain a corresponding target analysis result, and use the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time.
[0079] It can be seen that through the integration of the knowledge base module and the model fine-tuning technology, the adaptability of the system is significantly improved, making the process of analyzing images to be recognized in various fields more targeted, thereby improving the accuracy of image analysis; providing an interactive user interface, through which users can upload images, input requirements, and view the recognition and analysis results in real time, enhancing the user's sense of interaction; through in-depth analysis of the images to be recognized, feature information from multiple angles of the images is obtained, so as to deepen the model's understanding of the images, and further improve the accuracy of image analysis.
[0080] In some specific embodiments, the image cleaning module 11 may specifically include:
[0081] A first image denoising unit, configured to identify the type of noise in the image to be analyzed. If the type of noise in the image to be analyzed is Gaussian noise, the non-local means filtering algorithm is used to perform denoising processing on the image to be analyzed to obtain a corresponding denoised image;
[0082] A second image denoising unit, configured to, if the noise type in the image to be analyzed is salt-and-pepper noise, perform denoising processing on the image to be analyzed by using an adaptive median filtering algorithm to obtain a corresponding denoised image;
[0083] A third image denoising unit, configured to, if the noise type in the image to be analyzed is motion blur noise, perform denoising processing on the image to be analyzed by using a blind deconvolution algorithm to obtain a corresponding denoised image;
[0084] A brightness adjustment unit, configured to adjust the image brightness and image contrast of the denoised image to obtain the cleaned image corresponding to the image to be analyzed.
[0085] In some specific embodiments, the image analysis module 13 may specifically include:
[0086] A parameter adjustment unit, configured to adjust the model parameters of the preset image analysis large model according to the target scene category corresponding to the cleaned image to obtain a corresponding adjusted image analysis large model;
[0087] An image analysis unit, configured to analyze the cleaned image by using the target knowledge data, the feature information, the image analysis requirement, and the adjusted image analysis large model.
[0088] In some specific embodiments, the analysis result display module 14 may specifically include:
[0089] An information elimination unit, configured to determine target information corresponding to the image analysis requirement from each of the initial analysis results, determine redundant information in each of the initial analysis results based on the target information, and eliminate the redundant information in each of the initial analysis results to obtain a corresponding first analysis result;
[0090] A result screening unit, configured to determine the confidence level corresponding to each of the first analysis results, and screen out second analysis results with a confidence level greater than a preset confidence threshold from each of the first analysis results;
[0091] A semantic optimization unit, configured to perform semantic optimization on each of the second analysis results, and convert each of the second analysis results after semantic optimization into a target output format corresponding to the image analysis requirement to obtain a corresponding target analysis result.
[0092] In some specific embodiments, the large model-based image analysis device further includes:
[0093] A data acquisition module, configured to set a target time interval and acquire new knowledge data corresponding to the target knowledge data from a target data source based on the target time interval;
[0094] A data discarding module, configured to determine whether there is data error in the new knowledge data, and discard the new knowledge data if there is data error in the new knowledge data;
[0095] A data adding module, configured to add the new knowledge data to the target knowledge base to update the target knowledge base if there is no data error in the new knowledge data.
[0096] In some specific embodiments, the analysis result display module 14 further includes:
[0097] A result adjustment unit, configured to acquire feedback information input by the target user based on the interactive interface, adjust the target analysis result in real time according to the feedback information, and display the corresponding adjusted analysis result in real time using the interactive interface.
[0098] In some specific embodiments, the analysis result display module 14 further includes:
[0099] A file format determination unit, configured to determine a target file export format according to the actual application requirements of the target user, and generate a target file corresponding to the adjusted analysis result according to the target file export format;
[0100] A link display unit, configured to display a download link corresponding to the target file to the target user using the interactive interface, so that the target user can export the target file using the download link.
[0101] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 5 which is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation on the scope of use of the present application.
[0102] Figure 5 This is a schematic structural diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the image analysis method based on a large model disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0103] In this embodiment, the power supply 23 is used to provide operating voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.
[0104] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be transient storage or permanent storage.
[0105] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the large model-based image analysis method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.
[0106] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the large model-based image analysis method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0107] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0108] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0109] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be disposed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0110] Finally, it should also be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0111] The technical solutions provided in this application have been introduced in detail above. Specific examples are used herein to illustrate the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. An image analysis method based on a large model, characterized in that, Including: Obtain the image to be analyzed input by the target user based on the interactive interface and the image analysis requirement corresponding to the image to be analyzed, and clean the image to be analyzed to obtain the cleaned image corresponding to the image to be analyzed; wherein, the data type of the image analysis requirement includes text data and voice data, and the image to be analyzed includes a dynamic image; Analyze the cleaned image to obtain the feature information corresponding to the cleaned image; wherein, the feature information includes the texture feature of the cleaned image, the edge detection result of the cleaned image, the corner points and key points in the cleaned image, and the target object and the behavior pattern of the target object in the cleaned image; Obtain the target scene category corresponding to the cleaned image, obtain the corresponding target knowledge data from the target knowledge base associated with the preset image analysis large model according to the target scene category, and analyze the cleaned image based on the target knowledge data, the feature information, the image analysis requirement and the preset image analysis large model to obtain the corresponding initial analysis result; Optimize the initial analysis result to obtain the corresponding target analysis result, and use the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user in real time.
2. The image analysis method based on a large model according to claim 1, wherein The cleaning the image to be analyzed to obtain the cleaned image corresponding to the image to be analyzed includes: Identify the noise type in the image to be analyzed. If the noise type in the image to be analyzed is Gaussian noise, use the non-local mean filtering algorithm to perform noise reduction processing on the image to be analyzed to obtain the corresponding denoised image; If the noise type in the image to be analyzed is salt-and-pepper noise, use the adaptive median filtering algorithm to perform noise reduction processing on the image to be analyzed to obtain the corresponding denoised image; If the noise type in the image to be analyzed is motion blur noise, use the blind deconvolution algorithm to perform noise reduction processing on the image to be analyzed to obtain the corresponding denoised image; Adjust the image brightness and image contrast of the denoised image to obtain the cleaned image corresponding to the image to be analyzed.
3. The image analysis method based on a large model according to claim 1, wherein The analyzing the cleaned image based on the target knowledge data, the feature information, the image analysis requirement and the preset image analysis large model includes: Adjust the model parameters of the preset image analysis large model according to the target scene category corresponding to the cleaned image to obtain the corresponding adjusted image analysis large model; Analyze the cleaned image by using the target knowledge data, the feature information, the image analysis requirement and the adjusted image analysis large model.
4. The image analysis method based on a large model according to claim 1, wherein The optimizing the initial analysis result to obtain the corresponding target analysis result includes: Determine the target information corresponding to the image analysis requirement from each of the initial analysis results, determine the redundant information in each of the initial analysis results based on the target information, and remove the redundant information in each of the initial analysis results to obtain the corresponding first analysis result; Determine the confidence level corresponding to each of the first analysis results, and screen out the second analysis results with a confidence level greater than the preset confidence threshold from each of the first analysis results; Perform semantic optimization on each of the second analysis results, and convert each of the second analysis results after semantic optimization into the target output format corresponding to the image analysis requirement, so as to obtain the corresponding target analysis results.
5. The method for image analysis based on a large model according to any one of claims 1 to 4, characterized in that It further includes: Set a target time interval, and obtain new knowledge data corresponding to the target knowledge data from a target data source based on the target time interval; Judge whether there is data error in the new knowledge data. If there is data error in the new knowledge data, discard the new knowledge data; If there is no data error in the new knowledge data, add the new knowledge data to the target knowledge base to update the target knowledge base.
6. The image analysis method based on a large model according to claim 1, wherein After using the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis results to the target user in real time, it further includes: Obtain the feedback information input by the target user based on the interactive interface, adjust the target analysis results in real time according to the feedback information, and use the interactive interface to display the adjusted analysis results in real time.
7. The method for image analysis based on a large model according to claim 6, wherein After using the interactive interface to display the image analysis progress corresponding to the cleaned image and the target analysis results to the target user in real time, it further includes: Determine the target file export format according to the actual application requirements of the target user, and generate a target file corresponding to the adjusted analysis results according to the target file export format; Use the interactive interface to display the download link corresponding to the target file to the target user, so that the target user can export the target file using the download link.
8. An image analysis device based on a large model, characterized in that, It includes: An image cleaning module, configured to obtain a to-be-analyzed image input by a target user based on an interactive interface and an image analysis requirement corresponding to the to-be-analyzed image, and clean the to-be-analyzed image to obtain a cleaned image corresponding to the to-be-analyzed image; wherein, the data type of the image analysis requirement includes text data and voice data; A feature information acquisition module, configured to analyze the cleaned image to obtain feature information corresponding to the cleaned image; wherein, the feature information includes the texture feature of the cleaned image, the edge detection result of the cleaned image, the corner points and key points in the cleaned image, and the target object in the cleaned image and the behavior pattern of the target object; An image analysis module, configured to obtain the target scene category corresponding to the cleaned image, obtain corresponding target knowledge data from a target knowledge base associated with a preset image analysis large model according to the target scene category, and analyze the cleaned image based on the target knowledge data, the feature information, the image analysis requirement and the preset image analysis large model to obtain corresponding initial analysis results; An analysis result display module is used to optimize the initial analysis result to obtain a corresponding target analysis result, and use the interactive interface to real-time display the image analysis progress corresponding to the cleaned image and the target analysis result to the target user.
9. An electronic device, characterized in that, It includes: A memory for storing computer programs; A processor for executing the computer program to implement the large model-based image analysis method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, For storing a computer program, which when executed by a processor implements the large model-based image analysis method according to any one of claims 1 to 7.