Image analysis method and device, equipment and storage medium
By acquiring and fusing object position and identity information in the image, generating a target image, and inputting a multimodal model for analysis, the problem of difficulty in targeted analysis of specific objects in the image in the prior art is solved, and a more targeted image description is achieved.
Patent Information
- Application Number
- CN202510164735.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for the prior art to analyze and describe specific objects in images in a targeted manner.
By obtaining the area mask image corresponding to the image to be analyzed, including the image position information and identity information of the object, and performing image fusion processing to generate a target image. Then, the target image is input into the pre-trained multimodal model for analysis, and the object's image description information, including identity information and behavioral information.
It realizes targeted analysis and description of specific objects in the image, and outputs more targeted and valuable image description information.
Smart Images

Figure CN120071453A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to an image analysis method, apparatus, device, and storage medium. Background Art
[0002] With the development of computer technology, multi-modal data such as pictures, texts, and videos are everywhere, and single-modal text or picture information can no longer meet people's needs. Especially in the field of monitoring, image analysis and processing of monitoring images can provide users with more valuable information. However, in related technologies, only the entire image can be analyzed and understood, and it is powerless when it comes to the detailed information of specific objects. Therefore, it is particularly necessary to provide a better image analysis and processing method. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide an image analysis method, apparatus, device, and storage medium, so as to solve the problem in related technologies that it is difficult to perform targeted analysis and description of specific objects in an image.
[0004] To solve the above technical problems, the embodiments of this application are implemented as follows: On the one hand, the embodiments of this application provide an image analysis method, including: Obtain a region mask image corresponding to the image to be analyzed; the region mask image includes the image position information and identity information of each object in the image to be analyzed; Perform image fusion processing on the image to be analyzed and the region mask image to obtain a target image; Input the target image into a pre-trained multi-modal model for image analysis to obtain image description information of at least one object in the image to be analyzed; the image description information includes identity information and behavior information.
[0005] On the other hand, the embodiments of this application provide an image analysis apparatus, including: A first acquisition module, configured to obtain a region mask image corresponding to the image to be analyzed; the region mask image includes the image position information and identity information of each object in the image to be analyzed; A first image fusion module, configured to perform image fusion processing on the image to be analyzed and the region mask image to obtain a target image; An image analysis module, configured to input the target image into a pre-trained multi-modal model for image analysis to obtain image description information of at least one object in the image to be analyzed; the image description information includes identity information and behavior information.
[0006] In another aspect, an embodiment of the present application provides an image analysis device, including a processor; and a memory arranged to store computer-executable instructions, the computer-executable instructions being configured to be executed by the processor, and the computer-executable instructions being executed by the processor to implement the above-mentioned image analysis method.
[0007] In another aspect, an embodiment of the present application provides a storage medium for storing computer-executable instructions, the computer-executable instructions implementing the above-mentioned image analysis method when executed by a processor.
[0008] In another aspect, an embodiment of the present application provides a computer program product, including a computer program, the computer program implementing the above-mentioned image analysis method when executed by a processor.
[0009] By adopting the technical solution of the embodiment of the present application, by obtaining a region mask image corresponding to the image to be analyzed and performing image fusion processing on the image to be analyzed and the region mask image to obtain a target image. Since the obtained region mask image includes the image position information and identity information of each object in the image to be analyzed, the target image carries the image position information and identity information of each object. Thus, by inputting the target image into a pre-trained multi-modal model for image analysis, image description information of at least one object in the image to be analyzed can be obtained, and the image description information includes identity information and behavior information. It can be seen that this technical solution can use the target image including the image position information and identity information of each object as the input data of the multi-modal model, which enables the multi-modal model to perform targeted analysis and description of specific objects in the image by combining the image position information and identity information of each object, so as to better align the identity information and behavior information of each object and output more targeted and valuable image description information. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0011] Figure 1 is a schematic flowchart of an image analysis method according to an embodiment of the present application; Figure 2 is a schematic diagram of the implementation principle of the training process of a feature fusion network according to an embodiment of the present application; Figure 3 is a schematic diagram of the implementation principle of the training process of a multi-modal model according to an embodiment of the present application; Figure 4 It is a schematic diagram of a target sample image according to an embodiment of the present application; Figure 5 It is a schematic flowchart of an image analysis method according to another embodiment of the present application; Figure 6 It is a schematic flowchart of an image analysis method according to another embodiment of the present application; Figure 7 It is a schematic block diagram of an image analysis device according to an embodiment of the present application; Figure 8 It is a schematic structural diagram of an image analysis device according to an embodiment of the present application. Specific embodiments
[0012] The embodiments of the present application provide an image analysis method, device, equipment and storage medium, which are used to solve the problem that it is difficult to perform targeted analysis and description on specific objects in an image in the related art.
[0013] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0014] Figure 1 It is a schematic flowchart of an image analysis method according to an embodiment of the present application. As Figure 1 shown, the method includes: S102, obtaining a region mask image corresponding to the image to be analyzed, where the region mask image includes image position information and identity information of each object in the image to be analyzed.
[0015] Optionally, the image to be analyzed can be an image in any color mode such as an RGB image, a grayscale image, a bitmap image, etc., and the embodiments of the present application do not limit this. In the region mask image, the image position information of each object can be represented by different position coordinates, and the identity information of each object can be represented by different pixel values. For example, for the two objects, Xiaoming and Xiaohong, in the image to be analyzed, if the pixel value 1 is used to represent that the identity information of an object is Xiaoming, and the pixel value 2 is used to represent that the identity information of an object is Xiaohong, then in the region mask image, the pixel values at the image positions where Xiaoming is located are all 1, the pixel values at the image positions where Xiaohong is located are all 2, and the pixel values of the background area can be set to 0. Thus, the region mask image can reflect both the image position information of each object and the identity information of each object.
[0016] It can be understood that the identity information mentioned in this embodiment includes information that can distinguish different objects, such as name, code number, nationality, etc. How to represent the specific identity information in the region mask image can be set in advance in practical applications, and the embodiments of the present application do not limit this. Among them, for the same kind of identity information of different objects (such as the names of different objects), different objects can be represented by setting pixel values of different sizes. For different kinds of identity information of different objects (such as the names and nationalities of different objects), different objects can be represented by setting pixel values of different types or different value ranges, and are represented in the region mask image through identification values. The identification value is a set of pixel values corresponding to different identity information of the same object. For example, digital-form pixel values can be used to represent objects with different names, and letter-form pixel values can be used to represent the above-mentioned objects with different nationalities. For example, the name of the object is Xiaoming, and the nationality is A. The identity information of Xiaoming is represented by the pixel value 1, and the nationality A is represented by the pixel value A. Then, in the region mask image, the identity information of the object can be represented by the identification value (1, A), that is, Xiaoming of nationality A.
[0017] S104, perform image fusion processing on the image to be analyzed and the region mask image to obtain a target image.
[0018] Among them, when the image fusion processing methods are different, the number of image channels of the target image is different. In the embodiments of the present application, an appropriate image fusion processing method can be selected according to the requirements of the multimodal model for the number of image channels to obtain a target image with the desired number of image channels.
[0019] S106, input the target image into a pre-trained multimodal model for image analysis to obtain image description information of at least one object in the image to be analyzed. The image description information includes identity information and behavior information.
[0020] Optionally, the image analysis includes: identifying each object in at least one object from a target image according to the identity information of at least one object, and analyzing the behavior information of each object based on the target image.
[0021] By adopting the technical solution of the embodiment of the present application, by obtaining a region mask image corresponding to an image to be analyzed, and performing image fusion processing on the image to be analyzed and the region mask image to obtain a target image. Since the obtained region mask image includes the image position information and identity information of each object in the image to be analyzed, the target image carries the image position information and identity information of each object. Thus, inputting the target image into a pre-trained multi-modal model for image analysis can obtain image description information of at least one object in the image to be analyzed, and the image description information includes identity information and behavior information. It can be seen that this technical solution can use the target image including the image position information and identity information of each object as the input data of the multi-modal model, which enables the multi-modal model to perform targeted analysis and description of specific objects in the image by combining the image position information and identity information of each object, so as to better align the identity information and behavior information of each object and output more targeted and valuable image description information.
[0022] In one embodiment, before obtaining the region mask image corresponding to the image to be analyzed (i.e., S102), the image to be analyzed can be input into a feature extraction network for feature extraction to obtain the image position information and identity information of each object in the image to be analyzed, and thus, according to the image position information and identity information of each object, a region mask image corresponding to the image to be analyzed is generated.
[0023] The feature extraction network can be any network with a feature extraction function, such as a ReID (Person Re-identification) network, and the embodiment of the present application does not limit this.
[0024] In this embodiment, by performing feature extraction on the image to be analyzed to obtain the image position information and identity information of each object, and thus generating a corresponding region mask image, the region mask image includes the image position information and identity information of each object in the image to be analyzed, providing a data basis for subsequent image analysis and facilitating more targeted analysis and description of the image.
[0025] In one embodiment, to perform image fusion processing on the image to be analyzed and the region mask image to obtain a target image (i.e., S104), it can be executed as: performing image stitching processing on the image to be analyzed and the region mask image to obtain a target image.
[0026] It can be understood that the region mask image is similar to a binary image. Optionally, when the image to be analyzed is a three-channel image (such as an RGB image), after image stitching processing, the obtained target image is a four-channel image.
[0027] In this embodiment, by performing image stitching processing on the image to be analyzed and the region mask image to obtain the target image, the region of each object is directly represented at the pixel level through the mask in the direct stitching method. This not only enables the target image, which is the input data of the multimodal model, to carry the image position information and identity information of each object, facilitating more targeted analysis and description of the target image by the multimodal model, but also makes the number of image channels of the target image meet the requirements of the multimodal model's image channels, ensuring that the multimodal model can accurately analyze and process the target image.
[0028] In one embodiment, performing image fusion processing on the image to be analyzed and the region mask image to obtain the target image (i.e., S104) can be executed as: inputting the image to be analyzed and the region mask image into a feature fusion network for feature fusion to obtain the target image.
[0029] Among them, the number of image channels of the target image is the same as that of the image to be analyzed. The feature fusion network can be a network capable of performing fusion processing on feature-level information, such as a convolutional network, and the embodiments of the present application do not limit this.
[0030] In this embodiment, when the adopted feature fusion network is different, the input data of the feature fusion network may also be different. In one case, the input data of the feature fusion network is two images, namely the above-mentioned image to be analyzed and the region mask image. In this case, the feature fusion network will first perform image stitching processing on these two images to obtain a stitched image, and then perform feature fusion on the stitched image to obtain the target image. In another case, the input data of the feature fusion network is one image. Then, before inputting the image to be analyzed and the region mask image into the feature fusion network, it is necessary to first perform image stitching processing on these two images to obtain a stitched image, and then input this stitched image into the feature fusion network for feature fusion to obtain the target image.
[0031] In this embodiment, by inputting the image to be analyzed and the region mask image into the feature fusion network for feature fusion to obtain the target image, not only enables the target image, which is the input data of the multimodal model, to carry the image position information and identity information of each object, facilitating more targeted analysis and description of the target image by the multimodal model, but also makes the number of image channels of the target image meet the requirements of the multimodal model's image channels, without changing the existing model structure, having stronger adaptability, and ensuring that the multimodal model can accurately analyze and process the target image.
[0032] In one embodiment, a feature fusion network that meets the requirements of the embodiments of the present application can be pre-trained using sample data. It can be easily seen from the previous embodiment that the feature fusion network ultimately performs feature fusion processing on a stitched image. The steps of image stitching can be performed outside the feature fusion network or within the feature fusion network. Based on this, the feature fusion network can be trained according to the following steps A1 - A2, where the steps of image stitching are performed outside the feature fusion network, as Figure 2 shown, which shows the specific training process of the feature fusion network, including: Step A1: Obtain a plurality of sample stitched images and the corresponding sample fusion images for each sample stitched image.
[0033] Among them, before obtaining a plurality of sample stitched images, a plurality of second sample images can be first obtained; secondly, for each second sample image, the second sample image is input into the feature extraction network for feature extraction to obtain the sample image position information and sample identity information of each sample object in the second sample image, and according to the sample image position information and sample identity information of each sample object, a sample region mask image corresponding to the second sample image is generated; thirdly, image stitching processing is performed on the second sample image and the sample region mask image to obtain a sample stitched image.
[0034] Step A2: Input the sample stitched images and the sample fusion images into the feature fusion network to be trained for iterative training to obtain the trained feature fusion network.
[0035] In one embodiment, step A2 can be performed as the following steps A21 - A22: Step A21: For each sample stitched image, perform feature fusion processing on the sample stitched image through the feature fusion network to be trained to obtain the predicted fusion image corresponding to the sample stitched image.
[0036] Step A22: According to the sample fusion image and the predicted fusion image, iteratively adjust the network parameters of the feature fusion network until the iterative termination condition is met.
[0037] Optionally, the network parameters can be iteratively adjusted according to the difference between the sample fusion image and the predicted fusion image, as well as the loss function corresponding to the feature fusion network, until the iterative termination condition is met. The iterative termination condition can include reaching a preset number of iterations and / or the loss function converging.
[0038] In this embodiment, by training the feature fusion network, a model foundation is provided for quickly and accurately obtaining the target image based on the image to be analyzed and the region mask image.
[0039] In one embodiment, a multi-modal model can be trained through the following steps B1 - B2, as Figure 3 shown, which shows the specific training process of the multi-modal model, including: Step B1, obtain the target sample image and the sample data set corresponding to the target sample image.
[0040] Among them, the target sample image includes the sample image position information and sample identity information of each sample object in the target sample image. The sample data set includes the image description information of each sample object. The image description information includes sample identity information and sample behavior information.
[0041] Exemplarily, in the case of the target sample image as Figure 4 shown, the image description information of each sample object may include: Xiaohong, wearing a red and black plaid shirt, is using a mobile phone; Zhang San, wearing a blue denim shirt, is holding a laptop; Li Si, wearing a plaid shirt, is looking at a mobile phone; Wang Er, wearing a pink shirt, is holding a laptop; Xiaomei, wearing a blue dress, is holding a laptop. It can be understood that Figure 4 the one shown is a grayscale image, so the clothing features of each object are not reflected.
[0042] Optionally, the image description information of each sample object can also be in the form of questions and answers. Continuing with the Figure 4 target sample image shown, the image description information in the form of questions and answers can include: Question: What is each person in the picture doing? Answer: There are five people in the picture. From left to right, they are: Xiaohong, wearing a red and black plaid shirt, is using a mobile phone; Zhang San, wearing a blue denim shirt, is holding a laptop; Li Si, wearing a plaid shirt, is looking at a mobile phone; Wang Er, wearing a pink shirt, is holding a laptop; Xiaomei, wearing a blue dress, is holding a laptop. Question: Who is the person wearing a red and black plaid shirt in the picture? Answer: Xiaohong. Question: Where is Xiaohong in the picture? Answer: Xiaohong is sitting on the far left of the sofa. Question: What is Li Si doing? Answer: Li Si is looking at a mobile phone.
[0043] Step B2, input the target sample image and the sample data set into the multi-modal model to be trained for iterative training to obtain the trained multi-modal model.
[0044] In one embodiment, step B2 can be executed as the following steps B21 - B23: Step B21, through the multi-modal model to be trained, perform image feature extraction processing on the target sample image to obtain the image feature information of the target sample image.
[0045] Among them, the image feature information may include the sample image position information and sample identity information of each sample object. In addition, the image feature information may further include the visual information of each sample object, such as the clothing information, appearance information, hairstyle information, etc. of each sample object, which are visually observed information.
[0046] Step B22: Generate predicted image description information for at least one sample object in the target sample image according to the image feature information.
[0047] Among them, the predicted image description information may include predicted identity information and predicted behavior information. It can be understood that the predicted image description information may further include predicted visual information.
[0048] Step B23: Iteratively adjust the model parameters of the multimodal model according to the predicted image description information and the sample data set until the iteration termination condition is satisfied.
[0049] Optionally, the model parameters can be iteratively adjusted according to the difference between the predicted image description information and the sample data set, as well as the loss function corresponding to the multimodal model, until the iteration termination condition is satisfied. The iteration termination condition may include reaching a preset number of iterations, and / or the loss function converges.
[0050] In this embodiment, by training a multimodal model, it provides a model basis for quickly and accurately combining the image position information and identity information of each object in the input image to perform targeted analysis and description of the specific objects in the image, so as to better align the identity information and behavior information of each object and output more targeted and valuable image description information. The trained multimodal model can already very well understand the image position information and identity information of each object and can correspond well. In this way, the multimodal model can have very rich identity information when generating image description information. For example, previously the multimodal model could only output: A man is watching TV. Now it can output: Zhang San is watching TV.
[0051] In one embodiment, before obtaining the target sample image and the sample data set corresponding to the target sample image (i.e., step B1), the target sample image can be obtained based on the first sample image through the following steps C1 - C4.
[0052] Step C1: Obtain a plurality of first sample images.
[0053] Step C2: For each first sample image, input the first sample image into a feature extraction network for feature extraction to obtain the sample image position information and sample identity information of each sample object in the first sample image.
[0054] Among them, the feature extraction network can be any network with feature extraction capabilities, such as a ReID network, and the embodiments of the present application do not limit this.
[0055] Step C3: Generate a sample region mask image corresponding to the first sample image according to the sample image position information and sample identity information of each sample object.
[0056] Step C4: Perform image fusion processing on the first sample image and the sample region mask image to obtain a target sample image.
[0057] Among them, the specific implementation manner of obtaining the target sample image through image fusion processing is similar to the specific implementation manner of obtaining the target image through image fusion processing in the above embodiments, and will not be elaborated here.
[0058] In one embodiment, before obtaining the region mask image corresponding to the image to be analyzed (i.e., S102), the following steps D1 - D2 can be executed: Step D1: Receive a request for analyzing the behavior information of the target object. The request for analyzing the behavior information includes the identity information of the target object.
[0059] For example, if the target object is Xiaohong, the request for analyzing the behavior information of the target object can be what color clothes Xiaohong is wearing? What is Xiaohong doing? And so on.
[0060] Step D2: Based on the request for analyzing the behavior information, obtain an image containing the target object as the image to be analyzed.
[0061] In this embodiment, inputting the target image into a pre - trained multi - modal model for image analysis to obtain the image description information of at least one object in the image to be analyzed (i.e., S106) can be performed as: inputting the identity information of the target object and the target image into the multi - modal model for image analysis to obtain the image description information of the target object.
[0062] Among them, image analysis includes: identifying the target object from the target image according to the identity information of the target object, so as to analyze the behavior information of the target object based on the target image.
[0063] In this embodiment, the multi - modal model realizes the effect of targeted search and analysis in the image according to the identity information. This overcomes the previous difficulty of being unable to find out what a person has done. For example, by inputting: Please help me check what Zhang San has done today? The multi - modal model can extract all the image data containing Zhang San and analyze the corresponding behavior information, giving the user a very comprehensive summary.
[0064] Since traditional multimodal models cannot identify the specific location of people in the image, let alone identity information, they can only analyze and understand the entire image, and are powerless when it comes to specific areas or detailed information. Even with location and identity information, the format of location + name, such as (Zhang San: [100, 32, 564, 343]), is difficult for traditional multimodal models to understand, so multimodal models still cannot match specific location information with people's identities. Coordinates are a set of discrete values that represent the boundaries of the area, and the model needs to be calculated and reasoned to be associated with specific image areas. If it is a dense scene (such as multiple people standing or occluded), the alignment of coordinates and visual features may have obvious deviations. To this end, this technical solution focuses on allowing multimodal models to better understand the relationship between identity and location. This technical solution can more efficiently integrate location information into the multimodal model by embedding the mask into the original image (rather than simply inputting coordinates). The mask directly marks the location of the object at the pixel level, and the regional information is presented in the form of spatial distribution. The multimodal model can directly perceive the shape and distribution of the mask area through operations such as convolution. Each pixel value of the mask corresponds to a specific position in the original image. The multimodal model can easily associate the mask value with the visual information of the corresponding area during the visual feature extraction stage. The pixel value of the mask area provides clear object information (such as encoding 1 to represent object 1), which can directly guide the multimodal model to focus on the target object area. At the same time, for multimodal tasks, such as generating text descriptions or answering questions, the mask can guide the multimodal model to focus on specific areas and combine visual features to avoid unnecessary background interference.
[0065] In addition, the image analysis method provided in the embodiment of the present application can be applied to a variety of image analysis scenarios. For example, it can be applied in real-time monitoring scenarios, so that the message prompt in the monitoring can no longer be "a person walked through the door", but can be changed to "Zhang San walked through the door", which is more valuable. And when there are multiple characters in the monitoring, it can also provide richer guidance. For example, when you see that everyone else is playing and only Xiaohong is slumped on the sofa, you can remind "Xiaohong is slumped on the sofa, looks very tired, a little lonely, do you need a phone call to comfort her?" You can even link other devices to dial directly, which was not possible before when there was no identity information. It can be applied in company monitoring scenarios, such as "Help me check when Zhang Gong came to the company today? What clothes are you wearing?" and "Help me see where Zhang Gong went today and what he did?" The multimodal model can focus on specific people based on identity information and generate targeted descriptions and understandings.
[0066] Figure 5 is a schematic flow chart of an image analysis method according to another embodiment of the present application, such as Figure 5 As shown, the method includes: S501. Receive a request for analyzing the behavior information of a target object, where the request for analyzing the behavior information includes the identity information of the target object.
[0067] S502. Based on the request for analyzing the behavior information, obtain an image including the target object as the image to be analyzed.
[0068] S503. Input the image to be analyzed into a feature extraction network for feature extraction to obtain the image position information and identity information of each object in the image to be analyzed.
[0069] S504. Generate a region mask image corresponding to the image to be analyzed according to the image position information and identity information of each object.
[0070] S505. Perform image fusion processing on the image to be analyzed and the region mask image to obtain a target image.
[0071] Optionally, image stitching processing can be performed on the image to be analyzed and the region mask image to obtain a target image. Or, input the image to be analyzed and the region mask image into a feature fusion network for feature fusion to obtain a target image; the number of image channels of the target image is the same as that of the image to be analyzed.
[0072] S506. Input the identity information of the target object and the target image into a multi-modal model for image analysis to obtain image description information of the target object.
[0073] Among them, the image description information includes identity information and behavior information. The image analysis includes: identifying the target object from the target image according to the identity information of the target object; analyzing the behavior information of the target object based on the target image.
[0074] The specific processes of the above S501 - S506 have been described in detail in the above embodiments and will not be elaborated here.
[0075] By adopting the technical solution of the embodiment of the present application, when a request for analyzing the behavior information of a target object is received, an image to be analyzed including the target object is acquired, and a corresponding region mask image is obtained. The image to be analyzed and the region mask image are subjected to image fusion processing to obtain a target image. Since the region mask image includes the image position information and identity information of each object in the image to be analyzed, the target image carries the image position information and identity information of each object. Thus, by inputting the identity information of the target object and the target image into a multimodal model for image analysis, image description information of the target object can be obtained, and the image description information includes identity information and behavior information. It can be seen that this technical solution can use the target image including the image position information and identity information of each object and the identity information of the target object as the input data of the multimodal model, which enables the multimodal model to perform targeted analysis and description of the target object in the image by combining the image position information and identity information of each object, so as to better align the identity information and behavior information of the target object and output more targeted and valuable image description information.
[0076] Figure 6 is a schematic flowchart of an image analysis method according to another embodiment of the present application, as Figure 6 shown, the method includes: S601, input the image to be analyzed into a feature extraction network for feature extraction to obtain the image position information and identity information of each object in the image to be analyzed.
[0077] S602, generate a region mask image corresponding to the image to be analyzed according to the image position information and identity information of each object.
[0078] S603, perform image fusion processing on the image to be analyzed and the region mask image to obtain a target image.
[0079] S604, input the target image into a pre-trained multimodal model for image analysis to obtain image description information of at least one object in the image to be analyzed.
[0080] Wherein, the image description information includes identity information and behavior information.
[0081] The specific processes of S601 - S604 above have been described in detail in the above embodiments and will not be elaborated here.
[0082] Adopting the technical solution of the embodiment of the present application, by obtaining the region mask image corresponding to the image to be analyzed and performing image fusion processing on the image to be analyzed and the region mask image to obtain the target image. Since the obtained region mask image includes the image position information and identity information of each object in the image to be analyzed, the target image carries the image position information and identity information of each object. Thus, inputting the target image into a pre-trained multimodal model for image analysis can obtain the image description information of at least one object in the image to be analyzed, and the image description information includes identity information and behavior information. It can be seen that this technical solution can use the target image including the image position information and identity information of each object as the input data of the multimodal model, which enables the multimodal model to perform targeted analysis and description of specific objects in the image by combining the image position information and identity information of each object, so as to better align the identity information and behavior information of each object and output more targeted and valuable image description information.
[0083] In summary, specific embodiments of the present subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing may be advantageous.
[0084] The above is the image analysis method provided by the embodiment of the present application. Based on the same idea, the embodiment of the present application also provides an image analysis device.
[0085] Figure 7 is a schematic block diagram of an image analysis device according to an embodiment of the present application. Please refer to Figure 7 , the image analysis device may include: The first acquisition module 710 is configured to acquire the region mask image corresponding to the image to be analyzed; the region mask image includes the image position information and identity information of each object in the image to be analyzed; The first image fusion module 720 is configured to perform image fusion processing on the image to be analyzed and the region mask image to obtain the target image; The image analysis module 730 is configured to input the target image into a pre-trained multimodal model for image analysis to obtain the image description information of at least one object in the image to be analyzed; the image description information includes identity information and behavior information.
[0086] In one embodiment, the image fusion module 720 includes: The image splicing unit is configured to perform image splicing processing on the image to be analyzed and the region mask image to obtain the target image.
[0087] In one embodiment, the image fusion module 720 includes: A feature fusion unit, configured to input the image to be analyzed and the region mask image into a feature fusion network for feature fusion to obtain a target image; the number of image channels of the target image is the same as that of the image to be analyzed.
[0088] In one embodiment, the image analysis device further includes: A second acquisition module, configured to acquire a target sample image and a sample data set corresponding to the target sample image; the target sample image includes sample image position information and sample identity information of each sample object in the target sample image; the sample data set includes image description information of each sample object; the image description information includes sample identity information and sample behavior information; A model training module, configured to input the target sample image and the sample data set into a multi-modal model to be trained for iterative training to obtain a trained multi-modal model.
[0089] In one embodiment, the image analysis device further includes: A third acquisition module, configured to acquire a plurality of first sample images before acquiring the target sample image and the sample data set corresponding to the target sample image; A feature extraction module, configured to input each first sample image into a feature extraction network for feature extraction to obtain sample image position information and sample identity information of each sample object in the first sample image; A generation module, configured to generate a sample region mask image corresponding to the first sample image according to the sample image position information and sample identity information of each sample object; A second image fusion module, configured to perform image fusion processing on the first sample image and the sample region mask image to obtain a target sample image.
[0090] In one embodiment, the model training module includes: An image feature extraction unit, configured to perform image feature extraction processing on the target sample image through the multi-modal model to be trained to obtain image feature information of the target sample image; the image feature information includes sample image position information and sample identity information of each sample object; A generation unit, configured to generate predicted image description information of at least one sample object in the target sample image according to the image feature information; An iterative training unit, configured to iteratively adjust the model parameters of the multi-modal model according to the predicted image description information and the sample data set until an iterative termination condition is met.
[0091] In one embodiment, the image analysis device further includes: A receiving module, configured to receive a request for analyzing behavior information of a target object before obtaining a region mask image corresponding to an image to be analyzed; the request for analyzing behavior information includes identity information of the target object. A fourth obtaining module, configured to obtain, based on the request for analyzing behavior information, an image including the target object as the image to be analyzed.
[0092] In one embodiment, the image analysis module 730 includes: An image analysis unit, configured to input the identity information of the target object and the target image into a multi-modal model for image analysis to obtain image description information of the target object. The image analysis includes: identifying the target object from the target image according to the identity information of the target object; analyzing the behavior information of the target object based on the target image.
[0093] By adopting the technical solution of the embodiment of the present application, by obtaining a region mask image corresponding to the image to be analyzed, and performing image fusion processing on the image to be analyzed and the region mask image to obtain a target image. Since the obtained region mask image includes the image position information and identity information of each object in the image to be analyzed, the target image carries the image position information and identity information of each object. Thus, inputting the target image into a pre-trained multi-modal model for image analysis can obtain image description information of at least one object in the image to be analyzed, and the image description information includes identity information and behavior information. It can be seen that this technical solution can use the target image including the image position information and identity information of each object as the input data of the multi-modal model, which enables the multi-modal model to perform targeted analysis and description of specific objects in the image in combination with the image position information and identity information of each object, so as to better align the identity information and behavior information of each object and output more targeted and valuable image description information.
[0094] Those skilled in the art should understand that Figure 7 the image analysis device in can be used to implement the image analysis method described above, and the detailed description thereof should be similar to that described in the method part above. To avoid redundancy, it will not be described in detail here.
[0095] Based on the same idea, the embodiment of the present application further provides an image analysis device, such as Figure 8As shown. Image analysis devices can vary significantly due to differences in configuration or performance, and may include one or more processors 801 and a memory 802. One or more application programs or data may be stored in the memory 802. Among them, the memory 802 can be for transient storage or persistent storage. The application programs stored in the memory 802 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions for the image analysis device. Further, the processor 801 can be set to communicate with the memory 802 and execute a series of computer-executable instructions in the memory 802 on the image analysis device. The image analysis device may also include one or more power supplies 803, one or more wired or wireless network interfaces 804, one or more input / output interfaces 805, and one or more keyboards 806.
[0096] Specifically, in this embodiment, the image analysis device includes a memory and one or more application programs, where one or more application programs are stored in the memory, and one or more application programs may include one or more modules, and each module may include a series of computer-executable instructions for the image analysis device, and is configured to be executed by one or more processors. The one or more application programs include the following computer-executable instructions: Obtain a region mask image corresponding to the image to be analyzed; the region mask image includes the image position information and identity information of each object in the image to be analyzed; Perform image fusion processing on the image to be analyzed and the region mask image to obtain a target image; Input the target image into a pre-trained multi-modal model for image analysis to obtain image description information of at least one object in the image to be analyzed; the image description information includes identity information and behavior information.
[0097] Adopting the technical solution of the embodiment of the present application, by obtaining the region mask image corresponding to the image to be analyzed and performing image fusion processing on the image to be analyzed and the region mask image to obtain the target image. Since the obtained region mask image includes the image position information and identity information of each object in the image to be analyzed, the target image carries the image position information and identity information of each object. Thus, by inputting the target image into a pre-trained multimodal model for image analysis, the image description information of at least one object in the image to be analyzed can be obtained, and the image description information includes identity information and behavior information. It can be seen that this technical solution can use the target image including the image position information and identity information of each object as the input data of the multimodal model, which enables the multimodal model to perform targeted analysis and description of the specific objects in the image by combining the image position information and identity information of each object, so as to better align the identity information and behavior information of each object and output more targeted and valuable image description information.
[0098] The embodiment of the present application also proposes a storage medium, which stores one or more computer programs. The one or more computer programs include computer executable instructions. When the computer executable instructions are executed by an electronic device including multiple application programs, they can enable the electronic device to execute each process of the above-mentioned image analysis method embodiment and achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0099] The embodiment of the present application also proposes a computer program product, which includes a computer program. When the computer program is executed by a processor, it realizes each process of the above-mentioned image analysis method embodiment and achieves the same technical effect. To avoid repetition, it will not be elaborated here.
[0100] The system, device, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or entity, or by a product with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0101] For the convenience of description, when describing the above device, it is divided into various units according to functions and described separately. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0102] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0103] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0104] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0106] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0107] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0108] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0109] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0110] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0111] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0112] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An image analysis method, characterized in that: include: Obtain a regional mask image corresponding to the image to be analyzed; The regional mask image includes image position information and identity information of each object in the image to be analyzed; Performing image fusion processing on the image to be analyzed and the regional mask image to obtain a target image; The target image is input into a pre-trained multimodal model for image analysis to obtain image description information of at least one object in the image to be analyzed; the image description information includes identity information and behavior information.
2. The method according to claim 1, characterized in that The step of performing image fusion processing on the image to be analyzed and the regional mask image to obtain a target image includes: The image to be analyzed and the regional mask image are subjected to image stitching processing to obtain the target image.
3. The method according to claim 1, characterized in that The step of performing image fusion processing on the image to be analyzed and the regional mask image to obtain a target image includes: The image to be analyzed and the regional mask image are input into a feature fusion network for feature fusion to obtain the target image; the target image has the same number of image channels as the image to be analyzed.
4. The method according to claim 1, characterized in that: The method further comprises: Acquire a target sample image and a sample data set corresponding to the target sample image; the target sample image includes sample image position information and sample identity information of each sample object in the target sample image; the sample data set includes image description information of each sample object; the image description information includes the sample identity information and sample behavior information; The target sample image and the sample data set are input into the multimodal model to be trained for iterative training to obtain a trained multimodal model.
5. The method according to claim 4, characterized in that Before acquiring the target sample image and the sample data set corresponding to the target sample image, the method further includes: acquiring a plurality of first sample images; For each of the first sample images, inputting the first sample image into a feature extraction network to perform feature extraction, and obtaining sample image position information and sample identity information of each sample object in the first sample image; generating a sample region mask image corresponding to the first sample image according to the sample image position information and the sample identity information of each sample object; Performing image fusion processing on the first sample image and the sample region mask image to obtain the target sample image.
6. The method according to claim 5, characterized in that The step of inputting the target sample image and the sample data set into the multimodal model to be trained for iterative training to obtain the trained multimodal model comprises: By using the multimodal model to be trained, performing image feature extraction processing on the target sample image to obtain image feature information of the target sample image; the image feature information includes the sample image position information and the sample identity information of each sample object; generating predicted image description information of at least one sample object in the target sample image according to the image feature information; According to the predicted image description information and the sample data set, the model parameters of the multimodal model are iteratively adjusted until an iteration termination condition is met.
7. The method according to claim 1, characterized in that Before obtaining the regional mask image corresponding to the image to be analyzed, the method further includes: Receiving a behavior information analysis request for a target object; the behavior information analysis request includes identity information of the target object; Based on the behavior information analysis request, an image containing the target object is acquired as the image to be analyzed.
8. The method according to claim 7, characterized in that The step of inputting the target image into a pre-trained multimodal model for image analysis to obtain image description information of at least one object in the image to be analyzed includes: Inputting the identity information of the target object and the target image into the multimodal model for image analysis to obtain the image description information of the target object; The image analysis includes: identifying the target object from the target image according to the identity information of the target object; and analyzing the behavior information of the target object based on the target image.
9. An image analysis device, characterized in that: include: A first acquisition module is used to acquire a regional mask image corresponding to the image to be analyzed; The regional mask image includes image position information and identity information of each object in the image to be analyzed; A first image fusion module, used for performing image fusion processing on the image to be analyzed and the regional mask image to obtain a target image; The image analysis module is used to input the target image into a pre-trained multimodal model for image analysis to obtain image description information of at least one object in the image to be analyzed; the image description information includes identity information and behavior information.
10. An image analysis device, characterized in that: include: processor; as well as A memory arranged to store computer executable instructions, wherein the computer executable instructions are configured to be executed by the processor, and the computer executable instructions are executed by the processor to implement the image analysis method according to any one of claims 1 to 8.
11. A storage medium, characterized in that: The storage medium is used to store computer executable instructions, and when the computer executable instructions are executed by a processor, the image analysis method according to any one of claims 1 to 8 is implemented.
12. A computer program product, characterized in that The invention comprises a computer program, which implements the image analysis method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Patent Citations
Multi-modal model training method, device and equipment and readable storage medium
CN116561570A
Pentograph model training method and device, equipment and storage medium
CN117173504A
Data set analysis method and system, electronic equipment and medium
CN117633454A
Student classroom attention analysis system based on facial multi-modal emotion fusion
CN119251886A
Complex dynamic environment-oriented image multi-modal fusion method and system
CN119339201A
Cited By
Image analysis method, apparatus and device, and storage medium and computer program product
WO2026170980A1