Model training method, image processing method, electronic device, storage medium and program product
By screening image groups with scene image intersections to construct training data and train the model, the problem of low scene image retrieval accuracy is solved, and the effective use of image scene features and the improvement of information processing accuracy are achieved.
Patent Information
- Application Number
- CN202510900334.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
In the prior art, the accuracy of scene image retrieval is low. Especially when the scene information in the image is limited, it is difficult to effectively use the limited scene information in the image for accurate retrieval.
By obtaining a first image group of multiple shooting locations, screening out a second image group with scene screen intersection, constructing training data and training the first model, using image scene features for information processing, using the DINOv2 sub-model to extract image features and using the triplet loss function for model training.
It improves the accuracy of scene image retrieval, enhances the model's ability to focus on image scene features, and improves the accuracy of information processing.
Smart Images

Figure CN120804357A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer, and particularly, to a model training method, an image processing method, an electronic device, a storage medium and a program product. BACKGROUND
[0002] In some tasks, there is a demand for scene image retrieval. However, in the related art, the accuracy of scene image retrieval is low. SUMMARY
[0003] Embodiments of the present disclosure provide a model training method, an image processing method, an electronic device, a storage medium and a program product to improve the accuracy of information processing based on image scene features, such as improving the accuracy of scene image retrieval.
[0004] In a first aspect, embodiments of the present disclosure provide a model training method, comprising:
[0005] obtaining a first image group of multiple shooting locations, wherein the first image group includes multiple first images, the first images located in the same first image group have the same shooting location, and the first images located in different first image groups have different shooting locations;
[0006] performing image screening on the first image group to obtain a second image group, wherein the second image group includes multiple second images, and the scene pictures between any two second images located in the same second image group have an intersection;
[0007] constructing training data based on the second image group, wherein the training data includes multiple query image samples, and a positive sample image set and a negative sample image set of the query image samples;
[0008] training a first model using the training data to obtain a trained first model, wherein the trained first model is used for information processing based on image scene features.
[0009] In a second aspect, embodiments of the present disclosure also provide an image processing method, comprising:
[0010] in response to a processing instruction for a third image, using a pre-trained first model to obtain information matching the image scene features of the third image based on the image scene features of the third image, wherein the pre-trained first model is trained based on image scene feature matching samples, and the image scene feature matching samples are at least part of the sample set;
[0011] outputting a processing result of the third image based on the matching information.
[0012] In a third aspect, the embodiments of the present disclosure further provide a model training apparatus, comprising:
[0013] an image group obtaining module, configured to obtain a first image group of a plurality of shooting locations, wherein the first image group comprises a plurality of first images, the first images in the same first image group have the same shooting location, and the first images in different first image groups have different shooting locations;
[0014] an image group screening module, configured to perform image screening on the first image group to obtain a second image group, wherein the second image group comprises a plurality of second images, and the scene pictures between any two second images in the same second image group have an intersection;
[0015] a training data constructing module, configured to construct training data based on the second image group, wherein the training data comprises a plurality of query image samples, and a positive sample image set and a negative sample image set of the query image samples;
[0016] a model training module, configured to train a first model based on the training data to obtain a trained first model, wherein the trained first model is used for information processing based on image scene features.
[0017] In a fourth aspect, the embodiments of the present disclosure further provide an image processing apparatus, comprising:
[0018] an information obtaining module, configured to, in response to a processing instruction for a third image, obtain information matching an image scene feature of the third image based on the image scene feature of the third image by using a pre-trained first model, wherein the pre-trained first model is trained based on image scene feature matching samples, and the image scene feature matching samples are at least part of samples in a sample set;
[0019] a result output module, configured to output a processing result of the third image based on the matching information.
[0020] In a fifth aspect, the embodiments of the present disclosure further provide an electronic device, comprising:
[0021] one or more processors;
[0022] a memory, configured to store one or more programs,
[0023] when the one or more programs are executed by the one or more processors, the one or more processors implement the model training method or the image processing method according to the embodiments of the present disclosure.
[0024] In a sixth aspect, the present disclosure also provides a computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the model training method or the image processing method according to the present disclosure.
[0025] In a seventh aspect, the present disclosure also provides a computer program product, when the computer program product is executed by a computer, causing the computer to implement the model training method or the image processing method according to the present disclosure.
[0026] The model training method, the image processing method, the electronic device, the storage medium and the program product provided by the present disclosure can improve the accuracy of the first model in information processing based on image scene features, and further improve the accuracy of the generated processing result, such as improving the accuracy of scene image retrieval. BRIEF DESCRIPTION OF DRAWINGS
[0027] The above and other features, advantages, and aspects of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0028] Figure 1 A flowchart of a model training method according to an embodiment of the present disclosure is shown in FIG. 1;
[0029] Figure 2 A structure diagram of a first model according to an embodiment of the present disclosure is shown in FIG. 2;
[0030] Figure 3 A structure diagram of an image localizability prediction model according to an embodiment of the present disclosure is shown in FIG. 3;
[0031] Figure 4 A flowchart of another model training method according to an embodiment of the present disclosure is shown in FIG. 4;
[0032] Figure 5 A training process diagram of a first model according to an embodiment of the present disclosure is shown in FIG. 5;
[0033] Figure 6 A flowchart of an image processing method according to an embodiment of the present disclosure is shown in FIG. 6;
[0034] Figure 7A schematic diagram of a process for determining position information of an image scene is provided for an embodiment of the present disclosure.
[0035] Figure 8 A structural block diagram of a model training apparatus is provided for an embodiment of the present disclosure.
[0036] Figure 9 A structural block diagram of an image processing apparatus is provided for an embodiment of the present disclosure.
[0037] Figure 10 A structural schematic diagram of an electronic device is provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0038] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure will be shown and described below, it is to be understood that the present disclosure can be embodied in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided as part of the disclosure to convey the principles and aspects of the present disclosure to practitioners in the art.
[0039] It should be understood that each of the steps in the method embodiments of the present disclosure can be performed in a different order and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.
[0040] The term “comprising” and variations thereof as used herein are used inclusively, i.e., “comprising but not limited to.” The term “based on” means “based, at least in part, on.” The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments.” Related terms are defined in the description that follows.
[0041] It should be noted that the terms “first”, “second”, and the like used in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0042] It should be noted that the terms “one”, “multiple”, and the like used in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that “one or more” should be understood unless otherwise explicitly stated in the context.
[0043] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0044] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0045] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.
[0046] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of pop-up window, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0047] It can be understood that the above notification and obtaining of user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0048] It can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the present technical solution should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0049] Figure 1 A flowchart of a model training method provided by the embodiments of the present disclosure is shown. The method can be executed by a model training device, which can be realized by software and / or hardware and can be configured in an electronic device, typically, in a computer, a mobile phone or a tablet computer. The model training method provided by the embodiments of the present disclosure is suitable for the scene of training a first model, such as the scene of training a first model for information processing based on image scene features.
[0050] In some tasks, such as the scene of determining the location of an accident, there is a need for scene image retrieval. Therefore, it is necessary to start from the image scene and perform scene image retrieval.
[0051] Corresponding to the scene image retrieval is the foreground image retrieval, the difference is that the foreground image retrieval targets the pictures consistent with the foreground of a certain image, while the scene image retrieval targets the pictures with the same background as a certain image. For example, the image is taken of "a sculpture in a certain park", then all the pictures of the same sculpture are the correct results of the foreground image retrieval of this image, and all the pictures of "the same park" are the correct results of the scene image retrieval of this image.
[0052] In actual application scenarios, in many cases, the image scene exists occlusion (such as building occlusion, etc.), and the shooting location and other related information in the scene may only account for a small part of the image. The existing image retrieval model is difficult to effectively utilize the limited scene information in the image, resulting in that the retrieved pictures may all be the pictures with the foreground content as the image main body, which greatly affects the effect of the image retrieval model.
[0053] In view of this, the embodiment of the present disclosure provides a model training method, which is specially for model training in the scene image retrieval scene, so that the trained model can pay attention to the limited scene information in the image, and improve the accuracy of the scene image retrieval.
[0054] As shown in Figure 1 The model training method provided by the embodiment can include:
[0055] S101, a first image group of a plurality of shooting locations is obtained, wherein the first image group includes a plurality of first images, the first images located in the same first image group have the same shooting location, and the first images located in different first image groups have different shooting locations.
[0056] The shooting location of the image can be the shooting place of the image or the shooting position of the image, that is, the place shot in the image picture. The first image group can be understood as an initially obtained image group. The first image group can include a plurality of first images. The first images located in the same first image group can have the same shooting location, and the first images located in different first image groups can have different shooting locations. The first image can be understood as an image contained in the first image group, such as a picture contained in the first image group.
[0057] Specifically, a plurality of first image groups corresponding to different shooting locations can be obtained, for example, a plurality of shooting locations are determined, and for each shooting location or part of the shooting locations in the plurality of shooting locations, the first image of the shooting location is obtained to form the first image group of the shooting location.
[0058] In the embodiment, the acquisition method of the first image group is not limited. For example, the first image group of each shooting location in the plurality of shooting locations can be obtained through image acquisition, or the first image group of the plurality of shooting locations can be obtained based on a pre-constructed image library.
[0059] In some examples, the first image set of the plurality of shooting locations can be obtained based on a constructed point of interest (POI) image library, i.e., the first image set is obtained from the POI library. The POI refers to a geographical object that can be abstracted as a point, such as a shopping mall and / or a store, etc. The POI image library generally records the POI number corresponding to the shooting location of each POI image, so that at least part of the POI images in the POI image library can be taken as the first images, the first images are grouped according to the POI number, and the obtained image set is taken as the first image set. Thus, a large number of POI images for training can be obtained without image acquisition, which can further improve the accuracy of the first model obtained by training based on image scene features for information processing, and reduce the human resources required for image acquisition.
[0060] S102, image screening is performed on the first image set to obtain a second image set, wherein the second image set includes a plurality of second images, and scene pictures between two second images in the same second image set have intersection.
[0061] The second image set can be understood as an image set obtained by screening the first image set. The second image set can filter out non-scene images in the first image set and / or part of the scene images whose scene pictures do not all have intersection. The second image can be understood as an image included in the second image set, and the scene pictures between two second images in the same second image set have intersection. The scene picture can be understood as a picture of an image scene, and the image scene can be understood as a scene or environment presented in the image. The intersection of the scene pictures can be understood as the intersection of the picture content of the scene pictures, i.e., the overlapping of the shooting angles of the scene pictures.
[0062] In this embodiment, considering that there can be non-scene images in the first image set, which do not contain scene information and interfere with the retrieval effect of the scene images, and even for the scene images, the images taken at the same shooting location can not contain scene intersection and cannot help the first model relying on image scene to train, and even affect the effect of the first model, the first image set can be screened to filter out non-scene images in the first image set and / or part of the scene images whose scene pictures do not all have intersection, so as to improve the accuracy of the first model obtained by subsequent training based on image scene features for information processing.
[0063] Specifically, for each obtained image group, multiple images with intersection of scene pictures between each other are obtained from the first image group as second images, so as to obtain a second image group corresponding to the first image group; or, image cleaning is performed on the first image group, non-scene images and at least part of non-scene images without intersection between each other are deleted from the first image group, and the first image group after image cleaning is determined as the second image group. Thus, multiple second image groups can be obtained, and the second images in different first image groups have different shooting locations. The second images can be screened by using a second model or other manners, which are not limited in the embodiment.
[0064] In S103, training data is constructed based on the second image groups, and the training data includes multiple image samples to be queried and a positive sample image set and a negative sample image set of the image samples to be queried.
[0065] The training data can be data used for model training. The training data includes multiple image samples to be queried and a positive sample image set and a negative sample image set of each image sample to be queried. The image sample to be queried can be understood as an image sample to be queried when the model is trained. The image to be queried can be an image to be queried for matching other images. The positive sample image set of the image sample to be queried can be understood as a set of positive sample images of the image sample to be queried, and the positive sample image set of the image sample to be queried includes multiple positive sample images of the image to be queried. The positive sample image of the image sample to be queried can be an image with the same shooting location as the image to be queried. The negative sample image set of the image sample to be queried can be understood as a set of negative sample images of the image sample to be queried, and the negative sample image set of the image sample to be queried includes multiple negative sample images of the image to be queried. The negative sample image of the image sample to be queried can be an image with a different shooting location from the image to be queried.
[0066] In the embodiment, after obtaining the second image groups, the training data of the first model can be constructed based on the second image groups, such as constructing multiple image samples to be queried, and at least part of the image samples to be queried have different shooting locations; for each image sample to be queried, multiple second images with the same shooting location as the image sample to be queried are obtained from one or more second image groups as positive sample images of the image to be queried, so as to obtain a positive sample image set of the image to be queried; and / or, multiple second images with different shooting locations from the image sample to be queried are obtained from one or more second image groups as negative sample images of the image to be queried, so as to obtain a negative sample image set of the image to be queried.
[0067] S104, training the first model by using the training data to obtain a trained first model, wherein the trained first model is used for information processing based on image scene features.
[0068] Specifically, the first model can be trained by using the constructed training data to obtain a trained first model, so as to subsequently perform information processing based on image scene features by using the trained first model.
[0069] After the training of the first model, the first model can be used for information processing based on image scene features. The image scene features can be understood as features possessed by the image scene. The information processing based on the image scene features can be, for example, extracting image scene features, image retrieval based on image scene features, or other information processing operations, which are not limited in the embodiment.
[0070] The model structure of the first model is not limited, and an exemplary model structure is shown in FIG. 1. Figure 2 As shown in FIG. 1, taking the first model used at least for extracting image scene features of an input image as an example, the first model can include an image feature extraction module and a feature compression module. The image feature extraction module can be used for extracting image features of the input image, and the feature compression module can be used for compressing the image features of the input image to obtain image scene features in the image features of the input image. The type of the image feature extraction module is not limited, and an exemplary image feature extraction module can be a self-supervised sub-model capable of image feature extraction, such as a DINOv2 sub-model. The DINOv2 sub-model is a self-supervised learning visual model capable of learning visual features from a large amount of unlabeled images. Moreover, using a ViT architecture, the DINOv2 sub-model can better capture long-distance dependencies in images through a self-attention mechanism, and can have excellent performance in semantic segmentation, image classification, depth estimation, and other tasks as a visual backbone model. The DINOv2 sub-model is suitable for extracting high-quality image features. In the embodiment, the DINOv2 is used as a backbone model to extract image features, and model training is performed on this basis, which can improve training efficiency and training effect.
[0071] The feature compression module can be a sub-model capable of extracting image scene features in the image features output by the image feature extraction module.
[0072] In training the first model, the constructed training data can be used to train one of the image feature extraction module and the feature compression module, or the constructed training data can be used to train both the image feature extraction module and the feature compression module. In some examples, a pre-trained or existing image feature extraction model (such as a DINOv2 model) can be used as the image feature extraction module in the first model, and the constructed training data can be used to train the feature compression module to be trained in the first model. For example, when the constructed training data is used to train the first model, the parameters of the image feature extraction module in the first model are frozen, and only the parameters of the feature compression module are trained to further improve the training speed of the first model.
[0073] In addition, the loss function when training the first model can be flexibly set as needed. For example, the first model can be trained using a triplet loss function. For example, for each query image sample, the image scene features can be extracted by the first model, and the triplet loss function can be used as the loss function of the first model for contrastive learning training to reduce the feature distance between the query image sample and the positive sample image, and to increase the feature distance between the query image sample and the negative sample image. The triplet loss function is specifically:
[0074] Loss(anchor, positive, negative) = max{d(anchor, positive) - d(anchor, negative) + margin, 0}
[0075] +margin, 0}
[0076] wherein anchor, positive and negative are the image scene features of the query image sample, the positive sample image and the negative sample image, respectively, margin is a hyperparameter, and the specific value of the hyperparameter is not limited. The hyperparameter can ensure that the distance difference between the positive sample image and the negative sample image is at least margin, thereby enhancing the distinguishing ability of the first model.
[0077] The model training method provided in this embodiment obtains a first image group of multiple shooting locations, wherein the first image group includes multiple first images, first images in the same first image group have the same shooting location, and first images in different first image groups have different shooting locations; performs image screening on the first image group to obtain a second image group, wherein the second image group includes multiple second images, and the scene scenes between two second images in the same second image group have an intersection; constructs training data based on the second image group, and the training data includes multiple query image samples and a positive sample image set and a negative sample image set for each query image sample; uses the training data to train a first model to obtain a trained first model, and the trained first model is used to perform information processing based on image scene features. This embodiment utilizes the above technical solution to construct training data based on multiple image groups whose scene scenes between two images have an intersection, and trains a first model that performs information processing based on image scene features. This can improve the accuracy of the first model's information processing based on image scene features, thereby improving the accuracy of scene image retrieval.
[0078] In some embodiments, the image screening of the first image group to obtain the second image group includes: using a pre-trained second model to perform scene picture detection on the first image in the first image group to determine the scene image in the first image group, in which the scene picture is presented; and selecting multiple second images from the scene images in the first image group to generate a second image group.
[0079] The second model may be a model for detecting whether a scene is present in an image. Exemplarily, the second model may be an image locatability prediction model, which may be used to predict whether an image is locatable, such as detecting whether an image contains a scene for determining a shooting location, and predicting that the image is locatable when the detection determines that the image contains a scene; and / or predicting that the image is not locatable when the detection determines that the image does not contain a scene.
[0080] Taking the second model as an image localizability prediction model as an example, the model structure of the image localizability prediction model is not limited. In some examples, such as Figure 3 As shown, the image localizability prediction model can include a contrastive language-image pre-training (CLIP) module and a multi-layer perceptron (MLP) module. The CLIP module can be used to extract semantic features of the image. The MLP module can be used to predict the localizability of the image based on the semantic features of the image.
[0081] Scene images can be understood as images containing scene images, such as images that are predicted to be locatable. For example, scene images can be taken outdoors, such as images of shopping malls, storefronts, natural scenery, parks, squares, and / or street scenes. Non-scene images can be understood as images that do not contain scene images, such as images that are predicted to be non-locatable. For example, non-scene images can include, but are not limited to, trademark images, advertising poster images, dish images, text-only images, and / or interior images.
[0082] Exemplarily, for each first image group or part of the first image group, each first image in the first image group can be input into a pre-trained second model, and scene picture detection can be performed on each image in the first image group through the second model to generate a detection result indicating whether the first image presents a scene picture; after the detection result is generated, the first image presenting the scene picture in the first image group can be obtained according to the detection result as the scene image in the first image group; thereafter, image matching can be performed on the scene images in the first image group, and based on the matching results generated by the image matching, multiple scene images in which the scene pictures of each pair have intersections are selected as second images, and a second image group containing the second image is generated.
[0083] In this embodiment, the method for performing image matching on the scene images in the first image group is not limited. Optionally, selecting multiple second images from the scene images in the first image group to generate the second image group includes: performing image matching on the scene images in the first image group and calculating the number of matching points between each pair of the scene images; selecting multiple scene images from the first image group as second images based on the number of matching points, and generating the second image group based on the second images, wherein the scene images between each pair of the multiple scene images all have an intersection.
[0084] Specifically, image matching can be performed between each scene image in the first image group, and the number of matching points between each scene image can be counted; based on this number of matching points, multiple scene images in which the scene scenes between each two intersect are obtained as second images; and a second image group containing the obtained second images is generated.
[0085] In this embodiment, the manner of image matching of the scene images is not limited. For example, different scene images can be matched based on a scale invariant feature transform (SIFT). The SIFT is mainly used for image feature extraction and matching, and can detect feature points of an image at different scales, match SIFT features of a to-be-identified image with SIFT features of a known target, and determine whether the to-be-identified image contains a target object. In this application, SIFT can be used to clean training data, so as to ensure that the cleaned images are taken at the same location and have intersecting scene images.
[0086] After the number of matching points between each two scene images in a first image group is determined, a plurality of scene images with intersecting scene images between each two scene images can be obtained based on the number of matching points as second images.
[0087] For example, a scene image with a number of matching points greater than a preset number of matching points can be determined as a scene image with intersecting scene images, a subset of scene images with intersecting scene images between each two scene images in the first image group can be determined based on this, and a subset of scene images with a number of contained scene images satisfying a preset condition can be obtained. The images in the subset of scene images are taken as second images, and the subset of scene images is taken as a second image group.
[0088] For another example, each scene image in the first image group can be clustered based on the number of matching points between each two scene images, and a plurality of clustering clusters are obtained. A clustering cluster with a number of contained scene images satisfying a preset condition is obtained, the images in the clustering cluster are taken as second images, and the clustering cluster is taken as a second image group.
[0089] The preset condition can be set as required, for example, the preset condition can be that the number of contained scene images is maximum, and / or the number of contained scene images is greater than a preset image number threshold, and the like.
[0090] In this embodiment, the scene image detection model trained in advance is used to screen scene images in the first image group, and a plurality of images with intersecting scene images between each two scene images are selected from the scene images in the first image group as second images in the second image group, which can further improve the accuracy of the selected second images, and further improve the training effect of the first model.
[0091] In some embodiments, the constructing training data based on the second image groups comprises: obtaining a query image sample of at least part of the second image groups; determining a positive sample image set of the query image sample based on the second images in the same second image group as the query image sample; and determining a negative sample image set of the query image sample based on the second images in different second image groups from the second image group of the query image sample.
[0092] For example, a query image sample of each second image group of at least part of the second image groups can be obtained; for each query image sample, at least part of the second images in the second image group corresponding to the query image sample are taken as positive sample images of the query image sample, thereby obtaining a positive sample image set of the query image sample; and at least part of the second images in the second image groups other than the second image group corresponding to the query image sample are taken as negative sample images of the query image sample, thereby obtaining a negative sample image set of the query image sample.
[0093] In some examples, the query image sample can be a scene occlusion image or a non-scene occlusion image. For example, at least part of the obtained query image samples can be scene occlusion images, and / or at least part of the obtained query image samples can be non-scene occlusion images.
[0094] For example, a scene occlusion image can be understood as an image in which a scene picture is occluded, such as an image in which a shooting object picture and a scene picture are presented, and at least part of the area of the scene picture is occluded by the shooting object picture. The shooting object picture can be understood as a picture of a shooting target in the image, such as a foreground picture in the image. A non-scene occlusion image can be understood as an image in which a scene picture is not occluded, such as an image in which only a scene picture is presented. The scene occlusion picture and the non-scene occlusion picture can be obtained in various ways.
[0095] Optionally, the obtaining of the query image sample of at least part of the second image groups comprises: obtaining at least one second image from the second image group as the query image sample of the second image group; and / or obtaining a scene occlusion image matching the shooting location corresponding to the second image group as the query image sample of the second image group, wherein the scene occlusion image presents a shooting object picture and a scene picture, and the shooting object picture causes occlusion to part of the area of the scene picture.
[0096] Exemplarily, for each of at least part of the second image groups, a second image can be randomly or in other manners obtained from the second image group as the to-be-queried image sample corresponding to the second image group. And / or, for one or more second image groups other than the at least part of the second image groups, a second image is randomly or in other manners obtained from the second image group, the second image is image-processed, a shooting object picture causing occlusion to a partial region of a scene picture is added in the second image, and a scene occlusion image obtained through the image processing is taken as the to-be-queried image sample corresponding to the second image group. Or, according to a shooting location corresponding to the second image group, at least one scene occlusion image shot at the shooting location is obtained in an image acquisition manner or in a manner of being obtained from a published and disclosed work, as the to-be-queried image sample corresponding to the second image group.
[0097] It should be noted that, for the manner of obtaining the scene occlusion image from the published and disclosed work, the scene occlusion image is obtained from the work which is authorized and agreed by the publisher and other related personnel for model training.
[0098] In the embodiment, the second image group is used to determine the positive sample image set and the negative sample image set of the to-be-queried image sample, which can reduce the calculation amount required for determining the positive sample image set and the negative sample image set, and further improve the training speed of the first model.
[0099] Figure 4 Another flowchart of a model training method provided by the embodiment of the present disclosure is provided. The scheme in the embodiment can be combined with one or more optional schemes in the above-mentioned embodiments. Optionally, the training of the first model based on the training data comprises: determining simple negative sample images and difficult negative sample images in the negative sample image set based on the currently trained first model; performing at least one round of training on the first model based on the to-be-queried image sample, the positive sample image set of the to-be-queried image sample, the simple negative sample images and the difficult negative sample images, and returning to perform the operation of determining the simple negative sample images and the difficult negative sample images in the negative sample image set based on the currently trained first model until the training of the first model is completed.
[0100] Correspondingly, as shown in Figure 4 the model training method provided by the embodiment can comprise:
[0101] S201, obtaining a first image group of multiple shooting locations, wherein the first image group comprises multiple first images, the first images located in the same first image group have the same shooting location, and the first images located in different first image groups have different shooting locations.
[0102] S202, image screening is performed on the first image group to obtain a second image group, wherein the second image group includes a plurality of second images, and scene pictures between any two second images in the same second image group have intersections.
[0103] S203, training data is constructed based on the second image group, and the training data includes a plurality of query image samples, and a positive sample image set and a negative sample image set of the query image samples.
[0104] S204, based on the first model currently trained, simple negative sample images and difficult negative sample images in the negative sample image set are determined.
[0105] The simple negative sample image of the query image sample can be a negative sample image that has a large difference with the scene picture of the query image sample and is easily correctly distinguished by the first model. The difficult negative sample image of the query image sample can be a negative sample image that has a small difference with the scene picture of the query image sample and is not easily correctly distinguished by the first model.
[0106] Specifically, the simple negative sample images and the difficult negative sample images in the negative sample image set of the query image sample can be determined based on the first model currently trained. For example, based on at least part of the modules in the first model currently trained, image features of the query image sample and the negative sample images of the query image sample are extracted respectively, and based on the image features, a negative sample image similar to the query image is obtained as the difficult negative sample image of the query image sample; and / or, at least part of the negative sample images in the negative sample image set of the query image except the difficult negative sample image are taken as the simple negative sample images of the query image sample.
[0107] In some embodiments, based on the first model currently trained, the simple negative sample images and the difficult negative sample images in the negative sample image set are determined, including: respectively extracting first image features of the query image sample and second image features of each negative sample image in the negative sample image set of the query image sample by using the first model currently trained; calculating similarity between the negative sample images of the query image sample and the query image sample according to the first image features and the second image features; and determining the simple negative sample images and the difficult negative sample images in the negative sample image set based on the similarity.
[0108] The first image feature can be understood as an image feature of the query image. The second image feature can be an image feature of a negative sample image of the query image. The first image feature and the second image feature can be an image feature of the whole corresponding image, or can be an image scene feature of the corresponding image, which is not limited in the embodiment. For example, when the difficult negative sample image of the query image sample is determined before the first model is trained for the first time, the first image feature and the second image feature can be an image feature of the whole corresponding image, which can be extracted by the image feature extraction module in the first model. When the difficult negative sample image of the query image sample is determined again after the first model is trained for at least one time, the first image feature and the second image feature can be an image scene feature of the corresponding image, which can be extracted by the first model after each round of training is performed, so as to further improve the accuracy of the determined difficult negative sample image, and then improve the training effect of the first model.
[0109] For example, when the difficult negative sample of the query image is determined for the first time before the first model is trained, the image feature extraction module in the first model can be used to extract the image feature of the query image sample as the first image feature, and the image feature extraction module in the first model can be used to extract the image feature of each negative sample image of the query image sample as the second image feature. When the difficult negative sample of the query image is determined again after the first model is trained for at least one time, the image scene feature of the query image sample can be extracted by the first model after the last round of training is performed as the first image feature, and the image scene feature of each negative sample image of the query image sample can be extracted by the first model after the last round of training is performed as the second image feature.
[0110] After the first image feature of the query image sample and the second image feature of each negative sample image of the query image sample are obtained, the similarity between the negative sample image of the query image sample and the query image sample can be calculated according to the first image feature of the query image and the second image feature of each negative sample image of the query image.
[0111] After the similarity between the negative sample image of the to-be-queried image sample and the to-be-queried image sample is obtained, a negative sample image satisfying a first similarity condition in similarity can be obtained from the negative sample set of the to-be-queried image sample, such as a negative sample image with a similarity greater than a first similarity threshold, or a first preset number of negative sample images in descending order of similarity are obtained as difficult negative sample images of the to-be-queried image; and / or, at least part of the negative sample images in the negative sample set of the to-be-queried image except the difficult negative sample images are taken as simple negative sample images of the to-be-queried image sample.
[0112] In addition, for the case that the proportion of the determined difficult negative sample image in the negative sample set of the to-be-queried image sample is small, the embodiment can also directly determine each negative sample image in the negative sample set of the to-be-queried image sample as a simple negative sample image of the to-be-queried image sample, without excluding the difficult negative sample image in the negative sample set, to further reduce the operation required to determine the simple negative sample of the to-be-queried image, and further improve the construction efficiency of the training data.
[0113] S205, using the to-be-queried image sample, the positive sample image set, the simple negative sample image and the difficult negative sample image of the to-be-queried image sample, at least one round of training is performed on the first model.
[0114] Specifically, one or more rounds of training can be performed on the first model using the to-be-queried image sample, the positive sample image in the positive sample image set of the to-be-queried image sample, the simple negative sample image of the to-be-queried image sample, and the difficult negative sample image of the to-be-queried image sample.
[0115] In the embodiment, when training the first model, whether to use the difficult negative sample image of the to-be-queried image sample for training or to use the simple negative sample image of the to-be-queried image sample for training in each round of training can be determined based on the use probability p of the difficult negative sample image of the to-be-queried image sample. For example, for each to-be-queried image sample, there is a probability p of using the difficult negative sample image of the to-be-queried image sample for training in each round of training, and there is a probability of 1-p of using the simple negative sample image of the to-be-queried image sample for training. The use probability p of the difficult negative sample can remain unchanged during the training of the first model; or it can be updated according to the training of the first model, which is not limited in the embodiment.
[0116] In some embodiments, the employing the to-be-queried image sample, and the positive sample image set, the simple negative sample image, and the difficult negative sample image of the to-be-queried image sample, and performing at least one round of training on the first model, comprises: employing the to-be-queried image sample, and the positive sample image set, the simple negative sample image, and the difficult negative sample image of the to-be-queried image sample, and performing at least one round of training on the first model according to the current use probability of the difficult negative sample image.
[0117] The current use probability can be understood as the current use probability of the difficult negative sample. Before the use probability of the difficult negative sample is updated, the current use probability can be the initial use probability of the difficult negative sample. After the use probability of the difficult negative sample is updated, the current use probability can be the updated use probability.
[0118] Specifically, the first model can be trained at least one round of training according to the current use probability of the difficult negative sample image, the to-be-queried image sample, and the positive sample image set, the simple negative sample image, and the difficult negative sample image of the to-be-queried image sample. For example, in each round of training, a positive sample image is obtained from the positive sample image set of the to-be-queried image sample as the current positive sample image of the to-be-queried image sample in this round of training, and according to the current use probability p of the difficult negative sample image, a simple negative sample image of the to-be-queried image sample or a difficult negative sample image of the to-be-queried image sample is obtained as the current negative sample image of the to-be-queried image sample in this round of training, and each to-be-queried image sample, the current positive sample image of each to-be-queried image sample, and the current negative sample image of each to-be-queried image sample are used to train the first model in this round of model training.
[0119] Figure 5 A first model training process diagram provided by the embodiment is shown in FIG. 1. Figure 5As shown, the first model can determine a to-be-queried image sample when performing each round of training, such as sampling a to-be-queried image sample from a plurality of to-be-queried image samples to determine the to-be-queried image sample participating in training in the current round. For each to-be-queried image sample determined, a current positive sample image of the to-be-queried image sample participating in training in the current round is sampled from the positive sample image set of the to-be-queried image sample; whether the to-be-queried image sample uses a difficult negative sample image in the current round of training is determined according to the current use probability p of the difficult negative sample image, if yes, a current negative sample image of the to-be-queried image sample participating in the current round of training is sampled from each difficult negative sample image of the to-be-queried image sample; if not, a current negative sample image of the to-be-queried image sample participating in the current round of training is sampled from the negative sample image set of the to-be-queried image sample. After obtaining the current positive sample image and the current negative sample image of each to-be-queried image sample, the first model can be trained in the current round by using each to-be-queried image sample, the current positive sample image of each to-be-queried image sample, and the current negative sample image of each to-be-queried image sample, the training loss is calculated, the model parameters are updated, and the next round of training of the first model is continued until the first model training is completed.
[0120] In this embodiment, the current use probability of the difficult negative sample image can also be updated, such as gradually increasing the current use probability of the difficult negative sample image to gradually optimize the information processing effect of the first model based on image scene features. In this case, the model training method provided by the embodiment can also include: gradually increasing the current use probability of the difficult negative sample image in the training process of the first model.
[0121] In some examples, the current use probability of the difficult negative sample image can be updated by the corresponding personnel according to the training of the first model. Thus, the current application can update the current use probability of the difficult negative sample image based on the probability update instruction of the corresponding personnel.
[0122] In some examples, the current use probability of the difficult negative sample image can be automatically updated. For example, after each round of training is completed, whether the current use probability of the difficult negative sample image needs to be increased can be determined based on the training loss of the first model on the difficult negative sample image. For example, after each round of training is completed, the training loss of the first model on the difficult negative sample image in the current round of training can be calculated, if the training loss is greater than or equal to a preset threshold, the current use probability of the difficult negative sample image is kept unchanged, and the next round of training of the first model is continued using the current use probability; and / or if the training loss is less than the preset threshold, the current use probability of the difficult negative sample image can be increased by a preset value to obtain a new current use probability, and the next round of training of the first model is continued based on the new current use probability.
[0123] It can be understood that the probability range of the current use probability of the difficult negative sample image can be set in advance, for example, the current use probability of the difficult negative sample image can be set to be within a probability range of [0.1, 0.7], and the initial value of the current use probability can be set to be the minimum value in the probability range; with the training of the first model, the current use probability of the difficult negative sample image can be gradually increased within the probability range, and when the current use probability of the difficult negative sample image is increased to the maximum value of the probability range, the current use probability of the difficult negative sample image can no longer be increased.
[0124] In S206, it is determined whether the first model is trained. If yes, S207 is performed; if no, S204 is performed.
[0125] In S207, the trained first model is obtained, wherein the trained first model is used for information processing based on image scene features.
[0126] In this embodiment, the difficult negative sample image of the to-be-queried image sample can be updated once after at least one round of training of each pair of first models, so as to improve the accuracy of the determined difficult negative sample image. The number of the at least one round of training can be set in advance, for example, k rounds of training of each pair of first models can be set to update the difficult negative sample image of the to-be-queried image sample once; or the number of the at least one round of training can be determined based on the latest training loss of the first model (such as the positive sample training loss and / or the negative sample training loss, etc.), which is not limited in this embodiment. Wherein, k is a positive integer, and the specific value of k is not limited, for example, k can be set to 5, 10 or 15, etc.
[0127] Specifically, after at least one round of training of each pair of first models, it can be determined whether the first model is trained, for example, it can be determined whether the preset training end condition of the first model is met at present, if yes, the training of the first model can be ended, and the first model obtained by the latest training is obtained as the trained first model; if no, S204 is performed to update the difficult negative sample image of the to-be-queried image sample and continue at least one round of training of the first model after the update is completed, until the first model is trained.
[0128] In this embodiment, the determination condition of whether the first model is trained can be set as needed. Taking k rounds of training as an example of the above at least one round of training, after k rounds of training of each pair of first models, the current difficult negative sample loss of the last round of training in this k rounds of training and the last difficult negative sample loss of the last round of training in the last k rounds of training can be obtained; the reduction ratio of the current difficult negative sample loss relative to the last difficult negative sample loss is calculated, if the reduction ratio is less than a preset proportion threshold, it is determined that the first model is trained; otherwise, it can be determined that the first model has not been trained.
[0129] The model training method provided in this embodiment can dynamically update the difficult negative sample image of the to-be-queried image sample when training the model, and can train the to-be-queried image sample by using a simple negative sample image in some training rounds and by using a difficult negative sample image in other training rounds, so that the training effect of the first model can be further improved.
[0130] Figure 6 A flowchart of an image processing method provided in this embodiment of the present disclosure is shown. The method can be executed by an image processing device, where the device can be implemented by software and / or hardware, and can be configured in an electronic device, typically, in a computer, a mobile phone, or a tablet computer. The image processing method provided in this embodiment of the present disclosure is suitable for an image processing scenario, such as a scene image retrieval scenario. Figure 6 As shown in the figure, the image processing method provided in this embodiment can include:
[0131] S301, in response to a processing instruction for a third image, a first model trained in advance is used to obtain information matched with the image scene feature of the third image based on the image scene feature of the third image, where the first model trained in advance is trained based on image scene feature matched samples, and the image scene feature matched samples are at least part of the samples in a sample set.
[0132] The processing instruction for an image can be an instruction for instructing information processing based on the image scene feature of the image, such as a scene image retrieval instruction and / or an associated information obtaining instruction, etc. The third image can be an image corresponding to the processing instruction.
[0133] The first model can be a model for performing information processing based on the image scene feature of an input image (such as the third image, etc.). The training method of the first model can be set as needed, such as the first model can be trained based on image scene matched samples, which can include at least part of the samples in the sample set of the first model. Optionally, the first model trained in advance is trained by using the model training method provided in this embodiment of the present disclosure.
[0134] The information matched with the image scene feature of the third image can be information obtained based on the image feature of the third image, such as an image whose image scene feature is matched with the image scene feature of the third image and / or associated information of the image scene of the third image, etc.
[0135] Exemplarily, in a case that the processing instruction for the third image is received, the third image can be input into the pre-trained first model, and the pre-trained first model performs information processing based on the image scene feature of the third image, for example, the pre-trained first model extracts the image scene feature of the third image, and / or based on the image scene feature of the third image, obtains information matched with the image scene feature of the third image, and the like.
[0136] S302, output the processing result of the third image based on the matched information.
[0137] The processing result of the third image can be understood as a processing result generated in response to the processing instruction for the third image, and the content included in the processing result is set as needed. Optionally, the processing result includes at least one of the following: a fourth image, the image scene feature of the fourth image matches the image scene feature of the third image; scene association information of the third image, the scene association information includes at least one of position information of the image scene, description information of the image scene, operation control associated with the image scene, or operation prompt information associated with the image scene.
[0138] The fourth image can be an image obtained by image retrieval, such as an image obtained by scene image retrieval. The image scene feature of the fourth image matches the image scene feature of the third image.
[0139] The scene association information of the third image can be understood as information associated with the image scene of the third image. Exemplarily, the scene association information can include but is not limited to position information of the image scene, description information of the image scene, operation control associated with the image scene, and / or operation prompt information associated with the image scene, and the like. The position information of the image scene can be used to indicate the position of the image scene, such as the geographic position information of the image scene. The description information of the image scene can be used to describe the image scene, such as the scene type of the image scene, the content in the image scene, and / or the scene characteristics of the image scene, and the like. The operation control associated with the image scene can be used to execute the triggering operation associated with the image scene. In some examples, the operation control can be used to trigger the generation of operation instructions associated with the image scene, and in response to the operation instructions, execute the operation associated with the image scene. Taking the safety rescue scene as an example, exemplarily, the operation control can be configured to click to navigate to the position requiring safety rescue. The operation prompt information associated with the image scene can be used to prompt the user to execute the operation associated with the image scene. Still taking the safety rescue scene as an example, exemplarily, the operation prompt information can be used to inquire when and what kind of rescue is needed, and the like.
[0140] In this embodiment, after obtaining the information matching the image scene feature of the third image, the processing result of the third image can be generated based on the matching information, and the processing result is output.
[0141] For example, the processing result of the third image includes a fourth image. Illustratively, the image scene feature of the third image can be extracted by the first model. According to the image scene feature of the third image and the image scene features of the images stored in the preset image library, the similarity between each image in the preset image library and the third image is calculated, and at least one image in the preset image library that satisfies the second similarity condition with the third image is obtained as the fourth image, such as an image in the preset image library that has a similarity greater than a second similarity threshold with the third image is obtained as the fourth image, or a second preset number of images in the order of similarity from large to small with the third image are obtained as the fourth image, and the like. After obtaining the fourth image matching the image scene feature of the third image, the processing result of the third image can be generated based on the fourth image, and the processing result is output for the user sending the processing instruction of the third image to view.
[0142] In this embodiment, the manner of generating the processing result of the third image based on the fourth image is not limited. For example, in the image retrieval scene, such as in the scene where the user wants to obtain an image matching the image scene feature of the third image, the processing result containing the fourth image and / or image information (such as image identification information, etc.) of the fourth image can be generated to facilitate the user to view the fourth image matching the image scene feature of the third image; in the scene association information determination scene, such as in the scene where the user wants to determine the position information of the image scene in the third image, one or more possible position information of the image scene in the third image can be determined based on the position information of the image scene in the at least one fourth image, and the processing result containing the position information of the third image can be generated to facilitate the user to view the possible position information of the image scene in the third image.
[0143] For example, the processing result of the third image includes the associated scene information of the third image. Illustratively, the image scene feature of the third image can be extracted by the first model. According to the image scene feature of the third image, the scene association information of the third image is obtained, such as the position information, the description information, the operation control associated with the image scene of the third image, and / or the operation prompt information associated with the image scene of the third image, and the like, and the processing result containing the scene association information of the third image can be generated and output for the user to view the scene association information or perform corresponding operations based on the scene association information.
[0144] Figure 7 A determination process diagram of the position information of the image scene provided in this embodiment is as follows:Figure 7 As shown, the process of determining the scene position information of the image scene in the query image (such as the third image) can be described as: inputting the query image into the first model, and obtaining the image scene features of the query image output by the first model; retrieving an image whose image scene matches the query image from a preset image library based on the image scene features; and determining the position information of the image scene in the query image based on the scene position information of the image scene in the matching image.
[0145] The image processing method provided in this embodiment responds to a query instruction for a third image by using a pre-trained first model to obtain information matching the image scene features of the third image based on the image scene features of the third image, wherein the pre-trained first model is trained based on a sample set matching the image scene features; and outputs a processing result for the third image based on the obtained matching information. This embodiment utilizes the above-mentioned technical solution, using the pre-trained first model to perform information processing on the input image based on the image scene features of the input image, obtain information matching the image scene features of the input image, and generate a processing result for the input image based on this matching information. This can improve the accuracy of information processing based on image scene features, thereby improving the accuracy of the generated processing results, such as improving the accuracy of scene image retrieval.
[0146] Figure 8 This is a structural block diagram of a model training device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and can be configured in an electronic device, typically a computer, a mobile phone or a tablet computer, and can train a first model by executing a model training method, such as training a first model for information processing based on image scene features. Figure 8 As shown, the model training device provided in this embodiment may include: an image group acquisition module 801, an image group screening module 802, a training data construction module 803 and a model training module 804, wherein:
[0147] An image group acquisition module 801 is configured to acquire a first image group of multiple shooting locations, wherein the first image group includes multiple first images, the first images in the same first image group have the same shooting location, and the first images in different first image groups have different shooting locations;
[0148] An image group screening module 802 is configured to screen the first image group to obtain a second image group, wherein the second image group includes a plurality of second images, and scene images between any two second images in the same second image group have an intersection;
[0149] The training data construction module 803 is configured to construct training data based on the second image group, and the training data comprises a plurality of query image samples and a positive sample image set and a negative sample image set of the query image sample.
[0150] The model training module 804 is configured to train the first model by using the training data, to obtain a trained first model, wherein the trained first model is configured to perform information processing based on image scene features.
[0151] The model training apparatus provided in the embodiment is configured to acquire a plurality of first image groups of shooting locations by using the image group acquisition module, wherein the first image group comprises a plurality of first images, the first images in the same first image group have the same shooting location, and the first images in different first image groups have different shooting locations; the image group filtering module is configured to perform image filtering on the first image group, to obtain a second image group, wherein the second image group comprises a plurality of second images, and the scene pictures between any two second images in the same second image group have intersections; the training data construction module is configured to construct training data based on the second image group, and the training data comprises a plurality of query image samples and a positive sample image set and a negative sample image set of each query image sample; and the model training module is configured to train the first model by using the training data, to obtain a trained first model, and the trained first model is configured to perform information processing based on image scene features. In the embodiment, the training data is constructed based on a plurality of image groups in which the scene pictures between any two contained images have intersections, and the first model for extracting image scene features is trained, so that the accuracy of information processing based on image scene features of the first model is improved, and the accuracy of scene image retrieval is further improved.
[0152] Optionally, the image group filtering module 802 comprises a scene image determination unit configured to perform scene picture detection on the first images in the first image group by using a pre-trained second model, to determine scene images in the first image group, and the scene images present scene pictures; and an image group generation unit configured to select a plurality of second images from the scene images in the first image group, to generate a second image group.
[0153] Optionally, the image group generation unit can be specifically configured to perform image matching on the scene images in the first image group, to calculate the number of matching points between any two scene images; and select a plurality of scene images from the first image group as second images according to the number of matching points, and generate a second image group based on the second images, wherein the scene pictures between any two of the plurality of scene images have intersections.
[0154] Optionally, the training data construction module 803 comprises: a sample acquisition unit, configured to acquire a to-be-inquired image sample of at least part of the second image group; and an image set determination unit, configured to determine, based on a second image located in the same second image group as the to-be-inquired image sample, a positive sample image set of the to-be-inquired image sample, and determine, based on a second image located in a different second image group from the to-be-inquired image sample, a negative sample image set of the to-be-inquired image sample.
[0155] Optionally, the sample acquisition unit is specifically configured to: acquire at least one second image from the second image group as the to-be-inquired image sample of the second image group; and / or acquire a scene occlusion image matching a shooting location corresponding to the second image group as the to-be-inquired image sample of the second image group, wherein the scene occlusion image presents a shooting object picture and a scene picture, and the shooting object picture causes occlusion to a part of the scene picture.
[0156] Optionally, the model training module 804 comprises: a negative sample image determination unit, configured to determine, based on the first model currently trained, simple negative sample images and difficult negative sample images in the negative sample image set; and a model training unit, configured to perform at least one round of training on the first model by using the to-be-inquired image sample, the positive sample image set of the to-be-inquired image sample, the simple negative sample images and the difficult negative sample images, and return to perform the operation of determining, based on the first model currently trained, the simple negative sample images and the difficult negative sample images in the negative sample image set until the first model training is completed.
[0157] Optionally, the negative sample image determination unit is specifically configured to: extract, by using the first model currently trained, a first image feature of the to-be-inquired image sample and a second image feature of each negative sample image in the negative sample image set of the to-be-inquired image sample, respectively; calculate a similarity between the negative sample image of the to-be-inquired image sample and the to-be-inquired image sample according to the first image feature and the second image feature; and determine, based on the similarity, the simple negative sample images and the difficult negative sample images in the negative sample image set.
[0158] Optionally, the model training unit is specifically configured to: perform at least one round of training on the first model by using the to-be-inquired image sample, the positive sample image set of the to-be-inquired image sample, the simple negative sample images and the difficult negative sample images according to the current use probability of the difficult negative sample image; and the model training apparatus can further comprise: a probability updating module, configured to gradually increase the current use probability of the difficult negative sample image in the training process of the first model.
[0159] Optionally, the first image group is obtained from a point of interest image library.
[0160] The model training apparatus provided by the embodiments of the present disclosure can execute the model training method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of executing the model training method. Technical details not described in detail in the embodiments can be referred to the model training method provided by any of the embodiments of the present disclosure.
[0161] Figure 9 A structural block diagram of an image processing apparatus provided by the embodiments of the present disclosure is provided. The apparatus can be implemented by software and / or hardware, and can be configured in an electronic device, typically, in a computer, a mobile phone or a tablet computer, and can perform image processing, such as scene image retrieval, by executing an image processing method. As shown in the figure, the image processing apparatus provided by the embodiments of the present disclosure can include a feature extraction module 901, an image query module 902 and a result output module 903, wherein, Figure 9
[0162] The information acquisition module 901 is configured to, in response to a processing instruction for a third image, acquire information matched with an image scene feature of the third image based on the image scene feature of the third image by using a first model trained in advance, wherein the first model trained in advance is trained based on image scene feature matched samples, and the image scene feature matched samples are at least part of the sample set;
[0163] The result output module 902 is configured to output a processing result of the third image based on the matched information.
[0164] The image processing apparatus provided by the embodiments of the present disclosure acquires information matched with an image scene feature of a third image based on the image scene feature of the third image by using a first model trained in advance in response to a query instruction for the third image by the information acquisition module, wherein the first model trained in advance is trained based on a sample set of image scene feature matching; and outputs a processing result of the third image based on the acquired matched information by the result output module. By using the above technical solution, the embodiments of the present disclosure acquire information matched with an image scene feature of an input image based on the image scene feature of the input image by using a first model trained in advance to perform information processing on the input image, and generate a processing result of the input image based on the matched information, which can improve the accuracy of information processing based on the image scene feature, and further improve the accuracy of the generated processing result, such as improving the accuracy of scene image retrieval.
[0165] Optionally, the first model trained in advance is trained by using the model training method provided by the embodiments of the present disclosure.
[0166] Optionally, the processing result comprises at least one of: a fourth image, an image scene feature of the fourth image matching the image scene feature of the third image; and scene association information of the third image, the scene association information comprising at least one of position information of an image scene, description information, an operation control associated with the image scene, or operation prompt information associated with the image scene.
[0167] The image processing apparatus provided by the embodiments of the present disclosure can execute the image processing method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and advantages of executing the image processing method. Technical details not described in the embodiments can be referred to the image processing method provided by any of the embodiments of the present disclosure.
[0168] Reference is made below in conjunction with Figure 10 which shows a structural schematic diagram of an electronic device (e.g., a server or a terminal device) 1000 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (e.g., a car navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. Figure 10 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0169] As shown in Figure 10 , the electronic device 1000 can include a processing apparatus (e.g., a central processor, a graphics processor, etc.) 1001, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 1002 or loaded into a random access memory (RAM) 1003 from a storage apparatus 1008. Various programs and data required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processing apparatus 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0170] Generally, the following apparatuses can be connected to the I / O interface 1005: an input apparatus 1006 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output apparatus 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage apparatus 1008 including, for example, a magnetic tape, a hard disk, and the like; and a communication apparatus 1009. The communication apparatus 1009 can allow the electronic device 1000 to communicate with other devices wirelessly or through wires to exchange data. Although Figure 10The electronic device 1000 is illustrated with various means, but it is understood that not all of the illustrated means need be present in every embodiment. A greater or lesser number of means can alternatively be implemented.
[0171] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from the network by the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0172] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium, for example, can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used or used in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take on various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, a radio frequency (RF) link, or the like, or any suitable combination thereof.
[0173] In some embodiments, the client, server, can communicate using any known or later developed network protocols such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., communication networks). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any then-current or later developed networks.
[0174] The computer readable medium described above can be included in the electronic device described above; or can exist separately, without being assembled into the electronic device.
[0175] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a first image group of multiple shooting locations, wherein the first image group includes multiple first images, the first images located in the same first image group have the same shooting location, and the first images located in different first image groups have different shooting locations; perform image screening on the first image group to obtain a second image group, wherein the second image group includes multiple second images, and the second images located in the same second image group have an intersection between each other in terms of scene pictures; construct training data based on the second image group, wherein the training data includes multiple query image samples, and a positive sample image set and a negative sample image set of the query image sample; train a first model using the training data to obtain a trained first model, wherein the trained first model is used for information processing based on image scene features. Or
[0176] In response to a processing instruction for a third image, a pre-trained first model is used to acquire information matching the image scene features of the third image based on the image scene features of the third image, wherein the pre-trained first model is trained based on image scene feature matching samples, and the image scene feature matching samples are at least part of the sample set; and a processing result of the third image is output based on the matching information.
[0177] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0178] The computer program instructions can also be loaded onto a computer or other programmable information processing apparatus to cause a series of operations to be performed on the computer or other programmable information processing apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable information processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0179] The units described in the embodiments of the present disclosure can be implemented by hardware, software, or a combination of hardware and software. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0180] The functions described in the above description herein can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0181] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include a lined- up electrical connection, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0182] The above description is only preferred embodiments of the present disclosure and the explanation of the applied technical principles. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and also covers other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0183] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0184] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A model training method, characterized in that: include: Acquire a first image group at a plurality of shooting locations, wherein the first image group includes a plurality of first images, the first images in the same first image group have the same shooting location, and the first images in different first image groups have different shooting locations; Performing image screening on the first image group to obtain a second image group, wherein the second image group includes a plurality of second images, and scene images between any two second images in the same second image group have an intersection; constructing training data based on the second image group, the training data including a plurality of image samples to be queried and a positive sample image set and a negative sample image set of the image samples to be queried; The training data is used to train the first model to obtain a trained first model, wherein the trained first model is used to perform information processing based on image scene features.
2. The method according to claim 1, characterized in that The performing image screening on the first image group to obtain a second image group includes: Using a pre-trained second model to perform scene picture detection on a first image in the first image group, and determining a scene image in the first image group, wherein the scene picture is presented in the scene image; A plurality of second images are selected from the scene images of the first image group to generate a second image group.
3. The method according to claim 2, characterized in that The step of selecting a plurality of second images from the scene images of the first image group to generate a second image group includes: performing image matching on the scene images in the first image group, and calculating the number of matching points between any two of the scene images; According to the number of matching points, multiple scene images are selected from the first image group as second images, and a second image group is generated based on the second images, wherein scene scenes between any two of the multiple scene images have intersections.
4. The method according to claim 1, wherein The constructing of training data based on the second image group includes: Obtaining at least part of the query image samples of the second image group; Based on a second image located in the same second image group as the image sample to be queried, a positive sample image set of the image sample to be queried is determined; and based on a second image located in a different second image group than the image sample to be queried, a negative sample image set of the image sample to be queried is determined.
5. The method according to claim 4, characterized in that The obtaining of at least part of the query image samples of the second image group includes: Acquire at least one second image from the second image group as a query image sample of the second image group; and / or Obtain a scene occlusion image that matches the shooting location corresponding to the second image group as a query image sample of the second image group, wherein the scene occlusion image presents a shooting object picture and a scene picture, and the shooting object picture causes occlusion of a partial area of the scene picture.
6. The method according to claim 1, characterized in that The step of training the first model using the training data includes: Determining, based on the currently trained first model, simple negative sample images and difficult negative sample images in the negative sample image set; The first model is trained for at least one round using the image sample to be queried, as well as the positive sample image set, simple negative sample images, and difficult negative sample images of the image sample to be queried, and the operation of determining the simple negative sample images and difficult negative sample images in the negative sample image set based on the currently trained first model is returned to be executed until the training of the first model is completed.
7. The method according to claim 6, characterized in that The determining, based on the currently trained first model, the simple negative sample images and the difficult negative sample images in the negative sample image set includes: respectively extracting a first image feature of the image sample to be queried and a second image feature of each negative sample image in a negative sample image set of the image sample to be queried using the currently trained first model; Calculating the similarity between the negative sample image of the query image sample and the query image sample based on the first image feature and the second image feature; Based on the similarity, simple negative sample images and difficult negative sample images in the negative sample image set are determined.
8. The method according to claim 6, characterized in that The method of using the query image sample and the positive sample image set, the simple negative sample image, and the difficult negative sample image of the query image sample to perform at least one round of training on the first model includes: According to the current usage probability of the difficult negative sample image, the first model is trained for at least one round using the query image sample and the positive sample image set, the simple negative sample image and the difficult negative sample image of the query image sample; The method further comprises: During the training process of the first model, the current usage probability of the difficult negative sample image is gradually increased.
9. The method according to any one of claims 1 to 8, characterized in that: The first image group is obtained from a point of interest image library.
10. An image processing method, characterized in that: include: In response to a processing instruction for a third image, using a pre-trained first model, based on image scene features of the third image, to obtain information matching the image scene features of the third image, wherein the pre-trained first model is trained based on samples matching the image scene features, and the samples matching the image scene features are at least part of the samples in the sample set; The processing result of the third image is output based on the matched information.
11. The method according to claim 10, characterized in that The pre-trained first model is trained using the model training method described in any one of claims 1-9.
12. The method according to claim 10, characterized in that The processing result includes at least one of the following: a fourth image, wherein image scene features of the fourth image match image scene features of the third image; The scene association information of the third image includes at least one of position information and description information of the image scene, operation controls associated with the image scene, or operation prompt information associated with the image scene.
13. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model training method described in any one of claims 1-9 or the image processing method described in any one of claims 10-12.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable a processor to implement the model training method described in any one of claims 1 to 9 or the image processing method described in any one of claims 10 to 12 when executed.
15. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the model training method described in any one of claims 1 to 9 or the image processing method described in any one of claims 10 to 12.
Citation Information
Patent Citations
Method and device for generating image recognition model
CN108898185A
Workshop scene recognition method and device, electronic equipment and storage medium
CN113516090A
Image feature extraction model training method, image retrieval method and related equipment
CN115131570A
Training method and device of automatic driving model, electronic equipment and storage medium
CN115713749A
Open set image scene matching method and device based on deep learning
CN116012841A