Model training method, apparatus, and device based on scene clustering, and storage medium

By clustering and model optimization of multiple target scenarios, the problem of poor recognition effect of a single algorithm model in multiple large-variant scenarios is solved, and more efficient scene recognition is achieved.

WO2025113058A1PCT designated stage expired Publication Date: 2025-06-05E SURFING VISION TECHNOLOGY CO LTD

Patent Information

Application Number
PCT/CN2024/128416
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-10-30
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When the number of scenes is relatively large and the differences between scenes are relatively large, the feature performance ability of a single algorithm model is limited and cannot adapt to a large number of scenes, resulting in poor scene recognition effect.

Method used

By obtaining the initial images of each target scene, computing its feature vectors, and clustering these feature vectors, each target scene is clustered into multi-class target scenes. Then, for each type of target scenario, an initial recognition model is built, a training set is established to train the model, and optimize it to obtain the target scenario recognition model.

Benefits of technology

Through clustering and model optimization, it can adapt to scenes with large number of scenes and great differences, and improve the effect and accuracy of scene recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128416_05062025_PF_FP_ABST
    Figure CN2024128416_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a model training method, apparatus, and device based on scene clustering, and a storage medium. The method comprises: obtaining initial images of target scenes, and carrying out feature computation on each initial image to obtain a feature vector corresponding to the initial image; clustering the feature vectors so as to cluster the target scenes into multiple types of target scenes; for each type of target scenes, constructing an initial recognition model corresponding to the type of target scenes; establishing a training set to train the initial recognition model to obtain a first recognition model; and then performing optimization to obtain a target scene recognition model corresponding to the type of target scenes, so as to obtain target scene recognition models respectively corresponding to the types of target scenes. According to the solution, target scenes can be clustered into multiple types of scenes, and a corresponding initial recognition model is constructed for each type of target scenes; the present application is suitable for a large number of scenes having large differences, and improves the scene recognition effect.
Need to check novelty before this filing date? Find Prior Art

Description

Model training method, device, equipment and storage medium based on scene clustering Technical Field

[0001] The present application relates to the field of model training technology, and specifically to a model training method, apparatus, device and storage medium based on scene clustering. Background Art

[0002] Currently, the training process for multi-scene algorithms requires collecting images from all scenes and aggregating them into a total dataset. Data augmentation techniques are then used to perform operations such as cutout, random erasing, grid masking, and mixup on the dataset to improve data diversity. An algorithm model is then designed based on requirements, relevant hyperparameters are configured, and the model is trained using the data-enhanced dataset. The target model is obtained when the training reaches the end condition.

[0003] However, when the number of scenes is large and the differences between the scenes are also large, the feature representation ability of a single algorithm model is limited and cannot adapt to a large number of scenes, resulting in poor scene recognition effect.

[0004] Summary of the Invention

[0005] In view of this, the present application provides a model training method, apparatus, device and storage medium based on scene clustering, which is used to solve the problem that when the number of scenes is relatively large and the differences between the scenes are also relatively large, the feature expression ability of a single algorithm model is limited and it cannot adapt to a large number of scenes, resulting in poor scene recognition effect.

[0006] To achieve the above objectives, the following solutions are proposed:

[0007] First, a model training method based on scene clustering, comprising:

[0008] Obtain the initial images of each target scene, and perform feature calculation on each initial image to obtain the feature vector corresponding to each initial image;

[0009] Clustering each feature vector to cluster each target scene into multiple target scenes;

[0010] For each type of target scene, an initial recognition model corresponding to the target scene is constructed;

[0011] Establishing a training set corresponding to the target scene of this type, and using the training set to train the initial recognition model to obtain a first recognition model;

[0012] The first recognition model is optimized to obtain a target scene recognition model corresponding to the target scene type, thereby obtaining target scene recognition models corresponding to each type of target scene.

[0013] Preferably, clustering the feature vectors to cluster the target scenes into multiple target scenes includes:

[0014] Get the difference between every two eigenvectors;

[0015] Randomly select two eigenvectors, and if the difference between the two eigenvectors is less than a preset first difference threshold, then the two eigenvectors are regarded as eigenvectors of the same type;

[0016] If the difference between the two feature vectors is greater than a preset second difference threshold, the two feature vectors are regarded as feature vectors of different classes;

[0017] Traverse the difference between each two eigenvectors to determine each type of eigenvector;

[0018] Each type of feature vector corresponds to each target scene, so as to cluster each target scene into multiple types of target scenes.

[0019] Preferably, establishing a training set corresponding to the target scene includes:

[0020] Collecting a sample image of each target scene in the target scene category at intervals of a preset first time period to obtain sample images;

[0021] Performing deduplication processing on all the sample images to obtain first images;

[0022] performing data enhancement processing on each of the first images to obtain each second image;

[0023] The second images are aggregated to obtain a training set corresponding to the target scene.

[0024] Preferably, the optimizing the first recognition model to obtain a target scene recognition model corresponding to the target scene type includes:

[0025] Obtain video data of the target scene;

[0026] Inputting the video data into the first recognition model to obtain scene recognition information of the target scene;

[0027] Comparing the scene recognition information with the real scene information in the target scene to determine a problem data set;

[0028] The first recognition model is optimized using the problem data set to obtain a target scene recognition model corresponding to the target scene type.

[0029] Preferably, the optimizing the first recognition model by using the problem data set to obtain a target scene recognition model corresponding to the target scene type includes:

[0030] Determining the size of the problem dataset;

[0031] Determining whether the size of the problem data set is greater than a preset first automatic learning threshold;

[0032] If so, obtaining the union of the first data set and the problem data set;

[0033] Using the union to train the first recognition model to obtain a second recognition model;

[0034] Optimizing the second recognition model to obtain a third recognition model;

[0035] Calculating the accuracy and recall of the third recognition model;

[0036] If the accuracy of the third recognition model is greater than a preset accuracy threshold, and the recall rate of the third recognition model is greater than a preset recall rate threshold, the third recognition model is used as the target scene recognition model corresponding to this type of target scene.

[0037] Preferably, it also includes:

[0038] If the size of the problem data set is not greater than the first automatic learning threshold, new video data of this type of target scene is reacquired, and the step of inputting the video data into the first recognition model to obtain scene recognition information of the target scene is returned to, until the number of times the video data is reacquired reaches a preset acquisition number threshold, and the currently obtained first recognition model is used as the target scene recognition model corresponding to this type of target scene.

[0039] Preferably, comparing the scene recognition information with the real scene information in the target scene to determine the problem data set includes:

[0040] Determining each real object in the real scene information of the target scene;

[0041] Matching each piece of identification information in the scene identification information with each of the real objects one by one;

[0042] The identification information of the corresponding errors is summarized into the problem dataset.

[0043] In a second aspect, a model training device based on scene clustering includes:

[0044] The feature calculation module is used to obtain the initial image of each target scene and perform feature calculation on each initial image to obtain the feature vector corresponding to each initial image;

[0045] A clustering module, used to cluster each feature vector to cluster each target scene into multiple target scenes;

[0046] A model building module is used to build an initial recognition model corresponding to each type of target scene;

[0047] A model training module is used to establish a training set corresponding to the target scene of this type, and use the training set to train the initial recognition model to obtain a first recognition model;

[0048] The optimization module is used to optimize the first recognition model to obtain a target scene recognition model corresponding to the target scene type, thereby obtaining target scene recognition models corresponding to each type of target scene.

[0049] In a third aspect, a model training device based on scene clustering includes a memory and a processor;

[0050] The memory is used to store programs;

[0051] The processor is used to execute the program to implement the various steps of the model training method based on scene clustering as described in the first aspect.

[0052] In a fourth aspect, a storage medium stores a computer program thereon, which, when executed by a processor, implements the various steps of the model training method based on scene clustering as described in the first aspect.

[0053] It can be seen from the above technical solution that the present application obtains the initial images of each target scene and calculates the features of each initial image to obtain the feature vector corresponding to each initial image; clusters the feature vectors to cluster the target scenes into multiple categories of target scenes; for each category of target scene, constructs an initial recognition model corresponding to the target scene; establishes a training set corresponding to the target scene, uses the training set to train the initial recognition model to obtain a first recognition model; optimizes the first recognition model to obtain a target scene recognition model corresponding to the target scene, thereby obtaining each target scene recognition model corresponding to each category of target scene. After obtaining the initial images of each target scene, this solution calculates the feature vectors and clusters the feature vectors. It can cluster the target scenes into multiple categories of scenes, and then constructs a corresponding initial recognition model for each category of target scene. It is suitable for scenes with a large number of scenes and relatively large differences, and improves the effect of scene recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0055] FIG1 is a schematic diagram of a scene sample distribution provided by an embodiment of the present application;

[0056] FIG2 is a schematic diagram of a model training process provided in an embodiment of the present application;

[0057] FIG3 is an optional flow chart of a model training method based on scene clustering provided in an embodiment of the present application;

[0058] FIG4 is a schematic diagram of a process of scene clustering provided in an embodiment of the present application;

[0059] FIG5 is a schematic diagram of a model training process based on scene clustering provided in an embodiment of the present application;

[0060] FIG6 is a schematic diagram of a model optimization process provided in an embodiment of the present application;

[0061] FIG7 is a schematic structural diagram of a model training device based on scene clustering provided in an embodiment of the present application;

[0062] FIG8 is a structural diagram of a model training device based on scene clustering provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0064] Nowadays, cameras are required for surveillance and security in all walks of life. However, image analysis on large-scale video surveillance camera platforms with over 10 million devices is subject to significant sample imbalance and long-tail distribution issues, as shown in Figure 1. The same algorithm needs to cover all scenarios. This imbalance in samples during training prevents the algorithm model from learning the characteristics of scenes with small sample sizes, leading to poor performance during inference for scenes with small sample sizes.

[0065] In addition, to adapt to ever-changing monitoring scenarios, the traditional algorithm iteration method is time-consuming and cannot quickly respond to business development. Currently, for multi-scenario algorithms, the training process requires collecting images from all scenarios and aggregating all images into a total data set. Then, data enhancement technology is used to perform operations such as cutout, random erasing, grid masking, and mixup on the data set to improve data diversity. Then, the algorithm model is designed according to the needs, the relevant hyperparameters are configured, and the model is trained using the above data-enhanced data set. When the training reaches the end condition, the target model is obtained, as shown in Figure 2.

[0066] However, when the number of scenes is large and the differences between the scenes are also large (non-independent and identically distributed), the feature representation ability of a single algorithm model is limited and cannot adapt to a large number of scenes, resulting in poor scene recognition effect.

[0067] To address the above-mentioned deficiencies in the prior art, an embodiment of the present invention provides a model training method based on scenario clustering. This method can be applied to various computer terminals or smart terminals, and its execution subject can be a processor or server of the computer terminal or smart terminal. The method flow chart of the method is shown in FIG3 , which specifically includes:

[0068] S1: Obtain the initial images of each target scene, and perform feature calculation on each initial image to obtain the feature vector corresponding to each initial image.

[0069] In this application, a variety of scenes can be selected as target scenes, and the differences between the target scenes vary, with some target scenes having large differences and some target scenes having small differences. Therefore, this application first obtains an initial image of each target scene, and then performs feature calculation on each initial image to obtain each feature vector. The feature vector is used to cluster the target scenes, which can improve the accuracy of clustering.

[0070] S2: Clustering each feature vector to cluster each target scene into multiple categories of target scenes.

[0071] Since this application targets multiple target scenes and there are differences between the target scenes, the differences can be used to cluster these target scenes, and the clusters can be divided into multiple categories of target scenes. Then, a scene recognition model is constructed for each category of target scenes. This can ensure that each category of target scenes has a corresponding scene recognition model, thereby improving the recognition efficiency of each category of target scenes.

[0072] S3: For each type of target scene, construct an initial recognition model corresponding to the target scene.

[0073] In this step, a corresponding initial recognition model is constructed for each type of target scene, and the initial recognition model is trained using the training set. The trained model can accurately recognize the target scene.

[0074] S4: Establish a training set corresponding to the target scene of this type, and use the training set to train the initial recognition model to obtain a first recognition model.

[0075] In addition to clustering the target scenes, it is also necessary to train the model constructed corresponding to each type of target scene to optimize the model parameters and make the model recognition effect better. Therefore, for each type of target scene, it is necessary to establish a training set corresponding to this type of target scene. The training set is a data set related to this type of target scene. Then, the corresponding model is trained using the data set of this type of target scene to obtain the first recognition model.

[0076] S5: Optimizing the first recognition model to obtain a target scene recognition model corresponding to the target scene type, thereby obtaining target scene recognition models corresponding to each type of target scene.

[0077] It is understandable that the training process in step S4 above may not produce a model with excellent results. Therefore, in this step, the first recognition model is optimized to obtain each target scene recognition model corresponding to each type of target scene. Thus, for each type of target scene, there is a corresponding target scene recognition model. Compared with the existing technology that only builds a single model for all target scenes, this solution can identify scenes by type and use the corresponding target scene recognition model to identify the scene, thereby improving the recognition accuracy.

[0078] It can be seen from the above technical solution that the present application obtains the initial images of each target scene and calculates the features of each initial image to obtain the feature vector corresponding to each initial image; clusters the feature vectors to cluster the target scenes into multiple categories of target scenes; for each category of target scene, constructs an initial recognition model corresponding to the target scene; establishes a training set corresponding to the target scene, uses the training set to train the initial recognition model to obtain a first recognition model; optimizes the first recognition model to obtain a target scene recognition model corresponding to the target scene, thereby obtaining each target scene recognition model corresponding to each category of target scene. After obtaining the initial images of each target scene, this solution calculates the feature vectors and clusters the feature vectors. It can cluster the target scenes into multiple categories of scenes, and then constructs a corresponding initial recognition model for each category of target scene. It is suitable for scenes with a large number of scenes and relatively large differences, and improves the effect of scene recognition.

[0079] In the method provided in the embodiment of the present invention, the process of clustering each feature vector to cluster each target scene into multiple target scenes is specifically described as follows:

[0080] Get the difference between every two eigenvectors;

[0081] Randomly select two eigenvectors, and if the difference between the two eigenvectors is less than a preset first difference threshold, then the two eigenvectors are regarded as eigenvectors of the same type;

[0082] If the difference between the two feature vectors is greater than a preset second difference threshold, the two feature vectors are regarded as feature vectors of different classes;

[0083] Traverse the difference between each two eigenvectors to determine each type of eigenvector;

[0084] Each type of feature vector corresponds to each target scene, so as to cluster each target scene into multiple types of target scenes.

[0085] Specifically, after obtaining the difference value between each two eigenvectors, each target scene is clustered according to the difference value. In order to distinguish target scenes with large differences as much as possible, this solution sets a first difference threshold and a second difference threshold, traverses the difference value between each two eigenvectors, and divides the target scenes with a difference value less than the first difference threshold into the same category of target scenes, thereby dividing them into multiple categories of target scenes, and the differences between each category of target scenes should be as large as possible. In this application, target scenes with differences greater than the second difference threshold are regarded as different categories of target scenes. In other words, the differences between target scenes in each category are less than the preset first difference threshold, and the differences between target scenes of different categories are greater than the preset second difference threshold.

[0086] In one example, the camera used to obtain the initial image in each target scene can be coded and each camera is assigned a unique ID name, denoted as camera n Then, a monitoring screen picture is obtained from each camera, that is, the initial image, recorded as pic n , then perform feature calculation. Optionally, this application uses the CLIP network model to perform feature calculation on each initial image. The CLIP network model is an open-source large-scale (400 million image-text pairs) image-text training model based on contrastive learning developed by OpenAI. This model has very strong text semantics and image semantic representation capabilities. After academic and engineering experiments, it has a strong migration effect for non-preset labels, that is, it has zero-shot recognition capabilities. After feature calculation, the feature vector of each initial image is obtained, which is recorded as feature n , then perform K-means clustering on the N feature vectors, where K centers correspond to K different scenes, recorded as scene k ,Therefore, after clustering, N cameras are divided into K categories, each of which contains a different number of cameras (target scenes), as shown in Figure 4.

[0087] The following describes in detail the process of establishing a training set corresponding to this type of target scene in this application.

[0088] Collecting a sample image of each target scene in the target scene category at intervals of a preset first time period to obtain sample images;

[0089] Performing deduplication processing on all the sample images to obtain first images;

[0090] performing data enhancement processing on each of the first images to obtain each second image;

[0091] The second images are aggregated to obtain a training set corresponding to the target scene.

[0092] Specifically, the initial recognition model constructed for each type of target scene is recorded as algorithm k , collect sample images of each target scene in this type of target scene at every preset first time period to obtain each sample image, obtain each first image after deduplication, and summarize each first image to obtain a data set corresponding to this type of target scene, recorded as D k , and then enhance each data set to obtain each second image. The data set composed of the second images is recorded as DA k , which is the training set corresponding to this type of target scene.

[0093] Next, we use the above training set DA k training algorithm k , get the final first recognition model mod el k , understandably, model el k It is only necessary to perform reasoning analysis on the camera of scene K to meet the recognition effect of the scene. The above process can be shown in Figure 5. In the above scheme, after N non-independent and identically distributed target scenes are converted into K independent and identically distributed scenes, K algorithm models are allocated to the K independent target scenes for parallel independent training. The data sets collected by the K independent target scenes are used independently. For small sample target scenes trained independently, data enhancement can be used to enrich data diversity, avoiding the problem that the recognition effect of small sample target scenes cannot be improved due to the uneven merging of data from each target scene into model training. The scheme adopted by the present invention to cluster each target scene and then design an algorithm for each type of scene can solve the problem that the generalization ability of a single algorithm model in the prior art is limited and cannot adapt to all non-independent and identically distributed scenes.

[0094] The above scheme describes the process of establishing a training set corresponding to this type of target scene in this application. The following embodiment describes in detail the process of optimizing the first recognition model in this scheme to obtain a target scene recognition model corresponding to this type of target scene.

[0095] Obtain video data of the target scene;

[0096] Inputting the video data into the first recognition model to obtain scene recognition information of the target scene;

[0097] Comparing the scene recognition information with the real scene information in the target scene to determine a problem data set;

[0098] The first recognition model is optimized using the problem data set to obtain a target scene recognition model corresponding to the target scene type.

[0099] Specifically, the process of optimizing the first recognition model using the problem data set to obtain a target scene recognition model corresponding to the target scene type includes:

[0100] Determining the size of the problem dataset;

[0101] Determining whether the size of the problem data set is greater than a preset first automatic learning threshold;

[0102] If the size of the problem data set is greater than the first automatic learning threshold, obtaining a union of the first data set and the problem data set;

[0103] Using the union to train the first recognition model to obtain a second recognition model;

[0104] Optimizing the second recognition model to obtain a third recognition model;

[0105] Calculating the accuracy and recall of the third recognition model;

[0106] If the accuracy of the third recognition model is greater than a preset accuracy threshold, and the recall rate of the third recognition model is greater than a preset recall rate threshold, the third recognition model is used as the target scene recognition model corresponding to the target scene type;

[0107] If the size of the problem data set is not greater than the first automatic learning threshold, new video data of this type of target scene is reacquired, and the step of inputting the video data into the first recognition model to obtain scene recognition information of the target scene is returned to, until the number of times the video data is reacquired reaches a preset acquisition number threshold, and the currently obtained first recognition model is used as the target scene recognition model corresponding to this type of target scene.

[0108] In the above solution, the first recognition model model el k Access the corresponding target scene to obtain the video data of the target scene, input the video data into the first recognition model to obtain the scene recognition information of the target scene, and then make judgments based on the obtained scene recognition information to collect the misjudgment data set, that is, the problem data set ΔD k, and determine the size of the problem data set, which can be understood here as the size of the memory occupied. If its size is greater than the preset first automatic learning threshold, it means that the probability of misjudgment is relatively large, so the first recognition model needs to be further optimized. Merging the problem data set with the first data set, taking the union, and then using the union to optimize the training of the first recognition model can increase the amount of data in the training set, and locally increase the problematic data, which can make the training effect of the first recognition model better. This application also sets an accuracy threshold and a recall rate threshold to determine whether the final model meets the requirements, so as to continuously improve the accuracy of the algorithm for each scenario. The more iterations, the higher the accuracy; moreover, the parallel processing process can improve the efficiency of model training and shorten the iteration cycle of the model. The above process can be shown in Figure 6. Due to the long-tail distribution characteristics of each scene, the training data that can be collected is also unbalanced. Therefore, small sample scenes are not learned enough during the algorithm training process, which easily leads to poor recognition effect in the corresponding scene. The present invention uses a method of clustering first and then dividing the scene into different processing methods. The algorithm training is carried out independently of each other on the training sample data of each scene, and the number of individual scenes is continuously accumulated. This can improve the recognition effect of all target scenes and solve the problem that the recognition effect of small sample scenes cannot be improved. In addition, after multiple automatic learning of each target scene, the recognition accuracy of the target scene recognition model will gradually improve. Compared with the mainstream single model training scheme, the algorithm recognition effect is better and the optimization effect is more significant.

[0109] Furthermore, the process of comparing the scene recognition information with the real scene information in the target scene to determine the problem data set may specifically include:

[0110] Determining each real object in the real scene information of the target scene;

[0111] Matching each piece of identification information in the scene identification information with each of the real objects one by one;

[0112] The identification information of the corresponding errors is summarized into the problem dataset.

[0113] Corresponding to the method shown in FIG3 , an embodiment of the present invention further provides a model training device based on scene clustering, which is used to implement the method shown in FIG3 . The model training device based on scene clustering provided in an embodiment of the present invention can be used in a computer terminal or various mobile devices. In conjunction with FIG7 , the model training device based on scene clustering is introduced. As shown in FIG7 , the device may include:

[0114] The feature calculation module 10 is used to obtain the initial image of each target scene and perform feature calculation on each initial image to obtain the feature vector corresponding to each initial image;

[0115] A clustering module 20 is used to cluster each feature vector to cluster each target scene into multiple target scenes;

[0116] A model building module 30 is used to build an initial recognition model corresponding to each type of target scene;

[0117] A model training module 40 is configured to establish a training set corresponding to the target scene, and train the initial recognition model using the training set to obtain a first recognition model;

[0118] The optimization module 50 is configured to optimize the first recognition model to obtain a target scene recognition model corresponding to the target scene type, thereby obtaining target scene recognition models corresponding to each type of target scene.

[0119] It can be seen from the above technical solution that the present application obtains the initial images of each target scene and calculates the features of each initial image to obtain the feature vector corresponding to each initial image; clusters the feature vectors to cluster the target scenes into multiple categories of target scenes; for each category of target scene, constructs an initial recognition model corresponding to the target scene; establishes a training set corresponding to the target scene, uses the training set to train the initial recognition model to obtain a first recognition model; optimizes the first recognition model to obtain a target scene recognition model corresponding to the target scene, thereby obtaining each target scene recognition model corresponding to each category of target scene. After obtaining the initial images of each target scene, this solution calculates the feature vectors and clusters the feature vectors. It can cluster the target scenes into multiple categories of scenes, and then constructs a corresponding initial recognition model for each category of target scene. It is suitable for scenes with a large number of scenes and relatively large differences, and improves the effect of scene recognition.

[0120] Furthermore, an embodiment of the present application provides a model training device based on scene clustering. Optionally, FIG8 shows a hardware structure block diagram of the model training device based on scene clustering. Referring to FIG8 , the hardware structure of the model training device based on scene clustering may include: at least one processor 01, at least one communication interface 02, at least one memory 03, and at least one communication bus 04.

[0121] In the embodiment of the present application, the number of the processor 01 , the communication interface 02 , the memory 03 , and the communication bus 04 is at least one, and the processor 01 , the communication interface 02 , and the memory 03 communicate with each other through the communication bus 04 .

[0122] The processor 01 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0123] The memory 03 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0124] The memory stores a program, and the processor can call the program stored in the memory, and the program is used to execute the following model training method based on scene clustering, including:

[0125] Obtain the initial images of each target scene, and perform feature calculation on each initial image to obtain the feature vector corresponding to each initial image;

[0126] Clustering each feature vector to cluster each target scene into multiple target scenes;

[0127] For each type of target scene, construct an initial recognition model corresponding to the target scene;

[0128] Establishing a training set corresponding to the target scene of this type, and using the training set to train the initial recognition model to obtain a first recognition model;

[0129] The first recognition model is optimized to obtain a target scene recognition model corresponding to the target scene type, thereby obtaining target scene recognition models corresponding to each type of target scene.

[0130] Optionally, the refined and extended functions of the program may refer to the description of the model training method based on scenario clustering in the method embodiment.

[0131] The present application also provides a storage medium that can store a program suitable for execution by a processor. When the program is executed, the device where the storage medium is located is controlled to perform the following model training method based on scene clustering, including:

[0132] Obtain the initial images of each target scene, and perform feature calculation on each initial image to obtain the feature vector corresponding to each initial image;

[0133] Clustering each feature vector to cluster each target scene into multiple target scenes;

[0134] For each type of target scene, construct an initial recognition model corresponding to the target scene;

[0135] Establishing a training set corresponding to the target scene of this type, and using the training set to train the initial recognition model to obtain a first recognition model;

[0136] The first recognition model is optimized to obtain a target scene recognition model corresponding to the target scene type, thereby obtaining target scene recognition models corresponding to each type of target scene.

[0137] Specifically, the storage medium may be a computer-readable storage medium, and the computer-readable storage medium may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk or a ROM.

[0138] Optionally, the refined and extended functions of the program may refer to the description of the model training method based on scenario clustering in the method embodiment.

[0139] In addition, the functional modules in the various embodiments of the present disclosure can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for a computer device (which can be a personal computer, a live broadcast device, or a network device, etc.) to perform all or part of the steps of the methods of the various embodiments of the present disclosure.

[0140] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0141] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0142] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A model training method based on scene clustering, characterized in that: include: Obtaining the initial images of each target scene, and performing feature calculation on each initial image to obtain the feature vector corresponding to each initial image; Clustering each feature vector to cluster each target scene into multiple target scenes; For each type of target scene, an initial recognition model corresponding to the target scene is constructed; Establishing a training set corresponding to the target scene of this type, and using the training set to train the initial recognition model to obtain a first recognition model; The first recognition model is optimized to obtain a target scene recognition model corresponding to the target scene of this type, thereby obtaining target scene recognition models corresponding to each type of target scene.

2. The method according to claim 1, characterized in that The clustering of each feature vector to cluster each target scene into multiple target scenes includes: Get the difference between every two eigenvectors; Randomly select two feature vectors, and if the difference between the two feature vectors is less than a preset first difference threshold, the two feature vectors are regarded as feature vectors of the same type; If the difference between the two feature vectors is greater than a preset second difference threshold, the two feature vectors are regarded as feature vectors of different classes; Traverse the difference between every two eigenvectors to determine each type of eigenvector; Each type of feature vector corresponds to each target scene, so as to cluster each target scene into multiple types of target scenes.

3. The method according to claim 1, characterized in that The step of establishing a training set corresponding to the target scene includes: Collecting a sample image of each target scene in the target scene category at intervals of a preset first time period to obtain each sample image; Performing deduplication processing on all the sample images to obtain first images; Performing data enhancement processing on each of the first images to obtain each second image; The second images are aggregated to obtain a training set corresponding to the target scene of this type.

4. The method according to claim 1, characterized in that: The step of optimizing the first recognition model to obtain a target scene recognition model corresponding to the target scene of the type includes: Obtain video data of the target scene; Inputting the video data into the first recognition model to obtain scene recognition information of the target scene; Comparing the scene recognition information with the real scene information in the target scene to determine a problem data set; The first recognition model is optimized using the problem data set to obtain a target scene recognition model corresponding to the target scene of this type.

5. The method according to claim 4, characterized in that The step of optimizing the first recognition model by using the problem data set to obtain a target scene recognition model corresponding to the target scene of the type includes: Determining the size of the problem data set; Determining whether the size of the problem data set is greater than a preset first automatic learning threshold; If yes, obtaining the union of the first data set and the problem data set; Using the union to train the first recognition model to obtain a second recognition model; Optimizing the second recognition model to obtain a third recognition model; Calculating the accuracy and recall of the third recognition model; If the accuracy of the third recognition model is greater than a preset accuracy threshold, and the recall rate of the third recognition model is greater than a preset recall rate threshold, the third recognition model is used as the target scene recognition model corresponding to this type of target scene.

6. The method according to claim 5, characterized in that Also includes: If the size of the problem data set is not greater than the first automatic learning threshold, new video data of this type of target scene is reacquired, and the step of inputting the video data into the first recognition model to obtain scene recognition information of the target scene is returned to execute, until the number of times the video data is reacquired reaches a preset acquisition number threshold, and the currently obtained first recognition model is used as the target scene recognition model corresponding to this type of target scene.

7. The method according to claim 4, characterized in that The comparing the scene recognition information with the real scene information in the target scene to determine the problem data set includes: Determine each real object in the real scene information of the target scene; Matching each piece of identification information in the scene identification information with each of the real objects one by one; The identification information of the corresponding errors is summarized into the problem dataset.

8. A model training device based on scene clustering, characterized in that: include: A feature calculation module is used to obtain the initial image of each target scene, and perform feature calculation on each initial image to obtain a feature vector corresponding to each initial image; A clustering module, used for clustering each feature vector to cluster each target scene into multiple target scenes; A model building module is used to build an initial recognition model corresponding to each type of target scene. A model training module, used to establish a training set corresponding to the target scene of this type, and use the training set to train the initial recognition model to obtain a first recognition model; The optimization module is used to optimize the first recognition model to obtain a target scene recognition model corresponding to the type of target scene, thereby obtaining target scene recognition models corresponding to each type of target scene.

9. A model training device based on scene clustering, characterized in that: including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the field-based The various steps of the model training method for scene clustering.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the model training method based on scene clustering as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Training method for image recognition model, image recognition method and related device

    CN109376781A

  • Image classification model training method and system capable of automatically generating training data set

    CN110598752A

  • Model training method, scene recognition method, computing equipment and medium

    CN113822130A

  • Model training method and device based on scene clustering, equipment and storage medium

    CN117635992A

Cited By

  • Battlefield target intelligent identification plotting system and method based on deep learning

    CN120833468A

  • Scene text recognition method, device, equipment, storage medium and program product

    CN122510881A

  • Scene text recognition methods, devices, equipment, storage media, and software products

    CN122510881B