Graph type identification method and device, electronic equipment and readable storage medium
Patent Information
- Application Number
- CN202380010444.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-08-31
AI Technical Summary
The existing graphic recognition model cannot be recognized when facing new graphics, which reduces the user experience.
The graph type is determined by obtaining the feature vector of the graph to be identified and matching the similarity with the candidate feature vectors in the preset database. This method does not require retraining the model when a new type of graph appears.
The recognition of new graphics is realized, the efficiency of graphics type recognition is improved, and the cost of model retraining is reduced.
Smart Images

Figure CN119948536A_ABST
Abstract
Description
Graphic type recognition method, device, electronic device and readable storage medium Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a graphic type recognition method, device, electronic device, and readable storage medium. Background Art
[0002] In the field of graphic type recognition, the graphic recognition model can identify the target type of graphics after training. For example, in the smart screen application scenario, users can draw graphics on the smart screen, and the graphic recognition model can then recognize the drawn graphics, making it easier to use.
[0003] However, when drawing graphics, users may draw new types of graphics, such as mathematical graphics, which may cause the graphics recognition model to be unable to recognize them, thereby reducing the user experience.
[0004] Summary of the Invention
[0005] The present disclosure provides a graphic type recognition method, device, electronic device and readable storage medium to address the deficiencies of related technologies.
[0006] According to a first aspect of an embodiment of the present disclosure, a method for identifying a graphic type is provided, the method comprising:
[0007] Acquire an initial image to be recognized, wherein the initial image includes a graphic to be recognized;
[0008] Acquiring a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model;
[0009] Obtaining similarities between the feature vector and each candidate feature vector in a preset database to obtain at least one target candidate feature vector;
[0010] The type of the to-be-identified graphic is determined according to the graphic type of the at least one target candidate feature vector.
[0011] Optionally, obtaining a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model includes:
[0012] Obtaining a preset graphic vector extraction model; the input data of the graphic vector extraction model is the image to be identified and the output data is the feature vector of the graphic to be identified;
[0013] The initial image is input into the graphic vector extraction model to obtain a feature vector of the graphic to be identified output by the graphic vector extraction model.
[0014] Optionally, the graphic vector extraction model is a network model obtained by removing the fully connected layer after the graphic recognition model training is completed.
[0015] Optionally, the training process of the graphic recognition model includes:
[0016] Acquire a graphic training sample set, wherein the graphic training samples in the graphic training sample set include at least one target graphic and a graphic type of the target graphic;
[0017] When the loss value is greater than the preset loss threshold, each graphic training sample is input into the graphic recognition model to obtain the recognition result;
[0018] Inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a preset loss function to obtain a loss value of the graphic recognition model;
[0019] When the loss value is less than or equal to a preset loss threshold, training the graphic recognition model is stopped.
[0020] Optionally, the preset loss function includes a classification loss function, a metric loss function and a distribution loss function.
[0021] Optionally, inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a preset loss function to obtain a loss value of the graphic recognition model includes:
[0022] Inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a classification loss function to obtain a first loss value;
[0023] Inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a metric loss function to obtain a second loss value;
[0024] Inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a distribution loss function to obtain a third loss value;
[0025] Obtain a weighted sum of the first loss value, the second loss value, and the third loss value as the loss value of the graphic recognition model.
[0026] Optionally, the at least one target candidate feature vector is obtained by selecting at least one candidate feature vector starting from the maximum similarity according to a size relationship.
[0027] Optionally, determining the type of the to-be-identified graphic according to the graphic type of the at least one target candidate feature vector includes:
[0028] Obtaining the frequency of occurrence of each graphic type in the graphic types of the at least one target candidate feature vector;
[0029] In response to obtaining the maximum value of the frequency of each graphic type, the graphic type corresponding to the maximum value of the frequency is determined as the type of the graphic to be identified.
[0030] Optionally, the method further includes:
[0031] In response to obtaining that the frequencies of at least two graphic types are both maximum values, the graphic type of the target candidate feature vector with the greatest similarity is determined as the type of the graphic to be identified.
[0032] Optionally, the method further includes:
[0033] In response to the similarity of each candidate feature vector being less than the preset similarity threshold, generating prompt information of the registration graphic and displaying it in the operation interface;
[0034] In response to an input operation in the operation interface, acquiring an associated image and a graphic type of the initial image;
[0035] Obtaining a feature vector of a graphic within the associated image;
[0036] Establishing a matching relationship between the feature vector of the associated image and the graphic type;
[0037] The feature vector, graphic type and matching relationship of the associated image are stored in the preset database.
[0038] Optionally, the method further includes:
[0039] In response to the similarity of each candidate feature vector being less than the preset similarity threshold or in response to detecting a registration operation, generating prompt information of a registration graphic and displaying it in an operation interface;
[0040] In response to an input operation in the operation interface, acquiring an associated image and a graphic type of the initial image;
[0041] Obtaining a feature vector of a graphic within the associated image and an average vector of the feature vectors;
[0042] Establishing a matching relationship between the average vector and the graphic type;
[0043] The average vector, graphic type and matching relationship are stored in the preset database.
[0044] Optionally, the associated image is obtained by:
[0045] Randomly scaling the size of the initial image to obtain multiple scaled images;
[0046] Performing blur processing on each zoomed image to serve as the associated image;
[0047] and / or,
[0048] Randomly rotating the direction of the initial image to obtain multiple rotated images;
[0049] Mosaic processing is performed on each of the rotated images to form the associated images.
[0050] According to a second aspect of an embodiment of the present disclosure, a device for identifying a graphic type is provided, the device comprising:
[0051] An initial image acquisition module, configured to acquire an initial image to be identified, wherein the initial image includes a graphic to be identified;
[0052] A feature vector acquisition module, configured to acquire a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model;
[0053] a target vector acquisition module, configured to obtain similarities between the feature vector and each candidate feature vector in a preset database, and obtain at least one target candidate feature vector;
[0054] The graphic type determination module is used to determine the type of the graphic to be identified according to the graphic type of the at least one target candidate feature vector.
[0055] Optionally, obtaining a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model includes:
[0056] Obtaining a preset graphic vector extraction model; the input data of the graphic vector extraction model is the image to be identified and the output data is the feature vector of the graphic to be identified;
[0057] The initial image is input into the graphic vector extraction model to obtain a feature vector of the graphic to be identified output by the graphic vector extraction model.
[0058] Optionally, the graphic vector extraction model is a network model obtained by removing the fully connected layer after the graphic recognition model training is completed.
[0059] Optionally, a model training module is further included, for training the graphic recognition model, including:
[0060] A sample set acquisition submodule is configured to acquire a graphic training sample set, wherein the graphic training samples in the graphic training sample set include at least one target graphic and a graphic type of the target graphic;
[0061] The recognition result acquisition submodule is used to input each graphic training sample into the graphic recognition model to obtain the recognition result when the loss value is greater than the preset loss threshold;
[0062] a loss value acquisition submodule, configured to input the recognition result and the graphic type of the target graphic in the graphic training sample into a preset loss function to obtain a loss value of the graphic recognition model;
[0063] The loss value judgment submodule is used to stop training the graphic recognition model when the loss value is less than or equal to a preset loss threshold.
[0064] Optionally, the preset loss function includes a classification loss function, a metric loss function and a distribution loss function.
[0065] Optionally, the loss value acquisition submodule includes:
[0066] A first loss value obtaining unit, configured to input the recognition result and the graphic type of the target graphic in the graphic training sample into a classification loss function to obtain a first loss value;
[0067] A second loss value obtaining unit, configured to input the recognition result and the graphic type of the target graphic in the graphic training sample into a metric loss function to obtain a second loss value;
[0068] A third loss value obtaining unit, configured to input the recognition result and the graphic type of the target graphic in the graphic training sample into a distribution loss function to obtain a third loss value;
[0069] A loss value acquisition unit is used to acquire a weighted sum of the first loss value, the second loss value and the third loss value as the loss value of the graphic recognition model.
[0070] Optionally, the at least one target candidate feature vector is obtained by selecting at least one candidate feature vector starting from the maximum similarity according to a size relationship.
[0071] Optionally, the graphic type determination module includes:
[0072] a type frequency acquisition submodule, configured to acquire the frequency of occurrence of each graphic type in the graphic type of the at least one target candidate feature vector;
[0073] The graphic type identification submodule is configured to, in response to obtaining the maximum value of the frequencies of various graphic types, determine the graphic type corresponding to the maximum frequency value as the type of the graphic to be identified.
[0074] Optionally, the graphic type determination module is further configured to, in response to obtaining that the frequencies of at least two graphic types are both maximum, determine the graphic type of the target candidate feature vector with the greatest similarity as the type of the graphic to be identified.
[0075] Optionally, the device further comprises:
[0076] a prompt information generating module, configured to generate prompt information of the registration graphic and display it in the operation interface in response to the similarity of each candidate feature vector being less than the preset similarity threshold;
[0077] An associated image acquisition module, configured to acquire associated images and graphic types of the initial image in response to an input operation in the operation interface;
[0078] A feature vector acquisition module, used to acquire a feature vector of a graphic within the associated image;
[0079] A matching relationship establishing module, configured to establish a matching relationship between the feature vector of the associated image and the graphic type;
[0080] The storage module is used to store the feature vector, graphic type and matching relationship of the associated image into the preset database.
[0081] Optionally, the method further includes:
[0082] a prompt information generating module, configured to generate prompt information of the registration graphic and display it in the operation interface in response to the similarity of each candidate feature vector being less than the preset similarity threshold or in response to detecting a registration operation;
[0083] An associated image acquisition module, configured to acquire associated images and graphic types of the initial image in response to an input operation in the operation interface;
[0084] An average vector acquisition module, used to acquire the characteristic vectors of the graphics in the associated image and the average vector of the characteristic vectors;
[0085] a matching relationship establishing module, configured to establish a matching relationship between the average vector and the graphic type;
[0086] The storage module is used to store the average vector, graphic type and matching relationship in the preset database.
[0087] Optionally, the associated image acquisition module includes:
[0088] A scaled image acquisition submodule, configured to randomly scale the size of the initial image to obtain a plurality of scaled images;
[0089] An associated image acquisition submodule, configured to perform blur processing on each zoomed image to obtain the associated image;
[0090] and / or,
[0091] A rotation image acquisition submodule, configured to randomly rotate the direction of the initial image to obtain multiple rotated images;
[0092] The associated image acquisition submodule is used to perform mosaic processing on each rotated image to obtain the associated image.
[0093] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, comprising
[0094] processor;
[0095] a memory for storing a computer program executable by the processor;
[0096] The processor is configured to execute the computer program in the memory to implement the method described in the first aspect.
[0097] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which, when an executable computer program in the storage medium is executed by a processor, can implement the method described in the first aspect.
[0098] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:
[0099] As can be seen from the above embodiments, the solution provided by the embodiments of the present disclosure obtains an initial image to be identified, wherein the initial image includes a graphic to be identified; then, based on a preset graphic vector extraction model, a feature vector of the graphic to be identified in the initial image is obtained; thereafter, the similarity between the feature vector and each candidate feature vector in a preset database is obtained to obtain at least one target candidate feature vector; finally, the type of the graphic to be identified is determined based on the graphic type of the at least one target candidate feature vector. In this way, in this embodiment, the graphic type of the graphic to be identified is determined by obtaining the feature vector of the graphic to be identified and combining it with the similarity. This solves the problem that the existing graphic recognition model needs to be retrained and re-identified when the graphic to be identified is of a new type without having to retrain the model. This facilitates expansion and reduces costs, and is conducive to improving the efficiency of graphic type recognition.
[0100] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0102] Fig. 1 is a flow chart showing a method for identifying a graphic type according to an exemplary embodiment.
[0103] Fig. 2 is a flow chart showing a method of obtaining a feature vector according to an exemplary embodiment.
[0104] Fig. 3 is a schematic diagram showing a feature classification according to an exemplary embodiment.
[0105] Fig. 4 is a schematic diagram showing another feature classification according to an exemplary embodiment.
[0106] Fig. 5 is a schematic diagram showing a distribution direction according to an exemplary embodiment.
[0107] Fig. 6 is a block diagram showing a new graphic registration process according to an exemplary embodiment.
[0108] Fig. 7 is a block diagram showing a device for identifying graphic types according to an exemplary embodiment. DETAILED DESCRIPTION
[0109] Exemplary embodiments will be described in detail herein, with examples shown in the accompanying drawings. When the following description refers to the drawings, identical numbers in different drawings represent identical or similar elements, unless otherwise indicated. The exemplary embodiments described below do not represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices consistent with certain aspects of the present disclosure, as detailed in the appended claims. It should be noted that, unless there is a conflict, the features of the following embodiments and implementations may be combined with each other.
[0110] To address the above technical issues, embodiments of the present disclosure provide a method for identifying graphic types, which can be applied to electronic devices, including but not limited to smartphones, tablets, smart displays, or electronic whiteboards. Figure 1 is a flowchart illustrating a method for identifying graphic types, according to an exemplary embodiment.
[0111] Referring to FIG1 , a method for identifying a graphic type includes steps 11 to 14:
[0112] In step 11, an initial image to be recognized is obtained, where the initial image includes a graphic to be recognized.
[0113] In this embodiment, the processor of the electronic device can obtain an initial image to be recognized. For example, the display screen of the electronic device can display a drawing interface. A user can draw a graphic within the drawing interface, which is subsequently referred to as the graphic to be recognized. The display screen can obtain an image that matches the display screen and transmit it to the processor. At this point, the processor can obtain the initial image, which, as will be understood, includes the graphic to be recognized.
[0114] In this embodiment, the graphics may include but are not limited to straight lines, straight arrows, wavy lines, squares, rectangles, rhombuses, trapezoids, circles, ellipses, parallelograms, pentagons, five-pointed stars and other plane graphics, and may also include spheres, cylinders, cubes and other three-dimensional graphics. They can be set according to the specific scene and are not limited here.
[0115] In step 12, a feature vector of the graphic to be identified in the initial image is obtained based on a preset graphic vector extraction model.
[0116] In this embodiment, the processor of the electronic device may obtain the feature vector of the graphic to be identified in the initial image based on a preset graphic vector extraction model, as shown in FIG. 2 , which includes steps 21 and 22 .
[0117] In step 21, the processor may obtain a preset graphic vector extraction model; the input data of the graphic vector extraction model is the image to be identified and the output data is the feature vector of the graphic to be identified.
[0118] In this step, the electronic device may store a graphic vector extraction model, whose input data is the image to be recognized and whose output data is the feature vector of the graphic to be recognized. In one example, the graphic vector extraction model is a network model obtained by removing the fully connected layer after training the graphic recognition model, and includes at least one convolutional layer and at least one fully connected layer. The graphic recognition model can be implemented using a neural network model, such as a DenseNet model, which can be selected based on the specific scenario and is not limited here.
[0119] In one possible embodiment, the process of acquiring the graphic vector extraction model, i.e., the training process of the graphic recognition model, includes:
[0120] The processor can obtain a graphic training sample set, wherein the graphic training samples in the graphic training sample set include at least one target graphic and the graphic type of the target graphic. In one example, each type of training sample in the graphic training sample set can include an actually collected graphic, such as a hand-drawn or printed graphic, and a synthetic graphic synthesized based on the actually collected graphic, such as randomly scaling the size of the actually collected graphic to obtain multiple scaled images; blurring each scaled image to obtain a synthetic graphic; or randomly rotating the direction of the actually collected graphic to obtain multiple rotated images; and mosaicking each rotated image to obtain a synthetic graphic. In one example, the ratio of the actual image to the synthetic image is 1:4, thereby ensuring the number of candidate feature vectors in the preset database and improving the accuracy of image recognition.
[0121] When the loss value is greater than a preset loss threshold, the processor may input each graphic training sample into the graphic recognition model to obtain a recognition result. The processor may input the recognition result and the graphic type of the target graphic in the graphic training sample into a preset loss function to obtain a loss value for the graphic recognition model. When the loss value is less than or equal to the preset loss threshold, the processor may stop training the graphic recognition model.
[0122] In one example, the preset loss function may include a classification loss function, a metric loss function, and a distribution loss function. Among them, the classification loss function is implemented using cross entropy loss (nn.CrossEntropyLoss), which is used to ensure that the graphic type of the graphic training sample and the probability of the type in the output result of the graphic recognition model are maximized. The cross entropy loss function is a loss function that is almost always used in classification tasks. Compared with the absolute error (absolute error, that is, directly looking at whether the network prediction is correct, either the prediction is correct and the loss is 0, or the prediction is wrong and the loss is 1), its loss changes more smoothly, which can reduce the fluctuation of the network during training.
[0123] For example, in the MNIST handwritten digit recognition dataset, each image has a number between 0 and 9, and each image can only have one fixed label (the number above). For a single sample, the cross entropy loss function is calculated as:
[0124]
[0125] In formula (1), y i represents the true distribution, represents the predicted distribution and n represents the number of categories.
[0126] In the handwritten digit recognition task, if an image is of the number "5", then the true distribution should be: [0,0,0,0,0,1,0,0,0,0], where only the label corresponding to the number "5" is 1, and the rest of the positions 0 to 9 are 0.
[0127] If the distribution of the network output is: [0.1, 0.1, 0, 0, 0.7, 0, 0.1, 0, 0], at this moment the score corresponding to position 5 is higher, and the scores of other positions are lower, then the loss function calculated by the cross entropy function is: Loss = -0*log0.1-0*log0.1-0*log0-0*log0-0*log0 -1*log0.7-0*log0-0*log0.1-0*log0-0*log0≈0.3567.
[0128] If the distribution of the network output is: [0.2, 0.3, 0.1, 0, 0, 0.3, 0.1, 0, 0, 0], then the loss function is calculated as: Loss = -0*log0.2-0*log0.3-0*log0.1-0*log0-0*log0 -1*log0.3-0*log0.1-0*log0-0*log0-0*log0≈1.2040.
[0129] Comparing the two cases above, the calculated loss of 0.3567 for the first distribution is significantly lower than the loss of 1.2040 for the second distribution, indicating that the first distribution is closer to the true distribution. This is indeed the case, as for an image that actually contains the number "5," the first network's output confidence for the position of the 5 is 0.7, while the second output is only 0.3. This demonstrates that using cross-entropy loss is a good measure of whether the network's output distribution closely matches the true situation. A larger loss indicates a poorer prediction, thus imposing a larger penalty on the network. By training the network in this way, its predictions gradually become increasingly accurate, ensuring that the network has basic classification capabilities.
[0130] The metric loss function uses a triplet loss, TripletLoss, to obtain the distance between two shapes. Specifically, the distance is smaller when the two shapes are similar and larger when they are dissimilar. Metric learning can be used to classify different shapes by calculating the feature distance between them. In this case, the features extracted by the network must be as discriminative as possible. This means that the features must not only be able to classify pairs, but also ensure that the distances within the same class are as close as possible, while the distances between different classes are as large as possible.
[0131] For example, see Figure 3. The solid and hollow dots in Figure 3 correspond to the features of the image, and can be separated by the dotted line in the middle. See Figure 4. The dotted line in the middle can separate the solid and hollow dots. By comparison, it can be seen that the distance between the solid and hollow dots in Figure 4 is greater than the distance between the solid and hollow dots in Figure 3.
[0132] In this example, we hope that the model will learn the features shown in Figure 4, that is, small intra-class distance and large inter-class distance. Therefore, in this example, the metric loss function uses TripletLoss.
[0133] In this embodiment, considering that each graphic can be projected in six directions, such as horizontally and vertically, the differences in distribution in the six directions are counted to determine the degree of difference between the skeleton structure of the actual graphic and the generated graphic. Since different graphics have different skeleton structures, they can also be projected in different directions to obtain different distributions. These distributions can also be used to determine whether the model-predicted graphic and the real graphic match. When using distribution loss, the graphic to be predicted and the standard graphic need to be normalized. If the network predicts that the current graphic is a square, the distribution loss between the graphic input to the network and the square will be calculated, that is, the structural similarity will be calculated to determine the correctness of the prediction. If the prediction is correct, the cumulative distribution loss in the six directions will be small. If the prediction is wrong, the cumulative loss will be relatively large.
[0134] The distribution loss function is implemented using formula (1) to obtain the edge distribution loss of the graph in all directions.
[0135]
[0136] In formula (1), P represents the number of distributions. See Figure 5, P represents the number of six directions; Y0 and Y1 represent the two graphs to be compared, and L represents the KL dispersion.
[0137] The preset loss function of this embodiment may include a classification loss function, a metric loss function, and a distribution loss function, which can make the model have strong classification capabilities while having good inter-class discrimination, so that it is more robust and accurate when directly using the feature vector to calculate the distance in the future.
[0138] During each training, the processor can input the recognition result and the graphic type of the target graphic in the graphic training sample into the classification loss function to obtain a first loss value; the processor can input the recognition result and the graphic type of the target graphic in the graphic training sample into the metric loss function to obtain a second loss value; the processor can input the recognition result and the graphic type of the target graphic in the graphic training sample into the distribution loss function to obtain a third loss value; the processor can obtain the weighted sum of the first loss value, the second loss value and the third loss value as the loss value of the graphic recognition model.
[0139] Among them, the weight value of each loss function can be obtained through the following steps:
[0140] First, a subset of training samples is extracted from the image training sample set as a validation set. Weights for the classification loss function, metric loss function, and distribution loss function are set separately to obtain multiple sets of weights. The image recognition models are trained in parallel based on these multiple sets of weights. After training the image recognition models corresponding to each set of weights is complete, the set of weights corresponding to the image recognition model with the highest recognition accuracy is determined as the weights for the classification loss function, metric loss function, and distribution loss function. The training process is then repeated.
[0141] Taking the DenseNet-264 image recognition model as an example, the label format of each image training sample is as follows: { 0: {'000000000139.jpg':['. / data / aaa / train / 1 / 000000000139.jpg',0], '000000000285.jpg':['. / data / aaa / train / 1 / 000000000285.jpg',0], '000000000785.jpg':['. / data / aaa / train / 1 / 000000000785.jpg',0], '000000000802.jpg':['. / data / aaa / train / 1 / 000000000802.jpg',0]...}, 1: {'000000001490.jpg':['. / data / aaa / train / 2 / 000000001490.jpg',1], '000000001503.jpg':['. / data / aaa / train / 2 / 000000001503.jpg',1]...} ...}.
[0142] Moreover, the triplet sample of each batch includes three sample data (two positive samples and one negative sample), as shown below: [['. / data / aaa / train / 2 / 000000001503.jpg',1], ['. / data / aaa / train / 2 / 000000000285.jpg',1], ['. / data / aaa / train / 1 / 000000000139.jpg',0]].
[0143] During the training process, the processor can increase the learning rate of the graphic recognition model from 0 to the initial learning rate, and then set it in a cosine decrease manner, thereby improving training efficiency.
[0144] Based on the above method, the graphic recognition model can be trained, and then the graphic vector extraction model can be obtained by deleting the fully connected layer of the graphic recognition model.
[0145] In this step, the processor may call the above-mentioned graphic vector extraction model.
[0146] In step 22, the processor may input the initial image into the graphic vector extraction model to obtain the feature vector of the graphic to be identified output by the graphic vector extraction model.
[0147] In step 13, the similarity between the feature vector and each candidate feature vector in a preset database is obtained to obtain at least one target candidate feature vector; the at least one target candidate feature vector is obtained by selecting at least one candidate feature vector starting from the maximum similarity according to the size relationship.
[0148] In one embodiment, the electronic device stores a preset database containing feature vectors, graphic types, and matching relationships for a number of graphics. In other words, each feature vector corresponds to a specific graphic type. This allows the processor to determine the similarity between the feature vector of the graphic to be identified and each candidate feature vector in the preset database. The processor can then sort the similarities based on the magnitude of the similarities to obtain at least one target candidate feature vector. Specifically, the at least one target candidate feature vector is obtained by selecting at least one candidate feature vector starting from the one with the highest similarity based on the magnitude of the similarity. This ensures that the target candidate feature vector is obtained in each comparison process, facilitating consistent recognition of subsequent types.
[0149] In another embodiment, the processor can also sort according to the size relationship of the similarity to obtain at least one target candidate feature vector whose similarity is greater than or equal to the similarity threshold, that is, at least one target candidate feature vector is obtained by selecting at least one candidate feature vector starting from the maximum similarity according to the size relationship, which is conducive to improving the accuracy of type recognition.
[0150] In step 14, the type of the to-be-identified graphic is determined according to the graphic type of the at least one target candidate feature vector.
[0151] In this step, considering that the graphic types of the above-mentioned target candidate feature vectors may be the same, the processor also counts the frequency of occurrence of each graphic type in the graphic type of at least one target candidate feature vector, and uses the graphic type corresponding to the maximum frequency as the type of the graphic to be identified. For example, if there are 5 target candidate feature vectors, of which the frequency of the rectangular type is 3 and the frequency of the pentagon type is 2, then the graphic type of the graphic to be identified is determined to be the rectangular type. In another example, the processor also determines that the image type of the target candidate feature vector with the greatest similarity is the type of the graphic to be identified in response to obtaining that the frequencies of at least two graphic types are both maximum values. For example, if there are 5 target candidate feature vectors, of which the frequency of the rectangular type is 2, the frequency of the pentagon type is 2, and the frequency of the triangle type is 1, then the image type of the target candidate feature vector with the greatest similarity (such as 0.85) is found to be the type of the graphic to be identified.
[0152] In one embodiment, when the similarity is less than a preset similarity threshold, it indicates that the recognition accuracy of the image to be recognized is low or that the image to be recognized cannot be recognized, or that the image to be recognized is a new image. The processor may register the new image, generating a prompt message for registering the image and displaying it in the user interface, such as, "The image to be recognized is new. Please follow the instructions to register the image." The user may enter the new image and its type in the user interface, i.e., the initial image in step 11.
[0153] The processor can obtain associated images of the initial image. For example, the processor can randomly scale the initial image to obtain multiple scaled images; blur each scaled image to obtain the associated images. For another example, the processor can randomly rotate the initial image to obtain multiple rotated images; mosaic each rotated image to obtain the associated images. In this way, the processor can obtain multiple associated images of the initial image.
[0154] The processor can obtain the feature vector of the graphic within the associated image (see Figure 6) and establish a matching relationship between the feature vector of the associated image and the graphic type. Finally, the processor can store the feature vector of the associated image, the graphic type, and the matching relationship in the preset database, thereby completing the registration of the new graphic. This allows the processor to obtain the graphic type of the new graphic during subsequent recognition processes.
[0155] In one embodiment, a user may need to register a new graphic to increase the number of graphic types in a preset database. Upon detecting a registration operation to register a new graphic, the electronic device may display a prompt message on the display interface to generate the registration graphic and display it on the operation interface, such as "Please follow the prompts to register a new graphic," "Please enter a new graphic," or "Please enter a graphic type." For example, a first prompt message may be displayed, "Please enter a new graphic." The user may hold the new graphic and place it in the preview scene of the electronic device's camera, at which point the electronic device may use the camera to capture the preview scene. Alternatively, the user may draw the new graphic on the display interface, and the electronic device may capture an image of the display interface and detect the new graphic in the screenshot. Alternatively, the user may move an image containing the new graphic from a designated location (local storage, external mobile storage, or the cloud) to the display interface, and the electronic device may capture the new graphic after capturing the image. In this way, the electronic device can obtain the initial image in a variety of ways. Then, "Please enter a graphic type" may be displayed, and the electronic device may obtain the graphic type of the new graphic in the initial image, i.e., the initial image in step 11.
[0156] The processor can then obtain associated images of the initial image. For example, the processor can randomly scale the initial image to obtain multiple scaled images; blur each scaled image to obtain the associated images. For another example, the processor can randomly rotate the initial image to obtain multiple rotated images; mosaic each rotated image to obtain the associated images. In this way, the processor can obtain several associated images of the initial image.
[0157] Finally, the processor can obtain the feature vector of the graphic within the associated image (see Figure 6 ), and establish a matching relationship between the feature vector of the associated image and the graphic type. Finally, the processor can store the feature vector, graphic type, and matching relationship of the associated image in the preset database, thereby completing the registration of the new graphic. This allows the processor to obtain the graphic type of the new graphic during subsequent recognition processes.
[0158] In each of the above embodiments, when registering a new graphic, associated images of the initial image will be obtained, and feature vectors of each associated image will be obtained, which can enrich the number of candidate feature vectors in the preset database and be applicable to various graphics to be identified of the same graphic type. In one embodiment, the electronic device can obtain the average vector of feature vectors of the same graphic type, and use the average vector as the feature vector of the graphic type, wherein the average vector refers to the vector obtained by averaging the data at the same position in the feature vector and replacing the data at that position with the average value. In this example, each graphic type corresponds to a feature vector, thereby reducing the number of candidate feature vectors in the preset database, which is beneficial to improving the efficiency and accuracy of identifying graphics.
[0159] In this way, in this embodiment, the graphic type of the graphic to be identified is determined by obtaining the feature vector of the graphic to be identified and combining it with the similarity. There is no need to retrain the model when the graphic to be identified is of a new type, which can solve the problem that the existing graphic recognition model needs to be retrained and re-identified when the graphic to be identified is of a new type, which is beneficial to improving the efficiency of graphic type recognition.
[0160] Based on a graphic type recognition method provided in an embodiment of the present disclosure, this embodiment further provides a graphic type recognition device. Referring to FIG7 , the device includes:
[0161] An initial image acquisition module 71 is used to acquire an initial image to be recognized, wherein the initial image includes a graphic to be recognized;
[0162] A feature vector acquisition module 72 is configured to acquire a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model;
[0163] A target vector acquisition module 73 is configured to obtain similarities between the feature vector and each candidate feature vector in a preset database, and obtain at least one target candidate feature vector;
[0164] The graphic type determination module is used to determine the type of the graphic to be identified according to the graphic type of the at least one target candidate feature vector.
[0165] In one embodiment, obtaining a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model includes:
[0166] Obtaining a preset graphic vector extraction model; the input data of the graphic vector extraction model is the image to be identified and the output data is the feature vector of the graphic to be identified;
[0167] The initial image is input into the graphic vector extraction model to obtain a feature vector of the graphic to be identified output by the graphic vector extraction model.
[0168] In one embodiment, the graphic vector extraction model is a network model obtained by removing the fully connected layer after the graphic recognition model training is completed.
[0169] In one embodiment, a model training module is further included, for training the graphic recognition model, including:
[0170] A sample set acquisition submodule is configured to acquire a graphic training sample set, wherein the graphic training samples in the graphic training sample set include at least one target graphic and a graphic type of the target graphic;
[0171] The recognition result acquisition submodule is used to input each graphic training sample into the graphic recognition model to obtain the recognition result when the loss value is greater than the preset loss threshold;
[0172] a loss value acquisition submodule, configured to input the recognition result and the graphic type of the target graphic in the graphic training sample into a preset loss function to obtain a loss value of the graphic recognition model;
[0173] The loss value judgment submodule is used to stop training the graphic recognition model when the loss value is less than or equal to a preset loss threshold.
[0174] In one embodiment, the preset loss function includes a classification loss function, a metric loss function, and a distribution loss function.
[0175] In one embodiment, the loss value acquisition submodule includes:
[0176] A first loss value obtaining unit, configured to input the recognition result and the graphic type of the target graphic in the graphic training sample into a classification loss function to obtain a first loss value;
[0177] A second loss value obtaining unit, configured to input the recognition result and the graphic type of the target graphic in the graphic training sample into a metric loss function to obtain a second loss value;
[0178] A third loss value obtaining unit, configured to input the recognition result and the graphic type of the target graphic in the graphic training sample into a distribution loss function to obtain a third loss value;
[0179] A loss value acquisition unit is used to acquire a weighted sum of the first loss value, the second loss value and the third loss value as the loss value of the graphic recognition model.
[0180] In one embodiment, the at least one target candidate feature vector is obtained by selecting at least one candidate feature vector starting from the maximum similarity according to the size relationship.
[0181] In one embodiment, the graphic type determination module includes:
[0182] a type frequency acquisition submodule, configured to acquire the frequency of occurrence of each graphic type in the graphic type of the at least one target candidate feature vector;
[0183] The graphic type identification submodule is configured to, in response to obtaining the maximum value of the frequencies of the various graphic types, determine the graphic type corresponding to the maximum value of the frequencies as the type of the graphic to be identified.
[0184] In one embodiment, the graphic type determination module is further configured to, in response to obtaining that the frequencies of at least two graphic types are both maximum, determine the graphic type of the target candidate feature vector with the greatest similarity as the type of the graphic to be identified.
[0185] In one embodiment, the graphic type determination module includes:
[0186] The graphic type determining unit is configured to determine the graphic type of the target candidate feature vector with the greatest similarity as the type of the graphic to be identified.
[0187] In one embodiment, the apparatus further comprises:
[0188] a prompt information generating module, configured to generate prompt information of the registration graphic and display it in the operation interface in response to the similarity of each candidate feature vector being less than the preset similarity threshold or in response to detecting a registration operation;
[0189] An associated image acquisition module, configured to acquire associated images and graphic types of the initial image in response to an input operation in the operation interface;
[0190] A feature vector acquisition module, used to acquire a feature vector of a graphic within the associated image;
[0191] A matching relationship establishing module, configured to establish a matching relationship between the feature vector of the associated image and the graphic type;
[0192] The storage module is used to store the feature vector, graphic type and matching relationship of the associated image into the preset database.
[0193] In one embodiment, the apparatus further comprises:
[0194] a prompt information generating module, configured to generate prompt information of the registration graphic and display it in the operation interface in response to the similarity of each candidate feature vector being less than the preset similarity threshold or in response to detecting a registration operation;
[0195] An associated image acquisition module, configured to acquire associated images and graphic types of the initial image in response to an input operation in the operation interface;
[0196] An average vector acquisition module, used to acquire the characteristic vectors of the graphics in the associated image and the average vector of the characteristic vectors;
[0197] a matching relationship establishing module, configured to establish a matching relationship between the average vector and the graphic type;
[0198] The storage module is used to store the average vector, graphic type and matching relationship in the preset database.
[0199] In one embodiment, the associated image acquisition module includes:
[0200] A scaled image acquisition submodule, configured to randomly scale the size of the initial image to obtain a plurality of scaled images;
[0201] An associated image acquisition submodule, configured to perform blur processing on each zoomed image to obtain the associated image;
[0202] and / or,
[0203] A rotation image acquisition submodule, configured to randomly rotate the direction of the initial image to obtain multiple rotated images;
[0204] The associated image acquisition submodule is used to perform mosaic processing on each rotated image to obtain the associated image.
[0205] It should be noted that the device embodiment shown in this embodiment matches the content of the above-mentioned method embodiment. You can refer to the content of the above-mentioned method embodiment and will not repeat it here.
[0206] In an exemplary embodiment, an electronic device is also provided, including:
[0207] Display screen;
[0208] processor;
[0209] a memory for storing a computer program executable by the processor;
[0210] The processor is configured to execute the computer program in the memory to implement the above method.
[0211] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including an executable computer program. The executable computer program can be executed by a processor to implement the method of the above embodiment. The computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0212] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the disclosure herein. This disclosure is intended to cover any variations, uses, or adaptations that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0213] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for identifying a graphic type, characterized in that: The method comprises: Acquire an initial image to be recognized, wherein the initial image includes a graphic to be recognized; Acquire a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model; Obtaining the similarity between the feature vector and each candidate feature vector in a preset database to obtain at least one target candidate feature vector; The type of the to-be-recognized graphic is determined according to the graphic type of the at least one target candidate feature vector.
2. The method according to claim 1, characterized in that Acquiring a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model includes: Obtaining a preset graphic vector extraction model; the input data of the graphic vector extraction model is the image to be identified and the output data is the feature vector of the graphic to be identified; The initial image is input into the graphic vector extraction model to obtain a feature vector of the graphic to be identified output by the graphic vector extraction model.
3. The method according to claim 2, characterized in that The graphic vector extraction model is a network model after the graphic recognition model training is completed and the fully connected layer is removed.
4. The method according to claim 3, characterized in that: The training process of the graphic recognition model includes: Acquire a graphic training sample set, wherein the graphic training samples in the graphic training sample set include at least one target graphic and a graphic type of the target graphic; Input each graphic training sample into the graphic recognition model to obtain the recognition result; Inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a preset loss function to obtain a loss value of the graphic recognition model; When the loss value is less than or equal to a preset loss threshold, the training of the graphic recognition model is stopped.
5. The method according to claim 4, characterized in that The preset loss functions include a classification loss function, a metric loss function and a distribution loss function.
6. The method according to claim 5, characterized in that Inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a preset loss function to obtain the loss value of the graphic recognition model includes: Inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a classification loss function to obtain a first loss value; Inputting the recognition result and the graphic type of the target graphic in the graphic training sample into a metric loss function to obtain a second loss value; Input the recognition result and the graphic type of the target graphic in the graphic training sample into the distribution loss function, Get the third loss value; A weighted sum of the first loss value, the second loss value, and the third loss value is obtained as a loss value of the graphic recognition model.
7. The method according to claim 5, characterized in that The at least one target candidate feature vector is obtained by selecting at least one candidate feature vector starting from the maximum similarity according to the size relationship.
8. The method according to claim 7, characterized in that Determining the type of the to-be-recognized graphic according to the graphic type of the at least one target candidate feature vector includes: Obtaining the frequency of occurrence of each graphic type in the graphic type of the at least one target candidate feature vector; In response to obtaining the maximum value of the frequency of each graphic type, the graphic type corresponding to the maximum value of the frequency is determined as the type of the graphic to be identified.
9. The method according to claim 8, characterized in that The method further comprises: In response to obtaining that the frequencies of at least two graphic types are both maximum values, the graphic type of the target candidate feature vector with the greatest similarity is determined as the type of the graphic to be identified.
10. The method according to claim 1, characterized in that Determining the type of the to-be-recognized graphic according to the graphic type of the at least one target candidate feature vector includes: The graphic type of the target candidate feature vector with the greatest similarity is determined as the type of the graphic to be identified.
11. The method according to claim 1, characterized in that: The method further comprises: In response to the similarity of each candidate feature vector being less than the preset similarity threshold or in response to detecting a registration operation, generating prompt information of a registration graphic and displaying it in an operation interface; In response to an input operation in the operation interface, acquiring an associated image and a graphic type of the initial image; Obtaining a feature vector of a graphic in the associated image; Establishing a matching relationship between the feature vector of the associated image and the graphic type; The feature vector, graphic type and matching relationship of the associated image are stored in the preset database.
12. The method according to claim 1, characterized in that The method further comprises: In response to the similarity of each candidate feature vector being less than the preset similarity threshold or in response to detecting a registration operation, generating prompt information of a registration graphic and displaying it in an operation interface; In response to an input operation in the operation interface, acquiring an associated image and a graphic type of the initial image; Obtaining a feature vector of a graphic in the associated image and an average vector of the feature vector; Establishing a matching relationship between the average vector and the graphic type; The average vector, graphic type and matching relationship are stored in the preset database.
13. The method according to claim 11 or 12, characterized in that: The associated image is obtained by: Randomly scaling the size of the initial image to obtain multiple scaled images; Performing blur processing on each zoomed image to serve as the associated image; and / or, Randomly rotating the direction of the initial image to obtain multiple rotated images; Mosaic processing is performed on each of the rotated images to obtain the associated images.
14. A graphic type recognition device, characterized in that: The device comprises: An initial image acquisition module, used to acquire an initial image to be recognized, wherein the initial image includes a graphic to be recognized; A feature vector acquisition module, used for acquiring a feature vector of a graphic to be identified in the initial image based on a preset graphic vector extraction model; A target vector acquisition module, used to obtain the similarity between the feature vector and each candidate feature vector in a preset database, and obtain at least one target candidate feature vector; The graphic type determination module is used to determine the type of the graphic to be identified according to the graphic type of the at least one target candidate feature vector.
15. An electronic device, characterized in that: include: processor; a memory for storing a computer program executable by the processor; The processor is configured to execute the computer program in the memory to implement the method according to any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that: When the executable computer program in the storage medium is executed by a processor, the method according to any one of claims 1 to 13 can be implemented.
Citation Information
Patent Citations
Face recognition method, device, computer device and storage medium
CN109241868A
Target recognition model training method and device and electronic equipment
CN112990432A
Identification method and device, training method and device, electronic equipment and storage medium
CN113642481A
Image classification method and device, computer equipment and medium
CN114741581A
Image recognition method and device, electronic equipment and computer readable storage medium
CN115331062A