An identification content identification method and device for an image, a computer device, and a medium
By selecting positive and negative sample images from candidate label images for feature fusion, the problem of low label content recognition efficiency is solved, and more efficient label content recognition is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2022-04-11
- Publication Date
- 2026-05-15
AI Technical Summary
During map navigation, the poor quality of the sign images and the uneven coverage of the sign content result in low sign recognition efficiency.
By selecting positive sample images that meet the similarity criteria with the target label image from the candidate label images and randomly selecting negative sample images, feature fusion is performed to obtain fused features that match the target label image. Matching images that meet the matching criteria are then selected to determine the label content.
It improves the accuracy and efficiency of identifying signage content, and solves the problem of low identification efficiency.
Smart Images

Figure CN116955676B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, computer-readable storage medium, and computer program product for recognizing the content of an image. Background Technology
[0002] With the development of map navigation technology, users' demand for road information during map navigation is gradually increasing. Therefore, it is necessary to identify and mark various signs in the road scene used for map navigation so that users can have a comprehensive understanding of the road conditions and scene status during map navigation.
[0003] However, during the scene data collection process, due to limitations in the collection range and scene conditions, problems such as poor image quality and uneven coverage of the label content may occur, resulting in low recognition efficiency when identifying the label content in the image. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for identifying the content of signage images, which can improve the recognition efficiency when identifying the content of signage images, in order to address the above-mentioned technical problems.
[0005] Firstly, this application provides a method for recognizing the content of an image. The method includes:
[0006] From the candidate icon images, select positive sample images that meet the similarity condition with the target icon image, and randomly select negative sample images;
[0007] The target features of the target identification image, the positive relationship features of the target identification image relative to the positive sample image, and the negative relationship features of the target identification image relative to the negative sample image are fused to obtain a fused feature that matches the target identification image.
[0008] From the candidate identifier images, select matching images whose image features and fused features meet the matching conditions;
[0009] The identifier content contained in the matching image is determined as the identifier content of the target identifier image.
[0010] Secondly, this application also provides a device for recognizing the content of an image. The device includes:
[0011] The image selection module is used to select positive sample images from the candidate label images that meet the similarity condition with the target label image, and to randomly select negative sample images.
[0012] The feature fusion module is used to fuse the target features of the target identification image, the positive relationship features of the target identification image relative to the positive sample image, and the negative relationship features of the target identification image relative to the negative sample image to obtain fused features that match the target identification image.
[0013] The image matching module is used to filter out matching images from the candidate identifier images whose image features and fused features meet the matching conditions.
[0014] The result determination module is used to determine the identifier content contained in the matching image as the identifier content of the target identifier image.
[0015] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described above.
[0016] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0017] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the above-described method.
[0018] The aforementioned method, apparatus, computer equipment, computer-readable storage medium, and computer program product for identifying the content of identified images, by selecting positive sample images from candidate imagery that meet the similarity criteria with the target image, and randomly selecting negative sample images, fuses the target features of the target image, the positive relationship features of the target image relative to the positive sample images, and the negative relationship features of the target image relative to the negative sample images to obtain fused features that match the target image. This allows for spatial correction of the target features using positive and negative relationship features, enriching the semantic information of the target features and thus improving matching accuracy. Then, by selecting matching images from the candidate imagery whose image features and fused features meet the matching criteria, and determining the identification content contained in the matching images as the identification content of the target image, the efficiency of identifying identification content is improved, thereby increasing the recognition efficiency when identifying the identification content of imagery. Attached Figure Description
[0019] Figure 1This is an application environment diagram of a method for recognizing the content of an image in one embodiment;
[0020] Figure 2 This is a flowchart illustrating a method for recognizing the identifier content of an image in one embodiment;
[0021] Figure 3 This is a schematic diagram illustrating the determination of a target identifier image from a target image in one embodiment;
[0022] Figure 4 This is a schematic diagram of the candidate identifier region image obtained in one embodiment;
[0023] Figure 5 This is a schematic diagram of a method for recognizing the content of an image in a specific embodiment;
[0024] Figure 6 This is a schematic diagram of the target image and the target identifier image in a specific embodiment;
[0025] Figure 7 This is a schematic diagram of a candidate identifier image in a specific embodiment;
[0026] Figure 8 This is a structural block diagram of a device for recognizing the content of an image in one embodiment;
[0027] Figure 9 This is an internal structural diagram of a computer device in one embodiment;
[0028] Figure 10 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0030] It should be noted that the image data involved in this application, including but not limited to candidate logo images, target logo images and target images, are all information and data that have been fully authorized by all parties, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0031] In one embodiment, the method for recognizing the content of an identifier image provided in this application can be applied to, for example... Figure 1In the illustrated application environment, both terminal 102 and server 104 are involved. In some embodiments, terminal 106 is also involved. Terminals 102 and 106 can communicate with server 104 via a network or other means. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located in the cloud or on another server.
[0032] Specifically, server 104 obtains the target identifier image from terminal 102. Then, server 104 selects positive sample images from candidate identifier images that meet the similarity criteria with the target identifier image, and randomly selects negative sample images. The candidate identifier images can be images pre-stored by server 104, or images pre-obtained by server 104 from terminal 102 and / or terminal 106. Next, server 104 fuses the target features of the target identifier image, the positive relationship features of the target identifier image relative to the positive sample images, and the negative relationship features of the target identifier image relative to the negative sample images to obtain fused features that match the target identifier image. Server 104 then filters out matching images from the stored candidate identifier images whose image features match the fused features. The identifier content contained in the matching image is determined as the identifier content of the target identifier image. Thus, server 104 can send the identifier content of the target identifier image to terminal 102 and / or terminal 106 for display on terminal 102 and / or terminal 106.
[0033] In one embodiment, when the data processing capability of terminal 102 meets the data processing conditions, the application environment may only involve terminal 102. Specifically, terminal 102 pre-stores candidate identifier images. After obtaining the target identifier image, terminal 102 selects positive sample images from the candidate identifier images that meet the similarity condition with the target identifier image, and randomly selects negative sample images. Then, it processes these images to obtain fusion features that match the target identifier image. Next, it selects matching images from the stored candidate identifier images whose matching degree between the image features and the fusion features meets the matching condition. Thus, the identifier content contained in the matching image is determined as the identifier content of the target identifier image, and the identifier content is displayed.
[0034] Terminals 102 and 106 may be equipped with camera devices to capture images of the target identifier. These devices include, but are not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, smart voice interaction devices, in-vehicle terminals, aircraft, and portable wearable devices. IoT devices may include smart TVs and smart in-vehicle devices. Portable wearable devices may include smartwatches, smart bracelets, and head-mounted devices. Server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers. This invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0035] In one embodiment, such as Figure 2 As shown, a method for recognizing the labeled content of an image is provided, which can be applied to... Figure 1 Taking server 104 as an example, the following steps are included:
[0036] Step S202: Select positive sample images from the candidate label images that meet the similarity condition with the target label image, and randomly select negative sample images.
[0037] Here, a labeled image refers to an image containing labeled content. A target labeled image refers to an image containing labeled content to be identified, and a candidate labeled image refers to an image containing known labeled content. The known labeled content can be attribute information such as the name, quantity, and type of the labeled content.
[0038] Specifically, candidate label images can be obtained from publicly available datasets, or they can be obtained by manually labeling the recognition content contained in the image or by using a pre-trained neural network model. Candidate label images can be images acquired in real time before each execution of step S202, or they can be images that have been pre-stored; there are no restrictions on this.
[0039] In specific applications, the type of labeling content can vary depending on the actual application scenario. The type of labeling content should be matched with the application scenario and determined based on actual technical needs. For example, in a road scenario for map navigation, the labeling content can include buildings, vehicles, roads, and road-related facilities, including but not limited to streetlights, road signs, traffic lights, and traffic cameras. As another example, in a customer information collection scenario, the labeling content can include a customer's facial image, shopping basket, and items in the shopping basket.
[0040] Similarity refers to the degree of similarity between images, and it can be a specific numerical value. The similarity between a target image and a candidate image can be determined in any possible way. For example, a pre-trained neural network model can be used to determine the type of the target image, and the similarity can be determined based on the type of the target image and the type of the candidate image. The more consistent the types, the more similar the target image and the candidate image. Alternatively, at least one of the following can be calculated between the target image and the candidate image: Euclidean distance, cosine similarity, or correlation coefficient. The smaller the Euclidean distance, the closer the cosine similarity is to 1, and the closer the correlation coefficient is to 1, the more similar the target image and the candidate image are.
[0041] Positive sample images refer to candidate label images whose similarity to the target label image meets the similarity condition, while negative samples refer to randomly selected candidate label images. Specifically, from the candidate label images, positive sample images that meet the similarity condition with the target label image are selected, and negative sample images are randomly selected. The similarity condition can be set according to actual technical needs; for example, it can be determined according to the method of calculating similarity. Taking the calculation of cosine similarity between the target label image and candidate label images as an example, the similarity condition can be set to a cosine similarity greater than a similarity threshold. That is, candidate label images with a cosine similarity greater than the similarity threshold are determined as positive sample images.
[0042] It should be noted that the number of positive sample images can be one or more, forming a set of positive sample images. If there is only one positive sample image, this image can be used in subsequent processing steps. If there are multiple positive sample images, the most similar positive sample image must be used in subsequent processing steps. That is, the maximum similarity is determined based on the similarity scores of the multiple positive sample images, and the positive sample image corresponding to the maximum similarity score is taken as the most similar positive sample image for subsequent processing steps.
[0043] In practical applications, taking a road scenario in map navigation as an example, where the marker content is traffic elements, the candidate marker image is the candidate traffic element image, and the target marker image is the target traffic element image. From the candidate marker images, positive sample images that meet the similarity condition with the target marker image are selected, and negative sample images are randomly selected. This can specifically include: from the candidate traffic element images, selecting positive sample images that meet the similarity condition with the target traffic element image, and randomly selecting negative sample images.
[0044] Step S204: The target features of the target image, the positive relationship features of the target image relative to the positive sample image, and the negative relationship features of the target image relative to the negative sample image are fused to obtain fused features that match the target image.
[0045] The features of the target label image are referred to as target features. The image features jointly determined by the target label image and the positive sample image are called the positive relationship features of the target label image relative to the positive sample image, and the image features jointly determined by the target label image and the negative sample image are called the negative relationship features of the target label image relative to the negative sample image.
[0046] Feature fusion refers to the optimized combination of different image features. The image features obtained by fusing target features, positive relation features, and negative relation features are called fused features that match the target image. Through feature fusion, the distance between the target features of the target image and the positive relation feature vectors of positive sample images is reduced, while the distance between the target features and the negative relation feature vectors of negative sample images is increased. This enriches the semantic information of the target features of the target image, resulting in richer fused features and thus improving the accuracy of subsequent feature matching based on these fused features.
[0047] It should be noted that, in order to ensure the correctness of feature fusion, the feature dimensions of the target feature, the positive relation feature, and the negative relation feature should be the same.
[0048] Step S206: From the candidate identifier images, select matching images whose image features and fused features meet the matching conditions.
[0049] Matching degree refers to the degree of matching between images, and it can be a specific numerical value. Matching degree can be the matching degree between the image features of the candidate label image and the fused features of the target label image, or it can be the matching degree between the image features of the positive sample images in the positive sample image set and the fused features of the target label image. The specific choice needs to be determined based on actual technical requirements. For example, it can be determined based on matching time and matching rate. Since the positive sample image set is a part of the candidate label images, meaning the number of positive sample images is less than or equal to the number of candidate label images, theoretically, the matching time and matching rate corresponding to traversing the image features of all candidate label images are greater than or equal to the matching time and matching rate corresponding to traversing the image features of all positive sample images in the positive sample image set. When there are no requirements for matching time and matching rate, either matching method can be chosen.
[0050] Specifically, the matching degree between image features and fused features can be determined in any possible way. For example, at least one of cosine similarity or correlation coefficient between image features and fused features can be calculated. The closer the cosine similarity is to 1 and the closer the correlation coefficient is to 1, the higher the matching degree between image features and fused features. A matched image refers to a candidate label image whose matching degree between image features and fused features meets the matching conditions. The matching conditions can be set according to actual technical needs; for example, they can be determined according to the method of calculating the matching degree. Taking the calculation of cosine similarity between image features and fused features as an example, the matching condition can be set to a cosine similarity greater than a preset threshold. That is, candidate label images with a cosine similarity greater than the preset threshold are determined as matched images.
[0051] It should be noted that the number of matching images can be determined to be one based on actual technical needs, so that the identifier content of the target identifier image can be determined based on the identifier content contained in the matching images.
[0052] Step S208: The identifier content contained in the matching image is determined as the identifier content of the target identifier image.
[0053] Since candidate signage images contain known signage content, after determining the matching image, the signage content contained in the matching image can be identified as the signage content of the target signage image. Taking a road scenario applied in map navigation, where the signage content is traffic elements, as an example, the candidate signage image is the candidate traffic element image, the target signage image is the target traffic element image, and the matching image is the matching traffic element image. Identifying the signage content contained in the matching image as the signage content of the target signage image can specifically include: identifying the traffic element content contained in the matching traffic element image as the traffic element content of the target traffic element image.
[0054] In the aforementioned method for recognizing the content of a labeled image, positive sample images that meet the similarity criteria with the target labeled image are selected from candidate labeled images, and negative sample images are randomly selected. The target features of the target labeled image, the positive relationship features of the target labeled image relative to the positive sample images, and the negative relationship features of the target labeled image relative to the negative sample images are fused to obtain fused features that match the target labeled image. This allows for spatial correction of the target features using positive and negative relationship features, enriching the semantic information of the target features and thus improving matching accuracy. Then, by selecting matching images from the candidate labeled images whose image features and fused features meet the matching criteria, and identifying the labeled content contained in the matching images as the labeled content of the target labeled image, the efficiency of identifying labeled content is improved, thereby increasing the recognition efficiency when recognizing the labeled content of a labeled image.
[0055] In one embodiment, the negative sample image is an image randomly selected from the candidate label images, while step S202, which selects positive sample images from the candidate label images that meet the similarity condition with the target label image, may include steps S302 to S306:
[0056] Step S302: Obtain the target features of the target identification image and the image features of the candidate identification image.
[0057] In this process, feature extraction is performed on the target identification image to obtain its target features. If the candidate identification image is an image acquired in real time before each execution of step S202, the same or different feature extraction methods as the target identification image can be used to extract features from the candidate identification image to obtain its image features. If the candidate identification image is a pre-stored image, its image features can be stored accordingly when storing the candidate identification image. That is, the association between the candidate identification image and its image features is established and stored, so that the image features of the candidate identification image can be determined based on the association.
[0058] In one embodiment, the method for extracting features from the target label image or candidate label image may be to use a scale-invariant feature transform (SIFT) feature extraction algorithm, a histogram of oriented gradients (HOG) feature extraction algorithm, or a pre-trained neural network model for feature extraction. The specific method can be selected and determined according to the actual technical needs.
[0059] Step S304: Calculate the similarity between the target feature and each image feature to obtain the similarity results.
[0060] Specifically, the image features of all candidate identifier images are traversed, and the similarity between the image features and the target features is calculated to obtain various similarity results. The similarity calculation can be performed using any feasible method, such as calculating the cosine similarity or correlation coefficient between the target features and each image feature to obtain the various similarity results.
[0061] Step S306: The candidate identifier image corresponding to the maximum similarity result among the similarity results is determined as the positive sample image.
[0062] Specifically, the number of positive sample images is determined to be one, and the candidate label image corresponding to the maximum similarity result among the similarity results is determined as the positive sample image so that the target label image can be processed subsequently.
[0063] In this embodiment, by setting candidate identifier images and selecting positive and negative sample images of the target identifier image from the candidate identifier images, subsequent image processing steps are performed based on the ternary image data structure composed of the target identifier image, positive sample image, and negative sample image. This can solve the problem of poor recognition accuracy caused by uneven coverage of the identifier content in the target identifier image.
[0064] In one embodiment, the target identification image can be obtained from a target image. In a specific application, taking a road scenario applied to map navigation as an example, where the identification content is traffic elements, the target identification image can be a target traffic element image, and the target image can be a target road image.
[0065] The target image can be an image captured in real time by the camera device, or an image captured historically. If the camera device captures video, the target image can be an image obtained by extracting frames from the video at a preset frame extraction frequency, or an image obtained by extracting frames from the video at a preset driving distance. The preset frame extraction rate and preset driving distance can be set according to actual technical needs; for example, it can be set to extract 2 frames per second or 1 frame every 5 meters traveled.
[0066] Since the target road image contains various traffic elements, it is necessary to determine a target signage image from the target image. Specifically, the method of obtaining the target signage image may include the following steps S402 to S406:
[0067] Step S402: Perform feature point recognition on the target image to determine each feature point of the target image.
[0068] Feature point identification is performed on the target image to determine its various feature points. This feature point identification can be performed using a pre-trained neural network model, such as a convolutional neural network (CNN).
[0069] Specifically, please refer to Figure 3The target image is input into a convolutional neural network (CNN), where feature maps are extracted through feature extraction layers. These feature maps represent the characteristics of the target image as a series of feature points. The feature extraction layers consist of convolutional layers, batch normalization layers, and activation layers. The activation function can be a Rectified Linear Function (RELU). The convolutional layers extract basic features such as edge textures from the target image. The batch normalization layers normalize the features extracted by the convolutional layers according to a normal distribution, filtering out noisy features and accelerating model training convergence. The activation layers perform non-linear mapping on the features extracted by the convolutional layers, enhancing the model's generalization ability.
[0070] Step S404: According to the region division parameters that match the feature points, the image is divided into multiple candidate label regions corresponding to each feature point, with each feature point as the center.
[0071] Based on the extracted feature points, region segmentation parameters matching the feature points are determined. These parameters include the region segmentation ratio and the region segmentation scale. Then, according to these region segmentation parameters, multiple candidate labeling regions corresponding to each feature point are obtained, centered on each feature point. The region segmentation parameters can be determined based on the position of the feature point in the target image, or they can be pre-determined based on the labeling content to be identified in the application scenario.
[0072] In a specific application, taking a road scenario in map navigation as an example, where the identified content is traffic elements, please refer to [link to relevant documentation]. Figure 3 After determining the feature points of the target image, region segmentation is required. Please refer to [link / reference]. Figure 4 Based on the known length-to-width ratio of traffic elements, the region division ratio in the region division parameters can be set to three types: 1:1, 2:1, and 1:2. The region division scale can be set to 1, 2, and 3 feature points. That is, for each feature point, 9 candidate identification region images are obtained.
[0073] It should be noted that the above-mentioned region division parameters are predetermined based on the identification content to be identified in the application scenario. The region division ratio and region division scale involved are only examples. In actual determination of region division parameters, the candidate identification region image obtained by division can be larger or smaller.
[0074] Step S406: Based on the scores corresponding to each candidate labeling region image, select the target labeling region image from each candidate labeling region image, and determine the target labeling region image as the target labeling image.
[0075] After dividing the image into multiple candidate labeling regions corresponding to each feature point, a score is predicted for each candidate labeling region. Based on the scores of each candidate labeling region, the target labeling region image can be selected. A score threshold can be set, and candidate labeling region images with scores greater than the threshold can be identified as the target labeling region image. Alternatively, the candidate labeling region image with the highest score can be selected as the target labeling region image.
[0076] It should be noted that since the target road image contains various traffic elements, if multiple target sign images can be identified from the target image, the processing steps of this application embodiment need to be performed for each target sign image. When performing the processing steps of this application embodiment for each target sign image, they can be performed sequentially or simultaneously, depending on actual technical needs.
[0077] In this embodiment, by determining the target identification image from the target image, subsequent data processing of the target identification image can be performed in a targeted manner, thereby improving the efficiency of image processing of the target identification image.
[0078] In one embodiment, the positive relationship features of the target identification image relative to the positive sample image and the negative relationship features of the target identification image relative to the negative sample image need to be determined based on the positive sample features and negative sample features corresponding to the positive sample image and the negative sample image, respectively.
[0079] For ease of description, both positive and negative sample images are referred to as sample images. The determination of the positive sample features of the positive sample image and the negative sample features of the negative sample image for each of the positive and negative sample images can specifically include steps S502 to S506:
[0080] Step S502: Determine the sub-image features under each feature dimension of the sample image and the sub-target features under each feature dimension of the target identifier image; the image feature dimensions of the sample image are the same as the image feature dimensions of the target identifier image.
[0081] The sample images include positive and negative sample images. The image feature dimensions of the sample images are the same as those of the target label image; that is, the image feature dimensions of both the positive and negative sample images are identical to those of the target label image. The initial image features in each dimension of the sample image are called sub-image features, and the target features in each dimension of the target label image are called sub-target features. Specifically, the sub-image features in each feature dimension of the sample image and the sub-target features in each feature dimension of the target label image are determined.
[0082] Step S504: Determine the feature weight of each dimension corresponding to the sample image based on the sub-image features of the sample image and the sub-target features of the target identifier image under the same feature dimension.
[0083] Each dimension of the sample image has a corresponding feature weight. This weight can be determined by comparing the sub-image features of the sample image with the sub-target features of the target image within the same feature dimension. Alternatively, the feature weights for each dimension can be pre-set based on the image features of the sample image. Specifically, the feature weights for each dimension of the positive sample image are determined by comparing the sub-image features of the positive sample image with the sub-target features of the target image within the same feature dimension, and the feature weights for each dimension of the negative sample image are determined by comparing the sub-image features of the negative sample image with the sub-target features of the target image within the same feature dimension.
[0084] In one embodiment, the initial image features of the positive sample image are represented as [P1, P2…P ... m ] m∈Z Where Z represents the image feature dimension of the positive sample image, m represents the m-th dimension, and the sub-image features under each feature dimension of the positive sample image are P1, P2, ..., Pm, respectively. m The initial image features of the negative sample images can be represented as [N1, N2…N]. m ] m∈Z Where Z represents the image feature dimension of the negative sample image, and m represents the m-th dimension. The sub-image features under each feature dimension of the negative sample image are N1, N2, ..., N. m The target features of the target identification image can be represented as [K1, K2…K]. n ] n∈Z Where Z represents the image feature dimension of the target identification image, n represents the nth dimension, and m corresponds to n. The sub-target features under each feature dimension of the target identification image are K1, K2, ..., Kn. n The feature weights for each dimension corresponding to the positive sample images are represented as a1, a2, ..., a1. mThe value of a1 is determined based on K1 and P1, the value of a2 is determined based on K2 and P2, and so on. The feature weights for each dimension corresponding to the negative sample image are represented as b1, b2, ..., b... m The value of b1 is determined by K1 and N1, the value of b2 is determined by K2 and N2, and so on.
[0085] It should be noted that the feature weights for each dimension corresponding to the sample image are all positive numbers. If it is determined that there are non-positive feature weight values in a certain dimension, the feature weights in that dimension will be reset to 0, meaning that the sub-image features in that dimension will not participate in feature fusion in subsequent processes.
[0086] Step S506: Based on the image features of the sample image and the feature weight of each dimension corresponding to the sample image, the image features of the sample image are weighted to obtain the positive sample features of the positive sample image and the negative sample features of the negative sample image.
[0087] The weighted processing of the initial image features of the sample image refers to multiplying the sub-image features of each dimension of the sample image with the corresponding feature weights to obtain the positive sample features of the positive sample image and the negative sample features of the negative sample image.
[0088] In one embodiment, the positive sample features of a positive sample image are represented as [a1P1, a2P2…a ... m P m ] m∈Z The negative sample features of the negative sample image are represented as [b1N1, b2N2…b m N m ] m∈Z .
[0089] In this embodiment, the positive sample features of the positive sample image and the negative sample features of the negative sample image are determined by using the target features of the target identification image and the initial image features of the sample image. In the subsequent feature fusion process, the positive sample features and negative sample features can be used to spatially correct the target features, thereby more robustly and accurately identifying the identification content of the target identification image.
[0090] In one embodiment, feature fusion is performed on the target features of the target identification image, the positive relationship features of the target identification image relative to the positive sample image, and the negative relationship features of the target identification image relative to the negative sample image to obtain fused features that match the target identification image. This may include the following steps S602 to S608:
[0091] Step S602: Obtain the positive sample features of the positive sample image and the negative sample features of the negative sample image.
[0092] Specifically, the positive sample features of the positive sample image and the negative sample features of the negative sample image determined in step S508 are obtained.
[0093] Step S604: Determine the positive relationship features based on the positive fusion coefficient matched by the positive sample image and the positive sample features.
[0094] The fusion coefficient matched by the positive sample image is called the forward fusion coefficient. Specifically, the specific value of the forward fusion coefficient can be determined according to actual technical needs, and the forward fusion coefficient is a positive number. For example, in this embodiment, the forward fusion coefficient can be set to 1.
[0095] Step S606: Determine the reverse relationship features based on the reverse fusion coefficients matched by the negative sample images and the negative sample features.
[0096] The fusion coefficient matched with the negative sample image is called the reverse fusion coefficient. Specifically, the value of the reverse fusion coefficient can be determined according to actual technical needs, and the reverse fusion coefficient is negative. For example, in this embodiment, the reverse fusion coefficient can be set to -1. It should be noted that the forward fusion coefficient and the reverse fusion coefficient are independent of each other.
[0097] Step S608: Summation is performed on the target features, positive relation features, and negative relation features to obtain fused features that match the target identification image.
[0098] By summing the target features, forward fusion features, and reverse fusion features, a fusion feature matching the target identifier image can be obtained. Specifically, the formula for calculating the fusion feature is expressed as:
[0099]
[0100] Among them, F Z This represents the fusion feature that matches the target image, with a feature dimension of Z.
[0101] In this embodiment, since the labeled content of the positive and negative sample images is known, fusing the target features based on positive and negative relational features is equivalent to adding prior knowledge to the target features, which can improve the recognition effect of the labeled content. Through the above feature fusion method, the distance between the target features and the positive and negative relational features can be constrained. Specifically, the positive relational features generated by the positive sample images act as an attraction, bringing the target features closer to the ground truth; conversely, the negative relational features generated by the negative sample images act as a repulsion, distancing the target features from other non-ground truths, thus improving the accuracy of the fused features and thereby increasing the recognition efficiency of identifying the labeled content of the target image.
[0102] In one embodiment, after determining the fusion features that match the target identifier image, matching images whose image features and fusion features meet the matching conditions can be selected from the candidate identifier images. This can specifically include the following steps S702 to S704:
[0103] Step S702: Calculate the matching degree between the fused features and the image features of the candidate identifier image.
[0104] This can be achieved by traversing the image features of all candidate label images to select a matching image. Specifically, the matching degree between the fused features and the image features of the candidate label images is calculated. This can be done using a spatial norm matching algorithm, that is, by calculating the cosine similarity between the fused features and the image features of the candidate label images to determine each matching degree.
[0105] Step S704: Determine the maximum matching degree among all matching degrees, and use the candidate identifier image corresponding to the maximum matching degree as the matching image.
[0106] The number of matching images is determined to be one. Specifically, the maximum matching degree among all matching degrees is determined, and the candidate label image corresponding to the maximum matching degree is used as the matching image. Since the label content contained in the candidate label image is known, that is, the label content contained in the matching image is known, the label content of the target label image can be directly determined based on the label content contained in the matching image.
[0107] In this embodiment, by traversing all candidate label images, the matching image closest to the target label image is determined from the candidate label images based on the matching degree between the fused features and the image features of the candidate label images. Thus, the label content of the target label image can be determined, improving the accuracy and efficiency of the label content recognition of the target label image.
[0108] In one embodiment, after determining the fusion features that match the target identifier image, matching images whose image features and fusion features meet the matching conditions can be selected from the candidate identifier images. This can specifically include the following steps S802 to S808:
[0109] Step S802: Obtain the similarity results of the target features of the target image and the image features of the candidate image.
[0110] In this way, matching images can be selected from all positive sample images by traversing the image features of multiple positive sample images in all positive sample image sets, thereby further improving the matching efficiency.
[0111] Specifically, the similarity results of the target features of the target image and the image features of the candidate image are obtained, that is, the similarity calculation results obtained in step S304 are obtained without repeated calculation, so as to determine the set of positive sample images from the candidate image.
[0112] Step S804: From the various similarity results, identify multiple similarity results that meet the screening conditions, and determine the candidate identifier images corresponding to the multiple similarity results as candidate positive sample images.
[0113] Candidate positive sample images refer to candidate identifier images that meet the screening criteria, i.e., multiple positive sample images in the aforementioned set of positive sample images. The screening criteria can be set according to actual technical needs. For example, a preset similarity threshold can be set; if the similarity result is greater than or equal to the preset similarity threshold, then the similarity result is determined to meet the screening criteria. Specifically, from each similarity result, multiple similarity results that meet the screening criteria are identified, and the candidate identifier images corresponding to these multiple similarity results are determined as candidate positive sample images.
[0114] Step S806: Calculate the matching degree between the fused features and the image features of each candidate positive sample image.
[0115] Specifically, the matching image is determined by calculating the matching degree between the fused features and the image features of each candidate positive sample image. The matching degree can be calculated using a spatial norm matching algorithm, that is, by calculating the cosine similarity between the fused features and the image features of the candidate positive sample image.
[0116] Step S808: Determine the maximum matching degree among all matching degrees, and use the candidate positive sample image corresponding to the maximum matching degree as the matching image.
[0117] The number of matching images is determined to be one. Specifically, the maximum matching degree among all matching degrees is determined, and the candidate positive sample image corresponding to the maximum matching degree is used as the matching image. Since the label content contained in the candidate positive sample image is known, that is, the label content contained in the matching image is known, the label content of the target label image can be directly determined based on the label content contained in the matching image.
[0118] In this embodiment, by determining candidate positive sample images and traversing all candidate positive sample images, the matching image that is closest to the target label image is determined. Thus, the label content of the target label image can be determined, further improving the accuracy and efficiency of the label content recognition of the target label image.
[0119] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description, in conjunction with the accompanying drawings and a specific embodiment, will further illustrate this application. It should be understood that the specific embodiment described herein is merely illustrative and not intended to limit the scope of this application.
[0120] In practical applications, taking a road scenario in map navigation as an example, the labeled content is traffic elements, which can specifically be road signs, such as... Figure 5 The diagram illustrates a method for recognizing the content of a sign image. This method mainly includes a series of steps such as feature extraction, feature fusion, and feature matching on the target traffic element image, positive sample images, and negative sample images of the target traffic element. The specific steps are as follows:
[0121] Acquiring the target image: In one embodiment, a camera device is mounted at the front of the vehicle, and the camera device's shooting angle, shooting height, and other parameters remain fixed. The camera device can capture video of the road scene while the vehicle is in motion. The driving video can be captured in real time or from historical footage. The target image is an image obtained by extracting frames from the video according to a preset frame extraction frequency, or it can be an image obtained by extracting frames from the video according to a preset driving distance. The preset frame extraction rate and preset driving distance can be set according to actual technical needs; for example, it can be set to extract 2 frames per second or 1 frame every 5 meters traveled. Please refer to the schematic diagram of the target image. Figure 6 As shown in Figure 6-1, Figure 6-1 It includes various elements such as vehicles, streetlights, green belts, and signs.
[0122] The target image is input into a pre-trained convolutional neural network, which identifies feature points in the target image to determine each feature point. According to the region division parameters that match the feature points, the image is divided into multiple candidate traffic element region images corresponding to each feature point, with each feature point as the center. In one embodiment, the region division ratio in the region division parameters is determined to be 1:1, 2:1, and 1:2 based on the length-to-width ratio of the sign. The region division scale is 1, 2, and 3 feature points, that is, for each feature point, 9 candidate traffic element region images are obtained.
[0123] Based on the scores corresponding to each candidate traffic element region image, a target traffic element region image is selected from the candidate traffic element region images and determined as the target traffic element image; in one embodiment, please refer to Figure 6 As shown in Figure 6-2, the area marked by the dashed box is the target traffic element image that needs to be identified for traffic element content recognition.
[0124] Feature extraction is performed on the target traffic element image and the candidate traffic element images to obtain the target features of the target traffic element image and the image features of the candidate traffic element images; in one embodiment, please refer to Figure 7 Using a publicly available dataset, schematic diagrams of all signs within the traffic element are pre-acquired and stored, and these schematic diagrams are identified as candidate traffic element images. That is, the traffic elements in the candidate traffic element images are known. In some embodiments, to facilitate the storage of candidate traffic element images, they can also be categorized and stored according to the meaning represented by the signs, wherein... Figure 7 The image shows the danger type (7-1), the prohibition type (7-2), and the passable type (7-3). The corresponding number is shown below the image.
[0125] The similarity between the target feature and each image feature is calculated to obtain the similarity results. In one embodiment, the cosine similarity between the target feature and each image feature can be calculated to obtain the similarity results. The candidate traffic element image corresponding to the largest similarity result among the similarity results is determined as the positive sample image. Then, a negative sample image is randomly selected from the candidate traffic element images.
[0126] For each sample image in the positive and negative sample images, sub-image features under each feature dimension of the sample image and sub-target features under each feature dimension of the target traffic element image are determined; wherein, the image feature dimension of the sample image is the same as the image feature dimension of the target traffic element image, so as to perform feature fusion.
[0127] Based on the sub-image features of the sample image and the sub-target features of the target traffic element image under the same feature dimension, the feature weight of each dimension corresponding to the sample image is determined; in one embodiment, the initial image features of the positive sample image are represented as [P1, P2…P ... m ] m∈Z Where Z represents the image feature dimension of the positive sample image, m represents the m-th dimension, and the sub-image features under each feature dimension of the positive sample image are P1, P2, ..., Pm, respectively. m The initial image features of the negative sample images are represented as [N1, N2…N…]. m ] m∈Z Where Z represents the image feature dimension of the negative sample image, and m represents the m-th dimension. The sub-image features under each feature dimension of the negative sample image are N1, N2, ..., N. m The target features of the target traffic element image can be represented as [K1, K2…K]. n ] n∈ZWhere Z represents the image feature dimension of the target traffic element image, n represents the nth dimension, and m corresponds to n. The sub-target features under each feature dimension of the target traffic element image are K1, K2, ..., Kn. n The feature weights for each dimension corresponding to the positive sample images are represented as a1, a2, ..., a1. m The value of a1 is determined based on K1 and P1, the value of a2 is determined based on K2 and P2, and so on. The feature weights for each dimension corresponding to the negative sample image are represented as b1, b2, ..., b... m The value of b1 is determined by K1 and N1, the value of b2 is determined by K2 and N2, and so on.
[0128] Based on the image features of the sample images and the feature weights of each dimension corresponding to the sample images, the image features of the sample images are weighted to obtain the positive sample features of the positive sample images and the negative sample features of the negative sample images of the target traffic element images; in one embodiment, the positive sample features of the positive sample images of the target traffic element images are represented as [a1P1, a2P2…a ... m P m ] m∈Z The negative sample features of the negative sample image are represented as [b1N1, b2N2…b m N m ] m∈Z .
[0129] Determine the positive fusion coefficient matched by the positive sample image and the reverse fusion coefficient corresponding to the negative sample image; in one embodiment, the positive fusion coefficient and the reverse fusion coefficient can be set according to actual technical needs, and the positive fusion coefficient can be set to 1 and the reverse fusion coefficient can be set to -1.
[0130] Based on the forward fusion coefficient and positive sample features, forward relationship features are determined; based on the reverse fusion coefficient and negative sample features, reverse relationship features are determined. The target features, forward relationship features, and reverse relationship features are summed to obtain the fusion features that match the target traffic element image. In one embodiment, the calculation formula for the fusion features is expressed as:
[0131]
[0132] Among them, F Z This represents the fusion feature that matches the target traffic element image, with a feature dimension of Z.
[0133] The matching degree between the fused features and the image features of the candidate traffic element images is calculated; alternatively, the similarity results between the target features of the target traffic element image and the image features of the candidate traffic element images are obtained; from the similarity results, multiple similarity results that meet the screening conditions are determined, and the candidate traffic element images corresponding to the multiple similarity results are determined as candidate positive sample images; the matching degree between the fused features and the image features of each candidate positive sample image is calculated; in one embodiment, the screening conditions can be set according to actual technical needs. For example, a preset similarity threshold can be set, and if the similarity result is greater than or equal to the preset similarity threshold, then the similarity result is determined to meet the screening conditions.
[0134] Determine the maximum matching degree among all matching degrees, and use the candidate traffic element image corresponding to the maximum matching degree as the matched traffic element image; or, determine the maximum matching degree among all matching degrees, and use the candidate positive sample image corresponding to the maximum matching degree as the matched traffic element image.
[0135] The traffic element content contained in the matched traffic element image is determined as the traffic element content of the target traffic element image. Subsequently, driving information prompts can be given to vehicle drivers based on the traffic element content.
[0136] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0137] Based on the same inventive concept, this application also provides a device for recognizing the content of a sign image to implement the aforementioned method for recognizing the content of a sign image. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the device for recognizing the content of a sign image provided below can be found in the limitations of the method for recognizing the content of a sign image described above, and will not be repeated here.
[0138] In one embodiment, such as Figure 8As shown, a device for recognizing the content of an image is provided, comprising: an image selection module 810, a feature fusion module 820, an image matching module 830, and a result determination module 840, wherein:
[0139] The image selection module 810 is used to select positive sample images from the candidate label images that meet the similarity condition with the target label image, and to randomly select negative sample images.
[0140] The feature fusion module 820 is used to fuse the target features of the target identification image, the positive relationship features of the target identification image relative to the positive sample image, and the negative relationship features of the target identification image relative to the negative sample image to obtain fused features that match the target identification image.
[0141] The image matching module 830 is used to filter out matching images from the candidate identifier images whose image features and fused features meet the matching conditions.
[0142] The result determination module 840 is used to determine the identifier content contained in the matching image as the identifier content of the target identifier image.
[0143] In one embodiment, the image selection module 810 is further configured to acquire target features of the target identifier image and image features of the candidate identifier image; calculate the similarity between the target features and each of the image features to obtain each similarity result; and determine the candidate identifier image corresponding to the largest similarity result among the similarity results as the positive sample image.
[0144] In one embodiment, the image selection module 810 is further configured to perform feature point recognition on the target image to determine each feature point of the target image; according to the region division parameters matching the feature points, divide the target image into multiple candidate identifier region images centered on each feature point; based on the scores corresponding to each candidate identifier region image, select the target identifier region image from the candidate identifier region images, and determine the target identifier region image as the target identifier image.
[0145] In one embodiment, the feature fusion module 820 is configured to acquire positive sample features of the positive sample image and negative sample features of the negative sample image; determine the positive relationship feature based on the positive fusion coefficient matched by the positive sample image and the positive sample feature; determine the reverse relationship feature based on the reverse fusion coefficient matched by the negative sample image and the negative sample feature; and sum the target feature, the positive relationship feature, and the reverse relationship feature to obtain a fusion feature that matches the target identifier image.
[0146] In one embodiment, for each of the positive and negative sample images, the feature fusion module 820 is used to determine sub-image features under each feature dimension of the sample image and sub-target features under each feature dimension of the target identifier image; the image feature dimensions of the sample image are the same as those of the target identifier image; based on the sub-image features of the sample image and the sub-target features of the target identifier image under the same feature dimension, the feature weight of each dimension corresponding to the sample image is determined; based on the image features of the sample image and the feature weight of each dimension corresponding to the sample image, the image features of the sample image are weighted to obtain the positive sample features of the positive sample image and the negative sample features of the negative sample image.
[0147] In one embodiment, the image matching module 830 is used to calculate the matching degree between the fused features and the image features of the candidate identifier image; determine the maximum matching degree among the matching degrees, and use the candidate identifier image corresponding to the maximum matching degree as the matching image.
[0148] In one embodiment, the image matching module 830 is configured to acquire similarity results between the target features of the target identifier image and the image features of the candidate identifier image; determine multiple similarity results that meet the screening conditions from the multiple similarity results, and determine the candidate identifier images corresponding to the multiple similarity results as candidate positive sample images; calculate the matching degree between the fused feature and the image features of each candidate positive sample image; determine the maximum matching degree among the matching degrees, and use the candidate positive sample image corresponding to the maximum matching degree as the matching image.
[0149] In one embodiment, the candidate sign image includes candidate traffic element images, the target sign image includes a target traffic element image, and the matching image includes a matching traffic element image; the image selection module 810 is used to select positive sample images from the candidate traffic element images that meet the similarity condition with the target traffic element image; the result determination module 840 is used to determine the traffic element content contained in the matching traffic element image as the traffic element content of the target traffic element image.
[0150] Each module in the aforementioned image identification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0151] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores identification data for labeled images. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for identifying the identification content of labeled images.
[0152] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a method for recognizing the content of an identified image. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0153] Those skilled in the art will understand that Figure 9 and Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0154] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0155] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0156] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described above.
[0157] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0158] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0159] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for recognizing the content of an image, characterized in that, The method includes: From the candidate icon images, select positive sample images that meet the similarity condition with the target icon image, and randomly select negative sample images; Obtain the positive sample features of the positive sample image and the negative sample features of the negative sample image; Based on the positive fusion coefficient matched by the positive sample image and the positive sample features, the positive relationship features are determined; Based on the reverse fusion coefficient matched by the negative sample image and the negative sample features, the reverse relationship features are determined; The target features, the positive relationship features, and the negative relationship features of the target identification image are summed to obtain a fusion feature that matches the target identification image. From the candidate identifier images, select matching images whose image features and fused features meet the matching conditions; The identifier content contained in the matching image is determined as the identifier content of the target identifier image.
2. The method according to claim 1, characterized in that, The step of selecting positive sample images from candidate identifier images that meet the similarity condition with the target identifier image includes: Obtain the target features of the target identifier image and the image features of the candidate identifier images; Calculate the similarity between the target feature and each of the image features to obtain the similarity results; The candidate identifier image corresponding to the highest similarity result among the various similarity results is determined as the positive sample image.
3. The method according to claim 2, characterized in that, The method further includes: Feature point recognition is performed on the target image to determine each feature point of the target image; According to the region division parameters that match the feature points, the image of multiple candidate identifier regions corresponding to each feature point is divided with each feature point as the center. Based on the scores corresponding to each candidate identifier region image, a target identifier region image is selected from each candidate identifier region image, and the target identifier region image is determined as the target identifier image.
4. The method according to claim 1, characterized in that, For each of the positive and negative sample images, the method further includes: Determine the sub-image features under each feature dimension of the sample image, and the sub-target features under each feature dimension of the target identifier image; the image feature dimensions of the sample image and the image feature dimensions of the target identifier image are the same; Based on the sub-image features of the sample image and the sub-target features of the target identification image under the same feature dimension, determine the feature weight of each dimension corresponding to the sample image; Based on the image features of the sample image and the feature weight of each dimension corresponding to the sample image, the image features of the sample image are weighted to obtain the positive sample features of the positive sample image and the negative sample features of the negative sample image.
5. The method according to claim 1, characterized in that, The step of selecting matching images from the candidate identifier images whose image features and fused features meet the matching conditions includes: Calculate the matching degree between the fused features and the image features of the candidate identifier image; Determine the maximum matching degree among all the matching degrees, and use the candidate identifier image corresponding to the maximum matching degree as the matching image.
6. The method according to claim 2, characterized in that, The step of selecting matching images from the candidate identifier images whose image features and fused features meet the matching conditions includes: Obtain the similarity results of the target features of the target identifier image and the image features of the candidate identifier image; From the various similarity results, multiple similarity results that meet the screening conditions are identified, and the candidate identifier images corresponding to the multiple similarity results are identified as candidate positive sample images; Calculate the matching degree between the fused features and the image features of each candidate positive sample image; The maximum matching degree among the various matching degrees is determined, and the candidate positive sample image corresponding to the maximum matching degree is taken as the matching image.
7. The method according to claim 1, characterized in that, The candidate sign image includes candidate traffic element images, the target sign image includes target traffic element images, and the matching image includes matching traffic element images; The step of selecting positive sample images from candidate signage images that meet the similarity condition with the target signage image includes: selecting positive sample images from candidate traffic element images that meet the similarity condition with the target traffic element image. The step of determining the identifier content contained in the matched image as the identifier content of the target identifier image includes: determining the traffic element content contained in the matched traffic element image as the traffic element content of the target traffic element image.
8. A device for recognizing the content of an image, characterized in that, The device includes: The image selection module is used to select positive sample images from the candidate label images that meet the similarity condition with the target label image, and to randomly select negative sample images. The feature fusion module is used to acquire the positive sample features of the positive sample image and the negative sample features of the negative sample image; determine the positive relationship features based on the positive fusion coefficient matched by the positive sample image and the positive sample features; determine the reverse relationship features based on the reverse fusion coefficient matched by the negative sample image and the negative sample features; and sum the target features, the positive relationship features, and the reverse relationship features of the target identification image to obtain the fusion features that match the target identification image. The image matching module is used to filter out matching images from the candidate identifier images whose image features and fused features meet the matching conditions. The result determination module is used to determine the identifier content contained in the matching image as the identifier content of the target identifier image.
9. The identification content recognition device for identification images according to claim 8, characterized in that, The image selection module is further configured to acquire target features of the target identifier image and image features of the candidate identifier image; calculate the similarity between the target features and each of the image features to obtain each similarity result; and determine the candidate identifier image corresponding to the maximum similarity result among the similarity results as the positive sample image.
10. The identification content recognition device for identification images according to claim 9, characterized in that, The image selection module is also used to identify feature points in the target image, determine each feature point of the target image, and divide the target image into multiple candidate identifier regions corresponding to each feature point according to the region division parameters that match the feature points. Based on the scores corresponding to each candidate identifier region image, a target identifier region image is selected from each candidate identifier region image, and the target identifier region image is determined as the target identifier image.
11. The identification content recognition device for identification images according to claim 8, characterized in that, The feature fusion module is further configured to determine sub-image features under each feature dimension of the sample image and sub-target features under each feature dimension of the target identifier image; the image feature dimensions of the sample image are the same as those of the target identifier image; based on the sub-image features of the sample image and the sub-target features of the target identifier image under the same feature dimension, the feature weight of each dimension corresponding to the sample image is determined; based on the image features of the sample image and the feature weight of each dimension corresponding to the sample image, the image features of the sample image are weighted to obtain the positive sample features of the positive sample image and the negative sample features of the negative sample image.
12. The identification content recognition device for identification images according to claim 8, characterized in that, The image matching module is further configured to calculate the matching degree between the fused features and the image features of the candidate identifier image; determine the maximum matching degree among the matching degrees, and use the candidate identifier image corresponding to the maximum matching degree as the matching image.
13. The identification content recognition device for identification images according to claim 9, characterized in that, The image matching module is further configured to obtain the similarity results between the target features of the target identifier image and the image features of the candidate identifier image; determine multiple similarity results that meet the screening conditions from the multiple similarity results, and determine the candidate identifier images corresponding to the multiple similarity results as candidate positive sample images; Calculate the matching degree between the fused features and the image features of each candidate positive sample image; The maximum matching degree among the various matching degrees is determined, and the candidate positive sample image corresponding to the maximum matching degree is taken as the matching image.
14. The identification content recognition device for identification images according to claim 8, characterized in that, The candidate sign image includes candidate traffic element images, the target sign image includes a target traffic element image, and the matching image includes a matching traffic element image; the image selection module is used to select positive sample images from the candidate traffic element images that meet the similarity condition with the target traffic element image; the result determination module is used to determine the traffic element content contained in the matching traffic element image as the traffic element content of the target traffic element image.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.