Tire attribute information identification method, device, system and equipment and storage medium
By combining tire region detection and slicing with a text segmentation model, the problem of excessive memory usage in tire attribute information detection is solved, achieving high-precision adaptive recognition and accurate tire assembly.
Patent Information
- Application Number
- CN202510869311.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-28
AI Technical Summary
In existing technologies, tire attribute information detection suffers from excessive memory consumption, leading to increased hardware costs or decreased detection accuracy. Furthermore, low-resolution images cause blurring of minute semantic markers, affecting detection accuracy.
The tire region is located by a tire region detection model, the tire region image is sliced and processed, the text region is detected by a text segmentation model, and finally the tire attribute information is identified by a tire text recognition model. The phased and progressive recognition method reduces the demand for video memory and preserves image details.
It improves the adaptability and accuracy of tire attribute information detection, reduces hardware consumption, reduces the possibility of false detection and missed detection, and ensures the accuracy of tire assembly.
Smart Images

Figure CN120853174A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for identifying tire attribute information. Background Art
[0002] In the field of industrial automotive tire inspection, accurate identification of semantic information on the tire surface (such as brand, specifications, weight markings, 3C certification, etc.) is a core requirement to ensure assembly accuracy.
[0003] In related technologies, visual inspection is usually used to detect information from tire images. However, this method has certain requirements for video memory resources. For example, when using neural networks to detect information from high-resolution, large-size images, it will lead to excessive video memory usage, a sharp increase in hardware costs, and difficulty in real-time deployment. On the other hand, if low-resolution images are used to reduce costs, the image details are severely lost, and the tiny semantic marks on the tire surface are prone to edge blurring and texture loss, resulting in a decrease in detection accuracy. It can be seen that related technologies have the problem of low applicability for tire attribute information detection. Summary of the Invention
[0004] Therefore, it is necessary to provide a tire attribute information identification method, device, computer equipment, computer-readable storage medium, and computer program product that can improve the applicability of tire attribute information identification in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a method for identifying tire attribute information, including:
[0006] Obtain the raw tire image;
[0007] Based on the trained tire region detection model, tire regions in the original tire image are detected to obtain the location information of the tire regions. The tire region detection model is trained based on historical tire images carrying category labels and location information of the tire regions.
[0008] Based on the location information of the tire region, the tire region image is segmented from the original tire image, and the tire region image is sliced to obtain a slice image set.
[0009] Based on the trained text segmentation model, the text region in each slice image in the slice image set is detected to obtain the text region detection result. The text region detection result includes the detection box of the text region. The text segmentation model is trained based on the category label, location information and segmentation mask of the historical slice images carrying the ground truth box of the text.
[0010] Based on the text region detection results, the detection boxes of the text regions are merged to obtain the target text region detection results of the tire region image;
[0011] Based on the target text region detection results and the trained tire text recognition model, tire attribute information in the tire region image is identified, and the tire attribute information recognition result is obtained. The tire text recognition model is trained based on historical tire images carrying text labels.
[0012] Secondly, this application also provides a tire attribute information identification device, comprising:
[0013] Tire image acquisition module, used to acquire raw tire images;
[0014] The tire region detection module is used to detect tire regions in the original tire image based on the trained tire region detection model, and obtain the location information of the tire region. The tire region detection model is trained based on historical tire images carrying category labels and location information of the tire regions.
[0015] The tire image slicing module is used to segment the tire region image from the original tire image based on the location information of the tire region, and to slice the tire region image to obtain a set of sliced images.
[0016] The text region detection module is used to detect text regions in each slice image in the slice image set based on the trained text segmentation model, and obtain the text region detection result. The text region detection result includes the detection box of the text region. The text segmentation model is trained based on the category label, location information and segmentation mask of the historical slice images carrying the ground truth box of the text.
[0017] The label fusion module is used to merge the detection boxes of text regions based on the text region detection results to obtain the target text region detection results of the tire region image.
[0018] The text recognition module is used to identify tire attribute information in tire region images based on the target text region detection results and the trained tire text recognition model, and obtain the tire attribute information recognition results. The tire text recognition model is trained based on historical tire images carrying text labels.
[0019] Thirdly, this application provides a tire attribute information recognition system, including a processor and an image acquisition device:
[0020] The image acquisition device is used to acquire images of the tire to be identified, obtain the original tire image, and send the original tire image to the processor;
[0021] The processor is used to execute the steps in any of the above embodiments of the tire attribute information recognition method to obtain the tire attribute information recognition result of the tire to be identified.
[0022] Fourthly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in any of the above embodiments of the tire attribute information identification method.
[0023] Fifthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps in any of the above embodiments of the tire attribute information identification method.
[0024] Sixthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in any of the above-described embodiments of the tire attribute information identification method.
[0025] The aforementioned tire attribute information recognition methods, devices, computer equipment, computer-readable storage media, and computer program products address the issue of excessive video memory consumption in high-resolution tire image detection. First, a tire region detection model rapidly locates the tire region in the original tire image, obtaining its positional information. Based on this location information, the tire region image is segmented from the original tire image, reducing the data volume of subsequent models and thus lowering the video memory requirements during model detection. Second, the tire region image is sliced to obtain a sliced image set, further significantly reducing the image size. Simultaneously, the sliced image set retains the details of the original tire image, allowing text regions in the sliced image set to be detected based on a text segmentation model. This reduces the hardware consumption of model detection and eliminates the need to reduce tire image resolution to lower hardware consumption, reducing the possibility of false positives or false negatives caused by text detection on low-resolution tire images and improving adaptability to different tire attribute information detection scenarios. Based on the tire text detection results, the detection boxes of the text regions are merged to trace back the text detection results of the sliced image set to the tire region image before the sliced processing, so as to obtain the target text region detection results containing complete tire attribute information. Based on the target text region detection results and the trained tire text recognition model, the tire attribute information is identified. In this way, based on this phased and progressive recognition method, the adaptability to high-precision tire attribute information recognition in different detection fields is improved. Furthermore, it is beneficial to improve the accuracy of tire assembly based on the accurate tire attribute information identified. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a diagram illustrating the application environment of a tire attribute information recognition method in one embodiment.
[0028] Figure 2 This is a flowchart illustrating a tire attribute information identification method in one embodiment;
[0029] Figure 3 This is a schematic diagram of a tire region image in one embodiment;
[0030] Figure 4 This is a schematic diagram of the rotation slicing process in one embodiment;
[0031] Figure 5 This is a schematic diagram of a sliced image set in one embodiment;
[0032] Figure 6 This is a flowchart illustrating the tire attribute information identification method in another embodiment;
[0033] Figure 7 This is a schematic diagram of the text region detection results of a sliced image set in one embodiment;
[0034] Figure 8 This is a schematic diagram of the target text region detection results in one embodiment;
[0035] Figure 9 This is a schematic diagram of a tire region image after polar coordinate transformation in one embodiment;
[0036] Figure 10 This is a schematic diagram of the target image in one embodiment;
[0037] Figure 11 This is a structural block diagram of a tire attribute information recognition system in one embodiment;
[0038] Figure 12 This is a flowchart illustrating the tire assembly inspection process in one embodiment;
[0039] Figure 13 This is a structural block diagram of a tire attribute information recognition device in one embodiment;
[0040] Figure 14 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0042] Tire attribute information refers to the parameters of a tire that are identifiable on its surface. Tire attribute information includes, but is not limited to, tire section width, aspect ratio, tire inner diameter, speed rating, load index, and tire type designation. In industrial tire semantic information detection, the relevant technologies typically employ automated detection methods to identify tire surface attribute information. To preserve tire surface details, tire images captured in industrial settings are often high-resolution, such as 4K resolution. Traditional manual feature detection methods are not robust enough to noise and occlusion, frequently resulting in false positives, and the feature extraction process is time-consuming, failing to meet the cycle time requirements of actual production. Therefore, in actual production, more efficient and accurate deep learning methods are often employed.
[0043] However, deep learning-based detection methods place certain demands on video memory. If a deep learning model is directly used to infer information from the entire tire image, video memory overflow is likely to occur in actual production, preventing on-site inspection from proceeding normally. This forces companies to upgrade hardware or adopt distributed computing frameworks, significantly increasing costs. Furthermore, compressing image quality to reduce hardware costs results in the loss of pixel information for tiny markings on the tire surface, leading to blurred semantic edges and texture breaks. Consequently, the deep learning model cannot effectively extract features, increasing both false positive and false negative rates. Therefore, to improve the adaptability of tire attribute information recognition to different application scenarios, this application proposes a tire attribute information recognition method that enhances the adaptability of high-precision recognition in various detection scenarios.
[0044] The tire attribute information identification method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server.
[0045] Specifically, the operator can upload the acquired raw tire image to the server 104 via terminal 102, and then send a tire attribute information recognition message to the server 104 via terminal 102. The server 104 obtains the raw tire image data, and then, based on a trained tire region detection model, detects the tire regions in the raw tire image to obtain the location information of the tire regions. The tire region detection model is trained based on historical tire images carrying category labels and location information of the tire regions. Then, based on the location information of the tire regions, the tire region images are segmented from the raw tire image, and the tire region images are sliced to obtain a set of sliced images. Finally, based on the trained text... This segmentation model detects text regions in each slice of image within a slice image set, obtaining text region detection results. These results include bounding boxes for the text regions. The text segmentation model is trained on historical slice images carrying text bounding boxes, along with category labels, location information, and segmentation masks. Then, based on the text region detection results, the bounding boxes are merged to obtain the target text region detection results for the tire region image. Finally, based on the target text region detection results and the trained tire text recognition model, tire attribute information in the tire region image is identified, yielding tire attribute information recognition results. The tire text recognition model is trained on historical tire images carrying text labels.
[0046] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0047] In one exemplary embodiment, such as Figure 2 As shown, a method for identifying tire attribute information is provided, which can be applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps (hereinafter referred to as S): S100 to S600. Wherein:
[0048] S100, acquire the original tire image.
[0049] The original tire image can be an uncompressed image containing the vehicle's tires. The original tire image can be obtained by an imaging device acquiring an image of the tire to be identified.
[0050] In practical applications, operators can capture images of tires in the actual application scenario to obtain raw tire images, and then upload the captured raw tire images to a server. Alternatively, a monitoring device can capture raw tire images containing the tires in the actual application scenario, store them in a local database, and then the operator uploads the stored raw tire images to the server. This application does not limit the method of obtaining raw tire images. Practical application scenarios can include vehicle tire assembly, tire maintenance and inspection, tire inventory management, and tire production quality control, etc.
[0051] Taking the application scenario of detecting whether a vehicle's tire assembly meets pre-installation requirements as an example, a specific space can be pre-defined as a detection area within the factory. The vehicle drives into this area, and an image acquisition device deployed there captures images of the vehicle's tires, obtaining raw tire images. These raw tire images are then transmitted to a server. The server obtains the raw tire images. In other embodiments, after obtaining the raw tire images, the server performs edge cropping to ensure uniform image size and adapt to model input requirements.
[0052] S200, based on a trained tire region detection model, detects tire regions in the original tire image and obtains the location information of the tire regions. The tire region detection model is trained based on historical tire images carrying category labels and location information of the tire regions.
[0053] Among them, the tire region detection model is used to locate the tire region in the image.
[0054] In practical applications, during the model training phase, historical tire images collected over a historical time period are pre-acquired for model training. The method for acquiring these historical tire images can refer to the steps described above for acquiring original tire images, and will not be repeated here. Data annotation is performed on the historical tire images, labeling the category labels and location information of the ground truth bounding boxes of the tire regions. Annotation methods can include tool-based annotation or manual annotation. The category label can include whether it is a tire region. In other embodiments, after acquiring the historical tire images, data augmentation processing is performed on them to enhance model performance. Data augmentation processing includes, but is not limited to, rotation, flipping, scaling and cropping, brightness adjustment, and noise addition.
[0055] Then, an initial tire region detection model is constructed based on a pre-trained model. Pre-trained models include, but are not limited to, YOLO models, R-CNN (Region-based Convolutional Neural Networks) models, and SSD (Single Shot Multibox Detector). For example, in this embodiment, for tire foreground localization, an initial tire region detection model is constructed based on a YOLOV8 model to locate the tire foreground region in the image. The YOLOV8 model is an efficient single-stage object detection model that supports flexible size scaling and mixed-precision training, suitable for scenarios with high real-time and robustness requirements, such as industrial inspection and autonomous driving. The server inputs historical tire images into the constructed initial tire region detection model. Based on the tire region location information predicted by the initial tire region detection model and the location information of the labeled ground truth bounding boxes of the tire region, the model prediction error is determined through a preset loss function. The parameters of the initial tire region detection model are iteratively updated based on the error until a preset training termination condition is reached, at which point the training of the initial tire region detection model is stopped, resulting in a trained tire region detection model. The preset loss function can be CIoU (Complete Intersection over Union) or DFL (Distribution Focal Loss), etc.
[0056] In practice, the server can take the original tire image as input, call a trained tire region detection model, and use the tire region detection model to detect the tire region in the original tire image to obtain the location information of the tire region. The location information of the tire region can be the location information of the detection box, including the coordinates of the center point of the detection box and the length and width of the detection box.
[0057] S300, based on the location information of the tire region, segments the tire region image from the original tire image, performs slicing processing on the tire region image, and obtains a slice image set.
[0058] In practical applications, to reduce the amount of data and computational burden of a single detection by the model, a pre-set image slicing strategy can be used to slice the tire region image, obtaining a set of sliced images that retain the detailed features of the original tire image. This pre-set slicing strategy can include a slicing strategy based on a pre-set slice size. Specifically, a slicing strategy based on a pre-set slice size can involve pre-setting slicing parameters, such as the width and height of the slices, according to the dimensions of the image to be sliced, and then segmenting the image according to these pre-set parameters to obtain a series of non-overlapping small images.
[0059] The preset slicing strategy can also include a slicing strategy based on dynamically adjusting the slice size. Specifically, a slicing strategy based on dynamically adjusting the slice size can involve using smaller slice sizes for information-rich regions of the image and larger slice sizes for information-poor regions. For example, edge detection can be performed on the tire region image first to evaluate the information content of different regions of the image and obtain information content analysis results. Subsequently, based on the information content analysis results and a preset information content threshold, regions with less information and regions with more information are classified, and slicing is performed on the regions with less information and regions with more information respectively to obtain a sliced image set.
[0060] In practice, the server segments the tire region image from the original tire image based on the location information of the detection bounding box, such as... Figure 3 As shown, based on a preset slicing strategy, the tire region image is segmented into multiple slice images. It is understandable that, depending on the slicing strategy, the slice image set may include multiple non-overlapping slice images, or it may include multiple slice images with overlapping regions.
[0061] S400, based on a trained text segmentation model, detects text regions in each slice image in a slice image set and obtains text region detection results. The text region detection results include the detection boxes of the text regions. The text segmentation model is trained based on historical slice images carrying the category labels, location information, and segmentation masks of the ground truth boxes of the text.
[0062] The category label for the ground truth bounding box can include whether it is a text region. The segmentation mask is used to identify text regions in the image.
[0063] In practical applications, during the model training phase, historical tire images can be pre-sliced into segments. The slicing method can refer to the steps described above, which outline the implementation steps for slicing tire region images based on a pre-defined slicing strategy; these will not be repeated here. Data annotation is then performed on the historical sliced images, including the category labels, location information, and segmentation masks of the ground truth bounding boxes in the text regions. Annotation methods can include tool-based annotation or manual annotation.
[0064] Then, an initial text segmentation model is constructed based on the pre-trained model. Pre-trained models include, but are not limited to, PSENet (Progressive Scale Expansion Network), PAN (Path Aggregation Network), and MSR (Multi-Scale Shape Regression for Scene Text Detection). For example, in this embodiment, for pixel-level text region detection in segmented images, an initial text segmentation model is constructed based on a model such as YOLOv8-Seg to locate text regions in the image. The YOLOv8-Seg model is an extension of the YOLO series for instance segmentation tasks; it integrates pixel-level mask prediction capabilities on top of the detection model, achieving a unification of object detection and fine contour segmentation. The server inputs historical segmented images into the pre-built initial text segmentation model, iteratively trains the initial text segmentation model, and determines the prediction error of the model based on the location information of the text regions predicted by the initial text segmentation model, the category labels and segmentation masks, and the location information of the labeled ground truth bounding boxes, using a preset loss function. The parameters of the initial text segmentation model are iteratively updated based on the error until a preset training termination condition is met, at which point training of the initial text segmentation model stops, resulting in a trained tire region detection model. The preset loss function can be the CIoU loss function, the DFL loss function, or the Dice coefficient loss function.
[0065] In practice, the server can take a set of image segments as input, call a trained text segmentation model, and use the text segmentation model to detect text regions in the image segments, obtaining the category label, location information, and segmentation mask of the detection boxes for the text regions. As an example, the location information of the detection boxes can include the coordinates of the center point of the detection box and the length and width of the detection box.
[0066] S500, based on the text region detection results, merges the detection boxes of the text regions to obtain the target text region detection results of the tire region image.
[0067] The tire attribute information includes, but is not limited to, tire section width, tire aspect ratio, tire inner diameter, speed rating, load index, rim diameter, tire type code, tire brand, date, and weight markings.
[0068] The target text region detection result is the text region detection result corresponding to the tire region image. The target text region detection result includes the segmentation mask of the text line of complete tire attribute information in the tire region image, the category label of the detection box, and the location information.
[0069] In practical applications, because the tire region image is segmented, the segmented image set may contain multiple segmented images with overlapping areas. This can lead to overlapping detected text regions. Furthermore, tire attribute information in the image may be truncated due to segmentation. Therefore, to obtain accurate text detection results on the tire, it is necessary to fuse the detected labels belonging to the same text region. Specifically, based on the pixel mapping relationship between the segmented images and the tire region image, the text region detection results of the segmented images can be back-mapped to the tire region image. For text detection results with overlapping areas, the detection boxes and segmentation masks of the text regions are merged to obtain the target text region detection result containing complete tire attribute information.
[0070] S600, based on the target text region detection results and the trained tire text recognition model, identifies tire attribute information in the tire region image and obtains the tire attribute information recognition result. The tire text recognition model is trained based on historical tire images carrying text labels.
[0071] The tire attribute information recognition result may include the identified tire attribute information.
[0072] In practical applications, during the model training phase, historical tire images are pre-acquired. The acquisition method can refer to the steps described above for obtaining the original tire images, and will not be repeated here. Data annotation is then performed on the images in the historical tire image set, including text annotations. Annotation methods can include tool-based annotation or manual annotation.
[0073] Then, an initial tire text recognition model is constructed based on a deep neural network model. This deep neural network model includes, but is not limited to, CRNN (Convolutional Recurrent Neural Network) models, TextCNN (Text Convolutional Neural Network) models, and recurrent neural network models. For example, in this embodiment, the initial tire text recognition model is constructed based on a CRNN model. The server takes a set of historical tire images with text labels as input, calls the constructed initial tire text recognition model, and determines the prediction error of the model through a preset loss function based on the predicted tire attribute information and the labeled tire attribute information. The parameters of the initial tire text recognition model are iteratively updated according to the error until a preset training termination condition is reached, resulting in a trained tire text recognition model. The preset training termination condition can be that the loss function value is less than a preset loss threshold for a preset number of consecutive times. The preset loss function includes, but is not limited to, the mean squared error loss function, the cross-entropy loss function, and the smoothing L1 loss function.
[0074] In practice, the server marks the regions containing tire attribute information in the tire region image based on the target text region detection results. Using the marked tire region image as input, the trained tire text recognition model is called to recognize the text content within the marked regions and obtain the tire attribute information recognition results.
[0075] In the aforementioned tire attribute information recognition method, to address the issue of excessive GPU memory consumption in high-resolution tire image detection, firstly, a tire region detection model is used to quickly locate the tire region in the original tire image, obtaining the positional information of the tire region in the image. Based on the positional information of the tire region, the tire region image is segmented from the original tire image, thereby reducing the data volume of subsequent models and helping to reduce the GPU memory requirements during model detection. Secondly, the tire region image is sliced to obtain a sliced image set, which further significantly reduces the image size. At the same time, the sliced image set retains the details of the original tire image, thereby detecting the text region in the sliced image set based on a text segmentation model, obtaining the text region detection result, reducing the hardware consumption of model detection. Furthermore, it is not necessary to reduce the resolution of the tire image to reduce hardware consumption, reducing the possibility of false detections or false negatives caused by text detection on low-resolution tire images, and improving the adaptability to different tire attribute information detection scenarios. Based on the tire text detection results, the detection boxes of the text regions are merged to trace back the text detection results of the sliced image set to the tire region image before the sliced processing, so as to obtain the target text region detection results containing complete tire attribute information. Based on the target text region detection results and the trained tire text recognition model, the tire attribute information is identified. In this way, based on this phased and progressive recognition method, the adaptability to high-precision tire attribute information recognition in different detection fields is improved. Furthermore, it is beneficial to improve the accuracy of tire assembly based on the accurate tire attribute information identified.
[0076] To reduce the data pressure on model inference, in an exemplary embodiment, before detecting tire regions in the original tire image based on a trained tire region detection model and obtaining the location information of the tire regions, the method further includes:
[0077] If the resolution of the original tire image is greater than a preset resolution threshold, the original tire image is compressed.
[0078] Compression processing may include, but is not limited to, reducing image resolution using compression algorithms such as JPEG (Joint Photographic Experts Group) and WebP (Web Picture Format).
[0079] In practical applications, particularly in the visual inspection of tire information, the small size of tire attribute information necessitates the use of high-resolution cameras (e.g., 5 megapixels, 10 megapixels, or higher) to capture tire images and maintain detection accuracy. However, higher resolution translates to larger data volumes, increasing image processing time and computational resource requirements. To reduce the data processing burden on the model and improve the applicability of tire attribute information detection, a resolution threshold can be pre-set based on actual needs and the performance of the tire detection model. The system then determines whether to compress the original tire image based on the set resolution threshold and the original tire image's resolution. Specifically, the server compares the original tire image's resolution with the preset resolution threshold. If the original tire image's resolution exceeds the preset threshold, it compresses the original tire image according to a preset compression method. For example, the preset resolution threshold is 1920x1080 pixels. It is understood that in other embodiments, the preset resolution threshold can be any value other than 1920x1080 pixels, and this application does not limit this.
[0080] Based on a trained tire region detection model, tire regions are detected in the original tire image to obtain the location information of the tire regions, including:
[0081] Based on the trained tire detection model, the tire region in the compressed tire image is detected to obtain the location information of the tire region.
[0082] In practical applications, after image compression, the server uses the compressed tire image as input and calls a trained tire detection model. The model detects the region where the tire is located in the image and outputs the location information of the detection box for the tire region. The location information of the detection box can include the coordinates of the four vertices of the detection box.
[0083] Based on the compression relationship between the original tire image and the compressed tire image, the location information of the tire region in the original tire image is determined.
[0084] The compression relationship can be the size ratio between the compressed tire image and the original tire image.
[0085] In practical applications, the server pre-determines the size ratio between the compressed tire image and the original tire image. Specifically, the size ratio includes the length ratio and the width ratio. This ratio can be determined based on the length and width of the compressed tire image and the original tire image, respectively. The coordinate information of the tire region in the compressed tire image is detected by the model and mapped to the original tire image according to the length and width ratios, thus obtaining the position information of the tire region in the original tire image.
[0086] In this embodiment, by compressing the high-resolution original tire image, the number of images is reduced, thereby locating the tire region in the compressed tire image based on the tire detection model, which improves the model detection efficiency. The location information of the tire region is mapped to the original tire image based on the compression relationship. This indirect method of locating the tire region in the original tire image reduces the data processing pressure of the model.
[0087] To improve the applicability of tire attribute information recognition, in an exemplary embodiment, the tire region image is sliced to obtain a sliced image set, including:
[0088] The tire region image is sliced, and the target slice image is extracted from the multiple slice images obtained.
[0089] A reference slice image is selected from multiple slice images. The tire region image is rotated around the reference slice image at a preset rotation angle. The reference slice image is a slice image that includes the center point of the tire region image.
[0090] The rotated tire region image is sliced, and the target slice image is extracted from the multiple slice images obtained.
[0091] Return to the steps of selecting a reference slice image from multiple slice images, and rotating the tire region image with the reference slice image as the center according to a preset rotation angle until the sum of the rotation angles of the tire region image is equal to or greater than 360 degrees, to obtain a slice image set containing multiple target slice images.
[0092] In this embodiment, to preserve the original image's detailed features, a spatial segmentation and angle transformation strategy is used to convert the ultra-large tire image into a high-quality slice dataset suitable for deep learning models. Considering the tire's circular geometry, rotating the slices becomes possible. Therefore, this application designs a dynamic rotational slicing mechanism, which iteratively performs: slicing the tire region image, extracting slices, and rotating the tire region image to obtain a slice image set.
[0093] In practical applications, firstly, the number of iterations for the dynamic rotation slicing mechanism is determined based on a preset rotation angle. Specifically, the rotation angle is preset according to the resolution of the tire region image and the slicing strategy, used for rotating and slicing the image. The number of iterations is the number of rotations required for the image to rotate one full circle (360 degrees) according to the preset rotation angle. For example, if the preset rotation angle is 22.5 degrees, the number of iterations is 360 / 22.5 = 16.
[0094] The reference slice image can be a slice image containing the center point of the tire region image. The target slice images can be one or more. The slice image set includes multiple extracted target slice images.
[0095] Secondly, the server can slice the tire region image according to a preset slicing strategy, resulting in multiple slice images. The steps for slicing the tire region image according to the slicing strategy described above will not be repeated here. Then, the server can extract the target image from the multiple slice images according to a preset slice extraction strategy.
[0096] In one example, the preset slicing strategy may also include a horizontal slicing method based on a fixed width and a vertical slicing method based on a fixed length. Specifically, the horizontal slicing method based on a preset width can divide the image into several fixed-width slices along a horizontal direction according to a preset fixed width; the horizontal slicing method based on a preset length can divide the image into several fixed-length slices along a vertical direction according to a preset fixed length.
[0097] In one example, the preset slice extraction strategy must ensure that the extracted target slice image covers the identified areas of tire attribute information, such as the tire sidewall area. Specifically, the preset slice extraction strategy may be to extract a slice image from a fixed area of slice images from slice images in different divided areas, and then determine the extracted slice image as the target slice image. For example, a fixed area may be pre-set for different slice processing strategies to extract the target slice image. The fixed area can be randomly determined. For instance, in a horizontal slicing method based on a fixed width, the fixed area can be set to the area containing the first row of the horizontal division.
[0098] Then, the server determines the slice image containing the center point of the tire region image from the multiple slice images after division as the reference slice image. With the reference slice image as the center, the tire region image is rotated according to a preset rotation angle and a preset rotation direction (such as clockwise rotation).
[0099] Next, the server slices the rotated tire region image again using the same slicing strategy, extracting a fixed region from the resulting slices as the target slice image. The slicing and extraction steps described above are repeated here.
[0100] Then, return to the step of selecting a reference slice image from multiple slice images, and rotating the tire region image with the reference slice image as the center according to a preset rotation angle until the sum of the rotation angles of the tire region image is equal to or greater than 360 degrees, to obtain a slice image set containing multiple target slice images.
[0101] In other embodiments, to ensure that the sliced image set covers the area containing tire attribute information in the tire region image, the identification area of tire attribute information, such as the tire sidewall area, can be pre-determined based on the manufacturing standards of the tire to be identified. Subsequently, a mask is created for the identification area of tire attribute information in the tire region image, with initial values all set to 0. After extracting the target sliced image during loop execution, the pixel position corresponding to the target sliced image in the mask is marked as the target value (target value is 1). After obtaining the sliced image set, it is verified whether the mask is entirely set to 1. If the mask is not entirely set to 1, the slice extraction strategy or a preset rotation angle (such as reducing the rotation angle) needs to be adjusted to re-extract the sliced image set until it is determined that the mask is entirely set to 1, ensuring that the rotated and extracted sliced image set covers the area containing tire attribute information in the tire region image without missing detailed features.
[0102] In this embodiment, a rotation slicing mechanism for the image is designed by combining spatial segmentation and angle transformation strategies with the characteristics of the tire shape. This balances the preservation of detailed features of the original image with the amount of input data for a single model detection. This improves the applicability of tire attribute information recognition.
[0103] To reduce slicing, in one exemplary embodiment, the tire region image is sliced, and the target slice image is extracted from the multiple sliced images obtained, further comprising:
[0104] Based on a preset nine-grid slicing strategy, the tire region image is sliced, and the target slice image within a fixed region is extracted from the slice images within the divided nine-grid region.
[0105] The preset nine-grid slicing strategy involves dividing the tire area image into a 3×3 grid. For example... Figure 4 As shown, the black grid lines represent the dividing grid (similar to a tic-tac-toe grid). The fixed area can be the grid area in the first row horizontally and the second column vertically of the 3x3 grid. It is understood that in other embodiments, the fixed area can be any area other than the central area (i.e., the grid area in the second row horizontally and the second column vertically of the 3x3 grid).
[0106] To reduce the amount of input data required for a single detection by the model, this embodiment designs a composite slicing mechanism based on a nine-grid slicing strategy and dynamic rotation, taking into account the distribution characteristics of tire attribute information (such as tire attribute information typically being marked on the tire sidewall and wheel rim). The steps in this embodiment are explained below through an example:
[0107] First, the rotation angle is set in advance based on experience or testing requirements. In this embodiment, the preset rotation angle is 22.5 degrees. Figure 4As shown, the following steps are executed repeatedly until the sum of the rotation angles of the tire region image is equal to or greater than 360 degrees, at which point the following loop body is stopped, resulting in a set of sliced images:
[0108] Step 1: Based on the preset nine-grid slicing strategy, slice the tire region image. This can be done by dividing the tire region image into a 3×3 grid. At this point, the resolution of a single slice image is only 1 / 9 of the original image, but the microscopic clarity of the tire rim pattern is fully preserved. The preset extraction strategy can be to extract the slice image corresponding to the first horizontal row and the second vertical column of the nine-grid, and determine this slice image as the target slice image.
[0109] Step 2: The reference slice image can be determined by selecting the slice image corresponding to the second row horizontally and the second column vertically in a 3x3 grid (i.e., the fifth grid). Using the reference slice image as the center, rotate the tire region image 22.5 degrees clockwise. Specifically, construct an affine transformation matrix with the center point of the reference slice image as the origin, and rotate the image using this matrix.
[0110] Step 3: Perform slicing on the rotated tire region image using a nine-grid slicing strategy, extract the slice image corresponding to the first horizontal row and the second vertical column of the nine-grid, and determine the slice image as the target slice image.
[0111] Returning to step 2, the slice image corresponding to the second row horizontally and second column vertically in the 3x3 grid (i.e., the fifth grid cell) is determined as the reference slice image. Centered on the reference slice image, the tire region image is rotated 22.5 degrees clockwise. This process continues until the total rotation angle of the tire region image equals or exceeds 360 degrees, resulting in a slice image set containing multiple target slice images. After the loop completes, a slice image set containing 16 target slice images is obtained, as shown below. Figure 5 As shown.
[0112] In this embodiment, full coverage of the tire's circumference can be achieved through multiple iterations. This spiral progressive slicing strategy significantly reduces the amount of input data for a single model inference (compressing it from tens of thousands of pixels to thousands of pixels). Furthermore, it eliminates the need to compress high-resolution tire images before detecting text regions; the sliced image set can cover the complete detection information of the tire, thus achieving information segmentation of the entire tire image.
[0113] In one exemplary embodiment, the text region detection result includes a segmentation mask of the tire text, category labels of the detection boxes, and location information, such as... Figure 6 As shown, S500 includes S520 to S560. Wherein:
[0114] S520 maps the text region detection results of the slice image set to the coordinate space of the tire region image through inverse affine transformation.
[0115] In practical applications, following the example above, after slicing the tire region image into a set of 16 slice images, the text region detection result for each slice image is obtained through a text segmentation model, such as... Figure 7 As shown, to merge the detection boxes for the text regions, in this embodiment, the text region detection results of the sliced image set are first mapped to the coordinate space of the tire region image through an inverse affine transformation. Specifically, during each iteration of the above-mentioned slicing process for the tire region image, the server records the affine transformation matrix of the image rotation. After completing the segmentation prediction of the sliced image set, for each slice image, the server transforms the local coordinate system of the text region detection results of the sliced image to the global coordinate system of the rotated tire region image based on the rotation angle and the resolution of the tire region image. Based on the rotation angle and affine transformation matrix recorded when the slice image was extracted, the text region detection results of the slice image are mapped back to the global coordinate system of the unrotated tire region image through an inverse affine transformation matrix.
[0116] S540: For any two detection boxes with the same category label, determine the intersection-union ratio between the two detection boxes based on the position information of the detection boxes.
[0117] In practical applications, due to the pre-defined rotation slicing mechanism of the loop, overlapping areas exist between sliced images in the sliced image set. When mapping the text region detection results of the sliced images back to the tire region image, redundant label clusters with multi-angle predictions are formed in the coordinate space of the tire region image. To address this, this application designs an adaptive label aggregation strategy based on the cross-union ratio (CUP). Specifically, for two detection boxes with the same category of labels, the CUP between the two detection boxes is determined based on their position information, and label aggregation is then performed based on the CUP.
[0118] S560, if the cross-union ratio is greater than the preset cross-union ratio threshold, merge the segmentation masks of the tire text in the two detection boxes to obtain the target text region detection result of the tire region image.
[0119] The target text region detection result can be the text region detection result corresponding to the tire region image.
[0120] In practical applications, the operator can pre-set the intersection-union (IU) threshold based on experience. For example, the IU threshold can be set to 0.1. The server compares the detection boxes with the IU ratio and the preset IU threshold. When the IU ratio of two detection boxes with the same category of labels exceeds the threshold, the two detection boxes are merged based on their minimum bounding rectangle, overlapping segmentation masks are merged, and the center of the merged detection box is taken as the new label anchor point. Figure 8 As shown, the text region detection results of the segmented image set are backtracked to the tire region image, thereby obtaining the target text region detection results containing complete tire attribute information.
[0121] In this embodiment, the "divide and conquer-aggregate" solution can effectively solve two major problems in high-resolution ring object processing: on the one hand, the tens of thousands of pixel images are decomposed into GPU memory-friendly subtasks by rotating slices; on the other hand, the robustness of detecting small information is improved by utilizing the complementarity of multi-view prediction.
[0122] In an exemplary embodiment, after obtaining the detection result of the target text region containing tire attribute information, the method further includes:
[0123] The image of the tire region is transformed by polar coordinates to adjust the direction of the tire attribute information to the horizontal direction, and the detection results of the target text region are updated.
[0124] Based on the detection boxes in the updated target text region detection results, the target image is cropped from the transformed tire image. The target image is the region image within the detection boxes.
[0125] In industrial tire inspection scenarios, character recognition is a crucial step for quality traceability and specification verification. However, the distribution of characters on the tire surface is not a linear structure like ordinary text. When the character regions of the tire area image are located using a text segmentation model, the directly cropped information blocks will have an arc-shaped structure, which can easily lead to recognition errors due to geometric distortion during subsequent text recognition. Therefore, considering the circular distribution characteristics of tire attribute information, a polar coordinate transformation is performed on the tire area image before image cropping.
[0126] Polar coordinate transformation is the process of transforming points on a plane. Convert to polar coordinates .in, It is the distance from the point to the origin. It is the angle between the point and the x-axis. The transformation formula can be... ; .
[0127] In practical applications, the server performs polar coordinate transformation on the tire region image according to the polar coordinate transformation formula. Simultaneously, it performs polar coordinate transformation on the mask region and detection box in the target text region detection results. This is to adjust the character orientation in the image to horizontal, such as... Figure 9 As shown.
[0128] After polar coordinate transformation, based on the position information of the detection boxes in the updated target text region detection results, the region image within the detection boxes is cropped from the transformed tire image, which is the target image. Figure 10 As shown, this image displays the region within a cropped detection box.
[0129] Based on the target text region detection results and the trained tire text recognition model, tire attribute information in the target text region detection results is identified to obtain tire attribute information recognition results, including: inputting the target image into the trained tire text recognition model to obtain tire attribute information recognition results.
[0130] In practical applications, after the server crops the target image, it inputs the target image into a trained tire text recognition model. The tire text recognition model then identifies the text content in the target image to obtain tire attribute information. The implementation method for obtaining tire attribute information recognition results based on the trained tire text recognition model is described here, referring to the steps in the above embodiment of identifying text content in the target text region detection results based on the target text region detection results and the trained tire text recognition model to obtain tire attribute information recognition results. These steps will not be repeated here.
[0131] In this embodiment, based on the character distribution characteristics of the annular tire surface, polar coordinate transformation is performed on the merged semantic region to convert the arc-shaped distributed characters into a linear arrangement, which significantly improves the recognition accuracy of the subsequent text recognition model.
[0132] In an exemplary embodiment, before performing polar coordinate transformation on the tire region image, the method further includes:
[0133] Using the center point of the tire region image as the polar coordinate expansion center and the horizontal axis of the image as the polar coordinate horizontal expansion line, rotate the tire region image until the mask region in the target text region detection result does not intersect with the polar coordinate horizontal expansion line, provided that the mask region in the target text region detection result intersects with the polar coordinate horizontal expansion line.
[0134] In practical applications, when performing polar coordinate transformation on a tire region image, the image center point is used as the polar coordinate expansion center, the horizontal axis containing the polar coordinate expansion center is used as the expansion x-axis, and the vertical axis of the image is used as the expansion y-axis. Before performing the polar coordinate transformation, to avoid truncating the complete semantic labels, the tire region image needs to be rotated. Specifically, the server determines the coordinates of each point on the polar coordinate horizontal expansion line based on the y-coordinate of the tire region image center point. It then checks whether the coordinates of the mask region's pixels and the coordinates of each point on the preset horizontal expansion line intersect. If an intersection exists, it is considered that the tire region image has truncated complete semantic labels after the polar coordinate transformation. Therefore, the tire region image needs to be rotated to ensure that the complete semantic labels are not truncated after the polar coordinate transformation. Specifically, a baseline rotation angle (e.g., 22.5 degrees) can be pre-set based on experience. Based on this preset baseline rotation angle, the mask region of the target text region detection result in the rotated tire region image is calculated to determine whether it intersects with the polar coordinate horizontal unfolding line. If they still intersect, the baseline rotation angle is superimposed on the current rotation angle, and the intersecting process continues until the mask region no longer intersects with the polar coordinate horizontal unfolding line, thus obtaining the target rotation angle that moves the polar coordinate horizontal unfolding line away from the complete semantic label. The segmented tire region image is then rotated according to this dynamically calculated target rotation angle.
[0135] In this embodiment, the rotation angle is dynamically calculated using preset horizontal expansion lines and mask areas. The tire region image is then rotated according to the rotation angle to reduce the possibility of polar coordinate transformation truncating the text content in the tire region image, thereby improving the accuracy of text recognition.
[0136] To provide a clearer explanation of the tire attribute information identification method provided in this application, a specific embodiment is described below, which includes the following steps:
[0137] S1, Obtain the original tire image.
[0138] S2, if the resolution of the original tire image is greater than a preset resolution threshold, the original tire image is compressed. Based on the trained tire detection model, the tire region in the compressed tire image is detected to obtain the location information of the tire region. Based on the compression relationship between the original tire image and the compressed tire image, the location information of the tire region in the original tire image is determined. The tire region detection model is trained based on historical tire images carrying category labels and location information of the tire region.
[0139] S3, Based on the location information of the tire region, segment the tire region image from the original tire image.
[0140] S4, based on the preset nine-grid slicing strategy, performs slicing processing on the tire area image, and extracts the target slice image within a fixed area from the slice images within the divided nine-grid area.
[0141] S5. Select a reference slice image from multiple slice images. Rotate the tire region image with the reference slice image as the center according to a preset rotation angle. The reference slice image is a slice image that contains the center point of the tire region image.
[0142] S6, slice the rotated tire region image and extract the target slice image from the multiple slice images obtained.
[0143] Return to S5 until the sum of the rotation angles of the tire region images is equal to or greater than 360 degrees, resulting in a slice image set containing multiple target slice images.
[0144] S7. Based on the trained text segmentation model, detect the text region in each slice image in the slice image set and obtain the text region detection result. The text region detection result includes the detection box of the text region. The text segmentation model is trained based on the category label, location information and segmentation mask of the historical slice images carrying the ground truth box of the text.
[0145] S8. Through inverse affine transformation, the text region detection results of the slice image set are mapped to the coordinate space of the tire region image. For any two detection boxes with the same category label, the cross-union ratio (CUP) between the two detection boxes is determined based on the position information of the detection boxes. If the CUP is greater than the preset CUP threshold, the segmentation masks of the tire text in the two detection boxes are merged to obtain the target text region detection results of the tire region image.
[0146] S9. Using the center point of the tire region image as the polar coordinate expansion center and the horizontal axis of the image as the polar coordinate horizontal expansion line, rotate the tire region image until the mask region in the target text region detection result does not intersect with the polar coordinate horizontal expansion line, provided that the mask region in the target text region detection result intersects with the polar coordinate horizontal expansion line.
[0147] S8. Perform polar coordinate transformation on the tire region image to adjust the direction of the tire attribute information to the horizontal direction and update the target text region detection result; based on the detection box in the updated target text region detection result, crop the target image from the transformed tire image. The target image is the region image within the detection box.
[0148] S9. Input the target image into the trained tire text recognition model to identify the tire attribute information in the target image and obtain the tire attribute information recognition result. The tire text recognition model is trained based on historical tire images with text labels.
[0149] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0150] Based on the same inventive concept, in an exemplary embodiment, such as Figure 11 As shown, a tire attribute information recognition system is provided, including a processor 120 and an image acquisition device 140:
[0151] The image acquisition device 140 is used to acquire raw tire images and send the raw tire images to the processor 120.
[0152] The processor 120 is used to execute the steps in any of the above embodiments of the tire attribute information recognition method to obtain the tire attribute information recognition result.
[0153] The tire attribute information recognition system in this embodiment can be applied to scenarios such as vehicle tire assembly, tire maintenance and inspection, tire inventory management, and tire production quality control. Taking the application scenario of detecting whether a vehicle's tire assembly meets pre-installation requirements as an example, a specific space can be pre-defined in the factory as a detection area. When a vehicle enters the preset detection area, an image acquisition device 140 deployed in the detection area acquires images of the tires of the vehicle to be identified, obtaining raw tire images. These raw tire images are then transmitted to the processor 120. The processor 120 acquires the raw tire images and, based on executing the steps in any of the above embodiments of the tire attribute information recognition method, obtains the tire attribute information recognition result. Based on the tire attribute information recognition result, it determines whether the tire assembly is correct and decides whether to proceed to the next process step based on the determination result. Figure 12 As shown.
[0154] In one exemplary embodiment, such as Figure 13As shown, a tire attribute information recognition device 600 is provided, including: a tire image acquisition module 610, a tire region detection module 620, a tire image slicing module 630, a text region detection module 640, a label fusion module 650, and a text recognition module 660, wherein:
[0155] Tire image acquisition module 610 is used to acquire raw tire images;
[0156] The tire region detection module 620 is used to detect tire regions in the original tire image based on a trained tire region detection model, and obtain the location information of the tire regions. The tire region detection model is trained based on historical tire images carrying category labels and location information of the tire regions.
[0157] The tire image slicing module 630 is used to segment the tire region image from the original tire image based on the location information of the tire region, and to slice the tire region image to obtain a set of sliced images.
[0158] The text region detection module 640 is used to detect text regions in each slice image in the slice image set based on the trained text segmentation model, and obtain text region detection results. The text region detection results include the detection boxes of the text regions. The text segmentation model is trained based on the category labels, location information and segmentation masks of the historical slice images carrying the ground truth boxes of the text.
[0159] The label fusion module 650 is used to merge the detection boxes of the text regions based on the text region detection results to obtain the target text region detection results of the tire region image.
[0160] The text recognition module 660 is used to identify tire attribute information in tire region images based on the target text region detection results and the trained tire text recognition model, and obtain tire attribute information recognition results. The tire text recognition model is trained based on historical tire images carrying text labels.
[0161] In an exemplary embodiment, the tire attribute information recognition device 600 further includes an image compression module 670, used to compress the original tire image when the resolution of the original tire image is greater than a preset resolution threshold.
[0162] The tire region detection module 620 is also used to detect tire regions in the compressed tire image based on the trained tire detection model, and obtain the location information of the tire regions; and to determine the location information of the tire regions in the original tire image based on the compression relationship between the original tire image and the compressed tire image.
[0163] In an exemplary embodiment, the tire image slicing module 630 is further configured to slice the tire region image, extract a target slice image from the multiple slice images obtained from the division; select a reference slice image from the multiple slice images, and rotate the tire region image with the reference slice image as the center according to a preset rotation angle, wherein the reference slice image is a slice image containing the center point of the tire region image; slice the rotated tire region image, extract the target slice image from the multiple slice images obtained from the division; return to the step of selecting a reference slice image from the multiple slice images and rotating the tire region image with the reference slice image as the center according to a preset rotation angle, until the sum of the rotation angles of the tire region image is equal to or greater than 360 degrees, thereby obtaining a slice image set containing multiple target slice images.
[0164] In an exemplary embodiment, the tire image slicing module 630 is further configured to slice the tire region image based on a preset nine-grid slicing strategy, and extract the target slice image within a fixed region from the slice images within the divided nine-grid region.
[0165] In an exemplary embodiment, the label fusion module 650 is further configured to map the text region detection results of the slice image set to the coordinate space of the tire region image through inverse affine transformation; for any two detection boxes with the same category label, determine the cross-union ratio (CUP) between the two detection boxes based on the position information of the detection boxes; if the CUP is greater than a preset CUP threshold, merge the segmentation masks of the tire text in the two detection boxes to obtain the target text region detection result of the tire region image.
[0166] In an exemplary embodiment, the tire attribute information recognition device 600 further includes a polar coordinate transformation module 680, which performs polar coordinate transformation on the tire region image, adjusts the direction of the tire attribute information to the horizontal direction, and updates the detection box in the target text region detection result; based on the updated target text region detection result, a target image is cropped from the transformed tire region image, and the target image is the region image within the detection box.
[0167] The text recognition module 660 is also used to input the target image into the trained tire text recognition model to obtain the tire attribute information recognition result.
[0168] In an exemplary embodiment, the polar coordinate transformation module 680 is further configured to rotate the tire region image with the center point of the tire region image as the polar coordinate expansion center and the horizontal axis of the image as the polar coordinate horizontal expansion line, until the mask region in the target text region detection result intersects with the polar coordinate horizontal expansion line, and so on.
[0169] In an exemplary embodiment, the tire attribute information recognition device 600 further includes a slice image verification module 690, which is used to update the mask of the identification area of tire attribute information in the tire region image according to the target slice image, and after obtaining the slice image set, verify whether the updated mask is all the target value.
[0170] The tire image slicing module 630 is also used to re-extract the slice image set based on the adjusted slicing extraction strategy or the adjusted preset rotation angle, provided that the mask of the updated marked area has an initial value.
[0171] Each module in the aforementioned tire attribute information recognition device 600 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0172] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 14 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When executed by the processor, the computer program implements a tire attribute information recognition method.
[0173] Those skilled in the art will understand that Figure 14 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0174] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in any of the above embodiments of the tire attribute information recognition method.
[0175] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps in any of the above embodiments of the tire attribute information identification method.
[0176] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the tire attribute information recognition method.
[0177] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0178] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0179] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0180] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for identifying tire attribute information, characterized in that, The method includes: Obtain the raw tire image; Based on the trained tire region detection model, tire regions in the original tire image are detected to obtain the location information of the tire regions. The tire region detection model is trained based on historical tire images carrying category labels and location information of the tire regions. Based on the location information of the tire region, the tire region image is segmented from the original tire image, and the tire region image is sliced to obtain a sliced image set; Based on the trained text segmentation model, the text region in each slice image in the slice image set is detected to obtain the text region detection result. The text region detection result includes the detection box of the text region. The text segmentation model is trained based on the category label, location information and segmentation mask of the historical slice images carrying the ground truth box of the text. Based on the text region detection results, the detection boxes of the text regions are merged to obtain the target text region detection results of the tire region image. Based on the target text region detection results and the trained tire text recognition model, tire attribute information in the tire region image is identified to obtain tire attribute information recognition results. The tire text recognition model is trained based on historical tire images carrying text labels.
2. The method according to claim 1, characterized in that, Before obtaining the tire region location information by detecting the tire region in the original tire image based on the trained tire region detection model, the method further includes: If the resolution of the original tire image is greater than a preset resolution threshold, the original tire image is compressed. The trained tire region detection model detects tire regions in the original tire image to obtain the location information of the tire regions, including: Based on the trained tire detection model, the tire region in the compressed tire image is detected to obtain the location information of the tire region. Based on the compression relationship between the original tire image and the compressed tire image, the location information of the tire region in the original tire image is determined.
3. The method according to claim 2, characterized in that, The step of slicing the tire region image to obtain a sliced image set includes: The tire region image is sliced, and the target slice image is extracted from the multiple slice images obtained; A reference slice image is selected from multiple slice images. The tire region image is rotated around the reference slice image by a preset rotation angle. The reference slice image is a slice image that includes the center point of the tire region image. The rotated tire region image is sliced, and the target slice image is extracted from the multiple slice images obtained. Returning to the step of selecting a reference slice image from multiple slice images, and rotating the tire region image with the reference slice image as the center according to a preset rotation angle, until the sum of the rotation angles of the tire region image is equal to or greater than 360 degrees, a slice image set containing multiple target slice images is obtained.
4. The method according to claim 3, characterized in that, The tire region image is sliced, and the target slice image is extracted from the multiple slice images obtained, including: Based on a preset nine-grid slicing strategy, the tire area image is sliced, and the target slice image within a fixed area is extracted from the slice images within the divided nine-grid area.
5. The method according to claim 3, characterized in that, The text region detection results include the segmentation mask of the tire text, the category label of the detection box, and the location information; Based on the text region detection results, the detection boxes of the text regions are merged to obtain the target text region detection results of the tire region image, including: By using inverse affine transformation, the text region detection results of the slice image set are mapped to the coordinate space of the tire region image; For any two bounding boxes with the same category label, determine the intersection-union ratio between the two bounding boxes based on their position information; If the cross-union ratio is greater than a preset cross-union ratio threshold, the segmentation masks of the tire text in the two detection frames are merged to obtain the target text region detection result of the tire region image.
6. The method according to claim 5, characterized in that, After obtaining the target text region detection result of the tire region image, the method further includes: The image of the tire region is transformed by polar coordinates to adjust the direction of the tire attribute information to the horizontal direction, and the detection result of the target text region is updated. Based on the detection boxes in the updated target text region detection results, the target image is cropped from the transformed tire region image, and the target image is the region image within the detection boxes; The method, based on the target text region detection result and the trained tire text recognition model, identifies tire attribute information in the target text region detection result to obtain tire attribute information recognition results, including: The target image is input into a trained tire text recognition model to obtain tire attribute information recognition results.
7. The method according to claim 6, characterized in that, Before performing polar coordinate transformation on the tire region image, the method further includes: Using the center point of the tire region image as the polar coordinate expansion center and the horizontal axis of the image as the polar coordinate horizontal expansion line, when the mask region in the target text region detection result intersects with the polar coordinate horizontal expansion line, rotate the tire region image until the mask region in the target text region detection result does not intersect with the polar coordinate horizontal expansion line.
8. A tire attribute information identification device, characterized in that, The device includes: Tire image acquisition module, used to acquire raw tire images; The tire region detection module is used to detect tire regions in the original tire image based on a trained tire region detection model, and obtain the location information of the tire regions. The tire region detection model is trained based on historical tire images carrying category labels and location information of the tire regions. The tire image slicing module is used to segment the tire region image from the original tire image based on the location information of the tire region, and to slice the tire region image to obtain a set of sliced images; The text region detection module is used to detect text regions in each slice image of the slice image set based on the trained text segmentation model, and obtain text region detection results. The text region detection results include detection boxes of the text regions. The text segmentation model is trained based on historical slice images carrying the category labels, location information and segmentation masks of the ground truth boxes of the text. The label fusion module is used to merge the detection boxes of the text regions based on the text region detection results to obtain the target text region detection results of the tire region image. The text recognition module is used to identify tire attribute information in the tire region image based on the target text region detection result and the trained tire text recognition model, and obtain the tire attribute information recognition result. The tire text recognition model is trained based on historical tire images carrying text labels.
9. A tire attribute information recognition system, characterized in that, The system includes a processor and an image acquisition device: The image acquisition device is used to acquire an image of the tire to be identified, obtain an original tire image, and send the original tire image to the processor; The processor is used to execute the tire attribute information recognition method as described in any one of claims 1 to 7, and obtain the tire attribute information recognition result of the tire to be identified.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.