Training method for image recognition model and image recognition model of the same

The method for training an image recognition model by enlarging object images to fit within predetermined areas addresses the challenge of recognizing objects smaller than the minimum recognizable area, enhancing recognition rates and safety in applications like autonomous driving.

JP2025090486AActive Publication Date: 2025-06-17WHETRON ELECTRONICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024005573
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-01-17
Publication Date
2025-06-17
Estimated Expiration
2044-01-17

AI Technical Summary

Technical Problem

Conventional image recognition models struggle to recognize objects smaller than their minimum recognizable area, leading to incomplete or missing recognition results, which can pose safety risks in applications like autonomous driving.

Method used

A method for training an image recognition model that involves determining whether an object image can be completely covered by a predetermined area without exceeding its boundaries. If not, the object image is enlarged to an adjusted size that can be completely covered by one or more minimum recognition areas, allowing for effective training and recognition.

Benefits of technology

The trained image recognition model achieves better recognition ability, including the ability to recognize objects smaller than the minimum recognition area, thereby improving recognition rates and reducing safety risks in applications like autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025090486000001_ABST
    Figure 2025090486000001_ABST
Patent Text Reader

Abstract

To provide a training method for an image recognition model that achieves the effect of imparting further superior recognition performance to the image recognition model, and the image recognition model of the same.SOLUTION: A training method for an image recognition model comprises: an object image enlargement step for, if it is determined that an image area enclosed by an object image cannot be completely covered by a predetermined surface area without exceeding this image area, enlarging the entire object image to an adjusted object image such that the image area enclosed by the adjusted object image can be completely covered by the predetermined surface area without exceeding this image area; and a step for acquiring a trained image recognition model by inputting the image processed through the object image enlargement step into the image recognition model to be trained.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image processing technology, and in particular, to a method for training an image recognition model and the image recognition model thereof.

Background Art

[0002] In a conventional image recognition model, particularly in an image recognition model constructed by a semantic segmentation model, training is performed by inputting a recognition target image having various object tags. By the way, since this image recognition model has a minimum recognizable area, when the image area surrounded by the object images included in the recognition target image for model training cannot be completely covered by one minimum recognizable area without exceeding this image area, even an object having a corresponding object tag (Tag), the image recognition model cannot recognize an object smaller than the minimum recognizable area of the recognition model from the recognition target image during training and after training is completed. Such a situation where recognition is impossible deteriorates even more when the overall image size is reduced in advance for the recognition target image before training or recognition in order to improve the training or recognition speed of the image recognition model.

[0003] The problems occurring in the above conventional image recognition model are shown in FIG. 8. Reference numeral M' denotes a conventional image recognition model having a minimum recognizable area Am'. Reference numeral I' denotes a recognition target image, and this recognition target image I' has a pixel point p' in the first axis direction and a pixel point q' in the second axis direction, and includes at least one object image IO'. And this object image IO' may be classified into a first type object image IOa', a second type object image IOb', and a third type object image IOc'. Furthermore, since the image area enclosed by the first type of object image IOa’ can be completely covered without exceeding the image area by one or more minimum recognition areas Am’, it can be recognized and marked with the corresponding object display frame FOa’. Inside this object display frame FOa’, it is a completely recognized first type of object. Since the entire image area enclosed by the second type of object image IOb’ cannot be completely covered without exceeding the image area by one minimum recognition area Am’, it cannot be recognized (marked with the corresponding object display frame). Since a part of the image area enclosed by the third type of object image IOc’ cannot be completely covered without exceeding the image area by one minimum recognition area Am’, it cannot be recognized. However, since another part of the image area enclosed by the third type of object image IOc’ can be completely covered without exceeding the image area by one or more minimum recognition areas Am’, it can be recognized and marked with the corresponding object display frame Foc’. In this case, it may be considered that the third type of object image IOc’ is not completely recognized. In other words, if all or part of the image area corresponding to the object image IO’ cannot accommodate the minimum recognition area Am’, all or part of the object image IO’ cannot be recognized.

[0004] In particular, in the image recognition technology, that the object image IO’ can be completely recognized may be understood as that by appropriately changing the position and angle of the minimum recognition area Am’, the image area enclosed by the object image IO’ can be reproduced without exceeding any boundary of the object image IO’ through overlapping or non - overlapping joining. On the other hand, that all or part of the object image IO’ is unrecognizable may be understood as that no matter how the position and angle of the minimum recognition area Am’ are changed, at least one of the boundaries of the object image IO’ will be exceeded.

[0005] For example, as shown in FIG. 9, the minimum recognition area Am’ has a minimum recognition pixel size pm’×qm’ consisting of the number of minimum recognizable pixel points pm’ in the first axis direction and the number of minimum recognizable pixel points qm’ in the second axis direction. Note that pm’ and qm’ are usually the same. Taking the case where the minimum recognition pixel size is 2 pixel points (pm’)×2 pixel points (qm’) as an example, for the image corresponding to the first type of object image IOa’, the number of pixel points pa’ in the first axis direction is 4, the number of pixel points qa’ in the second axis direction is 40, and correspondingly, when it has a pixel size of 4 pixel points (pa’)×40 pixel points (qa’), the image area defined by the first type of object image IOa’ can be completely covered by one or more minimum recognition areas Am’ without exceeding the image area. Therefore, the first type of object image IOa’ can be completely recognized.

[0006] For the image corresponding to the second type of object image IOb’, the number of pixel points pb’ in the first axis direction is 30, the number of pixel points qb’ in the second axis direction is 1, and correspondingly, when it has a pixel size of 30 pixel points (pb’)×1 pixel point (qb’), the entire image area defined by the second type of object image IOb’ cannot accommodate the minimum recognition area Am’ (that is, this minimum recognition area Am’ exceeds the image area defined by the second type of object image IOb’ no matter where it is located). Therefore, this second type of object image IOb’ cannot be recognized at all.

[0007] For the pixel size of the image corresponding to the third type of object image IOc’, for example, a threshold edge CB’ is included. When the number of pixel points of this threshold edge CB’ is exactly equal to the smaller number of pixel points in the minimum recognizable pixel size and the number of pixel points of the threshold edge CB’ is defined as the number of threshold pixel points, the third type of object image IOc’ is divided into a first partial image IOc1’ and a second partial image IOc2’ on both sides of the side with the number of threshold pixel points of the threshold edge CB’ greater than or equal to the number of threshold pixel points and the side smaller than the number of threshold pixel points of the threshold edge CB’, respectively. In the first partial image IOc1’, the number of pixel points on one side of the threshold edge CB’ is equal to or greater than the number of threshold pixel points, and the image area defined by the first partial image IOc1’ can be completely covered without exceeding the image area by one or more of the minimum recognition areas Am’. Therefore, a part of the third-class object image IOc’ (the part corresponding to the first partial image IOc1’) can be recognized. In the second partial image IOc2’, the number of pixel points on the other side of the threshold edge CB is less than the number of threshold pixel points, and the entire image area defined by the second partial image IOc2’ cannot accommodate the minimum recognition area Am’. Therefore, another part of the third-class object image IOc’ (the part corresponding to the second partial image IOc2’) cannot be recognized. In other words, when the number of maximum pixel points of the boundary of a partial area of the object image IO’ is less than the number of the smallest pixel points in the minimum recognizable pixel size, this partial area of the object image IO’ cannot be recognized.

[0008] Here, if the entire second-class object image IOb’ cannot be recognized, or if a part of the third-class object image IOc’ cannot be recognized, the recognition result will be missing / incomplete, which may lead to safety issues. For example, in applications where the image recognition function is used in driving, such as when performing the driving or parking function of a transportation means, if an obstacle cannot be completely recognized, there may be a lack of information used in the determination of the recognition result by the driving assistance system or the autonomous driving system of the transportation means, increasing the risk of danger due to unexpected collisions. Alternatively, if the auxiliary markings in the space (such as lanes and parking lines) cannot be completely recognized, the occurrence rate of unexpected deviations may increase, and the possibility of potential danger may increase.

[0009] In view of this, the conventional image processing technology definitely has room for improvement.

Prior Art Documents

Patent Documents

[0010]

Patent Document 1

Summary of the Invention

[0011] In order to solve the above problems, an object of the present invention is to provide a method for training an image recognition model that improves the recognition ability of the constructed image recognition model.

[0012] Another object of the present invention is to provide an image recognition model that improves the recognition rate of object images smaller than the minimum recognition area.

[0013] The quantifiers "1" and "1 piece" used for the elements and members described throughout the present invention are used only for convenience of use and to provide the normal meaning within the scope of the present invention. In the present invention, unless otherwise specified, it should be interpreted as including one or at least one, and also including a plurality in the concept of a single number.

[0014] Throughout the present invention, the above-mentioned "image recognition model" and its training are executed by a "computer". The above-mentioned "computer" may include at least one "processor". As can be understood by those skilled in the art, the above-mentioned processor is a hardware or various data processing devices realized as hardware and software having specific functions for processing and analyzing information and / or generating corresponding control information. Examples include electronic controllers, servers, cloud platforms, virtual devices, desktop computers, notebook computers, tablet computers, smartphones, and the like. In addition, for receiving or transmitting necessary data, a corresponding data receiving or transmitting unit may be included. Furthermore, for reading and storing data, a corresponding database or memory cell (especially a non-temporary memory cell) may be included. In particular, unless excluded or contradictory, the processor may be a set of multiple processors based on a distributed system structure for including or displaying the process, mechanism, and results of information streaming processing among the multiple processors.

[0015] The method for training an image recognition model of the present invention includes an image input step of obtaining a recognition target image including a sub-region including an object image and having an image area equal to or larger than the image area constituting the object image, and determining whether the image region surrounded by the object image can be completely covered by one or more predetermined areas without exceeding the image region, which includes an object image determination step. When it is determined that the whole or a part of the image region surrounded by the object image cannot be completely covered by a predetermined area without exceeding the image region, the whole object image is enlarged to an adjusted object image, and the image region surrounded by the adjusted object image can be completely covered by one or more predetermined areas without exceeding the image region, which includes an object image enlargement step. The method further includes a model training step of inputting the recognition target image including the adjusted object image into the image recognition model to be trained and performing training to obtain a trained image recognition model. The image recognition model to be trained has a minimum recognition area, and the image region surrounded by the adjusted object image can be completely covered by one or more minimum recognition areas without exceeding the image region, and the predetermined area is equal to or larger than the minimum recognition area.

[0016] According to this, the method for training an image recognition model of the present invention can achieve the effect that the image recognition model after training has better recognition ability by using the adjusted object image obtained by the object image determination step and the object image enlargement step for training and constructing the corresponding model. In particular, it can achieve the effect of recognizing an object image smaller than the minimum recognition area.

[0017] The predetermined area is equal to the minimum recognition area. Thus, the method for training an image recognition model is applicable whether the corresponding input object image undergoes dimensionality reduction processing before the object image enlargement step or not, and also applicable whether the corresponding input object image undergoes dimensionality reduction processing or not between the object image enlargement step and the model training step, and an effect that the trained image recognition model has better recognition ability can be achieved.

[0018] The image to be recognized has an image pixel size defined by the product of the number of pixel points in the first axis direction and the number of pixel points in the second axis direction. The predetermined area has a predetermined pixel size defined by the product of the number of pixel points in the first axis direction and the number of pixel points in the second axis direction. One or both of the number of pixel points in the first axis direction and the number of pixel points in the second axis direction of the predetermined area are values defined between 0.008 times and 0.08 times the number of pixel points in the first axis direction or the number of pixel points in the second axis direction of the image to be recognized. Thus, by the relationship between the image pixel size and the predetermined pixel size, the predetermined pixel size can be set within a reasonable range, and an effect that the trained image recognition model has better recognition ability can be achieved.

[0019] The minimum recognition area has a minimum recognition pixel size defined by the product of the number of pixel points in the first axis direction and the number of pixel points in the second axis direction. The number of pixel points in the first axis direction and the number of pixel points in the second axis direction are each at least 1 pixel point. Thus, by defining the range of the minimum pixel size, an effect that the trained image recognition model has better recognition ability can be achieved.

[0020] Before executing the model training step, reduce the image pixel size of the entire image to be recognized. Thus, it has the effect of improving the efficiency of the subsequent image training process.

[0021] The image recognition model to be trained is a semantic segmentation model. In this way, since this image recognition model has an appropriate model structure, the effect that the image recognition model after training has better recognition ability can be achieved.

[0022] The image recognition model of the present invention includes a minimum recognition area defined as the minimum resolution that the image recognition model can recognize. The image recognition model is for recognizing an object image in the image to be recognized. The object image includes a non-primitively recognizable object image, or the object image includes a primitively recognizable object image and a non-primitively recognizable object image. The primitively recognizable object image is defined such that the entire image area surrounded thereby is completely covered by one or more minimum recognition areas without exceeding the image area surrounded thereby. The non-primitively recognizable object image is defined such that the entire or part of the image area surrounded thereby cannot be completely covered by the minimum recognition area without exceeding the image area surrounded thereby.

[0023] According to this, the image recognition model of the present invention has the ability to recognize an object image smaller than the minimum recognition area (i.e., a non-primitively recognizable object image which is a type 2 object image and / or a type 3 object image), and the effect of improving the recognition rate of the object image smaller than the minimum recognition area can be achieved.

[0024] When the image recognition model recognizes an object image, it generates a corresponding object display frame on the outer periphery of the image area of the object image. In this way, by forming the object display frame, the boundary position information of the object to be recognized can be limited. Furthermore, when the image recognition model is used for the driving purpose of a transportation means, the effects of obstacle identification and collision avoidance can be achieved.

[0025] When the image recognition model recognizes an object corresponding to an original recognizable object image, it generates a corresponding original object display frame on the outer periphery of the image area of the original recognizable object image. There is a first distance between the original object display frame and the boundary of the image area of the original recognizable object image. When the image recognition model recognizes an object corresponding to a non-original recognizable object image, it generates a corresponding non-original object display frame on the outer periphery of the image area of the non-original recognizable object image. There is a second distance between the non-original object display frame and the boundary of the image area of the non-original recognizable object image, and the second distance is greater than the first distance. In this way, the image recognition model of the present invention has the ability to recognize an object image smaller than the minimum recognition area (that is, a non-original recognizable object image that is a type 2 object image and / or a type 3 object image), and by giving a relatively large range of non-original object display frames to the non-original recognizable object image, the effect of reducing the possibility of collision with the object corresponding to the non-original recognizable object image can be achieved.

[0026] The image recognition model is constructed based on the training method of the above image recognition model. In this way, the effect of ensuring the ability of the image recognition model to recognize an object image smaller than the minimum recognition area (that is, a non-original recognizable object image that is a type 2 object image and / or a type 3 object image) can be achieved.

Brief Description of the Drawings

[0027]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0028] In order to more clearly clarify the above and other objects, features, and advantages of the present invention, hereinafter, preferred embodiments of the present invention will be given, and the present invention will be described in detail with reference to the accompanying drawings. Note that those denoted by the same reference numerals in different drawings are the same, and the description thereof will be omitted.

[0029] Referring to FIG. 1, FIG. 1 is a training procedure of a training method for an image recognition model of the present invention. This training method includes, by a computer, an image input step S1, an object image determination step S2, an object image enlargement step S3, a selectable overall image dimension reduction step S3A, and a model training step S4, and executes the steps.

[0030] In particular, referring to FIGS. 2 to 5, in order to more clearly explain the content such as image processing and recognition mechanism, schematic diagrams of the image input step S1, the object image determination step S2, the object image enlargement step S3, and the model training step S4 are shown.

[0031] Corresponding to FIG. 2, in the image input step S1, a recognition target image I is acquired. A sub-region in this recognition target image I has an object image IO, and the image area constituting this object image IO is less than or equal to the image area of the recognition target image I. Specifically, this image I to be recognized has an image pixel size defined by the product of the number p of pixel points in the first axis direction and the number q of pixel points in the second axis direction.

[0032] Corresponding to FIG. 3, in the object image determination step S2, it is determined whether the image area can be completely covered by one or a plurality of preset areas Ap without exceeding the image area surrounded by the object image IO. The preset area Ap has a preset pixel size defined by the product of the number pp of pixel points in the first axis direction and the number qp of pixel points in the second axis direction. In particular, one or both of the number pp of pixel points in the first axis direction and the number qp of pixel points in the second axis direction of this preset area Ap may be values defined between 0.008 times and 0.08 times the number p of pixel points in the first axis direction or the number q of pixel points in the second axis direction of the image I to be recognized.

[0033] Based on this determination result, the object image IO may be classified into a first - type object image IOa, a second - type object image IOb, and a third - type object image IOc. And the image area surrounded by this first - type object image IOa can be completely covered by a plurality of preset areas Ap without exceeding the range of the image area. Also, the entire image area surrounded by the second - type object image IOb cannot be completely covered by the preset area Ap without exceeding the range of the image area. To more clearly explain the above situation, the entire image area is shown as a grayscale area in FIG. 3. In other words, the preset area Ap exceeds at least one boundary of the image area at any position (especially at any angle) in the image area of the entire second - type object image IOb. Furthermore, a part of the image area surrounded by the third - type object image IOc cannot be completely covered by the preset area Ap without exceeding the range of the image area. To more clearly explain the above situation, a part of the image area is shown as a grayscale area in FIG. 3. In other words, the predetermined area Ap exceeds at least one boundary of the image area at any position (especially at any angle) in a partial image area of the Class 3 object image IOc.

[0034] When it is determined that the entire image area (i.e., corresponding to the Class 2 object image IOb) or a part (i.e., corresponding to the Class 3 object image IOc) surrounded by the object image IO corresponding to FIG. 4 cannot be completely covered without exceeding the range of the image area, the object image enlargement step S3 is executed. In this object image enlargement step S3, the entire object image IO is enlarged to an adjusted image of object (AIO). And the image area surrounded by this adjusted object image AIO can be completely covered by one or more predetermined areas Ap without exceeding this image area. In other words, after executing the object image enlargement step S3, the object images IO (especially, the signs IOb and IOc) in the recognition target image I are replaced with the adjusted object image AIO (especially, the adjusted Class 2 object image represented by the sign AIOb and the adjusted Class 3 object image represented by the sign AIOc). Furthermore, the image area of this adjusted object image AIO is larger than the image area of the corresponding object image IO and is less than or equal to the image area of the recognition target image I.

[0035] In the image processing technology, since there are already many methods for enlarging the entire image that can be understood by those skilled in the art, in the present invention, the corresponding enlargement methods will not be repeatedly described. Also, the enlargement result shown in FIG. 4 is merely a specific example for explanation, and the present invention can achieve a result such that the image area surrounded by the adjusted object image AIO can be completely covered by one or more predetermined areas Ap without exceeding this image area by using an image enlargement method already proposed at present or proposed in the future.

[0036] The object image IO, particularly the adjusted object image AIO, has a tag for recording the classification of the object corresponding to the object image IO. Preferably, the tag of the adjusted object image AIO further records magnification information about the object image IO before adjustment, such as the magnification ratio or the magnification method.

[0037] Optionally, following the object image magnification step S3, by executing this overall image dimension reduction step S3A, the image pixel size / image area of the entire recognition target image I is reduced, and in particular, the image pixel size of the entire recognition target image I is reduced based on the reduction rate value Vr. Thereby, the efficiency of the subsequent image training process is improved. Specifically, in the overall image dimension reduction step S3A, while the image pixel size of the entire recognition target image I is reduced based on the reduction rate Vr, the adjusted object pixel size corresponding to the adjusted object image AIO (including the adjusted second-class object image AIOb or the adjusted third-class object image AIOc) is also reduced based on the reduction rate Vr. When the recognition target image I has another object image IO (i.e., one that does not require the object image magnification step S3, such as the first-class object image IOa), the object pixel size corresponding to the object image IO is also reduced based on the reduction rate Vr. Note that in image processing technology, there are already many methods for reducing the dimension of the entire image that can be understood by those skilled in the art, so in the present invention, the corresponding dimension reduction methods will not be repeatedly described. In particular, the overall image dimension reduction step S3A is not limited to being after the object image magnification step S3, and may also be any step before the execution of the model training step S4.

[0038] Following the object image magnification step S3 or the overall image dimension reduction step S3A, the model training step S4 is executed, and by inputting the recognition target image I including the adjusted object image AIO into the image recognition model to be trained for training, a trained image recognition model is obtained. This image recognition model is a model constructed using machine learning technology. In particular, in the specific training process, the present invention realizes corresponding results by means of an Image Semantic Segmentation Model. Note that the above semantic segmentation model is used only as an image recognition model training method for more clearly explaining the operation based on the present invention. Therefore, as long as the above steps are satisfied, the method of the image recognition model of the present invention is not limited thereto.

[0039] As shown in FIG. 5, the image recognition model to be trained has a minimum recognition area Am, the minimum recognition area Am has a minimum recognition pixel size, and the minimum recognition pixel size may be defined by the product of the number of pixel points pm in the first axis direction and the number of pixel points qm in the second axis direction. The minimum recognition pixel size is at least 1 pixel point (pm) × 1 pixel point (qm). Referring to FIGS. 4 and 5, the image area surrounded by the adjusted object image AIO (which may or may not be the image obtained by the processing of the overall image dimension reduction step S3A) can be completely covered by one or more minimum recognition areas Am without exceeding the range of the image area. Thus, the adjusted object image AIO (shown in FIG. 4) can be recognized by the image recognition model during training or after training, and a corresponding object display frame FO (shown in FIG. 5) is generated. Therefore, for the adjusted second-class object image AIOb and the adjusted third-class object image AIOc, during and after the training of the image recognition model, the image recognition model can generate corresponding second-class object image display frames FOb and third-class object image display frames FOc respectively. Note that the first-class object image IOa can be recognized regardless of whether the object image enlargement step S3 is performed, and has a corresponding first-class object display frame FOa. In particular, after the object image enlargement step S3, the adjusted second-class object image AIOb and the adjusted third-class object image AIOc may be regarded as having the characteristic that the image regions surrounded by the first-class object IOa are all recognizable. In other words, in the present invention, for the recognition target image I, after the object image enlargement step S3, regardless of whether the overall image dimension reduction step S3A is performed, the features of the corresponding objects in all object images IO and all adjusted object images AIO are retained and can be completely recognized by a trained or training image recognition model.

[0040] Note that the predetermined area Ap is equal to or greater than the minimum recognition area Am. When the adjusted object image AIO is generated by the object image enlargement step S3 and then input into the image recognition model for training without performing the overall image dimension reduction step S3A, the predetermined area Ap is equal to or greater than the minimum recognition area Am, and preferably, it is equal to the minimum recognition area Am. When the adjusted object image AIO is generated by the object image enlargement step S3, then the overall image dimension reduction step S3A is performed, and the result is input into the image recognition model for training, the value obtained by multiplying the predetermined area Ap by the reduction magnification Vr is equal to or greater than the minimum recognition area Am, and preferably, it is equal to the minimum recognition area Am.

[0041] For the actually image recognition using the image recognition model trained by the above training method, the recognition results obtained by inputting the same recognition target image I into the comparison model and the image recognition model of the present invention respectively are shown in FIGS. 6 and 7. The recognition target image I is a peripheral image of a yacht 1 moored at a quay, and in particular, it is a peripheral image captured and synthesized by an imaging device attached to the yacht 1. The recognition target image I includes a quay passage 2 and a plurality of columns / bars 3 on the water surface for defining the area for mooring ships. The comparison model and the image recognition model of the present invention have the same minimum recognition area, and the entire image area corresponding to a plurality of these columns 3 (corresponding to the second type of object image IOb) cannot be completely covered by the minimum recognition area without exceeding the image area. A part of the image area of one of these columns 3 (corresponding to the third type of object image IOc) cannot be completely covered by the minimum recognition area without exceeding the image area.

[0042] As shown in FIG. 6, when the recognition target image I is input to the comparison model and the comparison model directly recognizes the recognition target image I, only a relatively large-area partial image area corresponding to the quay passage 2 and only one column 3 is recognized, and the corresponding object display frame FO is generated. Specifically, during the training process, the comparison model directly performed training using the unprocessed recognition target image I. In particular, the unprocessed recognition target image I is one in which the object image determination step S2 and the object image enlargement step S3 of the present invention have not been performed.

[0043] As shown in FIG. 7, when the recognition target image I is input to the image recognition model of the present invention and the image recognition model directly recognizes the recognition target image I, the quay passage 2 and all the columns 3 can be recognized, and the corresponding object display frame FO is generated. The quay passage 2 corresponds to the first type of object image IOa, and these columns 3 correspond to the second type of object image IOb and the third type of object image IOc. These object display frames FO may be regarded as the corresponding first object display frame FOa, second object display frame FOb, and third object display frame Foc. In other words, in the image recognition model constructed based on the training method provided by the present invention, during the training process of this image recognition model, based on the same minimum recognition area as the model illustrated in FIG. 6, and based on the recognition target images I for training such as the type-2 object image IOb and the type-3 object image IOc, these object images are first processed by the object image enlargement step S3 and then used for the training and construction of the image recognition model. As a result of image recognition, an unexpected effect is produced. Even in the case of the same recognition input information and the minimum recognition area of the model, an object that cannot be recognized by the comparison model (corresponding to the recognition result in FIG. 6) can be recognized by the image recognition model of the present invention (corresponding to the recognition result in FIG. 7). In other words, without changing the minimum recognizable resolution (corresponding to the minimum recognition area) of the image recognition model of the present invention, the image recognition model of the present invention can recognize an object that cannot be recognized by a conventional model. The image recognition model constructed by the training method provided by the present invention is particularly used for recognizing elongated objects, but is not limited thereto. In addition, for an object image that cannot be completely covered by the minimum recognition area without exceeding the image area of the object (i.e., corresponding to the type-2 object image IOb and the type-3 object image IOc) (i.e., an object that cannot be recognized by the comparison model but can be recognized by the image recognition system of the present invention), the boundary / area of the object display frame FO generated by the present invention (i.e., corresponding to the type-2 object display frame FOb and the type-3 object display frame FOc) has the characteristic that it is larger than the boundary / area of these object images. In particular, a small boundary / area is located inside a large boundary / area.

[0044] In particular, based on the above, in the image recognition model provided by the present invention, object images can be classified into a primitive recognizable object image (corresponding to the type-1 object image IOa) and a non-primitive recognizable object image (corresponding to the type-2 object image IOb and the type-3 object image IOc). The primitive recognizable object image is defined as follows. The entire image area of the original recognizable object image can be completely covered by one or more minimum recognition areas without exceeding the image area surrounded by the original recognizable object image. The image recognition model is for recognizing the object corresponding to the original recognizable object image and generating a corresponding original object display frame (corresponding to the first-class object display frame FOa as shown in FIG. 5) on the outer periphery of the image area of the original recognizable object image. In particular, there may be a first distance between the original object display frame and the boundary of the image area of the original recognizable object image. The non-original recognizable object image is defined as follows. The whole or part of the image area of the non-original recognizable object image cannot be completely covered by the minimum recognition area without exceeding the image area surrounded by the non-original recognizable object image. The image recognition model is for recognizing the object corresponding to the non-original recognizable object image and generating a corresponding non-original object display frame (corresponding to the second-class object display frame FOb and the third-class object display frame F Oc as shown in FIG. 5) on the outer periphery of the image area of the non-original recognizable object image. In particular, there is a second distance greater than the first distance between the non-original object display frame and the boundary of the image area of the non-original recognizable object image.

[0045] To more clearly explain the training method of the image recognition model of the present invention and the remarkable technical features and effects of the image recognition model of the present invention, the above content is summarized in Table 1 below.

[0046]

Table 1

[0047] Here, in Table 1, the differences between the image recognition model of the present invention (the "recognition model of the present invention" in Table 1, an example of the recognition result is shown in FIG. 7) and the conventional comparative image recognition model (the "conventional comparative recognition model" in Table 1, an example of the recognition result is shown in FIG. 6) are compared. In the comparison of the training process, in the image recognition model of the present invention, the processing method for the first-class object image IOa may be the same as the processing methods for the first-class object image IOa, the second-class object image IOb, and the third-class object image IOc by the conventional comparative image recognition model. On the contrary, as described above, in the present invention, for the second-class object image IOb and the third-class object image IOc, before performing the corresponding model training, first, "expansion by a specific method" is performed. The "expansion by a specific method" means expanding the entire corresponding object image to an adjusted object image so that the image area is completely covered by one or more predetermined areas without exceeding the image area surrounded by this adjusted object image. In this way, the image recognition model constructed based on the above training method of the present invention can completely identify the entire contour of the corresponding object from the original or non-primitive identifiable object image.

[0048] Summarizing the above, the training method of the image recognition model of the present invention uses the adjusted object image obtained by the object image determination step and the object image expansion step for the training and construction of the corresponding model, so that the trained image recognition model has better recognition ability. In addition, by using the image recognition model constructed based on the above training method, the recognition rate of recognizing an object image smaller than the minimum recognition area can be improved.

[0049] It should be noted that the term "completely cover" described in the present invention (in particular, completely cover the entire corresponding image area without exceeding the image area of the object image corresponding to a specific object by one or more minimum recognition areas) includes or is equivalent to that the object contour corresponding to the specific object can be completely identified by "substantially covering" the image area corresponding to the specific object. Here, in the case of actual "substantially covering", there is a very small amount of image area that is ignored or not covered. However, from the perspective of the limits and recognition accuracy of image recognition, the features within the recognized range in the object image corresponding to the specific object already satisfy the feature conditions (size, color, proportion, pattern / arrangement of parts, etc.) necessary to identify the entire object. Thus, when the generated display frame can include the contour of the corresponding object image, the terms "completely cover" or "substantially cover" described in the present invention may be regarded as such. Therefore, those skilled in the art can understand that the term "completely cover" is for clearly defining and explaining the gist and spirit of the technical features of this application and is equivalent to the aforementioned "substantially cover".

[0050] Although the present invention has been disclosed by the above embodiments, the present invention is not limited thereto. Various changes and modifications made to the above embodiments by those skilled in the art without departing from the spirit and scope of the present invention still fall within the technical scope protected by the present invention. Therefore, the protection scope of the present invention should include the literal meaning described in the appended claims and all modifications within the equivalent scope thereof. In addition, when the above multiple embodiments can be combined, the present invention includes all combination embodiments.

Explanation of Reference Numerals

[0051] 1: Yacht 2: Dock Passage 3: Column AIO: Adjusted Object Image AIOb: Adjusted Object Image of the Second Type AIOc: Adjusted Object Image of the Third Type Am: Minimum Recognition Area Ap: Predetermined Area FO: Object Display Frame FOa: Object Display Frame of the First Type FOb: Object Display Frame of the Second Type FOc: Object Display Frame of the Third Type I: Image to be Recognized IO: Object Image IOa: Object Image of the First Type IOb: Image of the second type of object IOc: Image of the third type of object p, pm, pp: Number of pixel points in the first axis direction q, qm, qp: Number of pixel points in the second axis direction S1: Image input step S2: Object image determination step S3: Object image enlargement step S3A: Overall image dimension reduction step S4: Model training step Am’: Minimum recognition area FO’: Object display frame FOa’: First type of object display frame FOc’: Third type of object display frame I’: Image to be recognized IO’: Object image IOa’: First type of object image IOb’: Second type of object image IOc’: Third type of object image p’, pm’: Number of pixel points in the first axis direction q’, qm’: Number of pixel points in the second axis direction

Claims

1. An image input step including acquiring a recognition target image including a sub-region including an object image and having an image area equal to or larger than an image area constituting the object image; an object image determination step including determining whether an image area enclosed by the object image can be completely covered by one or more predetermined areas without exceeding the image area; an object image enlargement step including, when it is determined that the whole or part of the image area surrounded by the object image cannot be completely covered by the predetermined area without exceeding the image area, enlarging the entire object image into an adjusted object image so that the image area surrounded by the adjusted object image can be completely covered by one or more of the predetermined areas without exceeding the image area; A model training step includes inputting the recognition target image including the adjusted object image into an image recognition model to be trained, thereby obtaining a trained image recognition model; the training target image recognition model has a minimum recognition area and can completely cover the image area with one or more minimum recognition areas without exceeding the image area enclosed by the adjusted object image; A method for training an image recognition model, wherein the predetermined area is greater than or equal to the minimum recognition area.

2. The method of training an image recognition model of claim 1 , wherein the predetermined area and the minimum recognition area are equal.

3. the recognition target image has an image pixel size defined by the product of the number of pixel points in the first axis direction and the number of pixel points in the second axis direction; the predetermined area has a predetermined pixel size defined by a product of a number of first axial pixel points and a number of second axial pixel points; 2. The method for training an image recognition model according to claim 1, wherein one or both of the number of pixel points in the first axis direction and the number of pixel points in the second axis direction of the specified area are values ​​defined between 0.008 and 0.08 times the number of pixel points in the first axis direction or the number of pixel points in the second axis direction of the image to be recognized.

4. the minimum recognition area has a minimum recognition pixel size defined by the product of a number of pixel points in a first axis direction and a number of pixel points in a second axis direction; The method of training an image recognition model according to claim 1 , wherein the number of pixel points in the first axis direction and the number of pixel points in the second axis direction are each at least one pixel point.

5. The method for training an image recognition model according to claim 1 , further comprising the step of reducing an image pixel size of the entire recognition target image before performing the model training step.

6. The method for training an image recognition model according to claim 1 , wherein the image recognition model to be trained is a semantic segmentation model.

7. An image recognition model including a minimum recognition area defined as the minimum recognizable resolution, the image recognition model is for recognizing an object image in a recognition target image, The object image includes a non-original identifiable object image, or the object image includes a original identifiable object image and the non-original identifiable object image; The original identifiable object image is defined such that the entire image area is completely covered by one or more minimum recognition areas without exceeding the image area enclosed thereby; The non-original identifiable object image is defined such that no part or all of its image area can be completely covered by the minimum recognition area without exceeding the image area enclosed thereby, an image recognition model.

8. The image recognition model according to claim 7 , wherein, when the image recognition model recognizes the object image, the image recognition model generates a corresponding object display frame on a periphery of an image area of ​​the object image.

9. When the image recognition model recognizes an object corresponding to the original identifiable object image, a corresponding original object display frame is generated around the periphery of an image area of ​​the original identifiable object image; The original object display frame and a boundary of an image area of ​​the original discernible object image have a first distance; When the image recognition model recognizes an object corresponding to the non-original identifiable object image, a corresponding non-original object display frame is generated around the periphery of an image area of ​​the non-original identifiable object image; The image recognition model of claim 8 , wherein the non-primitive object representation frame and a boundary of an image region of the non-primitive discernible object image have a second distance that is greater than the first distance.

10. The image recognition model according to any one of claims 7 to 9, which is constructed based on the training method for an image recognition model according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method, apparatus, computer device and computer program for training a detection model

    JP2022518939A

  • TW762562