A multi-target small object recognition method based on category loss and difference detection
By adopting a method based on category loss and difference detection in Chinese chess recognition, combined with deep learning neural network and image difference detection, the problem of difficult to take into account the speed and accuracy of chess recognition in the existing technology is solved, and an efficient and robust recognition effect is achieved.
Patent Information
- Application Number
- CN202111318533.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-09
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-09
AI Technical Summary
The prior art is difficult to achieve fast and accurate recognition in Chinese chess recognition, especially when the target is small, the number is large, and the types are large, the existing deep learning networks are difficult to take into account both speed and accuracy, and are not robust enough.
A multi-object small object recognition method based on category loss and difference detection is adopted. The deep learning neural network ChessColor-Net and cross-entropy loss function plus penalty term are trained. Combined with image difference detection and threshold automatic adjustment, high-speed and high-precision classification recognition of object chess boards are achieved.
It has achieved rapid and accurate identification of Chinese chess boards, with an identification accuracy of 97.8%, and has high robustness in chess detection, and can maintain detection effect under different environmental conditions.
Smart Images

Figure CN114429584B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of object recognition, and relates to the technology of deep learning, image classification and OpenCV image processing, and specifically to a multi-target small object recognition method based on category loss and difference detection. Background Art
[0002] Image recognition is a hot topic in current artificial intelligence research. There has been extensive research on the use of convolutional neural networks for detection, classification, recognition, segmentation and other tasks. Existing open source convolutional neural networks, such as VGG, EfficientDent, MobileNet, ResNet and other networks, perform well on mainstream data sets. When the network is trained, a loss function is used to measure the quality of the training results. For example, functions such as Softmax and cross-entropy loss are used to compare the difference between the network output and the real target in classification tasks. The network parameters are updated according to the calculation results of the loss function to guide the network training, so that the network loss function gradually decreases and remains at a small value. The existing backbone network and loss function can certainly perform well on large data sets, but when it comes to specific recognition tasks, the effect is often unsatisfactory, and it is difficult to strike a balance between speed and accuracy.
[0003] Currently, in the field of Chinese chess recognition, the existing backbone network and loss function are not applicable due to the small size, large number and variety of targets.
[0004] Therefore, a new technical solution is needed to solve this problem. Summary of the invention
[0005] Purpose of the invention: In order to overcome the problems in the prior art that the Chinese chess recognition targets are small, large in number, many in types, and have high robustness requirements, and the existing deep learning network is difficult to quickly and accurately recognize Chinese chess, a multi-target small object recognition method based on category loss and difference detection is provided, which is particularly suitable for the recognition of chess. In terms of the classification and recognition of chess, a Chinese chess detection and recognition method based on chess type loss and difference detection is provided. For targets such as Chinese chess, high-speed and high-precision classification and recognition are achieved, and it has high robustness in chess move detection.
[0006] Technical solution: To achieve the above purpose, the present invention provides a multi-target small object recognition method based on category loss and difference detection, comprising the following steps:
[0007] S1: Collect the initial chessboard image, calibrate the chessboard image coordinates, and segment the chessboard image according to the coordinates;
[0008] S2: Preprocess the segmented chessboard image, classify the processed image using the built deep learning network model, and convert the classification result into chess game information;
[0009] S3: Collect chessboard image P1, and after the chess player makes a move, collect chessboard image P2 again;
[0010] S4: Based on the chessboard images P1 and P2, a difference detection method is used to judge the rationality of the starting point and the end point of the chess move. If reasonable, the coordinates of the starting point and the end point are converted into chess move information, and the chess game information is updated.
[0011] Furthermore, the step S1 specifically includes the following steps:
[0012] A1: Fix the chessboard position, set up a camera above the chessboard and fix the camera position to capture images of the chessboard and chess pieces;
[0013] A2: Calibrate the pixel coordinates of the points A, B, and C at the lower left corner, upper left corner, and lower right corner of the chessboard. Based on these three coordinates, calculate the pixel coordinates of the remaining 87 points on the chessboard.
[0014] A3: Capture images of all placement points.
[0015] Furthermore, the method for calculating the pixel coordinates of the drop point in step A2 is:
[0016] Calculate the pixel length and width of each chessboard point, as shown in formulas (1), (2), (3), and (4):
[0017]
[0018]
[0019]
[0020]
[0021] In the formula, x A 、x B 、x C is the horizontal coordinate of points A, B, and C, y A ,y B ,y C is the ordinate of points A, B, and C, dW x , dW y dL is the difference between the horizontal and vertical coordinates of each point on the width of the chessboard. x ,dL yis the difference between the horizontal and vertical coordinates of each point on the length of the chessboard. The 9 and 10 in the denominator are because the chessboard is 9 squares long and 10 squares wide.
[0022] Taking point A as the reference point, the coordinates of other drop points are calculated using formulas (5) and (6):
[0023] S x =x A +i L ×D x -i W ×D x (5)
[0024] S y =y A +i L ×D y -i W ×D y (6)
[0025] In the formula, S x and S y is the horizontal and vertical coordinates of any drop point, i L is the column where the currently calculated drop point is located, with a value range from 0 to 8, i W is the row where the currently calculated drop point is located, and the value range is from 0 to 9; i L Each time a value is taken, i W The values are calculated from 0 to 9, so the pixel coordinates of all 90 drop points can be calculated.
[0026] Furthermore, the deep learning network model in step S2 is a deep learning neural network ChessColor-Net, which uses an enhanced feature extraction structure and a convolution module to perform feature extraction;
[0027] The enhanced feature extraction structure first feeds the feature map into a depth-wise separable convolution for feature extraction, then performs a normalization operation, and then uses the ReLU activation function to improve the nonlinear ability of the network, and then performs a 1×1 (convolution kernel size) convolution to adjust the number of channels, and finally performs a normalization and ReLU activation function;
[0028] The convolution module consists of a convolution layer plus a normalization and ReLU activation function to extract features and adjust the number of channels.
[0029] Furthermore, in step S2, the process of deep learning neural network ChessColor-Net classifying the image is as follows: first, the image is passed into a convolution module to output a feature map; then, the feature map is sent to 5 enhanced feature extraction structures, the step size of the first enhanced feature extraction structure is 1, and the number of channels is set to 64, so the output feature map is 112×112 pixels and the number of channels is 64; the step size of the second enhanced feature extraction structure is 2, the number of channels is 128, and the output feature map is 56×56 pixels and the number of channels is 128; the step size of the third enhanced feature extraction structure is also 2, the number of channels is 256, and the output The feature map is 28×28 pixels with 256 channels. The stride of the fourth enhanced feature extraction structure is still 2, the number of channels is 512, and the output is a feature map of 14×14 pixels with 512 channels. The stride of the fifth enhanced feature extraction structure is 2, the number of channels is 512, and the output is a feature map of 7×7 pixels with 512 channels. The feature map is then globally average pooled to obtain a feature map of 1×1 pixels with 512 channels. Finally, it passes through a convolution layer with a convolution kernel of 1×1 to adjust the number of channels to 15 for classification. Because there are 14 kinds of chess pieces in total, plus a category without chess pieces.
[0030] Furthermore, since red and black pieces must be accurately classified, the color of the pieces needs to be highlighted in the design of the loss function. The deep learning neural network ChessColor-Net in step S2 is provided with a category loss function for optimizing the judgment of the color of the pieces. The category loss function adopts a cross entropy loss function and adds a penalty item for judging black and red pieces.
[0031] Furthermore, the cross entropy loss function in the category loss function is shown in formula (7):
[0032]
[0033] Where, L cross is the value calculated by the cross entropy loss function, n is the number of samples, m is the number of categories, and y i is the label of category j to which sample i belongs, which is 0 or 1; f(x ij ) is the probability that sample i is predicted as class j; the penalty term for black and red chess judgment is shown in formula (8):
[0034]
[0035] Where, L color is the value calculated for the penalty item for judging black and red pieces. Similarly, n is the number of samples, k ranges from 1 to 3, representing whether it is red, black or blank, and h ik is the label of whether sample i belongs to red chess, black chess or blank chess; g(x ik) is the probability of the sample i being classified as red, black or blank, g(x ik ) is calculated as shown in formula (9):
[0036] g(x ik )=∑p(x i ) (9)
[0037] In the formula, p(x i ) is the predicted probability of sample i, when g(x ik ) When the desired position is red, the sum is the predicted probability of all red pieces. Similarly, the probability of black pieces and blanks can be obtained.
[0038] The total loss function value L is shown in formula (10):
[0039]
[0040] When the sample is red, h ik The label representing the red chess piece is 1, and the others are 0. ik ) Only the predicted probability of red chess is valid. If the network predicts black chess, then g(x ik ) is close to 0, logg(x ik ) will be a very large negative number, L color Will become very large, and so will the total loss L, so that the network can adjust parameters more effectively and is particularly sensitive to the color of the chess pieces in the image.
[0041] According to the results of neural network classification, the classification results of 90 chess pieces on the chessboard are converted into chess game information. The correspondence between chess pieces and letters adopts the correspondence in the general interface engine protocol of chess, namely, King: K, Red Soldier: A, Red Horse: N, Red Rook: R, Red Cannon: C, Bishop: B, Pawn: P, General: k, Black Soldier: a, Black Horse: n, Black Rook: r, Black Cannon: c, Bishop: b, Pawn: p, and Blank: 0.
[0042] Furthermore, the step S3 specifically includes: using the camera at a fixed position in step A1 to take a chessboard picture P1 before making a move, the chess player makes a move, and after the move is completed, using the camera at a fixed position in step A1 to take another chessboard picture P2.
[0043] Furthermore, the specific process of the difference detection method in step S4 is as follows:
[0044] B1: grayscale processing of chessboard images P1 and P2;
[0045] B2: Subtract the grayscale image of P1 from the grayscale image of P2 to obtain image P3, and perform binarization on P3;
[0046] B3: According to the image P3, the coordinates are enumerated and judged by the set judgment threshold. When the output result is two coordinates, the coordinates are used as the starting point and the end point of the move. If the output result is not two coordinates, the judgment threshold is iterated until the output result is two coordinates.
[0047] B4: Convert the starting point and end point coordinates into chess move information and update the chess game information.
[0048] Furthermore, the step B2 is specifically as follows: using the subtraction function subtract of OpenCV (computer vision and machine learning software library) to subtract the grayscale image of P1 from the grayscale image of P2 to obtain image P3, and binarizing P3 so that the pixel values of P3 are only 0 and 255. Since the shooting positions of P1 and P2 are fixed, only the pixel contents of two placement points of P1 and P2 have changed, and regardless of whether the piece is captured or not, after the subtraction and binarization processing, the pixel values of the same parts of P1 and P2 become 0, and the pixel values of the changed parts become 255.
[0049] Furthermore, the step B3 is specifically as follows:
[0050] C1: A threshold judgment rectangle is set for each of the 90 chess pieces according to the size of the chessboard. The center of the rectangle is the coordinate of the chess piece. The rectangle is slightly larger than the chess piece, and there is no overlap between adjacent rectangles.
[0051] C2: Use a threshold to measure the number of points with a pixel value of 255 in the rectangle. When the 255 pixel points of a certain drop point exceed this threshold, this drop point is regarded as the target drop point. Traverse 90 drop points and find all points that exceed the threshold.
[0052] C3: When the number of given placement points exceeds 2 or is less than 2, it is judged to be unreasonable and the threshold needs to be adjusted. Repeat step C2 using the adjusted threshold;
[0053] C4: The threshold adjustment algorithm is shown in formula (11):
[0054] T=T0+T inc *x inc -T dec *x dec (11)
[0055] Among them, T is the threshold used for judgment, and T0 is the set initial threshold, which is generally set to a smaller value. inc The number of threshold increases each time. Generally, it is set to 5 to increase the threshold quickly. dec The number of times the threshold is reduced each time is generally set to 2. When the increased threshold is too high, the threshold can be gradually reduced, and it can be iterated to any threshold when it is set to 5 and 2.inc is the number of times the threshold is increased, the initial value is 0, x dec is the number of times the threshold is reduced, the initial value is 0;
[0056] C5: The initial threshold is used for the first judgment. When the number of points given by the judgment is greater than 2, it means that the threshold is set too small and some noise points are also counted. Therefore, the threshold needs to be increased. inc Each time the threshold is increased by 1, the judgment is made again; when the number of drop points given is less than 2, it means that the threshold is too large and the correct target drop point is also discarded, so the threshold needs to be reduced, and x dec Add 1 each time to reduce the threshold and make the judgment again;
[0057] C6: The final threshold judgment will get two target drop points. According to the obtained chess game information, the chess piece information contained in the two target drop points is obtained. The chess piece information corresponding to the current chess player (for example, the drop point is red chess, and the current chess player is red) is the starting point, and the other drop point is the end point;
[0058] C7: Update the chess game information according to the starting point and the end point. The information of the starting point is changed to empty, and the information of the end point is changed to the chess piece at the starting point to complete the chess move detection.
[0059] The method of the present invention independently designs a lightweight backbone neural network ChessColor-Net, and trains it with a category loss function that adds a penalty term for red and black chess pieces. It only takes 2 seconds on an embedded device to classify and recognize a complete chess board, and the recognition accuracy reaches 97.8%. In the chess move detection part, an image difference detection with automatic iteration of thresholds is designed to quickly screen the placement points, and the detection effect is not affected by environmental factors.
[0060] Based on the above scheme, the innovative points of the present invention can be summarized as follows:
[0061] (1) A method for calibrating the pixel coordinates of the chessboard image placement point is provided. (2) A method for building a neural network for chess game detection is provided, as well as a loss function calculation formula that adds chess features. (3) A chess move detection method that uses image difference for detection and automatically adjusts the threshold is provided. (4) It is capable of high-speed and high-precision classification and recognition of chess games. (5) Highly robust chess move recognition is not affected by chess piece type, chessboard appearance, lighting and other conditions.
[0062] Beneficial effects: Compared with the prior art, the present invention realizes the detection and recognition of Chinese chess from the beginning to each move, reads the chess game information of the chessboard image and updates it in real time, achieves high-speed and high-precision classification and recognition, and has high robustness in chess move detection. It has the advantages of fast speed and high robustness, and is less affected by light, and can be extended to the detection of other chess types, solving the problem that the existing methods are difficult to achieve fast and accurate recognition of Chinese chess, and can perform small target, large quantity, and multi-type recognition, which is of great significance in human-computer interaction in Chinese chess games and popularization of artificial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a schematic diagram of the workflow of a multi-target small object recognition method based on category loss and difference detection provided by an embodiment of the present invention;
[0064] Figure 2 An image of chess board coordinate calibration provided by an embodiment of the present invention;
[0065] Figure 3 It is an overall block diagram of a deep learning neural network system based on chess loss provided by an embodiment of the present invention;
[0066] Figure 4 This is an effect image of the embodiment of the present invention applied to the classification and recognition result of the color image of a complete Chinese chess board;
[0067] Figure 5 The embodiment of the present invention is applied to images captured before and after a move in a Chinese chess game;
[0068] Figure 6 This is an effect image of the embodiment of the present invention applied to difference detection during a Chinese chess game. DETAILED DESCRIPTION
[0069] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0070] The present invention provides a multi-target small object recognition method based on category loss and difference detection, and its main recognition process is: after the initial chessboard is set up, a deep learning neural network based on chess type loss is used for classification and recognition to obtain initial chess game information, and then a chessboard picture is taken at a fixed position before and after each chess move, and the images before and after the chess move are sent to a difference detection algorithm to detect the starting point and end point of the chess move.
[0071] like Figure 1 As shown, the above method specifically includes the following steps:
[0072] S1: Collect the initial chessboard image, calibrate the chessboard image coordinates, and segment the chessboard image according to the coordinates;
[0073] S2: Preprocess the segmented chessboard image, classify the processed image using the built deep learning network model, and convert the classification result into chess game information;
[0074] S3: Collect chessboard image P1, and after the chess player makes a move, collect chessboard image P2 again;
[0075] S4: Based on the chessboard images P1 and P2, a difference detection method is used to judge the rationality of the starting point and the end point of the chess move. If reasonable, the coordinates of the starting point and the end point are converted into chess move information, and the chess game information is updated.
[0076] In this embodiment, step S1 specifically includes the following steps:
[0077] A1: Fix the chessboard position, set up a camera above the chessboard and fix the camera position to capture a color image of the chessboard and all chess pieces;
[0078] A2: Calibrate the pixel coordinates of the points A, B, and C at the lower left corner, upper left corner, and lower right corner of the chessboard. Based on these three coordinates, calculate the pixel coordinates of the remaining 87 points on the chessboard.
[0079] The calculation method of the pixel coordinates of the drop point in this embodiment is:
[0080] Calculate the pixel length and width of each chessboard point, as shown in formulas (1), (2), (3), and (4):
[0081]
[0082]
[0083]
[0084]
[0085] In the formula, x A 、x B 、x C is the horizontal coordinate of points A, B, and C, y A ,y B ,y C is the ordinate of points A, B, and C, dW x , dW y dL is the difference between the horizontal and vertical coordinates of each point on the width of the chessboard. x ,dL yis the difference between the horizontal and vertical coordinates of each point on the length of the chessboard. The 9 and 10 in the denominator are because the chessboard is 9 squares long and 10 squares wide.
[0086] Taking point A as the reference point, the coordinates of other drop points are calculated using formulas (5) and (6):
[0087] S x =x A +i L ×D x -i W ×D x (5)
[0088] S y =y A +i L ×D y -i W ×D y (6)
[0089] In the formula, S x and S y is the horizontal and vertical coordinates of any drop point, i L is the column where the currently calculated drop point is located, with a value range from 0 to 8, i W is the row where the currently calculated drop point is located, and the value range is from 0 to 9; i L Each time a value is taken, i W The values are calculated from 0 to 9, so that the pixel coordinates of all 90 drop points can be calculated, and rectangles are drawn on the calculated coordinates for visualization, as shown below: Figure 2 As shown;
[0090] A3: The pixel coordinates are the center of the square, dW y is the side length of the square, and images of all 90 placement points are captured.
[0091] Step S2 of this embodiment specifically includes:
[0092] 1. Perform grayscale processing on the obtained placement point image and unify the image size to 224×224 pixels;
[0093] 2. The deep learning network model is a deep learning neural network ChessColor-Net, which uses a self-designed enhanced feature extraction structure and convolution module for feature extraction;
[0094] The enhanced feature extraction structure first feeds the feature map into a depth-wise separable convolution for feature extraction, then performs a normalization operation, and then uses the ReLU activation function to improve the nonlinear ability of the network, and then performs a 1×1 (convolution kernel size) convolution to adjust the number of channels, and finally performs a normalization and ReLU activation function;
[0095] The convolution module consists of a convolution layer plus a normalization and ReLU activation function to extract features and adjust the number of channels.
[0096] like Figure 3 As shown in the figure, the overall framework of the deep learning neural network ChessColor-Net is given. The figure shows the feature map size of each convolution module and its corresponding output. Figure 3 , the process of deep learning neural network ChessColor-Net to classify images is as follows: first, the 224×224 pixel image is passed into a convolution module with a convolution kernel size of 3×3 pixels and a step size of 2 for feature extraction, and a 112×112 pixel feature map with 32 channels is output; then the feature map is sent to 5 enhanced feature extraction structures, the step size of the first enhanced feature extraction structure is 1, and the number of channels is set to 64, so the output feature map is 112×112 pixels and the number of channels is 64; the step size of the second enhanced feature extraction structure is 2, the number of channels is 128, and the output feature map is 56×56 pixels with 128 channels; the third enhanced feature extraction structure is The step size of the extraction structure is also 2, the number of channels is 256, and the output is a feature map of 28×28 pixels and 256 channels; the step size of the fourth enhanced feature extraction structure is still 2, the number of channels is 512, and the output is a feature map of 14×14 pixels and 512 channels; the step size of the fifth enhanced feature extraction structure is 2, the number of channels is 512, and the output is a feature map of 7×7 pixels and 512 channels; the obtained feature map is then globally average pooled to obtain a feature map of 1×1 pixels and 512 channels, and finally passed through a convolution layer with a convolution kernel of 1×1 to adjust the number of channels to 15 for classification, because there are a total of 14 kinds of chess pieces, and an additional classification of no chess pieces is added.
[0097] Since red and black pieces must be accurately classified, the color of the pieces needs to be highlighted in the design of the loss function. A category loss function is set in the deep learning neural network ChessColor-Net to optimize the judgment of the color of the pieces. The category loss function uses the cross entropy loss function and adds a penalty item for judging black and red pieces.
[0098] The cross entropy loss function in the category loss function is shown in formula (7):
[0099]
[0100] Where, L cross is the value calculated by the cross entropy loss function, n is the number of samples, m is the number of categories, and y i is the label of category j to which sample i belongs, which is 0 or 1; f(x ij ) is the probability that sample i is predicted as class j; the penalty term for black and red chess judgment is shown in formula (8):
[0101]
[0102] Where, L color is the value calculated for the penalty item for judging black and red pieces. Similarly, n is the number of samples, k ranges from 1 to 3, representing whether it is red, black or blank, and h ik is the label of whether sample i belongs to red chess, black chess or blank chess; g(x ik ) is the probability of the sample i being classified as red, black or blank, g(x ik ) is calculated as shown in formula (9):
[0103] g(x ik )=∑p(x i ) (9)
[0104] In the formula, p(x i ) is the predicted probability of sample i, when g(x ik ) When the desired position is red, the sum is the predicted probability of all red pieces. Similarly, the probability of black pieces and blanks can be obtained.
[0105] The total loss function value L is shown in formula (10):
[0106]
[0107] When the sample is red, h ik The label representing the red chess piece is 1, and the others are 0. ik ) Only the predicted probability of red chess is valid. If the network predicts black chess, then g(x ik ) is close to 0, logg(x ik ) will be a very large negative number, L color Will become very large, and so will the total loss L, so that the network can adjust parameters more effectively and is particularly sensitive to the color of the chess pieces in the image.
[0108] 3. According to the classification result of the neural network, the classification result of the 90 chessboard points is converted into chess game information. The correspondence between chess pieces and letters adopts the correspondence in the chess general interface engine protocol, namely, General: K, Red Soldier: A, Red Horse: N, Red Rook: R, Red Cannon: C, Bishop: B, Soldier: P, General: k, Black Soldier: a, Black Horse: n, Black Rook: r, Black Cannon: c, Bishop: b, Pawn: p, and Blank: 0. The recognition result of the complete chessboard in this embodiment is as follows: Figure 4 shown.
[0109] Step S3 of this embodiment specifically includes: using the camera at the fixed position of step A1 to take a chessboard picture P1 before the chess player makes a move, and after the chess player makes a move, using the camera at the fixed position of step A1 to take another chessboard picture P2. The pictures P1 and P2 are specifically as follows: Figure 5 shown.
[0110] The specific process of the difference detection method in step S4 of this embodiment is as follows:
[0111] B1: grayscale processing is performed on the chessboard images P1 and P2 to convert the three-channel RGB images into grayscale images with pixel values between 0 and 255;
[0112] B2: Use the subtract function of OpenCV (computer vision and machine learning software library) to subtract the grayscale image of P1 from the grayscale image of P2 to obtain image P3. Binarize P3 so that the pixel values of P3 are only 0 and 255. Since the shooting positions of P1 and P2 are fixed, only the pixel contents of two points of P1 and P2 have changed, and no matter whether the pieces are captured or not, after subtracting and binarizing, the pixel values of the same parts of P1 and P2 become 0, and the pixel values of the changed parts become 255. Although the shooting positions of the two pictures are fixed, they cannot be exactly the same. There are always some noise interferences in some places of P3 after processing, so it is necessary to set the judgment threshold.
[0113] B3: According to the image P3, the coordinates are enumerated and judged by the set judgment threshold. When the output result is two coordinates, the coordinates are used as the starting point and the end point of the move. If the output result is not two coordinates, the judgment threshold is iterated until the output result is two coordinates.
[0114] B4: Convert the starting point and end point coordinates into chess move information and update the chess game information.
[0115] Step B3 of this embodiment is specifically as follows:
[0116] C1: According to the size of the chessboard, a threshold judgment rectangle is set for the 90 placement points. The center of the rectangle is the coordinate of the placement point. The size of the rectangle is slightly larger than the chess piece, and the maximum size does not exceed dW y, and there is no overlap between adjacent rectangles;
[0117] C2: Use a threshold to measure the number of points with a pixel value of 255 in the rectangle. When the 255 pixel points of a certain drop point exceed this threshold, this drop point is regarded as the target drop point. Traverse 90 drop points and find all points that exceed the threshold.
[0118] C3: When the number of given placement points exceeds 2 or is less than 2, it is judged to be unreasonable and the threshold needs to be adjusted. Repeat step C2 using the adjusted threshold;
[0119] C4: The threshold adjustment algorithm is shown in formula (11):
[0120] T=T0+T inc *x inc -T dec *x dec (11)
[0121] Among them, T is the threshold used for judgment, and T0 is the set initial threshold, which is generally set to a smaller value. inc The number of times the threshold value is increased each time is set to 5 in this embodiment, which can quickly increase the threshold value. dec The number of times the threshold is reduced each time is set to 2 in this embodiment. When the increased threshold is too high, the threshold can be gradually reduced, and it can be set to 5 and 2 to iterate to any threshold. inc is the number of times the threshold is increased, the initial value is 0, x dec The number of times the threshold is reduced, the initial value is 0, by setting x inc and x dec The value of can make T any integer;
[0122] C5: The initial threshold is used for the first judgment. When the number of points given by the judgment is greater than 2, it means that the threshold is set too small and some noise points are also counted. Therefore, the threshold needs to be increased. inc Each time the threshold is increased by 1, the judgment is made again; when the number of drop points given is less than 2, it means that the threshold is too large and the correct target drop point is also discarded, so the threshold needs to be reduced, and x dec Add 1 each time to reduce the threshold and make the judgment again;
[0123] C6: The final threshold judgment will get two target drop points. The detection effect is as follows Figure 6 As shown, according to the obtained chess game information, the chess piece information contained in the two target placement points is obtained, and the chess piece information corresponding to the current chess player (for example, the placement point is a red chess piece, and the current chess player is the red player) is the starting point, and the other placement point is the end point;
[0124] C7: Update the chess game information according to the starting point and the end point. The information of the starting point is changed to empty, and the information of the end point is changed to the chess piece at the starting point to complete the chess move detection.
[0125] In order to verify the effect of the above solution, this embodiment applies the above solution to the classification and recognition of Chinese chess games, and the recognition accuracy rate can reach 97.8% on the verification set.
Claims
1. A multi-target small object recognition method based on category loss and difference detection, characterized in that: The steps include: S1: Collect the initial chessboard image, calibrate the chessboard image coordinates, and segment the chessboard image according to the coordinates; S2: Preprocess the segmented chessboard image, classify the processed image using the built deep learning network model, and convert the classification result into chess game information; S3: Collect chessboard image P1, and after the chess player makes a move, collect chessboard image P2 again; S4: Based on the chessboard images P1 and P2, a difference detection method is used to judge the rationality of the starting point and the end point of the chess move. If they are reasonable, the coordinates of the starting point and the end point are converted into chess move information, and the chess game information is updated; The specific process of the difference detection method in step S4 is as follows: B1: grayscale processing of chessboard images P1 and P2; B2: Subtract the grayscale image of P1 from the grayscale image of P2 to obtain image P3, and perform binarization on P3; B3: According to the image P3, the coordinates are enumerated and judged by the set judgment threshold. When the output result is two coordinates, the coordinates are used as the starting point and the end point of the move. If the output result is not two coordinates, the judgment threshold is iterated until the output result is two coordinates. B4: Convert the starting point and end point coordinates into chess move information and update the chess game information; Step B3 is specifically as follows: C1: Set a threshold judgment rectangle for 90 placement points according to the size of the chessboard, and the center of the rectangle is the coordinate of the placement point; C2: Use a threshold to measure the number of points with a pixel value of 255 in the rectangle. When the 255 pixel points of a certain drop point exceed this threshold, this drop point is regarded as the target drop point. Traverse 90 drop points and find all points that exceed the threshold. C3: When the number of given placement points exceeds 2 or is less than 2, it is judged to be unreasonable and the threshold needs to be adjusted. Repeat step C2 using the adjusted threshold; C4: The threshold adjustment algorithm is shown in formula (11): T=T0+T inc *x inc -T dec *x dec (11) Among them, T is the threshold used for judgment, T0 is the initial threshold set, T inc is the number of thresholds added each time, T dec is the number of thresholds reduced each time, x inc The number of times the threshold is increased, the initial value is 0, x dec is the number of times the threshold is reduced, the initial value is 0; C5: The initial threshold is used for the first judgment. When the number of points given by the judgment is greater than 2, it means that the threshold is set too small and some noise points are also counted. Therefore, the threshold needs to be increased. inc Each time the threshold is increased by 1, the judgment is made again; when the number of drop points given is less than 2, it means that the threshold is too large and the correct target drop point is also discarded, so the threshold needs to be reduced, and x dec Add 1 each time to reduce the threshold and make the judgment again; C6: The final threshold judgment will get two target drop points. According to the obtained chess game information, the chess piece information contained in the two target drop points is obtained. The chess piece information corresponding to the current chess player is the starting point, and the other drop point is the end point. C7: Update the chess game information according to the starting point and the end point. The information of the starting point is changed to empty, and the information of the end point is changed to the chess piece at the starting point to complete the chess move detection.
2. According to claim 1, a multi-target small object recognition method based on category loss and difference detection is characterized in that: The step S1 specifically includes the following steps: A1: Fix the chessboard position, set up a camera above the chessboard and fix the camera position to capture images of the chessboard and chess pieces; A2: Calibrate the pixel coordinates of the points A, B, and C at the lower left corner, upper left corner, and lower right corner of the chessboard, and calculate the pixel coordinates of the remaining points on the chessboard based on these three coordinates; A3: Capture images of all placement points.
3. The multi-target small object recognition method based on category loss and difference detection according to claim 2 is characterized in that: The method for calculating the pixel coordinates of the drop point in step A2 is: Calculate the pixel length and width of each chessboard point, as shown in formulas (1), (2), (3), and (4): In the formula, x A 、x B 、x C is the horizontal coordinate of points A, B, and C, y A ,y B ,y C is the ordinate of points A, B, and C, dW x , dW y dL is the difference between the horizontal and vertical coordinates of each point on the width of the chessboard. x ,dL y is the difference between the horizontal and vertical coordinates of each point on the length of the chessboard. The 9 and 10 in the denominator are because the chessboard is 9 squares long and 10 squares wide. Taking point A as the reference point, the coordinates of other drop points are calculated using formulas (5) and (6): S x =x A +i L ×dL x -i W ×dW x (5) S y =y A +i L ×dL y -i W ×dW y (6) In the formula, S x and S y is the horizontal and vertical coordinates of any drop point, i L is the column where the currently calculated drop point is located, with a value range from 0 to 8, i W is the row where the currently calculated drop point is located, and the value range is from 0 to 9; i L Each time a value is taken, i W The values are calculated from 0 to 9, so the pixel coordinates of all the drop points can be calculated.
4. The multi-target small object recognition method based on category loss and difference detection according to claim 1 is characterized in that: The deep learning network model in step S2 is a deep learning neural network ChessColor-Net, which uses an enhanced feature extraction structure and a convolution module to perform feature extraction; The enhanced feature extraction structure first sends the feature map to a depth-separable convolution for feature extraction, then performs a normalization operation, and then uses the ReLU activation function to improve the nonlinear ability of the network, and then performs a convolution to adjust the number of channels, and finally performs a normalization and ReLU activation function; The convolution module consists of a convolution layer plus a normalization and ReLU activation function to extract features and adjust the number of channels.
5. The multi-target small object recognition method based on category loss and difference detection according to claim 4 is characterized in that: The process of deep learning neural network ChessColor-Net classifying the image in step S2 is as follows: first, the image is passed into a convolution module to output a feature map; then, the feature map is sent to 5 enhanced feature extraction structures, the step size of the first enhanced feature extraction structure is 1, the number of channels is set to 64, so the output feature map is 112×112 pixels and the number of channels is 64; the step size of the second enhanced feature extraction structure is 2, the number of channels is 128, and the output feature map is 56×56 pixels and the number of channels is 128; the step size of the third enhanced feature extraction structure is also 2, the number of channels is 256, and the output is 28× The fourth enhanced feature extraction structure still has a step size of 2 and a channel number of 512, and outputs a feature map of 14×14 pixels and 512 channels; the fifth enhanced feature extraction structure has a step size of 2 and a channel number of 512, and outputs a feature map of 7×7 pixels and 512 channels; the obtained feature map is then globally average pooled to obtain a feature map of 1×1 pixels and 512 channels, and finally passes through a convolution layer with a convolution kernel of 1×1 to adjust the number of channels to 15 for classification, because there are a total of 14 kinds of chess pieces, plus a category without chess pieces.
6. The multi-target small object recognition method based on category loss and difference detection according to claim 4 is characterized in that: The deep learning neural network ChessColor-Net of step S2 is provided with a category loss function for optimizing the judgment of chess piece colors. The category loss function adopts a cross entropy loss function and adds a penalty item for judging black and red chess pieces.
7. The multi-target small object recognition method based on category loss and difference detection according to claim 6 is characterized in that: The cross entropy loss function in the category loss function is shown in formula (7): Where, L cross is the value calculated by the cross entropy loss function, n is the number of samples, m is the number of categories, and y i is the label of category j to which sample i belongs, which is 0 or 1; f(x ij ) is the probability that sample i is predicted as class j; the penalty term for black and red chess judgment is shown in formula (8): Where, L color is the value calculated for the penalty item for judging black and red pieces. Similarly, n is the number of samples, k ranges from 1 to 3, representing whether it is red, black or blank, and h ik is the label of whether sample i belongs to red chess, black chess or blank chess; g(x ik ) is the probability of the sample i being classified as red, black or blank, g(x ik ) is calculated as shown in formula (9): g(x ik )=∑p(x i ) (9) In the formula, p(x i ) is the predicted probability of sample i, when g(x ik ) When the desired position is red, the sum is the predicted probability of all red pieces. Similarly, the probability of black pieces and blanks can be obtained. The total loss function value L is shown in formula (10):
8. The multi-target small object recognition method based on category loss and difference detection according to claim 1 is characterized in that: The step B2 is specifically as follows: using OpenCV's subtract function to subtract the grayscale image of P1 from the grayscale image of P2 to obtain image P3, and binarizing P3 so that the pixel values of P3 are only 0 and 255. Since the shooting positions of P1 and P2 are fixed, only the pixel contents of two placement points of P1 and P2 have changed, and regardless of whether a capture occurs, after the subtraction and binarization processing, the pixel values of the same parts of P1 and P2 become 0, and the pixel values of the changed parts become 255.
Citation Information
Patent Citations
Weiqi referee system based on MLP neural network and computer vision
CN110399888A
Go map reliable identification method capable of overcoming light reflection phenomenon
CN113298767A