Target detection method and system for visual image

By acquiring visual image data and extracting specific target images in the target area image based on the region growth method, the problems of poor target detection and low detection efficiency in the prior art are solved, and more efficient and universal target detection is achieved.

CN120070869APending Publication Date: 2025-05-30HUANENG YICHUN THERMAL POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234840.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the target detection is not very versatile, poor universality and low detection efficiency.

Method used

By acquiring visual image data, the image complexity is determined, and the specific target image in the target area image is extracted based on the region growth method, and the key point position coordinates are obtained for target action prediction.

Benefits of technology

It improves the accuracy and efficiency of target detection and enhances the universality and universality of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070869A_ABST
    Figure CN120070869A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, and particularly discloses a target detection method and system for a visual image, and the method comprises the steps: obtaining the visual image data of a target image, and determining the image complexity according to the visual image data of the target image; determining a target region image according to the image complexity based on a region growing method, and extracting a specific target image in the target region image; and obtaining position coordinates of the key points in the specific target image, and performing target action prediction according to the position coordinates of the key points in the specific target image. The method comprehensively considers the influence of the image complexity of the visual image data on the target detection, carries out the target detection of the to-be-detected target, predicts the target motion, and improves the target detection capability of a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of target detection, and more specifically, to a method and system for target detection of visual images. Background Art

[0002] Visual images refer to visual information such as images and videos generated through visual perception and processing. These information can be natural scenery, people, or artificially manufactured images and videos, etc. Visual images are finally formed by the reflection of light on objects or the transmission of light through objects, followed by the perception of the eyes and the processing of the brain. How to efficiently utilize a large amount of multi-source visual image data for target detection is a major application requirement in the field of target detection. Moving target detection has extensive applications in fields such as autonomous driving, intelligent transportation, and video surveillance.

[0003] In the prior art, since the target detection of visual image data requires determining the moving background in advance and has poor anti-interference ability, its application range is narrow, the generality is not strong, the universality is poor, and the detection efficiency is too low. Summary of the Invention

[0004] The present invention provides a method and system for target detection of visual images to solve the problems of poor generality, poor universality, and too low detection efficiency in target detection in the prior art. The method includes:

[0005] Obtain the visual image data of the target image, and determine the image complexity according to the visual image data of the target image;

[0006] Based on the region growing method, determine the target region image according to the image complexity, and extract the specific target image in the target region image;

[0007] Obtain the position coordinates of the key points in the specific target image, and perform target action prediction according to the position coordinates of the key points in the specific target image.

[0008] Further, the determining the image complexity according to the visual image data of the target image includes:

[0009] Perform grayscale processing on the visual image data of the target image to obtain a visual image grayscale image;

[0010] Obtain a preset grid, and divide the visual image grayscale image according to the preset grid to obtain visual image grayscale image blocks;

[0011] Obtain a preset grayscale range, and determine the occupied area of each preset grayscale range in the visual image grayscale image block according to the preset grayscale range to obtain the image complexity of each visual image grayscale image block;

[0012] Determine the complexity weight of each visual image grayscale image block according to the distance between the visual image grayscale image block and the central image block located at the center of the visual image grayscale image, and perform weighted summation on the image complexity of all visual image grayscale image blocks according to the complexity weight to obtain the image complexity of the visual image grayscale image.

[0013] Further, the determining the target region image according to the image complexity based on the region growing method includes:

[0014] Obtain the grayscale values of each pixel point in the visual image grayscale image, perform clustering processing on each pixel point according to the grayscale values, and determine the clustering center of each pixel point according to the clustering result;

[0015] Obtain the grayscale value of the target pixel point, calculate the difference between the grayscale value of the clustering center and the grayscale value of the target pixel point, and determine the clustering region corresponding to the difference less than the first preset threshold as the initial target region;

[0016] Set the clustering center pixel point corresponding to the initial target region as the initial seed point, calculate the similarity between the initial seed point and each pixel point in its preset neighborhood in the visual image grayscale image according to the image complexity, and set the pixel point with the similarity less than the second preset threshold as the new seed point;

[0017] Continue to detect the remaining pixel points according to the new seed points until the region can no longer grow to obtain the target region image.

[0018] Further, the performing clustering processing on each pixel point according to the grayscale values includes:

[0019] Establish a pixel point grayscale data set according to the grayscale values of each pixel point in the visual image grayscale image, and randomly select k initial clustering centers of the pixel point grayscale data set;

[0020] Calculate the Euclidean distance from the grayscale value of the pixel point in the pixel point grayscale data set to the initial clustering center, and divide each pixel point into the corresponding clustering cluster according to the Euclidean distance from the grayscale value of the pixel point in the pixel point grayscale data set to the initial clustering center;

[0021] Calculate the average value of the grayscale values of the pixel points in each clustering cluster, and re-determine the clustering center according to the average value of the grayscale values of the pixel points in each clustering cluster;

[0022] Repeat the above steps iteratively until the clustering center no longer changes or the number of iterations reaches the preset iteration threshold to obtain the clustering result of each pixel point.

[0023] Further, the calculating the similarity between the initial seed point and each pixel point in its preset neighborhood in the visual image grayscale image according to the image complexity includes:

[0024] Calculate the similarity between the initial seed points in the grayscale image of the visual image and each pixel point within its preset neighborhood according to the similarity calculation formula. The specific similarity calculation formula is as follows:

[0025]

[0026] where S ij is the similarity between the i-th seed point and the j-th pixel point within its preset neighborhood, D ij is the Euclidean distance between the i-th seed point and the j-th pixel point within its preset neighborhood, P is the image complexity, R is the preset range parameter, G i is the grayscale value of the i-th seed point, G j is the grayscale value of the j-th pixel point, C max is the maximum grayscale value of the cluster to which the i-th seed point belongs, C min is the minimum grayscale value of the cluster to which the i-th seed point belongs.

[0027] Furthermore, the extraction of the specific target image from the target region image includes:

[0028] Obtain the target region image of the historical visual image and the corresponding specific target image, and preprocess the target region image of the historical visual image and the corresponding specific target image;

[0029] Establish a training sample set according to the preprocessed target region image of the historical visual image and the corresponding specific target image, and establish an initial target detection model according to the training sample set;

[0030] Train the initial target detection model according to the training sample set to obtain the target detection model;

[0031] Input the current target region image into the trained target detection model to obtain the specific target image corresponding to the current target region image.

[0032] Furthermore, the target detection model is specifically an artificial convolutional neural network model, including 1 input layer, 1 output layer, 2 hidden layers, and 3 convolutional layers. The relu is used as the activation function, and the cross-entropy is used as the loss function.

[0033] Furthermore, the prediction of the target action according to the position coordinates of the key points in the specific target image includes:

[0034] Obtain the positions of the key points of the specific target image, count the position change data of the key points within a preset period, and determine the target motion trajectory according to the position change data of the key points within the preset period;

[0035] Determine the acceleration data of each key point according to the target motion trajectory, and calculate the average acceleration of all key points according to the acceleration data of each key point.

[0036] Calculate the difference between the acceleration of each key point and the average acceleration, and calculate the standard deviation according to the difference between the acceleration of each key point and the average acceleration;

[0037] Calculate the ratio of the standard deviation to the preset allowable standard deviation, determine the target stability coefficient according to the ratio of the standard deviation to the preset allowable standard deviation, and perform target action prediction according to the target stability coefficient.

[0038] Further, the performing target action prediction according to the target stability coefficient includes:

[0039] Obtain a preset target action library, and calculate the cosine similarity between the current target motion trajectory and the action trajectories in the preset target action library;

[0040] Correct the cosine similarity according to the target stability coefficient to obtain the corrected cosine similarity between the current target motion trajectory and the action trajectories in the preset target action library;

[0041] Screen out the action trajectories in the preset target action library whose corrected cosine similarity is greater than the third preset threshold to obtain the predicted target action.

[0042] To achieve the above object, the present invention also provides a target detection system for visual images, including:

[0043] A first module, configured to obtain visual image data of a target image and determine the image complexity according to the visual image data of the target image;

[0044] A second module, configured to determine a target region image based on the region growing method according to the image complexity, and extract a specific target image in the target region image;

[0045] A third module, configured to obtain the position coordinates of key points in the specific target image, and perform target action prediction according to the position coordinates of the key points in the specific target image.

[0046] The beneficial effects of the present invention are as follows:

[0047] By applying the above technical solutions, the present invention comprehensively considers the image complexity of visual image data, so as to be able to extract accurate target images, effectively improve the accuracy of target detection, and at the same time be able to perform motion prediction on the target images, greatly increasing the versatility and universality of target detection and improving the detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0049] Figure 1 It shows the overall flowchart of a target detection method for visual images proposed in an embodiment of the present invention;

[0050] Figure 2 It shows the structural schematic diagram of a target detection system for visual images proposed in an embodiment of the present invention. Detailed implementation manners

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0052] The embodiments of the present application provide a target detection method for visual images, as Figure 1 shown, the method includes:

[0053] S101, obtaining visual image data of a target image, and determining the image complexity according to the visual image data of the target image;

[0054] In some embodiments of the present application, the determining the image complexity according to the visual image data of the target image includes: performing grayscale processing on the visual image data of the target image to obtain a visual image grayscale image; obtaining a preset grid, dividing the visual image grayscale image according to the preset grid to obtain visual image grayscale image blocks; obtaining a preset grayscale range, determining the occupied area of each preset grayscale range in the visual image grayscale image blocks to obtain the image complexity of each visual image grayscale image block; determining the complexity weight of each visual image grayscale image block according to the distance between each visual image grayscale image block and the central image block located at the center of the visual image grayscale image, and performing weighted summation on the image complexity of all visual image grayscale image blocks according to the complexity weight to obtain the image complexity of the visual image grayscale image.

[0055] In this embodiment, since the target to be detected is usually placed at the center of the image when performing target detection on visual image data, the complexity weight of each visual image gray-scale image block is determined by the distance between the visual image gray-scale image block and the central image block located at the center of the visual image gray-scale image. The value range of the complexity weight is [0, 2]. The farther the distance from the central image block, the greater the corresponding complexity weight, thereby enhancing the detection of the complexity of the image background.

[0056] S102. Determine the target region image according to the image complexity based on the region growing method, and extract the specific target image in the target region image.

[0057] In some embodiments of the present application, the determining the target region image according to the image complexity based on the region growing method includes: obtaining the gray-scale values of each pixel point in the visual image gray-scale image, performing clustering processing on each pixel point according to the gray-scale values, and determining the clustering center of each pixel point according to the clustering result; obtaining the gray-scale value of the target pixel point, calculating the difference between the gray-scale value of the clustering center and the gray-scale value of the target pixel point, and determining the clustering region corresponding to the difference less than the first preset threshold as the initial target region; setting the clustering center pixel point corresponding to the initial target region as the initial seed point, calculating the similarity between the initial seed point and each pixel point in its preset neighborhood in the visual image gray-scale image according to the image complexity, and setting the pixel point with the similarity less than the second preset threshold as the new seed point; continuing to detect the remaining pixel points according to the new seed point until the region can no longer grow, and obtaining the target region image.

[0058] In this embodiment, since the target to be detected is usually placed at the center of the image, the pixel point at the very center of the visual image gray-scale image is set as the target pixel point. Several initial target regions with gray-scale values closest to the target are screened out through the target pixel point, and the clustering center pixel point of the initial target region is set as the initial seed point for region growing, so as to grow the target region image. The clustering center pixel point is specifically the pixel point with the gray-scale value closest to the clustering center value.

[0059] In some embodiments of the present application, the performing clustering processing on each pixel point according to the gray-scale values includes: establishing a pixel point gray-scale data set according to the gray-scale values of each pixel point in the visual image gray-scale image, and randomly selecting k initial clustering centers of the pixel point gray-scale data set; calculating the Euclidean distance from the gray-scale value of the pixel point in the pixel point gray-scale data set to the initial clustering center, and dividing each pixel point into the corresponding clustering cluster according to the Euclidean distance from the gray-scale value of the pixel point in the pixel point gray-scale data set to the initial clustering center; calculating the average value of the gray-scale values of the pixel points in each clustering cluster, and re-determining the clustering center according to the average value of the gray-scale values of the pixel points in each clustering cluster; repeating the above steps iteratively until the clustering center no longer changes or the number of iterations reaches the preset iteration threshold, and obtaining the clustering result of each pixel point.

[0060] In this embodiment, based on the k-means clustering algorithm, each pixel point is clustered according to the gray value. The value of k is determined by the number of pixel points. The larger the number of pixel points, the larger the corresponding value of k selected.

[0061] In some embodiments of the present application, calculating the similarity between the initial seed point in the visual image gray image and each pixel point in its preset domain according to the image complexity includes: calculating the similarity between the initial seed point in the visual image gray image and each pixel point in its preset domain according to the similarity calculation formula. The specific similarity calculation formula is

[0062]

[0063] where S ij is the similarity between the i-th seed point and the j-th pixel point in its preset domain, D ij is the Euclidean distance between the i-th seed point and the j-th pixel point in its preset domain, P is the image complexity, R is the preset range parameter, G i is the gray value of the i-th seed point, G j is the gray value of the j-th pixel point, C max is the maximum gray value of the clustering cluster to which the i-th seed point belongs, C min is the minimum gray value of the clustering cluster to which the i-th seed point belongs.

[0064] In some embodiments of the present application, extracting the specific target image in the target area image includes: obtaining the target area image of the historical visual image and the corresponding specific target image, and preprocessing the target area image of the historical visual image and the corresponding specific target image; establishing a training sample set according to the preprocessed target area image of the historical visual image and the corresponding specific target image, and establishing an initial target detection model according to the training sample set; training the initial target detection model according to the training sample set to obtain a target detection model; inputting the current target area image into the trained target detection model to obtain the specific target image corresponding to the current target area image.

[0065] In some embodiments of the present application, the target detection model is specifically an artificial convolutional neural network model, including 1 input layer, 1 output layer, 2 hidden layers, and 3 convolutional layers. The relu is used as the activation function, and the cross entropy is used as the loss function.

[0066] In this embodiment, the specific target in the target area image is detected by establishing a target detection model, and further a specific target image is obtained.

[0067] S103. Obtain the position coordinates of the key points in the specific target image, and predict the target action according to the position coordinates of the key points in the specific target image.

[0068] In some embodiments of the present application, the target action prediction based on the position coordinates of key points in a specific target image includes: obtaining the positions of key points in the specific target image, counting the position change data of the key points within a preset period, and determining the target motion trajectory according to the position change data of the key points within the preset period; determining the acceleration data of each key point according to the target motion trajectory, and calculating the average acceleration value of all key points according to the acceleration data of each key point; calculating the difference between the acceleration of each key point and the average acceleration value, and calculating the standard deviation according to the difference between the acceleration of each key point and the average acceleration value; calculating the ratio of the standard deviation to a preset allowable standard deviation, determining the target stability coefficient according to the ratio of the standard deviation to the preset allowable standard deviation, and performing target action prediction according to the target stability coefficient.

[0069] In this embodiment, the target stability coefficient is calculated by analyzing the motion trajectory of the key points of the target image to determine the stability of the target, so as to correct the predicted target action through the target stability coefficient, and achieve accurate motion prediction of the target image.

[0070] In some embodiments of the present application, the target action prediction based on the target stability coefficient includes: obtaining a preset target action library, and calculating the cosine similarity between the current target motion trajectory and the action trajectories in the preset target action library; correcting the cosine similarity according to the target stability coefficient to obtain the corrected cosine similarity between the current target motion trajectory and the action trajectories in the preset target action library; screening out the action trajectories in the preset target action library with the corrected cosine similarity greater than a third preset threshold to obtain the predicted target action.

[0071] In this embodiment, the motion prediction of the target is performed by the similarity between the current target motion trajectory and the target action trajectories in the preset target action library, and the similarity is corrected based on the target stability coefficient, so as to obtain the predicted target action.

[0072] Based on the same technical concept, as Figure 2 shown, the present invention also provides a target detection system for visual images, including:

[0073] A first module, configured to obtain visual image data of a target image and determine the image complexity according to the visual image data of the target image; a second module, configured to determine a target region image based on the region growing method according to the image complexity, and extract a specific target image from the target region image; a third module, configured to obtain the position coordinates of key points in the specific target image and perform target action prediction according to the position coordinates of the key points in the specific target image.

[0074] By applying the above technical solutions, the present invention obtains the visual image data of the target image, determines the image complexity according to the visual image data of the target image; determines the target region image based on the region growing method according to the image complexity, and extracts the specific target image in the target region image; obtains the position coordinates of the key points in the specific target image, and predicts the target action according to the position coordinates of the key points in the specific target image. The present invention comprehensively considers the influence of the image complexity of the visual image data on target detection, performs target detection on the target to be detected and predicts the target action at the same time, and improves the target detection ability for complex scenes.

[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present invention.

[0076] Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed to be located in one or more devices different from the present implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.

[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A method for detecting an object in a visual image, characterized in that: The method comprises: Acquire visual image data of a target image, and determine image complexity according to the visual image data of the target image; Based on the region growing method, the target region image is determined according to the image complexity, and the specific target image in the target region image is extracted; The position coordinates of the key points in the specific target image are obtained, and the target action is predicted according to the position coordinates of the key points in the specific target image.

2. The method for object detection using visual images according to claim 1, characterized in that: The determining of the image complexity according to the visual image data of the target image comprises: Grayscale processing is performed on the visual image data of the target image to obtain a visual image grayscale image; Obtain a preset grid, and divide the visual image grayscale image into blocks according to the preset grid to obtain visual image grayscale image blocks; Obtaining a preset grayscale range, determining the area occupied by each preset grayscale range in the visual image grayscale image block according to the preset grayscale range, and obtaining the image complexity of each visual image grayscale image block; The complexity weight of each visual image grayscale image block is determined according to the distance between each visual image grayscale image block and the central image block located at the center of the visual image grayscale image. The image complexity of all visual image grayscale image blocks is weightedly summed according to the complexity weight to obtain the image complexity of the visual image grayscale image.

3. The method for object detection using visual images according to claim 2, characterized in that: The method of determining the target region image according to the image complexity based on the region growing method includes: Obtain the grayscale value of each pixel in the grayscale image of the visual image, perform clustering processing on each pixel according to the grayscale value, and determine the cluster center of each pixel according to the clustering result; Obtaining the grayscale value of the target pixel, calculating the difference between the grayscale value of the cluster center and the grayscale value of the target pixel, and determining the cluster area corresponding to the difference being less than a first preset threshold as the initial target area; The cluster center pixel point corresponding to the initial target area is set as the initial seed point, and the similarity between the initial seed point and each pixel point in the preset area in the visual image grayscale image is calculated according to the image complexity, and the pixel point with a similarity less than a second preset threshold is set as a new seed point; Continue to detect the remaining pixels based on the new seed point until the area can no longer grow and the target area image is obtained.

4. The method for object detection using visual images according to claim 3, characterized in that: The clustering of each pixel point according to the gray value includes: A pixel grayscale data set is established according to the grayscale value of each pixel in the visual image grayscale image, and k initial clustering centers of the pixel grayscale data set are randomly selected; Calculate the Euclidean distance from the grayscale value of the pixel point in the pixel grayscale data set to the initial cluster center, and divide each pixel point into a corresponding cluster cluster according to the Euclidean distance from the grayscale value of the pixel point in the pixel grayscale data set to the initial cluster center; Calculate the average grayscale value of each pixel in each cluster, and re-determine the cluster center according to the average grayscale value of each pixel in each cluster; Repeat the above steps until the cluster center no longer changes or the number of iterations reaches a preset iteration threshold, and obtain the clustering results of each pixel.

5. The method for object detection using visual images according to claim 3, characterized in that: The method of calculating the similarity between the initial seed point in the visual image grayscale image and each pixel point in its preset area according to the image complexity includes: The similarity between the initial seed point in the visual image grayscale image and each pixel point in the preset area is calculated according to the similarity calculation formula. The similarity calculation formula is specifically: Among them, S ij is the similarity between the i-th seed point and the j-th pixel point in its preset area, D ij is the Euclidean distance between the i-th seed point and the j-th pixel point in the preset area, P is the image complexity, R is the preset range parameter, G i is the gray value of the i-th seed point, G j is the gray value of the j-th pixel, C max is the maximum gray value of the cluster to which the i-th seed point belongs, C min is the minimum grayscale value of the cluster to which the i seed point belongs.

6. The method for object detection of visual images according to claim 1, characterized in that: The step of extracting a specific target image from the target area image comprises: Acquire the target area image and the corresponding specific target image of the historical visual image, and pre-process the target area image and the corresponding specific target image of the historical visual image; Establish a training sample set based on the target area image of the preprocessed historical visual image and the corresponding specific target image, and establish an initial target detection model based on the training sample set; The initial target detection model is trained according to the training sample set to obtain a target detection model; The current target area image is input into the trained target detection model to obtain the specific target image corresponding to the current target area image.

7. The method for object detection using visual images according to claim 6, characterized in that: The target detection model is specifically an artificial convolutional neural network model, including 1 input layer, 1 output layer, 2 hidden layers, and 3 convolutional layers. Relu is used as an activation function and cross entropy is used as a loss function.

8. The method for object detection using visual images according to claim 1, characterized in that: The target action prediction according to the position coordinates of the key points in the specific target image includes: Obtain the key point positions of a specific target image, count the position change data of the key points within a preset period, and determine the target motion trajectory based on the position change data of the key points within the preset period; Determine the acceleration data of each key point according to the target motion trajectory, and calculate the average acceleration of all key points according to the acceleration data of each key point; Calculate the difference between the acceleration of each key point and the average acceleration, and calculate the standard deviation based on the difference between the acceleration of each key point and the average acceleration; The ratio of the standard deviation to the preset allowable standard deviation is calculated, the target stability coefficient is determined according to the ratio of the standard deviation to the preset allowable standard deviation, and the target action is predicted according to the target stability coefficient.

9. The method for object detection using visual images according to claim 8, characterized in that: The target action prediction according to the target stability coefficient includes: Obtain a preset target action library, and calculate the cosine similarity between the current target motion trajectory and the action trajectory in the preset target action library; The cosine similarity is corrected according to the target stability coefficient to obtain the cosine similarity between the corrected current target motion trajectory and the motion trajectory in the preset target motion library; The action trajectories in the preset target action library whose corrected cosine similarity is greater than the third preset threshold are screened out to obtain the predicted target action.

10. A visual image target detection system, characterized in that: include: The first module is used to obtain visual image data of the target image and determine the image complexity according to the visual image data of the target image; The second module is used to determine the target region image according to the image complexity based on the region growing method, and extract the specific target image in the target region image; The third module is used to obtain the position coordinates of the key points in the specific target image and predict the target action according to the position coordinates of the key points in the specific target image.