Scene character recognition and classification method and device based on multi-collaborative neural network

By using monocular cameras and multi-collaborative neural network technology in the substation operation and maintenance site, the problems of low recognition accuracy and high computational complexity of monocular cameras in complex environments are solved, and fast and accurate character recognition classification is achieved, and digital management and control of substation equipment is supported.

CN120220011APending Publication Date: 2025-06-27STATE GRID SHANDONG ELECTRIC POWER CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145565.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

At the substation operation and maintenance site, due to the lack of depth data, monocular cameras are difficult to accurately extract human features under complex backgrounds and lighting changes, resulting in low identification and classification accuracy and high calculation complexity, which makes it difficult to meet the real-time processing needs.

Method used

A scene character recognition classification method based on multi-collaborative neural network is adopted, and data augmentation is collected through a monocular camera, and residual neural network is designed based on tensorflow for training to generate a model for real-time identification classification.

Benefits of technology

The speed and accuracy of character identification and classification on site of substation operation and maintenance has been improved, digital management and control of substation equipment has been realized, the misjudgment rate has been reduced, and the real-time processing needs have been met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220011A_ABST
    Figure CN120220011A_ABST
Patent Text Reader

Abstract

The invention discloses a scene character recognition and classification method and device based on a multi-collaborative neural network, and belongs to the technical field of power transformation operation and maintenance and image processing, and the method comprises the steps: S1, employing the image data of a person at a power transformation operation and maintenance site through monocular camera equipment; s2, performing data enhancement processing on the acquired personnel image data to obtain processed personnel image data; s3, designing a residual neural network based on tensorflow, and performing network training by using the processed person image data to obtain a trained person recognition classification model; and S4, performing data enhancement processing by adopting personnel image data of a power transformation operation and maintenance site in real time, and inputting the trained character recognition and classification model to perform character recognition and classification. According to the invention, the personnel appearing in the power transformation operation and maintenance site are identified and classified through the personnel image data in the power transformation operation and maintenance site, and digital management and control of the power transformation equipment are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and device for scene person recognition and classification based on a multi - collaborative neural network, belonging to the technical fields of substation operation and maintenance inspection and image processing. Background Art

[0002] In the traditional substation operation and maintenance inspection work mode, personnel management mainly relies on manual records and on - site inspections, and equipment control is mostly based on experience judgment and simple instrument detection. This method is not only inefficient, but also prone to omissions when facing a large number of substation equipment and complex maintenance tasks, and its accuracy is difficult to guarantee, with misjudgments and missed judgments occurring frequently. With the rapid development of digital technology, the substation operation and maintenance field urgently needs to achieve digital transformation to improve management efficiency and maintenance quality. Person recognition and classification technology, as the core link to realize digital control of on - site operations in substation operation and maintenance, is of great importance.

[0003] With the rapid development of computer vision technology, human body recognition and classification technology based on image processing has become a research hotspot in many fields. Human body recognition and classification refers to determining the identity or behavior information of an individual by analyzing the appearance characteristics of the human body in an image, and it is widely used in scenarios such as security monitoring, behavior analysis, and driverless driving. However, existing person recognition methods face many challenges in complex environments such as substation operation and maintenance. Compared with traditional binocular or multi - binocular camera systems, monocular cameras have attracted attention due to their advantages such as low hardware cost and convenient deployment. However, since monocular cameras can only capture two - dimensional image information and lack depth data support, they face great technical challenges in realizing human body recognition and classification. First, since monocular cameras can only obtain planar images, it is more difficult to extract the boundaries of the target human body and build feature models in scenarios with complex backgrounds or changing lighting. Second, in the case of occlusion, pose changes, and viewing angle differences, monocular images are prone to cause blurring or loss of target features, thus affecting the recognition and classification accuracy. Finally, traditional monocular recognition and classification methods have bottlenecks in terms of computational complexity and real - time performance, and it is difficult to meet the requirements of efficient processing.

[0004] In recent years, the rise of deep learning technology has provided a new solution for human body recognition and classification of monocular cameras. By constructing a deep learning model based on a convolutional neural network (CNN), multi - level and multi - scale human body features can be extracted from monocular images, thus significantly improving the recognition and classification performance. In addition, by combining human pose estimation and target detection algorithms, the human body recognition and classification ability of monocular cameras in complex scenarios is further enhanced. Therefore, the present invention provides a method and device for scene person recognition and classification based on a multi - collaborative neural network, aiming to efficiently and accurately recognize and classify human bodies using monocular cameras, overcome the deficiencies of traditional technologies, and provide a low - cost and high - performance solution for the field of image processing. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a method and device for scene person recognition and classification based on a multi-collaborative neural network, which can improve the speed and accuracy of person recognition and classification in the complex environment of substation operation and maintenance, thereby assisting substation operation and maintenance inspection work.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0007] In a first aspect, a method for scene person recognition and classification based on a multi-collaborative neural network provided by an embodiment of the present invention includes the following steps:

[0008] Step S1, using a monocular camera device to collect personnel image data at the substation operation and maintenance site;

[0009] Step S2, performing data augmentation processing on the collected personnel image data to obtain processed personnel image data;

[0010] Step S3, designing a residual neural network based on tensorflow and using the processed personnel image data for network training to obtain a trained person recognition and classification model;

[0011] Step S4, collecting personnel image data at the substation operation and maintenance site in real time, performing data augmentation processing and inputting it into the trained person recognition and classification model for person recognition and classification.

[0012] As a possible implementation manner of this embodiment, in step S1, using a monocular camera device to collect personnel image data at the substation operation and maintenance site includes:

[0013] Placing a monocular camera that meets the specifications at a designated position at the substation operation and maintenance site;

[0014] Requesting to obtain the recording permission of the monocular camera;

[0015] Judging whether the recording permission is successfully obtained. If successful, extract 1 frame of image from the video every 2 seconds and save it. Otherwise, request to obtain the recording permission of the monocular camera again.

[0016] As a possible implementation manner of this embodiment, in step S2, performing data augmentation processing on the collected personnel image data to obtain processed personnel image data includes:

[0017] Calling the ImageDataGenerator function and setting the hyperparameters of rotation, translation, and scaling;

[0018] Performing rotation, translation, and scaling data augmentation processing on the collected personnel image data.

[0019] As a possible implementation of this embodiment, step S3 involves designing a residual neural network based on tensorflow and using the processed personnel image data for network training to obtain a trained person recognition and classification model, including:

[0020] Create a residual_block function;

[0021] Add a BatchNormalization layer and a "skip connection" to form a residual block;

[0022] Stack the residual blocks to form a residual network, and use the SGD gradient descent method to optimize the network hyperparameters;

[0023] Input the processed personnel image data into the residual network for training to obtain a trained person recognition and classification model.

[0024] As a possible implementation of this embodiment, step S4 involves real-time using the personnel image data at the substation operation and maintenance site, performing data augmentation processing, and inputting it into the trained person recognition and classification model for person recognition and classification, including:

[0025] Use the load_model class under tensorflow.keras.models to load the trained person recognition and classification model;

[0026] Use Tk under tkinter to create a main window and pop up a file selection dialog box to select a picture file;

[0027] Use Image.open() under the Pillow library to open the selected picture and adjust the image size to fit the model input;

[0028] Use the trained person recognition and classification model to convert the Pillow image object into a NumPy array and output the person recognition and classification result.

[0029] In a second aspect, an apparatus for scene person recognition and classification based on a multi-collaborative neural network provided by an embodiment of the present invention includes:

[0030] A data acquisition module for using a monocular camera device to acquire personnel image data at the substation operation and maintenance site;

[0031] A data processing module for performing data augmentation processing on the acquired personnel image data to obtain processed personnel image data;

[0032] A network training module for designing a residual neural network based on tensorflow and using the processed personnel image data for network training to obtain a trained person recognition and classification model;

[0033] A person recognition and classification module is used to adopt the personnel image data at the substation operation and maintenance site in real time, perform data enhancement processing, and input it into the trained person recognition and classification model for person recognition and classification.

[0034] As a possible implementation manner of this embodiment, the data acquisition module is specifically used for:

[0035] Place a monocular camera that meets the specifications at a designated position at the substation operation and maintenance site;

[0036] Request to obtain the recording permission of the monocular camera;

[0037] Judge whether the recording permission is obtained successfully. If successful, extract 1 frame of image from the video every 2 seconds and save it. Otherwise, request to obtain the recording permission of the monocular camera again.

[0038] As a possible implementation manner of this embodiment, the data processing module is specifically used for:

[0039] Call the ImageDataGenerator function and set the hyperparameters of rotation, translation, and scaling;

[0040] Perform data enhancement processing of rotation, translation, and scaling on the collected personnel image data.

[0041] As a possible implementation manner of this embodiment, the network training module is specifically used for:

[0042] Create the residual_block function;

[0043] Add the BatchNormalization layer and the "skip connection" to form a residual block;

[0044] Stack the residual blocks to form a residual network, and use the SGD gradient descent method to optimize the network hyperparameters;

[0045] Input the processed personnel image data into the residual network for training to obtain a trained person recognition and classification model.

[0046] As a possible implementation manner of this embodiment, the person recognition and classification module is specifically used for:

[0047] Use the load_model class under tensorflow.keras.models to load the trained person recognition and classification model;

[0048] Use Tk under tkinter to create a main window and pop up a file selection dialog box to select a picture file;

[0049] Open the selected image using Image.open() under the Pillow library and adjust the image size to fit the model input;

[0050] Convert the Pillow image object into a NumPy array using the trained person recognition and classification model and output the person recognition and classification results.

[0051] The beneficial effects produced by the technical solution of the embodiment of the present invention are as follows:

[0052] By obtaining the personnel image data at the substation operation and maintenance site and identifying and classifying the personnel appearing at the substation operation and maintenance site, the present invention realizes the digital management and control of substation equipment.

[0053] By obtaining the personnel image data at the substation operation and maintenance site, using the residual neural network designed based on tensorflow and related data processing technologies, processing and analyzing these image data, and realizing the accurate identification and classification of the personnel appearing at the substation operation and maintenance site. Based on this identification and classification result, establish the association between the personnel behavior and the state of the substation equipment, so as to achieve the digital management and control of the substation equipment and realize the intelligent and digital management and control of the on-site operation of substation operation and maintenance. Brief Description of the Drawings

[0054] Figure 1 is a flowchart of a method for scene person recognition and classification based on a multi-collaborative neural network shown according to an exemplary embodiment;

[0055] Figure 2 is a schematic structural diagram of a device for scene person recognition and classification based on a multi-collaborative neural network shown according to an exemplary embodiment. Detailed Embodiments

[0056] To more clearly illustrate the technical features of the solution of the present invention, the present invention will be described in detail below through specific embodiments and in conjunction with its drawings.

[0057] As Figure 1 shown, a method for scene person recognition and classification based on a multi-collaborative neural network provided by an embodiment of the present invention includes the following steps:

[0058] Step S1, use a monocular camera device to collect the personnel image data at the substation operation and maintenance site;

[0059] Step S2, perform data augmentation processing on the collected personnel image data to obtain the processed personnel image data;

[0060] Step S3, design a residual neural network based on tensorflow and use the processed personnel image data for network training to obtain a trained person recognition and classification model;

[0061] Step S4, in real time, adopt the personnel image data at the substation operation and maintenance site, perform data enhancement processing, and input it into the trained person recognition and classification model for person recognition and classification.

[0062] The present invention uses the person recognition and classification model to perform person recognition and classification on the personnel at the substation operation and maintenance site, assist in operation and maintenance and repair, realizes digital management and control of substation equipment, and supports digital management and control of on-site operations at the substation operation and maintenance site.

[0063] As a possible implementation manner of this embodiment, in step S1, the monocular camera device is used to adopt the personnel image data at the substation operation and maintenance site, including:

[0064] Place the monocular camera that meets the specifications at the designated position at the substation operation and maintenance site to ensure that people appearing in the required area can be clearly recorded without occlusion;

[0065] Request to obtain the recording permission of the monocular camera;

[0066] Judge whether the recording permission is successfully obtained. If successful, extract 1 frame of image from the video every 2 seconds and save it. Otherwise, request to obtain the recording permission of the monocular camera again.

[0067] As a possible implementation manner of this embodiment, in step S2, perform data enhancement processing on the collected personnel image data to obtain the processed personnel image data, including:

[0068] Call the ImageDataGenerator function and set the hyperparameters of rotation, translation and scaling;

[0069] For the collected personnel image data, use the tensorflow.keras.preprocessing.image.ImageDataGenerator: data augmentation class under TensorFlow / Keras to perform rotation, translation and scaling data augmentation processing to expand the diversity of training data.

[0070] As a possible implementation manner of this embodiment, in step S3, design a residual neural network based on tensorflow, and use the processed personnel image data for network training to obtain the trained person recognition and classification model, including:

[0071] Create a residual_block function;

[0072] Add a BatchNormalization layer and a "skip connection" to form a residual block;

[0073] Stack the residual blocks to form a residual network, and use the SGD gradient descent method to optimize the network hyperparameters;

[0074] Input the processed personnel image data into the residual network for training to obtain a trained person recognition classification model.

[0075] Use tensorflow.keras.models.Model: Used to create a Keras model object, representing a neural network. Model is the base class for building neural network models in Keras. tensorflow.keras.layers.Input: Represents the input layer of the model and can specify the shape of the input data. tensorflow.keras.layers.Conv2D: 2D convolutional layer for extracting image features. This layer uses convolutional operations and is usually used to build the convolutional layers in a CNN network. tensorflow.keras.layers.BatchNormalization: Batch normalization layer for accelerating training and improving model stability. tensorflow.keras.layers.Activation: Activation function layer for applying non-linear activation functions. The ReLU activation function is used in the code. tensorflow.keras.layers.Add: Performs element-wise addition operations and is usually used to implement residual connections. tensorflow.keras.layers.MaxPooling2D: 2D max pooling layer for pooling the image to reduce the size of the feature map. tensorflow.keras.layers.Flatten: Flattens multi-dimensional inputs into one dimension and is usually used to connect convolutional layers and fully connected layers. tensorflow.keras.layers.Dense: Fully connected layer for outputting classes or making regression predictions. Here it is used for the final classification output. tensorflow.keras.optimizers.SGD: Uses SGD (stochastic gradient descent) as the optimizer to control the way the network's weights are updated.

[0076] As a possible implementation of this embodiment, in step S4, the personnel image data at the substation operation and maintenance site is collected in real time, processed for data augmentation, and input into the trained person recognition classification model for person recognition and classification, including:

[0077] Use the load_model class under tensorflow.keras.models to load the trained person recognition classification model;

[0078] Use Tk under tkinter to create the main window and pop up a file selection dialog box to select a picture file;

[0079] Open the selected image using Image.open() under the Pillow library and adjust the image size to fit the model input;

[0080] Use the trained person recognition and classification model to convert the Pillow image object into a NumPy array and output the person recognition and classification results.

[0081] classify_image function: After the user clicks the button to select an image, the classify_image function will be triggered. Open a file dialog to select an image, which will be loaded and resized to 256x256 pixels. The selected model will be loaded according to the selection in the dropdown menu. Preprocess the selected image, including resizing, normalization ( / 255.0), and perform classification prediction using model.predict(). Display the predicted class, specifically used to find the corresponding person based on the content recognized and classified by the camera; achieve the display effect.

[0082] As Figure 2 shown, a scene person recognition and classification device based on a multi-collaborative neural network provided by an embodiment of the present invention includes:

[0083] A data acquisition module, used to collect personnel image data at the substation operation and maintenance site using a monocular camera device;

[0084] A data processing module, used to perform data enhancement processing on the collected personnel image data to obtain the processed personnel image data;

[0085] A network training module, used to design a residual neural network based on tensorflow and use the processed personnel image data for network training to obtain a trained person recognition and classification model;

[0086] A person recognition and classification module, used to collect personnel image data at the substation operation and maintenance site in real time, perform data enhancement processing, and input it into the trained person recognition and classification model for person recognition and classification.

[0087] As a possible implementation manner of this embodiment, the data acquisition module is specifically used for:

[0088] Place a monocular camera that meets the specifications at a designated position at the substation operation and maintenance site;

[0089] Request to obtain the recording permission of the monocular camera;

[0090] Judge whether the recording permission is successfully obtained. If successful, extract 1 frame of image from the video every 2 seconds and save it. Otherwise, request to obtain the recording permission of the monocular camera again.

[0091] As a possible implementation of this embodiment, the data processing module is specifically configured to:

[0092] Call the ImageDataGenerator function and set the hyperparameters for rotation, translation, and scaling;

[0093] Perform rotation, translation, and scaling data augmentation processing on the collected personnel image data.

[0094] As a possible implementation of this embodiment, the network training module is specifically configured to:

[0095] Create the residual_block function;

[0096] Add the BatchNormalization layer and the "skip connection" to form a residual block;

[0097] Stack the residual blocks to form a residual network, and use the SGD gradient descent method to optimize the network hyperparameters;

[0098] Input the processed personnel image data into the residual network for training to obtain a trained person recognition classification model.

[0099] As a possible implementation of this embodiment, the person recognition classification module is specifically configured to:

[0100] Use the load_model class under tensorflow.keras.models to load the trained person recognition classification model;

[0101] Use Tk under tkinter to create the main window and pop up a file selection dialog box to select a picture file;

[0102] Use Image.open() under the Pillow library to open the selected picture and adjust the image size to fit the model input;

[0103] Use the trained person recognition classification model to convert the Pillow image object into a NumPy array and output the person recognition classification result.

[0104] In some embodiments, the steps of using ImageDataGenerator for data augmentation (such as rotation, translation, scaling, etc.) to expand the diversity of training data include: using tensorflow.keras.preprocessing.image.ImageDataGenerator under TensorFlow / Keras: a data augmentation class for augmenting images (such as rotation, translation, scaling, etc.).

[0105] In some embodiments, the steps of training using a residual neural network designed based on tensorflow, obtaining a trained model, and saving it include: Using tensorflow.keras.models.Model: Used to create a Keras model object, representing a neural network. Model is the base class for building neural network models in Keras. tensorflow.keras.layers.Input: Represents the input layer of the model, and the shape of the input data can be specified. tensorflow.keras.layers.Conv2D: A 2D convolutional layer used to extract image features. This layer uses convolutional operations and is typically used to construct the convolutional layer in a CNN network. tensorflow.keras.layers.BatchNormalization: A batch normalization layer used to accelerate training and improve model stability. tensorflow.keras.layers.Activation: An activation function layer that applies a non-linear activation function. The ReLU activation function is used in the code. tensorflow.keras.layers.Add: Performs an element-wise addition operation and is typically used to implement a residual connection. tensorflow.keras.layers.MaxPooling2D: A 2D max pooling layer used to perform pooling operations on images, reducing the size of the feature map. tensorflow.keras.layers.Flatten: Flattens the multi-dimensional input into a one-dimensional vector, typically used to connect the convolutional layer and the fully connected layer. tensorflow.keras.layers.Dense: A fully connected layer used to output classes or perform regression predictions. Here, it is used for the final classification output. tensorflow.keras.optimizers.SGD: Uses SGD (Stochastic Gradient Descent) as the optimizer to control the way the network's weights are updated.

[0106] In some embodiments, the steps of the Tkinter-based image classification application for intercepting local pictures from the video obtained by the camera and performing person recognition and classification using a deep learning model include: The classify_image function: After the user clicks the button to select a picture, the classify_image function is triggered. A file dialog box is opened to select a picture, and the picture is loaded and resized to 256x256 pixels. The selected model is loaded according to the selection in the drop-down menu. The selected image is preprocessed, including resizing, normalization ( / 255.0), and classification prediction using model.predict(). The predicted class is displayed. Specifically, it is used to find the corresponding person based on the content recognized and classified from the camera to achieve the display effect.

[0107] The method of the present invention features lightweight and fast response. By optimizing the neural network structure and algorithm, the amount of calculation and memory occupancy are reduced, enabling rapid and accurate person recognition and classification in the complex environment of substation operation and maintenance sites, meeting the real-time requirements.

[0108] Through data augmentation and residual neural network training, the present invention effectively improves the accuracy and generalization ability of the model, enabling it to adapt to various complex situations at substation operation and maintenance sites and reducing the misjudgment rate.

[0109] The application program based on Tkinter of the present invention realizes convenient image classification and recognition operations. Operators can start using it with only simple training, greatly improving work efficiency, effectively assisting substation operation and maintenance and repair work, enhancing the digital management and control level of substation equipment, and strongly supporting the digital control of on-site operations in substation operation and maintenance.

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that modifications or equivalent replacements can still be made to the specific implementation manners of the present invention. Any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A scene character recognition and classification method based on multi-cooperative neural networks, characterized in that: The steps include: Step S1, using a monocular camera device to obtain personnel image data at a substation operation and maintenance site; Step S2, performing data enhancement processing on the collected personnel image data to obtain processed personnel image data; Step S3, designing a residual neural network based on tensorflow, and using the processed person image data to perform network training to obtain a trained person recognition classification model; Step S4, using the personnel image data at the substation operation and maintenance site in real time, performing data enhancement processing and inputting the trained character recognition and classification model for character recognition and classification.

2. The scene character recognition and classification method based on multi-cooperative neural networks according to claim 1 is characterized in that: The step S1, using a monocular camera device to obtain personnel image data at the substation operation and maintenance site, includes: Place the monocular camera that meets the specifications at the designated location of the substation operation and maintenance site; Request the recording permission of the monocular camera; Determine whether the recording permission is obtained successfully. If successful, extract 1 frame of image from the video every 2 seconds and save it. Otherwise, re-request the recording permission of the monocular camera.

3. The scene character recognition and classification method based on multi-cooperative neural networks according to claim 1 is characterized in that: The step S2, performing data enhancement processing on the collected personnel image data to obtain processed personnel image data, includes: Call the ImageDataGenerator function and set the hyperparameters of rotation, translation and scaling; The collected personnel image data is subjected to rotation, translation and scaling data enhancement processing.

4. The scene character recognition and classification method based on multi-cooperative neural networks according to claim 1 is characterized in that: The step S3, designing a residual neural network based on tensorflow, and using the processed personnel image data to perform network training to obtain a trained person recognition classification model, includes: Create residual_block function; Add BatchNormalization layer and "skip link" to form a residual block; The residual blocks are stacked to form a residual network, and the SGD gradient descent method is used to optimize the network hyperparameters; The processed person image data is input into the residual network for training to obtain a trained person recognition classification model.

5. The scene character recognition and classification method based on multi-cooperative neural networks according to any one of claims 1 to 4, characterized in that: The step S4 uses the personnel image data at the substation operation and maintenance site in real time, performs data enhancement processing and inputs the trained person recognition and classification model to perform person recognition and classification, including: Use the load_model class under tensorflow.keras.models to load the trained person recognition classification model; Use Tk under tkinter to create the main window, and pop up the file selection dialog box to select the image file; Use Image.open() from the Pillow library to open the selected image and resize the image to fit the model input; Use the trained person recognition classification model to convert the Pillow image object into a NumPy array and output the person recognition classification result.

6. A scene character recognition and classification device based on multi-cooperative neural networks, characterized in that: include: A data acquisition module is used to collect personnel image data at the substation operation and maintenance site using a monocular camera device; A data processing module is used to perform data enhancement processing on the collected personnel image data to obtain processed personnel image data; The network training module is used to design a residual neural network based on tensorflow and use the processed person image data for network training to obtain a trained person recognition classification model; The character recognition and classification module is used to use the personnel image data at the substation operation and maintenance site in real time, perform data enhancement processing, and input the trained character recognition and classification model for character recognition and classification.

7. The scene person recognition and classification device based on multiple collaborative neural networks according to claim 6 is characterized in that: The data acquisition module is specifically used for: Place the monocular camera that meets the specifications at the designated location of the substation operation and maintenance site; Request the recording permission of the monocular camera; Determine whether the recording permission is obtained successfully. If successful, extract 1 frame of image from the video every 2 seconds and save it. Otherwise, re-request the recording permission of the monocular camera.

8. The scene person recognition and classification device based on multiple collaborative neural networks according to claim 6 is characterized in that: The data processing module is specifically used for: Call the ImageDataGenerator function and set the hyperparameters of rotation, translation and scaling; The collected personnel image data is subjected to rotation, translation and scaling data enhancement processing.

9. The scene person recognition and classification device based on multiple collaborative neural networks according to claim 6, characterized in that: The network training module is specifically used for: Create residual_block function; Add BatchNormalization layer and "skip link" to form a residual block; The residual blocks are stacked to form a residual network, and the SGD gradient descent method is used to optimize the network hyperparameters; The processed person image data is input into the residual network for training to obtain a trained person recognition classification model.

10. The scene person recognition and classification device based on multiple collaborative neural networks according to any one of claims 6 to 9, characterized in that: The character recognition and classification module is specifically used for: Use the load_model class under tensorflow.keras.models to load the trained person recognition classification model; Use Tk under tkinter to create the main window, and pop up the file selection dialog box to select the image file; Use Image.open() from the Pillow library to open the selected image and resize the image to fit the model input; Use the trained person recognition classification model to convert the Pillow image object into a NumPy array and output the person recognition classification result.