Training method, classification method, detection method, device, system and equipment
By generating a multi-scale feature matrix from training videos and classification labels, and adjusting network parameters using marker points, the problem of poor accuracy in object behavior classification is solved, achieving more efficient object behavior recognition and improving the security and efficiency of financial escort tasks.
Patent Information
- Application Number
- CN202310576951.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-22
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2043-05-22
AI Technical Summary
Existing technologies have poor accuracy in classifying object behavior, making it impossible to accurately identify object behavior and affecting the security and efficiency of financial escort missions.
By acquiring multiple training videos and classification labels from the training set, a multi-scale feature matrix is extracted using the initial behavior classification model. Behavior classification is then performed based on the labeled points, and the network parameters are iteratively adjusted through the loss function to generate an object behavior classification model.
It improves the accuracy of object behavior classification, enabling more accurate identification of object behavior and enhancing the security and efficiency of financial escort tasks.
Smart Images

Figure CN116597516B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of behavior classification and finance, and in particular to a training method for an object behavior classification model, an object behavior classification method, a method for detecting the driving safety of transport vehicles, a training device for an object behavior classification model, an object behavior classification device, a vehicle monitoring system, electronic equipment, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the development of technology, financial business management has set higher standards for financial security and the quality and efficiency of financial services. Among these, financial escort, as an important component of financial operations, plays a vital role in maintaining the safety of property transportation, ensuring banking operations, improving the quality of financial services, and building order in the financial market.
[0003] In the financial escort business, traditional financial escort involves numerous manual operational procedures, which significantly impacts banks' financial operations and system management. The uncertainty, randomness, and mobility of transport vehicles and escort personnel are uncontrollable factors, primarily contributing to the reduced efficiency of financial escort operations. Furthermore, non-compliant behavior by escort personnel can compromise the security of transportation missions, leading to unnecessary security risks and property losses.
[0004] In realizing the present invention, the inventors discovered that the related technology has at least the following problems: when classifying the behavior of objects, the accuracy of the classification results is poor, and the behavior of objects cannot be accurately identified, thus affecting the escort mission. Summary of the Invention
[0005] In view of the above problems, this disclosure provides a training method for an object behavior classification model, an object behavior classification method, a method for detecting the driving safety of transport vehicles, a training device for an object behavior classification model, an object behavior classification device, a vehicle monitoring system, an electronic device, a computer-readable storage medium, and a computer program product.
[0006] According to a first aspect of this disclosure, a method for training an object behavior classification model is provided, comprising:
[0007] Obtain a training set, which includes multiple training videos and classification labels, wherein the videos include multiple image frames that are temporally correlated.
[0008] The training video is input into the initial behavior classification model, and multiple multi-scale feature matrices corresponding to the training video are output. Each of the multi-scale feature matrices includes multiple prediction regions including different key points of the object.
[0009] Based on the markers of the training video, behavior classification processing is performed on multiple prediction regions corresponding to the training video to obtain the behavior classification result of the training video. The markers represent the positions of different key points of the object in each image frame.
[0010] Input the classification result and classification label corresponding to each of the above training videos into the loss function, and output the loss result;
[0011] Based on the aforementioned loss results, the network parameters of the initial behavior classification model are iteratively adjusted to generate a trained object behavior classification model.
[0012] According to embodiments of this disclosure, when the number of multi-scale feature matrices is three, the above-mentioned input of the training video into the initial behavior classification model outputs multiple multi-scale feature matrices corresponding to the training video, including:
[0013] Based on the first preset step size, the feature extraction sub-model is used to perform channel adjustment and feature extraction processing on multiple of the above image frames to obtain the first image features;
[0014] The first image features are processed using the channel adjustment sub-model to obtain the second and third image features;
[0015] The first image feature, the second image feature, and the third image feature are processed using the first multi-scale sub-model, the second multi-scale sub-model, and the third multi-scale sub-model, respectively, to obtain three multi-scale feature matrices.
[0016] According to embodiments of this disclosure, the above-mentioned channel adjustment and feature extraction processing of multiple image frames based on a first preset step size, using a feature extraction sub-model to obtain first image features, includes:
[0017] Multiple first convolutional normalization layers are used to perform channel adjustment and feature extraction processing on multiple of the above image frames to obtain first intermediate features, wherein one of the above convolutional normalization layers corresponds to a first preset stride.
[0018] The first intermediate feature is processed by the first feature processing layer to perform channel adjustment and feature stacking to obtain the second intermediate feature;
[0019] The second intermediate feature is downsampled using the first downsampling layer to obtain the third intermediate feature;
[0020] The third intermediate feature is processed by channel adjustment and feature stacking using the second feature processing layer to obtain the first image feature.
[0021] According to embodiments of this disclosure, the above-mentioned processing of the first image features using a channel adjustment sub-model to obtain second and third image features includes:
[0022] The first image features are downsampled using a second downsampling layer to obtain a fourth intermediate feature.
[0023] The third feature processing layer is used to perform channel adjustment and feature extraction on the fourth intermediate feature to obtain the second image feature.
[0024] The second image features are downsampled using a third downsampling layer to obtain the fifth intermediate feature.
[0025] The fourth feature processing layer is used to perform channel adjustment and feature extraction on the fifth intermediate feature to obtain the third image feature.
[0026] According to embodiments of this disclosure, the first image feature, the second image feature, and the third image feature are processed using a first multi-scale sub-model, a second multi-scale sub-model, and a third multi-scale sub-model, respectively, to obtain three multi-scale feature matrices, including:
[0027] The first image feature and the first transition feature are processed using the first multi-scale sub-model described above, and a multi-scale feature matrix and a second transition feature are output.
[0028] The second image feature, the second transition feature, and the third transition feature are processed using the second multi-scale sub-model described above, and a multi-scale feature matrix, the first transition feature, and the fourth transition feature are output.
[0029] The third image feature and the fourth transition feature are processed using the third multi-scale sub-model described above, and a multi-scale feature matrix and the third transition feature are output.
[0030] According to embodiments of this disclosure, the above-mentioned processing of the first image feature and the first transition feature using the first multi-scale sub-model to output a multi-scale feature matrix and a second transition feature includes:
[0031] Based on the second preset stride, two second convolutional normalization layers are used to perform channel adjustment and feature extraction on the first image features and the first transition features to obtain the first channel features and the second channel features.
[0032] The second channel features are expanded using the first feature expansion layer to obtain the third channel features;
[0033] The first channel feature and the third channel feature are stacked using the first feature stacking layer to obtain the fourth channel feature;
[0034] The fifth feature processing layer is used to perform channel adjustment and feature extraction on the fourth channel features to obtain the fifth channel features, wherein the fifth channel features include two sub-channel features with a preset number of channels;
[0035] The second transition feature is obtained by downsampling one of the above-mentioned sub-channel features using the fourth downsampling layer.
[0036] The first convolutional stacking layer is used to perform convolution, normalization, and feature stacking on another of the above-mentioned sub-channel features to obtain the sixth channel feature;
[0037] Based on the second preset step size, the sixth channel features are processed by the third convolutional normalization layer to perform channel adjustment and feature extraction, resulting in the first multi-scale feature matrix. The first multi-scale feature matrix includes a first preset number of grids and a target number of channels.
[0038] According to embodiments of this disclosure, the above-described processing of the second image features, the second transition features, and the third transition features using the second multi-scale sub-model to output a multi-scale feature matrix, the first transition features, and the fourth transition features includes:
[0039] Based on the third preset stride, two fourth convolutional normalization layers are used to perform channel adjustment and feature extraction on the second image features and the third transition features to obtain the seventh channel features and the eighth channel features.
[0040] The eighth channel feature is expanded using the second feature expansion layer to obtain the ninth channel feature.
[0041] The seventh channel feature and the ninth channel feature are stacked using the second feature stacking layer to obtain the tenth channel feature.
[0042] The eleventh channel feature is obtained by using the sixth feature processing layer to perform channel adjustment and feature extraction on the tenth channel feature.
[0043] The eleventh channel feature and the second transition feature are stacked using the third feature stacking layer to obtain the twelfth channel feature.
[0044] The thirteenth channel feature is obtained by using the seventh feature processing layer to perform channel adjustment and feature extraction on the twelfth channel feature.
[0045] The thirteenth channel feature is downsampled using the fifth downsampling layer to obtain the fourth transition feature.
[0046] The thirteenth channel features are obtained by performing convolution, normalization, and feature stacking on the above-mentioned thirteenth channel features using the second convolution stacking layer;
[0047] Based on the third preset step size, the fourth channel feature is processed by the fifth convolutional normalization layer to perform channel adjustment and feature extraction, resulting in the second multi-scale feature matrix. The second multi-scale feature matrix includes a second preset number of grids and a target number of channels.
[0048] According to embodiments of this disclosure, the above-mentioned processing of the third image feature and the fourth transition feature using the third multi-scale sub-model to output a multi-scale feature matrix and the third transition feature includes:
[0049] The third image features are extracted, pooled, and stacked using a feature extraction stacking layer to obtain the third transition feature.
[0050] The third and fourth transition features are stacked using the fourth feature stacking layer to obtain the fifteenth channel feature.
[0051] The 15th channel feature is processed by the 8th feature processing layer to perform channel adjustment and feature extraction, resulting in the 16th channel feature.
[0052] The sixteenth channel feature is processed by convolution, normalization and feature stacking using the third convolution stacking layer to obtain the seventeenth channel feature;
[0053] Based on the fourth preset stride, the sixth convolutional normalization layer is used to perform channel adjustment and feature extraction processing on the seventeenth channel features to obtain the third multi-scale feature matrix. The third multi-scale feature matrix includes a third preset number of grids and a target number of channels.
[0054] According to embodiments of this disclosure, the above-mentioned behavior classification processing of multiple prediction regions corresponding to the training video based on the marker points of the training video to obtain the behavior classification result of the training video includes:
[0055] Based on the preset key point model, the position of the marked point in each of the above prediction regions is processed to obtain the state parameters of each of the above key points.
[0056] Based on the above state parameters, determine the behavioral state of the above object and the time and / or number of times it is in the above behavioral state;
[0057] If the above-mentioned behavioral state belongs to one of the classification lists and the preset conditions are met at the above-mentioned time or number of times, the above-mentioned training video will be classified as the first sub-class result.
[0058] If the above-mentioned behavioral state belongs to one of the above-mentioned classification lists and the above-mentioned time or number of times does not meet the preset conditions, the above-mentioned training video will be classified as the second classification sub-result.
[0059] If the aforementioned behavioral state does not belong to any of the aforementioned classification lists, the aforementioned training video is classified as a third sub-category result, wherein the aforementioned behavioral classification result includes the aforementioned first sub-category result, the aforementioned second sub-category result, and the aforementioned third sub-category result.
[0060] According to a second aspect of this disclosure, an object behavior classification method is provided, comprising:
[0061] Obtain the video to be classified, wherein the video to be classified includes multiple image frames to be classified that are temporally related;
[0062] Multiple image frames to be classified from the video to be classified are input into the object behavior classification model, and the predicted behavior classification result is output. The predicted behavior classification result represents the behavior posture of the object when the object exists in the video to be classified.
[0063] According to embodiments of this disclosure, the object behavior classification method further includes:
[0064] If the above prediction behavior classification results indicate that the behavior posture of the above object belongs to the preset behavior posture, the target information corresponding to the above preset behavior posture is determined from the information list.
[0065] The target information is presented to the aforementioned objects in a visual format.
[0066] According to a third aspect of this disclosure, a method for detecting the driving safety of a transport vehicle is provided, comprising:
[0067] While the aforementioned transport vehicle is in motion, the in-vehicle video of the aforementioned transport vehicle is captured in real time using the image acquisition device of the aforementioned transport vehicle, wherein the in-vehicle video includes multiple in-vehicle images that are sequentially related.
[0068] Multiple in-vehicle images from the aforementioned in-vehicle video are transmitted to a server, so that the server processes the in-vehicle video based on an object behavior classification model to obtain an in-vehicle behavior classification result, wherein the in-vehicle behavior classification result represents the behavioral posture of at least one object in the aforementioned in-vehicle video.
[0069] If the above in-vehicle behavior classification results indicate that the behavior of the above object belongs to a violation behavior posture, the first alarm information corresponding to the above violation behavior posture is determined from the alarm information list and transmitted to the above transport vehicle.
[0070] The first alarm information is displayed to the aforementioned objects in a visual format.
[0071] According to embodiments of this disclosure, the method for detecting the driving safety of transport vehicles further includes:
[0072] The vehicle status parameters of the aforementioned transport vehicle are collected using a sensor module, wherein the aforementioned vehicle status parameters include at least one of the following: tire pressure, vehicle interior temperature, engine status, range, vehicle speed, engine speed, and driving trajectory.
[0073] The vehicle status parameters are transmitted to the server so that if any of the vehicle status parameters meets the alarm conditions, the server will transmit the second alarm information corresponding to the parameter to the transport vehicle.
[0074] The second alarm information is displayed to the aforementioned objects in a visual format.
[0075] According to a fourth aspect of this disclosure, a training apparatus for an object behavior classification model is provided, comprising:
[0076] The first acquisition module is used to acquire a training set, wherein the training set includes multiple training videos and classification labels, and the videos include multiple image frames that are temporally correlated.
[0077] The multi-scale module is used to input the above training video into the initial behavior classification model and output multiple multi-scale feature matrices corresponding to the above training video. Each of the above multi-scale feature matrices includes multiple prediction regions including different key points of the object.
[0078] The first classification module is used to perform behavior classification processing on multiple prediction regions corresponding to the training video based on the marker points of the training video to obtain the behavior classification result of the training video, wherein the marker points represent the positions of different key points of the object in each image frame.
[0079] The loss module is used to input the classification result and classification label corresponding to each of the above training videos into the loss function and output the loss result.
[0080] The iterative module is used to iteratively adjust the network parameters of the initial behavior classification model based on the loss results, thereby generating a trained object behavior classification model.
[0081] According to a fifth aspect of this disclosure, an object behavior classification apparatus is provided, comprising:
[0082] The second acquisition module is used to acquire the video to be classified, wherein the video to be classified includes multiple image frames to be classified that are temporally associated.
[0083] The second classification module is used to input multiple image frames to be classified from the video to be classified into the object behavior classification model and output the predicted behavior classification result, wherein the predicted behavior classification result represents the behavior posture of the object when the object exists in the video to be classified.
[0084] According to a sixth aspect of this disclosure, a vehicle monitoring system is provided, comprising:
[0085] The server, as described above, is configured as follows:
[0086] The in-vehicle video is processed based on the object behavior classification model to obtain the in-vehicle behavior classification result, wherein the in-vehicle behavior classification result represents the behavior posture of at least one object in the in-vehicle video.
[0087] If the above in-vehicle behavior classification results indicate that the behavior posture of the above object belongs to the violation behavior posture, the first alarm information corresponding to the above violation behavior posture is determined from the alarm information list and transmitted to the alarm device.
[0088] Transport vehicles, including:
[0089] Vehicle body;
[0090] The image acquisition device is configured to acquire the in-vehicle video of the transport vehicle in real time and transmit it to the server while the main body of the vehicle is in motion, wherein the in-vehicle video includes multiple in-vehicle images that are sequentially associated.
[0091] The aforementioned alarm device is configured to display the first alarm information to the aforementioned object in a visual form.
[0092] According to embodiments of this disclosure, the vehicle monitoring system further includes:
[0093] The sensor module is configured as follows:
[0094] The vehicle status parameters of the aforementioned transport vehicles are collected and transmitted to the aforementioned server. The aforementioned vehicle status parameters include at least one of the following: tire pressure, vehicle interior temperature, engine status, range, vehicle speed, engine speed, and driving trajectory.
[0095] The server is also configured to transmit a second alarm message corresponding to any of the vehicle status parameters to the alarm device when any of the parameters meets the alarm conditions.
[0096] The aforementioned alarm device is also configured to display the second alarm information to the aforementioned object in a visual form.
[0097] A seventh aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the methods described above.
[0098] An eighth aspect of this disclosure also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the methods described above.
[0099] The ninth aspect of this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0100] According to embodiments of this disclosure, a multi-scale feature matrix is extracted from an image using an initial behavior classification model to determine the prediction region where the object is located as accurately as possible. The object behavior within the prediction region is classified using keypoint markers. The network parameters are iteratively adjusted based on the loss results determined by the markers of different keypoints and the classification results, thereby obtaining an object behavior classification model that can be used for behavior classification. Since the initial behavior classification model can continuously compress the image size and increase the number of image channels during the generation of the multi-scale feature matrix, and fuse different images to obtain the multi-scale feature matrix, using this multi-scale feature matrix for object behavior classification can yield more accurate classification results. This avoids the problem of inaccurate object behavior identification caused by the low classification accuracy of related technologies. Attached Figure Description
[0101] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0102] Figure 1 This illustration schematically shows a training method for an object behavior classification model or an application scenario diagram of an object behavior classification method according to embodiments of the present disclosure;
[0103] Figure 2 A flowchart illustrating a training method for an object behavior classification model according to an embodiment of the present disclosure is shown schematically.
[0104] Figure 3 A flowchart illustrating the processing of an object behavior classification model according to an embodiment of the present disclosure is shown.
[0105] Figure 4 A schematic diagram illustrating the structure of a CBM module according to an embodiment of the present disclosure is shown.
[0106] Figure 5A schematic diagram illustrating the structure of the ESCP1 module according to an embodiment of the present disclosure is shown.
[0107] Figure 6 A schematic diagram illustrating the structure of an ESCPM module according to an embodiment of the present disclosure is shown.
[0108] Figure 7 A schematic diagram illustrating the structure of the ESCP2 module according to an embodiment of the present disclosure is shown.
[0109] Figure 8 A schematic block diagram of a REPC module according to an embodiment of the present disclosure is shown.
[0110] Figure 9 A schematic diagram illustrating the structure of an SPPCM module according to an embodiment of the present disclosure is shown.
[0111] Figure 10 A schematic diagram illustrating the structure of a CBS module according to an embodiment of the present disclosure is shown.
[0112] Figure 11 This diagram illustrates the coordinate positions of key facial feature points according to an embodiment of the present disclosure.
[0113] Figure 12 A flowchart illustrating an object behavior classification method according to an embodiment of the present disclosure is shown schematically.
[0114] Figure 13 A flowchart illustrating a method for detecting the driving safety of a transport vehicle according to an embodiment of the present disclosure is shown schematically.
[0115] Figure 14 This schematic diagram illustrates a structural block diagram of a training apparatus for an object behavior classification model according to an embodiment of the present disclosure;
[0116] Figure 15 A schematic diagram illustrating the structure of an object behavior classification apparatus according to an embodiment of the present disclosure is shown.
[0117] Figure 16 A schematic block diagram of a vehicle monitoring system according to an embodiment of the present disclosure is shown; and
[0118] Figure 17 A block diagram schematically illustrates an electronic device suitable for implementing the above-described method according to an embodiment of the present disclosure. Detailed Implementation
[0119] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0120] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0121] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0122] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).
[0123] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of data (including but not limited to user personal information) comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals.
[0124] Embodiments of this disclosure provide a training method, a classification method, a detection method, a device system, an apparatus, and a medium. The training method includes acquiring a training set, wherein the training set includes multiple training videos and classification labels, and the videos include multiple temporally related image frames; inputting the training videos into an initial behavior classification model, outputting multiple multi-scale feature matrices corresponding to the training videos, wherein each multi-scale feature matrix includes multiple prediction regions including different key points of an object; performing behavior classification processing on the multiple prediction regions corresponding to the training videos based on the marker points of the training videos, obtaining behavior classification results for the training videos, wherein the marker points represent the positions of different key points of an object in each image frame; inputting the classification results and classification labels corresponding to each training video into a loss function, outputting a loss result; and iteratively adjusting the network parameters of the initial behavior classification model according to the loss result to generate a trained object behavior classification model.
[0125] Figure 1 The diagram illustrates a training method for an object behavior classification model or an application scenario of an object behavior classification method according to embodiments of the present disclosure.
[0126] like Figure 1 As shown, application scenario 100 according to this embodiment may include the behavior classification of drivers and passengers when bank escort vehicles perform escort tasks. Network 104 is used as a medium to provide a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. Network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0127] Users can interact with server 105 via network 104 using at least one of the first terminal device 101, second terminal device 102, and third terminal device 103 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, second terminal device 102, and third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0128] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0129] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0130] It should be noted that the training method or object behavior classification method of the object behavior classification model provided in this embodiment can generally be executed by server 105. Correspondingly, the training device or object behavior classification device of the object behavior classification model provided in this embodiment can generally be located in server 105. The training method or object behavior classification method of the object behavior classification model provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the training device or object behavior classification device of the object behavior classification model provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0131] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0132] Figure 2 A flowchart illustrating a training method for an object behavior classification model according to an embodiment of the present disclosure is shown.
[0133] like Figure 2 As shown, the training method of the object behavior classification model in this embodiment includes operations S210 to S250.
[0134] In operation S210, a training set is obtained, which includes multiple training videos and classification labels. The videos include multiple image frames that are temporally correlated.
[0135] In operation S220, the training video is input into the initial behavior classification model, and multiple multi-scale feature matrices corresponding to the training video are output. Each multi-scale feature matrix includes multiple prediction regions including different key points of the object.
[0136] In operation S230, based on the marker points of the training video, behavior classification processing is performed on multiple prediction regions corresponding to the training video to obtain the behavior classification result of the training video. Here, the marker points represent the positions of different key points of the object in each image frame.
[0137] In operation S240, the classification result and classification label corresponding to each training video are input into the loss function, and the loss result is output.
[0138] In operation S250, the network parameters of the initial behavior classification model are iteratively adjusted based on the loss results to generate a trained object behavior classification model.
[0139] According to embodiments of this disclosure, the training video includes multiple frames of images, where 2n (n≥1) frames include objects. Classification labels characterize the behaviors of the objects in the training video, such as smoking, using a mobile phone, not wearing a seatbelt, taking both hands off the steering wheel, drinking water, yawning, closing eyes, and other behaviors. Key points can be the eyes, mouth, hands, and abdomen of the human body, etc. The loss function can be the cross-entropy function.
[0140] According to embodiments of this disclosure, the prediction region can refer to dividing the generated multi-scale feature matrix into different numbers of grids, and selecting key points of an object within the grids using anchor boxes. The object can refer to the human body. It should be noted that in the technical solutions of this disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application of data (including but not limited to user personal information) comply with relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.
[0141] According to embodiments of this disclosure, each training video is input into an initial behavior classification model to obtain multiple multi-scale feature maps for each training video. Based on the marker points of the training videos, behavior classification processing is performed on multiple prediction regions corresponding to the training videos to obtain the behavior classification result of the training videos, for example, determining that the object in the training video is in a drinking state. The classification result and the classification label corresponding to the training video are input into a loss function to calculate the loss result. Then, the network parameters of the initial behavior classification model are iteratively adjusted according to the loss result to obtain the object behavior classification model.
[0142] According to embodiments of this disclosure, a multi-scale feature matrix is extracted from an image using an initial behavior classification model to determine the prediction region where the object is located as accurately as possible. The object behavior within the prediction region is classified using keypoint markers. The network parameters are iteratively adjusted based on the loss results determined by the markers of different keypoints and the classification results, thereby obtaining an object behavior classification model that can be used for behavior classification. Since the initial behavior classification model can continuously compress the image size and increase the number of image channels during the generation of the multi-scale feature matrix, and fuse different images to obtain the multi-scale feature matrix, using this multi-scale feature matrix for object behavior classification can yield more accurate classification results. This avoids the problem of inaccurate object behavior identification caused by the low classification accuracy of related technologies.
[0143] Figure 3 A flowchart illustrating the processing of an object behavior classification model according to an embodiment of the present disclosure is shown.
[0144] like Figure 3 As shown, when there are three multi-scale feature matrices, the training video is input into the initial behavior classification model, and the output is multiple multi-scale feature matrices corresponding to the training video, including:
[0145] Based on the first preset step size, the feature extraction sub-model is used to perform channel adjustment and feature extraction on multiple image frames to obtain the first image features;
[0146] The first image features are processed using a channel adjustment sub-model to obtain the second and third image features;
[0147] The first image feature, the second image feature, and the third image feature are processed using the first multi-scale sub-model, the second multi-scale sub-model, and the third multi-scale sub-model, respectively, to obtain three multi-scale feature matrices.
[0148] According to embodiments of this disclosure, the first preset step size can be specifically set according to actual needs, for example, it can be at least one or a combination of more of 1, 2, 3, etc.
[0149] According to embodiments of this disclosure, an example is a 3-channel image with a length and width of 640 and 640 pixels, respectively, i.e., the size of the image frame is 640×640×3. Based on a first preset step size, a feature extraction sub-model is used to perform channel adjustment and feature extraction processing on multiple image frames. For example, when the number of channels in multiple image frames is 3, the feature extraction sub-model can obtain a first image feature with 512 channels after channel adjustment of the multiple image frames. At the same time, the feature extraction sub-model can process the 640×640 pixel image frame into an 80×80 first image feature.
[0150] According to an embodiment of this disclosure, the channel adjustment sub-model further performs feature extraction and channel adjustment on the 80×80×512 first image feature, thereby obtaining a 40×40×1024 second image feature and a 20×20×1024 third image feature.
[0151] According to embodiments of this disclosure, based on a first image feature of 80×80×512, a second image feature of 40×40×1024, and a third image feature of 20×20×1024, the above image features are processed using a first multi-scale sub-model, a second multi-scale sub-model, and a third multi-scale sub-model to obtain three multi-scale feature matrices. The sizes of the three multi-scale feature matrices can be 13×13×39, 26×26×39, and 52×52×39, respectively.
[0152] According to embodiments of this disclosure, in the dimensions of the multi-scale feature matrix, 13×13, 26×26, and 52×52 represent that the image frame is divided into a grid of 13×13, 26×26, and 52×52, respectively, with each grid corresponding to three anchor boxes (i.e., prediction regions). When the center point of a keypoint is located within an anchor box, that anchor box is responsible for target detection. The initial behavior classification model adjusts the position of the anchor boxes by learning image features, making them eventually approach the true position of the keypoint, thereby obtaining the final predicted anchor boxes and achieving accurate localization and recognition of the keypoint. 39 is the product of 3 and 13, where 3 indicates that each grid will have three preset anchor boxes, and 13 represents the sum of 8, 4, and 1. Here, 8 is the number of behavior classification categories in this disclosure, 4 are the four position parameters of the anchor boxes, and 1 is the probability that the target is within the anchor box, i.e., the probability is 1 when the center point of the keypoint is within the anchor box, and 0 otherwise.
[0153] It should be noted that the dimensions of any of the image features mentioned above are merely illustrative and can be adjusted according to actual needs. The dimensions mentioned above are not intended to limit the scope of protection of this disclosure.
[0154] According to embodiments of this disclosure, referring to Figure 3 Based on a first preset step size, a feature extraction sub-model is used to perform channel adjustment and feature extraction processing on multiple image frames to obtain the first image features, including:
[0155] Using multiple first convolutional normalization layers (i.e. Figure 3 CBM in 1 Multiple image frames are processed for channel adjustment and feature extraction to obtain the first intermediate feature, wherein a convolutional normalization layer corresponds to a first preset stride.
[0156] Using the first feature processing layer (i.e. Figure 3 ESCP1 1The first intermediate feature is processed by channel adjustment and feature stacking to obtain the second intermediate feature;
[0157] Using the first downsampling layer (i.e. Figure 3 ESCPM in 1 The second intermediate feature is downsampled to obtain the third intermediate feature;
[0158] Utilizing the second feature processing layer (i.e. Figure 3 ESCP1 2 The third intermediate feature is processed by channel adjustment and feature stacking to obtain the first image feature.
[0159] According to embodiments of this disclosure, such as Figure 3 As shown, four first convolutional normalization layers are used as an example for illustration. Figure 3 In the feature extraction sub-model, the CBM is the first convolutional normalization layer. The first preset strides of the four first convolutional normalization layers are s=1, s=2, s=1, and s=2, respectively, and the kernel size is 3×3.
[0160] According to an embodiment of this disclosure, a 640×640×3 image frame is input to a first convolutional normalization layer, which outputs a 640×640×32 feature. This feature is then input to a second first convolutional normalization layer, which outputs a 320×320×64 feature. This feature is then input to a third first convolutional normalization layer, which outputs a 320×320×64 feature. This feature is then input to a fourth first convolutional normalization layer, which outputs a 160×160×128 first intermediate feature.
[0161] According to an embodiment of this disclosure, a first intermediate feature of 160×160×128 is input to a first feature processing layer, i.e. Figure 3 The ESCP1 sub-model in the middle feature extraction process performs channel adjustment and feature stacking, outputting a 160×160×256 second intermediate feature. The second intermediate feature is then downsampled using a first downsampling layer to obtain a 80×80×256 third intermediate feature.
[0162] According to an embodiment of this disclosure, a second feature processing layer is used to perform channel adjustment and feature stacking processing on the third intermediate feature of 80×80×256 to obtain a first image feature of 80×80×512.
[0163] Figure 4 A schematic block diagram of a CBM module according to an embodiment of the present disclosure is shown.
[0164] According to embodiments of this disclosure, such as Figure 4As shown, any one of the convolutional normalization layers (CBM modules) in the first convolutional normalization layer above, and the second, third, fourth, and fifth convolutional normalization layers below, includes a regular convolutional layer (Conv), a batch normalization layer (BN), and a Mish activation function.
[0165] The kernel sizes of ordinary convolutional layers include 1×1 and 3×3. The 1×1 kernel is used to adjust the number of feature map channels, and the 3×3 kernel is used for feature extraction. However, the distribution of key points in each image frame in the training set is unbalanced, which increases the cost of model training and inference. To address this, batch normalization is used to balance the data distribution, improve the model's convergence speed, and ensure that the feature information retains its original feature expression ability during transmission. Formula (1) is the Mish activation function. This activation function has the characteristics of being unbounded at the upper limit, bounded at the lower limit, and non-monotonic. It can not only improve the generalization ability of the model, but also better improve the nonlinear feature expression ability of the model.
[0166] Mish=x·tanh(ln(1+e x )) (1)
[0167] Figure 5 A schematic block diagram of the ESCP1 module according to an embodiment of the present disclosure is shown.
[0168] According to embodiments of this disclosure, any one of the first feature processing layer, the second feature processing layer mentioned above, and the third and fourth feature processing layers mentioned below can be a convolutional normalization layer. Figure 5 The ESCP1 module structure is shown.
[0169] According to embodiments of this disclosure, ESCP1 is a highly efficient feature extraction module that effectively controls the longest and shortest gradient paths, enabling the model to learn more effective features and thus improving model robustness. The ESCP1 module has feature transfer paths. The first path uses a CBM module with a 1×1 kernel size and a stride of 1 for channel adjustment. In the second path, a CBM module with a 1×1 kernel size and a stride of 1 is first used for channel adjustment (CBM module structure referenced). Figure 4 Then, four CBM modules with 3×3 kernels and a stride of 1 are used for feature extraction. A feature stacking layer (CONC) is then used to stack the multi-scale features from each path, meaning the number of channels in the stacked feature layer is the sum of the number of channels in each path, thus enriching the local feature information in the feature layer. Finally, a CBM module with 3×3 kernels and a stride of 1 is used to adjust the channels of the stacked feature layer, which is then used as the input features for the next layer.
[0170] Figure 6A schematic block diagram of the ESCPM module according to an embodiment of the present disclosure is shown.
[0171] According to embodiments of this disclosure, any one of the following downsampling layers—the first downsampling layer mentioned above, the second downsampling layer, the third downsampling layer, the fourth downsampling layer, and the fifth downsampling layer—can be […]. Figure 6 The ESCPM module structure is shown.
[0172] According to embodiments of this disclosure, the ECSPM module performs a downsampling operation, reducing the length and width of the feature layer to half its original size. The ECSPM module primarily consists of two paths. The first path uses a MaxPool pooling module with a 2×2 kernel to downsample the feature layer, followed by a CBM module with a 1×1 kernel and a stride of 1 to adjust the number of channels. In the second path, a CBM module with a 1×1 kernel and a stride of 1 is used to adjust the number of channels in the feature layer, followed by a CBM module with a 3×3 kernel and a stride of 2 to downsample the feature layer. Finally, the feature layers from both paths are stacked using a feature stacking layer (CONC), resulting in an output feature layer with a length and width half that of the input feature layer and twice the number of channels. This demonstrates how the ECSPM module performs the downsampling operation.
[0173] According to embodiments of this disclosure, referring to Figure 3 The first image features are processed using a channel adjustment sub-model to obtain the second and third image features, including:
[0174] Using the second downsampling layer (i.e. Figure 3 ESCPM in 2 The first image feature is downsampled to obtain the fourth intermediate feature;
[0175] Utilizing the third feature processing layer (i.e. Figure 3 ESCP1 3 The fourth intermediate feature is processed by channel adjustment and feature extraction to obtain the second image feature;
[0176] Using the third downsampling layer (i.e. Figure 3 ESCPM in 3 The second image features are downsampled to obtain the fifth intermediate feature;
[0177] Utilizing the fourth feature processing layer (i.e. Figure 3 ESCP1 4 The fifth intermediate feature is processed by channel adjustment and feature extraction to obtain the third image feature.
[0178] According to an embodiment of this disclosure, an 80×80×512 first image feature is input into a channel adjustment sub-model. After the second downsampling layer performs downsampling processing on the first image feature, it outputs a 40×40×512 fourth intermediate feature. The fourth intermediate feature is then processed by a third feature processing layer for channel adjustment and feature extraction to obtain a 40×40×1024 second image feature. The second image feature is then processed by a third downsampling layer for downsampling to obtain a 20×20×1024 fifth intermediate feature. Finally, the fifth intermediate feature is processed by a fourth feature processing layer for channel adjustment and feature extraction to obtain a 20×20×1024 third image feature.
[0179] According to embodiments of this disclosure, referring to Figure 3 The first image feature, second image feature, and third image feature are processed using the first multi-scale sub-model, the second multi-scale sub-model, and the third multi-scale sub-model, respectively, resulting in three multi-scale feature matrices, including:
[0180] The first image features and the first transition features are processed using the first multi-scale sub-model, and a multi-scale feature matrix and a second transition feature are output.
[0181] The second image feature, the second transition feature, and the third transition feature are processed using the second multi-scale sub-model, and a multi-scale feature matrix, the first transition feature, and the fourth transition feature are output.
[0182] The third multi-scale sub-model is used to process the third image features and the fourth transition features, and outputs a multi-scale feature matrix and the third transition features.
[0183] According to embodiments of this disclosure, a first multi-scale sub-model, a second multi-scale sub-model, and a third multi-scale sub-model are processed in parallel. The first multi-scale sub-model generates a multi-scale feature matrix based on the first image features output by the feature extraction sub-model and the first transition features output by the second multi-scale sub-model. The second multi-scale sub-model generates a multi-scale feature matrix based on the second transition features output by the first multi-scale sub-model, the second image features output by the channel adjustment sub-model, and the third transition features output by the third multi-scale sub-model. The third multi-scale sub-model generates a multi-scale feature matrix based on the third image features output by the channel adjustment sub-model and the fourth transition features output by the second multi-scale sub-model.
[0184] According to embodiments of this disclosure, referring to Figure 3 The first multi-scale sub-model is used to process the first image features and the first transition features, outputting a multi-scale feature matrix and a second transition feature, including:
[0185] Based on the second preset stride, two second convolutional normalization layers (i.e.) are used. Figure 3 CBM2 The first image feature and the first transition feature are processed by channel adjustment and feature extraction respectively to obtain the first channel feature and the second channel feature;
[0186] Using the first feature expansion layer (i.e. Figure 3 UPS 1 The second channel features are expanded using a feature layer to obtain the third channel features.
[0187] Utilizing the first feature stacking layer (i.e. Figure 3 CONC 1 The first and third channel features are stacked to obtain the fourth channel feature.
[0188] Utilizing the fifth feature processing layer (i.e. Figure 3 ESCP2 1 The fourth channel feature is processed by channel adjustment and feature extraction to obtain the fifth channel feature, which includes two sub-channel features with a preset number of channels.
[0189] Using the fourth downsampling layer (i.e. Figure 3 ESCPM 4 A second transition feature is obtained by downsampling a feature from one sub-channel.
[0190] Using the first convolutional stacking layer (i.e. Figure 3 REPC 1 The features of another sub-channel are convolved, normalized, and superimposed to obtain the features of the sixth channel.
[0191] Based on the second preset stride, the third convolutional normalization layer (i.e. Figure 3 CBS 1 The sixth channel features are processed by channel adjustment and feature extraction to obtain the first multi-scale feature matrix, which includes a first preset number of grids and the target number of channels.
[0192] According to embodiments of this disclosure, the second preset step size can be specifically set according to actual conditions, for example, the second preset step size s = 1.
[0193] According to an embodiment of this disclosure, a second convolutional normalization layer processes the first image feature of 80×80×512 and outputs the first channel feature of 80×80×128, and another second convolutional normalization layer processes the first transition feature of 40×40×256 and outputs the second channel feature of 40×40×128.
[0194] According to embodiments of this disclosure, a first feature expansion layer is used to expand the second channel feature to obtain a third channel feature of 80×80×128. A first feature stacking layer is used to stack the first and third channel features to obtain a fourth channel feature of 80×80×256. A fifth feature processing layer is used to adjust the channel and extract features from the fourth channel feature to obtain a fifth channel feature of 80×80×256. The fifth channel feature includes two sub-channel features with a preset number of channels, both of which are 80×80×128.
[0195] According to an embodiment of this disclosure, a fourth downsampling layer is used to downsample the feature of one sub-channel to obtain a second transition feature of 40×40×256. A first convolutional stacking layer is used to convolve, normalize, and stack the feature of another sub-channel to obtain a sixth channel feature. Based on a second preset stride, a third convolutional normalization layer is used to adjust the channel and extract features from the sixth channel feature to obtain a first multi-scale feature matrix.
[0196] Figure 7 A schematic block diagram of the ESCP2 module according to an embodiment of the present disclosure is shown.
[0197] According to embodiments of this disclosure, any one of the fifth feature processing layer mentioned above, and the sixth, seventh, and eighth feature processing layers mentioned below, includes the following: Figure 7 The ESCP2 module structure is shown. Similar to the ESCP1 module, the ESCP2 module mainly uses CBM modules with convolutional kernel sizes of 1×1 and 3×3 and a stride of 1 for channel adjustment and feature extraction, respectively. Unlike ESCP1, the ESCP2 module performs feature splitting and stacking after each 3×3 CBM module. This operation not only improves the efficiency of feature transfer in the network but also effectively enriches deep local features.
[0198] Figure 8 A schematic block diagram of a REPC module according to an embodiment of the present disclosure is shown.
[0199] According to embodiments of this disclosure, the first convolutional stacking layer mentioned above, the second convolutional stacking layer mentioned below, and the third convolutional stacking layer all include the following... Figure 8The REPC module shown primarily consists of a Conv ordinary convolution operation, BN batch normalization, and an Add weighted operation. The REPC module passes the input features through three paths. First, a 3×3 kernel ordinary convolution is used for feature extraction in the first path. Then, a 1×1 kernel ordinary convolution is used for feature smoothing in the second path. Finally, the input features are directly subjected to batch normalization in the last path. After processing, the features from the three paths are fused using an Add weighted operation. The number of channels remains unchanged after the Add weighted operation. Because the features are superimposed, the output features processed by the REPC module have more accurate target localization information.
[0200] According to embodiments of this disclosure, referring to Figure 3 The second multi-scale sub-model is used to process the second image features, the second transition features, and the third transition features, outputting a multi-scale feature matrix, the first transition feature, and the fourth transition feature, including:
[0201] Based on the third preset stride, two fourth convolutional normalization layers (i.e.) are used. Figure 3 CBM 4 The second image features and the third transition features are processed by channel adjustment and feature extraction respectively to obtain the seventh channel features and the eighth channel features;
[0202] Using the second feature expansion layer (i.e. Figure 3 UPS 2 The eighth channel feature is subjected to feature layer expansion processing to obtain the ninth channel feature;
[0203] Utilizing the second feature stacking layer (i.e. Figure 3 CONC 2 The features of the seventh and ninth channels are stacked to obtain the features of the tenth channel.
[0204] Utilizing the sixth feature processing layer (i.e. Figure 3 ESCP2 6 The features of the tenth channel are processed by channel adjustment and feature extraction to obtain the features of the eleventh channel;
[0205] Utilizing a third feature stacking layer (i.e. Figure 3 CONC 3 The eleventh channel feature and the second transition feature are stacked to obtain the twelfth channel feature.
[0206] Utilizing the seventh feature processing layer (i.e. Figure 3 ESCP2 7 Channel adjustment and feature extraction were performed on the twelfth channel features to obtain the thirteenth channel features;
[0207] Using the fifth downsampling layer (i.e. Figure 3 ESCPM 5 The thirteenth channel feature is downsampled to obtain the fourth transition feature;
[0208] Using the second convolutional stack (i.e. Figure 3 REPC 2 The thirteenth channel features are processed by convolution, normalization, and feature overlay to obtain the fourteenth channel features;
[0209] Based on the third preset stride, the fifth convolutional normalization layer (i.e. Figure 3 CBS 5 The fourteenth channel features are processed by channel adjustment and feature extraction to obtain the second multi-scale feature matrix, which includes a second preset number of grids and a target number of channels.
[0210] According to the embodiments of this disclosure, the third preset step size can also be specifically set according to the actual situation. This disclosure selects step size s = 1 as the third preset step size.
[0211] According to an embodiment of this disclosure, based on a third preset step size s=1, a fourth convolutional normalization layer processes the 40×40×1024 second image features output by the channel adjustment sub-model to obtain a 40×40×256 seventh channel feature. Another fourth convolutional normalization layer performs channel adjustment and feature extraction processing on the 20×20×512 third transition feature to obtain a 20×20×256 eighth channel feature. The eighth channel feature is then expanded using a second feature expansion layer to obtain a 40×40×256 ninth channel feature.
[0212] According to an embodiment of this disclosure, the seventh channel feature and the ninth channel feature are stacked using a second feature stacking layer to obtain a tenth channel feature of 40×40×512; the tenth channel feature is adjusted and extracted using a sixth feature processing layer to obtain an eleventh channel feature of 40×40×256, namely the first transition feature; the eleventh channel feature and the second transition feature are stacked using a third feature stacking layer to obtain a twelfth channel feature of 40×40×512.
[0213] According to embodiments of this disclosure, the 12th channel feature is processed by a seventh feature processing layer to perform channel adjustment and feature extraction, resulting in a 40×40×256 13th channel feature; the 13th channel feature is processed by a fifth downsampling layer to perform downsampling, resulting in a 20×20×512 fourth transition feature; the 13th channel feature is processed by a second convolution stacking layer to perform convolution, normalization, and feature stacking, resulting in a 14th channel feature; based on a third preset stride, the 14th channel feature is processed by a fifth convolution normalization layer to perform channel adjustment and feature extraction, resulting in a second multi-scale feature matrix, wherein the second multi-scale feature matrix includes a second preset number of grids and a target number of channels.
[0214] According to embodiments of this disclosure, the upsampling method for the first feature expansion layer and the second feature expansion layer is to expand the feature layer using the UPS module of the nearest neighbor interpolation algorithm, that is, to double the length and width of the feature layer while keeping the number of channels unchanged.
[0215] According to embodiments of this disclosure, referring to Figure 3 The third multi-scale sub-model is used to process the third image features and the fourth transition features, outputting a multi-scale feature matrix and the third transition features, including:
[0216] Using feature extraction stacking layers (i.e.) Figure 3 SPPCM (Simplified Chinese SPPCM) is used to extract, pool, and stack the third image features to obtain the third transition feature;
[0217] Utilizing the fourth feature stacking layer (i.e. Figure 3 CONC 4 The third and fourth transition features are stacked to obtain the fifteenth channel feature.
[0218] Using the eighth feature processing layer (i.e. Figure 3 ESCP2 8 The fifteenth channel features are processed by channel adjustment and feature extraction to obtain the sixteenth channel features;
[0219] Using the third convolutional stacking layer (i.e. Figure 3 REPC 3 The sixteenth channel features are processed by convolution, normalization, and feature overlay to obtain the seventeenth channel features;
[0220] Based on the fourth preset stride, the sixth convolutional normalization layer (i.e. Figure 3 CBS 6 Channel adjustment and feature extraction are performed on the seventeenth channel feature to obtain the third multi-scale feature matrix, which includes a third preset number of grids and the target number of channels.
[0221] According to the embodiments of this disclosure, the fourth preset step size can also be specifically set according to the actual situation. This disclosure selects step size s = 1 as the fourth preset step size.
[0222] According to an embodiment of this disclosure, a feature extraction stacking layer is used to perform feature extraction, pooling, and stacking processing on the third image features to obtain a third transition feature of 20×20×512; a fourth feature stacking layer is used to perform feature stacking processing on the third transition feature and the fourth transition feature to obtain a fifteenth channel feature of 20×20×1024.
[0223] According to an embodiment of this disclosure, the fifteenth channel feature is processed by the eighth feature processing layer to perform channel adjustment and feature extraction to obtain the sixteenth channel feature; the sixteenth channel feature is processed by the third convolution stacking layer to perform convolution, normalization and feature stacking to obtain the seventeenth channel feature; based on the fourth preset stride s=1, the seventeenth channel feature is processed by the sixth convolution normalization layer to perform channel adjustment and feature extraction to obtain the third multi-scale feature matrix, wherein the third multi-scale feature matrix includes a third preset number of grids and a target number of channels.
[0224] Figure 9 A schematic block diagram of an SPPCM module according to an embodiment of the present disclosure is shown.
[0225] According to embodiments of this disclosure, the feature extraction stacking layer includes, as follows: Figure 9 The SPPCM module shown is used to increase the receptive field of the model, making the object behavior classification model adaptable to images of different resolutions. The SPPCM module consists of two paths. In the first path, CBM modules with convolutional kernel sizes of 1×1 and 3×3 and a stride of 1 are used for channel adjustment and feature extraction, respectively. Pooling operations with kernel sizes of 5×5, 9×9, 13×13, and 1×1 are used to increase the model's receptive field for multi-scale objects, making the model more robust to multi-scale targets. In the second path, the input features are processed by a CBM module with a convolutional kernel size of 1×1 and a stride of 1 for channel adjustment. Then, the feature layers from the two paths are stacked by the CONC module, and the stacked feature layer is processed by the CBM module as the input features for the next layer.
[0226] Figure 10 A schematic block diagram of a CBS module according to an embodiment of the present disclosure is shown.
[0227] According to embodiments of this disclosure, the sixth convolutional normalization layer includes, as follows: Figure 10 The CBS module shown consists of Conv ordinary convolution operation, batch normalization operation and SiLU activation function.
[0228] Unlike the CBM module, the CBS module uses the SiLU activation function to smooth the output features of the three multi-scale feature layers. The SiLU activation function is shown in Equation (2).
[0229]
[0230] Figure 11 The diagram illustrates the coordinate positions of key facial feature points according to an embodiment of the present disclosure.
[0231] According to embodiments of this disclosure, based on the markers of the training video, behavior classification processing is performed on multiple prediction regions corresponding to the training video to obtain the behavior classification result of the training video, including:
[0232] The positions of the marked points in each prediction region are processed based on the preset key point model to obtain the state parameters of each key point;
[0233] Based on multiple state parameters, determine the object's behavioral state and the time and / or number of times it is in that behavioral state;
[0234] If the behavior state belongs to one of the categories and the preset conditions are met in terms of time or number of times, the training video will be classified as the first category sub-result.
[0235] If the behavior state belongs to one of the categories and the time or number of times does not meet the preset conditions, the training video will be classified as the second category sub-result.
[0236] If the behavior state does not belong to any of the classification lists, the training video is classified into the third sub-category, where the behavior classification results include the first sub-category, the second sub-category, and the third sub-category.
[0237] According to embodiments of this disclosure, the preset keypoint model may include an eye model. Using the coordinate positions of key eye feature points determined by the eye model, state parameters of the eye, such as the actual aspect ratio, are calculated. Based on a preset aspect ratio threshold, the behavioral state of the eye is determined, such as open or closed. Finally, the proportion of eye closure time or the number of closures between different image frames is calculated to classify the behavior of the object. Figure 11 As shown, these are the coordinates of key facial feature points.
[0238] According to an embodiment of this disclosure, taking the left eye as an example, mathematical modeling of the eye can yield an eye model as shown in formula (3).
[0239]
[0240] Where, p 38 p42 p 39 p 41 p 37 and p 40 They are respectively Figure 11 The coordinates of the left eye keypoint. When the object's eyes are open, the value of `eye` remains within a certain range. When the object's eyes are closed, `eye` approaches 0. Therefore, when `eye` is below a preset aspect ratio threshold, such as 0.3, the eyes are in a closed state. When `eye` decreases from a certain value to the preset aspect ratio threshold and then rapidly rises to above the preset aspect ratio threshold, it can be defined as one blink, i.e., one blink count.
[0241] According to embodiments of this disclosure, let t1 be the time when the eyes are 80% open when they are about to close, t2 be the time when the eyes are 20% open when they are about to close, t3 be the time when the eyes are 20% open when they are about to open, and t4 be the time when the eyes are 80% open when they are about to open. The percentage of time the eyes are in a closed state within a certain time period is defined as t. The time the object is in a behavioral state is determined by t, as shown in formula (4).
[0242]
[0243] According to embodiments of this disclosure, if t ≥ 0.1, the closing time of the object is considered to have exceeded a preset condition. If t < 0.1, the closing time of the object is considered to have not exceeded the preset condition. The mathematical models for the left and right eyes are similar.
[0244] According to embodiments of this disclosure, when the object's eye behavior falls under the closed state category in the classification list, if the closing time exceeds a preset time threshold (e.g., 0.1 seconds) or the number of closures per unit time exceeds a frequency threshold (e.g., 10 times per minute), the object's training video is classified as a first sub-category result, indicating that the object is in a drowsy state. If the object's eye behavior falls under the closed state category in the classification list, but neither the closing time nor the number of closures exceeds the preset threshold, the object's training video can be classified as a second sub-category result, indicating that the object occasionally closes its eyes but is not drowsy. If the behavior state does not fall under any category in the classification list, the training video is classified as a third sub-category result, indicating that the object's behavior is good.
[0245] According to embodiments of this disclosure, the preset key point model may further include a mouth model. The actual aspect ratio of the mouth is calculated based on the coordinate positions of the mouth feature key points determined by the mouth model, thereby determining whether the mouth is in an open or closed state based on a preset aspect ratio threshold. Finally, the proportion of mouth closure time between each image frame is calculated and compared with a yawning threshold to classify the behavior state of the object. The mouth model is shown in formula (5).
[0246]
[0247] In formula (5), p 51 p 59 p 53 p 57 p 55 and p 49 They are respectively Figure 11 The coordinates of key points in the middle of the mouth. If the mouth value is greater than or equal to 0.75, the object is considered to be yawning, and the yawn count is incremented by 1. In yawn detection, if the number of yawns or the duration of yawning detected within a preset time period (e.g., 30 seconds) exceeds a preset threshold (e.g., the threshold for the number of yawns is 2, and the threshold for the duration is 15 seconds), the object is classified as being in a state of fatigue.
[0248] Figure 12 A flowchart illustrating an object behavior classification method according to an embodiment of the present disclosure is shown schematically.
[0249] like Figure 12 As shown, the object behavior classification method in this embodiment includes operations S1210 to S1220.
[0250] In operation S1210, the video to be classified is acquired, wherein the video to be classified includes multiple image frames to be classified that are temporally related.
[0251] In operation S1220, multiple image frames to be classified from the video to be classified are input into the object behavior classification model, and the predicted behavior classification result is output. The predicted behavior classification result represents the behavior and posture of the object when the object exists in the video to be classified.
[0252] According to embodiments of this disclosure, the video to be classified can be captured by an image acquisition device, such as a camera or webcam.
[0253] According to embodiments of this disclosure, the behavior and posture of objects in the video to be classified can be classified by inputting the video to be classified into an object behavior classification model.
[0254] According to embodiments of this disclosure, a multi-scale feature matrix is extracted from the image frame to be classified using an object behavior classification model, so as to determine the prediction region where the object is located as much as possible. The object behavior within the prediction region is classified using the marker points of key points, thereby obtaining the classification result of the object's behavior posture. Since the object behavior classification model can continuously compress the image size and increase the number of image channels during the generation of the multi-scale feature matrix, and fuse different images to obtain the multi-scale feature matrix, the classification of object behavior using this multi-scale feature matrix can obtain a more accurate classification result, avoiding the problem of inaccurate identification of object behavior caused by the low classification accuracy of related technologies.
[0255] According to embodiments of this disclosure, the object behavior classification method further includes:
[0256] If the predicted behavior classification results indicate that the object’s behavior posture belongs to the preset behavior posture, the target information corresponding to the preset behavior posture is determined from the information list.
[0257] Present target information to the object in a visual format.
[0258] According to embodiments of this disclosure, preset behavioral postures may include smoking, using a mobile phone, not wearing a seatbelt, taking both hands off the steering wheel, drinking water, yawning, closing eyes, etc.
[0259] According to embodiments of this disclosure, the information list can pre-store information corresponding to different categories. For example, for smoking behavior, it can store "No smoking," and for yawning and closing eyes, it can store "You are fatigued, please do not drive while fatigued." It should be noted that the information in the information list can be specifically set according to actual needs; the above is merely an example.
[0260] According to embodiments of this disclosure, when an object exhibits the aforementioned behaviors, the corresponding target information can be displayed to the object through various visual means such as a display screen and a speaker.
[0261] According to embodiments of this disclosure, sending reminder messages to objects through visual display can prevent non-compliant behavior of objects from causing unnecessary safety hazards to driving or working in certain scenarios.
[0262] Figure 13 A flowchart illustrating a method for detecting the driving safety of a transport vehicle according to an embodiment of the present disclosure is shown.
[0263] like Figure 13 As shown, the method for detecting the driving safety of a transport vehicle in this embodiment includes operations S1310 to S1340.
[0264] When operating S1310, while the transport vehicle is in motion, the image acquisition device of the transport vehicle is used to acquire in-vehicle video in real time, wherein the in-vehicle video includes multiple in-vehicle images that are correlated in time.
[0265] In operation S1320, multiple in-vehicle images of the in-vehicle video are transmitted to the server so that the server processes the in-vehicle video based on the object behavior classification model to obtain the in-vehicle behavior classification result, wherein the in-vehicle behavior classification result represents the behavior posture of at least one object in the in-vehicle video.
[0266] In operation S1330, if the in-vehicle behavior classification result indicates that the object’s behavior posture is a violation posture, the first alarm information corresponding to the violation posture is determined from the alarm information list and transmitted to the transport vehicle.
[0267] In operation S1340, the first alarm information is displayed to the object in a visual form.
[0268] According to embodiments of this disclosure, the transport vehicle can refer to an escort vehicle. When the driver (i.e., the object) is performing an escort task, the driver's state largely determines the safety of the escort. If the driver is fatigued or not wearing a seat belt, a vehicle accident is likely to occur. Therefore, image acquisition devices such as cameras can be installed in the escort vehicle in advance. During the escort task, the escort vehicle transmits the in-vehicle video captured by the image acquisition device to the server in real time. The server classifies the driver's behavior and posture in the in-vehicle video in real time using an object behavior classification model to obtain the in-vehicle behavior classification result.
[0269] According to embodiments of this disclosure, if the in-vehicle behavior classification results indicate that the object's behavior posture is an illegal behavior posture, such as the driver being in a fatigued state with eyes closed or yawning, the first alarm information corresponding to the illegal behavior posture is determined from the alarm information list and transmitted to the transport vehicle. Thus, the first alarm information is sent to the driver through the alarm device installed in the escort vehicle, thereby minimizing the occurrence of escort safety accidents caused by the driver's illegal behavior posture.
[0270] It should be noted that the driving safety testing method disclosed herein can detect not only the driver's behavior and posture, but also the behavior and posture of passengers inside the vehicle.
[0271] According to embodiments of this disclosure, a multi-scale feature matrix is extracted from image frames of in-vehicle video using an object behavior classification model to determine the prediction region where the object is located as much as possible. The object behavior within the prediction region is classified using key point markers to determine whether the object inside the transport vehicle exhibits any illegal behavior or posture. An alarm device then sends an alarm message to the object. Since the object behavior inside the transport vehicle is detected in real time during the vehicle's operation, the safety and confidentiality of the transport vehicle during the transport mission are improved, avoiding the problem of reduced safety and confidentiality caused by illegal behavior of the object during the execution of the transport mission.
[0272] It should be noted that the object behavior classification model disclosed herein can not only classify the behavior of drivers and passengers in transport vehicles, but also classify the behavior of other objects. For example, it can classify the posture of animals in a zoo, such as determining whether an animal is in a static, running, or flying state.
[0273] According to embodiments of this disclosure, the method for detecting the driving safety of transport vehicles further includes:
[0274] The vehicle status parameters of the transport vehicle are collected using a sensor module. These vehicle status parameters include at least one of the following: tire pressure, interior temperature, engine status, range, vehicle speed, engine speed, and driving trajectory.
[0275] The vehicle status parameters are transmitted to the server so that if any of the vehicle status parameters meets the alarm conditions, the server will transmit the second alarm information corresponding to the parameter to the transport vehicle.
[0276] The second alarm information is displayed to the target audience in a visual format.
[0277] According to embodiments of this disclosure, the second alarm information can be set according to actual conditions. For example, when the tire pressure is low, the second alarm information can be "abnormal tire pressure, please check if the tire is damaged", etc.
[0278] According to embodiments of this disclosure, the sensor module includes a vehicle positioning module. This module collects the current operating status of the transport vehicle in real time through an onboard GPS positioning and navigation system, enabling data sharing between the escort personnel and the remote terminal. If the transport vehicle deviates from the predetermined driving trajectory, drives in the wrong direction, runs red lights, speeds, or drives at low speeds, the onboard terminal automatically sends an alarm signal to the application monitoring unit. The data center then shares the information through a big data cloud, and finally, the application monitoring unit achieves real-time monitoring and quickly provides effective supervision, scheduling, and warnings regarding the transport vehicle's operating status.
[0279] According to embodiments of this disclosure, the sensor module can also collect current vehicle operating parameters, such as tire pressure, interior temperature, engine status, fuel level, battery level, vehicle speed, engine speed, and other dashboard parameters. Furthermore, the sensor module may also include a collision sensor, which collects the vehicle collision coefficient to determine whether a collision or other emergency has occurred, thereby promptly issuing a corresponding second alarm message to the application monitoring unit and the driver after a collision.
[0280] According to embodiments of this disclosure, the driving safety detection method can also determine the identity information of the current driver (or passenger) through in-vehicle video. The server determines whether the identity information belongs to the intended driver (or intended passenger). If it is determined that the identity information does not belong to the intended driver (or intended passenger), the server can promptly send corresponding second alarm information to the application supervision unit and the driver.
[0281] Figure 14 A schematic block diagram of a training apparatus for an object behavior classification model according to an embodiment of the present disclosure is shown.
[0282] like Figure 14 As shown, the training device 1400 for the object behavior classification model in this embodiment includes a first acquisition module 1410, a multi-scale module 1420, a first classification module 1430, a loss module 1440, and an iteration module 1450.
[0283] The first acquisition module 1410 is used to acquire a training set, wherein the training set includes multiple training videos and classification labels, and the videos include multiple image frames that are temporally correlated.
[0284] The multi-scale module 1420 is used to input the training video into the initial behavior classification model and output multiple multi-scale feature matrices corresponding to the training video. Each multi-scale feature matrix includes multiple prediction regions that include different key points of the object.
[0285] The first classification module 1430 is used to perform behavior classification processing on multiple prediction regions corresponding to the training video based on the marker points of the training video to obtain the behavior classification result of the training video. The marker points represent the positions of different key points of the object in each image frame.
[0286] The loss module 1440 is used to input the classification result and classification label corresponding to each training video into the loss function and output the loss result.
[0287] Iteration module 1450 is used to iteratively adjust the network parameters of the initial behavior classification model based on the loss results, thereby generating a trained object behavior classification model.
[0288] According to embodiments of this disclosure, a multi-scale feature matrix is extracted from an image using an initial behavior classification model to determine the prediction region where the object is located as accurately as possible. The object behavior within the prediction region is classified using keypoint markers. The network parameters are iteratively adjusted based on the loss results determined by the markers of different keypoints and the classification results, thereby obtaining an object behavior classification model that can be used for behavior classification. Since the initial behavior classification model can continuously compress the image size and increase the number of image channels during the generation of the multi-scale feature matrix, and fuse different images to obtain the multi-scale feature matrix, using this multi-scale feature matrix for object behavior classification can yield more accurate classification results. This avoids the problem of inaccurate object behavior identification caused by the low classification accuracy of related technologies.
[0289] According to embodiments of this disclosure, when the number of multi-scale feature matrices is three, the multi-scale module 1420 includes a feature extraction submodule, a first obtaining submodule, and a second obtaining submodule.
[0290] The feature extraction submodule is used to perform channel adjustment and feature extraction processing on multiple image frames based on a first preset step size, and to obtain the first image features.
[0291] The first submodule is used to process the first image features using the channel adjustment submodel to obtain the second and third image features.
[0292] The second submodule is used to process the first image features, the second image features, and the third image features using the first multi-scale sub-model, the second multi-scale sub-model, and the third multi-scale sub-model, respectively, to obtain three multi-scale feature matrices.
[0293] According to embodiments of this disclosure, the feature extraction submodule includes a first obtaining unit, a second obtaining unit, a first downsampling unit, and a third obtaining unit.
[0294] The first obtaining unit is used to perform channel adjustment and feature extraction processing on multiple image frames using multiple first convolutional normalization layers to obtain first intermediate features, wherein one convolutional normalization layer corresponds to a first preset stride.
[0295] The second obtaining unit is used to perform channel adjustment and feature stacking processing on the first intermediate feature using the first feature processing layer to obtain the second intermediate feature;
[0296] The first downsampling unit is used to downsample the second intermediate feature using the first downsampling layer to obtain the third intermediate feature;
[0297] The third obtaining unit is used to perform channel adjustment and feature stacking processing on the third intermediate feature using the second feature processing layer to obtain the first image feature.
[0298] According to embodiments of this disclosure, the first obtaining submodule includes a second downsampling unit, a fourth obtaining unit, a third downsampling unit, and a fifth obtaining unit.
[0299] The second downsampling unit is used to downsample the first image features using the second downsampling layer to obtain the fourth intermediate feature;
[0300] The fourth obtaining unit is used to perform channel adjustment and feature extraction processing on the fourth intermediate feature using the third feature processing layer to obtain the second image feature;
[0301] The third downsampling unit is used to downsample the second image features using the third downsampling layer to obtain the fifth intermediate feature.
[0302] The fifth obtaining unit is used to perform channel adjustment and feature extraction processing on the fifth intermediate feature using the fourth feature processing layer to obtain the third image feature.
[0303] According to embodiments of this disclosure, the second output submodule includes a first output unit, a second output unit, and a third output unit.
[0304] The first output unit is used to process the first image features and the first transition features using the first multi-scale sub-model, and output a multi-scale feature matrix and a second transition feature.
[0305] The second output unit is used to process the second image features, the second transition features and the third transition features using the second multi-scale sub-model, and outputs a multi-scale feature matrix, the first transition feature and the fourth transition feature.
[0306] The third output unit is used to process the third image features and the fourth transition features using the third multi-scale sub-model, and outputs a multi-scale feature matrix and the third transition feature.
[0307] According to embodiments of this disclosure, the first output unit includes a first obtaining subunit, a second obtaining subunit, a third obtaining subunit, a fourth obtaining subunit, a first downsampling subunit, a fifth obtaining subunit, and a sixth obtaining subunit.
[0308] The first subunit is used to perform channel adjustment and feature extraction processing on the first image features and the first transition features respectively using two second convolutional normalization layers based on the second preset stride, to obtain the first channel features and the second channel features.
[0309] The second subunit is used to perform feature layer expansion processing on the second channel features using the first feature expansion layer to obtain the third channel features.
[0310] The third subunit is used to perform feature stacking processing on the first channel feature and the third channel feature using the first feature stacking layer to obtain the fourth channel feature.
[0311] The fourth sub-unit is used to perform channel adjustment and feature extraction on the fourth channel features using the fifth feature processing layer to obtain the fifth channel features, wherein the fifth channel features include two sub-channel features with a preset number of channels;
[0312] The first downsampling subunit is used to downsample the feature of a subchannel using the fourth downsampling layer to obtain the second transition feature;
[0313] The fifth sub-unit is used to perform convolution, normalization, and feature stacking on the features of another sub-channel using the first convolution stack layer, to obtain the features of the sixth channel.
[0314] The sixth sub-unit is used to perform channel adjustment and feature extraction processing on the sixth channel features using the third convolutional normalization layer based on the second preset stride, to obtain the first multi-scale feature matrix, wherein the first multi-scale feature matrix includes a first preset number of grids and the target number of channels.
[0315] According to embodiments of this disclosure, the second output unit includes a seventh obtaining subunit, an eighth obtaining subunit, a ninth obtaining subunit, a tenth obtaining subunit, an eleventh obtaining subunit, a twelfth obtaining subunit, a second downsampling subunit, a thirteenth obtaining subunit, and a fourteenth obtaining subunit.
[0316] The seventh subunit is used to perform channel adjustment and feature extraction on the second image features and the third transition features respectively using two fourth convolutional normalization layers based on the third preset stride, to obtain the seventh channel features and the eighth channel features.
[0317] The eighth subunit is used to perform feature layer expansion processing on the features of the eighth channel using the second feature expansion layer to obtain the features of the ninth channel.
[0318] The ninth subunit is used to perform feature stacking processing on the seventh and ninth channel features using the second feature stacking layer to obtain the tenth channel feature.
[0319] The tenth sub-unit is used to perform channel adjustment and feature extraction on the tenth channel features using the sixth feature processing layer to obtain the eleventh channel features.
[0320] The eleventh sub-unit is used to perform feature stacking processing on the eleventh channel feature and the second transition feature using the third feature stacking layer to obtain the twelfth channel feature.
[0321] The twelfth sub-unit is used to perform channel adjustment and feature extraction on the twelfth channel features using the seventh feature processing layer to obtain the thirteenth channel features;
[0322] The second downsampling subunit is used to downsample the thirteenth channel feature using the fifth downsampling layer to obtain the fourth transition feature;
[0323] The thirteenth sub-unit is used to perform convolution, normalization, and feature stacking on the thirteenth channel features using the second convolution stacking layer to obtain the fourteenth channel features.
[0324] The fourteenth sub-unit is used to perform channel adjustment and feature extraction on the features of the fourteenth channel using the fifth convolutional normalization layer based on the third preset step size, to obtain the second multi-scale feature matrix, wherein the second multi-scale feature matrix includes the second preset number of grids and the target number of channels.
[0325] According to embodiments of this disclosure, the third output unit includes a fifteenth obtaining subunit, a sixteenth obtaining subunit, a seventeenth obtaining subunit, an eighteenth obtaining subunit, and a nineteenth obtaining subunit.
[0326] The fifteenth subunit is used to perform feature extraction, pooling, and stacking on the third image features using the feature extraction stacking layer to obtain the third transition feature.
[0327] The sixteenth sub-unit is used to perform feature stacking processing on the third and fourth transition features using the fourth feature stacking layer to obtain the fifteenth channel feature.
[0328] The seventeenth sub-unit is used to perform channel adjustment and feature extraction on the fifteenth channel features using the eighth feature processing layer to obtain the sixteenth channel features;
[0329] The eighteenth subunit is used to perform convolution, normalization, and feature stacking on the sixteenth channel features using the third convolution stacking layer to obtain the seventeenth channel features;
[0330] The nineteenth sub-unit is used to perform channel adjustment and feature extraction on the features of the seventeenth channel using the sixth convolutional normalization layer based on the fourth preset step size, to obtain the third multi-scale feature matrix. The third multi-scale feature matrix includes the third preset number of grids and the target number of channels.
[0331] According to embodiments of this disclosure, the first classification module 1430 includes a processing submodule, a determination submodule, a first classification submodule, a second classification submodule, and a third classification submodule.
[0332] The processing submodule is used to process the position of the marker points in each prediction area based on the preset key point model to obtain the state parameters of each key point;
[0333] The determination submodule is used to determine the behavioral state of an object and the time and / or number of times it is in that behavioral state based on multiple state parameters;
[0334] The first classification submodule is used to classify training videos into the first classification sub-results when the behavior state belongs to one of the classification lists and the time or number of times meets the preset conditions.
[0335] The second classification submodule is used to classify training videos into the second classification sub-results when the behavior state belongs to one of the classification lists and the preset conditions are not met in terms of time or number of times.
[0336] The third classification submodule is used to classify training videos into third classification results when the behavior state does not belong to any of the classification lists. The behavior classification results include first classification results, second classification results, and third classification results.
[0337] Figure 15 A schematic block diagram of an object behavior classification apparatus according to an embodiment of the present disclosure is shown.
[0338] like Figure 15 As shown, the object behavior classification device 1500 of this embodiment includes a second acquisition module 1510 and a second classification module 1520.
[0339] The second acquisition module 1510 is used to acquire the video to be classified, wherein the video to be classified includes multiple image frames to be classified that are temporally associated.
[0340] The second classification module 1520 is used to input multiple image frames to be classified from the video to be classified into the object behavior classification model and output the predicted behavior classification result, wherein the predicted behavior classification result represents the behavior posture of the object when the object exists in the video to be classified.
[0341] According to embodiments of this disclosure, a multi-scale feature matrix is extracted from the image frame to be classified using an object behavior classification model, so as to determine the prediction region where the object is located as much as possible. The object behavior within the prediction region is classified using the marker points of key points, thereby obtaining the classification result of the object's behavior posture. Since the object behavior classification model can continuously compress the image size and increase the number of image channels during the generation of the multi-scale feature matrix, and fuse different images to obtain the multi-scale feature matrix, the classification of object behavior using this multi-scale feature matrix can obtain a more accurate classification result, avoiding the problem of inaccurate identification of object behavior caused by the low classification accuracy of related technologies.
[0342] According to embodiments of this disclosure, the object behavior classification device further includes a determination module and a display module.
[0343] The determination module is used to determine the target information corresponding to the preset behavior posture from the information list when the predicted behavior classification results indicate that the object's behavior posture belongs to the preset behavior posture.
[0344] The presentation module is used to display target information to objects in a visual format.
[0345] According to embodiments of this disclosure, any multiple modules among the first acquisition module 1410, multi-scale module 1420, first classification module 1430, loss module 1440, iteration module 1450, or second acquisition module 1510, second classification module 1520, can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the first acquisition module 1410, multi-scale module 1420, first classification module 1430, loss module 1440, iteration module 1450, or second acquisition module 1510, second classification module 1520 can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or any other reasonable means of integrating or packaging circuits, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the first acquisition module 1410, multi-scale module 1420, first classification module 1430, loss module 1440, iteration module 1450, or second acquisition module 1510, second classification module 1520 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0346] Figure 16 A schematic block diagram of a vehicle monitoring system according to an embodiment of the present disclosure is shown.
[0347] like Figure 16 As shown, the vehicle monitoring system 1600 of this embodiment includes:
[0348] Server 1610 is configured as follows:
[0349] The in-vehicle video is processed based on the object behavior classification model to obtain the in-vehicle behavior classification result, wherein the in-vehicle behavior classification result represents the behavior posture of at least one object in the in-vehicle video.
[0350] If the in-vehicle behavior classification results indicate that the object’s behavior posture is a violation posture, the first alarm information corresponding to the violation posture is determined from the alarm information list and transmitted to the alarm device.
[0351] Transport vehicle 1620, which includes:
[0352] Vehicle body;
[0353] The image acquisition device is configured to acquire in-vehicle video of the transport vehicle 1620 in real time and transmit it to the server 1610 while the main body of the vehicle is in motion. The in-vehicle video includes multiple in-vehicle images that are correlated in time.
[0354] The alarm device is configured to display the first alarm information to the object in a visual form.
[0355] According to embodiments of this disclosure, in-vehicle video can be uploaded to server 1610 via in-vehicle terminal 1630. Server 1610 transmits the classification results to big data cloud 1642 via data center 1641 in platform sharing unit 1640, thereby enabling application supervision unit 1650 to display in-vehicle video and classification results when objects exhibit illegal phase poses in real time through various display devices.
[0356] According to embodiments of this disclosure, a multi-scale feature matrix is extracted from image frames of in-vehicle video using an object behavior classification model to determine the prediction area where the object is located as much as possible. The object behavior within the prediction area is classified using key point markers to determine whether the object inside the transport vehicle 1620 exhibits any illegal behavior or posture. An alarm message is then sent to the object via an alarm device. Since the object behavior inside the transport vehicle 1620 is detected in real time during the transport vehicle's operation, the safety and confidentiality of the transport vehicle 1620 during the transport mission are improved, avoiding the problem of reduced safety and confidentiality caused by illegal behavior of the object during the execution of the transport mission.
[0357] According to embodiments of this disclosure, the vehicle monitoring system 1600 further includes:
[0358] The sensor module is configured as follows:
[0359] The system collects vehicle status parameters of transport vehicle 1620 and transmits the vehicle status parameters to server 1610. The vehicle status parameters include at least one of the following: tire pressure, vehicle interior temperature, engine status, range, vehicle speed, engine speed, and driving trajectory.
[0360] The server 1610 is also configured to transmit the second alarm information corresponding to any of the vehicle status parameters to the alarm device when any of the parameters meets the alarm conditions.
[0361] The alarm device is also configured to display a second alarm message to the target in a visual format.
[0362] According to embodiments of this disclosure, the vehicle monitoring system 1600 further includes a platform sharing unit 1640 and an application supervision unit 1650. The application supervision unit 1650 can communicate bidirectionally with the server 1610 and the transport vehicle 1620 in real time through the platform sharing unit 1640, thereby displaying in-vehicle video and vehicle status parameters in real time. This allows supervisors to understand the driving safety of the transport vehicle 1620 in a timely manner, ensuring monitoring and dispatching of the transport vehicle 1620 in the event of an emergency, and effectively guaranteeing transportation safety, standardization, and privacy.
[0363] Figure 17 A block diagram schematically illustrates an electronic device suitable for implementing the above-described method according to an embodiment of the present disclosure.
[0364] like Figure 17 As shown, an electronic device 1700 according to an embodiment of the present disclosure includes a processor 1701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1702 or a program loaded from a storage portion 1708 into a random access memory (RAM) 1703. The processor 1701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1701 may also include onboard memory for caching purposes. The processor 1701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0365] RAM 1703 stores various programs and data required for the operation of electronic device 1700. Processor 1701, ROM 1702, and RAM 1703 are interconnected via bus 1704. Processor 1701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 1702 and / or RAM 1703. It should be noted that the programs may also be stored in one or more memories other than ROM 1702 and RAM 1703. Processor 1701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0366] According to embodiments of this disclosure, the electronic device 1700 may further include an input / output (I / O) interface 1705, which is also connected to a bus 1704. The electronic device 1700 may also include one or more of the following components connected to the input / output (I / O) interface 1705: an input section 1706 including a keyboard, mouse, etc.; an output section 1707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1708 including a hard disk, etc.; and a communication section 1709 including a network interface card such as a LAN card, modem, etc. The communication section 1709 performs communication processing via a network such as the Internet. A drive 1710 is also connected to the input / output (I / O) interface 1705 as needed. A removable medium 1711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 1710 as needed so that computer programs read from it can be installed into the storage section 1708 as needed.
[0367] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0368] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 1702 and / or RAM 1703 and / or one or more memories other than ROM 1702 and RAM 1703 described above.
[0369] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this disclosure.
[0370] When the computer program is executed by the processor 1701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0371] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 1709, and / or installed from a removable medium 1711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0372] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1709, and / or installed from the removable medium 1711. When the computer program is executed by the processor 1701, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0373] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0374] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0375] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0376] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A training method for an object behavior classification model, comprising: Obtain a training set, wherein the training set includes multiple training videos and classification labels, and the videos include multiple image frames that are temporally correlated; The training video is input into the initial behavior classification model, and multiple multi-scale feature matrices corresponding to the training video are output. Each multi-scale feature matrix includes multiple prediction regions including different key points of the object. Based on the marker points of the training video, behavior classification processing is performed on multiple prediction regions corresponding to the training video to obtain the behavior classification result of the training video, wherein the marker points represent the positions of different key points of the object in each image frame; Input the classification result and classification label corresponding to each training video into the loss function, and output the loss result; The network parameters of the initial behavior classification model are iteratively adjusted based on the loss result to generate a trained object behavior classification model; Where the number of multi-scale feature matrices is three, the step of inputting the training video into the initial behavior classification model and outputting multiple multi-scale feature matrices corresponding to the training video includes: Based on a first preset step size, a feature extraction sub-model is used to perform channel adjustment and feature extraction processing on multiple image frames to obtain the first image features; The first image features are processed using a channel adjustment sub-model to obtain the second and third image features; The first image feature, the second image feature, and the third image feature are processed using the first multi-scale sub-model, the second multi-scale sub-model, and the third multi-scale sub-model, respectively, to obtain three multi-scale feature matrices. Specifically, the first, second, and third image features are processed using the first, second, and third multi-scale sub-models, respectively, resulting in three multi-scale feature matrices: The first multi-scale sub-model is used to process the first image features and the first transition features, and outputs a multi-scale feature matrix and a second transition feature. The second transition feature is obtained by sequentially performing channel adjustment, feature extraction, feature layer expansion, feature stacking and downsampling on the first image features and the first transition features of the first multi-scale sub-model. The second image feature, the second transition feature, and the third transition feature are processed using the second multi-scale sub-model to output a multi-scale feature matrix, a first transition feature, and a fourth transition feature. The first transition feature is obtained by sequentially performing channel adjustment, feature extraction, feature layer expansion, and feature stacking on the second image feature and the third transition feature using the second multi-scale sub-model. The fourth transition feature is obtained by performing feature stacking, channel adjustment, feature extraction, and downsampling on the first transition feature and the second transition feature. The third image features and the fourth transition features are processed using the third multi-scale sub-model, and a multi-scale feature matrix and the third transition features are output. The third transition features are obtained by performing feature extraction, pooling and stacking on the third image features using the third multi-scale sub-model.
2. The method according to claim 1, wherein, The first image feature is obtained by performing channel adjustment and feature extraction processing on multiple image frames based on a first preset step size using a feature extraction sub-model, including: Multiple first convolutional normalization layers are used to perform channel adjustment and feature extraction processing on multiple image frames to obtain first intermediate features, wherein one of the convolutional normalization layers corresponds to a first preset stride. The first intermediate feature is processed by channel adjustment and feature stacking using the first feature processing layer to obtain the second intermediate feature; The second intermediate feature is downsampled using the first downsampling layer to obtain the third intermediate feature; The third intermediate feature is processed by a second feature processing layer to perform channel adjustment and feature stacking to obtain the first image feature.
3. The method according to claim 1, wherein, The process of using a channel adjustment sub-model to process the first image features to obtain second and third image features includes: The first image features are downsampled using a second downsampling layer to obtain a fourth intermediate feature. The third feature processing layer is used to perform channel adjustment and feature extraction on the fourth intermediate feature to obtain the second image feature; The second image features are downsampled using a third downsampling layer to obtain the fifth intermediate feature; The fifth intermediate feature is processed by the fourth feature processing layer to perform channel adjustment and feature extraction to obtain the third image feature.
4. The method according to claim 1, wherein, The step of processing the first image features and the first transition features using the first multi-scale sub-model to output a multi-scale feature matrix and a second transition feature includes: Based on the second preset stride, two second convolutional normalization layers are used to perform channel adjustment and feature extraction on the first image features and the first transition features respectively, to obtain the first channel features and the second channel features. The second channel feature is expanded using the first feature expansion layer to obtain the third channel feature; The first channel feature and the third channel feature are stacked using a first feature stacking layer to obtain the fourth channel feature; The fifth feature processing layer is used to perform channel adjustment and feature extraction processing on the fourth channel feature to obtain the fifth channel feature, wherein the fifth channel feature includes two sub-channel features with a preset number of channels; The second transition feature is obtained by downsampling one of the sub-channel features using the fourth downsampling layer. The first convolutional stacking layer is used to perform convolution, normalization, and feature stacking on the feature of another sub-channel to obtain the feature of the sixth channel; Based on the second preset stride, the sixth channel features are processed by the third convolutional normalization layer to perform channel adjustment and feature extraction, resulting in the first multi-scale feature matrix, wherein the first multi-scale feature matrix includes a first preset number of grids and a target number of channels.
5. The method according to claim 1 or 4, wherein, The process of using the second multi-scale sub-model to process the second image features, the second transition features, and the third transition features, and outputting a multi-scale feature matrix, the first transition features, and the fourth transition features, includes: Based on the third preset stride, two fourth convolutional normalization layers are used to perform channel adjustment and feature extraction on the second image features and the third transition features respectively, to obtain the seventh channel features and the eighth channel features. The eighth channel feature is expanded using a second feature expansion layer to obtain the ninth channel feature. The seventh channel feature and the ninth channel feature are stacked using a second feature stacking layer to obtain the tenth channel feature. The eleventh channel feature is obtained by performing channel adjustment and feature extraction on the tenth channel feature using the sixth feature processing layer. The eleventh channel feature and the second transition feature are stacked using a third feature stacking layer to obtain the twelfth channel feature. The twelfth channel feature is processed by the seventh feature processing layer to perform channel adjustment and feature extraction, resulting in the thirteenth channel feature. The thirteenth channel feature is downsampled using the fifth downsampling layer to obtain the fourth transition feature; The thirteenth channel features are obtained by performing convolution, normalization, and feature stacking on the second convolution stacking layer; Based on the third preset step size, the fourteenth channel features are processed by the fifth convolutional normalization layer for channel adjustment and feature extraction to obtain the second multi-scale feature matrix, wherein the second multi-scale feature matrix includes a second preset number of grids and a target number of channels.
6. The method according to claim 1, wherein, The process of using the third multi-scale sub-model to process the third image feature and the fourth transition feature, and outputting a multi-scale feature matrix and the third transition feature, includes: The third image features are extracted, pooled, and stacked using a feature extraction stacking layer to obtain the third transition feature. The third transition feature and the fourth transition feature are stacked using a fourth feature stacking layer to obtain the fifteenth channel feature. The 15th channel feature is processed by the 8th feature processing layer to perform channel adjustment and feature extraction, resulting in the 16th channel feature. The sixteenth channel feature is processed by convolution, normalization and feature stacking using the third convolution stacking layer to obtain the seventeenth channel feature; Based on the fourth preset step size, the features of the seventeenth channel are processed by the sixth convolutional normalization layer to perform channel adjustment and feature extraction, resulting in the third multi-scale feature matrix. The third multi-scale feature matrix includes a third preset number of grids and a target number of channels.
7. The method according to claim 1, wherein, The step of performing behavior classification processing on multiple predicted regions corresponding to the training video based on the labeled points of the training video to obtain the behavior classification result of the training video includes: The positions of the marked points in each of the predicted regions are processed based on the preset key point model to obtain the state parameters of each key point; Based on multiple state parameters, the behavioral state of the object and the time and / or number of times it is in the behavioral state are determined; If the behavior state belongs to one of the classification lists and the time or number of times meets the preset conditions, the training video is classified as the first classification sub-result; If the behavior state belongs to one of the categories in the classification list and the time or number of times does not meet the preset conditions, the training video will be classified as a second category sub-result. If the behavior state does not belong to any of the classification lists, the training video is classified into a third sub-classification result, wherein the behavior classification result includes the first sub-classification result, the second sub-classification result, and the third sub-classification result.
8. A method for classifying object behavior, comprising: Obtain the video to be classified, wherein the video to be classified includes multiple image frames to be classified that are temporally associated; Multiple image frames to be classified from the video to be classified are input into the object behavior classification model, and the predicted behavior classification result is output. The predicted behavior classification result represents the behavior posture of the object when the object exists in the video to be classified. The object behavior classification model is trained using the method described in any one of claims 1 to 7.
9. The method according to claim 8, further comprising: If the predicted behavior classification result indicates that the object's behavior posture belongs to a preset behavior posture, the target information corresponding to the preset behavior posture is determined from the information list. The target information is presented to the object in a visual format.
10. A method for detecting the driving safety of a transport vehicle, comprising: While the transport vehicle is in motion, the in-vehicle video of the transport vehicle is captured in real time using the image acquisition device of the transport vehicle, wherein the in-vehicle video includes multiple in-vehicle images that are sequentially related. Multiple in-vehicle images from the in-vehicle video are transmitted to a server, so that the server processes the in-vehicle video based on an object behavior classification model to obtain an in-vehicle behavior classification result, wherein the in-vehicle behavior classification result represents the behavioral posture of at least one object in the in-vehicle video, and the object behavior classification model is trained using the method described in any one of claims 1 to 7. If the in-vehicle behavior classification result indicates that the object's behavior posture is a violation posture, the first alarm information corresponding to the violation posture is determined from the alarm information list and transmitted to the transport vehicle. The first alarm information is displayed to the object in a visual form.
11. The method of claim 10, further comprising: The vehicle status parameters of the transport vehicle are collected using a sensor module, wherein the vehicle status parameters include at least one of the following: tire pressure, interior temperature, engine status, range, vehicle speed, engine speed, and driving trajectory. The vehicle status parameters are transmitted to the server so that if any one of the vehicle status parameters meets the alarm condition, the server transmits the second alarm information corresponding to the parameter to the transport vehicle. The second alarm information is displayed to the object in a visual form.
12. A training device for an object behavior classification model, comprising: The first acquisition module is used to acquire a training set, wherein the training set includes multiple training videos and classification labels, and the videos include multiple image frames that are temporally correlated. A multi-scale module is used to input the training video into an initial behavior classification model and output multiple multi-scale feature matrices corresponding to the training video, wherein each multi-scale feature matrix includes multiple prediction regions including different key points of the object; The first classification module is used to perform behavior classification processing on multiple prediction regions corresponding to the training video based on the marker points of the training video to obtain the behavior classification result of the training video, wherein the marker points represent the positions of different key points of the object in each image frame; The loss module is used to input the classification result and classification label corresponding to each training video into the loss function and output the loss result. An iterative module is used to iteratively adjust the network parameters of the initial behavior classification model based on the loss result, thereby generating a trained object behavior classification model; In the case where there are three multi-scale feature matrices, the multi-scale module includes a feature extraction submodule, a first acquisition submodule, and a second acquisition submodule: The feature extraction submodule is used to perform channel adjustment and feature extraction processing on multiple image frames based on a first preset step size, and to obtain the first image features. The first submodule is used to process the first image features using the channel adjustment submodel to obtain the second and third image features. The second sub-module is used to process the first image features, the second image features, and the third image features using the first multi-scale sub-model, the second multi-scale sub-model, and the third multi-scale sub-model, respectively, to obtain three multi-scale feature matrices. The second sub-module includes: The first output unit is used to process the first image features and the first transition features using the first multi-scale sub-model, and output a multi-scale feature matrix and a second transition feature. The second transition feature is obtained by sequentially performing channel adjustment, feature extraction, feature layer expansion, feature stacking and downsampling on the first image features and the first transition feature using the first multi-scale sub-model. The second output unit is used to process the second image features, the second transition features, and the third transition features using the second multi-scale sub-model, and output a multi-scale feature matrix, a first transition feature, and a fourth transition feature. The first transition feature is obtained by sequentially performing channel adjustment, feature extraction, feature layer expansion, and feature stacking on the second image features and the third transition features using the second multi-scale sub-model. The fourth transition feature is obtained by performing feature stacking, channel adjustment, feature extraction, and downsampling on the first transition feature and the second transition feature. The third output unit is used to process the third image features and the fourth transition features using the third multi-scale sub-model, and output a multi-scale feature matrix and the third transition feature. The third transition feature is obtained by performing feature extraction, pooling and stacking processing on the third image features using the third multi-scale sub-model.
13. An object behavior classification device, comprising: The second acquisition module is used to acquire the video to be classified, wherein the video to be classified includes multiple image frames to be classified that are temporally associated. The second classification module is used to input multiple image frames to be classified from the video to be classified into the object behavior classification model and output the predicted behavior classification result, wherein the predicted behavior classification result represents the behavior posture of the object when there is an object in the video to be classified. The object behavior classification model is trained using the method described in any one of claims 1 to 7.
14. A vehicle monitoring system, comprising: The server is configured to: The in-vehicle video is processed based on the object behavior classification model to obtain the in-vehicle behavior classification result, wherein the in-vehicle behavior classification result represents the behavior posture of at least one object in the in-vehicle video, and the object behavior classification model is trained using the method as described in any one of claims 1 to 7. If the in-vehicle behavior classification result indicates that the object's behavior posture belongs to a violation behavior posture, the first alarm information corresponding to the violation behavior posture is determined from the alarm information list and transmitted to the alarm device. The transport vehicle includes: Vehicle body; An image acquisition device is configured to acquire and transmit in real-time in-vehicle video of the transport vehicle to the server while the main body of the vehicle is in motion, wherein the in-vehicle video includes multiple in-vehicle images that are sequentially correlated. The alarm device is configured to display the first alarm information to the object in a visual form.
15. The system of claim 14, further comprising: The sensor module is configured as follows: The vehicle status parameters of the transport vehicle are collected and transmitted to the server. The vehicle status parameters include at least one of the following: tire pressure, interior temperature, engine status, range, vehicle speed, engine speed, and driving trajectory. The server is also configured to transmit a second alarm message corresponding to any of the vehicle status parameters to the alarm device when any of the parameters meets the alarm conditions. The alarm device is also configured to display the second alarm information to the object in a visual form.
16. An electronic device comprising: One or more processors; Storage device for storing one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1 to 9.
17. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 9.
18. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Monitoring method and device for judging abnormal behavior based on deep learning, computer equipment and storage medium
CN113989540A
Branch site selection prediction method, device, equipment, medium and program product
CN114663165A