State recognition method for target object and training method for deep learning model

By processing video data and using deep learning models to identify the motion state and occlusion state of the target object, the problem of poor recognition effect in the prior art is solved, and a state recognition effect with high accuracy and robustness is achieved.

CN114998275BActive Publication Date: 2025-06-10BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210658887.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-09
Publication Date
2025-06-10
Estimated Expiration
2042-06-09

AI Technical Summary

Technical Problem

When identifying the motion state and occlusion state of the target object, the prior art is affected by external factors, the recognition effect is poor, the robustness is poor, and it is difficult to achieve continuous iterative training.

Method used

By processing video data associated with the target object, the image data to be identified containing time information and image data is extracted, and the image data is recognized using a deep learning model. At the same time, by processing historical video data, sample image data is generated, including motion state labels and occlusion state labels, and deep learning models are trained to improve recognition accuracy.

Benefits of technology

It realizes high accuracy identification of the moving state and occlusion state of the target object, improves the robustness and applicability of the identification system, and can work effectively in complex and diverse scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998275B_ABST
    Figure CN114998275B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for identifying the state of a target object, a method for training a deep learning model, an apparatus, a device, a medium, and a product, which relate to the field of artificial intelligence, specifically to technical fields such as computer vision, image processing, and deep learning, and are applicable to the scenario of identifying the state of a target object of industrial equipment. The method for identifying the state of a target object includes: processing video data associated with the target object to obtain image data to be recognized, where the image data to be recognized includes time information associated with the video data and image data associated with the target object; and identifying the image data to be recognized to obtain the motion state of the target object and the occlusion state of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, specifically to technical fields such as computer vision, image processing, and deep learning. More specifically, it relates to a method for recognizing the state of a target object, a method for training a deep learning model, an apparatus, an electronic device, a medium, and a program product. Background Art

[0002] In some scenarios, it is necessary to recognize the relevant state of a target object. The relevant state indicates whether the target object is in motion or at rest, or whether the target object is occluded by other objects. However, in the related art, due to the influence of external factors, the recognition effect is not good when recognizing the relevant state of the target object. Summary of the Invention

[0003] The present disclosure provides a method for recognizing the state of a target object, a method for training a deep learning model, an apparatus, an electronic device, a storage medium, and a program product.

[0004] According to one aspect of the present disclosure, there is provided a method for recognizing the state of a target object, including: processing video data associated with the target object to obtain image data to be recognized, where the image data to be recognized includes time information associated with the video data and image data associated with the target object; recognizing the image data to be recognized to obtain the motion state of the target object and the occlusion state of the target object.

[0005] According to another aspect of the present disclosure, there is provided a method for training a deep learning model, including: processing historical video data associated with the target object to obtain sample image data, where the sample image data includes time information associated with the historical video data and image data associated with the target object, and the sample image data includes a motion state label and an occlusion state label; using the deep learning model to be trained to process the sample image data to obtain the motion state recognition result of the target object and the occlusion state recognition result of the target object; training the deep learning model to be trained based on the motion state recognition result, the motion state label, the occlusion state recognition result, and the occlusion state label.

[0006] According to another aspect of the present disclosure, there is provided a device for recognizing the state of a target object, including: a processing module and a recognition module. The processing module is configured to process video data associated with the target object to obtain image data to be recognized, where the image data to be recognized includes time information associated with the video data and image data associated with the target object; the recognition module is configured to recognize the image data to be recognized to obtain the motion state of the target object and the occlusion state of the target object.

[0007] According to another aspect of the present disclosure, there is provided a training apparatus for a deep learning model, including: a first processing module, a second processing module, and a training module. The first processing module is configured to process historical video data associated with a target object to obtain sample image data, where the sample image data includes time information associated with the historical video data and image data associated with the target object, and the sample image data includes a motion state label and an occlusion state label; the second processing module is configured to process the sample image data by using a deep learning model to be trained to obtain a motion state recognition result of the target object and an occlusion state recognition result of the target object; and the training module is configured to train the deep learning model to be trained based on the motion state recognition result, the motion state label, the occlusion state recognition result, and the occlusion state label.

[0008] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor and a memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned state recognition method of the target object and / or the training method of the deep learning model.

[0009] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause the computer to execute the above-mentioned state recognition method of the target object and / or the training method of the deep learning model.

[0010] According to another aspect of the present disclosure, there is provided a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the above-mentioned state recognition method of the target object and / or the steps of the training method of the deep learning model are implemented.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0013] Figure 1 Schematically shows a system architecture of a method for state recognition of a target object and / or training of a deep learning model according to an embodiment of the present disclosure;

[0014] Figure 2Schematically shows a flowchart of a method for identifying the state of a target object according to an embodiment of the present disclosure;

[0015] Figure 3 Schematically shows a schematic diagram of a method for identifying the state of a target object according to an embodiment of the present disclosure;

[0016] Figure 4 Schematically shows a flowchart of a method for training a deep learning model according to an embodiment of the present disclosure;

[0017] Figure 5 Schematically shows a schematic diagram of a deep learning model according to an embodiment of the present disclosure;

[0018] Figure 6 Schematically shows a schematic diagram of a method for identifying the state of a target object and a method for training a deep learning model according to an embodiment of the present disclosure;

[0019] Figure 7 Schematically shows a block diagram of a device for identifying the state of a target object according to an embodiment of the present disclosure;

[0020] Figure 8 Schematically shows a block diagram of a device for training a deep learning model according to an embodiment of the present disclosure; and

[0021] Figure 9 Is a block diagram of an electronic device for implementing the embodiments of the present disclosure for performing the method of identifying the state of a target object and / or training a deep learning model. Detailed Description of the Invention

[0022] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0025] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0026] With the vigorous promotion of smart mines, technologies such as artificial intelligence and image algorithms can be used to intelligently monitor the production status of coal mines, so as to timely discover abnormal dynamics of coal mines and automatically generate and push alarm information. In addition, the data analysis results for coal mines can be transferred to the comprehensive information system through the safety supervision department, thereby realizing all-weather remote detection.

[0027] The construction of smart mines covers a variety of intelligent tasks, such as automatic identification of personnel entering the mine, camera occlusion, automatic identification of camera movement angles, identification of the operating status of transportation equipment, etc. Take the task of identifying the motion status of transportation equipment as an example. For example, it is necessary to determine whether the transportation equipment is in motion or stationary state, and at the same time, it is necessary to identify whether the transportation equipment is unloaded. Transportation equipment includes conveyor belts, for example. However, since the mining scene is easily affected by factors such as lighting, reflection, camera angle of view, and movement of people, the existing recognition algorithms are difficult to apply, the effect is unstable, the robustness is poor, and it is impossible to continuously iterate.

[0028] In one example, it is possible to determine whether the conveyor belt is moving by an external motion detection device or by the working current of the conveyor belt. However, by introducing an additional device to determine the state of the conveyor belt, the equipment cost is increased.

[0029] In another example, the motion state of the conveyor belt can be identified by traditional image processing methods. For example, by collecting video data for the conveyor belt, the background difference of adjacent video frames in the video data is obtained to obtain foreground information. Then, the single-threshold maximum inter-class variance method (OTSU) is used for binarization, morphological processing and other methods to obtain the approximate motion area in the video frame and the ratio between the area of ​​the motion area and the total area of ​​the video frame, and the ratio is compared with the pre-set area ratio to identify whether the conveyor belt is moving. The maximum inter-class variance method is a method for automatically obtaining thresholds that is suitable for bimodal situations proposed by Japanese scholar Nobuyuki Otsu in 1979, also known as the Otsu method. However, this method is difficult to solve problems such as large differences in camera viewing angles and reflection effects caused by changes in illumination, making the robustness of this method poor.

[0030] In another example, traditional image feature extraction methods such as the Perceptual hash algorithm (phash), Scale-invariant feature transform (SIFT), and Histogram of Oriented Gradient (HOG) can be used to extract the features of adjacent video frames in video data, calculate the distance between the features, and compare the distance with a set distance threshold to obtain the motion state of the conveyor belt. The extraction ability of video frame features can also be optimized through Convolutional Neural Networks (CNN) technology. However, this method is limited by the quality of the extracted image features, and the image features extracted by traditional algorithms have weak representativeness and poor robustness. Therefore, it is extremely vulnerable to the image differences caused by small image perturbations and is not suitable for complex and diverse mining area scenarios. Even by introducing a deep neural network to improve the image feature extraction ability, it is currently very difficult to achieve training optimization in actual mining area scenarios. In addition, when processing video data, this method only determines the feature differences between two frames and does not fully utilize the spatio-temporal information of the video. Therefore, its real-time performance and applicability are poor.

[0031] The above-mentioned various methods usually separate the task of identifying the motion state of the conveyor belt from the task of identifying whether the conveyor belt is empty (whether it is transporting coal), and process these two tasks separately, increasing the algorithm time consumption. In addition, the above methods are difficult to achieve continuous iterative training. Therefore, due to the influence of camera perspective, lighting, etc. in diverse mining area scenarios, it is difficult to achieve self-iteration of the algorithm, making it difficult to guarantee the generalization performance of the algorithm.

[0032] In view of this, the embodiments of the present disclosure provide an optimized method for identifying the state of a target object and a method for training a deep learning model.

[0033] Exemplarily, the method for identifying the state of a target object includes: processing video data associated with the target object to obtain image data to be recognized, where the image data to be recognized includes time information associated with the video data and image data associated with the target object. Then, the image data to be recognized is recognized to obtain the motion state of the target object and the occlusion state of the target object.

[0034] Exemplarily, a method for training a deep learning model includes: processing historical video data associated with a target object to obtain sample image data, where the sample image data includes time information associated with the historical video data and image data associated with the target object, and the sample image data includes a motion state label and an occlusion state label. Then, using the deep learning model to be trained to process the sample image data to obtain a motion state recognition result of the target object and an occlusion state recognition result of the target object. Next, based on the motion state recognition result, the motion state label, the occlusion state recognition result, and the occlusion state label, training the deep learning model to be trained.

[0035] Figure 1 Schematically shows a system architecture for state recognition of a target object and / or a method for training a deep learning model according to an embodiment of the present disclosure. It should be noted that, Figure 1 The shown is only an example of a system architecture to which the embodiments of the present disclosure can be applied to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0036] As Figure 1 shown, the system architecture 100 according to this embodiment may include a video data acquisition device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the video data acquisition device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0037] The video data acquisition device 101 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc. The video data acquisition device 101 of the embodiments of the present disclosure may, for example, run an application program.

[0038] The server 103 may be a server providing various services. For example, the server 103 may be a cloud server, that is, the server 103 has cloud computing capabilities.

[0039] It should be noted that the method for state recognition of a target object and / or the method for training a deep learning model provided by the embodiments of the present disclosure may be executed by the server 103. Correspondingly, the device for state recognition of a target object and / or the device for training a deep learning model provided by the embodiments of the present disclosure may be provided in the server 103.

[0040] In one example, for the historical video data associated with the target object collected by the video data acquisition device 101, the server 103 may train a deep learning model based on the historical video data.

[0041] In one example, the server 103 may process the video data collected by the video data acquisition device 101 associated with the target object, so as to obtain the motion state of the target object and the occlusion state of the target object. For example, the server 103 may process the video data by using a trained deep learning model, so as to obtain the motion state of the target object and the occlusion state of the target object.

[0042] It should be understood that Figure 1 the numbers of the video data acquisition device, the network, and the server in

[0043] are merely illustrative. According to the implementation requirements, there may be any number of video data acquisition devices, networks, and servers. Figure 1 The following describes the method for recognizing the state of a target object and / or the method for training a deep learning model according to an exemplary embodiment of the present disclosure in conjunction with Figures 2 to 6 the system architecture of Figure 1 and with reference to Figure 1 The server shown in

[0044] Figure 2 schematically shows a flowchart of the method for recognizing the state of a target object according to an embodiment of the present disclosure.

[0045] As Figure 2 shown, the method 200 for recognizing the state of a target object according to an embodiment of the present disclosure may include, for example, operation S210 to operation S220.

[0046] In operation S210, the video data associated with the target object is processed to obtain image data to be recognized, and the image data to be recognized includes time information associated with the video data and image data associated with the target object.

[0047] In operation S220, the image data to be recognized is recognized to obtain the motion state of the target object and the occlusion state of the target object.

[0048] Exemplarily, the video data associated with the target object, for example, indicates that the video data is data collected for the target object, and the video data includes relevant information of the target object. Alternatively, the video data associated with the target object may also indicate that the video data is data collected for other objects by using the target object, and the video data includes relevant information of the other objects.

[0049] Video data usually includes multiple consecutive video frames, and the multiple consecutive video frames include time information. By processing the video data, image data to be recognized can be obtained, and the image data to be recognized can include at least a part of the multiple consecutive video frames. The image data to be recognized includes time information and image data associated with the target object.

[0050] After obtaining the image data to be recognized, the image data to be recognized can be subjected to image recognition to simultaneously obtain the motion state and occlusion state of the target object. The motion state indicates whether the target object is moving, and the occlusion state indicates whether the target object is occluded.

[0051] According to an embodiment of the present disclosure, by processing video data, image data to be recognized including time information and image information can be obtained, and by recognizing the image data to be recognized, the motion state and occlusion state of the target object can be simultaneously obtained. Thus, after obtaining the image data to be recognized including time information and image information by processing video data, based on the image data to be recognized, the motion state recognition task and the occlusion state recognition task can be simultaneously executed, reducing the processing time of the tasks and achieving the effect of multi-task processing.

[0052] In an example of the present disclosure, the target object includes, for example, a conveying device, and the video data includes data obtained by video capturing of the conveying device. The motion state characterizes whether the conveying device is moving. The occlusion state characterizes whether there is an object to be conveyed on the conveying device. The object to be conveyed includes, for example, coal, and the occlusion state includes, for example, whether there is coal on the conveyor belt. If there is coal on the conveyor belt, the occlusion state is occluded at this time, and if there is no coal on the conveyor belt, the occlusion state is unoccluded at this time.

[0053] In another example of the present disclosure, the target object includes, for example, a video data acquisition device, and the video data includes data obtained by using the video data acquisition device to perform video capturing on other objects. The motion state characterizes whether the video data acquisition device is moving, for example, characterizes whether the video data acquisition device has jitter, rotation, etc. The video data acquisition device includes, for example, a camera. The occlusion state characterizes whether the video data acquisition device is occluded, for example, characterizes whether the front of the camera is occluded when collecting data. In one example, when there is a stain on the camera, it indicates that the camera is occluded.

[0054] Due to the differences in the target object, the state recognition method of the target object in the embodiments of the present disclosure is applicable to various different scenarios, improving the applicable range of the state recognition method of the target object.

[0055] Figure 3 Schematically shows the principle diagram of the state recognition method of the target object according to an embodiment of the present disclosure.

[0056] As shown Figure 3 in the figure, multiple frames of images can be extracted from the video data 310, and then the multiple frames of images can be processed in various ways to obtain the image data to be recognized that includes multiple channels. The image data to be recognized includes, for example, the first image data to be recognized 341, the second image data to be recognized 342, and the third image data to be recognized 343. The first image data to be recognized 341 includes, for example, three first channels, the second image data to be recognized 342 includes, for example, three second channels, and the third image data to be recognized 343 includes, for example, three third channels.

[0057] In one way, the multiple frames of images include, for example, the first frame of image 321 and the second frame of image 322. The acquisition time of the second frame of image 322 is, for example, after the acquisition time of the first frame of image 321. For example, the second frame of image 322 is the current frame of image, and the first frame of image 321 is the previous frame of the current frame of image.

[0058] Based on the first frame of image 321 and the second frame of image 322, a first difference image 331 is obtained. For example, the pixels of the first frame of image 321 are subtracted from the corresponding pixels in the second frame of image 322 one by one, or the pixels of a local area in the first frame of image 321 are subtracted from the corresponding local pixels in the second frame of image 322 one by one, so as to obtain the first difference image 331. The local area of the image includes the area where the target object is located or an important area. In the figure, the local area of the image is indicated by a dashed box, and the dashed box can also be called an electronic fence. The first difference image 331 contains the differences between the first frame of image 321 and the second frame of image 322. Therefore, the first difference image 331 contains the time information associated with the video data 310.

[0059] Next, the first frame of image 321, the second frame of image 322, and the first difference image 331 are combined to obtain the first image data to be recognized 341 that includes three first channels.

[0060] It can be understood that the embodiments of the present disclosure are combined with the constraint of the electronic fence area. When the deep learning model 350 infers and predicts the motion state and occlusion state of the current frame of image, the current frame of image and the previous frame of image are cropped according to the electronic fence coordinates, a difference image is obtained based on the two frames of images, and the two frames of images and the difference image are merged to obtain the first image data to be recognized 341 that includes three first channels and is input into the deep learning model 350 for recognition. At the same time, by setting a lower frame extraction frequency, such as extracting 1 frame per second, while ensuring performance, the video processing speed is improved, which is convenient for the deployment of the deep learning model 350 at the edge.

[0061] In another way, the multi-frame images include, for example, a third frame image 323 and a fourth frame image 324. The acquisition time of the fourth frame image 324 is, for example, after the acquisition time of the third frame image 323. For example, the fourth frame image 324 is the current frame image, and the third frame image 323 is the previous frame image of the current frame image.

[0062] A second difference image 332 is obtained based on the third frame image 323 and the fourth frame image 324. For example, the pixels of the third frame image 323 are subtracted from the corresponding pixels in the fourth frame image 324 one by one, or the pixels in a local area of the third frame image 323 are subtracted from the corresponding pixels in a local area of the fourth frame image 324 one by one to obtain the second difference image 332.

[0063] Based on the third frame image 323 and the fourth frame image 324, a first optical flow image 333 associated with a first direction and a second optical flow image 334 associated with a second direction are obtained, or based on a local area of the third frame image 323 and a local area of the fourth frame image 324, a first optical flow image 333 associated with a first direction and a second optical flow image 334 associated with a second direction are obtained. The first direction is, for example, the horizontal direction (x direction) of the image, and the second direction is, for example, the vertical direction (y direction) of the image. Optical flow can represent the apparent motion of the image brightness pattern, and the optical flow expresses the change of the image and contains the motion information of the objects in the image.

[0064] The second difference image 332, the first optical flow image 333, and the second optical flow image 334 contain the differences between the third frame image 323 and the fourth frame image 324. Therefore, the second difference image 332, the first optical flow image 333, and the second optical flow image 334 contain the time information associated with the video data 310.

[0065] Next, the second difference image 332, the first optical flow image 333, and the second optical flow image 334 are combined to obtain second image data to be recognized 342 including three second channels.

[0066] In another way, the multi-frame images, for example, include a fifth frame image 325, a sixth frame image 326, and a seventh frame image 327. The acquisition time of the sixth frame image 326 is, for example, after the acquisition time of the fifth frame image 325, and the acquisition time of the seventh frame image 327 is, for example, after the acquisition time of the sixth frame image 326. For example, the sixth frame image 326 is the current frame image, the fifth frame image 325 is the previous frame image of the current frame image, and the seventh frame image 327 is the next frame image of the current frame image. The fifth frame image 325, the sixth frame image 326, and the seventh frame image 327 are images at different times in the video data 310. Therefore, the fifth frame image 325, the sixth frame image 326, and the seventh frame image 327 contain time information associated with the video data 310.

[0067] Next, the fifth frame image 325, the sixth frame image 326, and the seventh frame image 327 are combined to obtain a third image data to be recognized 343 including three third channels.

[0068] After obtaining the first image data to be recognized 341, the second image data to be recognized 342, or the third image data to be recognized 343, the first image data to be recognized 341, the second image data to be recognized 342, or the third image data to be recognized 343 can be input into the deep learning model 350. The deep learning model 350 is used to perform image recognition on any one of the first image data to be recognized 341, the second image data to be recognized 342, and the third image data to be recognized 343 to obtain the motion state and occlusion state 360 of the target object. The motion state and occlusion state 360 are, for example, the prediction results for the current frame image.

[0069] According to an embodiment of the present disclosure, multiple video frames in the video data 310 contain time information. The video data 310 is processed in various ways to obtain the image data to be recognized including time information and image information. Recognition is performed based on the image data to be recognized including time information, fully considering the time information in the video data, and improving the recognition accuracy of the motion state while recognizing the occlusion state based on the image information.

[0070] According to an embodiment of the present disclosure, after obtaining the motion state and occlusion state of the target object, an alarm can be generated based on the motion state and occlusion state. For example, when the motion state indicates that the target object moves or remains stationary for a long time, an alarm can be generated. Or, when the occlusion state indicates that the target object is occluded or not occluded for a long time, an alarm can be generated.

[0071] For example, the first alarm data can be generated based on the motion state, and the second alarm data can be generated based on the occlusion state.

[0072] Exemplarily, generating the first alarm data based on the motion state includes the following operations.

[0073] After obtaining the state value corresponding to the motion state of the current frame image, the obtained state value can be stored in the state list. When it is determined that the motion states in the state list include M 1 state values, determine whether the number of the first state values among the M 1 state values is greater than or equal to M 2 , where M 1 is an integer greater than 0, and M 2 is less than or equal to M 1 .

[0074] If it is determined that the number of the first state values is greater than or equal to M 2 , determine whether the number of the second state values between two non-adjacent first state values is less than or equal to M 3 , where M 3 is less than M 1 .

[0075] If it is determined that the number of the second state values is less than or equal to M 3 , modify the second state values to the first state values to obtain the updated motion state.

[0076] Then, based on the number of the first state values or the number of the second state values included in the updated motion state, generate the first alarm data.

[0077] For example, taking M 1 = 5, M 2 = 3, and M 3 = 1 as an example. The state list is, for example, [A A A B A], where A represents the first state value and B represents the second state value. The first state value represents, for example, stillness, and the second state value represents, for example, motion. Or, the first state value represents, for example, motion, and the second state value represents, for example, stillness. The number of state values in the state list, 5, is greater than or equal to M 1 = 5. Further determine that the number of the first state values A, 4, is greater than or equal to M 2 = 3. Then, determine that the number of the second state values (the fourth B) between two non-adjacent first state values (the third A and the fifth A) is less than or equal to M 3 = 1, indicating that the second state value B is likely to be an unstable outlier. At this time, modify the second state value B in the state list to the first state value A to obtain the updated motion state [A A A A A].

[0078] After obtaining the updated motion state [A A A A A], it is possible to determine whether to generate the first warning data. For example, when the target object is a conveyor belt, the first state value A represents, for example, being stationary, and the updated motion state indicates that the conveyor belt has been stationary for a long time (not working), and at this time, the first warning information can be generated. When the target object is a camera, the first state value A represents, for example, being in motion, and the updated motion state indicates that the camera has been in a state of shaking or rotating for a long time, and at this time, the first warning information can be generated. After generating the first warning information, the status list can be cleared, and the warning list can be updated based on the generated first warning information.

[0079] Exemplarily, generating the second warning data based on the occlusion state includes the following operations.

[0080] If it is determined that the occlusion state includes N 1 status values, determine whether the number of the third status values among the N 1 status values is greater than or equal to N 2 , where N 1 is an integer greater than 0, and N 2 is less than or equal to N 1 .

[0081] If it is determined that the number of the third status values is greater than or equal to N 2 , determine whether the number of the fourth status values between two non-adjacent third status values is less than or equal to N 3 , where N 3 is less than N 1 .

[0082] If it is determined that the number of the fourth status values is less than or equal to N 3 , modify the fourth status values to the third status values to obtain the updated occlusion state.

[0083] Generate the second warning data based on the number of the third status values or the number of the fourth status values included in the updated occlusion state.

[0084] It can be understood that the principle of generating the second warning data is similar to that of generating the first warning data, and will not be elaborated here.

[0085] When the target object is a conveyor belt, if the updated occlusion state indicates that the conveyor belt has been in a non-occluded state for a long time, it means that there is no coal mine on the conveyor belt and it is in an empty-load state, and at this time, the second warning information can be generated. When the target object is a camera, if the updated occlusion state indicates that the camera has been in an occluded state for a long time, it means that the camera may have stains or be externally occluded, affecting the data acquisition effect, and at this time, the second warning information can be generated.

[0086] According to an embodiment of the present disclosure, an alarm smoothing judgment strategy can be obtained by setting multiple thresholds, and alarms can be made through the alarm smoothing judgment strategy. For example, after obtaining the motion state and occlusion state of a target object, an alarm can be made based on the motion state and occlusion state, so as to take relevant measures in time when the target object is abnormal and ensure the normal operation of the target object.

[0087] Figure 4 Schematically shows a flowchart of a method for training a deep learning model according to an embodiment of the present disclosure.

[0088] As Figure 4 shown, the state recognition method 400 of the target object in the embodiment of the present disclosure may include, for example, operation S410 to operation S430.

[0089] In operation S410, historical video data associated with the target object is processed to obtain sample image data. The sample image data includes time information associated with the historical video data and image data associated with the target object. The sample image data includes a motion state label and an occlusion state label.

[0090] In operation S420, the sample image data is processed using the deep learning model to be trained to obtain a motion state recognition result of the target object and an occlusion state recognition result of the target object.

[0091] In operation S430, the deep learning model to be trained is trained based on the motion state recognition result, the motion state label, the occlusion state recognition result, and the occlusion state label.

[0092] Exemplarily, the historical video data is similar to the above-mentioned video data and will not be elaborated here. The historical video data includes, for example, a motion state label indicating whether the target object is moving and an occlusion state label indicating whether the target object is occluded. The label for the historical video data can be used as the label for the sample image data.

[0093] The deep learning model to be trained includes, for example, a classification model. The deep learning model can identify the sample image data to obtain a classification result, and the classification result includes, for example, a motion state recognition result of the target object and an occlusion state recognition result of the target object.

[0094] After obtaining the motion state recognition result and the occlusion state recognition result, the deep learning model to be trained can be trained based on the difference between the motion state recognition result and the motion state label and the difference between the occlusion state recognition result and the occlusion state label to obtain a trained deep learning model.

[0095] According to an embodiment of the present disclosure, sample image data including time information and image information can be obtained by processing historical video data, and a deep learning model is trained using the sample image data, so that the deep learning model has the ability to simultaneously identify the motion state and the occlusion state. It can be seen that after obtaining the sample image data including time information and image information by processing the historical video data, the deep learning model is trained based on the sample image data, realizing that the deep learning model simultaneously performs the motion state recognition task and the occlusion state recognition task, reducing the processing time of the task, and achieving the effect of multi-task processing.

[0096] In an example of the present disclosure, the target object includes, for example, a conveyor device, and the historical video data includes data obtained by video capturing the conveyor device. The motion state recognition result indicates whether the conveyor device is in motion. The occlusion state recognition result indicates whether there is an object to be conveyed on the conveyor device. The object to be conveyed includes, for example, coal, and the occlusion state recognition result, for example, indicates whether there is coal on the conveyor belt. If there is coal on the conveyor belt, the occlusion state recognition result is occluded at this time. If there is no coal on the conveyor belt, the occlusion state recognition result is unoccluded at this time.

[0097] In another example of the present disclosure, the target object includes, for example, a video data acquisition device, and the historical video data includes data obtained by using the video data acquisition device to video capture other objects. The motion state recognition result indicates whether the video data acquisition device is in motion, for example, indicates whether the video data acquisition device shakes, rotates, etc. The video data acquisition device includes, for example, a camera. The occlusion state recognition result indicates whether the video data acquisition device is occluded, for example, indicates whether the front of the camera is occluded when collecting data. In one example, when there is a stain on the camera, it means the camera is occluded.

[0098] When training a deep learning model, a large amount of sample image data is usually required for training. Taking the conveyor belt as the target object as an example, a large amount of historical video data is usually data about the conveyor belt in motion and occluded by coal, while the data of the conveyor belt being stationary and not occluded by coal is less, resulting in unbalanced samples for training the deep learning model. Therefore, the embodiments of the present disclosure can process the historical video data to expand the diversity of the samples.

[0099] For example, multiple frames of images are extracted from historical video data, and the multiple frames of images are respectively extracted to obtain multiple frames of local images that correspond one by one to the multiple frames of images. In other words, local images are randomly extracted from the multiple frames of images as samples, which increases the number of samples. In one example, local images of a preset size can be directly and randomly extracted from the entire frame image. Alternatively, the region of interest (ROI) in the frame image can be marked. The region of interest is, for example, the region where the target object is located or an important region, and then local images of a preset size are randomly cropped from the region of interest. The preset size is, for example, 128*128, 256*256, 344*344, 512*512, and so on.

[0100] In addition, due to uneven illumination and the influence of factors such as specular reflection in the actual scene, color enhancement strategies such as adjusting the image color through HSV color space transformation and non-linearly mapping the pixel values of the image through gamma transformation can be used to expand the number of samples and sample diversity. At the same time, it ensures that the samples in model training contain multi-scale information and improves the generalization of the model.

[0101] After obtaining the multiple frames of local images, the multiple frames of local images can be processed to obtain sample image data including multiple channels. Obtaining the sample image data of multiple channels is, for example, similar to the process of obtaining the multi-channel image data to be recognized mentioned above.

[0102] In one way, the multiple frames of local images include, for example, a first frame image and a second frame image. Based on the first frame image and the second frame image, a first difference image is obtained, and the first frame image, the second frame image, and the first difference image are combined to obtain first sample image data including three first channels.

[0103] In another way, the multiple frames of local images include, for example, a third frame image and a fourth frame image. Based on the third frame image and the fourth frame image, a second difference image is obtained, and based on the third frame image and the fourth frame image, a first optical flow image associated with a first direction and a second optical flow image associated with a second direction are obtained. The second difference image, the first optical flow image, and the second optical flow image are combined to obtain second sample image data including three second channels.

[0104] In another way, the multiple frames of local images include, for example, a fifth frame image, a sixth frame image, and a seventh frame image. The fifth frame image, the sixth frame image, and the seventh frame image are combined to obtain third sample image data including three third channels.

[0105] Figure 5 A schematic diagram of a deep learning model according to an embodiment of the present disclosure is schematically shown.

[0106] AsFigure 5 As shown, the deep learning model to be trained includes, for example, a convolutional neural network model, and the convolutional neural network model includes, for example, a resnet18 model. The deep learning model to be trained includes, for example, a pre-order network layer 510, a first post-order network layer 520, and a second post-order network layer 530, and the first post-order network layer 520 and the second post-order network layer 530 are in parallel. For example, the first post-order network layer 520 is connected to the pre-order network layer 510, and the second post-order network layer 530 is connected to the pre-order network layer 510.

[0107] For example, the pre-order network layer 510 is used to extract features from the sample image data to obtain pre-order features. The first post-order network layer 520 is used to process the pre-order features to obtain the recognition result of the motion state of the target object. The second post-order network layer 530 is used to process the pre-order features to obtain the recognition result of the occlusion state of the target object.

[0108] In one example, the first post-order network layer 520 includes, for example, a fully connected layer, and the second post-order network layer 530 includes, for example, a fully connected layer.

[0109] Figure 5 In the parameters of each layer, k represents the convolution kernel size, s represents the convolution calculation stride, and p represents the padding size. For example, for an n×n image, if a convolution is performed with an f×f filter, the output dimension is (n - f + 1)×(n - f + 1), and the size of the output feature map becomes smaller. When we do not want the convolution scale to become smaller each time, padding can be used for expansion to increase the size.

[0110] In order to reduce the model parameters of the deep learning model and the time consumption of deployment and prediction caused by multi-model cascading, the embodiment of the present disclosure can achieve end-to-end classification through the classification network resnet18 of multi-task training. At the output of the resnet18 model, two fully connected layers are constructed in parallel in combination with the multi-task learning principle. One fully connected layer is used to output the conveyor belt motion stationary information, and the other fully connected layer is used to output the information of whether there is coal on the conveyor belt. In addition, a weighted cross-entropy loss function is introduced at the output of the two fully connected layers to optimize and constrain the model training process, greatly reducing the impact of data imbalance between the two tasks. For example, the weights of different data categories are different, and the contribution degrees to the loss function are different. The weighted ratio is inversely proportional to the data volume of the data category. In this way, the loss weight of the category with a large data volume is small, which is equivalent to balancing the influence of different data categories. The quantity categories include stationary sample categories, motion sample categories, coal (occluded) sample categories, coal-free (unoccluded) sample categories, etc. Through the joint training of multi-tasks, the processing performance of the two tasks is improved synchronously, and the feature extraction ability of the model is improved.

[0111] Figure 6Schematically shows a schematic diagram of a method for recognizing the state of a target object and a method for training a deep learning model according to an embodiment of the present disclosure.

[0112] As Figure 6 shown, in the training stage of the deep learning model 640, by extracting frames from the historical video data 610, video frames 620 are obtained. The ROI regions in the video frames 620 are marked, and the marked video frames 620 are stored in the image library 630. For the video frames stored in the image library 630, local images are cropped from the ROI regions of the video frames for processing to obtain three-channel sample image data (first sample image data, second sample image data, or third sample image data). The three-channel sample image data is input into the deep learning model 640 for processing to obtain a prediction result 650 for the current frame image, and the deep learning model 640 is trained based on the difference between the prediction result 650 for the current frame image and the label of the sample image data.

[0113] In the usage stage of the trained deep learning model 640, by extracting frames from the video data 610, video frames 620 are obtained. The video frames 620 are stored in the image library 630. For the video frames stored in the image library 630, the video frames are processed to obtain three-channel image data to be recognized (first image data to be recognized, second image data to be recognized, or third image data to be recognized), and the trained deep learning model 640 is used to predict the local region indicated by the electronic fence in the image data to be recognized to obtain a prediction result 650 for the current frame image. The prediction result for the current frame image is stored in the status list 660, and an alarm is made and the alarm list 670 is updated based on the status list 660.

[0114] It can be understood that the embodiments of the present disclosure can be applied to the scenario of recognizing the state of a conveyor belt in a mining area, and can be further extended to similar scenarios such as camera rotation and occlusion in mines and safety production scenarios, improving the applicability of different scenarios. With the increase in video data, iteratively training and optimizing the multi-task classification model can further improve the effect of the model, effectively solving problems such as multi-task under-coupling, poor model generalization, and poor accuracy caused by illumination and perspective changes.

[0115] Figure 7 Schematically shows a block diagram of a device for recognizing the state of a target object according to an embodiment of the present disclosure.

[0116] As Figure 7 shown, the device 700 for recognizing the state of a target object according to an embodiment of the present disclosure includes, for example, a processing module 710 and a recognition module 720.

[0117] The processing module 710 can be used to process video data associated with a target object to obtain image data to be recognized. The image data to be recognized includes time information associated with the video data and image data associated with the target object. According to an embodiment of the present disclosure, the processing module 710 can, for example, perform the operation S210 described above with reference to Figure 2 and will not be elaborated herein.

[0118] The recognition module 720 can be used to recognize the image data to be recognized to obtain the motion state of the target object and the occlusion state of the target object. According to an embodiment of the present disclosure, the recognition module 720 can, for example, perform the operation S220 described above with reference to Figure 2 and will not be elaborated herein.

[0119] According to an embodiment of the present disclosure, the processing module 710 includes: an extraction sub-module and a processing sub-module. The extraction sub-module is used to extract multiple frames of images from the video data; the processing sub-module is used to process the multiple frames of images to obtain image data to be recognized including multiple channels.

[0120] According to an embodiment of the present disclosure, the image data to be recognized includes first image data to be recognized, the multiple channels include three first channels; the multiple frames of images include a first frame image and a second frame image; the processing sub-module includes: a first obtaining unit and a first combining unit. The first obtaining unit is used to obtain a first difference image based on the first frame image and the second frame image; the first combining unit is used to combine the first frame image, the second frame image and the first difference image to obtain the first image data to be recognized.

[0121] According to an embodiment of the present disclosure, the image data to be recognized includes second image data to be recognized, the multiple channels include three second channels; the multiple frames of images include a third frame image and a fourth frame image; the processing sub-module includes: a second obtaining unit, a third obtaining unit and a second combining unit. The second obtaining unit is used to obtain a second difference image based on the third frame image and the fourth frame image; the third obtaining unit is used to obtain a first optical flow image associated with a first direction and a second optical flow image associated with a second direction based on the third frame image and the fourth frame image; the second combining unit is used to combine the second difference image, the first optical flow image and the second optical flow image to obtain the second image data to be recognized.

[0122] According to an embodiment of the present disclosure, the image data to be recognized includes third image data to be recognized, the multiple channels include three third channels; the multiple frames of images include a fifth frame image, a sixth frame image and a seventh frame image; the processing sub-module includes: a third combining unit, used to combine the fifth frame image, the sixth frame image and the seventh frame image to obtain the third image data to be recognized.

[0123] According to an embodiment of the present disclosure, the recognition module 720 is further configured to: use a deep learning model to recognize the image data to be recognized, and obtain a motion state and an occlusion state.

[0124] According to an embodiment of the present disclosure, the target object includes a conveying device, the video data includes data obtained by video capturing of the conveying device, the motion state indicates whether the conveying device is in motion, and the occlusion state indicates whether there is an object to be conveyed on the conveying device.

[0125] According to an embodiment of the present disclosure, the target object includes a video data capturing device, the video data includes data obtained by video capturing using the video data capturing device, the motion state indicates whether the video data capturing device is in motion, and the occlusion state indicates whether the video data capturing device is occluded.

[0126] According to an embodiment of the present disclosure, the device 700 may further include at least one of a first generation module and a second generation module. The first generation module is configured to generate first alarm data based on the motion state; the second generation module is configured to generate second alarm data based on the occlusion state.

[0127] According to an embodiment of the present disclosure, the first generation module includes: a first determination sub-module, a second determination sub-module, a first modification sub-module, and a first generation sub-module. The first determination sub-module is configured to, in response to determining that the motion state includes M 1 state values, determine whether the number of first state values among the M 1 state values is greater than or equal to M 2 , M 1 is an integer greater than 0, and M 2 is less than or equal to M 1 ; the second determination sub-module is configured to, in response to determining that the number of first state values is greater than or equal to M 2 , determine whether the number of second state values between two non-adjacent first state values is less than or equal to M 3 , M 3 is less than M 1 ; the first modification sub-module is configured to, in response to determining that the number of second state values is less than or equal to M 3 , modify the second state values to first state values to obtain an updated motion state; the first generation sub-module is configured to generate first alarm data based on the number of first state values or the number of second state values included in the updated motion state.

[0128] According to an embodiment of the present disclosure, the second generation module includes: a third determination sub-module, a fourth determination sub-module, a second modification sub-module, and a second generation sub-module. The third determination sub-module is configured to, in response to determining that the occlusion state includes N 1 state values, determine N 1whether the number of third state values among the state values is greater than or equal to N 2 , N 1 is an integer greater than 0, N 2 less than or equal to N 1 ; a fourth determination sub-module, configured to, in response to determining that the number of third state values is greater than or equal to N 2 , determine whether the number of fourth state values between two non-adjacent third state values is less than or equal to N 3 , N 3 less than N 1 ; a second modification sub-module, configured to, in response to determining that the number of fourth state values is less than or equal to N 3 , modify the fourth state values to third state values to obtain an updated occlusion state; a second generation sub-module, configured to generate second alarm data based on the number of third state values or the number of fourth state values included in the updated occlusion state.

[0129] Figure 8 Schematically shows a block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure.

[0130] As Figure 8 shown, the training apparatus 800 for a deep learning model according to an embodiment of the present disclosure includes, for example, a first processing module 810, a second processing module 820, and a training module 830.

[0131] The first processing module 810 may be configured to process historical video data associated with a target object to obtain sample image data, where the sample image data includes time information associated with the historical video data and image data associated with the target object, and the sample image data includes a motion state label and an occlusion state label. According to an embodiment of the present disclosure, the first processing module 810 may, for example, perform the operation S410 described above with reference to Figure 4 , which will not be elaborated herein.

[0132] The second processing module 820 may use the deep learning model to be trained to process the sample image data to obtain a motion state recognition result of the target object and an occlusion state recognition result of the target object. According to an embodiment of the present disclosure, the second processing module 820 may, for example, perform the operation S420 described above with reference to Figure 4 , which will not be elaborated herein.

[0133] The training module 830 may be configured to train the deep learning model to be trained based on the motion state recognition result, the motion state label, the occlusion state recognition result, and the occlusion state label. According to an embodiment of the present disclosure, the training module 830 may, for example, perform the operation S430 described above with reference to Figure 4 , which will not be elaborated herein.

[0134] According to an embodiment of the present disclosure, the first processing module 810 includes: a first extraction sub-module, a second extraction sub-module, and a first processing sub-module. The first extraction sub-module is configured to extract multiple frames of images from historical video data; the second extraction sub-module is configured to perform extraction on the multiple frames of images respectively to obtain multiple frames of local images corresponding to the multiple frames of images one by one; the first processing sub-module is configured to process the multiple frames of local images to obtain sample image data including multiple channels.

[0135] According to an embodiment of the present disclosure, the sample image data includes first sample image data, the multiple channels include three first channels; the multiple frames of local images include a first frame image and a second frame image; the first processing sub-module includes: a first obtaining unit and a first combining unit. The first obtaining unit is configured to obtain a first difference image based on the first frame image and the second frame image; the first combining unit is configured to combine the first frame image, the second frame image, and the first difference image to obtain the first sample image data.

[0136] According to an embodiment of the present disclosure, the sample image data includes second sample image data, the multiple channels include three second channels; the multiple frames of local images include a third frame image and a fourth frame image; the first processing sub-module includes: a second obtaining unit, a third obtaining unit, and a second combining unit. The second obtaining unit is configured to obtain a second difference image based on the third frame image and the fourth frame image; the third obtaining unit is configured to obtain a first optical flow image associated with a first direction and a second optical flow image associated with a second direction based on the third frame image and the fourth frame image; the second combining unit is configured to combine the second difference image, the first optical flow image, and the second optical flow image to obtain the second sample image data.

[0137] According to an embodiment of the present disclosure, the sample image data includes third sample image data, the multiple channels include three third channels; the multiple frames of local images include a fifth frame image, a sixth frame image, and a seventh frame image; the first processing sub-module includes: a third combining unit configured to combine the fifth frame image, the sixth frame image, and the seventh frame image to obtain the third sample image data.

[0138] According to an embodiment of the present disclosure, the deep learning model to be trained includes a pre-order network layer, a first post-order network layer, and a second post-order network layer. The first post-order network layer is connected to the pre-order network layer, and the second post-order network layer is connected to the pre-order network layer; wherein, the second processing module 820 includes: a third extraction sub-module, a second processing sub-module, and a third processing sub-module. The third extraction sub-module is configured to perform feature extraction on the sample image data by using the pre-order network layer to obtain pre-order features; the second processing sub-module is configured to process the pre-order features by using the first post-order network layer to obtain a recognition result of the motion state of the target object; the third processing sub-module is configured to process the pre-order features by using the second post-order network layer to obtain a recognition result of the occlusion state of the target object.

[0139] According to an embodiment of the present disclosure, the first subsequent network layer includes a fully connected layer, and the second subsequent network layer includes a fully connected layer.

[0140] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, disclosure, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and it does not violate public order and good customs.

[0141] In the technical solution of the present disclosure, before obtaining or collecting the user's personal information, the authorization or consent of the user is obtained.

[0142] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0143] According to an embodiment of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method for identifying the state of a target object described above.

[0144] According to an embodiment of the present disclosure, there is provided a computer program product including computer programs / instructions that, when executed by a processor, implement the method for identifying the state of a target object described above.

[0145] Figure 9 is a block diagram of an electronic device for implementing the method for identifying the state of a target object and / or training a deep learning model according to an embodiment of the present disclosure.

[0146] Figure 9 FIG. shows a schematic block diagram of an exemplary electronic device 900 that can be used to implement an embodiment of the present disclosure. The electronic device 900 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described herein and / or claimed.

[0147] Such as Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0148] Multiple components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0149] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the method for identifying the state of a target object and / or the method for training a deep learning model. For example, in some embodiments, the method for identifying the state of a target object and / or the method for training a deep learning model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the method for identifying the state of a target object and / or the method for training a deep learning model described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the method for identifying the state of a target object and / or the method for training a deep learning model in any other appropriate manner (e.g., by means of firmware).

[0150] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0151] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable target object's state recognition device and / or a training device of a deep learning model, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0152] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0153] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0154] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0155] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server combined with a blockchain.

[0156] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0157] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for identifying the state of a target object, comprising: extracting multiple frames of images from video data associated with the target object; processing the multiple frames of images to obtain image data to be recognized including multiple channels, wherein the image data to be recognized includes time information associated with the video data and image data associated with the target object; and identifying the image data to be recognized to obtain the motion state of the target object and the occlusion state of the target object, wherein the image data to be recognized includes second image data to be recognized, the multiple channels include three second channels; the multiple frames of images include a third frame of image and a fourth frame of image; the processing the multiple frames of images to obtain the image data to be recognized including multiple channels includes: obtaining a second difference image based on a local part of the third frame of image and a local part of the fourth frame of image, wherein the local part of the third frame of image includes a region of interest in the third frame of image, and the local part of the fourth frame of image includes a region of interest in the fourth frame of image; obtaining a first optical flow image associated with a first direction and a second optical flow image associated with a second direction based on the difference between the local part of the third frame of image and the local part of the fourth frame of image; and combining the second difference image, the first optical flow image and the second optical flow image to obtain the second image data to be recognized.

2. The method according to claim 1, wherein, the image data to be recognized includes first image data to be recognized, the multiple channels include three first channels; the multiple frames of images include a first frame of image and a second frame of image; the processing the multiple frames of images to obtain the image data to be recognized including multiple channels includes: obtaining a first difference image based on the first frame of image and the second frame of image; and combining the first frame of image, the second frame of image and the first difference image to obtain the first image data to be recognized.

3. The method according to claim 1, wherein, the image data to be recognized includes third image data to be recognized, the multiple channels include three third channels; the multiple frames of images include a fifth frame of image, a sixth frame of image and a seventh frame of image; the processing the multiple frames of images to obtain the image data to be recognized including multiple channels includes: combining the fifth frame of image, the sixth frame of image and the seventh frame of image to obtain the third image data to be recognized.

4. The method according to any one of claims 1-3, wherein, the identifying the image data to be recognized to obtain the motion state of the target object and the occlusion state of the target object includes: using a deep learning model to identify the image data to be recognized to obtain the motion state and the occlusion state.

5. The method according to any one of claims 1-3, wherein, The target object includes a conveying device, the video data includes data obtained by video capturing of the conveying device, the motion state characterizes whether the conveying device is in motion, and the occlusion state characterizes whether there is an object to be conveyed on the conveying device.

6. The method according to any one of claims 1-3, wherein, the target object includes a video data capturing device, the video data includes data obtained by video capturing using the video data capturing device, the motion state characterizes whether the video data capturing device is in motion, and the occlusion state characterizes whether the video data capturing device is occluded.

7. The method according to any one of claims 1-3 further includes at least one of the following: generating first alarm data based on the motion state; and generating second alarm data based on the occlusion state.

8. The method according to claim 7, wherein, generating the first alarm data based on the motion state includes: In response to determining that the motion state includes M 1 state values, determining whether the number of first state values among the M 1 state values is greater than or equal to M 2 , where M 1 is an integer greater than 0, and M 2 is less than or equal to M 1 ; In response to determining that the number of the first state values is greater than or equal to M 2 , determine whether the number of second state values between two non-adjacent first state values is less than or equal to M 3 , M 3 is less than M 1 ; In response to determining that the number of the second state values is less than or equal to M 3 , modify the second state value to the first state value to obtain an updated motion state; and generating the first alarm data based on the number of the first state values or the number of the second state values included in the updated motion state.

9. The method according to claim 8, wherein, generating the second alarm data based on the occlusion state includes: In response to determining that the occlusion state includes N 1 state values, determining whether the number of third state values among the N 1 state values is greater than or equal to N 2 , where N 1 is an integer greater than 0, and N 2 is less than or equal to N 1 ; In response to determining that the number of the third state values is greater than or equal to N 2 , determine whether the number of the fourth state values between two non-adjacent third state values is less than or equal to N 3 , N 3 is less than N 1 ; In response to determining that the number of the fourth state values is less than or equal to N 3 , modifying the fourth state value to the third state value to obtain an updated occlusion state; and generating the second alarm data based on the number of the third state values or the number of the fourth state values included in the updated occlusion state.

10. A training method for a deep learning model, including: extracting multiple frames of images from historical video data associated with a target object; respectively extracting the multiple frames of images to obtain multiple frames of local images corresponding to the multiple frames of images, wherein the local images are extracted from the regions of interest of the frame images; processing the multiple frames of local images to obtain sample image data including multiple channels, wherein the sample image data includes time information associated with the historical video data and image data associated with the target object, and the sample image data includes a motion state label and an occlusion state label; processing the sample image data using a deep learning model to be trained to obtain a motion state recognition result of the target object and an occlusion state recognition result of the target object; and training the deep learning model to be trained based on the motion state recognition result, the motion state label, the occlusion state recognition result, and the occlusion state label, wherein the sample image data includes second sample image data, the multiple channels include three second channels; the multiple frames of local images include a third frame image and a fourth frame image; and processing the multiple frames of local images to obtain the sample image data including multiple channels includes: obtaining a second difference image based on the third frame image and the fourth frame image; obtaining a first optical flow image associated with a first direction and a second optical flow image associated with a second direction based on the difference between the third frame image and the fourth frame image; and combining the second difference image, the first optical flow image, and the second optical flow image to obtain the second sample image data.

11. The method according to claim 10, wherein, the sample image data includes first sample image data, and the multiple channels include three first channels; the multiple frames of local images include a first frame image and a second frame image; the processing of the multiple frames of local images to obtain the sample image data including multiple channels includes: obtaining a first difference image based on the first frame image and the second frame image; and combining the first frame image, the second frame image, and the first difference image to obtain the first sample image data.

12. The method according to claim 10, wherein, the sample image data includes third sample image data, and the multiple channels include three third channels; the multiple frames of local images include a fifth frame image, a sixth frame image, and a seventh frame image; the processing of the multiple frames of local images to obtain the sample image data including multiple channels includes: combining the fifth frame image, the sixth frame image, and the seventh frame image to obtain the third sample image data.

13. The method according to claim 10, wherein, the deep learning model to be trained includes a pre-order network layer, a first post-order network layer, and a second post-order network layer, the first post-order network layer is connected to the pre-order network layer, and the second post-order network layer is connected to the pre-order network layer; wherein, the processing of the sample image data by using the deep learning model to be trained to obtain the motion state recognition result of the target object and the occlusion state recognition result of the target object includes: extracting pre-order features from the sample image data by using the pre-order network layer; processing the pre-order features by using the first post-order network layer to obtain the motion state recognition result of the target object; and processing the pre-order features by using the second post-order network layer to obtain the occlusion state recognition result of the target object.

14. The method according to claim 13, wherein, the first post-order network layer includes a fully connected layer, and the second post-order network layer includes a fully connected layer.

15. A state recognition device for a target object, comprising: an extraction sub-module for extracting multiple frames of images from video data associated with the target object; and a processing sub-module for processing the multiple frames of images to obtain image data to be recognized including multiple channels, wherein the image data to be recognized contains time information associated with the video data and image data associated with the target object; and a recognition module for recognizing the image data to be recognized to obtain the motion state of the target object and the occlusion state of the target object, wherein the image data to be recognized includes second image data to be recognized, and the multiple channels include three second channels; the multiple frames of images include a third frame image and a fourth frame image; the processing sub-module includes: A second acquisition unit, configured to obtain a second difference image based on a local part of the third frame image and a local part of the fourth frame image, where the local part of the third frame image includes a region of interest in the third frame image, and the local part of the fourth frame image includes a region of interest in the fourth frame image; A third acquisition unit, configured to obtain a first optical flow image associated with a first direction and a second optical flow image associated with a second direction based on a difference between the local part of the third frame image and the local part of the fourth frame image; and A second combination unit, configured to combine the second difference image, the first optical flow image, and the second optical flow image to obtain the second image data to be recognized.

16. The apparatus according to claim 15, wherein, the image data to be recognized includes first image data to be recognized, the multiple channels include three first channels; the multiple frame images include a first frame image and a second frame image; the processing sub-module includes: A first acquisition unit, configured to obtain a first difference image based on the first frame image and the second frame image; and A first combination unit, configured to combine the first frame image, the second frame image, and the first difference image to obtain the first image data to be recognized.

17. The apparatus according to claim 15, wherein, the image data to be recognized includes third image data to be recognized, the multiple channels include three third channels; the multiple frame images include a fifth frame image, a sixth frame image, and a seventh frame image; the processing sub-module includes: A third combination unit, configured to combine the fifth frame image, the sixth frame image, and the seventh frame image to obtain the third image data to be recognized.

18. The apparatus according to any one of claims 15-17, wherein, the recognition module is further configured to: Use a deep learning model to recognize the image data to be recognized to obtain the motion state and the occlusion state.

19. The apparatus according to any one of claims 15-17, wherein, the target object includes a conveying device, the video data includes data obtained by video capturing the conveying device, the motion state indicates whether the conveying device is in motion, and the occlusion state indicates whether there is an object to be conveyed on the conveying device.

20. The apparatus according to any one of claims 15-17, wherein, the target object includes a video data acquisition device, the video data includes data obtained by video capturing using the video data acquisition device, the motion state indicates whether the video data acquisition device is in motion, and the occlusion state indicates whether the video data acquisition device is occluded.

21. The apparatus according to any one of claims 15-17, further includes at least one of the following: A first generation module, configured to generate first alarm data based on the motion state; and A second generation module, configured to generate second alarm data based on the occlusion state.

22. The apparatus according to claim 21, wherein, the first generation module includes: The first determination sub-module is configured to, in response to determining that the motion state includes M 1 state values, determine whether the number of first state values among the M 1 state values is greater than or equal to M 2 , where M 1 is an integer greater than 0, and M 2 is less than or equal to M 1 ; A second determination sub-module, configured to, in response to determining that the number of the first state values is greater than or equal to M 2 , determine whether the number of second state values between two non-adjacent first state values is less than or equal to M 3 , M 3 is less than M 1 ; The first modification sub-module is configured to, in response to determining that the number of the second state values is less than or equal to M 3 , modify the second state value to the first state value to obtain an updated motion state; and A first generation sub-module, configured to generate the first warning data based on the number of the first state values or the number of the second state values included in the updated motion state.

23. The apparatus according to claim 21, wherein, the second generation module includes: A third determination sub-module, configured to determine whether the number of third state values among the N state values is greater than or equal to N in response to determining that the occlusion state includes N state values 1 where N is a positive integer, and N is less than or equal to N 1 2 1 2 1 ;​​​​ A fourth determination sub-module, configured to determine whether the number of fourth state values between two non-adjacent third state values is less than or equal to N in response to determining that the number of the third state values is greater than or equal to N 2 , where N 3 is less than N 3 ; 1 ​ A second modification sub-module, configured to, in response to determining that the number of the fourth state values is less than or equal to N 3 , modify the fourth state value to the third state value to obtain an updated occlusion state; and A second generation sub-module, configured to generate the second warning data based on the number of the third state values or the number of the fourth state values included in the updated occlusion state.

24. A training apparatus for a deep learning model, comprising: A first extraction sub-module, configured to extract multiple frames of images from historical video data associated with a target object; A second extraction sub-module, configured to respectively extract the multiple frames of images to obtain multiple frames of local images corresponding to the multiple frames of images, wherein the local images are extracted from regions of interest of the frame images; A first processing sub-module, configured to process the multiple frames of local images to obtain sample image data including multiple channels, wherein the sample image data includes time information associated with the historical video data and image data associated with the target object, and the sample image data includes a motion state label and an occlusion state label; A second processing module, configured to process the sample image data by using a deep learning model to be trained to obtain a motion state recognition result of the target object and an occlusion state recognition result of the target object; and A training module, configured to train the deep learning model to be trained based on the motion state recognition result, the motion state label, the occlusion state recognition result, and the occlusion state label, wherein the sample image data includes second sample image data, the multiple channels include three second channels; the multiple frames of local images include a third frame image and a fourth frame image; the first processing sub-module includes: A second obtaining unit, configured to obtain a second difference image based on the third frame image and the fourth frame image; A third obtaining unit, configured to obtain a first optical flow image associated with a first direction and a second optical flow image associated with a second direction based on a difference between the third frame image and the fourth frame image; and A second combining unit, configured to combine the second difference image, the first optical flow image, and the second optical flow image to obtain the second sample image data.

25. The apparatus according to claim 24, wherein, the sample image data includes first sample image data, the multiple channels include three first channels; the multiple frames of local images include a first frame image and a second frame image; the first processing sub-module includes: A first obtaining unit, configured to obtain a first difference image based on the first frame image and the second frame image; and A first combining unit, configured to combine the first frame image, the second frame image, and the first difference image to obtain the first sample image data.

26. The apparatus according to claim 24, wherein, the sample image data includes third sample image data, the multiple channels include three third channels; the multiple frames of local images include a fifth frame image, a sixth frame image, and a seventh frame image; The first processing sub-module includes: A third combining unit, configured to combine the fifth frame image, the sixth frame image, and the seventh frame image to obtain the third sample image data.

27. The apparatus according to claim 24, wherein, The deep learning model to be trained includes a pre-order network layer, a first post-order network layer, and a second post-order network layer. The first post-order network layer is connected to the pre-order network layer, and the second post-order network layer is connected to the pre-order network layer; wherein, the second processing module includes: A third extraction sub-module, configured to extract features from the sample image data by using the pre-order network layer to obtain pre-order features; A second processing sub-module, configured to process the pre-order features by using the first post-order network layer to obtain the recognition result of the motion state of the target object; and A third processing sub-module, configured to process the pre-order features by using the second post-order network layer to obtain the recognition result of the occlusion state of the target object.

28. The apparatus according to claim 27, wherein, The first post-order network layer includes a fully-connected layer, and the second post-order network layer includes a fully-connected layer.

29. An electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-14.

30. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.

31. A computer program product, including a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1-14 are implemented.

Citation Information

Patent Citations

  • Moving object detection method and device, medium and calculation device

    CN107507225A