Land mining activity video cognition method and device

By constructing engineering machinery behavior image datasets and improving deep learning models, combined with knowledge graphs, the problem of high false alarm rate in intelligent video surveillance methods is solved, and accurate identification and dynamic monitoring of land mining activities are achieved.

CN120356129APending Publication Date: 2025-07-22NAT GEOMATICS CENT OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510175784.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing intelligent video surveillance methods are prone to false alarms when identifying private mining and random mining activities of cultivated land, resulting in low recognition accuracy and inability to accurately identify land mining activities.

Method used

Build a construction machinery behavior image data set for land mining scenarios, condense the behavior state set through time-series knowledge, improve the deep learning model, combine knowledge graphs and knowledge reasoning mechanisms, and carefully capture the dynamic element characteristics of construction machinery, and achieve accurate identification of land mining activities.

Benefits of technology

It improves the accuracy of identification of land mining activities, reduces the false alarm rate, improves the degree of automation and interpretability of identification, and realizes accurate and dynamic identification of land mining activities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356129A_ABST
    Figure CN120356129A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a land mining activity video cognition method and device, and the method comprises the steps: constructing an engineering machinery behavior image data set facing a land mining scene at a data level, and comprehensively describing the interaction behavior characteristics between engineering machinery from the two dimensions of time and space. According to the method, on the algorithm level, aiming at the characteristics of frequent dynamic element movement and high timeliness requirement in land scene video data, time sequence knowledge is fused to carry out targeted improvement on a target detection model, and on the basis of a mixed intelligent element behavior recognition algorithm, subtle characteristic differences of the same element under different behaviors are carefully captured. According to the method, in a reasoning stage, video dynamic information and domain knowledge can be deeply fused by constructing a knowledge graph and a knowledge reasoning mechanism, so that the problem of incomplete identification of an element interaction relationship is overcome, and accurate identification of land mining activities is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a method and device for video cognition of land mining activities. Background Art

[0002] Cultivated land is an irreplaceable carrier for food production and also a basic resource and condition for human survival. However, with the rapid advancement of urbanization and industrialization, the contradiction between humans and land has become increasingly prominent. Illegal mining activities such as excavation, extraction, and related resource development within the scope of cultivated land are common, resulting in the conversion of cultivated land from food cultivation use to non-agricultural uses such as housing and commerce. Illegal mining of cultivated land refers to activities of unauthorized excavation, extraction, and related resource development within the scope of cultivated land without legal authorization or in violation of relevant regulations. Illegal mining activities of cultivated land not only cause continuous loss of cultivated land but also directly threaten food security and social stability.

[0003] In order to quickly identify and stop illegal mining activities of cultivated land, intelligent video monitoring methods can be used to focus on the typical features in the illegal mining scenarios of cultivated land. For example, construction machinery such as excavators and trucks in the earth-taking scenarios of illegal mining of cultivated land are typical illegal mining features. When typical illegal mining features are identified, the intelligent video monitoring system will perform early warning processing.

[0004] However, the above intelligent video monitoring method belongs to extensive monitoring, resulting in frequent false alarm problems in the intelligent video monitoring method. For example, when there is passing construction machinery in the monitored area, the intelligent video monitoring system will also extract illegal mining features from the video data, thus causing false alarms of illegal mining of cultivated land by the intelligent video monitoring system and reducing the recognition accuracy of illegal mining activities of cultivated land. Summary of the Invention

[0005] In view of this, embodiments of this application provide a method and device for video cognition of land mining activities to solve the problem of inaccurate identification of land mining activities.

[0006] According to one aspect of this application, a method for video cognition of land mining activities is provided. The method includes:

[0007] Obtain an image data set, where the image data set includes feature pictures extracted from video data, and the feature pictures include image targets corresponding to mining construction machinery;

[0008] Perform behavior annotation on the image data set according to a set of behavior states to obtain an annotated data set, where the set of behavior states is a set of states obtained by condensing the behavior process features of the mining construction machinery according to temporal sequence knowledge;

[0009] Train an object detection model using the labeled dataset, where the object detection model is a deep learning model improved based on the behavior recognition characteristics of mining construction machinery;

[0010] Input the video data of the land scene into the trained object detection model to obtain the recognition results output by the object detection model, where the recognition results include the dynamic elements contained in the current land scene, the element behaviors of the dynamic elements, and the pixel coordinates of the dynamic elements;

[0011] Generate a scene state sequence according to the recognition results, and perform inference calculations on the scene state sequence based on temporal knowledge to obtain the discrimination result of the land mining activity.

[0012] Optionally, obtain an image dataset, including:

[0013] Obtain the acquisition rules for guiding the video data acquisition of mining construction machinery, and the operation videos of the mining construction machinery;

[0014] Remove the invalid videos in the operation videos to obtain an initial dataset;

[0015] Perform video screening on the initial dataset based on the acquisition rules to obtain a video dataset that conforms to the land mining scene;

[0016] Perform frame extraction on the video dataset at a preset frame extraction interval to obtain the image dataset.

[0017] Optionally, perform behavior annotation on the image dataset according to the behavior state set to obtain a labeled dataset, including:

[0018] Summarize the behavior state set of the mining construction machinery from the aspects of element form and interaction mode;

[0019] Condense the behavior process characteristics of the mining construction machinery according to temporal knowledge;

[0020] Perform condensation on the behavior state set based on the behavior process characteristics to obtain a condensed behavior state set;

[0021] Perform behavior annotation on the image dataset according to the condensed behavior state set to obtain a labeled dataset.

[0022] Optionally, before training the object detection model using the labeled dataset, the method further includes:

[0023] Based on the behavioral recognition characteristics of the mining engineering machinery in the land mining scenario, construct temporal knowledge, which is used to perform recognition on the land mining scenario; the prior knowledge includes deep learning knowledge for model improvement, video image knowledge for element behavior classification, and geographical scenario knowledge for land mining activity judgment;

[0024] Combined with the temporal knowledge, determine an improvement scheme for the deep learning model, where the improvement scheme is used to improve at least one of the neck, head, and loss function of the deep learning model;

[0025] Improve the deep learning model according to the improvement scheme to obtain the object detection model.

[0026] Optionally, training the object detection model using the labeled dataset includes:

[0027] Read label data from the labeled dataset;

[0028] Input the labeled dataset into the object detection model;

[0029] Obtain the detection result data output by the object detection model;

[0030] Based on the loss function of the object detection model, calculate the training error between the detection result data and the label data;

[0031] If the training error is greater than the preset error threshold, adjust the model parameters of the object detection model according to the training error;

[0032] If the training error is less than or equal to the preset error threshold, output the model parameters of the object detection model to obtain the trained object detection model.

[0033] Optionally, generating a scene state sequence according to the recognition result includes:

[0034] Construct a state knowledge graph according to the scene element category, scene element behavior, and scene element spatial relationship of the land mining scenario;

[0035] Extract the dynamic elements, the element behaviors of the dynamic elements, and the pixel coordinates of the dynamic elements from the recognition result;

[0036] Obtain the pixel coordinates of the spatial range of static elements in the land mining scenario;

[0037] Perform spatial relationship calculation on the pixel coordinates of the dynamic elements and the pixel coordinates of the spatial range of the static elements to obtain the spatial relationship between the dynamic elements and the static elements;

[0038] Generate the scene state sequence according to the dynamic elements, the element behaviors of the dynamic elements, the spatial relationship between the dynamic elements and the static elements, and the state knowledge graph.

[0039] Optionally, generating the scene state sequence according to the dynamic elements, the element behaviors of the dynamic elements, the spatial relationship between the dynamic elements and the static elements, and the state knowledge graph includes:

[0040] Input the dynamic elements, the element behaviors, and the spatial relationship between the dynamic elements and the static elements into the state knowledge graph to obtain the scene state at the current moment;

[0041] Obtain the scene states at multiple moments in the complete time period corresponding to the video data according to the temporal relationship;

[0042] Fuse the scene states at multiple moments to generate the scene state sequence.

[0043] Optionally, perform inference calculation on the scene state sequence based on temporal knowledge to obtain the discrimination result of the land mining activity, including:

[0044] Condense the scene dynamic change characteristics presented in the process of land mining activities to obtain the temporal knowledge;

[0045] Perform inference calculation on the scene state sequence based on the temporal knowledge to obtain the discrimination result of the land mining activity;

[0046] Generate a prompt message according to the discrimination result.

[0047] Optionally, the discrimination result includes determination success, determination failure, and to be verified; performing inference calculation on the scene state sequence based on temporal knowledge to obtain the discrimination result of the land mining activity includes:

[0048] Determine the expected sequence according to the temporal knowledge;

[0049] Extract the behavior sequence from the scene state sequence;

[0050] If the behavior sequence appears completely according to the expected sequence, generate the discrimination result of determination success;

[0051] If the behavior sequence does not appear according to the expected sequence, generate the discrimination result of determination failure;

[0052] If the behavior sequence appears partially according to the expected sequence, generate the discrimination result of to be verified.

[0053] According to another aspect of the present application, there is provided a video cognition device for land mining activities, and the device includes:

[0054] An acquisition module, configured to acquire an image data set, where the image data set includes feature pictures extracted from video data, and the feature pictures include image targets corresponding to mining construction machinery;

[0055] A labeling module, configured to perform behavior labeling on the image data set according to a set of behavior states to obtain a labeled data set, where the set of behavior states is a set of states obtained by condensing the behavior process characteristics of the mining construction machinery according to temporal knowledge;

[0056] A training module, configured to train an object detection model using the labeled data set, where the object detection model is a deep learning model improved based on the behavior recognition characteristics of mining construction machinery;

[0057] A detection module, configured to input video data of a land scene into the trained object detection model to obtain an identification result output by the object detection model, where the identification result includes dynamic elements included in the current land scene, the element behaviors of the dynamic elements, and the pixel coordinates of the dynamic elements;

[0058] A discrimination module, configured to generate a scene state sequence according to the identification result, and perform inference calculation on the scene state sequence based on temporal knowledge to obtain a discrimination result of the land mining activity.

[0059] According to another aspect of the present application, there is provided a computer device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, where when the processor executes the program, the above-mentioned video cognition method for land mining activities is implemented.

[0060] According to still another aspect of the present application, there is provided a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned video cognition method for land mining activities is implemented.

[0061] By means of the above technical solutions, the embodiments of the present application provide a video cognition method and device for land mining activities. At the data level, the method comprehensively depicts the interaction behavior characteristics between construction machinery from two dimensions of time and space by constructing an image data set of construction machinery behaviors for a land mining scene. At the algorithm level, the method can, in view of the characteristics of frequent movement and high timeliness requirements of dynamic elements in land scene video data, fuse temporal knowledge to make targeted improvements to the object detection model, and based on a hybrid intelligent element behavior recognition algorithm, finely capture the subtle feature differences of the same element under different behaviors. At the inference stage, the method can, by constructing a knowledge graph and a knowledge inference mechanism, deeply integrate video dynamic information with domain knowledge, overcome the problem of incomplete recognition of element interaction relationships, and achieve accurate recognition of land mining activities.

[0062] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented in accordance with the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specific embodiments of this application are specifically given. Brief Description of the Drawings

[0063] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0064] Figure 1 It is a schematic structural diagram of an electronic system for identifying land mining activities provided by an embodiment of this application;

[0065] Figure 2 It is a flow chart of video recognition process of land mining activities provided by an embodiment of this application;

[0066] Figure 3 It is a schematic flow chart of the method for video recognition of land mining activities provided by an embodiment of this application;

[0067] Figure 4 It is a schematic flow chart of the process for generating an image data set provided by an embodiment of this application;

[0068] Figure 5 It is a schematic flow chart of the process for generating an annotation data set provided by an embodiment of this application;

[0069] Figure 6 It is a schematic flow chart of the improvement process of a deep learning model provided by an embodiment of this application;

[0070] Figure 7 It is a schematic flow chart of the training process of an object detection model provided by an embodiment of this application;

[0071] Figure 8 It is a schematic structural diagram of a knowledge graph of the scene state provided by an embodiment of this application;

[0072] Figure 9 It is a schematic diagram of the principle for identifying land mining activities provided by an embodiment of this application;

[0073] Figure 10 It is a schematic structural diagram of a device for video recognition of land mining activities provided by an embodiment of this application. Detailed Embodiments

[0074] In the following, this application will be described in detail with reference to the drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.

[0075] In the embodiment of the present application, the land exploitation activity refers to the process of developing and utilizing land resources in a specific monitoring area. Land exploitation activities will change the ecological environment and actual use of specific land resources. According to the relevant planning requirements of a specific area, land exploitation activities may include compliant land exploitation activities and non-compliant land exploitation activities.

[0076] In order to monitor whether there is land mining activity in a specific area, an electronic system for identifying land mining activities can be constructed according to the deployment of relevant monitoring equipment in the specific area. The system may include multiple electronic devices with different functions. For example, Figure 1 As shown, the system may include data acquisition equipment, data processing equipment, data display equipment, etc.

[0077] The data collection equipment can be used to collect data related to land mining activities in a specific area and send the collected data to the data processing equipment. For example, the data collection equipment may include a sensor for collecting image data, such as a camera, a remote sensing aircraft, a visual satellite, etc.

[0078] In some embodiments, the data collection device may include multiple sensors for collecting different types of data. In the process of identifying land mining activities, the data collection device may collect multiple different types of data through multiple sensors to obtain multi-dimensional data. For example, the data collection device may also include a vibration sensor for detecting ground vibration signals, a microphone for collecting sound signals in a specific area, and an optical heat dissipation sensor for monitoring dust concentration in a specific area.

[0079] It should be noted that in the embodiments of the present application, the land mining activity identification process is described by taking the collection of image data and the analysis of the image data as an example. It should be understood that the data collected by the data collection device is not limited to image data, but may also include vibration signals, sound signals, dust concentration and other data. In addition, in order to adapt to multi-dimensional data types, a more accurate and complex analysis process can be performed on the basis of the analysis of image data in combination with other forms of data. Therefore, other identification methods associated with those skilled in the art based on the image data collection and analysis described in the embodiments of the present application also fall within the scope of protection of the present application.

[0080] The data processing device is used to receive the data collected by the data collection device and analyze the collected data to extract signal features related to the land mining process from the collected data. Obviously, for different types of collected data, the expression form of the signal features is also different. For example, when the collected data is image data, the features related to the land mining process are a set of pixel points that can represent mining engineering machinery.

[0081] Among them, according to the different processes of land exploitation, the specific types of construction machinery for exploitation are also different. For example, when the land exploitation activity is illegal excavation and mining on cultivated land, the construction machinery for exploitation can be construction machinery such as excavators and trucks. Correspondingly, when the data processing device performs feature extraction, it can identify whether there is an illegal excavation and mining activity on cultivated land in the corresponding scenario by determining whether there are graphic features corresponding to construction machinery such as excavators and trucks in the image data.

[0082] By extracting the signal features related to the land exploitation process, the data processing device can determine whether there is a land exploitation activity in the corresponding area and generate a determination result. Then, the determination result is sent to the data display device. The data display device can display the determination result by means of displaying a picture, playing audio, triggering an alarm, etc. For example, the data display device can include a monitor for displaying a monitoring interface. When the data processing device generates a determination result, it can send the determination result information to the data display device to display the determination result on the monitoring interface displayed by the data display device.

[0083] For different purposes of the land exploitation process, the data display device can display the determination result in different display ways. For example, when the land exploitation activity recognition method is used to determine whether there is an illegal excavation and mining activity on cultivated land in a monitoring area, if the determination result is that there is an illegal excavation and mining activity in this area, the data display device can display the determination result through a display picture with an obvious warning effect to achieve a prompting or alarm effect.

[0084] In some embodiments, in order to quickly identify land exploitation activities, the intelligent video monitoring method can be used to focus on the typical features in the monitoring area. For example, construction machinery such as excavators and trucks in the scene of illegal excavation and mining on cultivated land are typical features. The intelligent video monitoring method can obtain an image of the cultivated land scene in the monitoring area through a monitoring camera to get the cultivated land scene image. And based on the principle of machine vision image recognition, it can identify whether there are typical features corresponding to construction machinery such as excavators and trucks in the cultivated land scene image. When the typical features in the monitoring area are identified, the intelligent video monitoring system will perform a warning process.

[0085] However, the intelligent video monitoring method belongs to extensive monitoring, resulting in frequent false alarm problems in the intelligent video monitoring method. For example, when there are passing construction machinery in the monitored area, the intelligent video monitoring system will also extract the illegal excavation and mining features in the video data, thus causing false alarms of the intelligent video monitoring system for illegal excavation and mining on cultivated land and reducing the recognition accuracy of illegal excavation and mining activities on cultivated land.

[0086] To solve the problem of inability to accurately identify land mining activities, some embodiments of the present application provide a video recognition method for land mining activities. The method can be applied to a data processing device, such as Figure 2 、 Figure 3 shown. The method includes:

[0087] S100. Obtain an image data set.

[0088] In the process of identifying land mining activities, an image data set can also be obtained first. The image data set includes feature pictures extracted from video data, and the feature pictures include image targets corresponding to mining construction machinery.

[0089] Such as Figure 4 shown. In some embodiments, in order to obtain an image data set, the data processing device can first obtain a collection rule for guiding the collection of video data of mining construction machinery. When identifying land mining activities, a set of collection rules for guiding the collection of video data of construction machinery can be refined for the land mining scenario. For example, the collection rules can include element categories, element behaviors, scene perspectives, scene scales, etc.

[0090] According to specific identification requirements, the element categories involved in the land mining scenario can include dynamic elements and static elements. Among them, static elements refer to targets whose positions do not move or whose moving speeds and moving positions are less than a preset threshold in the land mining scenario. Dynamic elements refer to targets whose moving speeds and / or moving positions are less than a preset threshold in the land mining scenario. For example, in the scenario of illegal excavation and mining in cultivated land, the element categories mainly include dynamic elements corresponding to mining construction machinery such as excavators and trucks, and static elements such as cultivated land, buildings, and plants.

[0091] Element behavior refers to the specific behavior of a target in the land mining scenario, that is, due to land mining activities, the movement of the target position or the change of the target form and action in the scenario. For example, in the scenario of illegal excavation and mining in cultivated land, the excavator mainly exhibits behavior characteristics such as excavation and loading, while the truck exhibits behavior characteristics such as docking, loading, and transportation. These dynamic behaviors are closely related to the cultivated land area, operation frequency, and interaction mode, and can be used for the identification and analysis of illegal operations.

[0092] Scene perspective refers to the shooting angle when collecting images in the land mining scenario. For example, in the scenario of illegal excavation and mining in cultivated land, the shooting perspective of the camera usually has the characteristics of wide-area coverage, elevated overlooking, or fixed monitoring point layout to ensure a comprehensive perception of excavators, trucks, and the operation area.

[0093] The scene scale refers to the scale characteristics of the image acquisition device during image shooting in the land mining scene, including shooting parameters, shooting resolution, shooting accuracy, etc. For example, in the scene of illegal excavation and mining of cultivated land, the shooting scale of the camera has multi-scale characteristics. It is necessary to cover a large area to monitor the overall operation dynamics, and also need to have a fine scale to capture the specific operation behaviors of excavators and trucks, so as to support the accurate identification and spatio-temporal analysis of illegal operations.

[0094] While obtaining the acquisition rules for guiding the video data acquisition of mining construction machinery, the initial data set can also be obtained by using the operation video of the mining construction machinery and removing the invalid videos in the operation video.

[0095] For example, the original video data containing the operation of equipment such as excavators and trucks can be collected from public Internet data sources and relevant construction machinery operation video platforms, and preliminary screening can be carried out on it. During the screening process, videos that are unclear, have too low resolution, and have serious picture jitter are removed to ensure the data quality and usability, and an initial video data set containing mining construction machinery content is constructed.

[0096] After obtaining the initial data set, the data processing device can perform video screening on the initial data set based on the acquisition rules to obtain a video data set that conforms to the land mining scene. That is, the data processing device can further screen the video data according to the initial data set a and the construction machinery data acquisition rules to obtain a high-quality video data set b of the illegal excavation and mining scene of cultivated land that meets the research requirements.

[0097] After obtaining the video data set, frame extraction can be performed on the video data set according to a preset frame extraction interval to obtain the image data set. For example, based on the time characteristics of the construction machinery operation behavior and the video frame rate characteristics in the data set b, a suitable time interval is selected to perform frame extraction on the video, and representative image frames are extracted. During the frame extraction process, the continuity of the operation behavior and the change of key actions can be considered to ensure that the extracted images can cover the complete characteristics of the excavator, truck and their interaction process. Finally, an image data set c containing mining construction machinery content is obtained.

[0098] S200. Perform behavior annotation on the image data set according to the behavior state set to obtain an annotated data set.

[0099] After acquiring the image data set, the data processing device can integrate the element behaviors of expert knowledge to perform data annotation, that is, to perform behavior annotation on the image data set. Among them, behavior annotation refers to the process of adding behavior labels to images in the image data set. In order to obtain the labeled data, the data processing device can perform behavior annotation according to the behavior state set, wherein the behavior state set is a state set obtained by condensing the behavior process characteristics of the mining engineering machinery according to the time series knowledge.

[0100] In some embodiments, when performing behavior annotation, the performance and explainability of the recognition process can be improved by condensing a variety of temporal knowledge, where temporal knowledge includes deep learning knowledge for model improvement, video image knowledge for feature behavior classification, and geographic scene knowledge for judging illegal mining of cultivated land. Figure 5 As shown, when integrating expert knowledge to perform element behavior annotation, the data processing device can first summarize the behavioral state set of the mining engineering machinery from the aspects of element form and interaction mode.

[0101] Based on the analysis of construction machinery behavior, the typical operation modes of excavators and trucks are sorted out to analyze their behavior characteristics in different scenarios. The operation behavior states of excavators and trucks are systematically summarized from the two dimensions of element form and interaction mode, and a preliminary behavior state set s is constructed.

[0102] The data processing equipment can comprehensively consider the element form, interaction mode and time series characteristics of the engineering machinery to generate the behavior state set of mining engineering machinery. Among them, mining engineering machinery includes excavators and trucks. The behaviors of excavators include digging, full shovel, loading, and empty shovel. The behaviors of trucks include empty, loading, loading, and unloading.

[0103] Among them, the element form refers to the specific form information of the target in the land mining scene. The interaction mode refers to the element information such as the relative movement and interaction between two or more dynamic element targets in the land mining scene. For example, the element form can include the boom state and bucket posture of the excavator, the loading state and docking method of the truck, etc. The interaction mode can include the loading interaction between the excavator and the truck, the transportation path of the truck, etc.

[0104] After constructing the behavior state set, the behavior process characteristics of the mining engineering machinery can be condensed according to the time series knowledge. The data processing equipment can combine the real scene of land mining activities, based on the actual operation process and the dynamic evolution characteristics of the equipment, and condense the behavior process characteristics of the mining engineering machinery in the land mining activities according to the time series knowledge.

[0105] Based on the behavioral process characteristics, the set of behavioral states is further refined to obtain a refined set of behavioral states. For example, starting from the time series relationship of operation start, excavation, loading, transportation, unloading, etc., the data processing device can analyze the logical associations and conversion rules of each behavioral state, and then extract behavioral characteristics with key recognition value. Based on the above analysis results, the set of behavioral states s is optimized and refined by removing redundant states, merging similar states, and enhancing key behavioral characteristics related to illegal operation modes, forming an optimized set of behavioral states s1.

[0106] Perform behavioral annotation on the image dataset according to the refined set of behavioral states to obtain an annotated dataset. That is, based on the constructed set of behavioral states s1, the image dataset c containing mining construction machinery is finely annotated. During the annotation process, the video timestamp and image content can be combined to identify the specific behavioral states of the mining construction machinery frame by frame and perform classification annotation to ensure the high quality and high consistency of the dataset.

[0107] In some embodiments, when performing behavioral annotation on the image dataset, a combination of manual verification and automatic annotation can also be used to ensure the accuracy and robustness of the annotation results, obtaining an annotated dataset d.

[0108] S300. Train a target detection model using the annotated dataset.

[0109] After obtaining the annotated dataset, the data processing device can use the annotated dataset for model training to obtain a target detection model that can identify dynamic elements, element behaviors, and pixel coordinates of elements in video data. Among them, the target detection model is a deep learning model improved based on the behavioral recognition characteristics of mining construction machinery.

[0110] In order to train the target detection model, in some embodiments, the data processing device can construct temporal knowledge based on the behavioral recognition characteristics of the mining construction machinery in the land mining scenario. By analyzing the characteristics of behavioral recognition of construction machinery in the land mining scenario and combining the model characteristics in the field of deep learning, temporal knowledge for identifying the private excavation and illegal mining scenario of cultivated land can be refined. Among them, the prior knowledge includes deep learning knowledge for model improvement, video image knowledge for element behavior classification, and geographical scenario knowledge for land mining activity judgment.

[0111] For deep learning knowledge, since the purpose of scene recognition during land mining activity is to timely warn of land mining behaviors, an appropriate model should be selected to achieve rapid recognition and segmentation of scene elements while meeting a certain accuracy. Selecting a suitable pre-trained model for feature extraction is the basis of element segmentation. Generally, larger models have stronger extraction capabilities but lower real-time inference performance.

[0112] To detect features of different scales to improve the detection effect, the detected features can be fused and enhanced, and an appropriate feature fusion model needs to be determined to make full use of the extracted features for the recognition and segmentation of small dynamic elements in outdoor scenes. It is also necessary to further extract high-level features from the fused features to obtain the category, location, and confidence of the segmentation. Since the same element in different scales in a large outdoor scene still has the same image features, a unified feature expression should be adopted. The Loss function of the object detection model is used to measure the difference between the prediction result and the ground truth. The recognition of the specific behavior of elements in the land mining scene is mainly based on the element contour. Therefore, the element contour information should be fully considered when designing the Loss function.

[0113] Regarding video image knowledge, there is spatial continuity between the previous frame and the next frame, enabling the spatial continuity of video frames to be used for performing element trajectory analysis and behavior analysis. Different elements can have different color and texture features, so different scene elements can be distinguished by different color and texture features on the video image. The same element may present different shape contour features. For example, for an excavator in operation, its shape contour will change within a certain range, and these shape contours become the key to identifying the behavior of the excavator. The pixel size of an element in a video image is not only related to its own size but also related to its distance from the camera. It is necessary to fully identify ground object elements at different scales, especially small dynamic elements in the land mining scene.

[0114] Regarding geographical scene knowledge, the spatial relationship between elements can reflect their interactions. By analyzing the topological spatial relationship between the excavator and various ground object elements, the mined elements can be judged. Different elements have different change frequencies. For example, elements such as cultivated land and water bodies have a lower change frequency; elements such as trucks and excavators have a higher change frequency. Different elements present different spatial feature laws. For example, elements such as cultivated land and water bodies can have arbitrary shapes and sizes, while elements such as trucks and excavators have aspect ratio limitations.

[0115] After constructing the temporal knowledge, the data processing device can combine the temporal knowledge to determine an improvement plan for the deep learning model and perform improvements on the deep learning model according to the improvement plan to obtain the object detection model. Among them, the improvement plan is used to improve at least one of the neck, head, and loss function of the deep learning model. The improvement plan can be optimized for the adaptability of the object detection model in the land mining scene. The improvement plan can be condensed starting from three core modules: the neck, the head, and the loss function, to solve problems such as large target scale changes, complex operation behaviors, and blurred features of distant targets.

[0116] As shown Figure 6 In some embodiments, based on the proposed improvement scheme, the Neck structure of the object detection model can be optimized to enhance the feature extraction and multi-scale fusion capabilities. When improving the Neck of the object detection model, any one of the feature fusion algorithms such as RepNCSPFPN, BiFPN, AFPN, GhostSlimFPN, EffQAFPN can be used.

[0117] For example, by introducing a more effective feature pyramid structure (such as BiFPN, PAFPN, etc.), optimizing the feature channel attention mechanism, and enhancing the feature cross-layer fusion ability, the model's perception ability for objects of different scales can be improved, especially the detection effect for distant and small-scale objects. After improving the Neck part of the model, the obtained object detection model is model-N.

[0118] Based on model-N, the Head structure of the object detection model can be optimized to further enhance the object classification and localization capabilities. When improving the Head of the object detection model, any one of the detection head enhancement algorithms such as ShareSepHead, LADH, AsDDet, SERDet, LiteShiftHead, DynamicHead can be used.

[0119] For example, when optimizing the Head structure of the object detection model, the optimization strategy can include adjusting the anchor-free design, improving the regression branch strategy, enhancing the key point detection ability, etc., to improve the fineness of object recognition and the accuracy of bounding box regression. At the same time, combined with the characteristics of the illegal mining scenario, a spatial attention mechanism is introduced to enhance the model's attention to key behavior parts such as the excavator bucket and the truck carriage. After optimizing the Head structure of the model, the obtained model is model-NH.

[0120] Based on model-NH, the Loss design of the object detection model can also be optimized to improve the training stability and detection robustness of the model. Considering the model's ability to recognize more features, when improving the model Loss calculation method, it can include but is not limited to calculation methods such as GWD, KLD, Repulsion, XIoU, ShapeIoU.

[0121] For example, by adopting Focal Loss to solve the problem of imbalance between positive and negative samples, adjusting the classification loss function, and introducing IoU-based Loss to improve the target localization accuracy, thereby optimizing the bounding box regression loss, and designing a specific loss function by combining the characteristics of illegal mining behaviors to enhance the model's recognition ability for key behavior categories. After optimizing the Loss of the model, the obtained model is model-NHL.

[0122] After improving the deep learning model to obtain an object detection model, the data processing device can use the constructed high-quality labeled dataset D to train and optimize the improved object detection module model-NHL to further enhance the model's ability to recognize construction machinery behaviors.

[0123] As Figure 7 shown, in some embodiments, to perform model training, when using the labeled dataset to train the object detection model, the data processing device can read the label data from the labeled dataset, then input the labeled dataset into the object detection model, and obtain the detection result data output by the object detection model. Then, based on the loss function of the object detection model, calculate the training error between the detection result data and the label data.

[0124] By comparing the calculated training error with a preset error threshold, if the training error is greater than the preset error threshold, adjust the model parameters of the object detection model according to the training error. When the training error is greater than the preset error threshold, it indicates that the current object detection model has not converged. Therefore, backpropagation can be performed based on the training error to adjust the model parameters of the object detection model.

[0125] After adjusting the model parameters, the data processing device can input the labeled dataset into the object detection model with adjusted model parameters again, obtain the detection result data output by the model, and recalculate the training error and compare it with the preset error threshold for iterative training. After multiple iterative trainings, if the training error is less than or equal to the preset error threshold, output the model parameters of the object detection model to obtain the trained object detection model.

[0126] During the training process, data augmentation strategies such as temporal augmentation and fuzzy object augmentation can also be adopted to improve the generalization ability of the model, and at the same time, optimize the hyperparameter configuration to improve the convergence speed and detection accuracy of the model. After sufficient training and verification, finally, an object detection model model-CMB for recognizing construction machinery behaviors in the scenario of illegal mining in cultivated land is obtained.

[0127] S400. Input the video data of the land scene into the trained target detection model to obtain the recognition result output by the target detection model.

[0128] After training the target detection model, the data processing device can apply the trained target detection model to identify land mining activities. To this end, the data processing device can obtain the video data of the land scene in the monitored area through the data acquisition device and input the video data of the land scene into the trained target detection model.

[0129] After inputting the video data into the trained target detection model, the target detection model can extract key frames from the video data and perform image recognition on each key frame image to identify dynamic elements and the associated information of the dynamic elements in the key frame image. That is, the target detection model can output the recognition result for the input video data. In some embodiments, when the target detection model identifies the behaviors of construction machinery in the land mining scene, it not only outputs the element category and element behavior, but also outputs the pixel coordinates after the element is imaged, so that the spatial topological relationship between elements can be calculated. That is, the recognition result includes the dynamic elements included in the current land scene, the element behaviors of the dynamic elements, and the pixel coordinates of the dynamic elements.

[0130] For example, in the scenario of detecting illegal excavation and mining in cultivated land, the core element categories, spatial relationships, and element behaviors that make up the scenario can be analyzed and sorted out. Among them, the element categories include dynamic elements such as excavators and trucks, and static elements such as cultivated land and roads. The spatial relationships include the intersection, inclusion, separation, etc. between dynamic elements and static elements; the element behaviors involve dynamic operation modes such as the excavation and loading of excavators, and the empty load and load of trucks.

[0131] Then, using the trained target detection model model-CMB, the input video frame can be subjected to target detection to automatically identify dynamic elements such as excavators and trucks included in the current scene, as well as specific operation behaviors such as excavation and loading.

[0132] S500. Generate a scene state sequence according to the recognition result, and perform inference calculation on the scene state sequence based on temporal knowledge to obtain the discrimination result of the land mining activity.

[0133] After obtaining the recognition result output by the target detection model, the data processing device can generate a scene state sequence according to the recognition result. Among them, the scene state sequence is used to characterize the specific behaviors of dynamic elements in the scene in terms of time series. Therefore, the scene state sequence is a combination of multiple behavior features arranged in time series.

[0134] In some embodiments, to generate a scene state sequence, the data processing device may construct a state knowledge graph based on the scene element categories, scene element behaviors, and scene element spatial relationships of the land mining scene. For example, as Figure 8 shown, by analyzing the core element categories, spatial relationships, and element behaviors that make up the scene, a knowledge graph KG representing the scene state at a certain moment is constructed. Among them, the element categories include dynamic elements such as excavators and trucks, and static elements such as cultivated land and roads; the spatial relationships include relationships such as intersection, inclusion, and separation between dynamic elements and static elements; the element behaviors involve dynamic operation modes such as excavation and loading of excavators, and empty load and load of trucks.

[0135] Then, extract the dynamic elements, the element behaviors of the dynamic elements, and the pixel coordinates of the dynamic elements from the recognition results. That is, use the trained object detection model model-CMB to perform object detection on the input video frame to automatically identify the dynamic elements included in the current scene, such as excavators and trucks. And the specific operation behaviors of the dynamic elements, such as excavation and loading.

[0136] Then, obtain the pixel coordinates of the spatial range of the static elements in the land mining scene, and perform spatial relationship calculation on the pixel coordinates of the dynamic elements and the pixel coordinates of the spatial range of the static elements to obtain the spatial relationship between the dynamic elements and the static elements.

[0137] For example, in combination with the recognition results of model-CMB, that is, the object detection results, further extract the pixel coordinate information of the dynamic elements, and perform spatial relationship calculation with the pixel coordinates of the pre-calculated spatial range of the static elements. When performing spatial relationship calculation, geometric methods such as point-polygon relationship calculation and nearest neighbor search can be used to determine the spatial relationship between dynamic elements such as excavators and trucks relative to static elements such as cultivated land and roads. The spatial relationship can include specific spatial relationships such as whether it enters the cultivated land area and whether it is close to the cultivated land, so as to obtain the interaction situation between each dynamic element and static element in the scene.

[0138] After obtaining the spatial relationship between the dynamic elements and the static elements, the scene state sequence can be generated according to the dynamic elements, the element behaviors of the dynamic elements, the spatial relationship between the dynamic elements and the static elements, and the state knowledge graph.

[0139] In some embodiments, the data processing device may input the dynamic elements, the element behaviors, and the spatial relationship between the dynamic elements and the static elements into the state knowledge graph to obtain the scene state at the current moment.

[0140] For example, a data processing device may input the identified element categories and element behaviors, as well as the calculated element spatial relationships, into the constructed knowledge graph (KG) to form the scene state SS (Scene State) at the current moment. By structuring the scene state, the scene information of the current frame can be represented, including the construction machinery participating in the operation, its operation state, and its positional relationship with static elements such as cultivated land, laying a foundation for subsequent scene dynamic modeling.

[0141] Repeat the above steps of using the trained object detection model model-CMB to perform object detection on the input video frame, extract the pixel coordinate information of dynamic elements, calculate the spatial relationship with the pixel coordinates of the pre-calculated static element spatial range, and form the scene state SS at the current moment. The complete video sequence can be processed frame by frame, and the scene state SS of each time frame can be constructed in turn. Then, according to the temporal relationship, the scene states at multiple moments in the complete time period corresponding to the video data are obtained, and the scene state sequence is generated by fusing the scene states at multiple moments.

[0142] For example, after constructing the scene state SS of each time frame by processing the complete video sequence frame by frame, a scene state sequence SSS (Scene State Sequence) of a complete video can be formed. This scene state sequence can reflect the evolution of operation elements, behavior changes, and spatial interaction patterns during the land mining process, providing data support for subsequent dynamic feature extraction and determination of illegal excavation and mining of cultivated land.

[0143] After generating the scene state sequence, the data processing device can perform inference calculations on the scene state sequence based on temporal knowledge to obtain the discrimination result of the land mining activity. That is, based on the constructed scene state sequence SSS, analyze the temporal evolution pattern of the illegal excavation and mining activities of cultivated land, and extract the typical dynamic change characteristics presented during the illegal excavation and mining activities. In some embodiments, key operation stages such as continuous operation of excavators and frequent loading and transportation of trucks and their transfer rules can be identified through temporal pattern mining such as behavior pattern analysis and movement trajectory analysis. By extracting the temporal knowledge SK (Temporal Knowledge) for determining land mining activities, a temporal discrimination rule is formed.

[0144] To this end, the data processing device can first distill the dynamic change characteristics of the scenes presented during the land mining activities to obtain the temporal knowledge; and then perform inference calculations on the scene state sequence based on the temporal knowledge to obtain the discrimination result of the land mining activities. That is, the data processing device can perform inference calculations based on the extracted temporal knowledge SK and the scene state sequence SSS to determine whether there is a land mining behavior. During the inference process, the data processing device can combine the state sequence in the knowledge graph and use methods such as rule matching, temporal pattern inference, and probability inference to comprehensively analyze the operation behavior and the evolution pattern of the spatial relationship of the mining construction machinery, so as to output the discrimination result of the land mining activities and provide intelligent support for the monitoring of land mining activities.

[0145] Among them, the discrimination result includes determination success, determination failure, and to-be-verified. The data processing device can determine the expected sequence according to the temporal knowledge and extract the behavior sequence from the scene state sequence. Then, the behavior sequence and the expected sequence are compared. If the behavior sequence appears completely according to the expected sequence, the discrimination result of determination success is generated. For example, as Figure 9 shown, by using the temporal knowledge to determine that the expected sequence of the illegal private excavation of cultivated land activity is "excavation - full shovel - loading", then when the behavior sequence determined according to the scene state sequence is also "excavation - full shovel - loading", it is determined that the behavior sequence appears completely according to the expected sequence. Therefore, the discrimination result of determination success can be generated.

[0146] If the behavior sequence does not appear according to the expected sequence, the discrimination result of determination failure is generated. For example, if the expected sequence is "excavation - full shovel - loading" and the behavior sequence determined according to the scene state sequence is "excavation - full shovel - excavation", it is determined that the behavior sequence does not appear according to the expected sequence. Therefore, the discrimination result of determination failure can be generated.

[0147] If the behavior sequence appears partially according to the expected sequence, the discrimination result of to-be-verified is generated. For example, if the expected sequence is "excavation - full shovel - loading" and the behavior sequence determined according to the scene state sequence is "excavation - full shovel", it is determined that the behavior sequence appears partially according to the expected sequence. Therefore, the discrimination result of to-be-verified is generated.

[0148] After obtaining the discrimination result of the land mining activity, prompt information can be generated according to the discrimination result, and different subsequent operations can be triggered according to different discrimination results. For example, for the identification of illegal excavation in cultivated land, when the discrimination result is determined to be successful, the current scene can be identified as an illegal excavation scene, and the start time of the violation can be returned to prompt the illegal excavation activity in cultivated land and provide intelligent support for the monitoring of illegal excavation in cultivated land. When the discrimination result is determined to be failed, the judgment information can be cleared and subsequent judgments can be made. When the discrimination result is to be verified, the data processing device can continuously perform behavior verification, that is, to obtain the final discrimination result of the judgment, continue to generate the scene state sequence in the manner provided in the above embodiment, and compare the behavior sequence corresponding to the scene state sequence with the expected sequence until a discrimination result of successful determination or failed determination is generated.

[0149] By applying the technical solutions of the above embodiments, the land mining activity video cognition method provided by the embodiments of the present application can solve the problems of low recognition accuracy, high missed detection rate, and low automation degree in video recognition monitoring. It combines natural intelligence and artificial intelligence, and performs data set annotation, model construction, and knowledge reasoning under the guidance of knowledge. By introducing behavior-level spatio-temporal element annotation at the data level and combining deep learning with domain knowledge at the algorithm level, the recognition method can more precisely capture the subtle feature differences of the same element under different behaviors and improve interpretability. The improved deep learning model performs excellently in the element behavior recognition task. At the same time, the introduction of the knowledge graph and knowledge reasoning provides deeper semantic support for the understanding of video dynamic information, thereby realizing the accurate dynamic recognition of land mining activities.

[0150] Furthermore, as a specific implementation of the land mining activity video cognition method described in the above embodiments, some embodiments of the present application also provide a land mining activity video cognition device, as Figure 10 shown. The device includes:

[0151] An acquisition module, configured to acquire an image data set, where the image data set includes feature pictures extracted from video data, and the feature pictures include image targets corresponding to mining engineering machinery;

[0152] A labeling module, configured to perform behavior labeling on the image data set according to a behavior state set to obtain a labeled data set, where the behavior state set is a state set obtained by condensing the behavior process characteristics of the mining engineering machinery according to temporal sequence knowledge;

[0153] A training module, configured to train a target detection model using the labeled data set, where the target detection model is a deep learning model improved based on the behavior recognition characteristics of mining engineering machinery;

[0154] A detection module, used for inputting the video data of the land scene into the trained target detection model to obtain a recognition result output by the target detection model, wherein the recognition result includes the dynamic elements contained in the current land scene, the element behavior of the dynamic elements, and the pixel coordinates of the dynamic elements;

[0155] The discrimination module is used to generate a scene state sequence according to the recognition result, and perform inference calculation on the scene state sequence based on time series knowledge to obtain a discrimination result of the land mining activity.

[0156] By applying the technical solutions of the above-mentioned embodiments, the land mining activity video recognition device provided in the embodiments of the present application, at the data level, constructs an engineering machinery behavior image dataset for land mining scenes, and comprehensively characterizes the interactive behavior characteristics between engineering machinery from the two dimensions of time and space. At the algorithm level, the device can integrate time series knowledge to make targeted improvements to the target detection model based on the characteristics of frequent movement and high timeliness requirements of dynamic elements in land scene video data, and accurately capture the subtle feature differences of the same element under different behaviors based on the element behavior recognition algorithm of hybrid intelligence. In the reasoning stage, the device can deeply integrate video dynamic information with domain knowledge by constructing a knowledge graph and knowledge reasoning mechanism, overcome the problem of incomplete recognition of element interaction relationships, and realize accurate recognition of land mining activities.

[0157] It should be noted that, for other corresponding descriptions of the functional units involved in a land mining activity video recognition device provided in an embodiment of the present application, reference can be made to the corresponding descriptions in the land mining activity video recognition method provided in the above embodiment, and will not be repeated here.

[0158] The embodiment of the present application also provides a computer device, which can be a personal computer, a server, a network device, etc. The computer device includes a bus, a processor, a memory and a communication interface, and can also include an input and output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.

[0159] Those skilled in the art can understand that the structure of the above computer device is only a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components, or combine some components, or have different component arrangements.

[0160] In one embodiment, a computer-readable storage medium is further provided. The computer-readable storage medium may be non-volatile or volatile, and has a computer program stored thereon. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0161] In one embodiment, a computer program product is further provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0162] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0163] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments.

[0164] Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc.

[0165] Volatile memory may include random access memory (RAM), external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0166] The databases involved in the embodiments provided in this application may include at least one of relational databases and non-relational databases. Non-relational databases may include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application may be general-purpose processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0167] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0168] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A video recognition method for land mining activities, characterized in that, The method includes: Obtaining an image dataset, where the image dataset includes feature pictures extracted from video data, and the feature pictures include image targets corresponding to mining construction machinery; Performing behavior annotation on the image dataset according to a set of behavior states to obtain an annotated dataset, where the set of behavior states is a set of states obtained by condensing the behavior process characteristics of the mining construction machinery according to temporal knowledge; Training an object detection model using the annotated dataset, where the object detection model is a deep learning model improved based on the behavior recognition characteristics of mining construction machinery; Inputting the video data of the land scene into the trained object detection model to obtain the recognition result output by the object detection model, where the recognition result includes dynamic elements included in the current land scene, the element behaviors of the dynamic elements, and the pixel coordinates of the dynamic elements; Generating a scene state sequence according to the recognition result, and performing inference calculation on the scene state sequence based on temporal knowledge to obtain a discrimination result of the land mining activity.

2. The method according to claim 1, characterized in that, Obtaining an image dataset, including: Obtaining a collection rule for guiding the collection of video data of mining construction machinery, and the operation video of the mining construction machinery; Removing invalid videos from the operation video to obtain an initial dataset; Performing video screening on the initial dataset based on the collection rule to obtain a video dataset that conforms to the land mining scene; Performing frame extraction on the video dataset at a preset frame extraction interval to obtain the image dataset.

3. The method according to claim 1, wherein Performing behavior annotation on the image dataset according to a set of behavior states to obtain an annotated dataset, including: Inducing the set of behavior states of the mining construction machinery from aspects of element form and interaction mode; Condensing the behavior process characteristics of the mining construction machinery according to temporal knowledge; Performing condensation on the set of behavior states based on the behavior process characteristics to obtain a condensed set of behavior states; Performing behavior annotation on the image dataset according to the condensed set of behavior states to obtain an annotated dataset.

4. The method according to claim 1, characterized in that Before training the object detection model using the annotated dataset, the method further includes: Constructing temporal knowledge based on the behavior recognition characteristics of the mining construction machinery in the land mining scene, where the temporal knowledge is used to recognize the land mining scene; the prior knowledge includes deep learning knowledge for model improvement, video image knowledge for element behavior classification, and geographical scene knowledge for land mining activity judgment; Combining the temporal knowledge to determine an improvement scheme for the deep learning model, where the improvement scheme is used to improve at least one of the neck, head, and loss function of the deep learning model; Improving the deep learning model according to the improvement scheme to obtain the object detection model.

5. The method according to claim 1, wherein Training the object detection model using the annotated dataset, including: Reading label data from the annotated dataset; Inputting the annotated dataset into the object detection model; Obtaining the detection result data output by the object detection model; Calculating the training error between the detection result data and the label data based on the loss function of the object detection model; If the training error is greater than a preset error threshold, adjust the model parameters of the target detection model according to the training error; If the training error is less than or equal to the preset error threshold, output the model parameters of the target detection model to obtain the trained target detection model.

6. The method according to claim 1, characterized in that Generate a scene state sequence according to the recognition result, including: Construct a state knowledge graph according to the scene element categories, scene element behaviors, and scene element spatial relationships of the land mining scene; Extract the dynamic elements, the element behaviors of the dynamic elements, and the pixel coordinates of the dynamic elements from the recognition result; Obtain the pixel coordinates of the spatial range of static elements in the land mining scene; Perform a spatial relationship calculation on the pixel coordinates of the dynamic elements and the pixel coordinates of the spatial range of the static elements to obtain the spatial relationship between the dynamic elements and the static elements; Generate the scene state sequence according to the dynamic elements, the element behaviors of the dynamic elements, the spatial relationship between the dynamic elements and the static elements, and the state knowledge graph.

7. The method according to claim 6, wherein Generate the scene state sequence according to the dynamic elements, the element behaviors of the dynamic elements, the spatial relationship between the dynamic elements and the static elements, and the state knowledge graph, including: Input the dynamic elements, the element behaviors, and the spatial relationship between the dynamic elements and the static elements into the state knowledge graph to obtain the scene state at the current moment; According to the time series relationship, obtain the scene states at multiple moments in the complete time period corresponding to the video data; Fuse the scene states at multiple moments to generate the scene state sequence.

8. The method according to claim 1, characterized in that Perform an inference calculation on the scene state sequence based on temporal knowledge to obtain a discrimination result of the land mining activity, including: Condense the scene dynamic change characteristics presented during the land mining activity process to obtain the temporal knowledge; Perform an inference calculation on the scene state sequence based on the temporal knowledge to obtain a discrimination result of the land mining activity; Generate a prompt message according to the discrimination result.

9. The method according to claim 8, wherein The discrimination result includes determination success, determination failure, and to be verified; Perform an inference calculation on the scene state sequence based on temporal knowledge to obtain a discrimination result of the land mining activity, including: Determine an expected sequence according to the temporal knowledge; Extract a behavior sequence from the scene state sequence; If the behavior sequence appears completely according to the expected sequence, generate the discrimination result of determination success; If the behavior sequence does not appear according to the expected sequence, generate the discrimination result of determination failure; If the behavior sequence appears partially according to the expected sequence, generate the discrimination result of to be verified.

10. A video recognition device for land mining activities, characterized in that, The device includes: An acquisition module, configured to acquire an image data set, where the image data set includes feature pictures extracted from video data, and the feature pictures include image targets corresponding to mining construction machinery; A labeling module, configured to perform behavior labeling on the image data set according to a behavior state set to obtain a labeled data set, where the behavior state set is a state set obtained by condensing the behavior process characteristics of the mining construction machinery according to temporal knowledge; A training module for training an object detection model using the labeled data set, where the object detection model is a deep learning model improved based on the behavior recognition characteristics of mining construction machinery; A detection module for inputting video data of a land scene into the trained object detection model to obtain an identification result output by the object detection model, where the identification result includes dynamic elements included in the current land scene, the element behaviors of the dynamic elements, and the pixel coordinates of the dynamic elements; A discrimination module for generating a scene state sequence according to the identification result and performing inference calculation on the scene state sequence based on temporal knowledge to obtain a discrimination result of the land mining activity.

Citation Information

Cited By

  • Land resource dynamic monitoring and evaluation system based on remote sensing image recognition

    CN121413784A