Multimodal large model training sample screening method and device

By combining small-model detection and multimodal large-model recognition, high-quality training samples of coal mine violations are selected, solving the problem of sample shortage, improving screening efficiency and accuracy, and meeting the needs of coal mine safety monitoring.

CN122391778APending Publication Date: 2026-07-14CHINA COAL RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA COAL RES INST
Filing Date
2026-04-01
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

The lack of relevant samples of violations by coal mine personnel makes it difficult to train multimodal large models. Existing screening methods have low coverage and low efficiency of manual screening, making it difficult to meet the needs of coal mine safety monitoring.

Method used

A small model is used for personnel detection to obtain candidate video clips. A multimodal large model is combined to identify violation information. Training samples are filtered through annotation information. The YOLOv11n model and the Prompt instruction are used to improve the efficiency and accuracy of video clip filtering.

Benefits of technology

It improves the coverage and accuracy of video screening, provides high-quality training samples, reduces the inefficiency and low accuracy of manual screening, and enhances the ability of multimodal large models to identify violations in coal mine scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391778A_ABST
    Figure CN122391778A_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computer, and particularly relates to a multi-modal large model training sample screening method. The method comprises the following steps: adopting a small model to perform personnel detection processing on collected video data, and obtaining a candidate personnel video segment set, wherein the video data is video collected for coal mines; adopting a multi-modal large model to identify the candidate personnel video segment set, and obtaining rule violation information of each candidate personnel video segment in the candidate personnel video segment set; according to first annotation information corresponding to the rule violation information of each candidate personnel video segment, performing screening processing on the candidate personnel video segment set, and obtaining training samples of the multi-modal large model. The present disclosure can improve the coverage of video screening, improve the screening accuracy, and provide high-quality training samples for multi-modal large model coal mine scene rule violation discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a method and apparatus for screening training samples for multimodal large models. Background Technology

[0002] With the development of science and technology, intelligent identification of personnel behavior violations in mine monitoring systems is crucial for maintaining mine safety. Current multimodal large-scale models will significantly improve the ability to identify such violations in coal mines. However, the lack of relevant samples regarding personnel behavior violations severely hinders the training of multimodal large-scale models. Summary of the Invention

[0003] This disclosure provides a method and apparatus for screening training samples for a multimodal large model, which can improve the coverage and accuracy of video screening, and provide high-quality training samples for identifying violations in coal mine scenes using a multimodal large model. The technical solution of this disclosure is as follows:

[0004] According to a first aspect of the present disclosure, a method for screening training samples for a multimodal large model is provided, comprising: A small model is used to perform personnel detection processing on the collected video data to obtain a set of video clips of candidate personnel, wherein the video data is video collected for coal mines; A multimodal large model is used to identify the set of candidate video segments and obtain violation information for each candidate video segment in the set of candidate video segments; Based on the first annotation information corresponding to the violation information of each candidate video segment, the candidate video segment set is filtered to obtain the training samples of the multimodal large model.

[0005] According to some embodiments, the step of using a multimodal large model to identify the candidate video segment set and obtaining violation information for each candidate video segment in the candidate video segment set includes: Obtain secondary annotation information for historically violating video clips; Based on the second annotation information, obtain the prompt word instruction; Based on the prompt word instruction, a multimodal large model is used to identify the set of candidate video clips and obtain the violation information of each candidate video clip in the set of candidate video clips.

[0006] According to some embodiments, the method further includes: Based on the requirements of the training samples, the segment information corresponding to the candidate video segment set is obtained, wherein the segment information includes the number of segments, the segment duration, and the segment interval duration, and the segment information is used to obtain the candidate video segment set.

[0007] According to some embodiments, the process of using a small model to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video clips includes: The small model is used to perform personnel detection processing on the collected video data. If it is determined that there is a target person in the first frame image, the candidate person video segment corresponding to the segment duration is obtained based on the first frame image. Iterate through the video data to obtain a set of video clips of the candidates.

[0008] According to some embodiments, the process of using a small model to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video clips includes: Based on the duration of the segment intervals, the YOLOv11n target detection model is used to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video segments.

[0009] According to some embodiments, the step of using a multimodal large model to identify the candidate video segment set and obtaining violation information for each candidate video segment in the candidate video segment set includes: Obtain scene information corresponding to the video data, and obtain the multimodal large model and the Prompt instruction set based on the scene information, wherein the Prompt instruction set includes preset violation behaviors corresponding to the scene information; Based on the Prompt instruction set, the multimodal large model is used to identify the candidate video segment set and obtain the violation information of each candidate video segment in the candidate video segment set.

[0010] According to some embodiments, the step of filtering the set of candidate video segments based on the first annotation information corresponding to the violation information of each candidate video segment to obtain training samples for the multimodal large model includes: Based on the first annotation information corresponding to the violation information of each candidate's video segment, the set of candidate video segments is filtered to obtain sample information corresponding to each candidate's video segment; If the sample information indicates that each candidate video segment is a violation video segment, then each candidate video segment is added to the training samples of the multimodal large model.

[0011] According to a second aspect of the present disclosure, a multimodal large model training sample screening device is provided, comprising: The set acquisition unit is used to perform personnel detection processing on the collected video data using a small model to obtain a set of candidate personnel video clips, wherein the video data is video collected for coal mines; The information acquisition unit is used to identify the candidate video segment set using a multimodal large model and acquire violation information of each candidate video segment in the candidate video segment set; The sample acquisition unit is used to filter the set of candidate video segments based on the first annotation information corresponding to the violation information of each candidate video segment, and obtain the training samples of the multimodal large model.

[0012] According to a third aspect of the present disclosure, an electronic device is provided, comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the multimodal large model training sample selection method described in any one of the preceding aspects.

[0013] According to a fourth aspect of the present disclosure, a storage medium is provided that, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the multimodal large model training sample screening method described in any of the preceding aspects.

[0014] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method described in any one of the preceding aspects.

[0015] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects: In some or related embodiments, a small model is used to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video clips, wherein the video data is video collected for coal mines; a multimodal large model is used to identify the set of candidate personnel video clips to obtain violation information of each candidate personnel video clip in the set; based on the first annotation information corresponding to the violation information of each candidate personnel video clip, the set of candidate personnel video clips is filtered to obtain training samples for the multimodal large model. Therefore, a coal mine violation behavior training sample filtering mechanism can be provided, which can use a small model to filter video data and a large model to filter samples, which can reduce the low efficiency and low accuracy of manual video filtering, improve the coverage and accuracy of video filtering, and provide high-quality training samples for the multimodal large model to identify violations in coal mine scenarios.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0018] Figure 1 This is a flowchart of the first method for selecting training samples for a multimodal large model provided in this disclosure embodiment; Figure 2 This is a flowchart of the second method for selecting training samples for multimodal large models provided in this disclosure embodiment; Figure 3 This is an example schematic diagram illustrating a method for selecting training samples for a multimodal large model provided in this disclosure. Figure 4 This is a block diagram illustrating a multimodal large model training sample selection device according to an exemplary embodiment; Figure 5 This is an example schematic diagram of an electronic device according to an exemplary embodiment. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0020] This disclosure provides a method, apparatus, electronic device, and storage medium for screening training samples of multimodal large models. In some embodiments, the terms "multimodal large model training sample screening method" and "information processing method" and "communication method" can be used interchangeably; the terms "multimodal large model training sample screening apparatus" and "information processing apparatus" and "communication apparatus" can be used interchangeably; and the terms "information processing system" and "communication system" can be used interchangeably.

[0021] This disclosure is not exhaustive, but merely illustrative of some embodiments, and is not intended to limit the scope of protection of this disclosure. Unless otherwise specified, each step in a particular embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a particular embodiment can also be implemented as an independent embodiment, and the order of the steps in a particular embodiment can be arbitrarily interchanged. Furthermore, the optional implementation methods in a particular embodiment can be arbitrarily combined; moreover, the embodiments can be arbitrarily combined, for example, some or all steps of different embodiments can be arbitrarily combined, and a particular embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0022] In each of the disclosed embodiments, unless otherwise specified or in case of logical conflict, the terminology and / or descriptions of the embodiments are consistent and can be referenced by each other. Technical features in different embodiments can be combined to form new embodiments based on their inherent logical relationships.

[0023] The terminology used in the embodiments of this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure.

[0024] In this disclosure, unless otherwise stated, elements expressed in the singular form, such as "a," "an," "the," "the," "the," "the," "the," "the," "this," etc., can mean "one and only one," or "one or more," "at least one," etc. For example, when using articles such as "a," "an," "the," etc. in translation, the noun following the article can be understood as either a singular or a plural expression.

[0025] In the embodiments disclosed herein, "multiple" refers to two or more.

[0026] In some embodiments, the terms “at least one of,” “one or more,” “a plurality of,” and “multiple” may be used interchangeably.

[0027] The prefixes "first," "second," etc., used in the embodiments of this disclosure are merely for distinguishing different descriptive objects and do not impose restrictions on the position, order, priority, quantity, or content of the descriptive objects. The description of the descriptive objects is found in the claims or the context of the embodiments, and the use of prefixes should not constitute unnecessary restrictions. For example, if the descriptive object is a "field," the ordinal numbers preceding "field" in "first field" and "second field" do not restrict the position or order of the "fields." "First" and "second" do not restrict whether the "fields" they modify are in the same message, nor do they restrict the order of "first field" and "second field." Similarly, if the descriptive object is a "level," the ordinal numbers preceding "level" in "first level" and "second level" do not restrict the priority between "levels." Furthermore, the number of descriptive objects is not limited by ordinal numbers and can be one or more. For example, in "first device," the number of "devices" can be one or more. Furthermore, the objects modified by different prefixes can be the same or different. For example, if the object being described is "device", then "first device" and "second device" can be the same device or different devices, and their types can be the same or different. Similarly, if the object being described is "information", then "first information" and "second information" can be the same information or different information, and their content can be the same or different.

[0028] In some embodiments, "terminal" or "terminal device" may be referred to as "user equipment (UE)," "user terminal," "mobile station (MS)," "mobile terminal (MT)," "subscriber station," "mobile unit," "subscriber unit," "wireless unit," "remote unit," "mobile device," "wireless device," "wireless communication device," "remote device," "mobile subscriber station," "access terminal," "mobile terminal," "wireless terminal," "remote terminal," "handset," "user agent," "mobile client," "client," etc.

[0029] In some embodiments, data, information, etc., may be obtained with the user's consent.

[0030] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0031] According to some embodiments, training samples for a multimodal large-scale model aimed at identifying violations by coal mine personnel can be obtained through methods such as manual screening or simple sampling. However, for the extremely sparse violation-related samples present in massive amounts of surveillance video, manual screening is too costly; simple sampling methods yield too few effective samples. This results in low efficiency in obtaining actual samples for automated coal mine operation scenarios.

[0032] In some embodiments, in coal mine fully mechanized mining and main haulage production scenarios, equipment typically needs to operate automatically for extended periods during production shifts. During automatic operation, personnel inspections are usually required, and miners may need to intervene to participate in production operations when necessary. In these production scenarios, miners may violate relevant regulations during production, potentially leading to safety accidents that threaten miners' lives and the mine's production order. Accurate identification and early warning of violations by the monitoring system, along with timely reminders to violators, is crucial for reducing mine violations and ensuring mine production safety. While some embodiments of the monitoring system can effectively warn of simple violations such as "not wearing a safety helmet," they struggle with complex violations, such as "personnel crossing the conveyor belt." The superior video understanding and analysis capabilities of multimodal large-scale artificial intelligence models hold promise for detecting relatively complex violations in mine videos. However, general multimodal large-scale models are typically trained on general scene image and video data, without targeted training for mine-specific scene images and videos, resulting in low accuracy in identifying violations in mine scenarios and failing to meet the actual needs of coal mine safety monitoring. Therefore, it is necessary to collect a certain amount of samples of violations by personnel in mine video images to train a multimodal large model.

[0033] In some implementations, monitoring videos from fully mechanized mining and main transportation scenarios in coal mines have the following characteristics: for most periods, the footage only includes equipment operating status and scene information, with no miners appearing in the footage. If only this type of miner-free video data is used to train a multimodal large-scale model, it cannot provide the model with effective features related to violations, and does not substantially help improve the model's ability to identify violations. Therefore, accurately selecting video clips from massive amounts of mine monitoring videos that "show people and may be engaging in violations" as training samples is a crucial prerequisite for improving the violation identification performance of multimodal large-scale models in coal mine scenarios.

[0034] In some embodiments, the video sample selection methods employ equal-interval sampling or random sampling strategies. These methods have extremely low coverage of video content, making it difficult to accurately capture valid segments where people are present and potentially engaging in violations. This results in the selection of a large number of unoccupied or obviously safe samples, leading to poor sample effectiveness. This hinders the efficient training of multimodal large-scale models and restricts the practical application of intelligent violation recognition technology in coal mine scenarios. Another method for selecting valid samples is manual review of videos. This method is time-consuming and labor-intensive, and for the small number of violation samples in a massive amount of video data, manual screening is difficult to achieve effectively.

[0035] Figure 1 This is a flowchart of the first method for selecting training samples for a multimodal large model provided in this disclosure, as shown in the embodiment. Figure 1As shown, this multimodal large model training sample selection method can be used in scenarios where training samples for a large model are obtained for detecting violations by coal mine personnel. The method includes the following steps: In step S11, a small model is used to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video clips, wherein the video data is video collected for coal mines; In some embodiments, the executing entity of this disclosure may be, for example, an electronic device. This electronic device does not specifically refer to a particular fixed electronic device. For example, when the device identifier changes, the electronic device may also change accordingly. For example, when the structure of the electronic device changes, the electronic device may also change accordingly. The electronic device of the executing entity in this disclosure may or may not be a target robot; this disclosure does not limit this.

[0036] According to some embodiments, small language models (SLMs) can be seen as an important direction of rapid development in the field of artificial intelligence in recent years. They refer to large language models with relatively few parameters and low computational resource requirements. The term "small model" does not specifically refer to a single, fixed model. For example, the small model can change when its model type changes. Similarly, the small model can change when its model parameters change.

[0037] In some embodiments, the video data may be data collected in a coal mine. The method of acquiring this video data is not limited. For example, the video data can be collected by cameras installed in the coal mine, or by cameras mounted on personnel. This video data does not refer to any specific, fixed data. For example, the video data may change if the time of acquisition changes. Similarly, the video data may change if the corresponding video content changes.

[0038] According to some embodiments, personnel detection processing can be, for example, a detection processing method for detecting the presence of personnel in video data. This personnel detection processing can be performed on all persons present in the video data, or it can be performed on a specific person; this disclosure does not limit the scope of the embodiments.

[0039] In some embodiments, the candidate video clip set may be, for example, a collection of video clips including individuals. This candidate video clip set may include, for example, a collection of video clips from which it is determined whether a violation has occurred. This candidate video clip set does not specifically refer to a fixed set. For example, the candidate video clip set may change when the video data changes. For example, the candidate video clip set may change when the method of obtaining the candidate video clip set changes.

[0040] According to some embodiments, the coal mine may be, for example, the coal mine to be inspected. This coal mine may also be referred to as the target coal mine, the coal mine to be inspected, etc. The coal mine does not specifically refer to a particular fixed coal mine. For example, when the identification of the coal mine changes, the coal mine may also change accordingly. For example, when the structure of the coal mine also changes, the coal mine may also change accordingly.

[0041] In some embodiments, a small model is used to perform personnel detection processing on the collected video data to obtain a set of video clips of candidate personnel, wherein the video data is video collected for coal mines.

[0042] In step S12, a multimodal large model is used to identify the candidate video clip set and obtain the violation information of each candidate video clip in the candidate video clip set; According to some embodiments, "large language models" can refer, for example, to pre-trained language models based on deep learning (primarily the Transformer architecture) with a huge number of parameters (typically billions to hundreds of billions). Such large models can also be called large language models. The term "large model" does not specifically refer to a particular fixed model. For example, when the parameters of a large model change, the large model can also change accordingly. In this context, a large model could be, for example, a model that, once trained, can be used to identify coal mine-related data and obtain information on the risks and hazards of the coal mine.

[0043] In some embodiments, the multimodal large model may be, for example, a large model capable of recognizing multimodal data.

[0044] According to some embodiments, the violation information in each candidate's video clip may be, for example, the violation information present in each candidate's video clip. This violation information does not specifically refer to any fixed information. For example, when the video content in each candidate's video clip changes, the violation information in that candidate's video clip may also change accordingly. For example, when a person's violation changes, the violation information in that candidate's video clip may also change accordingly.

[0045] In some embodiments, violation information may be, for example, behavioral information in video clips of candidates that does not meet coal mine operation requirements. This violation information is not specifically defined by any single fixed detail. For example, different types of coal mines may correspond to different preset violation information. For instance, when a coal mine changes or its type changes, the preset violation information may also change accordingly. The violation information in each candidate's video clip may also change accordingly.

[0046] In some embodiments, a multimodal large model can be used to identify the set of candidate video clips and obtain violation information for each candidate video clip in the set of candidate video clips.

[0047] In step S13, the candidate video segment set is filtered according to the first annotation information corresponding to the violation information of each candidate video segment to obtain the training samples of the multimodal large model.

[0048] According to some embodiments, the annotation information may be information used to annotate violation information. This annotation information includes, but is not limited to, click annotation information, voice annotation information, and text annotation information. The first annotation information may, for example, be annotation information corresponding to the violation information of each candidate's video clip. The "first" in the first annotation information is used to distinguish it from other annotation information and does not specifically refer to any fixed information. For example, when the violation information of each candidate's video clip changes, the first annotation information may also change accordingly.

[0049] In some embodiments, the training samples may be samples used to train a large multimodal model. These training samples do not specifically refer to a single, fixed sample. For example, the training samples may change accordingly when the first annotation information changes. Similarly, the training samples may change accordingly when the selection process changes. These training samples may also be referred to as a training sample set, etc.

[0050] According to some embodiments, the set of candidate video segments can be filtered based on the first annotation information corresponding to the violation information of each candidate video segment to obtain training samples for the multimodal large model.

[0051] In some or related embodiments, a small model is used to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video clips, wherein the video data is video collected for coal mines; a multimodal large model is used to identify the set of candidate personnel video clips to obtain violation information of each candidate personnel video clip in the set; based on the first annotation information corresponding to the violation information of each candidate personnel video clip, the set of candidate personnel video clips is filtered to obtain training samples for the multimodal large model. Therefore, a coal mine violation behavior training sample filtering mechanism can be provided, which can use a small model to filter video data and a large model to filter samples, which can reduce the low efficiency and low accuracy of manual video filtering, improve the coverage and accuracy of video filtering, and provide high-quality training samples for the multimodal large model to identify violations in coal mine scenarios.

[0052] Figure 2 This is a flowchart of the second method for selecting training samples for multimodal large models provided in this disclosure, as shown in the embodiments. Figure 2 As shown, it includes the following steps: In step S21, the small model is used to perform personnel detection processing on the collected video data. If it is determined that there is a target person in the first frame image, the first frame image is used as a reference to obtain a candidate person video segment corresponding to the segment duration. The video data is video collected for coal mines. In some embodiments, the executing entity of this disclosure may be, for example, an electronic device. This electronic device does not specifically refer to a particular fixed electronic device. For example, when the device identifier changes, the electronic device may also change accordingly. For example, when the structure of the electronic device changes, the electronic device may also change accordingly. The electronic device of the executing entity in this disclosure may or may not be a target robot; this disclosure does not limit this.

[0053] The relevant processes can be as described above, and will not be repeated here.

[0054] According to some embodiments, Figure 3 This is an example schematic diagram illustrating a method for selecting training samples for a multimodal large model provided in this disclosure, such as... Figure 3 As shown, it can include input video stream, small-interval frame extraction, YOLO detection, discarding when no one is present, and cutting segments when someone is present. It can input up to a multimodal large model for filtering, and can also input a Prompt. It can discard when the large model reports safety, and manually review when the large model reports unsafe. It can label illegal samples and highly suspicious samples, and discard worthless samples.

[0055] In some embodiments, high-value samples specific to an industry are defined as samples where the large model misjudges or where there are genuine violations within that specific industry scenario. These samples are crucial for improving the behavior identification capabilities of the large model in that specific scenario. In coal mine fully mechanized mining and main transportation monitoring videos, these high-value samples are very sparsely distributed, making manual screening difficult. This technical solution addresses the characteristics of long-term automatic operation and low frequency of personnel appearance in fully mechanized mining and main transportation scenarios by employing a highly efficient small model for pedestrian detection. Experiments have shown that this model can filter out over 95% of videos without personnel, significantly improving sample screening efficiency. For video clips with personnel, this technical solution further filters high-value samples using the following methods: First, a small amount of manually selected industry-labeled data is used to fine-tune the multimodal large model, giving it a certain understanding of the industry-specific scenario. Then, violations are defined according to coal mine safety regulations and incorporated into the prompt, providing a basis for the large model's sample screening. Finally, for the aforementioned video clips with personnel, the fine-tuned model and prompt instructions are combined to effectively detect video clips with violations or suspected violations, i.e., high-value samples.

[0056] In some embodiments, the first frame may be, for example, the first image frame in which a target person is present when performing person detection on video data. This first frame does not specifically refer to a fixed frame. For example, the first frame may change accordingly when the video data changes. Similarly, the first frame may change accordingly when the target person changes.

[0057] In some embodiments, the target personnel may refer to a fixed person or any existing person; this disclosure does not limit this.

[0058] According to some embodiments, the segment duration may be, for example, used to indicate the length of a video segment featuring a candidate. This segment duration is not specifically a fixed length. For example, the segment duration may change when the method of determining the segment duration changes. For example, the segment duration may also change when a modification instruction for the segment duration is received.

[0059] According to some embodiments, the YOLO object detection model can be used to perform personnel detection processing on the collected video data to obtain a set of candidate person video clips.

[0060] According to some embodiments, the process of using a small model to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video clips includes: Based on the segment interval duration, the YOLOv11n object detection model is used to process the collected video data for personnel detection, obtaining a set of candidate person video segments. Obtaining the candidate person video segment set based on the segment interval duration can improve the probability of missed detections, enhance the completeness of video data detection, and increase the accuracy of obtaining the candidate person video segment set.

[0061] In some embodiments, the segment duration can be, for example, 40 seconds. The video data can be, for example, surveillance video, specifically, surveillance video from a coal mine's fully mechanized mining or main transportation scenario. The YOLO target detection model is used to perform frame-level detection on the video, with a detection interval set to 3 seconds per frame (i.e., the segment interval duration can be 3 seconds per frame). This means that a video frame is extracted every 3 seconds for personnel target detection. When a personnel target is detected in a frame, a 40-second video segment is designated as a candidate personnel segment, using that frame as the time reference. The segment time range is set to [3 seconds before the detection frame's corresponding time, and 37 seconds after the detection frame's corresponding time]. After the candidate segment is designated, personnel detection continues at 3-second / frame intervals, starting from the end time of that segment (i.e., 37 seconds after the detection frame's corresponding time). Time overlap between adjacent candidate segments is allowed to ensure no personnel appearance is missed, thus improving the coverage of personnel segments. Therefore, by adopting the YOLO model to detect video frames every 3 seconds and combining it with a 40-second segment delineation strategy, and allowing adjacent segments to overlap, this method improves the capture efficiency of video segments in which people appear, reduces the probability of missed captures, avoids missing effective scenes, and significantly improves time coverage compared to traditional equal-interval sampling or random sampling methods. This solves the problem of low time coverage and difficulty in obtaining effective samples caused by the scattered time periods of people appearing in coal mine scenes and traditional simple sampling methods.

[0062] According to some implementations, for monitoring videos of coal mine fully mechanized mining and main transportation scenarios, such as one hour of historical monitoring video, the YOLO object detection model is used to perform frame-level detection. The object detection time interval can be set to 3 seconds per frame, meaning that every 3 seconds, one frame is extracted from the video stream and input into the YOLO model for personnel detection. If no pedestrian is detected in a given frame, that frame is discarded, and an image from 3 seconds later is selected and input into the YOLO model to continue personnel detection.

[0063] The YOLO model used could be, for example, the YOLOv11n model. This model is highly adaptable and meets the actual needs of coal mine scenarios. It boasts advantages such as high detection accuracy and speed, making it suitable for the real-time processing requirements of underground coal mine monitoring videos. Before detection, the YOLO model needs to be pre-trained using a small number of underground coal mine personnel image samples to improve its adaptability to the dimly lit environment of the mine and personnel's clothing. A complete solution can be designed specifically for the video characteristics and specifications of fully mechanized mining and main transportation scenarios in coal mines. The real-time detection capability of the YOLO model is adapted to the massive data processing needs of underground monitoring videos, and the pre-tuned multimodal large model is adapted to the understanding of coal mine scenarios. The overall solution can be directly integrated into the coal mine safety monitoring system, making it highly practical.

[0064] In step S22, the video data is traversed to obtain a set of candidate video clips; The relevant processes can be as described above, and will not be repeated here.

[0065] According to some embodiments, the method further includes: Based on the requirements of the training samples, segment information corresponding to the candidate video segment set is obtained. This segment information includes the number of segments, segment duration, and segment interval duration. This segment information is used to obtain the candidate video segment set. Therefore, obtaining the corresponding segment information based on the requirements improves the matching between segment information and training samples, thereby increasing the accuracy of multimodal large model training.

[0066] This fragment of information does not refer to any specific, fixed piece of information. For example, when the requirements change, this fragment of information may also change accordingly. For example, when the type and quantity of information included in the fragment of information change, the fragment of information may also change accordingly.

[0067] According to some embodiments, the number of segments may be used to indicate, for example, the number of segments included in the candidate video segment set. This number of segments does not specifically refer to a fixed number. The segment duration may, for example, be the duration corresponding to each candidate video segment.

[0068] In some embodiments, segment information can be determined based on device information of the electronic device and the context length of the multimodal large model analysis video. For example, 32 frames of the video segment can be selected for processing. The segment length can be set to 40 seconds, approximately one frame per second, which allows for better understanding of the video content. The relatively long video duration also facilitates faster filtering of long videos, balancing the accuracy and efficiency of video recognition.

[0069] In step S23, a multimodal large model is used to identify the candidate video clip set and obtain the violation information of each candidate video clip in the candidate video clip set; The relevant processes can be as described above, and will not be repeated here.

[0070] According to some embodiments, "large language models" can refer, for example, to pre-trained language models based on deep learning (primarily the Transformer architecture) with a huge number of parameters (typically billions to hundreds of billions). This large model can also be called a large language model. The term "large model" does not specifically refer to a particular fixed model. For example, when the model parameters corresponding to a large model change, the large model can also change accordingly. In this context, a large model could be, for example, a trained model that can be used to identify coal mine-related data and obtain information on the risks and hazards of the coal mine. The large model could be the Qwen3.5-27B model.

[0071] According to some implementation methods, a large model can be fine-tuned to obtain a multimodal large model. Fine-tuning may involve manually selecting a small amount of high-quality, industry-specific data, including data on personnel violations, normal personnel operations, and empty scenarios, with basic attributes annotated by coal mine experts. During annotation, equipment type, equipment attributes, personnel behavior, and human-machine positional relationships are defined, and question-answer pairs are generated from the large model based on the annotated basic attributes. A general multimodal large model, such as the Qwen3.5-27B model, is manually selected, and the aforementioned small amount of annotated data from the coal mine scenario and question-answer pairs are used for initial fine-tuning of the multimodal large model.

[0072] According to some embodiments, the step of using a multimodal large model to identify the candidate video segment set and obtaining violation information for each candidate video segment in the candidate video segment set includes: Obtain secondary annotation information for historically violating video clips; Based on the second annotation information, obtain the prompt word instruction; Based on the prompt words, a multimodal large model is used to identify the set of candidate video clips, obtaining violation information for each candidate video clip in the set. Therefore, identifying the set of candidate video clips based on prompt words can improve the accuracy of candidate video clip identification and the accuracy of violation information determination.

[0073] According to some embodiments, the second annotation information may be, for example, information used to annotate historical violation video clips. This historical violation video clip may be, for example, a collection of historical violation video clips. The second annotation information may be used to annotate violation information within these historical violation video clips.

[0074] According to some implementation methods, a small amount of high-quality, industry-specific data, including data on personnel violations, normal personnel operations, and empty scenarios, can be manually selected and annotated by coal mine experts. During annotation, equipment types, equipment attributes, personnel behaviors, and human-machine positional relationships can be defined, and question-and-answer pairs can be generated from a large model based on the annotations. A general multimodal large model is then manually selected, and preliminary fine-tuning of the multimodal large model is performed using the aforementioned small amount of annotated data from coal mine scenarios and question-and-answer pairs.

[0075] According to some embodiments, combined with the violation definition methods in coal mine safety regulations, targeted Prompt instructions can be obtained. These Prompt instructions can clearly define common violation types (such as unauthorized entry into hazardous work areas) and the judgment criteria in fully mechanized mining and main transportation scenarios. The obtained video clips of potential personnel can be input into a multimodal large-scale model. The Prompt instruction guides the model to analyze the content of the clips and determine whether they are suspicious clips containing potentially unsafe behaviors, thus obtaining violation information for each candidate personnel's video clip. Since this multimodal large-scale model has only undergone minor adjustments with a small number of samples, its identification results may not be accurate. However, clips identified as containing violations are essentially boundary samples for model inference. These samples carry key features of violations in coal mine scenarios and can provide effective knowledge increments for subsequent model training, making them high-value samples.

[0076] The design of Prompt commands should adhere to the principles of accuracy and scenario-based application. The following example illustrates this using a fully mechanized mining scenario.

[0077] According to some embodiments, the Prompt instruction set for integrated mining scenarios is as follows: (1) Setting up the main model roles and tasks: Assume you are a coal mining expert. Please understand the entire video as a whole, and in conjunction with relevant regulations, determine whether the people in the video have violated any rules, and output structured information (JSON format). For example: JSON_TEMPLATE= {"has_violation": False, "description": "The miner is walking next to the stationary coal mining machine, and there is no violation"}.

[0078] (2) Definitions of various violations: The following are the issues to check: (a) Hydraulic supports: ① Workers use hydraulic supports to lift heavy objects or clear coal flows. Specifically, this is shown in the image: a person appears, the hydraulic support is moving and lifting heavy objects, or the hydraulic support is shoveling and moving coal flows.

[0079] (b) Coal mining machines: ① During coal mining operations, workers enter the area within 5 meters of the coal face or the front and rear of the coal mining machine drum. Specifically, this means that while the coal mining machine is running, there are people within 5 meters of the front and rear of the coal mining machine drum.

[0080] ② Repairing or cleaning operations on an unlocked coal mining machine, specifically manifested as: miners repairing or cleaning coal on or near the coal mining machine while it is running or when there is a warning that it is not locked.

[0081] ③ Walking, shoveling, washing, or touching the operating coal mining machine. Specifically, this means that when the coal mining machine is running, miners walk, shovel, wash, or touch it.

[0082] ④ Workers climb on or jump off the coal mining machine.

[0083] ⑤ Workers sit or lie on the coal mining machine.

[0084] (c) Gas detection category: ① When taking samples, the gas inspector did not comply with the following requirements: extend the gas sampling tube to the designated position, slowly draw out the gas, and avoid disturbing the airflow.

[0085] (d) Scraper conveyors: ① Use scraper conveyors to transport people.

[0086] If no issues from the checklist appear in the video, return False; otherwise, return True and provide a detailed description of the behavior and scene in the video.

[0087] According to some embodiments, the step of using a multimodal large model to identify the candidate video segment set and obtaining violation information for each candidate video segment in the candidate video segment set includes: Obtain scene information corresponding to the video data, and obtain the multimodal large model and the Prompt instruction set based on the scene information, wherein the Prompt instruction set includes preset violation behaviors corresponding to the scene information; Based on the Prompt instruction set, the multimodal large model is used to identify the candidate video segment set and obtain the violation information of each candidate video segment in the candidate video segment set.

[0088] In step S24, the candidate video segment set is filtered according to the first annotation information corresponding to the violation information of each candidate video segment to obtain the training samples of the multimodal large model.

[0089] The relevant processes can be as described above, and will not be repeated here.

[0090] According to some embodiments, the step of filtering the set of candidate video segments based on the first annotation information corresponding to the violation information of each candidate video segment to obtain training samples for the multimodal large model includes: Based on the first annotation information corresponding to the violation information of each candidate's video segment, the set of candidate video segments is filtered to obtain sample information corresponding to each candidate's video segment; If the sample information indicates that each candidate video segment is a violation video segment, then each candidate video segment is added to the training samples of the multimodal large model.

[0091] According to some implementation methods, the suspicious segments selected can be manually reviewed and accurately labeled with the violation information of each candidate's video segments. Personnel with coal mine safety operation qualifications confirm whether there are real violations in the segments, the specific type of violations and the time of occurrence, eliminate misjudged segments, and retain real and valid violation segments and suspected violation segments. The manually labeled segments are used as training samples for subsequent iterative training and parameter optimization of the multimodal large model.

[0092] According to some embodiments, the training samples can be used to train a multimodal large model, thereby improving the accuracy of multimodal large model acquisition and significantly enhancing the efficiency and precision of monitoring and early warning.

[0093] In one or related embodiments, the small model is used to perform personnel detection processing on the collected video data. If a target person is determined to exist in the first frame image, a candidate person video segment corresponding to the segment duration is obtained based on the first frame image. Traversing the video data to obtain a set of candidate person video segments can improve sample coverage and sample acquisition rate. Furthermore, by combining a multimodal large model with a Prompt based on coal mine safety regulations, the selected boundary samples carry core features of violations. After manual annotation, these samples can provide accurate training basis for the multimodal large model, avoiding interference from segments without personnel or too many normal operation segments on model fine-tuning, significantly improving the training efficiency and accuracy of the model in the task of identifying violations in coal mine scenarios.

[0094] A block diagram illustrating a multimodal large model training sample selection device according to an exemplary embodiment. (Refer to...) Figure 4 The device 400 includes: The set acquisition unit 401 is used to perform personnel detection processing on the collected video data using a small model to acquire a set of candidate personnel video segments, wherein the video data is video collected for coal mines; The information acquisition unit 402 is used to identify the candidate video segment set using a multimodal large model and acquire violation information of each candidate video segment in the candidate video segment set; The sample acquisition unit 403 is used to filter the set of candidate video segments according to the first annotation information corresponding to the violation information of each candidate video segment, and obtain the training samples of the multimodal large model.

[0095] According to some embodiments, the information acquisition unit 402, when using a multimodal large model to identify the candidate video segment set and acquiring violation information of each candidate video segment in the candidate video segment set, is specifically used for: Obtain secondary annotation information for historically violating video clips; Based on the second annotation information, obtain the prompt word instruction; Based on the prompt word instruction, a multimodal large model is used to identify the set of candidate video clips and obtain the violation information of each candidate video clip in the set of candidate video clips.

[0096] According to some embodiments, the collection acquisition unit 401 is further specifically used for: Based on the requirements of the training samples, the segment information corresponding to the candidate video segment set is obtained, wherein the segment information includes the number of segments, the segment duration, and the segment interval duration, and the segment information is used to obtain the candidate video segment set.

[0097] According to some embodiments, the set acquisition unit 401, when performing personnel detection processing on the collected video data using a small model to acquire a set of candidate video clips, specifically performs the following: The small model is used to perform personnel detection processing on the collected video data. If it is determined that there is a target person in the first frame image, the candidate person video segment corresponding to the segment duration is obtained based on the first frame image. Iterate through the video data to obtain a set of video clips of the candidates.

[0098] According to some embodiments, the set acquisition unit 401, when performing personnel detection processing on the collected video data using a small model to acquire a set of candidate video clips, specifically performs the following: Based on the duration of the segment intervals, the YOLOv11n target detection model is used to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video segments.

[0099] According to some embodiments, the information acquisition unit 402, when using a multimodal large model to identify the candidate video segment set and acquiring violation information of each candidate video segment in the candidate video segment set, is specifically used for: Obtain scene information corresponding to the video data, and obtain the multimodal large model and the Prompt instruction set based on the scene information, wherein the Prompt instruction set includes preset violation behaviors corresponding to the scene information; Based on the Prompt instruction set, the multimodal large model is used to identify the candidate video segment set and obtain the violation information of each candidate video segment in the candidate video segment set.

[0100] According to some embodiments, the sample acquisition unit 403 is used to filter the set of candidate video segments based on the first annotation information corresponding to the violation information of each candidate video segment, and to obtain the training samples of the multimodal large model, specifically for: Based on the first annotation information corresponding to the violation information of each candidate's video segment, the set of candidate video segments is filtered to obtain sample information corresponding to each candidate's video segment; If the sample information indicates that each candidate video segment is a violation video segment, then each candidate video segment is added to the training samples of the multimodal large model.

[0101] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0102] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0103] like Figure 5As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0104] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0105] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above. For example, in some embodiments, the above methods can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the above methods by any other suitable means (e.g., by means of firmware).

[0106] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0107] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0108] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0110] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0111] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0112] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0113] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for selecting training samples for a multimodal large model, characterized in that, include: A small model is used to perform personnel detection processing on the collected video data to obtain a set of video clips of candidate personnel, wherein the video data is video collected for coal mines; A multimodal large model is used to identify the set of candidate video segments and obtain violation information for each candidate video segment in the set of candidate video segments; Based on the first annotation information corresponding to the violation information of each candidate video segment, the candidate video segment set is filtered to obtain the training samples of the multimodal large model.

2. The method according to claim 1, characterized in that, The process of using a multimodal large model to identify the candidate video clip set and obtaining violation information for each candidate video clip in the candidate video clip set includes: Obtain secondary annotation information for historically violating video clips; Based on the second annotation information, obtain the prompt word instruction; Based on the prompt word instruction, a multimodal large model is used to identify the set of candidate video clips and obtain the violation information of each candidate video clip in the set of candidate video clips.

3. The method according to claim 1, characterized in that, The method further includes: Based on the requirements of the training samples, the segment information corresponding to the candidate video segment set is obtained, wherein the segment information includes the number of segments, the segment duration, and the segment interval duration, and the segment information is used to obtain the candidate video segment set.

4. The method according to claim 1, characterized in that, The process of using a small model to perform personnel detection processing on the collected video data to obtain a set of candidate person video clips includes: The small model is used to perform personnel detection processing on the collected video data. If it is determined that there is a target person in the first frame image, the candidate person video segment corresponding to the segment duration is obtained based on the first frame image. Iterate through the video data to obtain a set of video clips of the candidates.

5. The method according to claim 1, characterized in that, The process of using a small model to perform personnel detection processing on the collected video data to obtain a set of candidate person video clips includes: Based on the duration of the segment intervals, the YOLOv11n target detection model is used to perform personnel detection processing on the collected video data to obtain a set of candidate personnel video segments.

6. The method according to claim 1, characterized in that, The process of using a multimodal large model to identify the candidate video clip set and obtaining violation information for each candidate video clip in the candidate video clip set includes: Obtain scene information corresponding to the video data, and obtain the multimodal large model and the Prompt instruction set based on the scene information, wherein the Prompt instruction set includes preset violation behaviors corresponding to the scene information; Based on the Prompt instruction set, the multimodal large model is used to identify the candidate video segment set and obtain the violation information of each candidate video segment in the candidate video segment set.

7. The method according to claim 1, characterized in that, The step of filtering the set of candidate video clips based on the first annotation information corresponding to the violation information of each candidate video clip to obtain training samples for the multimodal large model includes: Based on the first annotation information corresponding to the violation information of each candidate's video segment, the set of candidate video segments is filtered to obtain sample information corresponding to each candidate's video segment; If the sample information indicates that each candidate video segment is a violation video segment, then each candidate video segment is added to the training samples of the multimodal large model.

8. A multimodal large model training sample selection device, characterized in that, include: The set acquisition unit is used to perform personnel detection processing on the collected video data using a small model to obtain a set of candidate personnel video clips, wherein the video data is video collected for coal mines; The information acquisition unit is used to identify the candidate video segment set using a multimodal large model and acquire violation information of each candidate video segment in the candidate video segment set; The sample acquisition unit is used to filter the set of candidate video segments based on the first annotation information corresponding to the violation information of each candidate video segment, and obtain the training samples of the multimodal large model.

9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the multimodal large model training sample selection method as described in any one of claims 1 to 7.

10. A storage medium storing instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the multimodal large model training sample selection method as described in any one of claims 1 to 7.