Action detection method and device based on multi-modal large model

By applying a multimodal large model action detection method in railway engineering equipment inspection videos, combined with pre-checking of timing action detection models, the problem of low efficiency and accuracy of inspection video inspection in the existing technology is solved, and more efficient and accurate action detection is achieved.

CN119964230AInactive Publication Date: 2025-05-09QINGDAO HISENSE ELECTRONICS TECH CONSULTANCY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411845034.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The inspection efficiency and accuracy of railway engineering equipment inspection videos in the prior art are low, making it difficult to supervise the inspection quality of on-duty personnel, and easily causing train safety hazards.

Method used

The action detection method based on the multimodal large model is adopted. By obtaining the inspection video, pre-checking is first based on the timing action detection model to determine the probability value of the target action in the video segment. If the threshold is not met, the multimodal large model will be used for secondary detection to obtain more accurate detection results.

Benefits of technology

It effectively improves the efficiency and accuracy of inspection video inspection, solves the problems of low accuracy and low efficiency of traditional artificial quality inspection, and improves the accuracy of action detection through secondary judgment of multimodal large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964230A_ABST
    Figure CN119964230A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to an action detection method and equipment based on a multi-modal large model. Acquiring an inspection video recording the inspection process of the on-duty personnel, pre-inspecting the inspection video based on a time sequence action detection model, determining a first probability value of a target action existing in a video segment corresponding to each prediction time range, and determining a second probability value of the target action existing in a video segment corresponding to each prediction time range for each prediction time range; and if the first probability value corresponding to the prediction time range does not meet the preset threshold requirement, detecting a video segment corresponding to the prediction time range based on a multi-modal large model to obtain a detection result used for describing a target action existing in the corresponding video segment, the problems of low accuracy and low efficiency of traditional manual quality inspection are effectively solved, and the accuracy of action detection is improved by utilizing a multi-modal large model secondary judgment mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a motion detection method and device based on a multimodal large model. Background Art

[0002] In the railway engineering scene, it is necessary to inspect the rails regularly to check whether the railway engineering equipment meets the safety conditions, such as whether the screws are tightened and whether the limiters are tight. At present, the inspection of railway engineering equipment is mainly completed manually by railway inspectors, who use knocking tools to complete the scheduled inspection procedures one by one. However, an important problem with the current railway engineering is that it is difficult to control the inspection quality of railway on-duty personnel, that is, it is impossible to supervise the railway safety inspection. The problem of inadequate security inspection caused by personal carelessness and other factors is easy to cause hidden dangers to train safety, which in turn will cause hidden dangers to the safety of people's lives and property.

[0003] In order to solve this problem, one of the current solutions is to equip railway workers with a first-person monitoring device with storage and voice reception, and record the work process in real time when the railway inspection workers are working. The video is then transmitted to the background for verification by the supervisory staff. However, there are the following problems with manual review. First, manual efficiency is low and prone to errors. Relying on manual methods will inevitably lead to low efficiency and errors caused by subjective opinions of staff. Secondly, the amount of video is large, and manual video verification is costly. A single maintenance section has to process more than 400 videos on average every day, and if it is fully checked, it will require 5-10 people.

[0004] Therefore, how to improve the efficiency and accuracy of inspection videos has become an urgent problem to be solved. Summary of the invention

[0005] The embodiments of the present application provide a motion detection method and device based on a multimodal large model to solve the problems of low efficiency and accuracy in inspecting patrol videos in the prior art.

[0006] The present application provides an action detection method based on a multimodal large model, the method comprising:

[0007] Obtain inspection videos recording the inspection process of on-duty personnel;

[0008] Determine, based on the temporal action detection model, a first probability value of a target action existing in a video segment corresponding to each predicted time range;

[0009] For each predicted time range, if the first probability value corresponding to the predicted time range does not meet the preset threshold requirement, the video segment corresponding to the predicted time range is detected based on the multimodal large model to obtain a detection result, which is used to describe the target action in the corresponding video segment.

[0010] The present application provides an action detection device based on a multimodal large model, the device comprising:

[0011] The acquisition module is used to acquire the inspection video recording the inspection process of the on-duty personnel;

[0012] The detection module is used to determine the first probability value of the target action existing in the video segment corresponding to each predicted time range based on the temporal action detection model; for each predicted time range, if the first probability value corresponding to the predicted time range does not meet the preset threshold requirement, the video segment corresponding to the predicted time range is detected based on the multimodal large model to obtain a detection result, and the detection result is used to describe the target action existing in the corresponding video segment.

[0013] The present application also provides an electronic device, which includes a processor, and the processor is used to implement the steps of any of the above-mentioned action detection methods based on a multimodal large model when executing a computer program stored in a memory.

[0014] The present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned multimodal large model-based action detection methods.

[0015] In the embodiment of the present application, an inspection video recording the inspection process of the on-duty personnel is obtained, and the inspection video is first pre-inspected based on the temporal action detection model to determine the first probability value of the target action existing in the video segment corresponding to each predicted time range. Then, for each predicted time range, if the first probability value corresponding to the predicted time range does not meet the preset threshold requirement, the video segment corresponding to the predicted time range is detected based on the multimodal large model to obtain a detection result for describing the target action existing in the corresponding video segment, which effectively solves the problems of low accuracy and low efficiency of traditional manual quality inspection, and at the same time improves the accuracy of action detection by using the multimodal large model for secondary judgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solution of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 A schematic diagram of an action detection process based on a multimodal large model provided in an embodiment of the present application;

[0018] Figure 2aA schematic diagram of a turnout inspection provided in an embodiment of the present application;

[0019] Figure 2b A schematic diagram of a turnout inspection provided in an embodiment of the present application;

[0020] Figure 2c A schematic diagram of a turnout inspection provided in an embodiment of the present application;

[0021] Figure 3 A schematic diagram of a temporal action detection model architecture provided in an embodiment of the present application;

[0022] Figure 4 A schematic diagram of constructing a candidate time range provided in an embodiment of the present application;

[0023] Figure 5 A schematic diagram of an action detection process based on a multimodal large model provided in an embodiment of the present application;

[0024] Figure 6 A schematic diagram of the structure of a motion detection device based on a multimodal large model provided in an embodiment of the present application;

[0025] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the embodiment of the present application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiment is only a part of the embodiment of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of this application.

[0027] An embodiment of the present application provides a motion detection method and device based on a multimodal large model. The method obtains an inspection video that records the inspection process of on-duty personnel; based on a temporal motion detection model, a first probability value of a target action existing in a video segment corresponding to each predicted time range is determined; for each predicted time range, if the first probability value corresponding to the predicted time range does not meet a preset threshold requirement, the video segment corresponding to the predicted time range is detected based on the multimodal large model to obtain a detection result, which is used to describe the target action existing in the corresponding video segment.

[0028] Figure 1 A schematic diagram of an action detection process based on a multimodal large model provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the process includes the following steps:

[0029] S101: Obtaining an inspection video recording the inspection process of the on-duty personnel.

[0030] An action detection method based on a multimodal large model provided in an embodiment of the present application is applied to an electronic device, which may be a server, a PC, a smart terminal, etc.

[0031] In order to check the on-duty process of the on-duty personnel, in an embodiment of the present application, an inspection video recording the on-duty personnel's inspection process can be obtained. The inspection video can be shot by a terminal carried by the on-duty personnel. The inspection video can be a real-time file or an offline file. The inspection video can be checked later to determine whether the on-duty personnel have performed the corresponding actions in accordance with the specifications.

[0032] Taking the railway turnout inspection scenario as an example, typical verification actions include: shooting the turnout number, the height direction of the turnout, insulating joints, close contact between the point rail and the basic rail, close contact between the point rail and the top iron, gap between the point rail and the slide bed, top iron bolts, limiters (or spacer irons), turnout sleeper bolts, roller anti-jump limit device (optional), slide bed gap, guard rail bolts, filling in work records, etc. Figure 2a A schematic diagram of a turnout inspection provided in an embodiment of the present application is shown in FIG. Figure 2a As shown, railway staff are checking the spacing of the slider bed plates. Figure 2b A schematic diagram of a turnout inspection provided in an embodiment of the present application is shown in FIG. Figure 2b As shown, railway staff are checking the tightness of the point rail and the base rail. Figure 2c A schematic diagram of a turnout inspection provided in an embodiment of the present application is shown in FIG. Figure 2c As shown, railway staff are taking photos of the completed work records.

[0033] S102: Determine a first probability value of a target action existing in a video segment corresponding to each predicted time range based on a temporal action detection model.

[0034] After obtaining the temporal action detection model, the first probability value of the target action in the video segment corresponding to each predicted time range can be determined based on the temporal action detection model. That is to say, the inspection video is analyzed using the temporal action detection model to determine multiple predicted time ranges in the inspection video, and the first probability value of the target action in the video segment corresponding to each predicted time range is predicted. The temporal action detection model can be an existing model, such as SCNN (Spatial Convolutional Neural Network), TURN, CDC, etc.

[0035] S103: For each predicted time range, if the first probability value corresponding to the predicted time range does not meet the preset threshold requirement, the video segment corresponding to the predicted time range is detected based on the multimodal large model to obtain a detection result, and the detection result is used to describe the target action existing in the corresponding video segment.

[0036] In order to accurately perform action detection, in an embodiment of the present application, a preset threshold requirement may be pre-configured. After obtaining the first probability value corresponding to each predicted time range, it is possible to determine whether the first probability value corresponding to the predicted time range meets the preset threshold requirement for each predicted time range. If not, it means that the possibility of the target action existing in the video segment corresponding to the predicted time range is relatively low. The video segment corresponding to the predicted time range can be detected using a multimodal large model to determine whether the target action exists in the video segment. In an embodiment of the present application, the detection result output by the multimodal large model can be used as the target action existing in the video segment corresponding to the preset time range. In other words, if the first probability value of the target action existing in the video segment corresponding to any predicted time range determined by the temporal action detection model meets the preset threshold requirement, the target action is determined as the action existing in the corresponding video segment. It also means that the on-duty personnel performed the action during the inspection process. If the first probability value of the target action existing in the video segment corresponding to any predicted time range determined by the temporal action detection model does not meet the preset threshold requirement, it means that the credibility is not high. The multimodal large model can be used to re-check the video segment corresponding to the predicted time range to determine the target action existing in the video segment.

[0037] In the embodiment of the present application, an inspection video recording the inspection process of the on-duty personnel is obtained, and the inspection video is first pre-inspected based on the temporal action detection model to determine the first probability value of the target action existing in the video segment corresponding to each predicted time range. Then, for each predicted time range, if the first probability value corresponding to the predicted time range does not meet the preset threshold requirement, the video segment corresponding to the predicted time range is detected based on the multimodal large model to obtain a detection result for describing the target action existing in the corresponding video segment, which effectively solves the problems of low accuracy and low efficiency of traditional manual quality inspection, and at the same time improves the accuracy of action detection by using the multimodal large model for secondary judgment.

[0038] The action detection process based on the multimodal large model is described below with reference to a specific embodiment.

[0039] First, obtain the inspection video recording the inspection process of the on-duty personnel, and use the constructed temporal action detection model to determine the prediction time range and the category label a corresponding to each prediction time range [t1, t2] i,i=1,...N, where the category label represents the action to be identified, N represents the number of actions to be identified, and N can be any integer.

[0040] Determine the probability of the category label corresponding to each prediction time range one by one. If the probability is less than the threshold, take the video segment at the time [t1, t2] and input it into the multimodal large model. Use the multimodal large model to perform secondary action recognition. If the multimodal large model recognizes the corresponding action, retain the category label corresponding to the prediction time range; otherwise, remove the category label corresponding to the prediction time range, and use the action recognized by the multimodal large model as the category label corresponding to the prediction time range.

[0041] In order to further improve the accuracy of action recognition, based on the above embodiment, in the embodiment of the present application, the method of determining the first probability value of the target action existing in the video segment corresponding to each predicted time range based on the temporal action detection model includes:

[0042] The inspection video is input into the temporal action detection model, and the temporal action detection model selects M candidate time ranges within the prior time range corresponding to each preset action, where M is a positive integer; for each candidate time range, a second probability value of each preset action existing in the video segment corresponding to the candidate time range is determined;

[0043] The temporal action detection model extracts features from the inspection video to obtain a first feature matrix, and performs full connection processing on the first feature matrix to obtain a third probability value of each preset action existing in the video segment corresponding to each predicted time range, wherein the number of the predicted time ranges is K, and K is the product value of M and the number of the preset actions;

[0044] Determine, based on the second probability value of each preset action occurring in each candidate time range of the inspection video and the third probability value of each preset action occurring in each predicted time range, a fourth probability value of each preset action existing in the video segment corresponding to each predicted time range;

[0045] For each predicted time range, the preset action corresponding to the maximum fourth probability value is determined as the target action existing in the video segment of the predicted time range, and the maximum fourth probability value is determined as the first probability value corresponding to the preset time range.

[0046] Although existing temporal action detection models can detect actions, these models are general-purpose action detection for various general scenarios, and their accuracy in identifying professional actions in railway turnout inspection is relatively low. Through a large number of statistical analyses, it is found that there is special prior knowledge in the railway turnout inspection scenario, that is, the occurrence time of each action has a fixed range of values, and the order between some actions is relatively fixed. Table 1 shows the prior time ranges of typical preset actions and the actions corresponding to the prior time ranges.

[0047] Table 1

[0048] Preset Actions Prior time range Actions corresponding to the prior time range Switch shooting 0-40s Switch number photography, switch height and direction Turnout height direction 32-45s Switch number photography, switch height direction, insulating joints Insulation joint 42-51s Fork height direction, insulation joint The point rail is closely attached to the stock rail 47-64s Insulation joints, point rails and base rails are tightly fitted …… … … Fill in the work record 421-505s Background and Operation Inspection

[0049] It can be seen from Table 1 that there are intersections between the a priori time ranges corresponding to different preset actions. Of course, there may be no intersection between the a priori time ranges corresponding to different preset actions, and those skilled in the art may configure as needed.

[0050] Therefore, in the embodiment of the present application, a time-series action detection model dedicated to railway turnout inspection can be pre-trained. When performing action recognition, the inspection video can be input into the time-series action detection model, and the time-series action detection model determines the first probability value of the target action existing in the video segment corresponding to each predicted time range based on the prior knowledge of each preset action.

[0051] Since the on-duty personnel can perform corresponding actions in any time period within the theoretical prior time range during the actual on-duty process, based on this, in the embodiment of the present application, the temporal action detection model can select M candidate time ranges within the prior time range corresponding to each preset action, where M is a positive integer. The candidate time range can be understood as the video segment corresponding to the candidate time range, in which the on-duty personnel may have performed the preset action.

[0052] After each candidate time range is obtained, a second probability value of the preset action existing in the video segment corresponding to the candidate time range may be determined for each candidate time range.

[0053] In a possible implementation, the temporal action detection model may include an action recognition sub-model, and when determining the second probability value, recognition is performed based on the action recognition sub-model.

[0054] In a possible implementation manner, when determining the second probability value of each preset action existing in the video segment corresponding to the candidate time range, the determination may be performed based on the following formula:

[0055]

[0056] Among them, Pk represents the candidate time range; r i represents the prior time range corresponding to the preset action i; ai represents the probability, through the candidate time range P k The prior time range r corresponding to the i-th preset action i The intersection-and-union ratio can be used to measure the candidate time range P k The probability of belonging to category i; represents the second probability value of the preset action i in the video segment corresponding to the candidate time range; c represents the number of preset actions, j represents the jth preset action; a identifies the second probability value of each preset action in the video segment corresponding to the candidate time range.

[0057] After determining each second probability value based on prior knowledge, the time series action detection model can perform feature extraction on the inspection video to obtain a first feature matrix, and perform full connection processing on the first feature matrix to obtain a third probability value for the existence of each preset action in the video segment corresponding to each predicted time range. Among them, the number of predicted time ranges can be K, which is the product of M and the number of preset actions. The specific value of K can be set in the model building stage. It can also be that after each candidate time range is obtained, each candidate time range is identified in the inspection video, so that the time series action detection model refers to the identified prior knowledge when determining the predicted time range for prediction. In an embodiment of the present application,

[0058] When extracting features from inspection videos, 3D convolution, BN layers, and activation functions can be used to build a feature extraction module, which is repeated N times as a feature extraction network.

[0059] After obtaining each third probability value, the fourth probability value of each preset action in the video segment corresponding to each predicted time range can be determined based on the second probability value of each preset action appearing in each candidate time range of the inspection video and the third probability value appearing in each predicted time range. Specifically, the second probability value and the third probability value of the corresponding time range can be fused in sequence according to the order of the time range. For example, the second probability value corresponding to the first candidate time range is fused with the third probability value corresponding to the first predicted time range, for example, the second probability value is added to the third probability value to obtain the fourth probability value.

[0060] After obtaining each fourth probability value, for each predicted time range, the preset action corresponding to the maximum fourth probability value can be determined as the target action existing in the video segment of the predicted time range, and the maximum fourth probability value can be determined as the first probability value corresponding to the preset time range. In the embodiment of the present application, since the fourth probability value is obtained by adding multiple probability values, in order to unify the measurement, the fourth probability value can be normalized in the embodiment of the present application. Specifically, it can be based on the following formula:

[0061] p out =softmax(b+a)

[0062] Wherein, b represents the third probability value; a represents the second probability value; b+a represents the fourth probability value; softmax represents the normalization function; P out Represents the normalized probability value.

[0063] In an embodiment of the present application, when determining the first probability value, the temporal action detection model integrates the model prediction result with the prior knowledge, thereby improving the accuracy of the action detection.

[0064] In order to further improve the accuracy of action recognition, on the basis of the above embodiments, in the embodiment of the present application, the first feature matrix is ​​fully connected to obtain a third probability value of each preset action appearing in each predicted time range of the inspection video, including:

[0065] The temporal action detection model performs a first full connection process on the first feature matrix to obtain a second feature matrix;

[0066] Performing a second full connection process on the second feature matrix to obtain K prediction time ranges;

[0067] A third full connection process is performed on the second feature matrix to obtain a third probability value of each preset action being included in the video segment corresponding to each prediction time range.

[0068] In an embodiment of the present application, when the first feature matrix is ​​fully connected to obtain the third probability value of each preset action appearing in each predicted time range of the inspection video, the temporal action detection model can perform a first full connection process on the first feature matrix to obtain a second feature matrix. In order to determine the predicted time range corresponding to the inspection video, and the third probability value of the preset action contained in the video segment corresponding to each predicted time range, in an embodiment of the present application, the second feature matrix can be subjected to two different full connection processes. Specifically, the second feature matrix can be subjected to a second full connection process to obtain K predicted time ranges, and the second feature matrix can be subjected to a third full connection process to obtain the third probability value of each preset action contained in the video segment corresponding to each predicted time range. It should be noted that how to perform a full connection process on a known feature matrix to obtain a preset output result is a prior art, and the embodiment of the present application will not repeat this process.

[0069] In an embodiment of the present application, after extracting features from the inspection video to obtain a first feature matrix, a feature transformation is performed on the first feature matrix to obtain a second feature matrix. Two fully connected layer branches are then used for action background classification and time range regression, respectively, thereby effectively improving the accuracy of action detection.

[0070] The architecture of the temporal action detection model is described below with reference to a specific embodiment. Figure 3 A schematic diagram of a temporal action detection model architecture provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, after obtaining the input video, the video is first subjected to a priori candidate extraction and a priori category label construction. Among them, the a priori candidate extraction is the candidate time range determination process described in the above embodiments, and the a priori category label construction is the second probability value determination process described in the above embodiments. The 3D convolution, the batch normalization (BN) layer and the activation function in the temporal action detection model jointly construct a feature extraction module, which is repeated N times as a feature extraction network, and then the first feature matrix output by the feature extraction network is subjected to feature conversion through a fully connected layer (FC), and then the second feature matrix after the feature conversion is respectively input into two FC layer branches, which are used for action background classification and time range regression respectively. It should be noted that how to perform feature extraction and how to perform full connection processing are prior art.

[0071] In order to further improve the accuracy of action recognition, on the basis of the above embodiments, in the embodiment of the present application, the M candidate time ranges are selected within the prior time range corresponding to the preset action, including:

[0072] Select the occurrence duration from the duration range saved in advance for the preset action;

[0073] Within the prior time range corresponding to the preset action, M candidate time ranges that meet the occurrence duration are selected.

[0074] Although different preset actions have corresponding a priori time ranges, in the actual on-duty process, the on-duty personnel can perform the corresponding preset actions at any time within the corresponding a priori time range. In addition, different on-duty personnel may be faster or slower when performing a preset action. Therefore, in the embodiment of the present application, the shortest occurrence time and the longest occurrence time are saved in advance for each preset action, and the longest occurrence time can be the difference between the two endpoints of the a priori time range.

[0075] In the embodiment of the present application, when selecting a candidate time range corresponding to a preset action, the occurrence duration can be selected from the duration range pre-stored for the preset action. The duration range can be expressed as [d i ,t 2i -t 1i ], where d i Indicates the shortest duration of the preset action i, t 2i -t 1i Indicates the maximum duration of the preset action i, t 2i and t 1i represents the two endpoints of the prior time range corresponding to the preset action i, where t 2i >t 1i The number of selected occurrence durations can be one or more.

[0076] After the occurrence duration is determined, M candidate time ranges that meet the occurrence duration may be selected from the prior time range corresponding to the preset action.

[0077] Specifically, assuming that the candidate time range corresponding to the preset action A is [30s, 150s], it can be determined that the longest occurrence duration of the preset action A is 120 seconds, and the preset configuration of the shortest occurrence duration of the preset action A is 20 seconds. This means that in the actual on-duty process, the duration of the on-duty personnel performing the preset action A is at least 20 seconds and at most 120 seconds. The specific duration of the continuous occurrence depends on different people and can be any duration within the duration range of [20, 120]. Assuming that the selected occurrence duration within the duration range [20, 120] is 50 seconds, then M candidate time ranges with a duration of 50 seconds can be selected within the prior time range [30s, 150s] corresponding to the preset action A. Assuming that M is 2, the selected candidate time ranges can be the 40th to 90th seconds in the inspection video, and the 50th to 100th seconds in the inspection video.

[0078] In an embodiment of the present application, a reasonable occurrence duration is selected within the duration range of a preset action, and the subsequently selected candidate time range satisfies the occurrence duration, thereby ensuring the rationality of the selected candidate time range and improving the accuracy of action detection.

[0079] In order to further improve the accuracy of action recognition, based on the above embodiments, in the embodiment of the present application, the selecting of the occurrence duration within the duration range pre-saved for the preset action includes:

[0080] Determine the number of equal divisions according to the duration range corresponding to the preset action and K;

[0081] The duration range is equally divided according to the number of equal divisions, and the duration value corresponding to each division point after the equal division is determined as the occurrence duration.

[0082] In order to detect all actions in the inspection video as much as possible, in the embodiment of the present application, multiple durations are selected within the duration range, and in order to ensure that the number of candidate time ranges finally selected is K, in the embodiment of the present application, the number of equal divisions can be determined according to the duration range corresponding to the preset action and K, and the duration range is divided into equal parts according to the number of equal divisions, and the duration value corresponding to each divided point after the equal division is determined as the occurrence duration. For example, assuming that 3 candidate time ranges are selected for each occurrence duration in advance, then it can be determined

[0083] Specifically, assuming that the duration range is divided into equal parts, the duration values ​​corresponding to each division point are 25 seconds, 33 seconds and 42 seconds respectively, then it can be determined that the occurrence duration is 25 seconds, 33 seconds and 42 seconds, and the duration of the subsequent candidate time range selected needs to meet any one of 25 seconds, 33 seconds and 42 seconds.

[0084] In order to reasonably select the candidate time range, on the basis of the above embodiments, in the embodiment of the present application, within the prior time range corresponding to the preset action, M candidate time ranges that meet the occurrence duration are selected, including:

[0085] According to the prior time range corresponding to the preset action and the pre-saved key time point, M candidate time ranges that meet the occurrence duration are selected, wherein the starting time point or the ending time point of any candidate time range is the key time point.

[0086] In order to ensure the rationality of the selection of candidate time ranges, in an embodiment of the present application, key time points can be pre-saved. When selecting a candidate time range, M candidate time ranges that meet the occurrence duration can be selected based on the prior time range corresponding to the preset action and the pre-saved key time points, wherein the starting time point or the ending time point of any candidate time range is the key time point.

[0087] Specifically, the candidate time range may be selected based on the following formula:

[0088]

[0089] Among them, t 1i ,t 2i , Indicates the key time point; t 1i ,t 2i represents the left and right endpoints of the prior time range of the preset action i, where t 2i >t 1i ;d ik Indicates the duration of the kth occurrence of preset action i.

[0090] In order to further reasonably select the candidate time range, based on the above embodiments, in the embodiment of the present application, the method further includes:

[0091] The priori time range corresponding to the preset action is determined as the candidate time range.

[0092] Since in the actual duty process, the on-duty personnel may continue to perform a certain action for the longest duration, or even exceed the duration, in order to be more realistic, in an embodiment of the present application, when selecting the candidate time range, the prior time range corresponding to the preset action can be determined as the candidate time range.

[0093] In order to obtain a preset number of candidate time ranges, based on the above embodiments, in the embodiment of the present application, the method further includes:

[0094] Count the number of candidate time ranges currently selected;

[0095] If the number is less than K, then determine the difference between K and the number; randomly select a supplementary time range of the difference within the prior time range, and the supplementary time range does not overlap with the currently selected candidate time range;

[0096] The supplementary time range is determined as the candidate time range.

[0097] Since the number K of candidate time ranges to be obtained is predefined in the embodiments of the present application, based on the above embodiments, after all candidate time ranges are obtained, the number of currently selected candidate time ranges can be counted. If the number is less than K, it means that the number of obtained candidate time ranges is insufficient, and the difference between K and the number can be determined. And randomly select the difference supplementary time range within the prior time range. In order to ensure that the candidate time range is fully obtained, the supplementary time range may not overlap with the currently selected candidate time range. The supplementary time range can be determined as the candidate time range.

[0098] Specifically, for the remaining n = k-len (L) candidates, we can randomly select [t 1i ,t 2i ]Take n supplementary time ranges that do not overlap with the candidate time ranges already obtained as candidate time ranges.

[0099] The process of determining the candidate time range is described below with reference to a specific embodiment. Constructing a candidate time range (proposal) is a key step in temporal action detection. Traditional methods generally use methods such as sliding windows, video frame grouping, and unit regression. However, for railway engineering actions, this is not necessarily the best method. This is because there is a strong prior knowledge of railway engineering actions, that is, each action has a fixed prior time range. In order to make full use of this prior knowledge, a technical solution for determining the candidate time range using a prior time range is proposed in the embodiments of the present application. Specifically, for a certain action i, let its prior range time be [t 1i ,t 2i ], t1i and t 2i They represent the earliest start time and the longest end time of the ith action respectively, so the longest time range of the action is t 2i -t 1i ; The key to the temporal action detection model is to construct multiple candidate time ranges, set to k; set the shortest statistical duration of the i-th action to be d i , then the k candidate time ranges of the action are constructed as follows:

[0100] Step 1: Construct a candidate time range list L and initialize a fixed candidate time range: L = {[t 1i ,t 2i ]}.

[0101] Step 2: Add [d i ,t 2i -t 1i ] is divided into An integer, [·] means rounding down.

[0102] Step 3: For each integer d in step 2 ik (d i0 =d i ), calculate the following three candidate time ranges:

[0103] [t 1i ,t 1i +d ik ],[t 2i -d ik ,t 2i ],

[0104] After obtaining the above candidate time ranges, these candidate time ranges are added to the candidate time range list L.

[0105] Step 4: For the remaining n=k-len(L) candidates, randomly select 1i ,t 2i ]Take n time ranges that do not overlap with those in the list and add them to the time range list L, where len represents the candidate time range list.

[0106] Figure 4 A schematic diagram of constructing a candidate time range provided in an embodiment of the present application is shown in FIG. Figure 4As shown, the preset actions included in the inspection video are: turnout height direction, insulating joint, basic rail, etc., and the operation record is filled in. There is an intersection between the prior time ranges corresponding to different preset actions. The candidate time ranges determined based on the above strategy are of unequal lengths. Specifically, the candidate time ranges corresponding to the turnout height direction can be [32,38], [33,41], ... [35,38]; the candidate time ranges corresponding to the insulating joint can be [42,51], [44,50], ... [44,51]; ... the candidate time ranges corresponding to the operation record can be [421,505], [430,489], ... [450,498].

[0107] In order to further improve the accuracy of action detection, based on the above embodiments, in the embodiment of the present application, the fine-tuning training process of the multimodal large model includes:

[0108] Acquire an instruction data set, wherein the instruction data set includes video global description training data, video detail description training data, and video comparison training data;

[0109] The original multimodal large model is fine-tuned and trained based on the training data included in the instruction data set to obtain the multimodal large model.

[0110] The general multimodal large model of video understanding is suitable for general video recognition, but it is difficult to use in professional video analysis tasks such as railway turnout inspection. Therefore, it is necessary to construct training data to fine-tune the open source general multimodal large model (hereinafter referred to as the original multimodal large model) and embed the professional knowledge in a specific field into the original multimodal large model. In the embodiment of the present application, an instruction data set is designed to allow the original multimodal large model to learn professional knowledge in a specific field. After fine-tuning the original multimodal large model using the instruction data set, the obtained multimodal large model has the ability to recognize professional knowledge in a specific field, for example, the ability to recognize turnout inspection actions. It should be noted that the equipment used for fine-tuning training of the multi-model may be the same as the equipment involved in the above-mentioned embodiments, or it may be different.

[0111] In an embodiment of the present application, the original multimodal large model can be fine-tuned and trained based on the instruction data set to obtain a multimodal large model for action recognition. When performing fine-tuning training, an instruction data set can be obtained. Since the multimodal large model needs to have a strong video understanding ability, in an embodiment of the present application, the multimodal large model can be trained from various aspects. In an embodiment of the present application, the instruction data set may include video global description training data, video detail description training data, and video comparison training data. After obtaining the instruction data set, the original multimodal large model can be fine-tuned and trained based on the training data included in the instruction data set to obtain a multimodal large model.

[0112] Specifically, the instruction data set may be constructed according to the following example. Of course, those skilled in the art may also construct it as needed.

[0113] 1) Video global description training data: mainly describes the video content, video background and other relevant information in the video.

[0114] Exemplarily, the video global description training data may be:

[0115] Q: Please describe the content of the video in detail

[0116] A: The video shows a railway inspection scene where railway workers are knocking...

[0117] 2) Video detail description training data: mainly asks about the actions, background and other relevant information in the video.

[0118] For example, the video detail description training data may be:

[0119] Q: What are the workers doing in the video?

[0120] A: In the video, workers are testing the tightness of the point rail and the base rail.

[0121] 3) Video comparison training data: mainly asks the difference between two videos.

[0122] Exemplarily, the video comparison training data may be:

[0123] Q: What is the difference between video [1] and video [2]?

[0124] A: Video 1 shows workers recording their work, while Video 2 shows…

[0125] Through the above three methods, a professional instruction data set in the railway official field is constructed, and then the benchmark large model is trained. In the embodiment of the present application, an efficient parameter fine-tuning method can be used to obtain a multimodal large model of railway official vertical domain video understanding.

[0126] The following describes the action detection process based on a multimodal large model in conjunction with a specific embodiment. Figure 5 A schematic diagram of an action detection process based on a multimodal large model is provided in an embodiment of the present application, such as Figure 5As shown, first obtain the input inspection video, and then determine the first probability value of the target action in the video segment corresponding to each predicted time range based on the time series action detection model. For each predicted time range, if the first probability value corresponding to the predicted time range is greater than the threshold, the target action is directly output as the result. If the first probability value corresponding to the predicted time range is not greater than the threshold, the video segment corresponding to the predicted time range is detected based on the multimodal large model, and the detection result is obtained and output. In an embodiment of the present application, the final output result is used to describe the target action in the inspection video, that is, to detect what operations the on-duty personnel have performed.

[0127] In the embodiments of the present application, a new solution for railway switching inspection is proposed, which uses an artificial intelligence algorithm to automatically identify video inspection actions. Specifically, a time-series action detection model and a multimodal large model are combined for joint judgment. The time-series action detection network and the multimodal large model of video understanding for railway inspection video enhancement are improved respectively. The present application is a large model and visual model fusion solution that can improve the recognition speed while improving the accuracy. That is, the proposed time-series detection model is first used for action recognition, and then the large model is used for secondary judgment of low-probability actions.

[0128] Based on the above embodiments, Figure 6 A schematic diagram of the structure of a motion detection device based on a multimodal large model provided in an embodiment of the present application, the device comprising:

[0129] The acquisition module 601 is used to acquire the inspection video recording the inspection process of the on-duty personnel;

[0130] The detection module 602 is used to determine the first probability value of the target action existing in the video segment corresponding to each predicted time range based on the temporal action detection model; for each predicted time range, if the first probability value corresponding to the predicted time range does not meet the preset threshold requirement, the video segment corresponding to the predicted time range is detected based on the multimodal large model to obtain a detection result, and the detection result is used to describe the target action existing in the corresponding video segment.

[0131] In a possible implementation, the detection module 602 is specifically used to input the inspection video into the temporal action detection model, wherein the temporal action detection model selects M candidate time ranges within the prior time range corresponding to each preset action, where M is a positive integer; for each candidate time range, a second probability value of each preset action existing in the video segment corresponding to the candidate time range is determined; the temporal action detection model performs feature extraction on the inspection video to obtain a first feature matrix, and performs full connection processing on the first feature matrix to obtain a third probability value of each preset action existing in the video segment corresponding to each predicted time range, wherein the number of predicted time ranges is K, and K is the product of M and the number of preset actions; according to the second probability value of each preset action appearing in each candidate time range of the inspection video and the third probability value appearing in each predicted time range, a fourth probability value of each preset action existing in the video segment corresponding to each predicted time range is determined; for each predicted time range, the preset action corresponding to the maximum fourth probability value is determined as the target action existing in the video segment of the predicted time range, and the maximum fourth probability value is determined as the first probability value corresponding to the preset time range.

[0132] In one possible implementation, the detection module 602 is specifically used for the temporal action detection model to perform a first full-connection processing on the first feature matrix to obtain a second feature matrix; perform a second full-connection processing on the second feature matrix to obtain K prediction time ranges; perform a third full-connection processing on the second feature matrix to obtain a third probability value of each preset action contained in the video segment corresponding to each prediction time range.

[0133] In a possible implementation, the detection module 602 is specifically used to select an occurrence duration within a duration range pre-saved for the preset action; and within a priori time range corresponding to the preset action, select M candidate time ranges that meet the occurrence duration.

[0134] In a possible implementation, the detection module 602 is specifically used to determine the number of equal divisions based on the duration range corresponding to the preset action and K; divide the duration range into equal divisions according to the number of equal divisions, and determine the duration value corresponding to each dividing point after the equal division as the occurrence duration.

[0135] In a possible implementation, the detection module 602 is specifically used to select M candidate time ranges that meet the occurrence duration based on the prior time range corresponding to the preset action and the pre-saved key time point, wherein the starting time point or the ending time point of any candidate time range is the key time point.

[0136] In a possible implementation, the detection module 602 is specifically configured to determine a priori time range corresponding to the preset action as the candidate time range.

[0137] In one possible implementation, the detection module 602 is specifically used to count the number of candidate time ranges currently selected; if the number is less than K, determine the difference between K and the number; randomly select the difference number of supplementary time ranges within the prior time range, and the supplementary time range does not overlap with the currently selected candidate time range; and determine the supplementary time range as the candidate time range.

[0138] In a possible implementation, the acquisition module 601 is further used to acquire an instruction data set, wherein the instruction data set includes video global description training data, video detail description training data, and video comparison training data;

[0139] The training module 603 is used to fine-tune the original multimodal large model based on the training data included in the instruction data set to obtain the multimodal large model.

[0140] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, it includes: a processor 701, a communication interface 702, a memory 703 and a communication bus 704, wherein the processor 701, the communication interface 702, and the memory 703 communicate with each other through the communication bus 704;

[0141] The memory 703 stores a computer program. When the program is executed by the processor 701, the processor 701 executes the steps of the action detection method based on the multimodal large model described in the above embodiments.

[0142] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface 702 is used for communication between the above electronic device and other devices. The memory may include a random access memory (RAM) and may also include a non-volatile memory (NVM), such as at least one disk storage. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0143] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (Network Processor, NP), etc.; it can also be a digital signal processing processor (Digital Signal Processing, DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.

[0144] On the basis of the above embodiments, an embodiment of the present invention further provides a computer-readable storage medium, which stores a computer program executable by a processor. When the program runs on the processor, the processor implements the steps of the action detection method based on the multimodal large model described in the above embodiments.

[0145] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0146] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0147] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A motion detection method based on a multimodal large model, characterized in that: The method comprises: Obtain inspection videos recording the inspection process of on-duty personnel; Determine, based on the temporal action detection model, a first probability value of a target action existing in a video segment corresponding to each predicted time range; For each predicted time range, if the first probability value corresponding to the predicted time range does not meet the preset threshold requirement, the video segment corresponding to the predicted time range is detected based on the multimodal large model to obtain a detection result, which is used to describe the target action in the corresponding video segment.

2. The method according to claim 1, characterized in that The step of determining a first probability value of a target action existing in a video segment corresponding to each predicted time range based on a temporal action detection model includes: The inspection video is input into the temporal action detection model, and the temporal action detection model selects M candidate time ranges within the prior time range corresponding to each preset action, where M is a positive integer; for each candidate time range, a second probability value of each preset action existing in the video segment corresponding to the candidate time range is determined; The temporal action detection model extracts features from the inspection video to obtain a first feature matrix, and performs full connection processing on the first feature matrix to obtain a third probability value of each preset action existing in the video segment corresponding to each predicted time range, wherein the number of the predicted time ranges is K, and K is the product value of M and the number of the preset actions; Determine, based on the second probability value of each preset action occurring in each candidate time range of the inspection video and the third probability value of each preset action occurring in each predicted time range, a fourth probability value of each preset action existing in the video segment corresponding to each predicted time range; For each predicted time range, the preset action corresponding to the maximum fourth probability value is determined as the target action existing in the video segment of the predicted time range, and the maximum fourth probability value is determined as the first probability value corresponding to the preset time range.

3. The method according to claim 2, characterized in that The first feature matrix is ​​fully connected to obtain a third probability value of each preset action occurring in each predicted time range of the inspection video, including: The temporal action detection model performs a first full connection process on the first feature matrix to obtain a second feature matrix; Performing a second full connection process on the second feature matrix to obtain K prediction time ranges; A third full connection process is performed on the second feature matrix to obtain a third probability value of each preset action being included in the video segment corresponding to each prediction time range.

4. The method according to claim 2, characterized in that: The selecting of M candidate time ranges within the prior time range corresponding to the preset action includes: Select the occurrence duration from the duration range saved in advance for the preset action; Within the prior time range corresponding to the preset action, M candidate time ranges that meet the occurrence duration are selected.

5. The method according to claim 4, characterized in that The selecting of the occurrence duration within the duration range saved in advance for the preset action includes: Determine the number of equal divisions according to the duration range corresponding to the preset action and K; The duration range is equally divided according to the number of equal divisions, and the duration value corresponding to each division point after the equal division is determined as the occurrence duration.

6. The method according to claim 4, characterized in that In the prior time range corresponding to the preset action, M candidate time ranges satisfying the occurrence duration are selected, including: According to the prior time range corresponding to the preset action and the pre-saved key time point, M candidate time ranges that meet the occurrence duration are selected, wherein the starting time point or the ending time point of any candidate time range is the key time point.

7. The method according to claim 4, characterized in that The method further comprises: The priori time range corresponding to the preset action is determined as the candidate time range.

8. The method according to any one of claims 4 to 7, characterized in that: The method further comprises: Count the number of candidate time ranges currently selected; If the number is less than K, then determine the difference between K and the number; randomly select a supplementary time range of the difference within the prior time range, and the supplementary time range does not overlap with the currently selected candidate time range; The supplementary time range is determined as the candidate time range.

9. The method according to claim 1, characterized in that: The fine-tuning training process of the multimodal large model includes: Acquire an instruction data set, wherein the instruction data set includes video global description training data, video detail description training data, and video comparison training data; The original multimodal large model is fine-tuned and trained based on the training data included in the instruction data set to obtain the multimodal large model.

10. An electronic device, characterized in that: The electronic device comprises a processor, and the processor is used to implement the steps of the action detection method based on a multimodal large model as claimed in any one of claims 1 to 9 when executing a computer program stored in a memory.