Target position stage identification method, device, equipment and medium

By combining a pre-trained model with a bidirectional long short-term memory network, the accuracy and efficiency issues of target location stage recognition in videos are solved, achieving more efficient stage recognition and evaluation.

CN122454472APending Publication Date: 2026-07-24PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY) +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
Filing Date
2026-03-11
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies suffer from poor accuracy, transient jitter, uneven distribution of sample data, insufficient model generalization ability, and excessive consumption of computing resources when identifying target locations in videos.

Method used

This paper employs a combination of a pre-trained model and a bidirectional long short-term memory network. By acquiring video frames and performing spatial feature extraction and temporal analysis, the pre-trained model is mapped into a multi-dimensional spatial feature vector. The bidirectional long short-term memory network is then used to capture temporal dependencies and perform stage category prediction. Smoothing and padding processes are then used to improve recognition accuracy.

Benefits of technology

It improves the accuracy and efficiency of target location stage identification, enabling more accurate assessment of operational scenarios and meeting the practical needs of high efficiency and low cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454472A_ABST
    Figure CN122454472A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a target position stage recognition method, device, equipment and medium. It comprises: acquiring a plurality of to-be-processed video frames of a target position; using a preset pre-training model to process the plurality of to-be-processed video frames to obtain a first spatial feature vector of the target position; based on a pre-trained bidirectional long short-term memory network, processing the first spatial feature vector of the target position to obtain a first predicted stage category of the target position; and determining a target recognition stage of the target position according to the first predicted stage category of the target position. Thus, for a video of a certain position, the video frames of the video are first acquired, then the spatial features and stage categories are extracted using a pre-training model and a bidirectional long short-term memory network, and finally the final stage of the position is determined based on the stage category predicted by the model, thereby improving the stage recognition accuracy and effectively evaluating the work scene and assisting the further research of the work scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This disclosure claims priority to Chinese Patent Application No. 2025113923573, filed on September 26, 2025, entitled "Method, Apparatus, Device and Medium for Stage Identification of Target Location", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of deep learning technology, and in particular to a method, apparatus, device and medium for phase recognition of target location. Background Technology

[0003] With the development of information technology, identifying the stage of a specific body part at different times from videos corresponding to actual work scenarios can effectively evaluate the work scenario and assist in further research. For example, identifying the stage of a certain tissue or organ from medical imaging videos can effectively assess the quality of clinical work and assist in clinical research.

[0004] In related technologies, convolutional neural networks or attention-based classifiers, or processor-based decoding schemes, are often used to identify the stage of a certain part at different times from videos corresponding to actual work scenarios. However, this approach often suffers from stage identification errors and has poor accuracy, requiring improvement. Summary of the Invention

[0005] To address the aforementioned technical problems, this disclosure provides a method, apparatus, device, and medium for phase identification of target location.

[0006] Firstly, this disclosure provides a method for identifying the stage of a target location, including: Acquire multiple video frames to be processed at the target location; Using a pre-trained model, the multiple video frames to be processed are processed to obtain the first spatial feature vector of the target location; Based on a pre-trained bidirectional long short-term memory network, the first spatial feature vector of the target location is processed to obtain the first prediction stage category of the target location; The target identification stage of the target location is determined based on the first prediction stage category of the target location.

[0007] Secondly, this disclosure provides a target location phase identification device, the device comprising: The first acquisition module is used to acquire multiple video frames to be processed at the target location; The second acquisition module is used to process the multiple video frames to be processed using a preset pre-trained model to obtain the first spatial feature vector of the target location. The third acquisition module is used to process the first spatial feature vector of the target location based on a pre-trained bidirectional long short-term memory network to obtain the first prediction stage category of the target location. The first determining module is used to determine the target identification stage of the target location based on the first prediction stage category of the target location.

[0008] Thirdly, embodiments of this disclosure also provide an electronic device, the device comprising: One or more processors; Storage device for storing one or more programs. When one or more programs are executed by one or more processors, the one or more processors implement the methods provided in the first aspect.

[0009] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method provided in the first aspect.

[0010] The technical solution provided in this disclosure has the following advantages compared with the prior art: This disclosure discloses a method, apparatus, device, and medium for stage identification of a target location. The method involves acquiring multiple video frames to be processed at the target location; processing the multiple video frames using a pre-trained model to obtain a first spatial feature vector of the target location; processing the first spatial feature vector of the target location based on a pre-trained bidirectional long short-term memory network to obtain a first predicted stage category of the target location; and determining the target identification stage of the target location based on the first predicted stage category. Therefore, for a video at a certain location, by first acquiring the video frames, then sequentially extracting spatial features and stage categories using a pre-trained model and a bidirectional long short-term memory network, and finally determining the final stage of the location based on the stage category predicted by the model, the accuracy of stage identification can be effectively improved. This allows for effective evaluation of the work scenario and assists in further research on the work scenario. Attached Figure Description

[0011] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0012] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating a stage identification method for a target location provided in an embodiment of this disclosure; Figure 2 A flowchart illustrating a training method for a bidirectional long short-term memory network provided in an embodiment of this disclosure; Figure 3 A schematic diagram of the structure of a target location phase identification device provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0014] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0015] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0016] In related technologies, methods that rely on convolutional neural networks or attention-based classifiers, or processor-based decoding schemes, to identify the stage of a certain part at different times from videos corresponding to actual work scenarios, often suffer from the following problems: Problem 1: When switching between adjacent video frames, brief jitters or incorrect judgments often occur; the same stage is often misidentified, and there is a lack of targeted post-processing strategies. Question 2: The distribution of sample data is uneven in different stages. The sample size in some stages is small and easily overlooked or misjudged. Problem 3: The model has poor generalization ability across different datasets or job scenarios; Question 4: Long video processing takes a long time. The existing training process requires extracting a large number of video frames from long videos and performing feature extraction, which places extremely high demands on computing resources and storage overhead. Question 5: The deployment of existing models has stringent requirements for video memory and system memory, making it difficult to meet the requirements of high timeliness and low cost in operational scenarios.

[0017] To solve the above problems, the following will be combined with... Figure 1The method for identifying the stage of a target location provided in this disclosure is described below. In this disclosure, the method for identifying the stage of a target location can be executed by an electronic device. The electronic device may include devices with communication capabilities such as tablet computers, desktop computers, and laptop computers, or it may include devices simulated by virtual machines or simulators.

[0018] Figure 1 A flowchart illustrating a stage identification method for target location provided in an embodiment of this disclosure is shown.

[0019] like Figure 1 As shown, the phase identification method for the target location may include the following steps.

[0020] S110. Obtain multiple video frames to be processed at the target location.

[0021] In this embodiment, in a scenario where work is being performed on a target location, the electronic device acquires a video of the target location in that scenario and extracts video frames from the video to obtain multiple video frames to be processed at the target location.

[0022] The target location refers to the location in the video frame that needs to be identified.

[0023] In actual clinical settings, the target location can be a tissue (such as the duodenum) or an organ.

[0024] When the actual work site is a road surface scenario, the target location can be a lane location, an intersection, etc.

[0025] In the actual work scenario of home renovation, the target location can be a specific room.

[0026] In some embodiments, the specific implementation method of S110 includes, but is not limited to, the following methods: acquiring multiple videos to be processed from multiple different objects at the target location; performing streaming decoding on the multiple videos to be processed to obtain the decoding results of the multiple videos to be processed; and extracting keyframes from the decoding results of the multiple videos to be processed according to a preset frame rate to obtain multiple video frames to be processed.

[0027] Specifically, electronic devices can use cross-platform multimedia frameworks (such as FFmpeg) or cross-platform multimedia frameworks (such as CUDA) to incrementally decode and stream multiple videos to be processed, continuously outputting the original video frame byte stream, extracting one frame per second to obtain multiple video frames to be processed.

[0028] The process involves streaming decoding of multiple videos to be processed to obtain decoding results for multiple videos to be processed. Specifically, this includes: obtaining the sequence lengths corresponding to the multiple videos to be processed; aligning the sequence lengths of the multiple videos to be processed according to the maximum sequence length among the multiple videos to be processed to obtain multiple videos to be processed with the maximum sequence length; and streaming decoding of the multiple videos to be processed with the maximum sequence length to generate decoding results for multiple videos to be processed.

[0029] Specifically, for the temporal length differences that may exist among multiple different objects, first calculate the sequence length of each object, and then pad it by using a unified maximum length so that the feature tensors of multiple objects form an alignment matrix of equal length in the time dimension.

[0030] Thus, for videos with multiple objects that may have temporal length differences, the efficiency of frame extraction for multiple long videos can be improved by forming an alignment matrix of equal length in the time dimension using the feature tensors of multiple objects.

[0031] In some cases, electronic devices can also stitch together multiple videos of different objects for processing to make full use of the graphics processor's computing resources and main memory bandwidth. Furthermore, under the same graphics processor constraints, stitching can improve video processing efficiency.

[0032] By employing the above method, multiple video frames are extracted from multiple videos to be processed using a streaming frame extraction method, which significantly reduces the memory consumed in extracting video frames and improves the frame extraction efficiency of long videos.

[0033] S120. Using a pre-trained model, process multiple video frames to be processed to obtain the first spatial feature vector of the target location.

[0034] To improve the accuracy of stage recognition of target location, electronic devices can input multiple video frames to be processed into a preset pre-trained model. The pre-trained model can then be used to map each video frame to be processed into a multi-dimensional spatial feature vector (e.g., 1536 dimensions) to obtain the first spatial feature vector of the target location.

[0035] Optionally, the pre-trained model may include, but is not limited to, a pure convolutional neural network (e.g., the ConvNeXt base model).

[0036] In some embodiments, the specific implementation method of S120 includes, but is not limited to, the following methods: based on the normalization layer in the preset pre-trained model, normalize multiple video frames to be processed to obtain normalized features of multiple video frames to be processed; based on the convolution layer in the preset pre-trained model, convolve the normalized features of multiple video frames to be processed to obtain the first spatial feature vector of the target location.

[0037] Specifically, electronic devices introduce a pre-trained model as the backbone for spatial feature extraction, retain all its convolutional and normalization layers, remove the final classification head, and utilize its learned spatial representation capabilities to grasp a variety of detailed information in the operational scenario.

[0038] To improve the processing efficiency of video frames, electronic devices can also uniformly scale the resolution of the video frames to a single value (e.g., 796×448) and enable a decoder (such as an NVIDIA GPU decoder) to significantly reduce the processor's decoding overhead. The read video frame byte stream is then cached in a memory buffer. After accumulating up to, say, 32 frames, it is converted into a specific array (e.g., a NumPy array) and fed into a batch of tensors (e.g., PyTorch tensors). These batch tensors are then normalized and fed into a pre-trained model to extract spatial features, resulting in a multi-dimensional spatial feature vector, which is then concatenated. When the number of frames at the end of the video is insufficient (e.g., less than 32 frames), the remaining byte stream is still converted into a final tensor and its features are extracted. These features are then concatenated into a feature sequence for subsequent inference. For example, the tensors of each batch are directly fed into the graphics processor. For a video of duration T seconds, one frame is extracted every second, the original aspect ratio is maintained, and the image is scaled to a size of 796*448. This frame is used as the input to a pre-trained model. Each image outputs a high-dimensional feature vector of size 1536, and each video outputs a feature matrix of dimensions [T, 1536]. The results are saved as a .pt file in binary form. This process only needs to be done once, and these features are used as the input for subsequent processing.

[0039] In practice, the weights of each parameter in the pre-trained model can be stored in a weight file and labeled with a specific naming prefix (e.g., "teacher.backbone.*"). When the pre-trained model is called, this prefix is ​​automatically removed, and the weights are mapped to the corresponding convolutional and normalization layers. If the weight key name contains the prefix "module." (generated in a data-parallel environment), it is automatically removed to ensure compatibility between single-card and multi-card environments.

[0040] In this way, each video frame is mapped to a multi-dimensional spatial feature vector using a pre-trained model, providing a rich and generalizable representation for subsequent time series analysis.

[0041] S130. Based on a pre-trained bidirectional long short-term memory network, the first spatial feature vector of the target location is processed to obtain the first prediction stage category of the target location.

[0042] To improve the detection accuracy of switching boundaries between different video frames, the electronic device employs a pre-trained bidirectional long short-term memory network to simultaneously capture forward and reverse temporal dependencies in order to obtain the first-stage prediction category of the target location.

[0043] Optionally, the pre-trained bidirectional long short-term memory network includes, but is not limited to, Bi-LSTM.

[0044] In some embodiments, the specific implementation method of S130 includes, but is not limited to, the following method: based on the first long short-term memory network in the pre-trained bidirectional long short-term memory network, the first spatial feature vector of the target location is processed in chronological order to generate a first temporal feature; based on the second long short-term memory network in the pre-trained bidirectional long short-term memory network, the first spatial feature vector of the target location is processed in chronological order to generate a second temporal feature; the first temporal feature and the second temporal feature are concatenated to generate a first prediction stage category of the target location.

[0045] Specifically, the pre-trained bidirectional Long Short-Term Memory (LSTM) network consists of two LSM layers with a hidden dimension of 1024, bidirectional connections, and a total output dimension of 2048. These are then mapped to multiple categories through linear layers. During inference computation, frame-level feature tensors concatenated in sequence are directly input into the LSM network. The forward LSM network processes data from the beginning to the end of the sequence (from left to right), while the backward LSM network processes data from the end of the sequence to the beginning (from right to left). Finally, the output at each time step is the concatenation or summation of the outputs from both LSM networks, yielding a stage score for each time step, which serves as the first predicted stage category for the target location.

[0046] In some embodiments, the specific implementation method of S130 includes, but is not limited to, the following method: dividing the first spatial feature vector of the target location into a preset number of partial first spatial feature vectors, wherein the preset number is the number of processing modules, and each processing module deploys a bidirectional long short-term memory network; processing the partial first spatial feature vectors based on the bidirectional long short-term memory network deployed by each processing module to obtain the category recognition result of each processing module; and superimposing the category recognition results corresponding to all processing modules to obtain the first prediction stage category of the target location.

[0047] Specifically, when an electronic device processes multidimensional spatial features using a pre-trained bidirectional long short-term memory network, it first detects the number of processing modules (i.e., processors). If there are multiple processing modules, the electronic device copies the pre-trained bidirectional long short-term memory network to each module. After dividing the first spatial feature vector of the target location according to the number of processing modules, it distributes it to the pre-trained bidirectional long short-term memory networks of different processing modules for parallel computation. Then, it summarizes the computation results of multiple processing modules to obtain the first prediction stage category of the target location.

[0048] Optionally, the first prediction stage category is the category corresponding to different times in the actual operation scenario.

[0049] For example, in a clinical setting, the first prediction stage categories for the target location include tissue transection, pancreaticojejunostomy, choledochojejunostomy, gastrointestinal anastomosis, and transfer interval.

[0050] For example, when the actual operation scenario is a road surface scenario, the first prediction stage categories for the target location include peak period, off-peak period, and fault period.

[0051] In this way, by using a time-series model to process the multi-dimensional spatial features of each video frame, the inherent laws of the actual operation scenario process in the time dimension can be obtained, further enhancing the ability to identify the stage category of the target location.

[0052] S140. Determine the target identification stage of the target location based on the first prediction stage category of the target location.

[0053] In this embodiment, in order to improve the accuracy of target location identification, the electronic system further processes the first prediction stage category of the target location to obtain the target identification stage of the target location.

[0054] In some embodiments, the specific implementation method of S140 includes, but is not limited to, the following method: performing short segment smoothing and zero-value segment filling on the first prediction stage category of the target location to obtain the target recognition stage of the target location.

[0055] The short segment smoothing process is as follows: For the first prediction stage category output by the pre-trained bidirectional long short-term memory network, if there are short predictions with consecutive identical labels whose length is lower than the threshold (the number of frames corresponding to 125 seconds), the short predictions are replaced with the preceding or following label of the adjacent short segment to eliminate noise predictions caused by brief jitter of the model in adjacent frames.

[0056] The zero-value segment filling process is as follows: after the first prediction stage category is smoothed by short sequence, if there are multiple consecutive "no stage" label segments (zero values), and the stage labels before and after the segment are the same and are not "no stage", then the zero-value segment is replaced with the same stage label to avoid stage interruption caused by occlusion, noise or low network output probability.

[0057] In this way, by smoothing the initial category of the model output before filling it, we can preserve the true short-term signal while ensuring the consistency of long-term prediction.

[0058] This disclosure provides a method for stage identification of a target location. The method involves acquiring multiple video frames to be processed at the target location; processing these multiple video frames using a pre-trained model to obtain a first spatial feature vector of the target location; processing the first spatial feature vector of the target location using a pre-trained bidirectional long short-term memory network to obtain a first predicted stage category of the target location; and determining the target identification stage of the target location based on the first predicted stage category. Therefore, for a video at a certain location, by first acquiring the video frames, then sequentially extracting spatial features and stage categories using a pre-trained model and a bidirectional long short-term memory network, and finally determining the final stage of the location based on the stage category predicted by the model, the method can effectively improve stage identification accuracy, thereby effectively evaluating the operational scenario and assisting in further research on the operational scenario.

[0059] In another embodiment of this disclosure, a specific explanation is given of the training method for a bidirectional long short-term memory network.

[0060] Figure 2 A flowchart illustrating a training method for a bidirectional long short-term memory network provided in an embodiment of this disclosure is shown.

[0061] like Figure 2 As shown, the training method for this bidirectional long short-term memory network may include the following steps.

[0062] S210. Obtain multiple reference video frames at the target location.

[0063] In this embodiment, the electronic device acquires multiple reference videos containing the target location under the above-mentioned actual working scenario, and performs streaming decoding on the multiple reference videos to obtain multiple reference video frames of the target location.

[0064] In some cases, for reference videos with varying temporal lengths of multiple different objects, the feature tensors of the multiple objects can be aligned into an equal-length matrix along the temporal dimension. Specific methods can be found in the embodiments described above. This improves the frame extraction efficiency for multiple long videos, thereby enhancing the efficiency of stage recognition.

[0065] In some cases, electronic devices can also stitch together reference videos of multiple different objects for processing to make full use of the graphics processor's computing resources and main memory bandwidth. Furthermore, under the same graphics processor constraints, stitching can improve video processing efficiency.

[0066] By employing the above method, multiple video frames are extracted from multiple reference videos using a streaming frame extraction approach, which significantly reduces the memory consumed in extracting video frames and improves the frame extraction efficiency for long videos.

[0067] S220. Using a pre-trained model, process multiple reference video frames to obtain the second spatial feature vector of the target location.

[0068] In this embodiment, the electronic device can use a large number of video frames and multi-dimensional spatial features from real-world application scenarios to pre-train the model, enabling the model to have the ability to perceive real-world application scenarios and to extract features efficiently, thereby obtaining a preset pre-trained model.

[0069] Specifically, the electronic device first normalizes multiple video frames to be processed based on the normalization layer in the preset pre-trained model, and then performs convolution processing on the normalized features of the multiple video frames to be processed based on the convolution layer in the preset pre-trained model to obtain the first spatial feature vector of the target location.

[0070] In some cases, the weights of each parameter in a pre-trained model can be stored in a weight file and labeled with a specific naming prefix (e.g., "teacher.backbone.*"). When generating the pre-trained model, this prefix is ​​automatically removed, and the weights are mapped to the corresponding convolutional and normalization layers. If the weight key name contains the prefix "module." (generated in a data-parallel environment), it is automatically removed to ensure compatibility between single-GPU and multi-GPU environments.

[0071] S230. Based on the initial network, the second spatial feature vector of the target location is processed to obtain the second prediction stage category of the target location.

[0072] In this embodiment, the electronic device can use a temporal model as the initial network to continue extracting temporal features from the second spatial feature vector of the target location to obtain the second prediction stage category of the target location.

[0073] In some cases, when an electronic device processes the second spatial feature vector using an initial network, depending on the number of processing modules (i.e., processors) configured, if there are multiple processing modules, the electronic device copies the initial network to each module, divides the second spatial feature vector of the target location according to the number of processing modules, distributes it to the initial networks of different processing modules for parallel computation, and then summarizes the computation results of multiple processing modules to obtain the second prediction stage category of the target location.

[0074] In this way, a temporal model is used to process the multidimensional spatial features of each video frame to obtain the inherent laws of the actual operation scenario in the time dimension, further enhancing the initial model's ability to identify the stage category of the target location.

[0075] S240. Calculate the loss value of the initial network using the reference stage category and the second prediction stage category of multiple objects at the target location.

[0076] In this embodiment, the electronic device uses a preset loss function to calculate the loss of multiple objects at the target location in the reference stage category and the second prediction stage category, thereby obtaining the loss value of the initial network.

[0077] Optionally, the preset loss function may include, but is not limited to, the cross-entropy loss function, or other types of loss functions.

[0078] S250. Based on the loss value, iteratively train the initial network until the network reaches the preset cutoff condition after the current training iteration, thus obtaining a bidirectional long short-term memory network.

[0079] In some embodiments, the specific implementation method of S250 includes, but is not limited to, the following method: obtaining the target reference stage category and the second target stage category from the reference stage category and the second predicted stage category of the target location, where the loss value is greater than a preset loss threshold; iteratively training the initial network based on the target reference stage category and the second target stage category until the network reaches a preset cutoff condition for the current training iterations, thereby obtaining a bidirectional long short-term memory network.

[0080] Specifically, for each training batch, the classification loss of the forward output of the initial network and the corresponding position of the real label is calculated element by element. The samples with larger loss values ​​are selected for gradient backpropagation in order to achieve online mining of difficult samples.

[0081] Optionally, the proportion of the target reference stage category and the second target stage category whose loss value is greater than the preset loss threshold can be set between 10% and 50% to balance the needs of difficult sample mining depth and overall training stability.

[0082] In this way, by adopting an online hard sample mining mechanism, samples with larger loss values ​​are selected for gradient backpropagation. This allows for targeted optimization of hard samples, reducing the interference of most easily classified samples on the model gradient. This accelerates the convergence of the model when switching positions during the surgical phase and improves the accuracy and robustness of the classification boundary.

[0083] In some cases, to address the issue of uneven distribution of different stage categories in actual work scenarios, the frame-level proportion of each category in the labeled dataset is statistically analyzed in advance, and its reciprocal is normalized and used as the weight of the category-weighted cross-entropy loss function, thereby increasing the attention given to the minority class stages.

[0084] In some cases, other hyperparameters can be pre-set during model training. For example, the batch size can be set to 6, and the learning rate to 3×10⁻⁶. -4 The weight decay coefficient is set to 1×10. -4 .

[0085] In some embodiments, the specific implementation method of S250 includes, but is not limited to, the following method: based on the loss value, the parameters are updated and gradients are accumulated for the initial network iteration according to single precision, and the backpropagation and forward propagation are performed for the initial network iteration according to half precision until the network reaches the preset cutoff condition for the current training number, thereby obtaining a bidirectional long short-term memory network.

[0086] Specifically, during model training, the deep learning framework is used to dynamically schedule tensor computations using a hybrid approach of half-precision (FP16) and single-precision (FP32). Most computational tasks are converted to half-precision execution, while key nodes are kept in single precision to ensure numerical stability. Specifically, single precision is used for parameter updates and gradient accumulation, while half precision is used for backpropagation and forward propagation in the initial network iteration process.

[0087] In this way, by using mixed precision training, the memory usage during training can be reduced, allowing the hidden dimensions of the bidirectional long short-term memory network to be expanded under the same memory conditions, thereby improving the bidirectional long short-term memory network's ability to fit complex temporal patterns.

[0088] This embodiment provides a training method for a bidirectional long short-term memory network. During the model training process, a pre-trained model, a bidirectional long short-term memory network, a hard sample mining mechanism, a mixed precision training mechanism, and streaming deployment are introduced to achieve not only extremely high recognition accuracy in long-term, high-resolution real-world operation scenarios, but also improved end-to-end inference efficiency, thus balancing accuracy and timeliness and meeting the practical needs of real-world operation scenarios.

[0089] This disclosure also provides a target location stage identification device for implementing the above-described target location stage identification method. The following is in conjunction with... Figure 3 The following explanation is provided. In this embodiment, the target location phase identification device can be an electronic device. This electronic device can include devices with communication capabilities such as tablets, desktop computers, and laptops, or it can include devices simulated by virtual machines or simulators.

[0090] Figure 3 A schematic diagram of the structure of a target location phase identification device provided in an embodiment of the present disclosure is shown.

[0091] like Figure 3 As shown, the target location phase identification device 300 may include: The first acquisition module 310 is used to acquire multiple video frames to be processed at the target location; The second acquisition module 320 is used to process the multiple video frames to be processed using a preset pre-trained model to obtain the first spatial feature vector of the target location. The third acquisition module 330 is used to process the first spatial feature vector of the target location based on a pre-trained bidirectional long short-term memory network to obtain the first prediction stage category of the target location. The first determining module 340 is used to determine the target identification stage of the target location based on the first prediction stage category of the target location.

[0092] This disclosure discloses a target location stage identification device that acquires multiple video frames to be processed at a target location; processes the multiple video frames using a preset pre-trained model to obtain a first spatial feature vector of the target location; processes the first spatial feature vector of the target location based on a pre-trained bidirectional long short-term memory network to obtain a first predicted stage category of the target location; and determines the target identification stage of the target location based on the first predicted stage category. Therefore, for a video at a certain location, by first acquiring the video frames, then sequentially extracting spatial features and stage categories using a pre-trained model and a bidirectional long short-term memory network, and finally determining the final stage of the location based on the stage category predicted by the model, the device can effectively improve stage identification accuracy, thereby effectively evaluating the work scenario and assisting in further research on the work scenario.

[0093] In some embodiments of this disclosure, the first acquisition module 310 includes: The first acquisition unit is used to acquire multiple videos to be processed from multiple different objects at the target location; A streaming decoding unit is used to perform streaming decoding on the plurality of videos to be processed, and obtain the decoding results of the plurality of videos to be processed; The keyframe extraction unit is used to extract keyframes from the decoding results of the plurality of videos to be processed according to a preset frame rate, so as to obtain the plurality of video frames to be processed.

[0094] In some embodiments of this disclosure, the streaming decoding unit is specifically used for: Obtain the sequence lengths corresponding to the plurality of videos to be processed; According to the maximum sequence length among the multiple videos to be processed, the sequence lengths corresponding to the multiple videos to be processed are aligned to obtain multiple videos to be processed with the maximum sequence length; Stream decoding is performed on multiple videos of the maximum sequence length to generate decoding results for the multiple videos.

[0095] In some embodiments of this disclosure, the second acquisition module 320 includes: The normalization unit is used to perform normalization processing on the multiple video frames to be processed based on the normalization layer in the preset pre-trained model, so as to obtain the normalized features of the multiple video frames to be processed. A convolutional unit is used to perform convolution processing on the normalized features of the multiple video frames to be processed based on the convolutional layers in the preset pre-trained model, so as to obtain the first spatial feature vector of the target location.

[0096] In some embodiments of this disclosure, the third acquisition module 330 includes: The first processing unit is used to process the first spatial feature vector of the target location in chronological order based on the first long short-term memory network in the pre-trained bidirectional long short-term memory network, and generate the first temporal feature. The second processing unit is used to process the first spatial feature vector of the target location in the order from back to front, based on the second long short-term memory network in the pre-trained bidirectional long short-term memory network, to generate the second temporal feature. The splicing unit is used to splice the first time feature and the second time feature to generate the first prediction stage category of the target location.

[0097] In some embodiments of this disclosure, the third acquisition module 330 includes: A partitioning unit is used to divide the first spatial feature vector of the target location into a preset number of partial first spatial feature vectors, wherein the preset number is the number of processing modules, and each processing module deploys the bidirectional long short-term memory network. The third processing unit is used to process the partial first spatial feature vector based on the bidirectional long short-term memory network deployed in each processing module to obtain the category recognition result of each processing module. The overlay unit is used to overlay the category recognition results corresponding to all processing modules to obtain the first prediction stage category of the target location.

[0098] In some embodiments of this disclosure, the first determining module 340 is specifically used for: The first prediction stage category of the target location is subjected to short segment smoothing and zero-value segment filling to obtain the target recognition stage of the target location.

[0099] In some embodiments of this disclosure, the device further includes: The fourth acquisition module is used to acquire multiple reference video frames at the target location; The fifth acquisition module is used to process the multiple reference video frames using the preset pre-trained model to obtain the second spatial feature vector of the target location; The sixth acquisition module is used to process the second spatial feature vector of the target location based on the initial network to obtain the second prediction stage category of the target location; A calculation module is used to calculate the loss value of the initial network using the reference stage category of multiple objects at the target location and the second prediction stage category; The iterative training module is used to iteratively train the initial network based on the loss value until the network reaches a preset cutoff condition after the current training iteration, thereby obtaining the bidirectional long short-term memory network.

[0100] In some embodiments of this disclosure, the iterative training module includes: The second acquisition unit is used to acquire, from the reference stage category of the target location and the second prediction stage category, the target reference stage category and the second target ... stage category and the second target stage category, the target stage category and the second target stage category, respectively, the target reference stage category and the target stage category. The first iterative training unit is used to iteratively train the initial network based on the target reference stage category and the second target stage category until the network reaches a preset cutoff condition for the current training iteration, thereby obtaining the bidirectional long short-term memory network.

[0101] In some embodiments of this disclosure, the iterative training module includes: The second iterative training unit is used to update parameters and accumulate gradients for the initial network iteration based on the loss value, using single precision, and to perform backpropagation and forward propagation for the initial network iteration using half precision, until the network reaches a preset cutoff condition for the current training iteration, thus obtaining the bidirectional long short-term memory network.

[0102] It should be noted that, Figure 3 The stage identification device 300 for the target location shown can perform... Figures 1-2 The various steps in the method embodiment shown are implemented. Figures 1-2 The processes and effects in the method embodiments shown are not described in detail here.

[0103] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown.

[0104] like Figure 4 As shown, the electronic device may include a processor 401 and a memory 402 storing computer program instructions.

[0105] Specifically, the processor 401 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0106] Memory 402 may include a large-capacity storage for information or instructions. For example, and not limitingly, memory 402 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway device. In a particular embodiment, memory 402 is a non-volatile solid-state memory. In a particular embodiment, memory 402 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (Electrically Programmable ROM, EPROM), an electrically erasable programmable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0107] The processor 401 reads and executes computer program instructions stored in the memory 402 to perform the steps of the target location phase identification method provided in the embodiments of this disclosure.

[0108] In one example, the electronic device may also include a transceiver 403 and a bus 404. Wherein, as... Figure 4 As shown, the processor 401, memory 402 and transceiver 403 are connected via bus 404 and communicate with each other.

[0109] Bus 404 includes hardware, software, or both. For example, and not limitingly, a bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 404 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.

[0110] The following are embodiments of a computer-readable storage medium provided in this disclosure. This computer-readable storage medium and the target location stage identification method of the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the computer-readable storage medium, please refer to the embodiments of the target location stage identification method described above.

[0111] This embodiment provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a stage identification method for a target location, including: Acquire multiple video frames to be processed at the target location; Using a pre-trained model, the multiple video frames to be processed are processed to obtain the first spatial feature vector of the target location; Based on a pre-trained bidirectional long short-term memory network, the first spatial feature vector of the target location is processed to obtain the first prediction stage category of the target location; The target identification stage of the target location is determined based on the first prediction stage category of the target location.

[0112] Of course, the computer-executable instructions provided in the embodiments of this disclosure are not limited to the above-described method operations, but can also perform related operations in the target location stage identification method provided in any embodiment of this disclosure.

[0113] Based on the above description of the implementation methods, those skilled in the art can clearly understand that this disclosure can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer cloud platform (which may be a personal computer, server, or network cloud platform, etc.) to execute the target location stage identification method provided in the various embodiments of this disclosure.

[0114] Note that the above description is merely a preferred embodiment and the technical principles employed in this disclosure. Those skilled in the art will understand that this disclosure is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this disclosure. Therefore, although this disclosure has been described in detail through the above embodiments, it is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this disclosure, and the scope of this disclosure is determined by the scope of the appended claims.

Claims

1. A method for stage identification of target location, characterized in that, include: Acquire multiple video frames to be processed at the target location; Using a pre-trained model, the multiple video frames to be processed are processed to obtain the first spatial feature vector of the target location; Based on a pre-trained bidirectional long short-term memory network, the first spatial feature vector of the target location is processed to obtain the first prediction stage category of the target location; The target identification stage of the target location is determined based on the first prediction stage category of the target location.

2. The method according to claim 1, characterized in that, The acquisition of multiple video frames to be processed at the target location includes: Acquire multiple videos of different objects at the target location; The plurality of videos to be processed are stream-decoded to obtain the decoding results of the plurality of videos to be processed; According to a preset frame rate, keyframes are extracted from the decoding results of the multiple videos to be processed to obtain the multiple video frames to be processed.

3. The method according to claim 2, characterized in that, The step of streaming decoding the plurality of videos to be processed to obtain the decoding results of the plurality of videos to be processed includes: Obtain the sequence lengths corresponding to the plurality of videos to be processed; According to the maximum sequence length among the multiple videos to be processed, the sequence lengths corresponding to the multiple videos to be processed are aligned to obtain multiple videos to be processed with the maximum sequence length; Stream decoding is performed on multiple videos of the maximum sequence length to generate decoding results for the multiple videos.

4. The method according to claim 1, characterized in that, The step of processing the multiple video frames to be processed using a preset pre-trained model to obtain the first spatial feature vector of the target location includes: Based on the normalization layer in the preset pre-trained model, the multiple video frames to be processed are normalized to obtain the normalized features of the multiple video frames to be processed. Based on the convolutional layers in the preset pre-trained model, the normalized features of the multiple video frames to be processed are convolved to obtain the first spatial feature vector of the target location.

5. The method according to claim 1, characterized in that, The pre-trained bidirectional long short-term memory network processes the first spatial feature vector of the target location to obtain the first prediction stage category of the target location, including: Based on the first long short-term memory network in the pre-trained bidirectional long short-term memory network, the first spatial feature vector of the target location is processed in chronological order to generate the first temporal feature; Based on the second long short-term memory network in the pre-trained bidirectional long short-term memory network, the first spatial feature vector of the target location is processed in chronological order from back to front to generate the second temporal feature. The first time feature and the second time feature are concatenated to generate the first prediction stage category of the target location.

6. The method according to claim 1, characterized in that, The pre-trained bidirectional long short-term memory network processes the first spatial feature vector of the target location to obtain the first prediction stage category of the target location, including: The first spatial feature vector of the target location is divided into a preset number of partial first spatial feature vectors, wherein the preset number is the number of processing modules, and each processing module deploys the bidirectional long short-term memory network. Based on the bidirectional long short-term memory network deployed in each processing module, the partial first spatial feature vector is processed to obtain the category recognition result of each processing module; The category recognition results corresponding to all processing modules are superimposed to obtain the first prediction stage category of the target location.

7. The method according to claim 1, characterized in that, The step of determining the target identification stage of the target location based on the first prediction stage category of the target location includes: The first prediction stage category of the target location is subjected to short segment smoothing and zero-value segment filling to obtain the target recognition stage of the target location.

8. The method according to claim 1, characterized in that, The training method for the bidirectional long short-term memory network includes: Obtain multiple reference video frames at the target location; Using the preset pre-trained model, the multiple reference video frames are processed to obtain the second spatial feature vector of the target location; Based on the initial network, the second spatial feature vector of the target location is processed to obtain the second prediction stage category of the target location; The loss value of the initial network is calculated using the reference stage category of multiple objects at the target location and the second prediction stage category; Based on the loss value, the initial network is iteratively trained until the network reaches a preset cutoff condition after the current training iterations, thus obtaining the bidirectional long short-term memory network.

9. The method according to claim 8, characterized in that, The step of iteratively training the initial network based on the loss value until the network reaches a preset cutoff condition after the current training iterations to obtain the bidirectional long short-term memory network includes: From the reference stage category and the second prediction stage category of the target location, obtain the target reference stage category and the second target stage category whose loss value is greater than a preset loss threshold; Based on the target reference stage category and the second target stage category, the initial network is iteratively trained until the network reaches a preset cutoff condition after the current training iterations, thus obtaining the bidirectional long short-term memory network.

10. The method according to claim 8, characterized in that, The step of iteratively training the initial network based on the loss value until the network reaches a preset cutoff condition after the current training iterations to obtain the bidirectional long short-term memory network includes: Based on the loss value, the parameters are updated and gradients are accumulated for the initial network iteration according to single precision, and backpropagation and forward propagation are performed for the initial network iteration according to half precision until the network reaches the preset cutoff condition for the current training number, thus obtaining the bidirectional long short-term memory network.

11. A stage identification device for a target location, characterized in that, include: The first acquisition module is used to acquire multiple video frames to be processed at the target location; The second acquisition module is used to process the multiple video frames to be processed using a preset pre-trained model to obtain the first spatial feature vector of the target location. The third acquisition module is used to process the first spatial feature vector of the target location based on a pre-trained bidirectional long short-term memory network to obtain the first prediction stage category of the target location. The first determining module is used to determine the target identification stage of the target location based on the first prediction stage category of the target location.

12. An electronic device, characterized in that, include: processor; Memory, used to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the method of any one of claims 1-10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, The storage medium stores a computer program that, when executed by a processor, causes the processor to implement the method described in any one of claims 1-10.