Training method and device of image processing model and image processing method and device

By conducting local and global training on the image processing model, and using the timing task branch network to generate the target model, the accuracy and real-time problems of traffic accident recognition in the existing technology are solved, and the accurate distinction and identification of traffic accident types are achieved.

CN120220091APending Publication Date: 2025-06-27BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510330084.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately identify traffic accidents and distinguish different types of traffic accidents, especially in complex traffic scenarios, which are prone to false alarms.

Method used

Through a training method of an image processing model, local training is performed using the first sample video and the timing task branch network to generate a third image processing model, and global training is performed through the second sample video to obtain the target image processing model.

Benefits of technology

Accurate timely identification and type distinction of traffic accidents are achieved, and the flexibility of image processing models and the accuracy and reliability of training processes are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220091A_ABST
    Figure CN120220091A_ABST
Patent Text Reader

Abstract

The invention provides a training method and device of an image processing model and an image processing method and device, and relates to the field of artificial intelligence, in particular to the technical field of image processing, deep learning and intelligent traffic. Comprising the steps of performing local training on a first image processing model according to a first sample video and a time sequence task branch network to obtain a second image processing model; generating a third image processing model according to a downstream task branch network and the second image processing model; and performing global training on the third image processing model according to the second sample video to obtain a target image processing model.According to the method, optimization training is performed on the image processing models in a staged and progressive manner, so that the difficulty of obtaining the target image processing model is effectively reduced, and the efficiency of obtaining the target image processing model is improved. And the accuracy and the reliability in the image processing model training process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, specifically to image processing, deep learning, and intelligent transportation technologies, and particularly to a method and apparatus for training an image processing model, and an image processing method and apparatus. Background Art

[0002] In the context of the increasingly complex traffic conditions in today's society, the frequent occurrence of traffic accidents has caused profound and multi-dimensional impacts and hazards. Rapid identification and location of traffic accidents play a key role in modern traffic management. However, due to the complex and diverse traffic scenarios, scenarios such as congestion and parking are very similar to accidents. In related technologies, traffic accident identification schemes based on rule-based strategies are difficult to partition and are prone to false alarms. Moreover, since the forms of traffic accidents are difficult to enumerate, involving different types of traffic accidents, the number of traffic accident targets, collision methods, severity levels, etc., how to accurately and real-time identify accidents and identify different accident types has become one of the important research directions. Summary of the Invention

[0003] The present disclosure provides a method for training an image processing model, an image processing method, an apparatus for training an image processing model, an image processing apparatus, an electronic device, a storage medium, and a computer program product.

[0004] According to a first aspect of the present disclosure, there is provided a method for training an image processing model, including: locally training a first image processing model according to a first sample video and a temporal task branch network to obtain a second image processing model; generating a third image processing model according to a downstream task branch network and the second image processing model; globally training the third image processing model according to a second sample video to obtain a target image processing model.

[0005] According to a second aspect of the present disclosure, there is provided an image processing method, including: obtaining a video, and inputting a current frame in the video into a target image processing model for processing to obtain a prediction result of the current frame; wherein the target image processing model is a model trained by using the training method described in the first aspect.

[0006] According to a third aspect of the present disclosure, there is provided an apparatus for training an image processing model, including: a local training module for locally training a first image processing model according to a first sample video and a temporal task branch network to obtain a second image processing model; a model generation module for generating a third image processing model according to a downstream task branch network and the second image processing model; a global training module for globally training the third image processing model according to a second sample video to obtain a target image processing model.

[0007] According to a fourth aspect of the present disclosure, there is provided an image processing apparatus, including: a prediction module configured to obtain a video and input a current frame in the video into a target image processing model for processing to obtain a prediction result of the current frame; wherein the target image processing model is a model trained by using the training method described in the first aspect.

[0008] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the training method of the image processing model described in the first aspect of the present disclosure or the image processing method described in the second aspect.

[0009] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the training method of the image processing model described in the first aspect of the present disclosure or the image processing method described in the second aspect.

[0010] According to a seventh aspect of the present disclosure, there is provided a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, it implements the training method of the image processing model described in the first aspect of the present disclosure or the image processing method described in the second aspect.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0013] Figure 1 is a schematic flowchart of a training method of an image processing model according to an embodiment of the present disclosure;

[0014] Figure 2 is a schematic flowchart of a training method of an image processing model according to an embodiment of the present disclosure;

[0015] FIG. 3(a) is a schematic diagram of obtaining temporal features according to an embodiment of the present disclosure;

[0016] FIG. 3(b) is a schematic diagram of obtaining temporal features according to another embodiment of the present disclosure;

[0017] Figure 4 is a schematic diagram of a training process of an image processing model;

[0018] Figure 5 Flow schematic diagram of an image processing method according to an embodiment of the present disclosure;

[0019] Figure 6 Structural schematic diagram of a training device for an image processing model according to an embodiment of the present disclosure;

[0020] Figure 7 Structural schematic diagram of an image processing device according to an embodiment of the present disclosure;

[0021] Figure 8 Schematic block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners

[0022] The following makes an explanation of exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] The following briefly explains the technical fields involved in the solutions of the present disclosure:

[0024] Artificial Intelligence (AI) is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level technologies and software-level technologies. Artificial Intelligence hardware technologies generally include several aspects such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0025] Image Processing technology is a technology for processing image information by a computer. It mainly includes image digitization, image enhancement and restoration, image data encoding, image segmentation, and image recognition.

[0026] Deep learning (DL for short) is a new research direction in the field of machine learning (ML). It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies.

[0027] Intelligent transportation, the intelligent transportation system (ITS for short) is to effectively integrate advanced scientific and technological means (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing, strengthen the connection among vehicles, roads, and users, so as to form an integrated transportation system that ensures safety, improves efficiency, improves the environment, and saves energy.

[0028] The following describes a method for training an image processing model and an image processing method according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0029] Figure 1 It is a schematic flowchart of a method for training an image processing model according to an embodiment of the present disclosure. It should be noted that the execution subject of the method for training the image processing model in this embodiment is a training device for the image processing model. The training device for the image processing model can specifically be a hardware device or software in a hardware device. Among them, the hardware device is, for example, a terminal device, a server, etc.

[0030] As Figure 1 shown, the method for training the image processing model proposed in this embodiment includes the following steps:

[0031] S101. Locally train the first image processing model according to the first sample video and the temporal task branch network to obtain a second image processing model.

[0032] Optionally, video collection can be performed in a normal traffic scenario, for example: video collection is performed in the "normal vehicle passing state" to obtain a first video, and pseudo-labels are marked on the first video to obtain a first sample video.

[0033] It should be noted that due to the powerful reasoning ability of the large model, the large model can be used to label the pseudo-labels of the first video, so as to improve the speed of obtaining the first sample video and reduce the cost of obtaining the first sample video.

[0034] In the embodiment of the present disclosure, the large model can be fine-tuned based on the temporal task branch network i to obtain the first target large model, and the first sample image in the first sample video is input into the first target large model to obtain the first label information.

[0035] For example, when the temporal task branch network i is a temporal vehicle detection task network, the large model is fine-tuned through the temporal vehicle detection task network to obtain the first target large model. Each frame of image, that is, the first sample image, is extracted from the first video, and the first sample image is input into the first target large model to output the detection box (bbox) of the vehicle, that is, the first label information is obtained.

[0036] For example, when the temporal task branch network i is a vehicle position offset prediction task network, the large model is fine-tuned through the vehicle position offset prediction task network to obtain the first target large model. Each frame of image, that is, the first sample image, is extracted from the first video, and the first sample image is input into the first target large model to output the offset (pre offset) of the center position of the detection box of the vehicle in the previous frame relative to the center position of the detection box of the vehicle in the current frame, that is, the first label information is obtained, and the offset (post offset) of the center position of the detection box of the vehicle in the current frame relative to the center position of the detection box of the vehicle in the next frame is output, that is, the first label information is obtained.

[0037] For example, when the temporal task branch network i is a temporal masked image modeling (MIM) task network, the large model is fine-tuned through the temporal masked image modeling MIM task network to obtain the first target large model. Each frame of image, that is, the first sample image, is extracted from the first video, and the first sample image is input into the first target large model to output the masked image, that is, the first label information is obtained.

[0038] In the embodiment of the present disclosure, the first label information can be summarized, that is, the detection box (bbox) of the vehicle, the offset (pre offset), the offset (post offset), and the masked image are used as the final first label information of each frame of image in the first video, and the first label information is labeled on each frame of image in the first video to obtain the first sample video.

[0039] In an embodiment of the present disclosure, after obtaining the first sample video, the first image processing model can be locally trained according to the first sample video and the time-series task branch network to obtain a second image processing model.

[0040] Optionally, the first sample image in the first sample video can be input into the first image processing model. The first image processing model obtains the first sample feature information of the first sample image. According to the first sample feature information and the first time-series feature of the historical second sample image in the first sample video, the second time-series feature of the first sample image is obtained. According to the second time-series feature and the time-series task branch network, the backbone network and the time-series fusion network in the first image processing model are trained to obtain a second image processing model.

[0041] S102. Generate a third image processing model according to the downstream task branch network and the second image processing model.

[0042] In an embodiment of the present disclosure, after obtaining the second image processing model, the downstream task branch network can be spliced with the second image processing model to generate a third image processing model.

[0043] Among them, the number of downstream task branch networks can be single or multiple.

[0044] For example, if there are 3 downstream branch task networks, namely, an accident type prediction network (Accident type), an accident level prediction network (Accident level), and an accident status prediction network (Accident status), the accident type prediction network, the accident level prediction network, and the accident status prediction network can be respectively and juxtaposedly spliced behind the second image processing model to generate a third image processing model.

[0045] It should be noted that in order to improve the flexibility of the second image processing model, the downstream task branch network can be arbitrarily added or deleted in a plug-and-play manner.

[0046] In an embodiment of the present disclosure, a downstream task update request can be received. Based on the downstream task update request, the corresponding downstream task branch network can be added or deleted in the third image processing model.

[0047] For example, if the current downstream branch task network includes a traffic accident type prediction network, a traffic accident level prediction network, and a traffic accident form prediction network, a form prediction network for traffic accident objects (Accident object status) can be added. The traffic accident type prediction network, the traffic accident level prediction network, the traffic accident form prediction network, and the form prediction network for traffic accident objects are respectively concatenated side by side after the second image processing model to generate an updated third image processing model.

[0048] For example, if the current downstream branch task network includes a traffic accident type prediction network, a traffic accident level prediction network, and a traffic accident form prediction network, the traffic accident level prediction network can be deleted. The traffic accident type prediction network and the traffic accident form prediction network are respectively concatenated side by side after the second image processing model to generate an updated third image processing model.

[0049] S103. According to the second sample video, perform global training on the third image processing model to obtain a target image processing model.

[0050] Optionally, video can be collected in a normal traffic scenario and in a traffic accident scenario, for example: video can be collected in the "normal vehicle passing state" and the "abnormal vehicle passing state" to obtain a second video.

[0051] It should be noted that since the large model has powerful reasoning capabilities, the large model can be fine-tuned to achieve the recognition of traffic accident tasks, and the second video can be labeled with pseudo-labels through the fine-tuned large model to improve the speed of obtaining the second sample video and reduce the cost of obtaining the second sample video.

[0052] In the embodiment of the present disclosure, the large model can be fine-tuned and trained based on the downstream task branch network j to obtain a second target large model. The third sample image in the second sample video is input into the second target large model to obtain second label information.

[0053] For example, for the downstream task branch network j being a traffic accident type prediction task network, the large model is fine-tuned and trained through the traffic accident type prediction task network to obtain a second target large model. Each frame of the image, that is, the third sample image, is extracted from the second video, and the third sample image is input into the second target large model to output a traffic accident type detection result, that is, the second label information is obtained.

[0054] For example, for the downstream task branch network j which is a traffic accident level prediction task network, through the traffic accident level prediction task network, the large model is fine-tuned to obtain a second target large model. Each frame of image is extracted from the second video, i.e., the third sample image, and the third sample image is input into the second target large model to output the traffic accident level detection result, i.e., the second label information is obtained.

[0055] For example, for the downstream task branch network j which is a traffic accident form prediction task network, through the traffic accident form prediction task network, the large model is fine-tuned to obtain a second target large model. Each frame of image is extracted from the second video, i.e., the third sample image, and the third sample image is input into the second target large model to output the traffic accident form detection result, i.e., the second label information is obtained.

[0056] For example, for the downstream task branch network j which is a traffic accident target form prediction task network, through the traffic accident target form prediction task network, the large model is fine-tuned to obtain a second target large model. Each frame of image is extracted from the second video, i.e., the third sample image, and the third sample image is input into the second target large model to output the form detection result of the traffic accident target, i.e., the second label information is obtained.

[0057] In the embodiment of the present disclosure, the second label information can be summarized, that is, the traffic accident type detection result, the traffic accident level detection result, the traffic accident form detection result, and the form detection result of the traffic accident target are used as the second label information for each frame of image in the final second video, and the second label information is marked on each frame of image in the second video to obtain a second sample video.

[0058] In the embodiment of the present disclosure, after the second sample video is obtained, the third image processing model can be globally trained according to the second sample video to obtain a target image processing model.

[0059] Optionally, the third sample image in the second sample video can be input into the third image processing model. The third image processing model obtains the second sample feature information of the third sample image, and fuses the second sample feature information with the third temporal feature of the historical fourth sample image in the second sample video to obtain the fourth temporal feature of the third sample image. According to the fourth temporal feature, the third image processing model is adjusted to obtain a target image processing model.

[0060] A training method for an image processing model according to an embodiment of the present disclosure locally trains a first image processing model based on a first sample video and a temporal task branch network to obtain a second image processing model, generates a third image processing model based on a downstream task branch network and the second image processing model, and globally trains the third image processing model based on a second sample video to obtain a target image processing model. Thus, the present disclosure locally trains the first image processing model to obtain the second image processing model, and generates the third image processing model based on the downstream task branch network and the second image processing model, ensuring that the third image processing model has strong flexibility. Then, by globally training the third image processing model and optimizing and training the image processing model in a phased and progressive manner, the difficulty of obtaining the target image processing model is effectively reduced, and the accuracy and reliability in the training process of the image processing model are improved.

[0061] Figure 2 It is a schematic flowchart of a training method for an image processing model according to an embodiment of the present disclosure.

[0062] As Figure 2 shown, the training method for the image processing model proposed in this embodiment includes the following steps:

[0063] In the above-mentioned embodiment, S101, "locally train the first image processing model based on the first sample video and the temporal task branch network to obtain a second image processing model", may specifically include S201 and S203.

[0064] S201: Input the first sample image in the first sample video into the first image processing model, and the first image processing model obtains the first sample feature information of the first sample image.

[0065] Optionally, the first sample image in the first sample video, that is, each frame image in the first sample video, can be input into the first image processing model in sequence for basic feature extraction to obtain the first sample feature information of the first sample image.

[0066] For example, if the current frame in the first sample video is the first sample image It, input the first sample image It into the first image processing model, and the backbone network in the first image processing model extracts features from the first sample image It to obtain the first sample feature information Ft of the first sample image It.

[0067] It should be noted that the present disclosure does not limit the network structure of the backbone network.

[0068] Optionally, the network structure of the backbone network can be a series of Residual Neural Networks (ResNet), such as ResNet-18.

[0069] S202. Based on the first temporal feature of the second sample image in the historical sample video, fuse the first sample feature information and the first temporal feature to obtain the second temporal feature of the first sample image.

[0070] In the embodiments of the present disclosure, for the first sample image, if the image type indicates that the first sample image is a key frame, the second sample image in the history of the first sample video (the historical sample image associated with the first sample image) is at least one key frame adjacent to the first sample image; if the image type indicates that the first sample image is a non-key frame, the second sample image in the history of the first sample video (the historical sample image associated with the first sample image) is the previous frame adjacent to the first sample image.

[0071] Optionally, a preset interval indicating that the sample image is a key frame can be obtained, and according to the preset interval, an image type indication is generated.

[0072] For example, if the current frame in the first sample video is the first sample image It, if the first sample image It is a key frame, and if the preset interval is 5, the previous frame first sample image It-1 and the previous two-frame first sample image It-2 of the first sample image It are non-key frames, the first sample image It-6 that is 5 frames away from the first sample image It is a key frame, and the first sample image It-11 that is 5 frames away from the first sample image It-6 is a key frame.

[0073] In the embodiments of the present disclosure, the first sample feature information and the first temporal feature can be temporally fused to obtain the second temporal feature of the first sample image.

[0074] Optionally, the temporal fusion network in the first image processing model can be used to temporally fuse the first sample feature information and the first temporal feature, realizing the temporal modeling between the first sample feature information of consecutive frames in the first sample video to obtain the second temporal feature of the first sample image.

[0075] For example, if the current frame in the first sample video is the first sample image It, the temporal fusion network in the first image processing model can be used to temporally fuse the first sample feature information Ft and the first temporal feature, realizing the temporal modeling between the first sample feature information of consecutive frames in the first sample video to obtain the second temporal feature Mt of the first sample image, where the second temporal feature Mt can simulate temporal features such as vehicle movement.

[0076] It should be noted that the present disclosure does not limit the temporal fusion network. Optionally, the temporal fusion network can be a Convolutional Gate Recurrent Unit (abbreviated as CGRU).

[0077] The following explains the specific process of performing temporal fusion on the first sample feature information and the first temporal feature to obtain the second temporal feature of the first sample image.

[0078] For example, as shown in Figure 3, if it is determined that the first sample image It is a non-key frame and the second sample image is the previous frame It-1 adjacent to the first sample image, the first temporal feature Mt-1 of the cached second sample image It-1 is obtained, and the Convolutional Gate Recurrent Unit CGRU performs temporal modeling on the first sample feature information Ft and the first temporal feature Mt-1 to obtain the second temporal feature Mt of the first sample image.

[0079] For example, as shown in Figure 3(b), if it is determined that the first sample image It is a key frame and the second sample image is the key frame It-6 adjacent to the first sample image, the first temporal feature Mt-6 of the cached second sample image It-6 is obtained, and the Convolutional Gate Recurrent Unit CGRU performs temporal modeling on the first sample feature information Ft and the first temporal feature Mt-6 of the second sample image It-6 to obtain the second temporal feature Mt of the first sample image, where the first temporal feature Mt-6 is obtained by the Convolutional Gate Recurrent Unit CGRU performing temporal modeling on the first sample feature information Ft-6 and the first temporal feature Mt-11 of the second sample image It-11.

[0080] S203. Train the backbone network and the temporal fusion network in the first image processing model according to the second temporal feature and the temporal task branch network to obtain a second image processing model.

[0081] In the embodiment of the present disclosure, the second sample temporal feature can be input into the temporal task branch network to obtain the first sample prediction result of the temporal task branch network, and the backbone network and the temporal fusion network are adjusted according to the first sample prediction result and the first label information of the first sample image to obtain a second image processing model.

[0082] It should be noted that in order to hierarchically optimize the representation capabilities of the backbone network and the temporal fusion network, a temporal task branch network is specifically constructed. Through the temporal task branch network, it can help the second image processing model have strong representation capabilities and temporal modeling capabilities, making the training process of the third image processing model simpler.

[0083] For example, the temporal task branch network may include a temporal vehicle detection task branch network (bbox), a vehicle position offset prediction task branch network (pre offset), a vehicle position offset prediction task branch network (post offset), and a temporal mask image modeling MIM task network branch.

[0084] For example, the second sample temporal feature can be input into the temporal vehicle detection task branch network to obtain the first sample prediction result of the temporal vehicle detection task branch network (the detection box bbox1 of the predicted vehicle). The second sample temporal feature is input into the vehicle position offset prediction task branch network to obtain the first sample prediction result of the vehicle position offset prediction task branch network (the predicted offset pre offset1). The second sample temporal feature is input into the vehicle position offset prediction task branch network to obtain the first sample prediction result of the vehicle position offset prediction task branch network (the predicted offset post offset 1). The second sample temporal feature is input into the sequential mask image modeling MIM task branch network to obtain the first sample prediction result of the sequential mask image modeling MIM task branch network (the predicted mask image).

[0085] Optionally, the first branch loss of the task branch network i can be obtained according to the first sample prediction result corresponding to the temporal task branch network i and the corresponding first label information i, where i is an integer greater than or equal to 1. According to the first branch loss of each temporal task branch network, the first model loss is determined, and the backbone network and the temporal fusion network are adjusted based on the first model loss to obtain the second image processing model.

[0086] For example, according to the first sample prediction result of the temporal vehicle detection task branch network (the detection box bbox1 of the predicted vehicle) and the first label information (the labeled detection box bbox of the vehicle), the first branch loss of the temporal vehicle detection task branch network is obtained. According to the first sample prediction result of the vehicle position offset prediction task branch network (the predicted offset pre offset1) and the first label information (the labeled offset pre offset), the first branch loss of the vehicle position offset prediction task branch network is obtained. According to the first sample prediction result of the vehicle position offset prediction task branch network (the predicted offset post offset 1) and the first label information (the labeled offset post offset), the first branch loss of the vehicle position offset prediction task branch network is obtained. According to the first sample prediction result of the temporal mask image modeling MIM task branch network (the predicted mask image) and the first label information (the labeled mask image), the first branch loss of the temporal mask image modeling MIM task branch network is obtained.

[0087] In an embodiment of the present disclosure, the first branch losses of each temporal task branch network can be aggregated to obtain a first model loss, and the backbone network and the temporal fusion network are adjusted based on the first model loss to obtain a second image processing model.

[0088] Optionally, the backbone network and the temporal fusion network can be adjusted based on the first model loss, and the adjusted backbone network and temporal fusion network are continuously trained until the training ends to obtain a second image processing model.

[0089] It should be noted that the present disclosure does not limit the setting of the training end condition, which can be set according to the actual situation. Optionally, the training end condition can be set as the first model loss being less than a preset loss threshold.

[0090] S204. Generate a third image processing model according to the downstream task branch network and the second image processing model.

[0091] For the relevant content of step S201, reference can be made to the above embodiments, which will not be elaborated here.

[0092] S103 in the above embodiment, "globally train the third image processing model according to the second sample video to obtain a target image processing model" can specifically include S205 and S207.

[0093] S205. Input the third sample image in the second sample video into the third image processing model, and the third image processing model obtains the second sample feature information of the third sample image.

[0094] For example, if the current frame in the second sample video is the third sample image It’, input the third sample image It’ into the third image processing model, and the backbone network in the third image processing model extracts features from the third sample image It’ to obtain the second sample feature information Ft’ of the third sample image It’.

[0095] It should be noted that the specific process of obtaining the second sample feature information will not be elaborated here, and reference can be made to the above embodiments.

[0096] S206. Based on the third temporal feature of the historical fourth sample image in the second sample video, fuse the second sample feature information and the third temporal feature to obtain the fourth temporal feature of the third sample image.

[0097] In the embodiments of the present disclosure, for the third sample image, if the image type indicates that the third sample image is a key frame, then the historical fourth sample image (the historical sample image associated with the third sample image) in the second sample video is at least one key frame adjacent to the third sample image; if the image type indicates that the third sample image is a non-key frame, then the historical fourth sample image (the historical sample image associated with the first sample image) in the second sample video is the previous frame adjacent to the third sample image.

[0098] Optionally, a preset interval indicating that the sample image is a key frame can be obtained, and an image type indication is generated according to the preset interval.

[0099] For example, for the third sample image It’ in the second sample video, if the preset interval is 5, then the third sample image It’-1 in the second sample video is a non-key frame, and the third sample image It’-5 in the second sample video is a key frame, It-10... are key frames.

[0100] For example, if the current frame in the second sample video is the third sample image It’, if the third sample image It’ is a key frame, and if the preset interval is 5, the previous frame third sample image It’-1 and the previous two-frame third sample image It’-2 of the third sample image It’ are non-key frames, and the first sample image It’-6 that is 5 frames away from the third sample image It’ is a key frame, and the third sample image It’-11 that is 5 frames away from the third sample image It’-6 is a key frame.

[0101] Optionally, the temporal fusion network in the third image processing model can be used to fuse the second sample feature information and the third temporal feature to obtain the fourth temporal feature of the third sample image.

[0102] For example, if the current frame in the second sample video is the third sample image It, the temporal fusion network in the third image processing model can be used to perform temporal fusion on the second sample feature information Ft’ and the third temporal feature, so as to realize the temporal modeling between the second sample feature information of consecutive frames in the second sample video, and obtain the fourth temporal feature Mt’ of the third sample image.

[0103] It should be noted that the specific process of obtaining the fourth temporal feature will not be elaborated here, and reference can be made to the above embodiments.

[0104] S207. Adjust the third image processing model according to the fourth temporal feature to obtain a target image processing model.

[0105] In an embodiment of the present disclosure, based on a detection head and a self-attention mechanism decoder, the fourth temporal feature can be processed to obtain target information input into a downstream task branch network. The target information is input into the downstream task branch network to obtain a second sample prediction result of a third sample image under the downstream task branch network. According to the second sample prediction result and the second label information of the third sample image, the third image processing model is adjusted to obtain a target image processing model.

[0106] It should be noted that since traffic accident prediction depends on the feature associations between multiple vehicles and between vehicles and the environment, a self-attention mechanism Transformer encoder can be used as a basic module to extract the inter-object dependency relationships of the fourth temporal feature.

[0107] For example, inputting the target information into the downstream task branch network (Accident type) can obtain traffic accident type prediction results, such as vehicle-vehicle accident, single-vehicle accident, vehicle-pedestrian accident, vehicle-non-motor vehicle accident, etc.; inputting the target information into the downstream task branch network (Accident level) can obtain the traffic accident prediction level, where the traffic accident prediction level can be used to determine the severity of the traffic accident; inputting the target information into the downstream task branch network (Accident status) can obtain the traffic accident prediction form; inputting the target information into the downstream task branch network (Accident object status) can obtain the prediction form of the traffic accident target.

[0108] Optionally, according to the second sample prediction result of the downstream task branch network j and the corresponding second label information j, the second branch loss of the downstream task branch network j can be obtained, where j is an integer greater than or equal to 1. According to the second branch loss of each downstream task branch network, the second model loss is determined, and the third image processing model is adjusted based on the second model loss to obtain a target image processing model.

[0109] For example, according to the second sample prediction result (traffic accident type prediction result) of the downstream task branch network (Accident type) and the second label information (annotated traffic accident type detection result), obtain the second branch loss of the downstream task branch network (Accident type). According to the second sample prediction result (traffic accident prediction level) of the downstream task branch network (Accident level) and the second label information (annotated traffic accident level detection result), obtain the second branch loss of the downstream task branch network (Accident level). According to the second sample prediction result (traffic accident prediction form) of the downstream task branch network (Accidentstatus) and the second label information (annotated traffic accident form detection result), obtain the second branch loss of the downstream task branch network (Accident status). According to the second sample prediction result (predicted form of the traffic accident target) of the downstream task branch network (Accident object status) and the second label information (detected form of the traffic accident target), obtain the second branch loss of the downstream task branch network (Accidentobject status).

[0110] In the embodiments of the present disclosure, the second branch losses of each downstream task branch network can be aggregated to obtain a second model loss, and the third image processing model can be adjusted based on the second model loss to obtain a target image processing model.

[0111] Optionally, the third image processing model can be adjusted based on the second model loss, and the adjusted third image processing model can be continuously trained until the training ends to obtain a target image processing model.

[0112] It should be noted that the present disclosure does not limit the setting of the training end condition, which can be set according to the actual situation. Optionally, the training end condition can be set as the second model loss being less than a preset loss threshold.

[0113] The following explains the specific process of the training method of the image processing model proposed in the present disclosure.

[0114] For example, as Figure 4As shown, for the first sample image It, the first sample image It is input into the backbone network in the first image processing model to extract features of the first sample image It, and the first sample feature information Ft of the first sample image It is obtained. The sample feature information Ft-m... Ft-n of the cached historical second sample images and the second temporal features Mt-m... Mt-n are obtained. Among them, if the first sample image It is a key frame, the second sample image is at least one key frame adjacent to the first sample image. If the first sample image It is a non-key frame, the second sample image is the previous frame adjacent to the first sample image. The temporal fusion network in the first image processing model performs temporal fusion on the first sample feature information and the first temporal feature to obtain the second temporal feature Mt of the first sample image. The second sample temporal feature Mt is input into the temporal task branch network. The temporal task branch network may include a temporal vehicle detection task branch network (bbox), a vehicle position offset prediction task branch network (pre offset), a vehicle position offset prediction task branch network (post offset), and a temporal mask image modeling MIM task network branch to obtain the first sample prediction result of the temporal task branch network. According to the first sample prediction result and the first label information of the first sample image, the backbone network and the temporal fusion network are adjusted to complete the local training of the first image processing model, and the second image processing model is obtained. The downstream task branch network is spliced with the second image processing model to generate the third image processing model. The downstream task branch network includes, but is not limited to, a traffic accident type prediction network (Accident type), a traffic accident level prediction network (Accident level), a traffic accident status prediction network (Accident status), and a traffic accident object status prediction network (Accident object status). The second sample feature information and the fourth temporal feature of the third sample image in the second sample video are obtained. The specific process is not elaborated here and is not shown in the figure. Based on the detection head head and the self-attention mechanism decoder decoder, the fourth temporal feature is processed to obtain the target information input into the downstream task branch network. The target information is input into the downstream task branch network to obtain the second sample prediction result of the third sample image under the downstream task branch network. According to the second sample prediction result and the second label information of the third sample image, the third image processing model is adjusted to obtain the target image processing model.

[0115] A method for training an image processing model according to an embodiment of the present disclosure inputs a first sample image in a first sample video into a first image processing model. The first image processing model obtains first sample feature information of the first sample image. According to the first sample feature information and the first temporal feature of a historical second sample image in the first sample video, second temporal feature of the first sample image is obtained. According to the second temporal feature and a temporal task branch network, the backbone network and the temporal fusion network in the first image processing model are trained to obtain a second image processing model. According to a downstream task branch network and the second image processing model, a third image processing model is generated. A third sample image in a second sample video is input into the third image processing model. The third image processing model obtains second sample feature information of the third sample image. The second sample feature information is fused with the third temporal feature of a historical fourth sample image in the second sample video to obtain a fourth temporal feature of the third sample image. According to the fourth temporal feature, the third image processing model is adjusted to obtain a target image processing model. Thus, the present disclosure does not require calibration parameters, reduces the deployment pressure of the target image processing model, can use videos in accident scenarios as training data for iterative training, realizes the closed-loop of training data, and through a hierarchical decoupling design, can optimize the image processing model in a phased and progressive manner, can arbitrarily add or delete downstream task branch networks in a plug-and-play manner, realizes the flexible setting of the target image processing model, and by obtaining an end-to-end target image processing model, greatly reduces the development cost and inference cost of the target image processing model, significantly reduces the time consumption of the target image processing model, and improves the effect of the target image processing model.

[0116] Figure 5 FIG. 4 is a schematic flowchart of an image processing method according to an embodiment of the present disclosure. It should be noted that the execution subject of the image processing method in this embodiment is an image processing device, and the image processing device can specifically be a hardware device or software in a hardware device, etc. The hardware device is, for example, a terminal device, a server, etc.

[0117] As Figure 5 shown, the image processing method proposed in this embodiment includes the following steps:

[0118] S501. Obtain a video, and input the current frame in the video into a target image processing model for processing to obtain a prediction result of the current frame.

[0119] The target image processing model is a model obtained by using the training method in the first aspect.

[0120] It should be noted that the present disclosure does not limit the specific manner of obtaining the video. Optionally, the video can be obtained by real-time acquisition through a video acquisition device.

[0121] In an embodiment of the present disclosure, after obtaining a video, feature extraction can be performed on the current frame based on the backbone network in the target image processing model to obtain the feature information of the current frame. Based on the temporal fusion network in the target image processing model, the feature information of the previous frame and the first temporal feature of the historical frames are fused to obtain the second temporal feature of the current frame. According to the second temporal feature of the current frame, the target information of the corresponding branch network of the target downstream task is obtained, and based on the target information, the prediction result corresponding to the current frame under the target downstream task is determined.

[0122] For example, for the current frame Xt in the video, the current frame Xt is input into the target image processing model, and the backbone network in the target image processing model performs feature extraction on the current frame Xt to obtain the feature information Xt of the current frame Xt.

[0123] In an embodiment of the present disclosure, if the current frame is a key frame, the historical frame is at least one key frame adjacent to the current frame; if the current frame is a non - key frame, the historical frame is the previous frame adjacent to the current frame.

[0124] It should be noted that the specific process of determining whether the current frame is a key frame will not be elaborated here, and reference can be made to the above - mentioned embodiments.

[0125] For example, the temporal fusion network in the target image processing model can fuse the feature information Xt of the current frame Xt and the first temporal feature to obtain the second temporal feature Yt of the current frame Xt.

[0126] In an embodiment of the present disclosure, after obtaining the second temporal feature, the second temporal feature can be processed based on the detection head and the self - attention mechanism decoder to obtain the target information of the corresponding branch network of the target downstream task. The target information is input into the target downstream task branch network to obtain the prediction result of the current frame under the target downstream task branch network.

[0127] It should be noted that the corresponding branch network of the target downstream task includes but is not limited to: the target downstream task branch network (Accident type), the target downstream task branch network (Accident level), the target downstream task branch network (Accident status), and the target downstream task branch network (Accident object status).

[0128] For example, for the current frame Xt, the target information can be input into the downstream task branch network (Accident type) to obtain the prediction result of the traffic accident type of the current frame Xt, input into the downstream task branch network (Accident level) to obtain the predicted level of the traffic accident of the current frame Xt, input into the downstream task branch network (Accident status) to obtain the predicted form of the traffic accident of the current frame Xt, and input into the downstream task branch network (Accident object status) to obtain the predicted form of the traffic accident target of the current frame Xt.

[0129] It should be noted that in the related art, it is usually a step-by-step prediction scheme of obstacle detection + tracking + rule judgment. The input is the real-time streaming video data of the monitoring camera. First, the video is framed into images at a stable frame rate, and the images are input into a general obstacle detection model to obtain the categories and coordinates of all obstacles on the road. The targets include vehicles, pedestrians, cyclists, cones, etc. The target detection results will be input into the tracking module to complete the tracking identification matching and obtain the coordinate positions of each target in different frames. After obtaining the information such as the coordinates of the target in consecutive frames, the motion state of each target in the video can be obtained. Finally, the motion state of each target is analyzed by means of rule judgment. When the vehicle target does not move for a long time and a pedestrian target appears around the vehicle target for a period of time, it is considered that an accident has occurred to the vehicle target. In order to prevent false alarms of roadside parking as accidents, manual calibration parameters are also added to avoid the roadside area. However, the above scheme is based on rule prediction, and the accuracy seriously depends on the effects of detection and tracking. There are disadvantages such as too long optimization chain and difficult positioning problems. In addition, the method based on rule judgment requires prior knowledge of accident behaviors by humans. However, due to the variety of accident situations, it is difficult to enumerate all the rules, resulting in obvious limitations in the above scheme, such as: long algorithm iteration period, difficult positioning problems; difficult to enumerate all the strategy rules, limited accident recognition effect, and poor real-time performance, etc.

[0130] An image processing method according to an embodiment of the present disclosure obtains a video, extracts features of the current frame based on the backbone network in the target image processing model to obtain the feature information of the current frame, and fuses and processes the feature information of the current frame and the first temporal feature of the historical frame based on the temporal fusion network in the target image processing model to obtain the second temporal feature of the current frame. According to the second temporal feature of the current frame, the target information of the corresponding branch network of the target downstream task is obtained, and according to the target information, the prediction result corresponding to the current frame under the target downstream task is determined. Thus, the present disclosure can determine the prediction result corresponding to the current frame under the target downstream task through the target image processing model, improving the diversity and accuracy of obtaining the prediction result corresponding to the current frame, and can be applied to various scenarios. For example, in a traffic accident scenario, rules do not need to be defined, and end-to-end prediction of accidents can be achieved, improving the prediction speed of accidents. Through different target downstream tasks, various prediction results such as traffic accident type prediction results, traffic accident prediction levels, traffic accident prediction forms, and prediction forms of traffic accident targets can be obtained.

[0131] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0132] Corresponding to the training method of the image processing model provided in the above several embodiments, an embodiment of the present disclosure further provides a training device for an image processing model. Since the training device for the image processing model provided in the embodiment of the present disclosure corresponds to the training method of the image processing model provided in the above several embodiments, the implementation manners of the training method of the image processing model are also applicable to the training device for the image processing model provided in this embodiment and will not be described in detail in this embodiment.

[0133] Figure 6 It is a schematic structural diagram of a training device for an image processing model according to an embodiment of the present disclosure.

[0134] As Figure 6 shown, the training device 600 for the image processing model includes: a local training module 610, a model generation module 620, and a global training module 630.

[0135] The local training module 610 is configured to locally train the first image processing model according to the first sample video and the temporal task branch network to obtain the second image processing model;

[0136] The model generation module 620 is configured to generate a third image processing model according to the downstream task branch network and the second image processing model;

[0137] The global training module 630 is used to globally train the third image processing model according to the second sample video to obtain a target image processing model.

[0138] Among them, the local training module 610 is further used to: input the first sample image in the first sample video into the first image processing model, and the first image processing model obtains the first sample feature information of the first sample image; based on the first time series feature of the historical second sample image in the sample video, fuse the first sample feature information and the first time series feature to obtain the second time series feature of the first sample image; according to the second time series feature and the time series task branch network, train the backbone network and the time series fusion network in the first image processing model to obtain the second image processing model.

[0139] Among them, the local training module 610 is further used to: input the second sample time series feature into the time series task branch network to obtain the first sample prediction result of the time series task branch network; according to the first sample prediction result and the first label information of the first sample image, adjust the backbone network and the time series fusion network to obtain the second image processing model.

[0140] Among them, the local training module 610 is further used to: obtain the first branch loss of the implementation task branch i according to the first sample prediction result corresponding to the time series task branch i and the corresponding first label information i, where i is an integer greater than or equal to 1; determine the first model loss according to the first branch loss of each time series task branch network; based on the first model loss, adjust the backbone network and the time series fusion network to obtain the second image processing model.

[0141] Among them, the global training module 630 is further used to: input the third sample image in the second sample video into the third image processing model, and the third image processing model obtains the second sample feature information of the third sample image; based on the third time series feature of the historical fourth sample image in the second sample video, fuse the second sample feature information and the third time series feature to obtain the fourth time series feature of the third sample image; according to the fourth time series feature, adjust the third image processing model to obtain a target image processing model.

[0142] Among them, the global training module 630 is further configured to: process the fourth temporal feature based on the detection head and the self-attention mechanism decoder to obtain the target information input into the downstream task branch network; input the target information into the downstream task branch network to obtain the second sample prediction result of the third sample image under the downstream task branch network; and adjust the third image processing model according to the second sample prediction result and the second label information of the third sample image to obtain the target image processing model.

[0143] Among them, the global training module 630 is further configured to: obtain the second branch loss of the implementation task branch j according to the sample prediction result of the downstream task branch j and the corresponding first label information j, where j is an integer greater than or equal to 1; determine the second model loss according to the second branch loss of each downstream task branch network; and adjust the third image processing model based on the second model loss to obtain the target image processing model.

[0144] Among them, before globally training the third image processing model according to the second sample video to obtain the target image processing model, the apparatus 600 is further configured to: receive a downstream task update request; and add or delete a corresponding downstream task branch network in the third image processing model based on the downstream task update request.

[0145] Among them, for any one of the first sample image and the third sample image, if the image type indicates that the any one of the sample images is a key frame, the historical sample image associated with the any one of the sample images is at least one key frame adjacent to the any one of the sample images; if the image type indicates that the any one of the sample images is not a key frame, the historical sample image associated with the any one of the sample images is the previous frame adjacent to the any one of the sample images.

[0146] Among them, the apparatus 600 is further configured to: fine-tune and train the large model based on the temporal task branch network i to obtain the first target large model; input the first sample image in the first sample video into the first target large model to obtain the first label information; fine-tune and train the large model based on the downstream task branch network j to obtain the second target large model; and input the third sample image in the second sample video into the second target large model to obtain the second label information.

[0147] A training device for an image processing model according to an embodiment of the present disclosure locally trains a first image processing model based on a first sample video and a temporal task branch network to obtain a second image processing model, generates a third image processing model based on a downstream task branch network and the second image processing model, and globally trains the third image processing model based on a second sample video to obtain a target image processing model. Thus, the present disclosure locally trains the first image processing model to obtain the second image processing model, and generates the third image processing model based on the downstream task branch network and the second image processing model, ensuring that the third image processing model has strong flexibility. Then, the third image processing model is globally trained, and the image processing model is optimized and trained in a staged and progressive manner, effectively reducing the difficulty of obtaining the target image processing model and improving the accuracy and reliability in the training process of the image processing model.

[0148] Corresponding to the image processing methods provided in the above several embodiments, an embodiment of the present disclosure further provides an image processing device. Since the image processing device provided in the embodiment of the present disclosure corresponds to the image processing methods provided in the above several embodiments, the implementation manners of the image processing methods are also applicable to the image processing device provided in this embodiment and will not be described in detail in this embodiment.

[0149] Figure 7 It is a schematic structural diagram of an image processing device according to an embodiment of the present disclosure.

[0150] As Figure 7 shown, the image processing device 700 includes: a prediction module 710. Among them:

[0151] The prediction module 710 is configured to obtain a video and input the current frame in the video into the target image processing model for processing to obtain a prediction result of the current frame; wherein, the target image processing model is a model trained by the training method as described in the first aspect.

[0152] Among them, the prediction module 710 is further configured to: extract features of the current frame based on the backbone network in the target image processing model to obtain feature information of the current frame; perform a fusion process on the feature information of the current frame and the first temporal feature of the historical frame based on the temporal fusion network in the target image processing model to obtain a second temporal feature of the current frame; obtain target information of the corresponding branch network of the target downstream task according to the second temporal feature of the current frame; and determine a prediction result corresponding to the current frame under the target downstream task according to the target information.

[0153] Wherein, if the current frame is a key frame, the historical frame is at least one key frame adjacent to the current frame; if the current frame is not a key frame, the historical frame is the previous frame adjacent to the current frame.

[0154] According to the image processing apparatus of an embodiment of the present disclosure, a video is acquired, feature extraction is performed on a current frame based on a backbone network in a target image processing model to obtain feature information of the current frame, fusion processing is performed on the feature information of the current frame and first temporal features of a historical frame based on a temporal fusion network in the target image processing model to obtain second temporal features of the current frame, target information of a corresponding branch network of a target downstream task is obtained according to the second temporal features of the current frame, and a prediction result corresponding to the current frame in the target downstream task is determined according to the target information. Thus, according to the present disclosure, through the target image processing model, the prediction result corresponding to the current frame in the target downstream task can be determined, the diversity and accuracy of obtaining the prediction result corresponding to the current frame are improved, and it can be applied to various scenarios. For example, in a traffic accident scenario, rules do not need to be defined, and end-to-end prediction of the accident can be realized, improving the prediction speed of the accident. Through different target downstream tasks, various prediction results such as traffic accident type prediction results, traffic accident prediction levels, traffic accident prediction forms, and prediction forms of traffic accident targets can be obtained.

[0155] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0156] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0157] As Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 802 or computer programs loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0158] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0159] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the training method of an image processing model or an image processing method. For example, in some embodiments, the training of an image processing model or an image processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the training of the image processing model or the image processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the training method of the image processing model or the image processing method in any other appropriate manner (e.g., by means of firmware).

[0160] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0162] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0163] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0164] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.

[0165] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.

[0166] The present disclosure also provides a computer program product including a computer program which, when executed by a processor, implements the training method or the image processing method of the image processing model as described above.

[0167] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is imposed herein.

[0168] The above specific embodiments do not constitute a limitation to the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A training method for an image processing model, wherein: The method comprises: According to the first sample video and the temporal task branch network, the first image processing model is locally trained to obtain a second image processing model; Generate a third image processing model according to the downstream task branch network and the second image processing model; According to the second sample video, the third image processing model is globally trained to obtain a target image processing model.

2. The method according to claim 1, wherein: The method of locally training the first image processing model according to the first sample video and the temporal task branch network to obtain the second image processing model includes: Inputting a first sample image in a first sample video into a first image processing model, and obtaining first sample feature information of the first sample image by the first image processing model; Based on the first time series feature of the second sample image in the history of the sample video, the first sample feature information and the first time series feature are fused to obtain the second time series feature of the first sample image; According to the second time series feature and the time series task branch network, the backbone network and the time series fusion network in the first image processing model are trained to obtain the second image processing model.

3. The method according to claim 2, wherein: The step of training the backbone network and the time series fusion network in the first image processing model according to the second time series feature and the time series task branch network to obtain the second image processing model includes: Inputting the second time series feature into the time series task branch network to obtain a first sample prediction result of the time series task branch network; According to the first sample prediction result and the first label information of the first sample image, the backbone network and the time series fusion network are adjusted to obtain the second image processing model.

4. The method according to claim 2, wherein: The step of adjusting the backbone network and the time series fusion network according to the first sample prediction result and the first label information of the first sample image to obtain the second image processing model includes: According to the first sample prediction result corresponding to the time series task branch network i and the corresponding first label information i, obtaining the first branch loss of the time series task branch network i, where i is an integer value greater than or equal to 1; Determine a first model loss according to a first branch loss of each of the sequential task branch networks; The backbone network and the temporal fusion network are adjusted based on the first model loss to obtain the second image processing model.

5. The method according to any one of claims 1 to 4, wherein: The third image processing model is globally trained according to the second sample video to obtain a target image processing model: Inputting the third sample image in the second sample video into the third image processing model, and obtaining second sample feature information of the third sample image by the third image processing model; Based on the third time series feature of the fourth sample image in the history of the second sample video, the second sample feature information and the third time series feature are fused to obtain the fourth time series feature of the third sample image; According to the fourth timing feature, the third image processing model is adjusted to obtain a target image processing model.

6. The method according to claim 5, wherein: The step of adjusting the third image processing model according to the fourth time series feature to obtain a target image processing model includes: Based on the detection head and the self-attention mechanism decoder, the fourth temporal feature is processed to obtain target information input into the downstream task branch network; Inputting the target information into the downstream task branch network to obtain a second sample prediction result of the third sample image under the downstream task branch network; According to the second sample prediction result and the second label information of the third sample image, the third image processing model is adjusted to obtain the target image processing model.

7. The method according to claim 6, wherein: The adjusting the third image processing model according to the second sample prediction result and the second label information of the third sample image to obtain the target image processing model includes: According to the second sample prediction result of the downstream task branch network j and the corresponding second label information j, obtaining the second branch loss of the downstream task branch network j, where j is an integer value greater than or equal to 1; Determining a second model loss according to a second branch loss of each of the downstream task branch networks; The third image processing model is adjusted based on the second model loss to obtain the target image processing model.

8. The method according to claim 5, wherein: Before globally training the third image processing model according to the second sample video to obtain a target image processing model, the method further includes: Receive downstream task update requests; Based on the downstream task update request, a corresponding downstream task branch network is added or deleted in the third image processing model.

9. The method according to claim 5, wherein: For any sample image of the first sample image and the third sample image, if the any sample image is a key frame, the historical sample image associated with the any sample image is at least one key frame adjacent to the any sample image; If the image type indicates that any of the sample images is a non-key frame, the historical sample image associated with any of the sample images is a previous frame adjacent to the any of the sample images.

10. The method according to claim 4 or 7, wherein: The method further comprises: Based on the temporal task branch network i, fine-tune the large model to obtain a first target large model; Inputting a first sample image in a first sample video into a first target large model to obtain first label information; Based on the downstream task branch network j, the large model is fine-tuned and trained to obtain a second target large model; the third sample image in the second sample video is input into the second target large model to obtain second label information.

11. An image processing method, wherein: The method comprises: Acquire a video, and input a current frame in the video into a target image processing model for processing to obtain a prediction result of the current frame; The target image processing model is a model obtained by using the training method described in any one of claims 1-10.

12. The method according to claim 11, wherein: The step of inputting the current frame in the video into the target image processing model for processing to obtain a prediction result of the current frame includes: Based on the backbone network in the target image processing model, feature extraction is performed on the current frame to obtain feature information of the current frame; Based on the time series fusion network in the target image processing model, the feature information of the current frame and the first time series features of the historical frames are fused to obtain the second time series features of the current frame; Acquire target information of a branch network corresponding to a target downstream task according to a second temporal feature of the current frame; Determine, according to the target information, a prediction result corresponding to the current frame under the target downstream task.

13. The method according to claim 12, wherein: If the current frame is a key frame, the historical frame is at least one key frame adjacent to the current frame; If the current frame is not a key frame, the historical frame is a previous frame adjacent to the current frame.

14. A training device for an image processing model, wherein: The device comprises: A local training module, used for performing local training on the first image processing model according to the first sample video and the temporal task branch network to obtain a second image processing model; A model generation module, used to generate a third image processing model according to the downstream task branch network and the second image processing model; The global training module is used to perform global training on the third image processing model according to the second sample video to obtain a target image processing model.

15. An image processing device, wherein: The device comprises: A prediction module is used to obtain a video and input a current frame in the video into a target image processing model for processing to obtain a prediction result of the current frame; The target image processing model is a model obtained by using the training method described in any one of claims 1 to 9.

16. An electronic device, characterized in that: including a processor and a memory; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to implement the method according to any one of claims 1-10 or claims 11-13.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 10 or claims 11 to 13 is implemented.

18. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1-10 or claims 11-13.