Image selection model training method, target image selection method and device
By performing multi-dimensional reward and punishment calculations on the initial and target images in the image sequence, the image selection model is trained, and the problem of low efficiency and accuracy of keyframe selection in video sequences in the prior art is solved, and more efficient and accurate keyframe selection is achieved.
Patent Information
- Application Number
- CN202211718471.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-12-29
AI Technical Summary
The prior art is less efficient in keyframe selection in video sequences and has less accuracy in identification results, especially when targets move quickly or targets block each other.
By obtaining the initial image and target image of the image sequence, and using multiple reward and punishment dimensions to calculate the training reward and punishment value, the image selection model is trained. This model combines information from multiple dimensions in the image to accurately select keyframes with high definition and rich image information from multiple consecutive images.
The efficiency and accuracy of keyframe selection are improved, so that the trained image selection model can accurately select keyframes with high definition and rich information from the video sequence.
Smart Images

Figure CN116152704B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a training method for an image selection model, a target image selection method and a device. Background Art
[0002] In order to recognize video sequences, a common method is to recognize each frame of the video sequence to obtain the information contained in the video sequence. However, due to factors such as rapid movement of objects or mutual occlusion between objects, some images in a video sequence contain less information. If each frame of the image is recognized, the amount of calculation is large and the recognition result is less accurate.
[0003] Therefore, in order to improve the efficiency of video sequence recognition, key frames can be selected from the video sequence, and the information contained in the corresponding video sequence can be obtained by identifying the key frames. At present, the commonly used key frame selection technology is to measure the quality of images in the video sequence by using reference images, and use images with better quality as key frames. Due to the need to use reference images, this method is less efficient in actual use. Summary of the invention
[0004] The main technical problem solved by the present application is to provide a training method for an image selection model, a target image selection method and a device, which can improve the efficiency of key frame selection.
[0005] To solve the above technical problems, a technical solution adopted in the present application is: to provide a training method for an image selection model, comprising: obtaining an initial image corresponding to an image sequence; inputting the image sequence and the initial image into an image selection model to obtain a target image; wherein the target image is an image in the image sequence that is within a preset step threshold away from the initial image; based on the initial image and the target image, obtaining training reward and punishment values; wherein the training reward and punishment values are related to multiple reward and punishment evaluation dimensions; based on the training reward and punishment values, determining the loss value of the image selection model; using the loss value to adjust the parameters of the image selection model until the preset convergence conditions are met, thereby obtaining the trained image selection model.
[0006] In order to solve the above technical problems, another technical solution adopted in the present application is: to provide a target image selection method, comprising: obtaining a sequence of images to be processed; inputting the sequence of images to be processed into an image selection model to obtain a target image corresponding to the sequence of images to be processed; wherein the image selection model is obtained based on the training method of the image selection model in the above technical solution.
[0007] To solve the above technical problems, another technical solution adopted in the present application is: to provide an electronic device, comprising: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method described in any of the above technical solutions.
[0008] The beneficial effect of the present application is that, different from the prior art, the training method of the image selection model proposed in the present application obtains the initial image and the target image from the image sequence, and uses multiple reward and punishment dimensions to calculate the training reward and punishment values of the initial image and the target image to train the image selection model. This method trains the constructed image selection model by combining information of multiple dimensions in the image, so that the image selection model obtained after training can accurately select key frames with high clarity and rich image information from multiple continuous images. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0010] Figure 1 It is a flowchart of an implementation method of an image selection model of the present application;
[0011] Figure 2 is a flowchart of an implementation method corresponding to step S102;
[0012] Figure 3 is a schematic diagram of an implementation method corresponding to step S1021;
[0013] Figure 4 is a schematic diagram of an implementation method corresponding to step S1021;
[0014] Figure 5 is a flowchart of an implementation method corresponding to step S103;
[0015] Figure 6 It is a flowchart of another implementation method of the training method of the image selection model of the present application;
[0016] Figure 7 It is a structural schematic diagram of an implementation scheme of the image selection model training system of the present application;
[0017] Figure 8 It is a structural diagram of an implementation method of an image selection model training device of the present application. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0019] See also Figure 1 , Figure 1 : is a flow chart of an embodiment of an image selection model of the present application, the method comprising:
[0020] S101: Acquire an initial image corresponding to an image sequence.
[0021] In one embodiment, step S101 includes: obtaining multiple image sequences as training samples, each image sequence including a preset number of images, wherein multiple image sequences including a preset number of images can be obtained by obtaining multiple training videos and segmenting each training video.
[0022] In addition, the preset number of images in the above-mentioned image sequence can be set according to factors such as the frame rate of the training video or the corresponding time interval between images. For example, when the frame rate of the training video is higher, the preset number of images in the image sequence is greater; or, when the time interval between adjacent images in the training video is shorter, the preset number of images in the image sequence is greater.
[0023] Furthermore, for each image sequence, an initial image corresponding to the image sequence is obtained.
[0024] Specifically, in this embodiment, all images in the image sequence are sorted according to the corresponding time information, and the image in the middle of the image sequence is used as the initial image. The frame image is used as the initial image corresponding to the image sequence; when the number of images m contained in the image sequence is an even number, the two frame images located in the middle position of the image sequence are fused, and the image obtained after the image fusion is used as the initial image of the corresponding image sequence.
[0025] Optionally, in other embodiments, when the number m of images included in the image sequence is an even number, the first Frame image or The frame image is used as the initial image of the corresponding image sequence.
[0026] S102: Input the image sequence and the initial image into the image selection model to obtain a target image, wherein the target image is an image in the image sequence that is within a preset step length threshold from the initial image.
[0027] In one embodiment, see Figure 2 , Figure 2 FIG. 1 is a flow chart of an implementation method corresponding to step S102. Step S102 includes:
[0028] S1021: Input the image sequence and the initial image into the image selection model, so that the image selection model determines multiple candidate positions and sets a first confidence score for the candidate positions based on the position of the initial image in the image sequence and a preset step threshold.
[0029] See also Figure 3 , Figure 3 Schematic diagram of an implementation method corresponding to step S1021. Specifically, step S1021 includes: constructing an image selection model. The image selection model includes a backbone network 10 and an output layer 20.
[0030] Specifically, in this embodiment, the backbone network 10 of the image selection model is a residual neural network, including multiple residual blocks; the output layer 20 includes multiple fully connected layers. Optionally, in other embodiments, the backbone network 10 can also be other commonly used neural networks; for example, a convolutional neural network, etc.
[0031] Furthermore, the image sequence and the initial image are input into the backbone network 10, so that the backbone network 10 uses the position of the initial image in the image sequence as the initial position, and moves to at least one side of the initial position at a step interval until a preset step threshold or the first and last positions of the image sequence are reached, thereby obtaining a plurality of candidate positions, wherein the step interval increases from the initial position at a preset ratio.
[0032] Specifically, in this embodiment, the step interval of the above steps satisfies the following formula:
[0033] N=2 i
[0034] Wherein, N represents the interval step size, and i represents the number of moves. That is, the backbone network 10 in the image selection model takes the position of the initial image in the image sequence as the initial position, and determines multiple movement instructions according to the initial position and the number of images contained in the image sequence, and each movement instruction corresponds to a candidate position. The movement instructions include moving forward 1 frame, moving forward 2 frames, moving forward 4 frames, etc., or moving backward 1 frame, moving backward 2 frames, moving backward 4 frames, etc. Wherein, in response to moving forward an arbitrary number of frames to obtain a candidate position, the timestamp corresponding to the image generation at the candidate position is earlier than the timestamp corresponding to the initial image generation; in response to moving backward an arbitrary number of frames to obtain a candidate position, the timestamp corresponding to the image generation at the candidate position is later than the timestamp corresponding to the initial image generation.
[0035] In addition, when moving forward or backward by 1 frame with the initial position as the center, the step interval is 1 frame; when moving forward or backward by 2 frames with the initial position as the center, the step interval is 2 frames. That is, the step interval increases by 2 times each time from the initial position. By increasing the step interval from the initial position according to a preset ratio, when the number of images contained in the image sequence is large, fewer candidate positions are set at a distance from the initial position, which helps to reduce the amount of calculation in the training process and improve the efficiency of model training.
[0036] Optionally, in other implementations, the above-mentioned step intervals may also increase at other ratios, such as 3 or 4.
[0037] In addition, it should be noted that, in practical applications, the above-mentioned step interval is less than or equal to a preset step threshold, wherein the preset step threshold is the distance between the start position or the end position of the image sequence and the initial position.
[0038] Furthermore, all candidate positions are input into the output layer 20 in the image selection model, so that the output layer 20 outputs the first confidence score corresponding to each candidate position. The higher the first confidence score, the higher the probability that the image at the corresponding candidate position is a key frame.
[0039] In one specific embodiment, see Figure 4 , Figure 4 FIG. 1 is a schematic diagram of an implementation method corresponding to step S1021. Figure 4 As shown in Figure a, when the image sequence contains 9 frames, the 5th frame at the middle position is used as the initial image. Further, the image sequence and the 5th frame are input into the image selection model to use the position corresponding to the 5th frame as the initial position and determine the preset step threshold. Based on the fact that the image sequence contains 9 frames and the initial image is the position corresponding to the 5th frame, the preset step threshold is the distance between the initial position and the starting position or the ending position, that is, 4 frames.
[0040] Furthermore, if Figure 4 As shown in the middle figure b, multiple movement instructions are determined with the initial position as the center, that is, moving 1 frame, 2 frames, and 4 frames to both sides of the initial position, and multiple candidate positions are obtained. The obtained candidate positions are the 1st frame, 3rd frame, 4th frame, 6th frame, 7th frame, and 9th frame in the original image sequence. Furthermore, the image selection model outputs the first confidence score corresponding to each movement instruction, that is, the first confidence score corresponding to each candidate position is obtained.
[0041] Optionally, in another embodiment, obtaining multiple candidate positions in step S1021 may also include: moving at least one side of the initial position at a fixed interval until reaching a preset step length or the head and tail positions of the image sequence to obtain multiple candidate positions. The fixed interval may be 2 frames, 3 frames, or 4 frames, etc., which may be set according to actual conditions.
[0042] In yet another embodiment, in response to the initial image itself being possibly a key frame, the multiple candidate positions determined in step S1021 may also include an initial position corresponding to the initial image.
[0043] S1022: Determine a target position from all candidate positions based on the first confidence scores of the candidate positions, and obtain a target image from the image sequence using the target position.
[0044] In one embodiment, step S1022 includes: in response to obtaining multiple candidate positions and corresponding first confidence scores, taking the candidate position corresponding to the first confidence score with the largest value as the target position, and taking the image corresponding to the target position in the image sequence as the target image.
[0045] Specifically, the movement instruction corresponding to the first confidence score with the largest value is used as the target movement instruction. In response to the movement direction corresponding to the target movement instruction being backward, the target position can be calculated by the following formula:
[0046]
[0047] Where F represents the target position, V j m Represents the initial image in the image sequence V m y represents the initial position in the image, y represents the step interval corresponding to the target movement instruction, and m represents the number of images in the image sequence.
[0048] Alternatively, when the moving direction corresponding to the target moving instruction is forward, the target position can be calculated by the following formula:
[0049]
[0050] S103: Based on the initial image and the target image, a training reward and punishment value is obtained. The training reward and punishment system is related to multiple reward and punishment evaluation dimensions.
[0051] See also Figure 5 , Figure 5 FIG. 1 is a flowchart of an implementation method corresponding to step S103. In this implementation method, the above-mentioned multiple reward and punishment evaluation dimensions at least include image category, number of recognized targets and semantic segmentation information, and the above-mentioned step S103 specifically includes:
[0052] S1031: Determine a first score corresponding to the target image based on image category information corresponding to the initial image and the target image.
[0053] In one embodiment, step S1031 includes: obtaining a first similarity between the initial image and different image categories, and taking the first similarity with the largest value as the first category confidence score; and obtaining a second similarity between the target image and different image categories, and taking the second similarity with the largest value as the second category confidence score.
[0054] Specifically, the initial image is classified to obtain the first similarities of the initial image corresponding to multiple image categories, and the largest value is selected from the multiple first similarities as the first category confidence score; similarly, the target image is classified to obtain the second similarities of the target image corresponding to multiple image categories, and the largest value is selected from the multiple second similarities as the second category confidence score. Among them, image classification can be implemented through many open source algorithms, and the specific process is not elaborated in detail here.
[0055] Further, a first score is obtained based on a difference between the second category confidence score and the first category confidence score.
[0056] Specifically, the difference between the second category confidence score and the first category confidence score is taken as the first score.
[0057] Alternatively, in other implementations, the first score may be obtained by calculating using the following formula:
[0058]
[0059] in, represents the first score corresponding to the target image, P n represents the confidence score of the second category, P n-1 represents the confidence score of the first category. sgn represents the step function, that is, when P n -P n-1 When the value of is positive, sgn(P n -P n-1 ) is 1, when P n-P n-1 When the value of is negative, sgn(P n -P n-1 ) has a value of -1. In addition, r0 represents a preset parameter, so that a small change in category confidence can also obtain a corresponding training reward and penalty value, so that the image selection model obtained after training can accurately select the corresponding key frames for multiple image sequences with slight changes in image categories.
[0060] S1032: Determine a second score corresponding to the target image based on the number of recognized targets corresponding to the initial image and the target image.
[0061] In one embodiment, step S1032 includes: performing target recognition on the initial image and the target image to obtain a first number of recognized targets in the initial image and a second number of recognized targets in the target image.
[0062] Specifically, the identified targets include people, animals, plants, etc., and all the identified targets in the initial image and the target image are obtained by performing target recognition. The number of identified targets can be used to indicate the target occlusion in the corresponding image, that is, the more the number of identified targets, the smaller the probability of occlusion between different targets in the corresponding image, and the more information the corresponding image contains.
[0063] Furthermore, the difference between the second number and the first number is used as a second score. By obtaining the second score, it is helpful to obtain a corresponding training reward and penalty value according to the amount of information contained in the image, so that the trained image selection model can select an image with less occlusion from multiple images, thereby improving the accuracy of the image selection model in selecting key frames.
[0064] S1033: Determine a third score corresponding to the target image based on semantic segmentation information corresponding to the initial image and the target image.
[0065] In one embodiment, step S1033 includes: performing semantic segmentation on the initial image and the target image to obtain a first semantic confidence score corresponding to a first pixel in the initial image and a second semantic confidence score corresponding to a second pixel in the target image.
[0066] Specifically, the initial image is segmented into multiple regions by semantic segmentation technology, and a first semantic confidence score corresponding to the region where each first pixel in the initial image is located is obtained. Similarly, a second semantic confidence score corresponding to each second pixel in the target image is obtained by semantic segmentation technology.
[0067] Furthermore, the first semantic confidence score is compared with a predetermined threshold value to obtain a third number of first pixels whose first semantic confidence score is less than the predetermined threshold value; and the second semantic confidence score is compared with the predetermined threshold value to obtain a fourth number of second pixels whose second semantic confidence score is less than the predetermined threshold value. The difference between the fourth number and the third number is taken as the third score. Among them, the predetermined threshold value can be obtained by estimation or by reverse deduction based on multiple test results by relevant researchers. By comparing the semantic confidence score corresponding to the pixel point with the predetermined threshold value, it is helpful to screen out the pixels with lower clarity in the image, and determine the training reward and punishment value corresponding to the target image based on the number of pixels with lower clarity, so that the image selection model obtained after training can select images with higher clarity as key frames from multiple images.
[0068] Optionally, in other embodiments, in response to the fact that blur or segmentation errors are more likely to occur at the junction of different regions in an image with lower definition, the first pixel point may also be a pixel point at the edge of each segmented region in the initial image, and the second pixel point may be a pixel point at the edge of each segmented region in the target image. Compared with the above embodiment, obtaining the third score based on the pixel points at the edge of the segmented region in the initial image and the target image can save computing costs and improve model training efficiency.
[0069] S1034: Obtain training reward and punishment values based on the first score, the second score, and the third score.
[0070] In one embodiment, step S1034 includes: setting a corresponding first weight for the second score, setting a corresponding second weight for the third score, taking the product of the second score and the first weight as the first reference value, taking the product of the third score and the second weight as the second reference value; taking the sum of the first score, the first reference value, and the second reference value as the training reward and punishment value. The specific calculation formula of the training reward and punishment value is as follows:
[0071]
[0072] Among them, R represents the training reward and punishment value, Indicates the first score, represents the second score, α represents the first weight, represents the third score, and β represents the third weight. The first weight and the second weight can be set according to the actual situation. By setting corresponding weight values for the second score and the third score, it is prevented that one or both scores have too great an impact on the training reward and punishment value.
[0073] Optionally, in another embodiment, the training reward or punishment value may be obtained based on any one or two of the first score, the second score and the third score, and the specific process is as above.
[0074] Alternatively, in another embodiment, the process of obtaining the training reward and penalty value may further include: performing instance segmentation on the initial image and the target image, and determining a fourth score according to the number of corresponding instances, and obtaining the training reward and penalty value based on the first score, the second score, the third score, and the fourth score. The fourth score may be obtained through other target perception tasks, such as panoramic segmentation.
[0075] S104: Determine the loss value of the image selection model based on the training reward and penalty value.
[0076] In one embodiment, step S104 includes: using the training reward and penalty values obtained in step S103 to calculate the loss value of the image selection model. The calculation formula of the loss value is as follows:
[0077]
[0078] Where L represents the loss value, B represents the number of image sequences used for training, R represents the training reward and penalty value, s represents the position information of the initial image in the corresponding image sequence, and a represents the target movement instruction. The target position is obtained by executing the target movement instruction on the initial position of the selected model of the input image.
[0079] S105: Using the loss value to adjust the parameters of the image selection model until a preset convergence condition is met, thereby obtaining a trained image selection model.
[0080] In one embodiment, step S105 includes: adjusting parameters in the image selection model using the obtained loss function, and in response to the loss value of the image selection model converging, stopping training and obtaining a trained image selection model.
[0081] The training method of the image selection model proposed in the present application obtains an initial image and a target image from an image sequence, and uses multiple reward and punishment dimensions to calculate the training reward and punishment values of the initial image and the target image to train the image selection model. This method trains the constructed image selection model by combining information of multiple dimensions in the image, so that the image selection model obtained after training can accurately select key frames with high clarity and rich image information from multiple continuous images.
[0082] In yet another embodiment, see Figure 6 , Figure 6 This is a flow chart of another embodiment of the training method of the image selection model of the present application. The method specifically includes:
[0083] S201: Acquire an initial image corresponding to an image sequence.
[0084] In one implementation, step S201 includes: acquiring a plurality of image sequences including a preset number of images. The specific implementation process may refer to the above step S101 and will not be elaborated in detail herein.
[0085] S202: Input the image sequence and the initial image into the image selection model to obtain the target image.
[0086] In one embodiment, step S202 includes: inputting the obtained image sequence and the initial image in the image sequence into the constructed image selection model to obtain the target image. The specific implementation process and the specific structure of the image selection model can refer to step S102, which will not be elaborated in detail here.
[0087] S203: Obtain training reward and penalty values based on the initial image and the target image.
[0088] In one embodiment, step S203 includes: processing the initial image and the target image in combination with multiple reward and punishment evaluation dimensions to obtain the training reward and punishment value corresponding to the target image. The specific implementation process may refer to step S103.
[0089] S204: Detect whether the number of executions corresponding to the step of obtaining training reward and penalty values based on the initial image and the target image reaches a number threshold.
[0090] In one implementation scenario, step S204 includes: in response to the execution times not reaching the times threshold, updating the target image to the initial image, and returning to the step of inputting the image sequence and the initial image into the image selection model to obtain the target image, that is, returning to the above step S202, and sequentially executing steps S202 to S204. The times threshold can be set according to the number of images in the image sequence. When the number of images contained in the image sequence is more, the times threshold is larger. By setting the times threshold to conduct multiple tests on the same image sequence, the accuracy of the obtained training reward and punishment value is higher.
[0091] It should be noted that, in response to updating the target image to the initial image, the initial position corresponding to the updated initial image changes, and the multiple candidate positions obtained based on the updated initial position also change accordingly. For example, in response to the updated initial image being located at the starting position of the image sequence, the multiple candidate positions are moved to the initial position at step intervals until the preset compensation threshold or the end of the image sequence is reached, thereby obtaining multiple candidate positions.
[0092] In another implementation scenario, in response to the above execution times reaching a times threshold, step S205 is executed.
[0093] S205: Determine the loss value of the image selection model based on the training reward and penalty value.
[0094] In one embodiment, the implementation process of step S205 includes: in response to the number of times of executing the step of obtaining training reward and penalty values based on the initial image and the target image reaches a number threshold, that is, obtaining a number of training reward and penalty values corresponding to the number threshold. Based on all the obtained training reward and penalty values, a target reward and penalty value is obtained.
[0095] Specifically, the sum of all training reward and penalty values is taken as the target reward and penalty value. The calculation formula is as follows:
[0096]
[0097] Among them, R′ represents the target reward or punishment value, and K represents the number threshold.
[0098] Furthermore, the loss value of the image selection model is determined based on the target reward and penalty value obtained above. The specific calculation formula is as follows:
[0099]
[0100] Where L represents the loss value, B represents the number of image sequences used for training, and R k ′ represents the target reward and penalty value obtained at the kth iteration, s k represents the position information of the initial image in the corresponding image sequence at the kth iteration, a k represents the target movement instruction output by the image selection model at the kth iteration; π θ (s k ,a k ) indicates that the image selection model with parameter θ corresponds to the position information of the initial image s k When the target movement instruction a is output k probability.
[0101] Furthermore, the obtained loss value is used to adjust the parameters in the image selection model, and in response to the number of training rounds reaching a preset number of rounds, or the loss function of the image selection model converges, the training is stopped and the trained image selection model is obtained.
[0102] In one embodiment, the present application proposes a target image selection method, the method comprising: obtaining a sequence of images to be processed, wherein the video to be processed can be obtained by capturing a video to be processed, and the video to be processed can be divided into a plurality of sequences of images to be processed including a preset number of images.
[0103] Furthermore, the obtained image sequence to be processed is input into the image selection model to obtain a target image corresponding to the image sequence to be processed, and the target image is a key frame with high definition and rich image content in the corresponding image sequence to be processed. The image selection model is obtained based on the training method of the image selection model mentioned in the above corresponding implementation mode.
[0104] See also Figure 7 , Figure 7 The structure diagram of an embodiment of the image selection model training system of the present application is shown in FIG. Specifically, the system includes an acquisition module 30 , a first acquisition module 40 , a second acquisition module 50 , a third acquisition module 60 and a processing module 70 which are coupled to each other.
[0105] Specifically, the acquisition module 30 is used to acquire an initial image corresponding to the image sequence.
[0106] The first acquisition module 40 is used to input the image sequence and the initial image into the image selection model to obtain the target image, wherein the target image is an image in the image sequence that is within a preset step threshold from the initial image.
[0107] The second obtaining module 50 is used to obtain a training reward and punishment value based on the initial image and the target image, wherein the training reward and punishment value is related to a plurality of reward and punishment evaluation dimensions.
[0108] The third obtaining module 60 is used to determine the loss value of the image selection model based on the training reward and penalty value.
[0109] The processing module 70 is used to adjust the parameters of the image selection model using the loss value until the preset bracelet conditions are met to obtain the trained image selection model.
[0110] See also Figure 8 , Figure 8 This is a structural diagram of an embodiment of an image selection model training device of the present application, and the image selection model training device includes a memory 80 and a processor 82 coupled to each other, and program instructions are stored in the processor 82, and the processor 82 is used to execute the program instructions to implement the video switching method in any of the above-mentioned embodiments. Specifically, the processor 82 can also be called a CPU (Central Processing Unit). The processor 82 may be an integrated circuit chip with signal processing capabilities. The processor 82 can also be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field-programmable gate array (Field-Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 82 can be implemented by an integrated circuit chip.
[0111] It should be noted that the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present implementation scheme.
[0112] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each implementation method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0114] The above description is only an implementation method of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A training method for an image selection model, characterized in that: include: Obtain the initial image corresponding to the image sequence; Input the image sequence and the initial image into an image selection model to obtain a target image; wherein the target image is an image in the image sequence that is within a preset step threshold from the initial image; Based on the initial image and the target image, a training reward and punishment value is obtained; wherein the training reward and punishment value is related to a plurality of reward and punishment evaluation dimensions; Determining a loss value of the image selection model based on the training reward and penalty value; Using the loss value to adjust the parameters of the image selection model until a preset convergence condition is met, thereby obtaining the trained image selection model; The multiple reward and punishment evaluation dimensions include at least image category, number of recognized targets and semantic segmentation information; obtaining training reward and punishment values based on the initial image and the target image includes: determining a first score corresponding to the target image based on the image category information corresponding to the initial image and the target image; determining a second score corresponding to the target image based on the number of recognized targets corresponding to the initial image and the target image; determining a third score corresponding to the target image based on the semantic segmentation information corresponding to the initial image and the target image; obtaining the training reward and punishment value based on at least two of the first score, the second score and the third score.
2. The method according to claim 1, characterized in that The step of inputting the image sequence and the initial image into an image selection model to obtain a target image comprises: Inputting the image sequence and the initial image into the image selection model, so that the image selection model determines a plurality of candidate positions and sets a first confidence score for the candidate positions based on the position of the initial image in the image sequence and the preset step size threshold; Based on the first confidence scores of the candidate positions, a target position is determined from all the candidate positions, and the target image is obtained from the image sequence using the target position.
3. The method according to claim 2, characterized in that The image selection model comprises a backbone network and an output layer, and the image sequence and the initial image are input into the image selection model so that the image selection model determines a plurality of candidate positions and sets a first confidence score for the candidate positions based on the position of the initial image in the image sequence and the preset step threshold, comprising: Inputting the image sequence and the initial image into the backbone network, so that the backbone network takes the position of the initial image in the image sequence as the initial position, and moves to at least one side of the initial position at a step-by-step interval until reaching the preset step-by-step threshold or the head and tail positions of the image sequence, thereby obtaining a plurality of candidate positions; wherein the step-by-step interval increases from the initial position at a preset ratio; All the candidate positions are input into the output layer, so that the output layer outputs the first confidence score corresponding to each candidate position.
4. The method according to claim 1, characterized in that: The determining, based on the image category information corresponding to the initial image and the target image, a first score corresponding to the target image comprises: Obtaining first similarities of the initial image corresponding to different image categories, and using the first similarity with the largest value as the first category confidence score; and obtaining second similarities of the target image corresponding to different image categories, and using the second similarity with the largest value as the second category confidence score; The first score is obtained based on a difference between the second category confidence score and the first category confidence score.
5. The method according to claim 1, characterized in that The determining a second score corresponding to the target image based on the number of recognized targets corresponding to the initial image and the target image includes: Performing target recognition on the initial image and the target image to obtain a first number of recognized targets in the initial image and a second number of recognized targets in the target image; The difference between the second number and the first number is taken as the second score.
6. The method according to claim 1, characterized in that The determining, based on semantic segmentation information corresponding to the initial image and the target image, a third score corresponding to the target image comprises: Performing semantic segmentation on the initial image and the target image to obtain a first semantic confidence score corresponding to a first pixel in the initial image and a second semantic confidence score corresponding to a second pixel in the target image; Obtaining a third number of the first pixel points corresponding to the first semantic confidence score being less than a predetermined threshold; and obtaining a fourth number of the second pixel points corresponding to the second semantic confidence score being less than the predetermined threshold; The difference between the fourth number and the third number is used as the third score.
7. The method according to claim 1, characterized in that Before determining the loss value of the image selection model based on the training reward and penalty value, the method includes: Detecting whether the number of executions corresponding to the step of obtaining the training reward and penalty value based on the initial image and the target image reaches a number threshold; In response to the execution times not reaching the times threshold, updating the target image to the initial image, and returning to the step of inputting the image sequence and the initial image into the image selection model to obtain the target image; The step of determining the loss value of the image selection model based on the training reward and penalty value comprises: Based on all the obtained training reward and penalty values, a target reward and penalty value is obtained, and the loss value of the image selection model is determined using the target reward and penalty value.
8. A target image selection method, characterized in that: include: Obtaining a sequence of images to be processed; The image sequence to be processed is input into an image selection model to obtain a target image corresponding to the image sequence to be processed; wherein the image selection model is obtained based on the training method of the image selection model described in any one of claims 1-7.
9. A training device for an image selection model, characterized in that: include: A memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1-7 or 8.
Citation Information
Patent Citations
Video processing method and device, equipment and storage medium
CN111294646A
Video target detection method, device and equipment and storage medium
CN112101114A