Trusted artificial intelligence method based on continuous time difference shunt field

By correlating the vehicle driving video frames in chronological order and weighting summing the pooled feature map, the problem of low accuracy of convolutional neural network information recognition is solved, and the accuracy and robustness of the recognition results are improved.

CN120107932AActive Publication Date: 2025-06-06XIAMEN UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510257350.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

In the process of intelligent vehicle recognition, the accuracy of convolutional neural network information recognition is low, resulting in a decrease in recognition accuracy and robustness.

Method used

By correlating multiple video frames in the vehicle driving video in chronological order, and weighted summing of the pooled feature map output from the pooled layer in the convolutional neural network, feature transmission is enhanced and the robustness and stability of the pooled feature map are improved.

Benefits of technology

The accuracy of information recognition results is improved, making the pooled feature map of the video frame at the current time after the weighted sum is more robust and stable, and the recognition accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107932A_ABST
    Figure CN120107932A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and discloses a credible artificial intelligence method based on a continuous time difference shunting field, and the method comprises the steps: obtaining vehicle driving videos of a plurality of continuous moments including a current moment; splitting the vehicle driving video into a plurality of video frames carrying time information; inputting the plurality of video frames carrying the time information into a pre-trained time flow convolutional neural network according to a time sequence, and executing the following steps through the time flow convolutional neural network: sequentially extracting pooling feature maps of the plurality of video frames carrying the time information according to the time sequence; performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map; and determining an information identification result of the video frame at the current moment based on the third pooling feature map, and outputting the information identification result. The plurality of video frames in the vehicle driving video are associated through the time sequence, so that the reliability and the stability of the pooling feature map are enhanced, and the recognition result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a credible artificial intelligence method based on continuous time differential flow field. Background Art

[0002] Convolutional neural networks have been widely used in the field of deep learning, especially in image and vision tasks. Because convolutional neural networks significantly reduce computational complexity and improve feature extraction capabilities through parameter sharing and local receptive field design, they are particularly suitable for processing high-dimensional data such as images and speech. They are gradually being used in vehicle intelligent recognition, for example, identifying obstacles on the road while the vehicle is driving.

[0003] However, information loss is unavoidable during the convolution process, such as incomplete feature expression, reduced model generalization performance, loss of small targets and details, reduced recognition accuracy and reduced robustness, which leads to reduced recognition accuracy. Therefore, a trusted artificial intelligence method is very necessary. Summary of the invention

[0004] In order to solve the problem in the prior art that, in the process of intelligent vehicle identification, the accuracy of information recognition of convolutional neural networks is low, the present invention associates multiple video frames in a vehicle driving video through a time series, and weighted sums the pooled feature maps output by the pooling layer in the convolutional neural network to achieve the purpose of feature transfer, so that the pooled feature maps of the video frames at the current moment after the weighted summation are more robust and stable, and the information recognition results obtained based on the pooled feature maps at the current moment after the weighted summation are more accurate.

[0005] In a first aspect, an embodiment of the present invention provides a trusted artificial intelligence method based on a continuous time differential flow field, the method comprising:

[0006] During the driving of the vehicle, obtaining a vehicle driving video at a plurality of consecutive moments including the current moment, wherein the vehicle driving video is a video captured by a camera of the vehicle during the driving of the vehicle;

[0007] Splitting the vehicle driving video into multiple video frames carrying time information according to the time sequence of multiple video frames in the vehicle driving video;

[0008] The plurality of video frames carrying time information are input into a pre-trained time stream convolutional neural network in time order, so that the time stream convolutional neural network performs the following steps:

[0009] Extracting the pooled feature maps of the multiple video frames carrying time information in sequence according to time order;

[0010] Performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map; wherein the first pooling feature map is a pooling feature map of the video frame at the current moment, the plurality of second pooling feature maps are video frames at a plurality of consecutive moments before the video frame at the current moment, and the third pooling feature map is a pooling feature map of the video frame at the current moment after weighted summation;

[0011] Based on the third pooling feature map, determine the information recognition result of the video frame at the current moment, and output the information recognition result.

[0012] Optionally, if the rate of change of the vehicle driving data in the vehicle driving video is greater than a preset rate;

[0013] The step of performing weighted summation of the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map includes:

[0014] Determine target video frames that are located at a target number of consecutive moments before the current moment video frame, where the target number is less than a preset number;

[0015] Determine the pooled feature maps of the target number of target video frames as a plurality of second pooled feature maps;

[0016] The first pooling feature map and the determined multiple second pooling feature maps are weightedly summed to obtain a third pooling feature map.

[0017] Optionally, the process of determining weighted coefficients of the first pooling feature map and the plurality of second pooling feature maps includes:

[0018] Determining a first weighting coefficient of the first pooling feature map;

[0019] Respectively determine the time intervals between the video frames corresponding to the multiple second pooling feature maps and the video frame at the current moment;

[0020] Determine that the second weighting coefficients of the multiple second pooling feature maps are all smaller than the first weighting coefficient, and the second weighting coefficient is inversely proportional to the time interval.

[0021] Optionally, if the rate of change of the vehicle driving data in the vehicle driving video is less than a preset rate;

[0022] The step of performing weighted summation of the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map includes:

[0023] Determine the pooled feature maps of video frames at all times except the current time in the vehicle driving video as the second pooled feature map;

[0024] Perform a weighted summation on the first pooling feature map and the determined multiple second pooling feature maps.

[0025] Optionally, the process of determining weighted coefficients of the first pooling feature map and the plurality of second pooling feature maps includes:

[0026] Determine a weighting coefficient of the first pooling feature map as a first preset coefficient;

[0027] Respectively determine the time intervals between the video frames corresponding to the multiple second pooling feature maps and the video frame at the current moment;

[0028] For each second pooling feature map, determine that the weighting coefficient corresponding to the second pooling feature map is the Nth power of the second preset coefficient, wherein the second preset coefficient is less than the first preset coefficient, and N is the time interval between the video frame corresponding to the second pooling feature map and the video frame at the current moment.

[0029] Optionally, the temporal stream convolutional neural network includes a convolution layer, an activation layer and a pooling layer;

[0030] The extracting the feature graphs of the plurality of video frames carrying time information in sequence in time order comprises:

[0031] Inputting the multiple video frames carrying time information into the convolution layer in sequence according to the time order to obtain multiple convolution feature maps carrying time information;

[0032] Inputting the multiple convolutional feature maps carrying time information into the activation layer in sequence according to the time order to obtain multiple activation feature maps carrying time information;

[0033] The multiple activation feature maps carrying time information are input into the pooling layer in time order to obtain multiple pooling feature maps carrying time information.

[0034] In a second aspect, an embodiment of the present invention provides a trusted artificial intelligence device based on a continuous time differential flow field, the device comprising:

[0035] A vehicle driving video acquisition module is used to acquire vehicle driving videos at a plurality of consecutive moments including the current moment during the vehicle driving process, wherein the vehicle driving video is a video captured by a camera of the vehicle during the vehicle driving process;

[0036] A vehicle driving video splitting module is used to split the vehicle driving video into multiple video frames carrying time information according to the time sequence of multiple video frames in the vehicle driving video;

[0037] The video frame input module is used to input the multiple video frames carrying time information into the pre-trained time stream convolutional neural network in time order, so as to perform the following steps through the time stream convolutional neural network:

[0038] Extracting the pooled feature maps of the multiple video frames carrying time information in sequence according to time order;

[0039] Performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map; wherein the first pooling feature map is a pooling feature map of the video frame at the current moment, the plurality of second pooling feature maps are video frames at a plurality of consecutive moments before the video frame at the current moment, and the third pooling feature map is a pooling feature map of the video frame at the current moment after weighted summation;

[0040] Based on the third pooling feature map, determine the information recognition result of the video frame at the current moment, and output the information recognition result.

[0041] In a third aspect, an embodiment of the present invention provides an electronic device, including:

[0042] at least one processor;

[0043] a memory for storing the at least one processor-executable instruction;

[0044] The at least one processor is configured to execute the instructions to implement the method described in the first aspect.

[0045] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in the first aspect.

[0046] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0047] The technical solution provided by the embodiment of the present invention associates each current video frame in the vehicle driving video with the video frames at multiple consecutive moments before it through a time series, and weighted sums the pooled feature maps output by the pooling layer in the convolutional neural network to achieve the purpose of feature transfer, so that the pooled feature maps of the video frames at the current moment after the weighted summation are more robust and stable, and the information recognition results obtained based on the pooled feature maps at the current moment after the weighted summation are more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A schematic diagram of the overall technical solution provided by an embodiment of the present invention;

[0049] Figure 2 A schematic diagram of weighted summation of multiple pooling feature maps when the vehicle driving data in the vehicle driving video provided by the embodiment of the present invention changes rapidly;

[0050] Figure 3 A schematic diagram of weighted summation of multiple pooling feature maps when the vehicle driving data in the vehicle driving video provided by the embodiment of the present invention changes slowly;

[0051] Figure 4 A flow chart of a trusted artificial intelligence method based on continuous time differential flow field provided by an embodiment of the present invention;

[0052] Figure 5 A schematic diagram of the structure of a trusted artificial intelligence device based on continuous time differential flow field provided by an embodiment of the present invention;

[0053] Figure 6 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] The present invention will be described in detail below through examples.

[0055] Convolutional Neural Networks (CNN) have been widely used in the field of deep learning, especially in image and vision tasks. Since CNN significantly reduces computational complexity and improves feature extraction capabilities through parameter sharing and local receptive field design, it is particularly suitable for processing high-dimensional data such as images and speech. However, information loss is difficult to completely avoid during the convolution process, mainly from pooling operations, edge effects, nonlinear activation, and channel dimensionality reduction. Therefore, it is necessary to reduce information loss and recognition errors caused by information loss.

[0056] At present, convolutional neural networks have been gradually applied to intelligent vehicle recognition, for example, identifying obstacles on the road while the vehicle is driving. However, during the convolution process, the recognition accuracy rate will be reduced due to incomplete feature expression, reduced model generalization performance, loss of small targets and details, reduced recognition accuracy and reduced robustness.

[0057] The present invention improves the stability and robustness of the pooled feature map by optimizing the network design. Specifically, during the driving process of the vehicle, a continuous video stream is collected by a camera installed on the vehicle, and a pooled feature map with time information at each time point in the continuous video stream is extracted to form a time information stream, and the pooled feature map at each time point is weighted and summed in chronological order, thereby enhancing the reliability and temporal stability of the pooled feature map.

[0058] The present invention is effective for the video collected by the camera during vehicle driving, and can be used for online information recognition of the camera in the vehicle, for example, identifying vehicles and pedestrians on the road; the video in the vehicle camera will continue to have new video frames entering the recognition algorithm platform, and each video frame of the video in the vehicle camera has a strong correlation with the previous video frame, for example, the video clips of pedestrians walking or vehicles driving in the video, and the video frames with short time intervals will have a strong correlation. The present invention associates each video frame in the video through a time series, and performs weighted summation on multiple pooled feature maps with time information output by the pooling layer in the convolutional neural network, so that the recognition result is more robust and stable, and the recognition accuracy is improved.

[0059] In order to describe the solution clearly, the overall technical solution of the embodiment of the present invention will be described in detail first. Figure 1 shown.

[0060] The present invention is based on the time stream convolutional neural network. All feature extraction, fusion and processing operations are implemented in the time stream convolutional neural network structure. The purpose is to add the pooled feature map in the time stream convolutional neural network to the pooled feature map of the video frame in the past time to improve the stability and robustness of the pooled feature map. Figure 1 As described above, the temporal stream convolutional neural network includes a convolutional layer, an activation layer ( Figure 1 In the temporal stream convolutional neural network, there are generally multiple convolutional layers and pooling layers stacked to gradually extract features from low-level to high-level to form a pooling feature map. This method can be used for all pooling layers or only for more critical pooling layers. The specific implementation method should correspond to different recognition requirements.

[0061] The specific steps may include:

[0062] Step 1, obtain the video captured by the vehicle's camera during the vehicle's driving process. The video can be called the vehicle driving video. The vehicle driving video includes multiple video frames. The multiple video frames can be the video frame at the current moment, which can be recorded as the video frame at time T, and the video frames at multiple consecutive moments before the video frame at the current moment, which can be recorded as the video frame at time T-1, the video frame at time T-2, ..., the video frame at time TN.

[0063] Step 2: In the time sequence of time T, time T-1, time T-2, ..., time TN, multiple video frames included in the vehicle driving video are sequentially input into the time stream convolutional neural network.

[0064] Step 3: In the initial stage of the time stream convolutional neural network, multiple video frames are input into the convolution layer in chronological order, and after convolution operation is performed on each video frame, a convolution feature map with time information is output. Specifically, after the convolution operation is performed on the video frame at time T, the convolution feature map obtained carries the time information T; after the convolution operation is performed on the video frame at time T-1, the convolution feature map obtained carries the time information T-1; similarly, after the convolution operation is performed on the video frame at time TN, the convolution feature map obtained carries the time information TN.

[0065] Step 4: In the order of time T, time T-1, time T-2, ..., time TN, the convolution feature map with time information is input into the activation layer in sequence, and the activation function is applied to remove the negative values ​​in the convolution feature map and increase the nonlinearity, and the activation feature map with time information is output. Specifically, after the convolution feature map at time T is input into the activation layer, the activation feature map obtained carries the time information T; after the convolution feature map at time T-1 is input into the activation layer, the activation feature map obtained carries the time information T-1; similarly, after the convolution feature map at time TN is input into the activation layer, the activation feature map obtained carries the time information TN.

[0066] Step 5: According to the time sequence of time T, time T-1, time T-2, ..., time TN, the activation feature map with time information is sequentially input into the pooling layer for pooling operation. After the activation feature map is pooled, the size becomes smaller, but the original depth (i.e., the number of channels) is maintained. The pooling feature map output from the pooling layer also carries time information. Specifically, after the activation feature map at time T is input into the pooling layer, the resulting pooling feature map carries the time information T; after the activation feature map at time T-1 is input into the pooling layer, the resulting pooling feature map carries the time information T-1; similarly, after the activation feature map at time TN is input into the pooling layer, the resulting pooling feature map carries the time information TN.

[0067] Step 6: Perform weighted summation of the pooled feature map in time sequence to improve the stability of feature extraction, because the pooled feature map not only has complete feature information, but also has a smaller size, which reduces the computational burden. Since each video frame of the vehicle driving video has a strong correlation with the adjacent video frames, the video frames in the time stream can be transferred through this method, making the pooled feature map more robust and stable.

[0068] The pooling feature map is usually a three-dimensional matrix, with the three dimensions being height, width, and number of channels. The size and depth of the pooling feature maps output by the same pooling layer are the same, so the time-continuous pooling feature maps can be weighted and summed. The corresponding formula is as follows:

[0069] F * (T) = F(T)*ω T +F(T-1)*ω T-1 +…+F(TN)*ω T-N

[0070] Among them, the pooled feature map at time T is F(T), F * (T) is the weighted pooling feature map at time T, and the weighting coefficient of F(T) is ω T , the pooling feature map at time T-1 is F(T-1), the pooling feature map at time TN is F(TN), ω T-1 is the weighting coefficient corresponding to F(T-1), ω T-N is the weighting coefficient corresponding to F(TN).

[0071] Regarding the method for determining the weighted coefficients of each pooling feature map, the embodiment of the present invention proposes two different determination methods for different application scenarios.

[0072] The first application scenario is that the vehicle driving data in the vehicle driving video changes rapidly, for example, the vehicle is driving at a high speed. In this case, the correlation between video frames with slightly longer time intervals is not that strong. At this time, only the pooled feature maps of N time steps are extracted for weighting, such as Figure 2 As shown, N can be determined according to the actual situation and is not specifically limited here. In addition, it is assumed that the weighting coefficient of the pooling feature map at the current time T is ω T , the weighted coefficient corresponding to the pooling feature map with a longer step length T from the current moment is smaller, and the weighted coefficient corresponding to the pooling feature map with a shorter step length T from the current moment is larger. That is, ω T-1 >ω T-2 >ω T-N This is because the longer the time interval from the current moment T, the smaller the correlation between the pooling feature map and the pooling feature map at the current moment T, so the corresponding weighting coefficient is smaller. The shorter the time interval from the current moment T, the greater the correlation between the pooling feature map and the pooling feature map at the current moment T, so the corresponding weighting coefficient is larger.

[0073] The second application scenario is that the vehicle driving data in the vehicle driving video changes slowly, for example, the vehicle is driving at a low speed. The features between video frames with a long time span are still related. In this case, the weighted sum of the pooled feature maps of all time steps is performed, such as Figure 3 As shown, the specific formula is as follows:

[0074] F * (2) = F (2) + F (1) * ω

[0075] F * (3) = F (3) + F* (2)*ω=F(3)+F(2)*ω+F(1)*ω 2

[0077] F * (T) = F (T) + F * (T-1)*ω=F(T)+F(T-1)ω+F(T-2)ω 2 +…+F(1)ω T-1

[0078] Among them, ω<1, and the size of ω needs to be determined according to the actual situation. The larger ω is, the higher the weight of the pooled feature map at the previous moment is, and the smaller ω is, the lower the weight of the pooled feature map at the previous moment is, which can be determined according to different driving conditions. The computational complexity of this method is very small, because the calculation result of the previous step can still be used. This method is more suitable for vehicle driving videos with slower changes in vehicle driving data. As the time step gradually increases, the pooled feature map farther away from the current moment will be multiplied by a higher order ω, that is, the weighted coefficient of the pooled feature map with a larger time interval T from the current moment is smaller.

[0079] From the above description, it can be seen that in the above two application scenarios, when the weighted sum of the pooled feature map at the current moment T and the pooled feature map at the previous moment is performed, since the pooled feature map with a smaller time interval from the current moment has a greater correlation with the pooled feature map at the current moment, the corresponding weighting coefficient is larger; similarly, since the pooled feature map with a larger time interval from the current moment has a smaller correlation with the pooled feature map at the current moment, the corresponding weighting coefficient is smaller. Thus, the pooled feature map at the current moment after weighted summation is more accurate, more robust and more stable.

[0080] Step 7: Input the weighted summed pooled feature map at the current moment into the fully connected layer, expand it into a one-dimensional vector, and then perform another operation to obtain the final information recognition result, and the output layer outputs the final information recognition result. The information recognition result can be the recognized target type and the probability of the target type.

[0081] Since the pooled feature map at the current moment after weighted summation is more accurate, more robust and more stable, the information recognition result obtained based on the pooled feature map at the current moment after weighted summation is more accurate.

[0082] After the overall technical solution of the embodiment of the present invention is described in detail, the trusted artificial intelligence method based on the continuous time differential flow field provided by the embodiment of the present invention will be described in detail below. Among them, splitting the vehicle driving video into multiple video frames in time order can be understood as a time flow field, and performing weighted summation of the pooled feature maps of the multiple video frames obtained by splitting can be understood as a differential operation.

[0083] like Figure 4 As shown, the trusted artificial intelligence method based on continuous time differential flow field may include the following steps:

[0084] S410, during the driving of the vehicle, obtaining vehicle driving videos at a plurality of consecutive moments including the current moment.

[0085] The vehicle driving video is the video captured by the vehicle's camera during the vehicle's driving process.

[0086] Specifically, during the driving process of the vehicle, the video captured by the vehicle's camera can be called the vehicle driving video, and the vehicle driving video includes multiple video frames. The multiple video frames can be the video frames at the current moment, which can be recorded as the video frames at the moment T, and the video frames at multiple consecutive moments before the video frame at the current moment, which can be recorded as the video frames at the moment T-1, the video frames at the moment T-2, ..., the video frames at the moments TN.

[0087] S420, splitting the vehicle driving video into multiple video frames carrying time information according to the time sequence of multiple video frames in the vehicle driving video.

[0088] Specifically, the vehicle driving video is split into a plurality of video frames carrying time information in the time sequence of time T, time T-1, time T-2, ..., time TN.

[0089] S430, multiple video frames carrying time information are input into a pre-trained time stream convolutional neural network in time sequence, so as to perform the following steps S440 to S460 through the time stream convolutional neural network.

[0090] Specifically, after multiple video frames carrying time information are obtained by splitting in S420, the multiple video frames carrying time information are input into a pre-trained time stream convolutional neural network.

[0091] S440, extracting pooled feature maps of multiple video frames carrying time information in chronological order.

[0092] In one embodiment, the temporal stream convolutional neural network may include a convolutional layer, an activation layer, and a pooling layer;

[0093] At this time, S440, extracting the pooled feature maps of multiple video frames carrying time information in chronological order may include the following steps, namely, step a1 to step a3:

[0094] Step a1: input multiple video frames carrying time information into the convolution layer in chronological order to obtain multiple convolution feature maps carrying time information.

[0095] Step a2: input multiple convolutional feature maps carrying time information into the activation layer in chronological order to obtain multiple activation feature maps carrying time information.

[0096] Step a3: input multiple activation feature maps carrying time information into the pooling layer in chronological order to obtain multiple pooling feature maps carrying time information.

[0097] It should be noted that since the specific execution steps of the convolution layer, the activation layer and the pooling layer have been explained in detail in the above embodiment, they will not be repeated here.

[0098] S450, performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map.

[0099] Among them, the first pooling feature map is the pooling feature map of the video frame at the current moment, the multiple second pooling feature maps are video frames at multiple consecutive moments before the video frame at the current moment, and the third pooling feature map is the weighted pooling feature map of the video frame at the current moment.

[0100] The stability of feature extraction is improved by weighted summing of the pooled feature maps in chronological order, because the pooled feature maps not only have complete feature information, but also have a smaller size, which reduces the computational burden. In addition, since each video frame of the vehicle driving video has a strong correlation with the adjacent video frames, the video frames in the time stream can transfer features through this method, making the pooled feature map of the current video frame after weighted summing more robust and stable.

[0101] In order to make the description of the solution complete and clear, the specific implementation of S450 will be described in detail in the following embodiments.

[0102] S460, determining the information recognition result of the video frame at the current moment based on the third pooling feature map, and outputting the information recognition result.

[0103] Specifically, after obtaining the weighted summed pooling feature map at the current moment (the third pooling feature map), the third pooling feature map is input into the fully connected layer, expanded into a one-dimensional vector, and then the information recognition result of the video frame at the current moment is obtained after another operation, and the output layer outputs the information recognition result. The information recognition result may be the recognized target type and the probability of the target type.

[0104] Since the pooled feature map at the current moment after weighted summation is more accurate, more robust and more stable, the information recognition result obtained based on the pooled feature map at the current moment after weighted summation is more accurate.

[0105] The technical solution provided by the embodiment of the present invention associates each current video frame in the vehicle driving video with the video frames at multiple consecutive moments before it through a time series, and weighted sums the pooled feature maps output by the pooling layer in the convolutional neural network to achieve the purpose of feature transfer, so that the pooled feature maps of the video frames at the current moment after the weighted summation are more robust and stable, and the information recognition results obtained based on the pooled feature maps at the current moment after the weighted summation are more accurate.

[0106] In order to make the description of the solution complete and clear, the specific implementation of S450 will be explained in detail below.

[0107] exist Figure 4 Based on the above embodiment, in one implementation, if the change rate of the vehicle driving data in the vehicle driving video is greater than a preset rate, for example, the vehicle is driving at a high speed. In this case, the correlation between the video frames with slightly longer time intervals is not so strong, and only the pooled feature maps at N moments are extracted for weighting.

[0108] At this time, S450, performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map may include the following steps, which are steps b1 to b3:

[0109] Step b1, determining target video frames that are located at a target number of consecutive moments before the current video frame.

[0110] The target number is less than the preset number. The preset number is the above-mentioned N, and the specific size of N can be determined according to actual conditions and is not specifically limited here.

[0111] Step b2, determining the pooling feature maps of a target number of target video frames as a plurality of second pooling feature maps;

[0112] Step b3: Perform weighted summation on the first pooling feature map and the determined multiple second pooling feature maps to obtain a third pooling feature map.

[0113] Correspondingly, the process of determining the weighted coefficients of the first pooling feature map and the plurality of second pooling feature maps may include the following steps, namely, step c1 to step c3:

[0114] Step c1, determining a first weighting coefficient of the first pooling feature map. The first weighting coefficient is ω of the above embodiment. T , can be determined according to the actual situation, ω T size.

[0115] Step c2, respectively determining the time interval between the video frames corresponding to the plurality of second pooling feature maps and the video frame at the current moment. The time interval is the step length between the video frame corresponding to the second pooling feature map and the video frame at the current moment.

[0116] Step c3, determining that the second weighting coefficients of the plurality of second pooling feature maps are all smaller than the first weighting coefficient, and the second weighting coefficient is inversely proportional to the time interval.

[0117] Specifically, if the time interval between the video frame corresponding to a second pooling feature map and the video frame at the current moment is long, it means that the correlation between the video frame corresponding to the second pooling feature map and the video frame at the current moment is small, and therefore, the corresponding second weighting coefficient is determined to be small. If the time interval between the video frame corresponding to a second pooling feature map and the video frame at the current moment is short, it means that the correlation between the video frame corresponding to the second pooling feature map and the video frame at the current moment is large, and therefore, the corresponding second weighting coefficient is determined to be large.

[0118] It should be noted that this part of the content has been explained in detail in the first application scenario of the overall technical solution and will not be repeated here.

[0119] exist Figure 4 Based on the embodiment, in another implementation manner, if the change rate of the vehicle driving data in the vehicle driving video is less than the preset rate, that is, the vehicle driving data in the vehicle driving video changes slowly, for example, the vehicle driving speed is low. There is still information correlation between video frames with a long time span. In this case, the pooled feature maps of all video frames in the vehicle driving video are weighted summed.

[0120] At this time, S450, performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map may include the following steps, which are respectively steps d1 to d2:

[0121] Step d1, determining the pooled feature maps of the video frames at all times except the current time in the vehicle driving video as the second pooled feature map.

[0122] Step d2, performing weighted summation on the first pooling feature map and the determined multiple second pooling feature maps.

[0123] Correspondingly, the process of determining the weighted coefficients of the first pooling feature map and the plurality of second pooling feature maps may include the following steps, namely, step e1 to step e3:

[0124] Step e1, determining that the weighting coefficient of the first pooling feature map is a first preset coefficient. The first preset coefficient may be 1.

[0125] Step e2, respectively determine the time interval between the video frames corresponding to the multiple second pooling feature maps and the video frame at the current moment.

[0126] Step e3: for each second pooling feature map, determine that the weighting coefficient corresponding to the second pooling feature map is the Nth power of the second preset coefficient.

[0127] Among them, the second preset coefficient is smaller than the first preset coefficient, and N is the time interval between the video frame corresponding to the second pooling feature map and the video frame at the current moment.

[0128] Assuming that the first preset coefficient is 1 and the second preset coefficient is ω, then ω<1, and the size of ω needs to be determined according to the actual situation. The larger ω is, the greater the weight of the pooled feature map at the previous moment, and the smaller ω is, the lower the weight of the pooled feature map at the previous moment. This can be determined according to different driving conditions.

[0129] When determining the weighting coefficient corresponding to each second pooling feature map, first determine the time interval N between the video frame corresponding to the second pooling feature map and the current video frame. The weighting coefficient corresponding to the second pooling feature map is ω N . This is explained using the following formula as an example.

[0130] F * (T) = F (T) + F * (T-1)*ω=F(T)+F(T-1)ω+F(T-2)ω 2 +…+F(1)ω T-1

[0131] It can be seen from the above formula that the pooling feature map at time T is F(T), the weighting coefficient corresponding to F(T) is 1, the pooling feature map at time T-1 is F(T-1), and the time interval between time T-1 and time T is 1, then the weighting coefficient corresponding to F(T-1) is ω, which is ω to the power of 1; the pooling feature map at time T-2 is F(T-2), and the time interval between time T-2 and time T is 2, then the weighting coefficient corresponding to F(T-2) is ω 2, which is ω to the power of 2; similarly, the pooling feature map at the first moment is F(1), and the time interval between the first moment and the T moment is T-1, then the weighting coefficient corresponding to F(1) is ω T-1 , which is ω to the power of T-1.

[0132] From the above description, it can be seen that in the above two application scenarios, when the weighted sum of the pooled feature map at the current moment T and the pooled feature map at the previous moment is performed, since the pooled feature map with a smaller time interval from the current moment has a greater correlation with the pooled feature map at the current moment, the corresponding weighting coefficient is larger; similarly, since the pooled feature map with a larger time interval from the current moment has a smaller correlation with the pooled feature map at the current moment, the corresponding weighting coefficient is smaller. Thus, the pooled feature map at the current moment after weighted summation is more accurate, more robust and more stable.

[0133] In a second aspect, an embodiment of the present invention provides a trusted artificial intelligence device 50 based on a continuous time differential flow field, such as Figure 5 As shown, the device comprises:

[0134] The vehicle driving video acquisition module 510 is used to acquire the vehicle driving video at a plurality of consecutive moments including the current moment during the vehicle driving process, wherein the vehicle driving video is the video captured by the camera of the vehicle during the vehicle driving process;

[0135] The vehicle driving video splitting module 520 is used to split the vehicle driving video into multiple video frames carrying time information according to the time sequence of multiple video frames in the vehicle driving video;

[0136] The video frame input module 530 is used to input the multiple video frames carrying time information into the pre-trained time stream convolutional neural network in time order, so as to perform the following steps through the time stream convolutional neural network:

[0137] Extracting the pooled feature maps of the multiple video frames carrying time information in sequence according to time order;

[0138] Performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map; wherein the first pooling feature map is a pooling feature map of the video frame at the current moment, the plurality of second pooling feature maps are video frames at a plurality of consecutive moments before the video frame at the current moment, and the third pooling feature map is a pooling feature map of the video frame at the current moment after weighted summation;

[0139] Based on the third pooling feature map, determine the information recognition result of the video frame at the current moment, and output the information recognition result.

[0140] The technical solution provided by the embodiment of the present invention associates each current video frame in the vehicle driving video with the video frames at multiple consecutive moments before it through a time series, and weighted sums the pooled feature maps output by the pooling layer in the convolutional neural network to achieve the purpose of feature transfer, so that the pooled feature maps of the video frames at the current moment after the weighted summation are more robust and stable, and the information recognition results obtained based on the pooled feature maps at the current moment after the weighted summation are more accurate.

[0141] In a third aspect, an embodiment of the present invention provides an electronic device 600, such as Figure 6 As shown, including:

[0142] at least one processor 601;

[0143] a memory 602 for storing the at least one processor executable instruction;

[0144] The at least one processor is configured to execute the instructions to implement the method described in the first aspect.

[0145] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method described in the first aspect.

[0146] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.

[0147] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and intent of the present invention.

Claims

1. A trusted artificial intelligence method based on continuous time differential flow field, characterized in that: The method comprises: During the driving of the vehicle, obtaining a vehicle driving video at a plurality of consecutive moments including the current moment, wherein the vehicle driving video is a video captured by a camera of the vehicle during the driving of the vehicle; Splitting the vehicle driving video into multiple video frames carrying time information according to the time sequence of multiple video frames in the vehicle driving video; The plurality of video frames carrying time information are input into a pre-trained time stream convolutional neural network in time order, so that the time stream convolutional neural network performs the following steps: Extracting the pooled feature maps of the multiple video frames carrying time information in sequence according to time order; Performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map; wherein the first pooling feature map is a pooling feature map of the video frame at the current moment, the plurality of second pooling feature maps are video frames at a plurality of consecutive moments before the video frame at the current moment, and the third pooling feature map is a pooling feature map of the video frame at the current moment after weighted summation; Based on the third pooling feature map, determine the information recognition result of the video frame at the current moment, and output the information recognition result.

2. The method according to claim 1, characterized in that: If the change rate of the vehicle driving data in the vehicle driving video is greater than a preset rate; The step of performing weighted summation of the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map includes: Determine target video frames that are located at a target number of consecutive moments before the current moment video frame, where the target number is less than a preset number; Determine the pooled feature maps of the target number of target video frames as a plurality of second pooled feature maps; The first pooling feature map and the determined multiple second pooling feature maps are weightedly summed to obtain a third pooling feature map.

3. The method according to claim 2, characterized in that The process of determining the weighted coefficients of the first pooling feature map and the plurality of second pooling feature maps includes: Determining a first weighting coefficient of the first pooling feature map; Respectively determine the time intervals between the video frames corresponding to the multiple second pooling feature maps and the video frame at the current moment; Determine that the second weighting coefficients of the multiple second pooling feature maps are all smaller than the first weighting coefficient, and the second weighting coefficient is inversely proportional to the time interval.

4. The method according to claim 1, characterized in that: If the change rate of the vehicle driving data in the vehicle driving video is less than a preset rate; The step of performing weighted summation of the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map includes: Determine the pooled feature maps of video frames at all times except the current time in the vehicle driving video as the second pooled feature map; Perform a weighted summation on the first pooling feature map and the determined multiple second pooling feature maps.

5. The method according to claim 4, characterized in that The process of determining the weighted coefficients of the first pooling feature map and the plurality of second pooling feature maps includes: Determine a weighting coefficient of the first pooling feature map as a first preset coefficient; Respectively determine the time intervals between the video frames corresponding to the multiple second pooling feature maps and the video frame at the current moment; For each second pooling feature map, determine that the weighting coefficient corresponding to the second pooling feature map is the Nth power of the second preset coefficient, wherein the second preset coefficient is less than the first preset coefficient, and N is the time interval between the video frame corresponding to the second pooling feature map and the video frame at the current moment.

6. The method according to any one of claims 1 to 5, characterized in that: The temporal stream convolutional neural network includes a convolution layer, an activation layer and a pooling layer; The extracting the feature graphs of the plurality of video frames carrying time information in sequence in time order comprises: Inputting the multiple video frames carrying time information into the convolution layer in sequence according to the time order to obtain multiple convolution feature maps carrying time information; Inputting the multiple convolutional feature maps carrying time information into the activation layer in sequence according to the time order to obtain multiple activation feature maps carrying time information; The multiple activation feature maps carrying time information are input into the pooling layer in time order to obtain multiple pooling feature maps carrying time information.

7. A trusted artificial intelligence device based on continuous time differential flow field, characterized in that: The device comprises: A vehicle driving video acquisition module is used to acquire vehicle driving videos at a plurality of consecutive moments including the current moment during the vehicle driving process, wherein the vehicle driving video is a video captured by a camera of the vehicle during the vehicle driving process; A vehicle driving video splitting module is used to split the vehicle driving video into multiple video frames carrying time information according to the time sequence of multiple video frames in the vehicle driving video; The video frame input module is used to input the multiple video frames carrying time information into the pre-trained time stream convolutional neural network in time order, so as to perform the following steps through the time stream convolutional neural network: Extracting the pooled feature maps of the multiple video frames carrying time information in sequence according to time order; Performing weighted summation on the first pooling feature map and the plurality of second pooling feature maps to obtain a third pooling feature map; wherein the first pooling feature map is a pooling feature map of the video frame at the current moment, the plurality of second pooling feature maps are video frames at a plurality of consecutive moments before the video frame at the current moment, and the third pooling feature map is a pooling feature map of the video frame at the current moment after weighted summation; Based on the third pooling feature map, determine the information recognition result of the video frame at the current moment, and output the information recognition result.

8. An electronic device, characterized in that: include: at least one processor; a memory for storing the at least one processor-executable instruction; The at least one processor is configured to execute the instructions to implement the method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the method according to any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.

Citation Information

Patent Citations

  • Traffic scene risk assessment method and system based on multi-branch convolutional neural network

    CN112016499A

  • Behavior recognition method and device, electronic equipment and storage medium

    CN113177450A

  • Abnormal driving behavior identification method based on video stream neural network and related device

    CN114612884A

  • Systems and approaches for learning efficient representations for video understanding

    US20210357651A1