Video data processing method and device, storage medium and electronic equipment
By extracting static and dynamic features of financial business videos through a dual-stream neural network model, the problem of low efficiency of manual quality inspection is solved, and the efficiency and accuracy of automated quality inspection are achieved.
Patent Information
- Application Number
- CN202510811241.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
In the existing technology, it is necessary to manually perform quality inspection on videos of users handling financial business, resulting in low quality inspection efficiency.
A dual-stream neural network model is used to extract the static and dynamic features of the video. The static feature information is obtained through the first feature extraction model, and the dynamic feature information is obtained through the second feature extraction model. The two are combined for quality inspection to generate quality inspection results.
It realizes automated quality inspection, reduces reliance on manual quality inspection, improves quality inspection efficiency and accuracy, and can accurately determine whether the operations in the video meet compliance standards.
Smart Images

Figure CN120708126A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a method and device for processing video data, a storage medium, and an electronic device. Background Art
[0002] In the financial industry, with the acceleration of digital transformation, dual recording (audio and video) has become a standard process for ensuring business compliance and protecting consumer rights. Financial institutions' dual-recording video quality inspection technology, a key component of this process, aims to automatically detect and verify whether operator behavior in videos complies with regulations, thereby improving service quality and mitigating risks. Financial institutions need to conduct comprehensive and detailed quality inspections of dual-recording videos when handling transactions such as wealth management product sales, loan applications, and customer inquiries. This includes, but is not limited to, verifying whether operators display identification, guiding customers through the correct signing gestures, and ensuring sufficient footage of both people in the video. An efficient and accurate quality inspection system is crucial for reducing labor costs, accelerating business processes, and improving customer satisfaction. Despite the development of automated quality inspection technology, existing technologies often require manual quality inspection of videos of users conducting financial transactions, which is difficult to meet the efficient quality inspection needs of real-world businesses.
[0003] Currently, no effective solution has been proposed to address the problem that the quality inspection of videos of users handling financial business needs to be performed manually, resulting in relatively low efficiency of video quality inspection. Summary of the Invention
[0004] The main purpose of this application is to provide a video data processing method and device, storage medium and electronic device to solve the problem in related technologies that manual quality inspection of videos of users handling financial business is required, resulting in relatively low efficiency of video quality inspection.
[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for processing video data is provided. The method comprises: obtaining a target video to be detected, wherein the target video is a video of the process of providing financial services to a target object; extracting static features of the target video using a first feature extraction model in a dual-stream neural network model to obtain first feature information; extracting dynamic features of the target video using a second feature extraction model in the dual-stream neural network model to obtain second feature information; performing quality inspection on the target video based on the first feature information and the second feature information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to characterize whether the process of providing financial services to the target object meets preset requirements.
[0006] Furthermore, the static features of the target video are extracted by the first feature extraction model in the dual-stream neural network model to obtain the first feature information, including: decomposing the target video according to a preset frame rate to obtain multiple initial video frames; screening the multiple initial video frames to obtain multiple target video frames; performing feature extraction on the multiple target video frames to obtain initial feature information corresponding to each target video frame; and obtaining the first feature information based on the initial feature information.
[0007] Furthermore, the dynamic features of the target video are extracted by the second feature extraction model in the dual-stream neural network model to obtain the second feature information, including: decomposing the target video to obtain multiple target video frames; calculating the difference between two adjacent frames of the multiple target video frames to obtain multiple optical flow images; and performing feature extraction on the multiple optical flow images to obtain the second feature information.
[0008] Furthermore, the difference between two adjacent frames of images in the multiple target video frames is calculated to obtain multiple optical flow images, including: filtering the pixel points in the first frame of the two frames according to the brightness of the pixel points to obtain multiple first feature points; obtaining multiple second feature points in the second frame of the two frames corresponding to the multiple first feature points; and obtaining optical flow images corresponding to the two frames based on the multiple first feature points and the multiple second feature points.
[0009] Furthermore, based on the multiple first feature points and the multiple second feature points, obtaining the optical flow image corresponding to the two frames of images includes: for each first feature point, calculating the displacement vector between the first feature point and the corresponding second feature point to obtain multiple displacement vectors; based on the multiple displacement vectors, generating a two-dimensional optical flow field; and converting the two-dimensional optical flow field to obtain the optical flow image.
[0010] Furthermore, the difference between two adjacent frames of images in the multiple target video frames is calculated to obtain multiple optical flow images, including: performing feature extraction on the two frames of images through the optical flow generation layer in the dual-stream neural network model to obtain feature information corresponding to the first frame of the two frames and feature information corresponding to the second frame of the two frames; performing correlation calculation on the feature information corresponding to the first frame of the image and the feature information corresponding to the second frame of the image to obtain a correlation graph; performing calculation based on the correlation graph to obtain a displacement vector of the pixel point in the first frame of the image; and obtaining the multiple optical flow images based on the displacement vector.
[0011] Furthermore, calculating based on the correlation graph to obtain the displacement vector of the pixel point in the first frame image includes: calculating the correlation graph to obtain a probability distribution, wherein the probability distribution represents the matching probability between the pixel point in the first frame image and the pixel point in the second frame image; calculating the optical flow component in the horizontal direction and the optical flow component in the vertical direction of the pixel point in the first frame image based on the probability distribution; and obtaining the displacement vector based on the optical flow component in the horizontal direction and the optical flow component in the vertical direction.
[0012] In order to achieve the above-mentioned purpose, according to another aspect of the present application, a video data processing device is provided. The device includes: an acquisition unit, configured to acquire a target video to be detected, wherein the target video is a process video of providing financial services to a target object; a first extraction unit, configured to extract static features of the target video using a first feature extraction model in a dual-stream neural network model to obtain first feature information; a second extraction unit, configured to extract dynamic features of the target video using a second feature extraction model in the dual-stream neural network model to obtain second feature information; and a quality inspection unit, configured to perform quality inspection on the target video based on the first feature information and the second feature information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to characterize whether the process of providing financial services to the target object meets preset requirements.
[0013] Furthermore, the first extraction unit includes: a first decomposition subunit, used to decompose the target video according to a preset frame rate to obtain multiple initial video frames; a screening subunit, used to screen the multiple initial video frames to obtain multiple target video frames; a first extraction subunit, used to perform feature extraction on the multiple target video frames to obtain initial feature information corresponding to each target video frame; and a determination subunit, used to obtain the first feature information based on the initial feature information.
[0014] Furthermore, the second extraction unit includes: a second decomposition subunit, used to decompose the target video to obtain multiple target video frames; a calculation subunit, used to calculate the difference between two adjacent frames of images in the multiple target video frames to obtain multiple optical flow images; and a second extraction subunit, used to perform feature extraction on the multiple optical flow images to obtain the second feature information.
[0015] Furthermore, the computing subunit includes: a screening module for screening the pixel points in the first frame image of the two frames according to the brightness of the pixel points to obtain a plurality of first feature points; an acquisition module for acquiring a plurality of second feature points corresponding to the plurality of first feature points in the second frame image of the two frames; and a first determination module for obtaining the optical flow images corresponding to the two frames of images based on the plurality of first feature points and the plurality of second feature points.
[0016] Furthermore, the determination module includes: a first calculation submodule, used to calculate the displacement vector between each first feature point and the corresponding second feature point for each first feature point to obtain multiple displacement vectors; a generation submodule, used to generate a two-dimensional optical flow field based on the multiple displacement vectors; and a conversion submodule, used to convert the two-dimensional optical flow field to obtain the optical flow image.
[0017] Furthermore, the computing subunit includes: an extraction module, used to extract features from the two frames of images through the optical flow generation layer in the dual-stream neural network model, and obtain feature information corresponding to the first frame of the two frames and feature information corresponding to the second frame of the two frames; a first computing module, used to perform correlation calculation on the feature information corresponding to the first frame of the image and the feature information corresponding to the second frame of the image, and obtain a correlation graph; a second computing module, used to perform calculation based on the correlation graph, and obtain a displacement vector of the pixel point in the first frame of the image; a second determination module, used to obtain the multiple optical flow images based on the displacement vector.
[0018] Furthermore, the second calculation module includes: a second calculation submodule, used to calculate the correlation graph to obtain a probability distribution, wherein the probability distribution represents the matching probability between the pixel points in the first frame image and the pixel points in the second frame image; a third calculation submodule, used to calculate the optical flow component in the horizontal direction and the optical flow component in the vertical direction of the pixel points in the first frame image based on the probability distribution; a determination submodule, used to obtain the displacement vector based on the optical flow component in the horizontal direction and the optical flow component in the vertical direction.
[0019] According to another aspect of an embodiment of the present invention, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes any one of the above-mentioned methods for processing video data when running.
[0020] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided. The storage medium stores a program, wherein when the program is running, the device where the storage medium is located is controlled to execute any of the above-mentioned video data processing methods.
[0021] In an embodiment of the present application, the following steps are adopted: obtaining a target video to be detected, wherein the target video is a process video of providing financial services to a target object; extracting static features of the target video through a first feature extraction model in a dual-stream neural network model to obtain first feature information; extracting dynamic features of the target video through a second feature extraction model in the dual-stream neural network model to obtain second feature information; performing quality inspection on the target video based on the first feature information and the second feature information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to characterize whether the process of providing financial services to the target object meets preset requirements, thereby solving the technical problem in related technologies that manual quality inspection is required for videos of users handling financial services, resulting in relatively low efficiency of video quality inspection.
[0022] In this solution, the first feature extraction model in the dual-stream neural network model processes each frame of the target video, extracting its static feature information to obtain the first feature information, and the second feature extraction model captures the dynamic feature information to obtain the second feature information. The first feature information (static features) is combined with the second feature information (dynamic features) to perform quality inspection on the target video. Specifically, the target video is evaluated to see whether the required elements (such as work IDs and customer signatures) appear in the video. By comparing the video content with the preset business process requirements, it can automatically determine whether the operations in the video meet the compliance standards and generate a quality inspection report. The quality inspection report describes which aspects of the target subject's financial business process meet the preset requirements and which aspects deviate or do not meet the standards. By integrating dynamic and static information, the traditional quality inspection method's high reliance on manpower is effectively avoided. Automated processing significantly reduces the time spent on manual video review and improves the efficiency of the quality inspection process. The dynamic and static fusion strategy effectively overcomes the misjudgments or omissions that may occur due to single feature analysis, thereby achieving the technical effect of improving the accuracy of quality inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0024] Figure 1 A hardware structure block diagram of a computer terminal for implementing a method for processing video data is shown;
[0025] Figure 2 is a flowchart of a method for processing video data provided in an embodiment of the present application;
[0026] Figure 3 is a schematic diagram of an optical flow image provided according to an embodiment of the present application;
[0027] Figure 4 is a schematic diagram of a two-stream neural network model provided according to an embodiment of the present application;
[0028] Figure 5 is a schematic diagram of a video data processing device provided according to an embodiment of the present application;
[0029] Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation portals for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.
[0033] Example 1
[0034] According to an embodiment of the present application, an embodiment of a method for processing video data is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0035] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for processing video data. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0036] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0037] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the video data processing method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned video data processing method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0038] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0039] The display may be a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0040] Under the above operating environment, this application provides Figure 2 The video data processing method shown. Figure 2 1 is a flowchart of a method for processing video data according to Embodiment 1 of the present application. The method for processing video data includes:
[0041] Step S201: obtaining a target video to be detected, wherein the target video is a process video of providing financial services to a target object.
[0042] Optionally, when providing financial services to customers (i.e., the above-mentioned target objects), such as purchasing financial products, applying for loans, signing insurance contracts, etc., the staff of the financial institution will start the dual recording equipment to record the entire business process to obtain the above-mentioned target video to be detected.
[0043] Step S202: extract static features of the target video using the first feature extraction model in the dual-stream neural network model to obtain first feature information.
[0044] Optionally, the target video is input into a two-stream neural network model, and static features are extracted using the first feature extraction model of the two-stream neural network to obtain first feature information. For example, the target video is first decomposed into consecutive frame images, and each frame is preprocessed, including but not limited to resizing, normalization, and noise reduction. Each frame is then input into the first feature extraction model, which extracts static features from the image, such as edges, textures, and shapes.
[0045] Step S203: extract the dynamic features of the target video through the second feature extraction model in the dual-stream neural network model to obtain second feature information.
[0046] Optionally, dynamic features refer to content in a video that changes over time, such as the movement of objects, changes in a person's posture, hand movements, etc. They are not limited to the static information of a single frame, but rather span temporal changes across multiple frames. In the dual-recording quality inspection scenario of financial institutions, dynamic features are particularly important because they can help determine whether the behavior of financial managers or customers follows prescribed procedures, for example, whether the financial manager correctly displays his or her work ID or whether the customer signs at key points. The second feature extraction model in the two-stream neural network model extracts dynamic features of the target video, such as the displacement changes of pixels in the scene, to identify actions and behaviors in the video.
[0047] Step S204 : performing a quality inspection on the target video based on the first characteristic information and the second characteristic information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to indicate whether the process of providing financial services to the target object meets preset requirements.
[0048] Optionally, the first feature information and the second feature information can be fused through the classification subnet in the dual-stream neural network model to obtain a fused feature vector, and then the fused feature vector is input into the decision layer of the classification subnet, such as a fully connected layer (FC layer), for classification judgment. The decision layer uses pre-trained model weights based on comprehensive features to classify the video and determine whether it meets the preset compliance requirements (for example, whether the financial manager correctly displays the work permit, whether the customer signs at key points). Through the fusion analysis of dynamic and static features, the dual-stream neural network model can perform detailed and comprehensive intelligent quality inspection on the target video, reducing reliance on manual judgment and improving the efficiency and consistency of quality inspection.
[0049] In summary, the first feature extraction model in the two-stream neural network model processes each frame of the target video, extracting its static feature information to obtain the first feature information. The second feature extraction model captures the dynamic feature information to obtain the second feature information. The first feature information (static features) is combined with the second feature information (dynamic features) to perform quality inspection on the target video. Specifically, the target video is evaluated for the presence of specified elements (such as work IDs and customer signatures). By comparing the video content with the preset business process requirements, it can automatically determine whether the operations in the video meet the compliance standards and generate a quality inspection report. The quality inspection report describes which aspects of the target subject's financial business process meet the preset requirements and which aspects deviate or do not meet the standards. By integrating dynamic and static information, the traditional quality inspection method's high reliance on manpower is effectively avoided. Automated processing significantly reduces the time spent on manual video review and improves the efficiency of the quality inspection process. Furthermore, the dynamic and static fusion strategy effectively overcomes the misjudgments or omissions that may occur due to single feature analysis, thereby achieving the technical effect of improving quality inspection accuracy.
[0050] Optionally, in the video data processing method provided in the embodiment of the present application, the static features of the target video are extracted by the first feature extraction model in the dual-stream neural network model to obtain the first feature information, including: decomposing the target video according to a preset frame rate to obtain multiple initial video frames; screening the multiple initial video frames to obtain multiple target video frames; performing feature extraction on the multiple target video frames to obtain initial feature information corresponding to each target video frame; and obtaining the first feature information based on the initial feature information.
[0051] In an optional embodiment, the target video is decomposed into multiple initial video frames according to a preset frame rate, which is the number of video frames extracted per second. The preset frame rate is selected based on the dynamic nature of the video content and the density of the required information, ensuring that sufficient details are captured without inefficient processing due to excessive data.
[0052] In order to eliminate those frames that are not helpful or have low contribution to feature extraction, such as overly similar consecutive frames, or frames of poor quality, such as blur, overexposure, etc. The initial video frames are preprocessed and analyzed, such as contrast inspection, inter-frame difference analysis, etc., to screen out representative or information-rich target video frames to reduce unnecessary data processing and storage requirements while maintaining the integrity of key information. Then, the first feature extraction model is used to perform deep learning processing on the multiple screened target video frames to extract the static feature information of each frame, that is, to obtain the above-mentioned initial feature information. After completing the feature extraction of each target video frame, all the initial feature information needs to be integrated to obtain the overall static features of the target video (that is, the above-mentioned first feature information). For example, feature dimensionality reduction, clustering or weighted averaging unifies the diverse feature information from different frames to form a compact first feature information to comprehensively describe the static content of the video.
[0053] By decomposing the video, screening valid frames, performing deep feature extraction on the selected frames and final feature integration, the accurate capture of the static features of the video is ensured.
[0054] Optionally, in the video data processing method provided in the embodiment of the present application, the dynamic features of the target video are extracted by the second feature extraction model in the dual-stream neural network model to obtain the second feature information, including: decomposing the target video to obtain multiple target video frames; calculating the difference between two adjacent frames of images in the multiple target video frames to obtain multiple optical flow images; and performing feature extraction on the multiple optical flow images to obtain the second feature information.
[0055] In an optional embodiment, extracting dynamic features of a target video includes: decomposing the target video, and calculating the difference between two adjacent frames of the decomposed target video frames, for example, analyzing the pixel displacement between adjacent frames, i.e., optical flow. The optical flow image intuitively shows the movement trend and speed of an object or background between two frames. By calculating the optical flow images of two adjacent frames in the target video frame, dynamic changes and motion patterns in the video are revealed. For example, Figure 3 In the schematic diagram shown, a and b are two consecutive frames of images in the video, and c is an optical flow image, in which the two channels represent the movement of the corresponding pixels in the horizontal and vertical directions, respectively.
[0056] After obtaining multiple optical flow images, the model extracts features from these images to generate secondary features. The model identifies key patterns in the optical flow images, such as specific hand gestures or body movements, which reflect the essence of dynamic features. Through multiple layers of convolution, pooling, and fully connected operations, the model encodes dynamic behavior and ultimately generates secondary features.
[0057] The second feature extraction model in the two-stream neural network can accurately capture and quantify the dynamic features in the target video, thereby improving the accuracy of subsequent quality inspection.
[0058] Optionally, in the video data processing method provided in the embodiment of the present application, the difference between two adjacent frames of images in multiple target video frames is calculated to obtain multiple optical flow images, including: filtering the pixel points in the first frame of the two frames according to the brightness of the pixel points to obtain multiple first feature points; obtaining multiple second feature points in the second frame of the two frames corresponding to the multiple first feature points; and obtaining optical flow images corresponding to the two frames based on the multiple first feature points and the multiple second feature points.
[0059] In an optional embodiment, in two adjacent frames of images, the first frame of image is first selected as a reference, and based on the brightness change or contrast of the pixels, pixels with obvious features are selected to obtain the above-mentioned multiple first feature points. The multiple first feature points can be edges, corners or the center of other visually significant areas, which are easy to track and can reflect local motion information. The first feature points obtained after screening will be searched for corresponding second feature points in the next frame of image (i.e., the second frame of image). This can be achieved by calculating the similarity between the candidate area and the image blocks around the first feature points. The purpose of matching is to determine the relative movement of these feature points between the two frames, providing a basis for the optical flow estimation algorithm. Based on the results of the feature point matching, the optical flow images corresponding to the above-mentioned two frames of images are finally obtained based on the displacement changes between the multiple first feature points and the multiple second feature points.
[0060] Optical flow images not only capture instantaneous dynamic changes, but also retain the direction and speed information of movement. They can accurately capture the movement changes of staff and customers in the video, such as the consistency and standardization of key processes such as document display and document signing, thereby achieving the technical effect of improving quality inspection accuracy.
[0061] Optionally, in the video data processing method provided in the embodiment of the present application, obtaining optical flow images corresponding to two frames of images based on multiple first feature points and multiple second feature points includes: for each first feature point, calculating the displacement vector between the first feature point and the corresponding second feature point to obtain multiple displacement vectors; generating a two-dimensional optical flow field based on the multiple displacement vectors; and converting the two-dimensional optical flow field to obtain an optical flow image.
[0062] In an optional embodiment, after matching feature points in the current frame (the first image) with the next frame (the second image), the positional difference between each pair of matched feature points, known as a displacement vector, is calculated. This displacement vector, which includes both horizontal and vertical displacement components and quantitatively describes the change in position of a feature point between the two frames, is an essential component of optical flow analysis.
[0063] After obtaining the displacement vectors of all matching feature points, these displacement vectors are used to construct a two-dimensional optical flow field within the image's spatial coordinate system. The optical flow field is a two-dimensional grid, where each grid point corresponds to a location, and the vector at that location represents the optical flow information for that point—its direction and velocity in the next frame. Interpolation methods can be used to extend the displacement vectors of known feature points across the entire image, filling in the optical flow information at non-feature point locations to ensure the continuity and integrity of the optical flow field.
[0064] While the optical flow field provides an intuitive representation of optical flow information, it is not easy to visualize or further process. Therefore, it is necessary to convert the two-dimensional optical flow field into an optical flow image for easier understanding and application. This conversion can be performed using the HSV color space, where H (hue) represents the direction of the displacement vector, S (saturation), and V (brightness) represent the magnitude of the displacement. The color of each pixel reflects the displacement trend at that point, while the brightness indicates the magnitude of the displacement.
[0065] By calculating the displacement vectors between feature points, constructing a continuous two-dimensional optical flow field and ultimately converting it into an optical flow image, the model can capture and visualize dynamic behavior from consecutive frames of a video. By analyzing optical flow images, the model can more accurately identify dynamic activities in the video, such as the coherent movements of a financial manager presenting a document or the smoothness of a customer's signature, thereby improving the accuracy and efficiency of intelligent quality inspection.
[0066] Optionally, in the video data processing method provided in the embodiment of the present application, the difference between two adjacent frames of images in multiple target video frames is calculated to obtain multiple optical flow images, including: extracting features from the two frames of images through the optical flow generation layer in the dual-stream neural network model to obtain feature information corresponding to the first frame of the two frames and feature information corresponding to the second frame of the two frames; performing correlation calculation on the feature information corresponding to the first frame of the image and the feature information corresponding to the second frame of the image to obtain a correlation graph; performing calculation based on the correlation graph to obtain a displacement vector of the pixel point in the first frame of the image; and obtaining multiple optical flow images based on the displacement vector.
[0067] In an optional embodiment, the following steps can also be used to obtain the above-mentioned multiple optical flow images: the optical flow generation layer in the dual-stream neural network model first performs deep feature extraction on two adjacent frames of images (the first frame image and the second frame image). Unlike conventional image feature extraction, the extraction here focuses on capturing information that can reflect image changes and movements. Through the convolutional layer of the neural network, the texture, boundary, depth and other features of the image can be extracted to obtain the feature information of the first frame image and the second frame image for subsequent optical flow calculation.
[0068] Based on the feature information, the optical flow generation layer further calculates the correlation between the feature information of the first frame and the feature information of the second frame. This correlation calculation typically uses feature matching, comparing the feature similarities at different locations in the two frames to determine which feature points maintain a stable correlation between the two frames. The resulting correlation map is a two-dimensional matrix, where each element represents the similarity between a position in the first frame and the corresponding position in the second frame. Based on the correlation map, the displacement vector of each pixel in the first frame in the second frame can be calculated, and multiple optical flow images can be generated based on the displacement vectors.
[0069] The optical flow generation process within the two-stream neural network efficiently and accurately captures and calculates motion information between video frames through deep feature extraction and correlation calculation. In dual-recording scenarios at financial institutions, this approach can effectively identify the details of staff movements and the continuity of customer interactions, improving the accuracy and efficiency of automated quality inspections.
[0070] Optionally, in the video data processing method provided in the embodiment of the present application, calculating based on the correlation graph to obtain the displacement vector of the pixel point in the first frame image includes: calculating the correlation graph to obtain a probability distribution, wherein the probability distribution represents the matching probability between the pixel point in the first frame image and the pixel point in the second frame image; calculating the horizontal optical flow component and the vertical optical flow component of the pixel point in the first frame image based on the probability distribution; and obtaining the displacement vector based on the horizontal optical flow component and the vertical optical flow component.
[0071] In an optional embodiment, after obtaining the correlation graph, the correlation graph is calculated to obtain a probability distribution. Each element in the correlation graph represents the degree of match between a certain pixel point in the first frame image and all possible positions in the second frame image. The calculation of the probability distribution is actually to normalize these matching degrees so that the sum of the elements in each row (representing a pixel point in the first frame) is equal to 1. In this way, each row of elements actually represents the probability of the pixel point appearing in different positions in the second frame. In an optional embodiment, the correlation graph can be normalized by a Softmax function. The Softmax function converts each element into a probability value and ensures that the sum of the probability values of all rows is 1. In this way, the possible matching positions of each pixel point in the second frame have a clear probability representation.
[0072] Based on the probability distribution, we can calculate the expected position of each pixel in the second frame—that is, the most likely position for that pixel in the next frame. This expected position is obtained by taking the weighted average of all possible positions of the pixel in the second frame. Based on this expected position, the horizontal and vertical optical flow components of the pixel in the first frame are calculated. Finally, the horizontal and vertical optical flow components for each pixel are combined to form the displacement vector for that pixel. The displacement vector describes the direction and magnitude of the pixel's movement between the two frames.
[0073] By converting the correlation graph into a probability distribution, and then calculating the horizontal and vertical optical flow components based on the probability distribution, and finally generating the displacement vector of each pixel, the precise conversion from image features to optical flow information is achieved, providing a solid foundation for subsequent dynamic feature extraction and behavior analysis.
[0074] In an optional embodiment, the schematic diagram of the two-stream neural network model is as follows Figure 4 As shown, it consists of two parallel convolutional networks: a first feature extraction model and a second feature extraction model. The first feature extraction model extracts static background and outline features, while the second feature extraction model extracts dynamic motion features. The upper half corresponds to the feature fusion and classification subnetwork, which is a simple FC layer classification subnetwork. The two-stream neural network architecture effectively combines background and motion information in the video by processing static images and dynamic optical flow in parallel, improving the accuracy of financial transaction quality inspection.
[0075] The video data processing method provided in the embodiment of the present application obtains a target video to be detected, wherein the target video is a process video of providing financial services to the target object; extracts the static features of the target video through the first feature extraction model in the dual-stream neural network model to obtain first feature information; extracts the dynamic features of the target video through the second feature extraction model in the dual-stream neural network model to obtain second feature information; performs quality inspection on the target video based on the first feature information and the second feature information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to characterize whether the process of providing financial services to the target object meets preset requirements, thereby solving the technical problem in related technologies that manual quality inspection is required for videos of users handling financial services, resulting in relatively low efficiency of video quality inspection.
[0076] In this solution, the first feature extraction model in the dual-stream neural network model processes each frame of the target video, extracting its static feature information to obtain the first feature information, and the second feature extraction model captures the dynamic feature information to obtain the second feature information. The first feature information (static features) is combined with the second feature information (dynamic features) to perform quality inspection on the target video. Specifically, the target video is evaluated to see whether the required elements (such as work IDs and customer signatures) appear in the video. By comparing the video content with the preset business process requirements, it can automatically determine whether the operations in the video meet the compliance standards and generate a quality inspection report. The quality inspection report describes which aspects of the target subject's financial business process meet the preset requirements and which aspects deviate or do not meet the standards. By integrating dynamic and static information, the traditional quality inspection method's high reliance on manpower is effectively avoided. Automated processing significantly reduces the time spent on manual video review and improves the efficiency of the quality inspection process. The dynamic and static fusion strategy effectively overcomes the misjudgments or omissions that may occur due to single feature analysis, thereby achieving the technical effect of improving the accuracy of quality inspection.
[0077] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0078] Example 2
[0079] The embodiment of the present application further provides a video data processing device. It should be noted that the video data processing device of the embodiment of the present application can be used to execute the video data processing method provided in the embodiment of the present application. The video data processing device provided in the embodiment of the present application is introduced below.
[0080] According to an embodiment of the present application, a device for implementing the above-mentioned video data processing method is also provided. Figure 5 As shown, the device includes: an acquisition unit 501, a first extraction unit 502, a second extraction unit 503 and a quality inspection unit 504.
[0081] An acquisition unit 501 is configured to acquire a target video to be detected, wherein the target video is a video of a process of providing financial services to a target object;
[0082] A first extraction unit 502 is configured to extract static features of the target video using a first feature extraction model in the two-stream neural network model to obtain first feature information;
[0083] A second extraction unit 503 is configured to extract dynamic features of the target video using a second feature extraction model in the two-stream neural network model to obtain second feature information;
[0084] The quality inspection unit 504 is used to perform quality inspection on the target video based on the first feature information and the second feature information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to indicate whether the process of providing financial services to the target object meets preset requirements.
[0085] The video data processing device provided in the embodiment of the present application obtains the target video to be detected through the acquisition unit 501, wherein the target video is a process video of providing financial services to the target object; the first extraction unit 502 extracts the static features of the target video through the first feature extraction model in the dual-stream neural network model to obtain first feature information; the second extraction unit 503 extracts the dynamic features of the target video through the second feature extraction model in the dual-stream neural network model to obtain second feature information; the quality inspection unit 504 performs quality inspection on the target video based on the first feature information and the second feature information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to characterize whether the process of providing financial services to the target object meets the preset requirements, thereby solving the technical problem in the related art that manual quality inspection is required for videos of users handling financial services, resulting in relatively low efficiency of video quality inspection.
[0086] In this solution, the first feature extraction model in the dual-stream neural network model processes each frame of the target video, extracting its static feature information to obtain the first feature information, and the second feature extraction model captures the dynamic feature information to obtain the second feature information. The first feature information (static features) is combined with the second feature information (dynamic features) to perform quality inspection on the target video. Specifically, the target video is evaluated to see whether the required elements (such as work IDs and customer signatures) appear in the video. By comparing the video content with the preset business process requirements, it can automatically determine whether the operations in the video meet the compliance standards and generate a quality inspection report. The quality inspection report describes which aspects of the target subject's financial business process meet the preset requirements and which aspects deviate or do not meet the standards. By integrating dynamic and static information, the traditional quality inspection method's high reliance on manpower is effectively avoided. Automated processing significantly reduces the time spent on manual video review and improves the efficiency of the quality inspection process. The dynamic and static fusion strategy effectively overcomes the misjudgments or omissions that may occur due to single feature analysis, thereby achieving the technical effect of improving the accuracy of quality inspection.
[0087] Optionally, in the video data processing device provided in the embodiment of the present application, the first extraction unit includes: a first decomposition subunit, used to decompose the target video according to a preset frame rate to obtain multiple initial video frames; a screening subunit, used to screen the multiple initial video frames to obtain multiple target video frames; a first extraction subunit, used to perform feature extraction on the multiple target video frames to obtain initial feature information corresponding to each target video frame; and a determination subunit, used to obtain first feature information based on the initial feature information.
[0088] Optionally, in the video data processing device provided in the embodiment of the present application, the second extraction unit includes: a second decomposition subunit, used to decompose the target video to obtain multiple target video frames; a calculation subunit, used to calculate the difference between two adjacent frames of images in the multiple target video frames to obtain multiple optical flow images; and a second extraction subunit, used to perform feature extraction on the multiple optical flow images to obtain second feature information.
[0089] Optionally, in the video data processing device provided in the embodiment of the present application, the computing subunit includes: a screening module for screening the pixel points in the first frame image of the two frames according to the brightness of the pixel points to obtain multiple first feature points; an acquisition module for acquiring multiple second feature points corresponding to the multiple first feature points in the second frame image of the two frames; and a first determination module for obtaining optical flow images corresponding to the two frames of images based on the multiple first feature points and the multiple second feature points.
[0090] Optionally, in the video data processing device provided in the embodiment of the present application, the determination module includes: a first calculation sub-module, used to calculate the displacement vector between the first feature point and the corresponding second feature point for each first feature point, to obtain multiple displacement vectors; a generation sub-module, used to generate a two-dimensional optical flow field based on the multiple displacement vectors; and a conversion sub-module, used to convert the two-dimensional optical flow field to obtain an optical flow image.
[0091] Optionally, in the video data processing device provided in the embodiment of the present application, the computing subunit includes: an extraction module, which is used to extract features of the two frames of images through the optical flow generation layer in the dual-stream neural network model to obtain feature information corresponding to the first frame of the two frames and feature information corresponding to the second frame of the two frames; a first computing module, which is used to perform correlation calculation on the feature information corresponding to the first frame of the image and the feature information corresponding to the second frame of the image to obtain a correlation graph; a second computing module, which is used to perform calculations based on the correlation graph to obtain a displacement vector of a pixel point in the first frame of the image; and a second determination module, which is used to obtain multiple optical flow images based on the displacement vector.
[0092] Optionally, in the video data processing device provided in the embodiment of the present application, the second calculation module includes: a second calculation submodule, used to calculate the correlation graph to obtain a probability distribution, wherein the probability distribution represents the matching probability between the pixel points in the first frame image and the pixel points in the second frame image; a third calculation submodule, used to calculate the horizontal optical flow component and the vertical optical flow component of the pixel points in the first frame image based on the probability distribution; and a determination submodule, used to obtain a displacement vector based on the horizontal optical flow component and the vertical optical flow component.
[0093] It should be noted that the acquisition unit 501, the first extraction unit 502, the second extraction unit 503, and the quality inspection unit 504 described above correspond to steps S201 to S204 in the first embodiment. The four units and the corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in the first embodiment. It should be noted that the modules or units described above may be hardware components or software components stored in a memory (e.g., the memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The units described above may also be part of a device and run in the computer terminal 10 provided in the first embodiment.
[0094] Example 3
[0095] An embodiment of the present application may provide an electronic device, Figure 6 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 6 As shown, the electronic device may include: one or more ( Figure 6 Only one is shown) processor 602, memory 604, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0096] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0097] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain the target video to be detected, wherein the target video is a process video of providing financial services to the target object; extract the static features of the target video through the first feature extraction model in the dual-stream neural network model to obtain first feature information; extract the dynamic features of the target video through the second feature extraction model in the dual-stream neural network model to obtain second feature information; perform quality inspection on the target video based on the first feature information and the second feature information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to characterize whether the process of providing financial services to the target object meets the preset requirements.
[0098] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: extract the static features of the target video through the first feature extraction model in the dual-stream neural network model to obtain the first feature information, including: decomposing the target video according to a preset frame rate to obtain multiple initial video frames; screening the multiple initial video frames to obtain multiple target video frames; performing feature extraction on the multiple target video frames to obtain initial feature information corresponding to each target video frame; and obtaining the first feature information based on the initial feature information.
[0099] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: extract the dynamic features of the target video through the second feature extraction model in the dual-stream neural network model to obtain the second feature information, including: decomposing the target video to obtain multiple target video frames; calculating the difference between two adjacent frames of images in the multiple target video frames to obtain multiple optical flow images; and extracting features from the multiple optical flow images to obtain the second feature information.
[0100] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: calculating the difference between two adjacent frames of images in multiple target video frames to obtain multiple optical flow images, including: filtering the pixel points in the first frame of the two frames according to the brightness of the pixel points to obtain multiple first feature points; obtaining multiple second feature points in the second frame of the two frames corresponding to the multiple first feature points; and obtaining optical flow images corresponding to the two frames of images based on the multiple first feature points and the multiple second feature points.
[0101] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: based on multiple first feature points and multiple second feature points, obtain the optical flow image corresponding to the two frames of images, including: for each first feature point, calculate the displacement vector between the first feature point and the corresponding second feature point to obtain multiple displacement vectors; based on the multiple displacement vectors, generate a two-dimensional optical flow field; convert the two-dimensional optical flow field to obtain an optical flow image.
[0102] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: calculating the difference between two adjacent frames of images in multiple target video frames to obtain multiple optical flow images, including: extracting features from the two frames of images through the optical flow generation layer in the dual-stream neural network model to obtain feature information corresponding to the first frame of the two frames and feature information corresponding to the second frame of the two frames; performing correlation calculation on the feature information corresponding to the first frame of the image and the feature information corresponding to the second frame of the image to obtain a correlation graph; performing calculation based on the correlation graph to obtain a displacement vector of the pixel point in the first frame of the image; and obtaining multiple optical flow images based on the displacement vector.
[0103] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: calculating based on the correlation graph to obtain the displacement vector of the pixel point in the first frame image, including: calculating the correlation graph to obtain a probability distribution, wherein the probability distribution represents the matching probability between the pixel point in the first frame image and the pixel point in the second frame image; calculating the optical flow component in the horizontal direction and the optical flow component in the vertical direction of the pixel point in the first frame image based on the probability distribution; and obtaining the displacement vector based on the optical flow component in the horizontal direction and the optical flow component in the vertical direction.
[0104] It can be understood by those skilled in the art that Figure 6 The structure shown is for illustration only, and the electronic device may also be a terminal device such as a smart phone, a tablet computer, a PDA, a mobile Internet device (MID), or a PAD. Figure 6 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 6 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 6 Different configurations shown.
[0105] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0106] Example 4
[0107] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the video data processing method provided in the first embodiment.
[0108] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0109] The present application also provides a computer program product, which, when executed on a data processing device, is suitable for executing the program steps of the method for processing video data.
[0110] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0111] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0112] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0113] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0114] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0115] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0116] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for processing video data, characterized in that: include: Acquire a target video to be detected, wherein the target video is a video of a process of providing financial services to a target object; Extracting static features of the target video using a first feature extraction model in a two-stream neural network model to obtain first feature information; Extracting dynamic features of the target video using a second feature extraction model in the dual-stream neural network model to obtain second feature information; The target video is quality-checked based on the first feature information and the second feature information to obtain a quality-check result of the target video, wherein the quality-check result is used to indicate whether the process of providing financial services to the target object meets preset requirements.
2. The method according to claim 1, characterized in that Extracting static features of the target video using the first feature extraction model in the dual-stream neural network model to obtain first feature information includes: Decomposing the target video according to a preset frame rate to obtain a plurality of initial video frames; Screening the multiple initial video frames to obtain multiple target video frames; Performing feature extraction on the multiple target video frames to obtain initial feature information corresponding to each target video frame; The first feature information is obtained according to the initial feature information.
3. The method according to claim 1, characterized in that The second feature information obtained by extracting the dynamic features of the target video through the second feature extraction model in the dual-stream neural network model includes: Decomposing the target video to obtain multiple target video frames; Calculating the difference between two adjacent frames of the target video frames to obtain a plurality of optical flow images; Feature extraction is performed on the multiple optical flow images to obtain the second feature information.
4. The method according to claim 3, characterized in that Calculating the difference between two adjacent frames of the target video frames to obtain a plurality of optical flow images includes: Filtering pixel points in a first frame of the two frames of images according to the brightness of the pixel points to obtain a plurality of first feature points; Acquire a plurality of second feature points corresponding to the plurality of first feature points in a second frame image of the two frames of images; Optical flow images corresponding to the two frames of images are obtained based on the multiple first feature points and the multiple second feature points.
5. The method according to claim 4, characterized in that Obtaining optical flow images corresponding to the two frames of images based on the plurality of first feature points and the plurality of second feature points includes: For each first feature point, calculating a displacement vector between the first feature point and the corresponding second feature point to obtain a plurality of displacement vectors; generating a two-dimensional optical flow field according to the plurality of displacement vectors; The two-dimensional optical flow field is transformed to obtain the optical flow image.
6. The method according to claim 3, characterized in that Calculating the difference between two adjacent frames of the target video frames to obtain a plurality of optical flow images includes: Performing feature extraction on the two frames of images through the optical flow generation layer in the two-stream neural network model to obtain feature information corresponding to the first frame of the two frames of images and feature information corresponding to the second frame of the two frames of images; performing correlation calculation on feature information corresponding to the first frame image and feature information corresponding to the second frame image to obtain a correlation graph; Calculating according to the correlation graph to obtain a displacement vector of a pixel point in the first frame of image; The multiple optical flow images are obtained according to the displacement vector.
7. The method according to claim 6, characterized in that Calculating according to the correlation graph to obtain the displacement vector of the pixel point in the first frame image includes: Calculating the correlation graph to obtain a probability distribution, wherein the probability distribution represents a matching probability between a pixel point in the first frame image and a pixel point in the second frame image; Calculating, according to the probability distribution, a horizontal optical flow component and a vertical optical flow component of a pixel point in the first frame of image; The displacement vector is obtained according to the optical flow component in the horizontal direction and the optical flow component in the vertical direction.
8. A video data processing device, characterized in that: include: an acquisition unit, configured to acquire a target video to be detected, wherein the target video is a video of a process of providing financial services to a target object; A first extraction unit is configured to extract static features of the target video using a first feature extraction model in a two-stream neural network model to obtain first feature information; A second extraction unit is configured to extract dynamic features of the target video using a second feature extraction model in the dual-stream neural network model to obtain second feature information; A quality inspection unit is used to perform quality inspection on the target video based on the first feature information and the second feature information to obtain a quality inspection result of the target video, wherein the quality inspection result is used to indicate whether the process of providing financial services to the target object meets preset requirements.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the video data processing method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: a memory storing an executable program; A processor is used to run the program, wherein the program executes the video data processing method according to any one of claims 1 to 7 when running.