Video processing method and device based on large model and electronic equipment
Through a large-model-based video processing method, machine learning models are used to detect and analyze target objects and their correlations in video frames, which solves the security and compliance risks in dual-recording video scenarios, realizes real-time monitoring and risk warnings for users and business personnel, and improves the compliance and security of the business processing process.
Patent Information
- Application Number
- CN202511264429.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing dual-recording video scenarios lack the identity recognition and monitoring of users and business personnel, and are unable to analyze behaviors such as abnormal movements and voices in real time, resulting in greater security and compliance risks and making it difficult to meet monitoring and risk warning needs in complex scenarios.
A large-model-based video processing method is adopted to detect target objects in video frames through machine learning models, calculate the correlation between objects and audio signals, generate risk values to evaluate the operation probability of preset objects, and combine biometrics and semantic analysis to improve compliance and security.
It enables real-time monitoring and risk assessment of target objects in dual-recording video scenarios, improves compliance and user safety, and can identify potential risk behaviors and conduct intelligent analysis.
Smart Images

Figure CN120808240A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of large model, in particular to a video processing method and device based on large model and electronic equipment. BACKGROUND
[0002] With the increasing requirements for compliance and business authenticity in the fields of finance, government affairs, etc., business institutions usually need to perform double recording operations on business personnel and users handling business to ensure compliance and user security. However, in the current double recording scheme, only the video in the business handling scene of the user and the business personnel is usually collected, but the process of the business handling scene is not monitored and constrained in any way.
[0003] In other words, the prior art lacks the ability to identify and monitor the identities of users and business personnel and to analyze behaviors such as abnormal actions and voices in real time, and cannot identify potential risky behaviors or abnormal situations. This technical limitation makes it difficult for the double recording system to monitor the behavior of the operator and to issue risk warnings when faced with complex scenarios.
[0004] To solve the above problems, a video processing scheme based on large model is needed to solve the defects that there are still great security risks and compliance risks in the double recording video scene in the prior art. SUMMARY
[0005] The embodiments of the present application provide a video processing method and device based on large model and electronic equipment to solve the defects that there are still great security risks and compliance risks in the double recording video scene in the prior art.
[0006] To achieve the above technical purpose, the embodiments of the present application propose a video processing method based on large model, comprising: acquiring a first video frame of a target video stream; detecting a target object in the first video frame using a first machine learning model; When the number of detected target objects is two or more, the image in a predetermined range is acquired as the first image of the target object with the position of each detected target object as the center; Based on the object information of the target object in the first image of each target object, the first correlation between the target object and the corresponding target object is calculated; According to the audio signal corresponding to the first video frame, the second correlation between the audio signal and each target object is calculated; Based on the first correlation and the second correlation, the risk value of the first video frame is calculated, wherein the risk value indicates the probability that the first video frame contains a preset object and the preset object performs a predetermined operation.
[0007] In the video processing method, the step of obtaining the first video frame of the target video stream can include: obtaining a video stream from a camera of a target device; decoding the obtained video stream, and extracting a predetermined number of image frames from the decoded image frames as the first image frames.
[0008] In the video processing method, the step of detecting the target object in the first video frame using the first machine learning model can include: inputting each first image frame into the first machine learning model, and calling a corresponding weight according to image frame information of the first image frame; calculating, using the first machine learning model, a similarity between the target object in the first image frame and a preset object category; determining a category of the target object according to the similarity, and calculating a position of the determined target object.
[0009] In the video processing method, the method can further include: when the category of the target object is a human, extracting a biological feature of the target object; calculating a similarity between the extracted biological feature and a preset biological feature representing preset identity information; determining the identity information corresponding to the preset biological feature with the highest similarity as the identity information of the target object.
[0010] In the video processing method, the step of calculating the risk value of the first video frame based on the first correlation and the second correlation can include: based on the first correlation, determining a target object with the highest first correlation as a target object corresponding to the target object; based on the second correlation, determining an audio signal with the highest second correlation as a candidate audio corresponding to the target object, and calculating a semantic feature of the candidate audio; based on the semantic feature and an object feature of the target object, using a second machine learning model to encode the semantic feature and the object feature to generate an encoded semantic feature vector representing a semantic relationship between each candidate audio and the target object; using a fully connected layer to calculate, as the risk value, a probability that an operation intention of the target object to the target object under the audio signal belongs to at least one of at least two preset operations.
[0011] In the video processing method, the preset operation includes one of contract signing, reading a file, and introducing a product.
[0012] The embodiment of the present application also provides a video processing device based on a large model, comprising: an acquisition module, which is used for acquiring a first video frame of a target video stream; a target object detection module, which is used for detecting target objects in the first video frame by using a first machine learning model, and acquiring images in a predetermined range as first images of the target objects respectively, with the positions of the detected target objects as centers when the number of the detected target objects is two or more than two; a first correlation calculation module, which is used for calculating first correlations between target objects and corresponding target objects based on object information of the target objects in the first images of the target objects; a second correlation calculation module, which is used for calculating second correlations between audio signals corresponding to the first video frame and the target objects according to the audio signals; a risk value calculation module, which is used for calculating a risk value of the first video frame based on the first correlations and the second correlations, wherein the risk value indicates a probability that the first video frame contains a preset object and the preset object performs a predetermined operation.
[0013] The embodiment of the present application also provides an electronic device, comprising: a memory, which is used for storing a program; a processor, which is used for running the program stored in the memory to execute the video processing method based on a large model according to the embodiment of the present application.
[0014] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program executable by a processor, wherein the program is executed by the processor to implement the video processing method based on a large model provided by the embodiment of the present application.
[0015] According to the method and device for processing video based on a large model and the electronic device, the first machine learning model is used to detect a target object in a first video frame in a video stream obtained; when the number of detected target objects is two or more, images in a predetermined range centered on the positions of the detected target objects are obtained as first images of the target objects; based on object information of a target object in the first image of each target object, a first correlation between the target object and the corresponding target object is calculated; based on an audio signal corresponding to the first video frame, a second correlation between the audio signal and each target object is calculated; and based on the first correlation and the second correlation, a risk value of the first video frame is calculated. Therefore, the video processing scheme provided in the embodiments of the present application can detect whether a target object exists in the collected video frame during the dual recording process, and further detect a target object in a predetermined range centered on the position of each detected target object, so that the correlation between the detected target object and the target object in the range can be calculated based on the detected target object, and the audio signal is further collected to calculate the correlation between the audio signal and the target object, so that the risk value of the first video frame can be calculated based on the correlations, that is, whether the first video frame contains a preset object and the preset object performs a predetermined operation. Therefore, not only the recognition result of the object contained in the video frame can be considered, but also the object around the object in the video frame and the audio during video shooting can be fully considered to determine whether the collected video represents that the predetermined object performs the predetermined operation on the predetermined object, so that the compliance and safety of the user in the dual recording scene can be improved.
[0016] The above description is only a summary of the technical solutions of the present application. In order to make the technical means of the present application more clear, the content of the specification can be implemented, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0017] Various other advantages and benefits will become apparent to those of ordinary skill in the art upon reading the following detailed description of the preferred embodiments. The accompanying drawings are included to provide a description of the preferred embodiments, and are not considered limiting of the present application. Moreover, like reference numerals in the attached figures are intended to represent the same parts throughout the various drawings. In the drawings: Figure 1 Flow chart of an embodiment of the method for processing video based on a large model provided by the present application; Figure 2 Structure schematic diagram of an embodiment of the device for processing video based on a large model provided by the present application; Figure 3 Structure schematic diagram of an embodiment of the electronic device provided by the present application. DETAILED DESCRIPTION
[0018] Exemplary embodiments of the present application will be described in greater detail below with reference to the accompanying drawings. While exemplary embodiments of the present application are shown in the drawings, it is understood that the present application can be embodied in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that this application will be thoroughly and completely understood, and will fully convey the scope of the application to those skilled in the art.
[0019] Embodiment One The existing dual-recording system is mainly applied to financial, government and other business scenarios, and records the business handling process through audio and video recording to ensure the compliance and authenticity of the signing link. However, the traditional dual-recording system has a single function and can only realize basic audio and video recording, which cannot meet the higher requirements of user identity verification, dialogue content understanding, abnormal behavior detection and compliance review in complex business scenarios. In actual application, the traditional system cannot effectively lock the real identity and intention of the operator, and cannot intelligently analyze and review the key links in the signing process in real time, which makes the authenticity and effectiveness of the business handling process exist certain risks, and is difficult to meet the increasingly strict regulatory requirements.
[0020] With the increasing requirements of compliance and business authenticity in the fields of finance, government and others, the dual-recording system needs to be further upgraded to realize higher degree of intelligence and automation. In the aspect of user identity recognition, the existing technology usually relies on simple verification methods such as identity card scanning or static face comparison, which is difficult to prevent the risk of fake identity or using others' information; in the aspect of dialogue content understanding, the traditional system only records the voice simply, and cannot analyze the semantic of voice content or judge the authenticity of user intention; in the aspect of abnormal behavior detection, it lacks real-time analysis capability of user expression, action or voice abnormality, and cannot identify potential risks in time; in the aspect of compliance review, the system cannot automatically match relevant regulatory requirements according to business scenarios, resulting in low efficiency of review in business handling process, and easy to miss or make mistakes.
[0021] Therefore, the embodiments of the present application disclose a video processing method based on a large model, Figure 1 is a flowchart showing a video processing method based on a large model according to an embodiment of the present application. As shown in Figure 1 The video processing method based on a large model according to the embodiments of the present application can include: S101, acquiring a first video frame of a target video stream.
[0022] In step S101, a first video frame of a target video stream can be acquired. For example, in the embodiments of the present application, a dual-recording video can be collected by a camera of a terminal device used by a service personnel when a user handles a related service, so as to obtain a target video stream. In the embodiments of the present application, the target video stream can be a video stream collected by a predetermined device. For example, whether it is a video stream collected by a predetermined device can be determined by detecting a device identifier contained in metadata of the video stream, and when the device identifier is consistent with a device in a preset device list, the collected video stream can be taken as a target video stream, and a video frame of the target video stream is extracted as a first video frame in step S101. For example, a video stream can be acquired from a camera of a target device; the acquired video stream is decoded, and a predetermined number of image frames are extracted from the image frames obtained after decoding as first image frames. That is, a video decoder, such as FFmpeg or OpenCV, can be used to decode the video stream collected by the terminal, and compressed encoded data is converted into original image frames. Then, the first frame image with a timestamp of zero or closest to zero is extracted from the original image frames, and more video frames continuous in time with the first frame image can be further extracted to be taken as the first image frames together with the first frame image. In addition, in the embodiments of the present application, the extracted first frame image can also be converted into a standard image format (such as RGB or BGR) for subsequent processing.
[0023] S102, detecting a target object in the first video frame using a first machine learning model.
[0024] In step S102, the first machine learning model pre-trained can be used to detect the target object in the first video frame obtained in step S101. For example, in the embodiment of the present application, when the business personnel performs double recording according to the business to be handled by the user, the video collected can be detected in step S102 to determine whether the video contains the predetermined target object, such as the designated business personnel and the user handling the business. For example, in the embodiment of the present application, the first image frame obtained in each step S101 can be input to the first machine learning model, and the corresponding weight can be called according to the image frame information of the first image frame. For example, the first frame image with a timestamp of zero or closest to zero can be given a smaller weight, and the image frame at a subsequent predetermined time interval, such as 3 seconds or 5 seconds, can be given a larger weight, so as to take into account the preparation time after the business personnel starts the shooting device. Then, the first machine learning model can be used to calculate the similarity of the target object in the first image frame to the preset object category, determine the category of the target object according to the similarity, and calculate the position of the determined target object. For example, in step S102, the preset object category can include human beings, so that it can be detected in step S102 whether the video frame contains a person. And when the category of the target object is human, the biological features of the target object are extracted. For example, the biological features such as the face image or the iris of the target object can be extracted, and the similarity between the extracted biological features and the preset biological features representing the preset identity information is calculated; the identity information corresponding to the preset biological features with the highest similarity is determined as the identity information of the target object. Therefore, in step S102, when it is detected that the target image frame contains a person, the biological feature matching process can be further performed on the contained person to determine whether the person contained in the video frame is the business personnel and the user. In particular, in the embodiment of the present application, the determination rule can be pre-set. For example, the identity information and the biological feature information of the business personnel can be pre-input into the database in association with each other, and when the user handles the business, the biological features and the identity information of the user can also be input into the database in association with each other before step S102. Therefore, in step S102, when it is detected that there is a human in the first video frame, feature extraction can be further performed on the human object region detected from the first image frame to compare with the biological features stored in the database, and the identity information corresponding to the matched biological features is taken as the identity information of the detected person, so that it can be confirmed that there are business personnel and users handling the business in the first video frame.
[0025] S103, when the number of detected target objects is two or more, the image in a predetermined range centered on the position of each target object is obtained as the first image of the target object.
[0026] In step S103, when the number of target objects detected in step S102 is two or more, for example, when it is determined that the service personnel and the user are detected in step S102, the image of a predetermined range centered on the position of the target object detected in step S102 can be acquired as the surrounding image corresponding to the target object in step S103. For example, the image can be acquired with a circle of a predetermined radius as the predetermined range, with the coordinate indicated in the position information of the target object as the center.
[0027] In step S104, the first correlation between the target object and the corresponding target object is calculated based on the object information of the target object in the first image of each target object.
[0028] In step S104, the image around the target object detected in step S102 can be detected based on the surrounding image determined in step S103, and the object can be detected in the region, and the detected object is taken as the target object, so that the correlation between the target object and the target object can be calculated based on the object information of such target object. For example, the features in the image of the surrounding area of the target object determined in step S103, such as the service personnel, are extracted, and the similarity between the features and the predetermined object is calculated, so that it can be determined whether there is an object in the surrounding area. For example, it can be detected in step S103 that there are service personnel and users, so that in step S103, the position coordinates of the service personnel and the user (such as the center of gravity) can be taken as the center, and a predetermined range can be determined with a predetermined radius, so that it can be detected whether there is a target object, such as a contract or a service brochure, around the service personnel. Then in step S104, the correlation between the target object thus detected and the target object determined in step S102, or the identity information determined in step S103, can be calculated. For example, it can be detected in step S103 that there are two objects, a contract and a mirror, around the service personnel, so that in step S104, the correlation with the target object determined in step S102 can be calculated according to the object information of the two objects. In particular, in the embodiments of the present application, the correlation here can represent the correlation between the object and the target object in the predetermined business. The predetermined business here can be input by the service personnel in advance before step S101, or can be calculated by the server according to the collected video information.
[0029] In step S105, the second correlation between the audio signal corresponding to the first video frame and each target object is calculated according to the audio signal.
[0030] In step S105, an audio signal corresponding to the first video frame can be further obtained from the target video stream, and a second correlation between the audio signal and the target object determined in step S102 can be calculated. Similar to the first correlation, the second correlation can also represent the correlation between the audio and the target object. For example, in step S105, whether the audio existing at the time of the first video frame or the time of the video frames before and after the first video frame is the audio related to the service personnel or the user can be determined. For example, in step S105, based on the first correlation obtained in step S102, the target object with the highest first correlation can be determined as the target object corresponding to the target object; based on the second correlation, the audio signal with the highest second correlation can be determined as the candidate audio corresponding to the target object, and the semantic feature of the candidate audio can be calculated; based on the semantic feature and the object feature of the target object, the semantic feature and the object feature can be encoded using the second machine learning model to generate an encoded semantic feature vector representing the semantic relationship between each candidate audio and the target object; and the probability that the operation intention of the target object to the target object under the audio signal belongs to at least one of the at least two preset operations can be calculated using the fully connected layer as the risk value.
[0031] In step S106, based on the first correlation and the second correlation, the risk value of the first video frame can be calculated.
[0032] In step S106, based on the two correlations calculated as above, the risk value of the video frame can be further calculated. In the embodiments of the present application, the risk value can indicate the probability that the first video frame contains the preset object and the preset object performs the predetermined operation. For example, based on the correlation calculated as above, the identity information of the object and the target object can be combined to calculate the predetermined object and the correlation determined in step S105, so as to determine whether the service personnel is using the device terminal to record, or whether the service personnel is making an extensive business propaganda to the user, and the like.
[0033] According to the video processing method based on a large model provided in the embodiments of the present application, the target objects in the first video frame of the acquired video stream are detected by using the first machine learning model; when the number of the detected target objects is two or more, the images in a predetermined range centered on the positions of the detected target objects are acquired as the first images of the target objects respectively; the first correlation between the target object and the corresponding target object is calculated based on the object information of the target object in the first image of the target object; the second correlation between the audio signal corresponding to the first video frame and the target object is calculated according to the audio signal; and the risk value of the first video frame is calculated based on the first correlation and the second correlation. Therefore, the video processing scheme provided in the embodiments of the present application can detect whether there is a target object in the acquired video frame in the dual recording process, and further detect the target object in a predetermined range centered on the position of each detected target object, so as to calculate the correlation between the detected target object and the target object in the range, and further acquire the audio signal to calculate the correlation between the audio signal and the target object, so as to calculate whether the video frame indicates that the first video frame contains a preset object and the preset object performs a predetermined operation based on the correlations. Therefore, not only the recognition result of the object contained in the video frame can be considered, but also the object around the object in the video frame and the audio when the video is shot can be fully considered to determine whether the acquired video represents that the predetermined object performs the predetermined operation on the predetermined object, so as to improve the compliance and the safety of the user in the dual recording scene.
[0034] Embodiment two Figure 2 The structural schematic diagram of one embodiment of the video processing device based on a large model provided in the present application is shown in FIG. 1. Figure 2 As shown in FIG. 1, the video processing device provided in the embodiments of the present application can include an acquisition module 21, a target object detection module 22, a first correlation calculation module 23, a second correlation calculation module 24 and a risk value calculation module 25.
[0035] The acquisition module 21 can be used to acquire the first video frame of the target video stream.
[0036] The acquisition module 21 can acquire a first video frame of a target video stream. For example, in the embodiments of the present application, a dual-recording video can be collected by a camera of a terminal device used by a service personnel when a user handles a related service, so as to obtain a target video stream. In the embodiments of the present application, the target video stream can be a video stream collected by a predetermined device. For example, whether it is a video stream collected by a predetermined device can be determined by detecting a device identifier contained in metadata of the video stream, and when the device identifier is consistent with a device in a preset device list, the collected video stream can be taken as a target video stream, and the acquisition module 21 extracts a video frame of the target video stream as a first video frame. For example, the video stream can be acquired from a camera of a target device; the acquired video stream is decoded, and a predetermined number of image frames are extracted from the image frames obtained after decoding as first image frames. That is, a video decoder, such as FFmpeg or OpenCV, can be used to decode the video stream collected by the terminal, and compressed encoded data is converted into original image frames. Then, the first frame image with a timestamp of zero or closest to zero is extracted from the original image frames, and more video frames continuous in time with the first frame image can be further extracted to be taken as the first image frames together with the first frame image. In addition, in the embodiments of the present application, the extracted first frame image can also be converted into a standard image format (such as RGB or BGR) for subsequent processing.
[0037] The target object detection module 22 can be configured to detect a target object in the first video frame using a first machine learning model.
[0038] The target object detection module 22 can use a pre-trained first machine learning model to detect a target object in the first video frame obtained by the acquisition module 21. For example, in the embodiment of the present application, when the business staff performs double recording according to the business to be handled by the user, the target object detection module 22 can detect the collected video to determine whether the video contains a predetermined target object, such as a designated business staff and a user handling business. For example, in the embodiment of the present application, the first image frame obtained by the acquisition module 21 can be input to the first machine learning model, and the corresponding weight can be called according to the image frame information of the first image frame. For example, a smaller weight can be given to the first frame image with a timestamp of zero or closest to zero, and a larger weight can be given to the image frame at a subsequent predetermined time interval, such as 3 seconds or 5 seconds, to take into account the preparation time after the business staff starts the shooting device. Then, the first machine learning model can be used to calculate the similarity of the target object in the first image frame to the preset object category, determine the category of the target object according to the similarity, and calculate the position of the determined target object. For example, the target object detection module 22 can include a human in the preset object category, so that the target object detection module 22 can detect whether a person is included in the video frame. And when the category of the target object is human, the biological features of the target object are extracted. For example, the biological features such as the face image or the iris of the target object can be extracted, and the similarity between the extracted biological features and the preset biological features representing the preset identity information is calculated; the identity information corresponding to the preset biological features with the highest similarity is determined as the identity information of the target object. Therefore, when the target object detection module 22 detects that the target image frame contains a person, the biological feature matching process of the contained person can be further performed to determine whether the person contained in the video frame is a business staff and a user. In particular, in the embodiment of the present application, the determination rule can be pre-set. For example, the identity information and biological feature information of the business staff can be pre-input into the database in association with each other, and when the user handles the business, the target object detection module 22 can also input the biological features and identity information of the user into the database in association with each other. Therefore, when the target object detection module 22 detects the presence of a human in the first video frame, the feature extraction can be further performed from the human object region detected in the first image frame to compare with the biological features stored in the database, and the identity information corresponding to the matched biological features is taken as the identity information of the detected person, so that it can be confirmed that the business staff and the user handling the business exist in the first video frame. When the number of detected target objects is two or more, the target object detection module 22 takes the position of each detected target object as the center to obtain the image in a predetermined range as the first image of the target object.
[0039] The first correlation calculation module 23 can acquire, as the surrounding image corresponding to the target object, an image of a predetermined range centered on the position of the target object detected by the target object detection module 22 when the number of target objects detected by the target object detection module 22 is two or more, for example, when the target object detection module 22 detects a business person and a user. For example, an image can be acquired with a circle of a predetermined radius as the predetermined range, with the coordinate indicated in the position information of the target object as the center.
[0040] The first correlation calculation module 23 is configured to calculate the first correlation between the target object and the target object based on the object information of the target object in the first image of each target object.
[0041] The first correlation calculation module 23 can detect the image around the target object detected by the target object detection module 22 based on the surrounding image determined by the target object detection module 22, and can detect the object in the area and take the detected object as the target object, so as to calculate the correlation between the target object and the target object based on the object information of such target object. For example, the first correlation calculation module 23 can extract the features in the image of the surrounding area of the target object, such as the business person, determined by the second correlation calculation module 24, and calculate the similarity between the features and the predetermined object, so as to determine whether there is an object in the surrounding area. For example, the first correlation calculation module 23 can detect that there is a business person and a user, so that the first correlation calculation module 23 can take the position coordinate of the business person and the user (for example, the center of gravity) as the center and determine a predetermined range with a predetermined radius, so as to detect whether there is a target object, such as a contract or a business promotion material, around the business person. Then the correlation between the target object detected in this way and the target object determined by the target object detection module 22, or the identity information determined by the first correlation calculation module 23, can be calculated. For example, the first correlation calculation module 23 can detect that there are two objects, a contract and a mirror, around the business person, so the correlation with the target object determined by the target object detection module 22 can be calculated according to the object information of the two objects. In particular, in the embodiments of the present application, the correlation here can represent the correlation between the object and the target object in the predetermined business. The predetermined business here can be input by the business person in advance before the acquisition module 21, or can be calculated by the server according to the collected video information.
[0042] The second correlation calculation module 24 calculates the second correlation between the audio signal corresponding to the first video frame and each target object according to the audio signal.
[0043] The second correlation calculation module 24 can further obtain an audio signal corresponding to the first video frame from the target video stream, and calculate a second correlation between the audio signal and the target object determined by the target object detection module 22. Similar to the first correlation, the second correlation can also represent the correlation between the audio and the target object. For example, in step S105, whether the audio existing at the time of the first video frame or the time of the video frames before and after the first video frame is the audio related to the service personnel or the user. For example, the second correlation calculation module 24 can determine the target object corresponding to the target object based on the first correlation obtained by the target object detection module 22, determine the audio signal with the highest second correlation as the candidate audio corresponding to the target object based on the second correlation, calculate the semantic feature of the candidate audio, encode the semantic feature and the object feature of the target object using the second machine learning model based on the semantic feature and the object feature of the target object to generate an encoded semantic feature vector representing the semantic relationship between each candidate audio and the target object, and calculate the probability that the operation intention of the target object to the target object under the audio signal belongs to at least one of the at least two preset operations using a full connection layer as a risk value.
[0044] The risk value calculation module 25 is configured to calculate the risk value of the first video frame based on the first correlation and the second correlation.
[0045] The risk value calculation module 25 can further calculate the risk value of the video frame based on the two correlations calculated as above. In the embodiments of the present application, the risk value can indicate the probability that the first video frame contains a preset object and the preset object performs a predetermined operation. For example, the predetermined object and the determined correlation can be calculated based on the correlation calculated as above, in combination with the identity information of the object and the target object, so as to determine whether the service personnel is using the device terminal to record, or whether the service personnel is making an extensive business promotion to the user, and the like.
[0046] According to the video processing device based on the large model provided in the embodiments of the present application, the target objects in the first video frame in the acquired video stream are detected by using the first machine learning model; when the number of the detected target objects is two or more, the images in a predetermined range are acquired as the first images of the target objects respectively, with the positions of the detected target objects as the centers; the first correlation between the target object and the corresponding target object is calculated based on the object information of the target object in the first image of each target object; the second correlation between the audio signal corresponding to the first video frame and each target object is calculated according to the audio signal; and the risk value of the first video frame is calculated based on the first correlation and the second correlation. Therefore, the video processing scheme provided in the embodiments of the present application can detect whether there is a target object in the collected video frame in the dual recording process, and further detect the target object in a predetermined range with the position of each detected target object as the center, so that the correlation between the detected target object and the target object in the range can be calculated, and the audio signal is further collected to calculate the correlation between the audio signal and the target object, so that the risk value of the first video frame can be calculated based on the correlations, that is, whether the first video frame contains a preset object and the preset object performs a predetermined operation. Therefore, not only the recognition result of the object contained in the video frame can be considered, but also the object around the object in the video frame and the audio when the video is shot can be fully considered to determine whether the collected video represents that the predetermined object performs the predetermined operation on the predetermined object, so that the compliance in the dual recording scene and the safety of the user can be improved.
[0047] Embodiment three The internal functions and structures of the video processing device are described above, and the device can be implemented as an electronic device. Figure 3 The structural schematic diagram of the electronic device provided in the embodiments of the present application is shown in FIG. 1. Figure 3 As shown in FIG. 1, the electronic device includes a memory 31 and a processor 32.
[0048] The memory 31 is configured to store programs. In addition to the above programs, the memory 31 can also be configured to store various other data to support the operation on the electronic device. Examples of the data include instructions of any application program or method for operating on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.
[0049] The memory 31 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk or an optical disk.
[0050] The processor 32, which is not limited to a central processing unit (CPU), can also be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an embedded neural processing unit (NPU), or an artificial intelligence (AI) chip. The processor 32 is coupled to the memory 31 and executes a program stored in the memory 31, which, when executed, performs the video processing method of Embodiment 1.
[0051] Further, as shown in Figure 3 The electronic device can further include a communication component 33, a power component 34, an audio component 35, a display 36, and other components. Figure 3 Some components are only schematically shown in the electronic device, and it does not mean that the electronic device only includes Figure 3 the components shown.
[0052] The communication component 33 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 3G, 4G, or 5G, or a combination thereof. In an example embodiment, the communication component 33 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 33 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0053] The power component 34 provides power to various components of the electronic device. The power component 34 can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the electronic device.
[0054] The audio component 35 is configured to output and / or input audio signals. For example, the audio component 35 includes a microphone (MIC) configured to receive external audio signals when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 31 or transmitted via the communication component 33. In some embodiments, the audio component 35 also includes a speaker for outputting audio signals.
[0055] The display 36 includes a screen, which can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide and a gesture on the touch panel. The touch sensor can detect not only a boundary of a touching or a sliding action, but also a duration and intensity of the touching or the sliding action.
[0056] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware. The foregoing program can be stored in a computer readable storage medium. The program, when executed, performs steps including the above-mentioned method embodiments; and the foregoing storage medium includes various storage media that can store program codes, such as ROM, RAM, magnetic disk or optical disk.
[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A video processing method based on a large model, characterized in that: include: Get the first video frame of the target video stream; detecting a target object in the first video frame using a first machine learning model; When the number of detected target objects is two or more, taking the position of each detected target object as the center, acquiring an image within a predetermined range as the first image of the target object; calculating a first correlation between the target object and the corresponding target object based on the object information of the target object in the first image of each target object; calculating, based on the audio signal corresponding to the first video frame, a second correlation between the audio signal and each target object; A risk value of the first video frame is calculated based on the first correlation and the second correlation, wherein the risk value indicates a probability that a preset object is included in the first video frame and that the preset object performs a predetermined operation.
2. The video processing method based on a large model according to claim 1, characterized in that The obtaining of the first video frame of the target video stream includes: Get the video stream from the target device's camera; The acquired video stream is decoded, and a predetermined number of image frames are extracted from the image frames obtained after decoding as first image frames.
3. The video processing method based on a large model according to claim 2, characterized in that: The detecting the target object in the first video frame using the first machine learning model includes: Inputting each first image frame into a first machine learning model, and calling corresponding weights according to image frame information of the first image frame; Calculating, using the first machine learning model, a similarity between a target object in the first image frame and a preset object category; The category of the target object is determined according to the similarity, and the position of the determined target object is calculated.
4. The video processing method based on a large model according to claim 3, characterized in that: The method further comprises: When the category of the target object is human, extracting the biometric features of the target object; Calculating the similarity between the extracted biometric feature and a preset biometric feature representing preset identity information; The identity information corresponding to the preset biometric feature with the highest similarity is determined as the identity information of the target object.
5. The video processing method based on a large model according to claim 1, characterized in that: The calculating the risk value of the first video frame based on the first correlation and the second correlation includes: Based on the first correlation, determining the target object with the highest first correlation as the target object corresponding to the target object; Based on the second correlation, determining the audio signal with the highest second correlation as the candidate audio corresponding to the target object, and calculating the semantic features of the candidate audio; Based on the semantic features and the object features of the target object, feature encode the semantic features and the object features using a second machine learning model to generate an encoded semantic feature vector representing a semantic relationship between each candidate audio and the target object; A fully connected layer is used to calculate the probability that the target object's operation intention on the target object under the audio signal belongs to at least one of at least two preset operations for the encoded semantic feature vector, as the risk value.
6. The video processing method based on a large model according to claim 5, characterized in that: The preset operation includes one of contract signing, document reading, and product introduction.
7. A video processing device based on a large model, characterized in that: include: An acquisition module, configured to acquire a first video frame of a target video stream; a target object detection module, configured to detect target objects in the first video frame using a first machine learning model, and when two or more target objects are detected, obtain images within a predetermined range centered on the position of each detected target object as first images of the target object; a first correlation calculation module, configured to calculate a first correlation between a target object and a corresponding target object based on object information of the target object in the first image of each target object; a second correlation calculation module, configured to calculate a second correlation between the audio signal and each target object based on the audio signal corresponding to the first video frame; The risk value calculation module is configured to calculate a risk value of the first video frame based on the first correlation and the second correlation, wherein the risk value indicates a probability that a preset object is included in the first video frame and that the preset object performs a predetermined operation.
8. An electronic device, characterized in that: include: Memory, used to store programs; A processor, configured to run the program stored in the memory to execute the large model-based video processing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Target association method and device, electronic equipment and machine readable storage medium
CN112949538A
Method and device for detecting compliance of double-recording video file and medium
CN115858857A
Double-recording real-time risk identification method and device based on AI, equipment and medium
CN118135498A
Abnormity detection method, device and equipment for double-recording video and storage medium
CN119251900A
Method and device for detecting interaction event in recorded video
CN119942419A