Artificial Intelligence-Based Audio and Video Data Processing Method and System

By judging the privacy of audio and video production and comparing it with standard databases, comparing different audio and videos is solved, and efficient data storage and management are achieved.

CN116546259BActive Publication Date: 2025-08-05GUANGZHOU AIPILI INFORMATION TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310530524.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-11
Publication Date
2025-08-05
Estimated Expiration
2043-05-11

AI Technical Summary

Technical Problem

The prior art is difficult to realize the privacy data management and efficient data management of individual users in the field of audio and video.

Method used

By obtaining the operation data and action data of the audio and video management subject, we can judge whether the audio and video production is private audio and video. If so, we will compare it with the standard privacy database to generate a comparison and different audio and video, and decide whether to store or delete the data based on the comparison results. We will use the 3DCNN network for feature extraction and comparison to achieve efficient data management.

Benefits of technology

It realizes privacy management during audio and video production, while avoiding duplicate storage, improving the efficiency and privacy of data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116546259B_ABST
    Figure CN116546259B_ABST
Patent Text Reader

Abstract

This application relates to an audio-visual data processing method and system based on artificial intelligence, including obtaining management subject action data and obtaining current actual audio-visual data; determining whether the audio-visual production of the current audio-visual management subject is private audio-visual production. If it is determined that it is not private audio-visual production, the current actual audio-visual data is stored in a normal audio-visual repository. If it is determined that it is private audio-visual production, the current actual audio-visual data is compared with standard private audio-visual data based on a preset intelligent comparison module, and after the comparison is completed, comparison difference audio-visuals are generated respectively. According to each comparison difference audio-visual, it is determined whether the current actual audio-visual data is a stored video. If it is determined to be no, the current actual audio-visual data is stored. The present invention realizes the privacy during the audio-visual production process and at the same time can realize that the stored video is no longer stored intelligently, so as to realize efficient data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio and video processing, and particularly to an audio and video data processing method and system based on artificial intelligence. Background Art

[0002] Audio and video is a multimedia communication technology composed of two parts, audio and video, which can achieve interactive transmission and playback of audio and video.

[0003] Currently, artificial intelligence is gradually applied to the field of audio and video. For example, the invention patent with the publication number CN111787418A discloses a docking processing method for audio and video streams based on artificial intelligence AI, including: receiving an address acquisition request sent by a control platform; invoking a load balancing interface to determine the address of a target server that is currently in an idle state from multiple servers corresponding to an audio and video processing platform; sending the address to the control platform; receiving the URL address of an RTMP stream sent by the control platform; sending a screenshot and stream interception instruction to the control platform, where the screenshot and stream interception instruction is used to instruct the control platform to perform picture interception and audio and video file interception on the real-time RTMP stream of the client side indicated by the URL address, and send the intercepted target picture and target audio and video files to the target server.

[0004] Although the above patent document can be applied to the scenarios of smart government / smart community, thus promoting the construction of a smart city, it still has defects for individual users. For example, it obviously cannot implement the privacy data management and efficient management of individual users. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide an audio and video data processing method and system based on artificial intelligence that can achieve privacy and efficient data management.

[0006] The technical solution of the present invention is as follows:

[0007] An audio and video data processing method based on artificial intelligence, the method includes:

[0008] Obtain the audio-video startup operation data and management entity action data of the current audio-video management entity when starting audio-video production, and obtain the current actual audio-video data of the current audio-video management entity after the audio-video production ends; determine whether the audio-video production of the current audio-video management entity is private audio-video production according to the audio-video startup operation data and management entity action data. If it is determined that it is not private audio-video production, store the current actual audio-video data in a normal audio-video repository, where the normal audio-video repository is preset in advance; if it is determined that it is private audio-video production, based on a preset intelligent comparison module, compare the current actual audio-video data with the standard private audio-video data in a pre-stored standard private database, and generate comparison difference audio-video respectively after the comparison is completed. One current actual audio-video data and one standard private audio-video data are compared to generate one comparison difference audio-video respectively. The standard private database and the standard private audio-video data are both preset in advance; determine whether the current actual audio-video data is a stored video according to each comparison difference audio-video. If it is determined to be yes, delete the current actual audio-video data; if it is determined to be no, store the current actual audio-video data.

[0009] Specifically, determine whether the current actual audio-video data is a stored video according to each comparison difference audio-video. If it is determined to be yes, delete the current actual audio-video data; if it is determined to be no, store the current actual audio-video data. Specifically, it includes:

[0010] Based on each of the compared differential audio - videos, determine the actual video capacity of each of the compared differential audio - videos, aggregate the actual video capacities of each and generate the total capacity of the differential videos; determine whether the total capacity of the differential videos is greater than or equal to the standard video capacity. If it is determined that the total capacity of the differential videos is less than the standard video capacity, then determine that the current actual audio - video data is the stored video and delete the current actual audio - video data; if it is determined that the total capacity of the differential videos is greater than or equal to the standard video capacity, then determine that the current actual audio - video data is not the stored video, and mark the videos in the current actual audio - video data other than the compared differential audio - videos as actually identical audio - videos; extract the video connection nodes of the actually identical audio - videos and the compared differential audio - videos, where the number of the video connection nodes is multiple; establish the video association data of the actually identical audio - videos and the compared differential audio - videos according to the video connection nodes; compress the video association data and generate the associated compressed data; crop the current actual audio - video data, and crop out the actually identical audio - videos from the current actual audio - video data, bind the cropped actually identical audio - videos with the associated compressed data, and generate the cloud - stored audio - video data after binding; store the cloud - stored audio - video data in a preset cloud space; store the compared differential audio - videos remaining after cropping out the actually identical audio - videos from the current actual audio - video data in the standard privacy database.

[0011] Specifically, based on the audio - video start operation data and the management subject action data, determine whether the audio - video production of the current audio - video management subject is privacy audio - video production. If it is determined that it is not privacy audio - video production, then store the current actual audio - video data in a normal audio - video repository, where the normal audio - video repository is pre - set; specifically including:

[0012] Extract operation actions according to the audio - video start operation data, and obtain the actual trigger nodes of the current audio - video management entity after the operation action extraction, where the number of the actual trigger nodes is multiple; obtain the actual trigger time points of each of the actual trigger nodes, connect the actual trigger nodes in a curve according to each of the actual trigger time points, and generate a virtual operation curve after the curve connection; determine whether the virtual operation curve matches the privacy trigger curve. If it is determined that the virtual operation curve matches the privacy trigger curve, then determine that the audio - video production of the current audio - video management entity is privacy audio - video production; if it is determined that the virtual operation curve does not match the privacy trigger curve, then extract actions from the management entity action data, and generate actual operation actions after the action extraction, where the number of the actual operation actions is multiple; combine each of the actual operation actions and generate a combined operation action; determine whether the combined operation action matches the privacy operation action. If it is determined that the combined operation action matches the privacy operation action, then determine that the audio - video production of the current audio - video management entity is privacy audio - video production; if it is determined that the combined operation action does not match the privacy operation action, then determine that the audio - video production of the current audio - video management entity is not privacy audio - video production, and store the current actual audio - video data in a pre - set normal audio - video storage repository.

[0013] Specifically, if it is determined that it is privacy audio - video production, then compare the current actual audio - video data with the standard privacy audio - video data in a pre - stored standard privacy database based on a pre - set intelligent comparison module, and generate comparison - difference audio - videos respectively after the comparison is completed. Specifically, it includes:

[0014] If it is determined that it is privacy audio - video production, then split the current actual audio - video data and the standard privacy audio - video data in the pre - stored standard privacy database respectively. Among them, after the current actual audio - video data is split, multiple actual refined audio - videos are generated, and after the standard privacy audio - video data is split, multiple standard refined audio - videos are generated, and each standard privacy audio - video data corresponds to multiple standard privacy refinement data; input each of the actual refined audio - videos into the 3DCNN network in the intelligent comparison module, and obtain the actual feature representation output by the 3DCNN network, where one actual refined audio - video corresponds to one actual feature representation; input the standard privacy refinement data corresponding to the standard privacy audio - video data into the 3DCNN network in the intelligent comparison module, and obtain the standard feature representation output by the 3DCNN network, where one standard privacy refinement data corresponds to one standard feature representation; compare the actual feature representation and the standard feature representation, and generate comparison - difference audio - videos respectively after the comparison is completed.

[0015] Specifically, an audio-visual data processing system based on artificial intelligence, the system comprising:

[0016] A current data acquisition module, configured to acquire audio-visual start operation data and management entity action data of the current audio-visual management entity when starting audio-visual production, and acquire the current actual audio-visual data of the current audio-visual management entity after the audio-visual production ends;

[0017] A data judgment and storage module, configured to judge whether the audio-visual production of the current audio-visual management entity is private audio-visual production according to the audio-visual start operation data and the management entity action data, and if it is judged that it is not private audio-visual production, store the current actual audio-visual data in a normal audio-visual storage repository, wherein the normal audio-visual storage repository is pre-set;

[0018] An intelligent data processing module, configured to, if it is judged that it is private audio-visual production, compare the current actual audio-visual data with the standard private audio-visual data in a pre-stored standard private database based on a pre-set intelligent comparison module, and generate comparison difference audio-visuals respectively after the comparison is completed, wherein one comparison difference audio-visual is generated respectively after one current actual audio-visual data is compared with one standard private audio-visual data, and the standard private database and the standard private audio-visual data are both pre-set;

[0019] A storage judgment and processing module, configured to judge whether the current actual audio-visual data is a stored video according to each of the comparison difference audio-visuals, and if it is judged to be yes, delete the current actual audio-visual data; if it is judged to be no, store the current actual audio-visual data.

[0020] Specifically, the storage judgment and processing module is further configured to:

[0021] Judge the actual video capacity of each of the compared differential audio-visual videos, summarize the actual video capacities of each of the compared differential audio-visual videos and generate the total capacity of the differential videos; judge whether the total capacity of the differential videos is greater than or equal to the standard video capacity. If it is judged that the total capacity of the differential videos is less than the standard video capacity, then judge that the current actual audio-visual data is the stored video and delete the current actual audio-visual data; if it is judged that the total capacity of the differential videos is greater than or equal to the standard video capacity, then judge that the current actual audio-visual data is not the stored video, and mark the videos in the current actual audio-visual data except the compared differential audio-visual videos as actually identical audio-visual videos; extract the video connection nodes of the actually identical audio-visual videos and the compared differential audio-visual videos, wherein the number of the video connection nodes is multiple; establish the video association data of the actually identical audio-visual videos and the compared differential audio-visual videos according to the video connection nodes; compress the video association data and generate the associated compressed data; crop the current actual audio-visual data, and crop out the actually identical audio-visual videos in the current actual audio-visual data, bind the cropped actually identical audio-visual videos with the associated compressed data, and generate the cloud storage audio-visual data after binding; store the cloud storage audio-visual data in a preset cloud space; store the compared differential audio-visual videos remaining after cropping out the actually identical audio-visual videos from the current actual audio-visual data in the standard privacy database.

[0022] Specifically, the data judgment and storage module is further configured to:

[0023] Extract operation actions according to the audio-visual start operation data, and obtain the actual trigger nodes of the current audio-visual management entity after the operation action extraction, where the number of the actual trigger nodes is multiple; obtain the actual trigger time points of each of the actual trigger nodes, and connect the actual trigger nodes according to each of the actual trigger time points to generate a virtual operation curve after the curve connection; determine whether the virtual operation curve matches the privacy trigger curve. If it is determined that the virtual operation curve matches the privacy trigger curve, determine that the audio-visual production of the current audio-visual management entity is privacy audio-visual production; if it is determined that the virtual operation curve does not match the privacy trigger curve, extract actions from the management entity action data, and generate actual operation actions after the action extraction, where the number of the actual operation actions is multiple; combine each of the actual operation actions to generate a combined operation action; determine whether the combined operation action matches the privacy operation action. If it is determined that the combined operation action matches the privacy operation action, determine that the audio-visual production of the current audio-visual management entity is privacy audio-visual production; if it is determined that the combined operation action does not match the privacy operation action, determine that the audio-visual production of the current audio-visual management entity is not privacy audio-visual production, and store the current actual audio-visual data in a pre-set normal audio-visual repository.

[0024] Specifically, the intelligent data processing module is further configured to:

[0025] If it is determined that it is privacy audio-visual production, split the current actual audio-visual data and the standard privacy audio-visual data in the pre-stored standard privacy database respectively. After the current actual audio-visual data is split, multiple actual refined audio-visuals are generated. After the standard privacy audio-visual data is split, multiple standard refined audio-visuals are generated, and each standard privacy audio-visual data corresponds to multiple standard privacy refinement data; input each of the actual refined audio-visuals into the 3DCNN network in the intelligent comparison module, and obtain the actual feature representation output by the 3DCNN network, where one actual refined audio-visual corresponds to one actual feature representation; input the standard privacy refinement data corresponding to the standard privacy audio-visual data into the 3DCNN network in the intelligent comparison module, and obtain the standard feature representation output by the 3DCNN network, where one standard privacy refinement data corresponds to one standard feature representation; compare the actual feature representation and the standard feature representation, and generate comparison difference audio-visuals respectively after the comparison is completed.

[0026] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned audio-visual data processing method based on artificial intelligence are implemented.

[0027] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the above-mentioned audio and video data processing method based on artificial intelligence are implemented.

[0028] The technical effects achieved by the present invention are as follows:

[0029] The above-mentioned audio and video data processing method and system based on artificial intelligence sequentially obtain the audio and video start operation data and management subject action data of the current audio and video management subject when starting audio and video production, and obtain the current actual audio and video data of the current audio and video management subject after the end of audio and video production; judge whether the audio and video production of the current audio and video management subject is private audio and video production according to the audio and video start operation data and management subject action data. If it is judged that it is not private audio and video production, store the current actual audio and video data in a normal audio and video repository, where the normal audio and video repository is preset in advance; if it is judged that it is private audio and video production, compare the current actual audio and video data with the standard private audio and video data in the pre-stored standard private database based on a preset intelligent comparison module, and generate comparison difference audio and video respectively after the comparison is completed. One comparison difference audio and video is generated respectively after one current actual audio and video data is compared with one standard private audio and video data. The standard private database and the standard private audio and video data are both preset in advance; judge whether the current actual audio and video data is a stored video according to each comparison difference audio and video. If it is judged to be yes, delete the current actual audio and video data; if it is judged to be no, store the current actual audio and video data. By obtaining the audio and video start operation data and management subject action data of the current audio and video management subject when starting audio and video production, and obtaining the current actual audio and video data of the current audio and video management subject after the end of audio and video production, and judging whether the audio and video production of the current audio and video management subject is private audio and video production. If it is judged that it is not private audio and video production, store the current actual audio and video data in a normal audio and video repository. Then, in order to achieve non-repetitive storage of private data and efficient processing of data, further compare the current actual audio and video data with the standard private audio and video data in the pre-stored standard private database based on a preset intelligent comparison module, and generate comparison difference audio and video respectively after the comparison is completed. Finally, by judging whether it is a stored video, while realizing the privacy in the process of audio and video production, it can also realize that the intelligent stored video is no longer stored continuously, thus realizing efficient data management. Description of the Drawings

[0030] Figure 1 It is a schematic flowchart of an audio and video data processing method based on artificial intelligence in an embodiment;

[0031] Figure 2 It is a structural block diagram of an AI-based audio-visual data processing system in an embodiment;

[0032] Figure 3 It is an internal structure diagram of a computer device in an embodiment. Specific Embodiments

[0033] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0034] In one embodiment, a terminal is provided. The terminal is used to: obtain the audio-visual startup operation data and management subject action data of the current audio-visual management subject when starting audio-visual production, and obtain the current actual audio-visual data of the current audio-visual management subject after the audio-visual production ends; determine whether the audio-visual production of the current audio-visual management subject is private audio-visual production based on the audio-visual startup operation data and management subject action data. If it is determined that it is not private audio-visual production, store the current actual audio-visual data in a normal audio-visual repository, where the normal audio-visual repository is pre-set; if it is determined that it is private audio-visual production, compare the current actual audio-visual data with the standard private audio-visual data in a pre-stored standard private database based on a pre-set intelligent comparison module, and generate comparison difference audio-visuals respectively after the comparison is completed. One comparison difference audio-visual is generated after comparing one current actual audio-visual data with one standard private audio-visual data. Both the standard private database and the standard private audio-visual data are pre-set; determine whether the current actual audio-visual data is a stored video based on each comparison difference audio-visual. If it is determined to be yes, delete the current actual audio-visual data; if it is determined to be no, store the current actual audio-visual data.

[0035] The terminal can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices.

[0036] In one embodiment, as Figure 1 shown, an AI-based audio-visual data processing method is provided. The method includes:

[0037] Step S100: Obtain the audio-visual startup operation data and management subject action data of the current audio-visual management subject when starting audio-visual production, and obtain the current actual audio-visual data of the current audio-visual management subject after the audio-visual production ends;

[0038] Step S200: Determine whether the audio-video production of the current audio-video management entity is private audio-video production based on the audio-video startup operation data and management entity action data. If it is determined that it is not private audio-video production, store the current actual audio-video data in the normal audio-video repository, where the normal audio-video repository is preset in advance.

[0039] Step S300: If it is determined that it is private audio-video production, compare the current actual audio-video data with the standard private audio-video data in the pre-stored standard private database based on a preset intelligent comparison module, and generate comparison difference audio-video respectively after the comparison. One comparison difference audio-video is generated respectively after one current actual audio-video data is compared with one standard private audio-video data. The standard private database and the standard private audio-video data are both preset in advance.

[0040] Step S400: Determine whether the current actual audio-video data is a stored video based on each comparison difference audio-video. If it is determined that it is, delete the current actual audio-video data; if it is determined that it is not, store the current actual audio-video data.

[0041] In this embodiment, the current audio-video management entity is the entity that needs to perform privacy management and redundant data management of audio-video. When performing the current operation, the current audio-video management entity has already produced and stored multiple audio-video. These already produced and stored audio-video are set as standard private audio-video data, and the database formed by the standard private audio-video data is the standard private database. In order to accurately obtain whether the audio-video production operation currently performed by the current audio-video management entity is private audio-video production to improve privacy, the audio-video startup operation data and management entity action data of the current audio-video management entity when starting audio-video production are obtained, and the current actual audio-video data of the current audio-video management entity after the audio-video production is completed is obtained, and it is determined whether the audio-video production of the current audio-video management entity is private audio-video production. If it is determined that it is not private audio-video production, store the current actual audio-video data in the normal audio-video repository. Then, in order to achieve non-repetitive storage of private data and efficient processing of data, the current actual audio-video data is compared with the standard private audio-video data in the pre-stored standard private database based on a preset intelligent comparison module, and comparison difference audio-video are generated respectively after the comparison. Finally, by determining whether it is a stored video, while achieving privacy in the audio-video production process, it can also achieve no longer storing the intelligent stored video, so as to achieve efficient data management.

[0042] In one embodiment, step S400: Determine whether the current actual audio-video data is a stored video based on each of the comparison difference audio-video. If the determination is yes, delete the current actual audio-video data; if the determination is no, store the current actual audio-video data; specifically including:

[0043] Step S410: Determine the actual video capacity of each of the comparison difference audio-video based on each of the comparison difference audio-video, summarize each of the actual video capacities and generate a total difference video capacity;

[0044] Step S420: Determine whether the total difference video capacity is greater than or equal to the standard video capacity. If it is determined that the total difference video capacity is less than the standard video capacity, determine that the current actual audio-video data is a stored video and delete the current actual audio-video data;

[0045] Step S430: If it is determined that the total difference video capacity is greater than or equal to the standard video capacity, determine that the current actual audio-video data is not a stored video, and mark the videos in the current actual audio-video data except the comparison difference audio-video as actually identical audio-video;

[0046] In this embodiment, the standard video capacity is preset. The standard video capacity is the memory size occupied by the video, that is, the size of the standard video capacity. When the standard video capacity is small, it can be determined that the current actual audio-video data is basically the same as the previously stored video. Therefore, it can be determined that the current actual audio-video data is a stored video. That is, first, determine the actual video capacity of each of the comparison difference audio-video based on each of the comparison difference audio-video, summarize each of the actual video capacities and generate a total difference video capacity; determine whether the total difference video capacity is greater than or equal to the standard video capacity. If it is determined that the total difference video capacity is less than the standard video capacity, determine that the current actual audio-video data is a stored video and delete the current actual audio-video data; then, if it is determined that the total difference video capacity is greater than or equal to the standard video capacity, determine that the current actual audio-video data is not a stored video, and mark the videos in the current actual audio-video data except the comparison difference audio-video as actually identical audio-video. Of course, the setting of the standard video capacity is diverse, such as set to 0MB, or set to 0.2MB, etc.

[0047] Step S440: Extract the video connection nodes of the actually identical audio-video and the comparison difference audio-video, where the number of the video connection nodes is multiple;

[0048] Step S450: Establish video association data of the actually identical audio-video and the comparison difference audio-video based on the video connection nodes;

[0049] Step S460: Compress the video association data and generate associated compressed data;

[0050] Step S470: Crop the current actual audio - video data, cut out the actually identical audio - video in the current actual audio - video data, bind the cut - out actually identical audio - video with the associated compressed data, and generate cloud - stored audio - video data after binding;

[0051] Step S480: Store the cloud - stored audio - video data in a preset cloud space;

[0052] Step S490: Store the comparison - difference audio - video remaining after cutting out the actually identical audio - video from the current actual audio - video data in a standard privacy database.

[0053] In this embodiment, in order to maintain the integrity of data while avoiding affecting the memory capacity due to not storing the same data, first, extract the video connection nodes of the actually identical audio - video and the comparison - difference audio - video, and then establish the video association data of the actually identical audio - video and the comparison - difference audio - video according to the video connection nodes; then, compress the video association data and generate associated compressed data; then, crop the current actual audio - video data, cut out the actually identical audio - video in the current actual audio - video data, bind the cut - out actually identical audio - video with the associated compressed data, and generate cloud - stored audio - video data after binding. By storing the cloud - stored audio - video data in a preset cloud space; and then storing the comparison - difference audio - video remaining after cutting out the actually identical audio - video from the current actual audio - video data in a standard privacy database, only the data different from the previously stored data is retained, specifically, the remaining comparison - difference audio - video is saved, and for the data identical to the previous ones, data compression is achieved by generating associated compressed data, thus saving the capacity of the cloud space. In addition, by establishing the video association data of the actually identical audio - video and the comparison - difference audio - video through the video connection nodes, when the current audio - video management entity needs to obtain the original video later, the actually identical audio - video can be downloaded from the cloud space, and at the same time, the actually identical audio - video and the comparison - difference audio - video are combined and spliced based on the video association data of the video connection nodes, thereby restoring the current actual audio - video data, achieving the privacy in the process of audio - video production while also being able to intelligently avoid storing the already - stored videos continuously, thus realizing efficient data management.

[0054] In one embodiment, step S200: Determine whether the audio-video production of the current audio-video management entity is private audio-video production based on the audio-video startup operation data and the management entity action data. If it is determined that it is not private audio-video production, store the current actual audio-video data in a normal audio-video repository, where the normal audio-video repository is pre-set; specifically including:

[0055] Step S210: Extract operation actions according to the audio-video startup operation data, and obtain the actual trigger nodes of the current audio-video management entity after the operation action extraction, where the number of the actual trigger nodes is multiple;

[0056] Step S220: Obtain the actual trigger time points of each of the actual trigger nodes, connect the actual trigger nodes according to each of the actual trigger time points by a curve, and generate a virtual operation curve after the curve connection;

[0057] Step S230: Determine whether the virtual operation curve matches a privacy trigger curve. If it is determined that the virtual operation curve matches the privacy trigger curve, determine that the audio-video production of the current audio-video management entity is private audio-video production;

[0058] Step S240: If it is determined that the virtual operation curve does not match the privacy trigger curve, extract actions from the management entity action data, and generate actual operation actions after the action extraction, where the number of the actual operation actions is multiple;

[0059] Step S250: Combine each of the actual operation actions and generate a combined operation action;

[0060] Step S260: Determine whether the combined operation action matches a privacy operation action. If it is determined that the combined operation action matches the privacy operation action, determine that the audio-video production of the current audio-video management entity is private audio-video production;

[0061] Step S270: If it is determined that the combined operation action does not match the privacy operation action, determine that the audio-video production of the current audio-video management entity is not private audio-video production, and store the current actual audio-video data in the pre-set normal audio-video repository.

[0062] In this embodiment, to ensure privacy management, curve automatic connection is set, that is, the user can click anywhere on the interface. Specifically, first obtain the actual trigger time points of each actual trigger node, and connect the curves of each actual trigger node according to each actual trigger time point, and generate a virtual operation curve after curve connection; then, determine whether the virtual operation curve matches the privacy trigger curve. If it is determined that the virtual operation curve matches the privacy trigger curve, it is determined that the audio-video production of the current audio-video management entity is privacy audio-video production; then, if it is determined that the virtual operation curve does not match the privacy trigger curve, extract the actions from the management entity action data, and generate actual operation actions after action extraction, where the number of the actual operation actions is multiple; at this time, after further determining that the actions may prevent the user from performing privacy authentication in a hurry, the actual operation actions can also be combined to generate combined operation actions; then, determine whether the combined operation actions match the privacy operation actions. If it is determined that the combined operation actions match the privacy operation actions, it is determined that the audio-video production of the current audio-video management entity is privacy audio-video production; finally, if it is determined that the combined operation actions do not match the privacy operation actions, it is determined that the audio-video production of the current audio-video management entity is not privacy audio-video production, and the current actual audio-video data is stored in a pre-set normal audio-video storage repository, thus realizing the management of efficient privacy data.

[0063] In one embodiment, step S300: If it is determined that it is privacy audio-video production, based on a pre-set intelligent comparison module, compare the current actual audio-video data with the standard privacy audio-video data in the pre-stored standard privacy database, and generate comparison difference audio-video respectively after the comparison is completed; specifically including:

[0064] Step S310: If it is determined that it is privacy audio-video production, split the current actual audio-video data and the standard privacy audio-video data in the pre-stored standard privacy database respectively. Among them, after the current actual audio-video data is split, multiple actual refined audio-video are generated, and after the standard privacy audio-video data is split, multiple standard refined audio-video are generated, and each standard privacy audio-video data corresponds to multiple standard privacy refined data;

[0065] Step S320: Input each actual refined audio-video into the 3DCNN network in the intelligent comparison module, and obtain the actual feature representation output by the 3DCNN network, where one actual refined audio-video corresponds to one actual feature representation;

[0066] Step S330: Input the standard privacy refinement data corresponding to the standard privacy audio-visual data into the 3DCNN network in the intelligent comparison module, and obtain the standard feature representation output by the 3DCNN network. Among them, one standard privacy refinement data corresponds to one standard feature representation;

[0067] Step S340: Compare the actual feature representation with the standard feature representation, and generate comparison difference audio-visuals respectively after the comparison is completed.

[0068] In this embodiment, the 3DCNN network is pre-trained. When comparing one standard privacy audio-visual data with one standard privacy database, the following steps are adopted:

[0069] First, split the standard privacy audio-visual data and the standard privacy database into continuous video frames respectively, and form a 3D video dataset by arranging these frames in sequence to form an RGB image sequence, that is, the actual refined audio-visual and the standard refined audio-visual respectively.

[0070] Then, input each actual refined audio-visual into the 3DCNN network in the intelligent comparison module, and obtain the actual feature representation output by the 3DCNN network. Input the standard privacy refinement data corresponding to the standard privacy audio-visual data into the 3DCNN network in the intelligent comparison module, and obtain the standard feature representation output by the 3DCNN network.

[0071] Finally, compare the actual feature representation with the standard feature representation. When comparing, 3D convolutional layers can be used to compress and extract features by itself, and different convolutional kernel sizes and strides can be combined to enable the network to better capture spatial and temporal information in the video. Then, distance metric methods including but not limited to Euclidean distance and Manhattan distance are used to compare the differences between the feature representations of two videos, and similarity scoring or classification is performed. In this way, by using 3DCNN for video feature extraction and comparison, this method can avoid the problem that it is difficult for simple RGB image frames to capture the flow of video information, and can more effectively mine the spatio-temporal features in the video.

[0072] In one embodiment, as Figure 2 shown, a video and audio data processing system based on artificial intelligence is provided. The system includes:

[0073] A current data acquisition module, configured to acquire the audio-visual start operation data and the management body action data of the current audio-visual management body when starting audio-visual production, and acquire the current actual audio-visual data of the current audio-visual management body after the audio-visual production is completed;

[0074] A data judgment and storage module, configured to determine whether the audio-video production of the current audio-video management entity is private audio-video production according to the audio-video start operation data and the management entity action data. If it is determined that it is not private audio-video production, the current actual audio-video data is stored in a normal audio-video storage repository, where the normal audio-video storage repository is preset in advance;

[0075] An intelligent data processing module, configured to, if it is determined that it is private audio-video production, compare the current actual audio-video data with the standard private audio-video data in a pre-stored standard private database based on a preset intelligent comparison module, and generate comparison difference audio-video respectively after the comparison is completed. One comparison difference audio-video is generated respectively after one current actual audio-video data is compared with one standard private audio-video data. The standard private database and the standard private audio-video data are both preset in advance;

[0076] A storage judgment and processing module, configured to determine whether the current actual audio-video data is stored video according to each comparison difference audio-video. If it is determined that it is, the current actual audio-video data is deleted; if it is determined that it is not, the current actual audio-video data is stored.

[0077] In one embodiment, the storage judgment and processing module is further configured to:

[0078] Determine the actual video capacity of each comparison difference audio-video according to each comparison difference audio-video, summarize the actual video capacities of each comparison difference audio-video and generate a total difference video capacity; determine whether the total difference video capacity is greater than or equal to a standard video capacity. If it is determined that the total difference video capacity is less than the standard video capacity, it is determined that the current actual audio-video data is stored video, and the current actual audio-video data is deleted; if it is determined that the total difference video capacity is greater than or equal to the standard video capacity, it is determined that the current actual audio-video data is not stored video, and the videos in the current actual audio-video data except the comparison difference audio-video are marked as actually identical videos; extract the video connection nodes of the actually identical videos and the comparison difference audio-video, where the number of the video connection nodes is multiple; establish video association data of the actually identical videos and the comparison difference audio-video according to the video connection nodes; compress the video association data and generate associated compressed data; crop the current actual audio-video data, and crop out the actually identical videos in the current actual audio-video data, bind the cropped actually identical videos with the associated compressed data, and generate cloud storage audio-video data after binding; store the cloud storage audio-video data in a preset cloud space; store the comparison difference audio-video remaining after cropping out the actually identical videos from the current actual audio-video data in the standard private database.

[0079] In one embodiment, the data judgment and storage module is further configured to:

[0080] Extract operation actions according to the audio-visual start operation data, and obtain the actual trigger nodes of the current audio-visual management entity after the operation actions are extracted, wherein the number of the actual trigger nodes is multiple; obtain the actual trigger time points of each of the actual trigger nodes, and connect the actual trigger nodes in a curve according to each of the actual trigger time points, and generate a virtual operation curve after the curve connection; judge whether the virtual operation curve matches the privacy trigger curve, if it is judged that the virtual operation curve matches the privacy trigger curve, then judge that the audio-visual production of the current audio-visual management entity is privacy audio-visual production; if it is judged that the virtual operation curve does not match the privacy trigger curve, then extract actions from the management entity action data, and generate actual operation actions after the actions are extracted, wherein the number of the actual operation actions is multiple; combine each of the actual operation actions and generate a combined operation action; judge whether the combined operation action matches the privacy operation action, if it is judged that the combined operation action matches the privacy operation action, then judge that the audio-visual production of the current audio-visual management entity is privacy audio-visual production; if it is judged that the combined operation action does not match the privacy operation action, then judge that the audio-visual production of the current audio-visual management entity is not privacy audio-visual production, and store the current actual audio-visual data in a pre-set normal audio-visual storage repository.

[0081] In one embodiment, the intelligent data processing module is further configured to:

[0082] If it is judged that it is privacy audio-visual production, split the current actual audio-visual data and the standard privacy audio-visual data in the pre-stored standard privacy database respectively, wherein the current actual audio-visual data generates multiple actual refined audio-visuals after being split, the standard privacy audio-visual data generates multiple standard refined audio-visuals after being split, and each standard privacy audio-visual data corresponds to multiple standard privacy refinement data; input each of the actual refined audio-visuals into the 3DCNN network in the intelligent comparison module, and obtain the actual feature representation output by the 3DCNN network, wherein one actual refined audio-visual corresponds to one actual feature representation; input the standard privacy refinement data corresponding to the standard privacy audio-visual data into the 3DCNN network in the intelligent comparison module, and obtain the standard feature representation output by the 3DCNN network, wherein one standard privacy refinement data corresponds to one standard feature representation; compare the actual feature representation and the standard feature representation, and generate comparison difference audio-visuals respectively after the comparison is completed.

[0083] In one embodiment, as Figure 3As shown, a computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned audio-visual data processing method based on artificial intelligence are implemented.

[0084] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned audio-visual data processing method based on artificial intelligence are implemented.

[0085] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0086] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0087] The above-described embodiments merely represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limitations on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

Claims

1. An audio and video data processing method based on artificial intelligence, characterized in that: The method comprises: Obtaining audio and video startup operation data and management subject action data of the current audio and video management subject when starting audio and video production, and obtaining current actual audio and video data of the current audio and video management subject after the audio and video production is completed; determining whether the audio and video production of the current audio and video management subject is private audio and video production based on the audio and video startup operation data and management subject action data; if it is determined that it is not private audio and video production, storing the current actual audio and video data in a normal audio and video storage library, wherein the normal audio and video storage library is pre-set; specifically including: An operation action is extracted according to the audio and video startup operation data, and after the operation action is extracted, the actual trigger node of the current audio and video management subject is obtained, wherein the number of the actual trigger nodes is multiple; the actual trigger time point of each actual trigger node is obtained, and each actual trigger node is connected by a curve according to each actual trigger time point, and a virtual operation curve is generated after the curve connection; it is judged whether the virtual operation curve matches the privacy trigger curve, and if it is judged that the virtual operation curve matches the privacy trigger curve, it is judged that the audio and video production of the current audio and video management subject is a privacy audio and video production; if it is judged that the virtual operation curve does not match the privacy trigger curve, Then, action extraction is performed on the management subject action data, and actual operation actions are generated after the action extraction, wherein the number of the actual operation actions is multiple; each of the actual operation actions is combined to generate a combined operation action; it is determined whether the combined operation action matches the privacy operation action; if it is determined that the combined operation action matches the privacy operation action, it is determined that the audio and video production of the current audio and video management subject is a privacy audio and video production; if it is determined that the combined operation action does not match the privacy operation action, it is determined that the audio and video production of the current audio and video management subject is not a privacy audio and video production, and the current actual audio and video data is stored in a pre-set normal audio and video storage library; The audio and video start operation data is a virtual operation curve generated by connecting the actual trigger nodes according to the actual trigger time points of the actual trigger nodes obtained by the user clicking anywhere on the interface, and the management subject action data is a combination of the user's actual operation actions; If it is determined that the audio and video is private, the current actual audio and video data is compared with the standard private audio and video data in the pre-stored standard privacy database based on the preset intelligent comparison module, and after the comparison is completed, a comparison difference audio and video is generated respectively, wherein, after one current actual audio and video data is compared with one standard private audio and video data, a comparison difference audio and video is generated respectively, and the standard privacy database and the standard private audio and video data are both pre-set; according to each comparison difference audio and video, it is determined whether the current actual audio and video data is a stored video, if it is determined to be yes, the current actual audio and video data is deleted; if it is determined to be no, the current actual audio and video data is stored; the comparison difference audio and video is the audio and video data in the current actual audio and video data that is different from the standard private audio and video data.

2. The method for processing audio and video data based on artificial intelligence according to claim 1, characterized in that: Determining whether the current actual audio and video data is a stored video based on the compared audio and video data, and if so, deleting the current actual audio and video data; If the judgment is no, then the current actual audio and video data is stored; specifically including: Determine the actual video capacity of each of the compared difference audio and video according to each of the compared difference audio and video, summarize the actual capacity of each video and generate the total capacity of the difference video; determine whether the total capacity of the difference video is greater than or equal to the standard video capacity, if it is determined that the total capacity of the difference video is less than the standard video capacity, then determine that the current actual audio and video data is a stored video, and delete the current actual audio and video data; if it is determined that the total capacity of the difference video is greater than or equal to the standard video capacity, then determine that the current actual audio and video data is not a stored video, then mark the video in the current actual audio and video data except the compared difference audio and video as the actual same audio and video; extract the actual same audio and video and the compared difference audio and video A video connection node for a video, wherein the number of the video connection nodes is multiple; video association data of the actual identical audio and video and the comparative difference audio and video is established according to the video connection node; the video association data is compressed and associated compressed data is generated; the current actual audio and video data is cropped, and the actual identical audio and video in the current actual audio and video data is cropped out, the cropped actual identical audio and video is bound to the associated compressed data, and cloud storage audio and video data is generated after binding; the cloud storage audio and video data is stored in a preset cloud space; the comparative difference audio and video remaining after the actual identical audio and video is cropped out of the current actual audio and video data is stored in a standard privacy database.

3. The method for processing audio and video data based on artificial intelligence according to claim 1, wherein: If it is determined that the audio and video is private, the current actual audio and video data is compared with the standard private audio and video data stored in the standard privacy database based on the preset intelligent comparison module, and after the comparison is completed, the difference audio and video are generated respectively; specifically including: If it is determined that the audio and video is private, the current actual audio and video data and the standard private audio and video data in the pre-stored standard privacy database are split respectively, wherein the current actual audio and video data are split to generate multiple actual refined audio and video, and the standard private audio and video data are split to generate multiple standard refined audio and video, and each of the standard private audio and video data corresponds to multiple standard private refined data; each of the actual refined audio and video is input into the 3DCNN network in the intelligent comparison module, and the actual feature representation output by the 3DCNN network is obtained, wherein one actual refined audio and video corresponds to one actual feature representation; the standard private refined data corresponding to the standard private audio and video data is input into the 3DCNN network in the intelligent comparison module, and the standard feature representation output by the 3DCNN network is obtained, wherein one standard private refined data corresponds to one standard feature representation; the actual feature representation and the standard feature representation are compared, and after the comparison is completed, the comparison difference audio and video are generated respectively.

4. An audio and video data processing system based on artificial intelligence, characterized in that: The system comprises: The current data acquisition module is used to obtain the audio and video startup operation data and management subject action data of the current audio and video management subject when starting audio and video production, and to obtain the current actual audio and video data of the current audio and video management subject after the audio and video production is completed; A data judgment storage module is used to judge whether the audio and video production of the current audio and video management subject is a private audio and video production based on the audio and video startup operation data and the management subject action data. If it is judged that it is not a private audio and video production, the current actual audio and video data is stored in a normal audio and video storage library, wherein the normal audio and video storage library is pre-set; the data judgment storage module is also used to: extract the operation action based on the audio and video startup operation data, and obtain the actual trigger node of the current audio and video management subject after the operation action is extracted, wherein the number of the actual trigger nodes is multiple; obtain the actual trigger time point of each actual trigger node, and connect each actual trigger node with a curve according to each actual trigger time point, and generate a virtual operation curve after the curve connection; judge whether the virtual operation curve matches the privacy trigger curve, and if it is judged that If the virtual operation curve matches the privacy trigger curve, it is determined that the audio and video production of the current audio and video management subject is a privacy audio and video production; if it is determined that the virtual operation curve does not match the privacy trigger curve, action extraction is performed on the management subject action data, and an actual operation action is generated after the action extraction, wherein the number of the actual operation actions is multiple; the actual operation actions are combined to generate a combined operation action; it is determined whether the combined operation action matches the privacy operation action, and if it is determined that the combined operation action matches the privacy operation action, it is determined that the audio and video production of the current audio and video management subject is a privacy audio and video production; if it is determined that the combined operation action does not match the privacy operation action, it is determined that the audio and video production of the current audio and video management subject is not a privacy audio and video production, and the current actual audio and video data is stored in a pre-set normal audio and video storage library; The audio and video start operation data is a virtual operation curve generated by connecting the actual trigger nodes according to the actual trigger time points of the actual trigger nodes obtained by the user clicking anywhere on the interface, and the management subject action data is a combination of the user's actual operation actions; An intelligent data processing module, configured to compare the current actual audio and video data with standard privacy audio and video data stored in a pre-stored standard privacy database based on a preset intelligent comparison module if the data is determined to be private audio and video production, and generate a comparison difference audio and video after the comparison is completed, wherein each comparison difference audio and video is generated after comparing the current actual audio and video data with the standard privacy audio and video data, and the standard privacy database and the standard privacy audio and video data are both pre-set; the comparison difference audio and video is audio and video data in the current actual audio and video data that differs from the standard privacy audio and video data; The storage judgment processing module is used to judge whether the current actual audio and video data is a stored video based on the compared difference audio and video. If it is judged to be yes, the current actual audio and video data is deleted; if it is judged to be no, the current actual audio and video data is stored.

5. The audio and video data processing system based on artificial intelligence according to claim 4, characterized in that: The storage judgment processing module is further used for: Determine the actual video capacity of each of the compared difference audio and video according to each of the compared difference audio and video, summarize the actual capacity of each of the videos and generate a total difference video capacity; determine whether the total difference video capacity is greater than or equal to the standard video capacity; if it is determined that the total difference video capacity is less than the standard video capacity, determine that the current actual audio and video data is a stored video, and delete the current actual audio and video data; if it is determined that the total difference video capacity is greater than or equal to the standard video capacity, determine that the current actual audio and video data is not a stored video, and mark the videos in the current actual audio and video data other than the compared difference audio and video as actually identical audio and video; Extract the video connection nodes of the actual identical audio and video and the comparative difference audio and video, wherein the number of the video connection nodes is multiple; establish video association data of the actual identical audio and video and the comparative difference audio and video according to the video connection nodes; compress the video association data and generate associated compressed data; crop the current actual audio and video data, and crop the actual identical audio and video in the current actual audio and video data, bind the cropped actual identical audio and video with the associated compressed data, and generate cloud storage audio and video data after binding; store the cloud storage audio and video data in a preset cloud space; store the comparative difference audio and video remaining after cropping the actual identical audio and video from the current actual audio and video data in a standard privacy database.

6. The audio and video data processing system based on artificial intelligence according to claim 5, characterized in that: The intelligent data processing module is also used for: If it is determined that the audio and video is private, the current actual audio and video data and the standard private audio and video data in the pre-stored standard privacy database are split respectively, wherein the current actual audio and video data are split to generate multiple actual refined audio and video, and the standard private audio and video data are split to generate multiple standard refined audio and video, and each of the standard private audio and video data corresponds to multiple standard private refined data; each of the actual refined audio and video is input into the 3DCNN network in the intelligent comparison module, and the actual feature representation output by the 3DCNN network is obtained, wherein one actual refined audio and video corresponds to one actual feature representation; the standard private refined data corresponding to the standard private audio and video data is input into the 3DCNN network in the intelligent comparison module, and the standard feature representation output by the 3DCNN network is obtained, wherein one standard private refined data corresponds to one standard feature representation; the actual feature representation and the standard feature representation are compared, and after the comparison is completed, the comparison difference audio and video are generated respectively.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Audio and video stream docking processing method based on artificial intelligence AI and related equipment

    CN111787418A

  • Method and device for comparing video

    CN103902553A

  • Comparing system and method for data having different file formats

    CN106021301A

  • Private data protection method, mobile terminal and storage medium

    CN109271764A