A method, apparatus, electronic device, and storage medium for multimedia information processing
The method uses neural networks to analyze audio features in videos for copyright verification, addressing the challenge of audio modifications in video editing, enhancing speed and accuracy in copyright detection.
Patent Information
- Application Number
- CN202110036472.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-01-12
AI Technical Summary
The prior art is difficult to effectively identify videos edited by audio, especially videos with accelerated sound changes and multi-layer audio track superimposed, making it difficult for video infringing content to be accurately identified and reviewed.
By obtaining the target audio in the multimedia information, forming a Mel spectrogram that matches the time domain characteristics and the frequency domain characteristics, using the first sub-model network and the second sub-model network in the multimedia information processing model, the first audio feature vector and the second audio feature vector of the target audio are determined, and the audio type is determined based on these feature vectors.
It improves the accuracy and speed of audio type judgment, reduces the workload of manual audits, improves the efficiency and accuracy of multimedia information audits, and improves user experience.
Smart Images

Figure CN113539299B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to multimedia information processing technology, and in particular to a multimedia information processing method, apparatus, electronic device, and storage medium. Background Art
[0002] In related technologies, the forms of multimedia information are diverse, and the demand for multimedia information has shown an explosive growth. The quantity and types of multimedia information received by multimedia information servers are also increasing. Taking long videos as an example, video servers can identify the similarity relationships between videos through corresponding matching algorithms. However, with the popularization and development of video editing tools, there are more and more audio editing methods for videos, such as accelerating and changing voices, superimposing multiple audio tracks to avoid copyright review, and it has become increasingly difficult to distinguish videos through algorithms. For such audiotrack-edited videos, it is slow to manually review infringing content, which affects users' usage. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a multimedia information processing method, apparatus, electronic device, and storage medium, which can classify target audio in multimedia information, process the target audio using a multimedia information processing model, and determine the type of the target audio in the target multimedia information.
[0004] The technical solution of the embodiments of the present invention is implemented as follows:
[0005] Embodiments of the present invention provide a multimedia information processing method, including:
[0006] Obtain target multimedia information, and parse the target multimedia information to separate the target audio included in the multimedia information;
[0007] Perform conversion processing on the target audio to form a Mel spectrogram that matches the time domain features and frequency domain features of the target audio;
[0008] Through the first sub-model network in the multimedia information processing model, based on the Mel spectrogram that matches the time domain features and frequency domain features of the target audio, determine the first audio feature vector corresponding to the target audio;
[0009] Through the second sub-model network in the multimedia information processing model, based on the Mel spectrogram that matches the time domain features and frequency domain features of the target audio, determine the second audio feature vector corresponding to the target audio;
[0010] Based on the first audio feature vector and the second audio feature vector, determine the type of the target audio in the target multimedia information.
[0011] An embodiment of the present invention further provides a multimedia information processing device, characterized in that the device includes:
[0012] An information transmission module, configured to obtain target multimedia information, and parse the target multimedia information to separate the target audio included in the multimedia information;
[0013] An information processing module, configured to perform conversion processing on the target audio to form a Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio;
[0014] The information processing module is configured to determine a first audio feature vector corresponding to the target audio based on the Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio through a first sub-model network in the multimedia information processing model;
[0015] The information processing module is configured to determine a second audio feature vector corresponding to the target audio based on the Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio through a second sub-model network in the multimedia information processing model;
[0016] The information processing module is configured to determine the type of the target audio in the target multimedia information based on the first audio feature vector and the second audio feature vector.
[0017] In the above solution, the information processing module is configured to parse the target multimedia information to obtain the timing information of the target multimedia information;
[0018] The information processing module is configured to parse the video parameters corresponding to the target multimedia information according to the timing information of the target multimedia information to obtain the playback duration parameter and the audio track information parameter corresponding to the target multimedia information;
[0019] The information processing module is configured to extract the target multimedia information based on the playback duration parameter and the audio track information parameter corresponding to the target multimedia information to obtain the target audio corresponding to the target multimedia information.
[0020] In the above solution, the information processing module is configured to perform channel conversion processing on the target audio to form mono audio data;
[0021] The information processing module is configured to perform short-time Fourier transform on the mono audio data based on a window function corresponding to the multimedia information processing model to form a corresponding spectrogram;
[0022] The information processing module is configured to determine the duration parameter corresponding to the multimedia information processing model;
[0023] The information processing module is configured to process the spectrogram according to the duration parameter to form a Mel spectrogram that matches the time-domain features and frequency-domain features of the target audio.
[0024] In the above solution, the information processing module is configured to convert the Mel spectrogram that matches the time-domain features and frequency-domain features of the target audio into a corresponding grayscale image;
[0025] The information processing module is configured to extract the feature vector of the Mel spectrogram according to the grayscale image through the convolutional neural network in the first sub-model network of the multimedia information processing model;
[0026] The information processing module is configured to process the feature vector of the Mel spectrogram through the gated recurrent unit in the first sub-model network to determine the first audio feature vector corresponding to the target audio.
[0027] In the above solution, the information processing module is configured to determine the number of channels of the gated recurrent unit in the first sub-model network based on the number of the Mel spectrograms;
[0028] The information processing module is configured to determine the time series parameter according to the time-domain features and frequency-domain features of the target audio;
[0029] The information processing module is configured to determine the recurrent neural network in the first sub-model network based on the number of channels of the gated recurrent unit in the first sub-model network and the time series parameter;
[0030] The information processing module is configured to determine the first audio feature vector corresponding to the target audio through the recurrent neural network in the first sub-model network.
[0031] In the above solution, the information processing module is configured to determine the output information of the average pooling layer network through the residual network in the second sub-model network of the multimedia information processing model based on the Mel spectrogram that matches the time-domain features and frequency-domain features of the target audio;
[0032] The information processing module is configured to adjust the parameters of the image classification network in the second sub-model network according to the output information of the average pooling layer network;
[0033] The information processing module is configured to determine the second audio feature vector corresponding to the target audio through the image classification network in the second sub-model network based on the Mel spectrogram that matches the time-domain features and frequency-domain features of the target audio.
[0034] In the above solution,
[0035] The information processing module is configured to establish a data storage mapping according to the information source of the target multimedia information;
[0036] The information processing module is configured to adjust the file format of the target audio in response to the established data storage mapping to match the information source.
[0037] In the above solution, the training module is configured to obtain a first training sample set, where the first training sample set is an audio sample in the video information collected by the terminal;
[0038] The training module is configured to add noise to the first training sample set to form a corresponding second training sample set;
[0039] The training module is configured to process the second training sample set through a multimedia information processing model to determine the initial parameters of the multimedia information processing model;
[0040] The training module is configured to, in response to the initial parameters of the multimedia information processing model, process the second training sample set through the multimedia information processing model to determine the updated parameters of the multimedia information processing model;
[0041] Iteratively update the network parameters of the multimedia information processing model through the second training sample set according to the updated parameters of the multimedia information processing model.
[0042] In the above solution, the training module is configured to determine a dynamic noise type matching the usage environment of the multimedia information processing model;
[0043] The training module is configured to add noise to the first training sample set according to the dynamic noise type to change the background noise, volume, or sampling rate of the audio samples in the first training sample set to form a corresponding second training sample set.
[0044] In the above solution, the training module is configured to substitute different audio samples in the second training sample set into the loss functions corresponding to the first sub-model network and the second sub-model network of the multimedia information processing model;
[0045] The training module is configured to determine the parameters corresponding to the first sub-model network and the second sub-model network in the multimedia information processing model when the loss function satisfies the corresponding convergence condition;
[0046] The training module is configured to use the parameters corresponding to the first sub-model network and the second sub-model network respectively as the updated parameters of the multimedia information processing model.
[0047] In the above solution, the training module is used to determine the convergence conditions respectively matching the first sub-model network and the second sub-model network in the multimedia information processing model;
[0048] The training module is used to iteratively update the parameters respectively corresponding to the first sub-model network and the second sub-model network until the loss functions respectively corresponding to the first sub-model network and the second sub-model network meet the corresponding convergence conditions.
[0049] In the above solution, the information processing module is used to perform vector fusion processing on the first audio feature vector and the second audio feature vector;
[0050] The information processing module is used to determine the type of the target audio in the target multimedia information based on the result of the vector fusion processing, where the type of the target audio includes at least one of the following:
[0051] Compliant audio, accelerated and pitch-changed audio, and multi-layer audio track superimposed audio.
[0052] In the above solution, the information processing module is used to determine the source multimedia information corresponding to the target multimedia information;
[0053] The information processing module is used to determine a corresponding set of inter-frame similarity parameters through the first audio feature vector and the second audio feature vector based on the target audio of the target multimedia information and the source audio of the source multimedia information;
[0054] The information processing module is used to obtain the number of audio frames reaching the similarity threshold in the set of inter-frame similarity parameters;
[0055] The information processing module is used to determine the similarity between the target multimedia information and the source multimedia information based on the number of audio frames reaching the similarity threshold.
[0056] In the above solution, the information processing module is used to obtain the copyright information of the target multimedia information when it is determined that the target multimedia information is similar to the source multimedia information;
[0057] The information processing module is used to determine the legality of the target multimedia information through the copyright information of the target multimedia information and the copyright information of the source multimedia information;
[0058] The information processing module is used to send a warning message when the copyright information of the target multimedia information and the copyright information of the source multimedia information are inconsistent.
[0059] In the above solution, the information processing module is configured to add the target multimedia information to the multimedia information source when it is determined that the target multimedia information is not similar to the source multimedia information;
[0060] The information processing module is configured to sort the recall order of the multimedia information to be recommended in the multimedia information source;
[0061] The information processing module is configured to recommend multimedia information to the target user based on the sorting result of the recall order of the multimedia information to be recommended.
[0062] In the above solution, the information processing module is configured to send the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information to the blockchain network, so that
[0063] The nodes of the blockchain network fill the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information into a new block, and when consensus on the new block is reached, append the new block to the end of the blockchain.
[0064] An embodiment of the present invention further provides an electronic device, where the electronic device includes:
[0065] A memory for storing executable instructions;
[0066] A processor, when running the executable instructions stored in the memory, implements the foregoing multimedia information processing method.
[0067] An embodiment of the present invention further provides a computer-readable storage medium storing executable instructions, and when the executable instructions are executed by a processor, the foregoing multimedia information processing method is implemented.
[0068] The embodiments of the present invention have the following beneficial effects:
[0069] An embodiment of the present invention obtains target multimedia information, analyzes the target multimedia information to separate the target audio included in the multimedia information; performs conversion processing on the target audio to form a Mel spectrogram that matches the time-domain characteristics and frequency-domain characteristics of the target audio; through the first sub-model network in the multimedia information processing model, based on the Mel spectrogram that matches the time-domain characteristics and frequency-domain characteristics of the target audio, determines the first audio feature vector corresponding to the target audio; through the second sub-model network in the multimedia information processing model, based on the Mel spectrogram that matches the time-domain characteristics and frequency-domain characteristics of the target audio, determines the second audio feature vector corresponding to the target audio; based on the first audio feature vector and the second audio feature vector, determines the type of the target audio in the target multimedia information. Thus, the type of the target audio in the target multimedia information can be determined, reducing the workload of manual review, improving the speed and accuracy of multimedia information review, and enhancing the user experience. Description of the Drawings
[0070] Figure 1 is a schematic diagram of the usage environment of a multimedia information processing method provided by an embodiment of the present invention;
[0071] Figure 2 is a schematic diagram of the composition structure of an electronic device provided by an embodiment of the present invention;
[0072] Figure 3 is a schematic diagram of short video playback in the related art of an embodiment of the present invention;
[0073] Figure 4 is an optional flowchart of a multimedia information processing method provided by an embodiment of the present invention;
[0074] Figure 5 is a schematic diagram of the processing process of the multimedia information processing model for audio in an embodiment of the present invention;
[0075] Figure 6 is an optional flowchart of a multimedia information processing method provided by an embodiment of the present invention;
[0076] Figure 7 is an optional flowchart of a multimedia information processing method provided by an embodiment of the present invention;
[0077] Figure 8 is an optional schematic diagram of the type determination of the target audio in an embodiment of the present invention;
[0078] Figure 9 is a schematic diagram of the architecture of a blockchain network provided by an embodiment of the present invention;
[0079] Figure 10It is a schematic structural diagram of the blockchain in the blockchain network 200 provided by an embodiment of the present invention;
[0080] Figure 11 It is a schematic functional architecture diagram of the blockchain network 200 provided by an embodiment of the present invention;
[0081] Figure 12 It is a schematic diagram of the usage scenario of the multimedia information processing method provided by an embodiment of the present invention;
[0082] Figure 13 It is a schematic diagram of the usage process of the multimedia information processing method in an embodiment of the present invention. Detailed implementation manners
[0083] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0084] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0085] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are subject to the following explanations.
[0086] 1) In response to: used to represent the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more executed operations can be real-time or can have a set delay; without special instructions, there is no limitation on the execution order of the multiple executed operations.
[0087] 2) Target video: various forms of video information available on the Internet, such as video files and multimedia information presented in a client or intelligent device.
[0088] 3) Client: a carrier for implementing specific functions in a terminal. For example, a mobile client (APP) is a carrier for specific functions in a mobile terminal, such as performing functions of online live broadcast (video streaming) or playing online videos.
[0089] 4) Short-Time Fourier Transform: The short-time Fourier transform (STFT) is a mathematical transform related to the Fourier transform, used to determine the frequency and phase of sine waves in local regions of time-varying signals.
[0090] 5) Mel Bank Features: Since the obtained spectrogram is large, in order to obtain appropriate-sized sound features, it is usually passed through a Mel-scale filter bank to become Mel Bank Features.
[0091] 6) Information flow: A form of content organization arranged vertically and horizontally according to specific specifications. From the perspective of display sorting, common ones include chronological order, popularity, and algorithm sorting.
[0092] 7) Audio feature vector, that is, the audio 01 vector, is a binary feature vector generated based on audio.
[0093] 8) Transaction: Equivalent to the computer term "transaction", a transaction includes operations that need to be submitted to the blockchain network for execution, not just referring to transactions in a business context. Given that the term "transaction" is conventionally used in blockchain technology, the embodiments of the present invention follow this convention.
[0094] For example, a Deploy transaction is used to install a specified smart contract on nodes in the blockchain network and prepare it to be called; an Invoke transaction is used to append transaction records to the blockchain by calling a smart contract and operate on the state database of the blockchain, including update operations (including adding, deleting, and modifying key-value pairs in the state database) and query operations (i.e., querying key-value pairs in the state database).
[0095] 9) Block chain: An encrypted, chained storage structure formed by blocks.
[0096] For example, the header of each block can include the hash values of all transactions in the block, and at the same time, also include the hash values of all transactions in the previous block, so as to achieve anti-tampering and anti-forgery of transactions in the block based on the hash values; newly generated transactions are filled into the block and after being consensus by nodes in the blockchain network, they will be appended to the tail of the blockchain to form a chained growth.
[0097] 10) Block chain Network: A set of nodes that incorporate new blocks into the blockchain through consensus.
[0098] 11) Ledger: A collective term for a blockchain (also known as ledger data) and a state database synchronized with the blockchain.
[0099] Among them, the blockchain records transactions in the form of files in the file system; the state database records transactions in the blockchain in the form of different types of key-value pairs, which is used to support fast query of transactions in the blockchain.
[0100] 12) Smart Contracts: Also known as chain code or application code, it is a program deployed in the nodes of the blockchain network. The nodes execute the smart contracts called in the received transactions to perform operations such as updating or querying the key-value pair data in the ledger database.
[0101] 13) Consensus: It is a process in the blockchain network used to reach an agreement on the transactions in a block among multiple involved nodes. The block that reaches an agreement will be appended to the end of the blockchain. The mechanisms to achieve consensus include Proof of Work (PoW), Proof of Stake (PoS), Delegated Proof-of-Stake (DPoS), Proof of Elapsed Time (PoET), etc.
[0102] 14) Multimedia information: including but not limited to: long videos (videos uploaded by users), short videos (videos with a length less than 1 minute uploaded by users), audio (such as mv with fixed pictures or records).
[0103] Figure 1 For the schematic diagram of the usage environment of the multimedia information processing method provided by the embodiments of the present invention, see Figure 1, different clients capable of performing different functions are set on the terminals (including terminal 10-1 and terminal 10-2). Among them, the terminals (including terminal 10-1 and terminal 10-2) obtain different video information for browsing from the corresponding servers 200 through different service processes via the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two, and uses a wireless link to achieve data transmission. Among them, the types of multimedia information obtained by the terminals (including terminal 10-1 and terminal 10-2) from the corresponding servers 200 through the network 300 are not the same. The multimedia information includes, but is not limited to: long videos (such as videos uploaded by users or existing videos that users need to conduct copyright verification on), short videos (such as videos with a length less than 1 minute uploaded by users), and audio (such as music videos with fixed pictures or records). For example, the terminals (including terminal 10-1 and terminal 10-2) can either obtain long videos (i.e., videos carrying video information or corresponding video links) from the corresponding servers 200 through the network 300, or obtain short videos for browsing from the corresponding servers 400 through the same video client or WeChat mini-program using the network 300. Different types of videos can be stored in the servers 200 and 400. Among them, in this application, the playback environments of different types of videos are not distinguished. During this process, the video information pushed to the user's client should be video information that complies with copyright regulations. Therefore, for a large number of videos, it is necessary to determine which videos are similar, and further conduct compliance detection on the copyright information of similar videos to avoid pushing duplicate or infringing video information.
[0104] Taking short videos as an example, the multimedia information processing model provided by the present invention can be applied to short video playback. In short video playback, different short videos from different data sources are usually processed, and finally, the multimedia information to be recommended corresponding to the corresponding user is presented on the user interface UI (User Interface). If the recommended video is an illegally broadcast video that does not comply with copyright regulations, it will directly affect the user experience. The background database of video playback receives a large amount of video data from different sources every day. The different videos obtained for multimedia information recommendation to the target user can also be called by other application programs (for example, the recommendation results of the short video recommendation process are migrated to the long video information recommendation process or the news recommendation process). Of course, the multimedia information processing model matching the corresponding target user can also be migrated to different multimedia information recommendation processes (such as the web multimedia information recommendation process, the mini-program multimedia information recommendation process, or the multimedia information recommendation process of the long video client).
[0105] Among them, the multimedia information processing method provided in the embodiments of this application is implemented based on artificial intelligence. Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0106] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0107] In the embodiments of this application, the artificial intelligence software technologies mainly involved include the above-mentioned speech processing technology and machine learning and other directions. For example, it may involve Automatic Speech Recognition (ASR) in Speech Technology, which includes Speech signal preprocessing, Speech signal frequency analyzing, Speech signal feature extraction, Speech signal feature matching / recognition, Speech training, etc.
[0108] For example, it may involve Machine Learning (ML). Machine learning is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning usually includes technologies such as Deep Learning. Deep learning includes artificial neural networks, such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and Deep Neural Network (DNN).
[0109] The structure of the electronic device according to the embodiments of the present invention will be described in detail below. The electronic device can be implemented in various forms, such as a terminal with multimedia information processing functions, such as a mobile phone running a video client. The trained multimedia information processing model can be encapsulated in the storage medium of the terminal, or it can be a server or a server group with multimedia information processing functions. The trained multimedia information processing model can be deployed in the server, such as the server 200 described above Figure 1 in the above. Figure 2 FIG. is a schematic diagram of the composition structure of the electronic device provided by the embodiments of the present invention. It can be understood that Figure 2 only shows the exemplary structure of the electronic device rather than all structures, and the partial structure or all structures shown can be implemented according to needs Figure 2 shown.
[0110] The electronic device provided by the embodiments of the present invention may include: at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. Each component in the electronic device 20 is coupled together through a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2 all kinds of buses are labeled as the bus system 205.
[0111] Among them, the user interface 203 may include a display, a keyboard, a mouse, a trackball, a click wheel, a button, a button, a touchpad, or a touch screen, etc.
[0112] It can be understood that the memory 202 can be a volatile memory, a non-volatile memory, or can include both volatile and non-volatile memories. The memory 202 in the embodiments of the present invention is capable of storing data to support the operation of a terminal (such as 10-1). Examples of such data include: any computer programs for operating on the terminal (such as 10-1), such as an operating system and application programs. Among them, the operating system contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs can include various application programs.
[0113] In some embodiments, the multimedia information processing device provided by the embodiments of the present invention can be implemented in a combination of software and hardware. As an example, the multimedia information processing device provided by the embodiments of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the multimedia information processing method provided by the embodiments of the present invention. For example, a processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs, Application Specific Integrated Circuits), DSPs, programmable logic devices (PLDs, Programmable Logic Devices), complex programmable logic devices (CPLDs, Complex Programmable Logic Devices), field-programmable gate arrays (FPGAs, Field-Programmable Gate Arrays), or other electronic components.
[0114] As an example of the multimedia information processing device provided by the embodiments of the present invention implemented in a combination of software and hardware, the multimedia information processing device provided by the embodiments of the present invention can be directly embodied as a combination of software modules executed by the processor 201. The software modules can be located in a storage medium, and the storage medium is located in the memory 202. The processor 201 reads the executable instructions included in the software modules in the memory 202 and combines with necessary hardware (for example, including the processor 201 and other components connected to the bus system 205) to complete the multimedia information processing method provided by the embodiments of the present invention.
[0115] As an example, the processor 201 can be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0116] As an example of the multimedia information processing device provided by the embodiments of the present invention implemented in hardware, the device provided by the embodiments of the present invention can be directly implemented by using a processor 201 in the form of a hardware decoding processor. For example, it can be implemented by one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components to execute the multimedia information processing method provided by the embodiments of the present invention.
[0117] The memory 202 in the embodiments of the present invention is used to store various types of data to support the operation of the electronic device 20. Examples of these data include: any executable instructions for operating on the electronic device 20, such as executable instructions, and the program implementing the multimedia information processing method of the embodiments of the present invention can be included in the executable instructions.
[0118] In some other embodiments, the multimedia information processing device provided by the embodiments of the present invention can be implemented in software. Figure 2 Shown is a multimedia information processing device 2020 stored in the memory 202, which can be software in the form of a program and a plug-in, etc., and includes a series of modules. As an example of the program stored in the memory 202, it can include a multimedia information processing device 2020. The multimedia information processing device 2020 includes the following software modules: an information transmission module 2081 and an information processing module 2082. When the software modules in the multimedia information processing device 2020 are read into the RAM by the processor 201 and executed, the multimedia information processing method provided by the embodiments of the present invention will be implemented. The functions of each software module in the multimedia information processing device 2020 are introduced below:
[0119] The information transmission module 2081 is used to obtain target multimedia information and parse the target multimedia information to separate the target audio included in the multimedia information.
[0120] The information processing module 2082 is used to perform conversion processing on the target audio to form a Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio.
[0121] The information processing module 2082 is configured to determine a first audio feature vector corresponding to the target audio based on a Mel spectrogram that matches the time domain features and frequency domain features of the target audio through a first sub-model network in the multimedia information processing model.
[0122] The information processing module 2082 is configured to determine a second audio feature vector corresponding to the target audio based on a Mel spectrogram that matches the time domain features and frequency domain features of the target audio through a second sub-model network in the multimedia information processing model.
[0123] The information processing module 2082 is configured to determine the type of the target audio in the target multimedia information based on the first audio feature vector and the second audio feature vector.
[0124] According to Figure 2 The electronic device shown, in one aspect of the present application, the present application further provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes different embodiments and combinations of embodiments provided in various alternative implementations of the above multimedia information processing method.
[0125] Combined with Figure 2 The electronic device 20 shown to illustrate the multimedia information processing method provided by the embodiments of the present invention. Before introducing the multimedia information processing method provided by the present invention, the defects of the related art are first introduced. In this process, although the existing video server can identify the similarity relationship between videos through the corresponding matching algorithm, with the popularization and development of video editing tools, the types of video frame attacks have become more complex. Refer to Figure 3 , Figure 3 is a schematic diagram of short video playback in the related art of the embodiments of the present invention. In Figure 3In the cropped video shown, simply relying on video image fingerprints is difficult to solve the problem of duplicate / infringing content in some videos with significant changes to the video footage. The audio in the video is speeded up, pitch-shifted, and multiple audio tracks are superimposed, making it very difficult to identify. In related technologies, the audio fingerprint algorithm can be used to compare the audio information in the video to determine whether the videos are similar. However, accurate identification cannot be achieved in the usage scenarios where the audio is speeded up, pitch-shifted, and multiple audio tracks are superimposed. For example, in the usage scenario of pitch-shifting attacks, since the landmark relies on the frequency peak points, and the pitch-shifted video changes the audio frequency, the generated hash will be different, resulting in a failure in similarity retrieval. Similarly, for the usage scenarios of speed-up / slow-down attacks, since the combined hash in the landmark depends on dt (t2 - t1), the change in dt due to speed-up or slow-down will cause the generated hash to be different.
[0126] To overcome the above defects, refer to Figure 4 , Figure 4 which is an optional flowchart of the multimedia information processing method provided by an embodiment of the present invention. It can be understood that Figure 4 the steps shown can be executed by various electronic devices running the multimedia information processing device. For example, it can be a terminal, a server, or a server cluster with multimedia information processing functions. When the multimedia information processing device runs in the terminal, it can trigger the WeChat mini-program in the terminal to perform multimedia information similarity detection. When the multimedia information processing device runs in the long video copyright detection server or the music playback software server, it can detect the corresponding long video copyright or music information copyright. The following will describe Figure 4 the steps shown.
[0127] Step 401: The multimedia information processing device obtains the target multimedia information and parses the target multimedia information to separate the target audio included in the multimedia information.
[0128] Among them, a data storage mapping can be established according to the information source of the target multimedia information; in response to the established data storage mapping, the file format of the target audio is adjusted to match the information source.
[0129] In some embodiments of the present invention, obtaining the target multimedia information and parsing the target multimedia information to separate the target audio included in the target multimedia information can be achieved in the following ways:
[0130] Parse the target multimedia information to obtain the timing information of the target multimedia information; according to the timing information of the target multimedia information, parse the video parameters corresponding to the target multimedia information to obtain the playback duration parameter and the audio track information parameter corresponding to the target multimedia information; based on the playback duration parameter and the audio track information parameter corresponding to the target multimedia information, extract the target multimedia information to obtain the target audio corresponding to the target multimedia information. Among them, the multimedia information to be processed can be sent to the multimedia information processing device by the client first. If the multimedia information is a long video information, the audio synchronization packet in the video data can be obtained first; then the corresponding playback duration parameter and the audio track information parameter can be obtained by parsing the audio header decoding data AACDecoderSpecificInfo and the audio data configuration information AudioSpecificConfig in the audio synchronization packet. Among them, the audio data configuration information AudioSpecificConfig is used to generate ADST (including the sampling rate, the number of channels, and the frame length data in the audio data). Based on the audio track information, other audio packets in the video data are obtained, and the original audio data is parsed. Finally, the AAC ES stream is packed into the ADTS format through the AAC decoder with the audio data header. Among them, a 7-byte header file ADTSheader is added before the AAC ES stream to achieve extraction to obtain the target audio corresponding to the target multimedia information.
[0131] Step 402: The multimedia information processing device performs conversion processing on the target audio to form a Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio.
[0132] In some embodiments of the present invention, the conversion processing of the target audio to form a Mel spectrogram corresponding to the target audio can be implemented in the following manner:
[0133] The target audio is subjected to channel conversion processing to form mono audio data; based on a windowing function corresponding to a multimedia information processing model, the mono audio data is subjected to short-time Fourier transform to form a corresponding spectrogram; a duration parameter corresponding to the multimedia information processing model is determined; and according to the duration parameter, the spectrogram is processed to form a Mel-spectrogram corresponding to the target audio. Taking the audio information processing in the video as an example, the audio can be first resampled to 16KHz single-channel audio; then a 25ms Hann time window, a 10ms frame shift, and a periodic Hann window are used to perform short-time Fourier transform on the audio to obtain the corresponding spectrogram; the mel spectrum is calculated by mapping the spectrogram to a 64-order mel filter group, where the range of mel bins is 125-7500Hz; log(mel-spectrum+0.01) is calculated to obtain a stable mel spectrum, and the added 0.01 bias is to avoid taking the logarithm of 0; the obtained features are framed with 0.96s features, and there is no frame overlap. Each frame contains 64 mel frequency bands and a duration of 10ms (96 frames in total), thereby extracting the corresponding mel spectrum.
[0134] Furthermore, when the audio data is converted into data in a Mel-spectrogram, since the unit of frequency is Hertz (Hz), the frequency range that the human ear can hear is 20-20000Hz, but the human ear does not have a linear perception relationship with the scale unit of Hz. For example, if people are adapted to a 1000Hz tone, if the tone frequency is increased to 2000Hz, the ear can only perceive that the frequency has increased a little bit, and it cannot be perceived that the frequency has doubled at all. If the ordinary frequency scale is converted to the Mel frequency scale, the human ear's perception of frequency becomes a linear relationship. In other words, under the Mel scale, if the Mel frequencies of two segments of speech differ by twice, the pitch that the human ear can perceive will also differ by about twice. Thus, a beneficial technical effect of visualizing audio data can be achieved.
[0135] Step 403: The multimedia information processing apparatus determines a first audio feature vector corresponding to the target audio based on a Mel-spectrogram matching the time domain features and frequency domain features of the target audio through a first sub-model network in the multimedia information processing model.
[0136] In some embodiments of the present invention, determining the first audio feature vector corresponding to the target audio based on the Mel-spectrogram matching the time domain features and frequency domain features of the target audio through the first sub-model network in the multimedia information processing model can be implemented in the following manner:
[0137] Convert the Mel spectrogram that matches the time-domain and frequency-domain features of the target audio into a corresponding grayscale image; according to the grayscale image, extract the feature vector of the Mel spectrogram through the convolutional neural network in the first sub-model network of the multimedia information processing model; through the gated recurrent unit in the first sub-model network, process the feature vector of the Mel spectrogram to determine the first audio feature vector corresponding to the target audio. Among them, based on the number of Mel spectrograms, determine the number of channels of the gated recurrent unit in the first sub-model network; according to the time-domain and frequency-domain features of the target audio, determine the time series parameters; based on the number of channels of the gated recurrent unit in the first sub-model network and the time series parameters, determine the recurrent neural network in the first sub-model network; determine the first audio feature vector corresponding to the target audio through the recurrent neural network in the first sub-model network.
[0138] Reference Figure 5 , Figure 5 FIG. is a schematic diagram of the processing process of the multimedia information processing model for audio in the embodiments of the present invention. Feature extraction can be performed through the VGGish network. Among them, the feature extraction of the multimedia information processing model can be implemented through the Visual Geometry Group network (VGGish, Visual Geometry Group). For example, for the audio information in the video, the audio file can be extracted to obtain the audio file. For the audio file, the corresponding Mel spectrogram is obtained, and then for the Mel spectrogram, audio feature extraction is performed through the Vggish network, and the extracted vector is clustered and encoded through the NetVlad (Net Vector of locally aggregated descriptors) to obtain the audio feature vector. NetVlad can save the distance between each feature point and the nearest cluster center and use it as a new feature.
[0139] Continuing with the example of long video processing for illustration, the VGGish network supports extracting 128-dimensional embedding feature vectors with semantics from the corresponding audio information. Specifically, it includes: converting the audio segment into a triple sample of the Mel spectrogram as the input of the VGGish model. Specifically, it includes: calculating the spectrogram of the audio segment using the signal amplitude, mapping the spectrogram of the audio segment into a 64-order Mel filter bank to calculate the Mel spectrogram, and obtaining N triple samples mapped from Hz to the Mel spectrogram, with a feature dimension of N*96*64; then, using the VGGish model based on TensorFlow as the audio feature extractor, taking the triple sample as the input, and using the VGGish network model for feature extraction to obtain the audio feature vector corresponding to the audio segment of N*128.
[0140] Step 404: The multimedia information processing device determines a second audio feature vector corresponding to the target audio based on a Mel spectrogram that matches the time-domain features and frequency-domain features of the target audio through a second sub-model network in the multimedia information processing model.
[0141] In some embodiments of the present invention, determining a second audio feature vector corresponding to the target audio based on a Mel spectrogram that matches the time-domain features and frequency-domain features of the target audio through a second sub-model network in the multimedia information processing model can be achieved in the following manner:
[0142] Based on a Mel spectrogram that matches the time-domain features and frequency-domain features of the target audio, the output information of the average pooling layer network is determined through a residual network in the second sub-model network in the multimedia information processing model; according to the output information of the average pooling layer network, the parameters of the image classification network in the second sub-model network are adjusted; through the image classification network in the second sub-model network, a second audio feature vector corresponding to the target audio is determined based on a Mel spectrogram that matches the time-domain features and frequency-domain features of the target audio.
[0143] Step 405: The multimedia information processing device determines the type of the target audio in the target multimedia information based on the first audio feature vector and the second audio feature vector.
[0144] Among them, vector fusion processing can be performed on the first audio feature vector and the second audio feature vector; based on the result of the vector fusion processing, the type of the target audio in the target multimedia information is determined, where the type of the target audio includes at least one of the following:
[0145] Compliant audio, accelerated and distorted audio, and multi-layer audio track superimposed audio.
[0146] Continue to refer to Figure 6 , Figure 6 which is an optional flowchart of the multimedia information processing method provided by the embodiments of the present invention. It can be understood that Figure 6 the steps shown can be executed by various electronic devices running the multimedia information processing device, such as a terminal, a server, or a server cluster with multimedia information processing functions. When the multimedia information processing device runs in a terminal, it can trigger a WeChat mini-program in the terminal to perform multimedia information similarity detection. When the multimedia information processing device runs in a short video copyright detection server or a music playing software server, it can detect the corresponding short video copyright or music information copyright. The following describes Figure 6 the steps shown.
[0147] Step 601: Determine the source multimedia information corresponding to the target multimedia information.
[0148] Step 602: Based on the target audio of the target multimedia information and the source audio of the source multimedia information, determine a corresponding set of inter-frame similarity parameters through the first audio feature vector and the second audio feature vector.
[0149] Step 603: Obtain the number of audio frames in the set of inter-frame similarity parameters that reach the similarity threshold.
[0150] Step 604: Determine the number of audio frames that reach the similarity threshold and compare it with the quantity. When it exceeds the quantity threshold, execute Step 605; otherwise, execute Step 606.
[0151] Step 605: Determine that the target multimedia information is similar to the source multimedia information, and prompt to provide copyright information.
[0152] Step 606: Determine that the target multimedia information is different from the source multimedia information, and enter the corresponding recommendation process.
[0153] In some embodiments of the present invention, when it is determined that the target multimedia information is similar to the source multimedia information, obtain the copyright information of the target multimedia information; determine the legality of the target multimedia information through the copyright information of the target multimedia information and the copyright information of the source multimedia information; when the copyright information of the target multimedia information and the copyright information of the source multimedia information are inconsistent, issue a warning message.
[0154] In some embodiments of the present invention, when it is determined that the target multimedia information is not similar to the source multimedia information, add the target multimedia information to the multimedia information source; sort the recall order of the multimedia information to be recommended in the multimedia information source; based on the sorting result of the recall order of the multimedia information to be recommended, perform multimedia information recommendation to the target user.
[0155] Continue to combine Figure 2 The electronic device 20 shown to illustrate the multimedia information processing method provided by the embodiments of the present invention. Refer to Figure 7 , Figure 7 is an optional flowchart of the multimedia information processing method provided by the embodiments of the present invention. It can be understood that Figure 7The steps shown can be executed by various electronic devices running a multimedia information processing device. For example, it can be a terminal, a server, or a server cluster with multimedia information processing capabilities. When the multimedia information processing device runs in a long video copyright detection server or a music playback software server to detect the corresponding long video copyright or music information copyright, the trained multimedia information processing can be deployed in the server to detect the similarity of the uploaded videos to determine whether to perform compliance detection on the copyright information of the videos. Of course, before deploying the multimedia information processing model, the multimedia information processing model needs to be trained, which specifically includes the following steps:
[0156] Step 701: Obtain a first training sample set, where the first training sample set is an audio sample in the video information collected by the terminal.
[0157] Step 702: Add noise to the first training sample set to form a corresponding second training sample set.
[0158] In some embodiments of the present invention, adding noise to the first training sample set to form a corresponding second training sample set can be achieved in the following manner:
[0159] Determine the dynamic noise type matching the usage environment of the multimedia information processing model; according to the dynamic noise type, add noise to the first training sample set to change the background noise, volume, or sampling rate of the audio samples in the first training sample set, and form a corresponding second training sample set. Among them, since audio information attacks include, but are not limited to: attacks by changing the audio frequency, attacks by changing the video speed, therefore, in the process of constructing the training sample set, an audio enhancement data set can be made according to common audio attack types. Common audio enhancement forms include: voice change, adding background noise, volume change, sampling rate change, sound quality change, etc. By setting different parameters, different enhanced audios can be obtained. It should be noted that in some embodiments of the present invention, the construction of the training sample set does not use the situation where the video duration changes or there is frame shift resulting in uneven frame pairs.
[0160] Make a training sample set according to the audio enhancement data. For example, one original audio corresponds to 20 attack audios. Here, the duration of each attack audio is the same as that of the original audio and there is no frame shift (that is, the audio at the corresponding time points is the same). The audio duration is dur, and with 0.96s as the step, each group of audio (original audio + corresponding attack audio) will generate dur / 0.96 labels, and the labels at the same time points are the same.
[0161] Step 703: Process the second training sample set through the multimedia information processing model to determine the initial parameters of the multimedia information processing model.
[0162] Step 704: In response to the initial parameters of the multimedia information processing model, process the second training sample set through the multimedia information processing model to determine the updated parameters of the multimedia information processing model.
[0163] In some embodiments of the present invention, in response to the initial parameters of the multimedia information processing model, processing the second training sample set through the multimedia information processing model to determine the updated parameters of the multimedia information processing model can be achieved in the following manner:
[0164] Substitute different audio samples in the second training sample set into the loss functions corresponding to the first sub-model network and the second sub-model network of the multimedia information processing model respectively; determine the parameters corresponding to the first sub-model network and the second sub-model network in the multimedia information processing model when the loss functions satisfy the corresponding convergence conditions; use the parameters corresponding to the first sub-model network and the second sub-model network respectively as the updated parameters of the multimedia information processing model.
[0165] Step 705: According to the updated parameters of the multimedia information processing model, iteratively update the network parameters of the multimedia information processing model through the second training sample set.
[0166] Specifically, determine the convergence conditions respectively matching the first sub-model network and the second sub-model network in the multimedia information processing model; iteratively update the parameters corresponding to the first sub-model network and the second sub-model network respectively until the loss functions corresponding to the first sub-model network and the second sub-model network satisfy the corresponding convergence conditions.
[0167] Figure 8 It is an optional schematic diagram for determining the type of the target audio in the embodiments of the present invention. Among them, during the playback process of the video, the displayed picture area of the video changes over time along the time axis, and there are different video targets in the displayed picture area. By processing the audio information through the multimedia information processing model, it is determined whether there is accelerated voice change or multi-layer audio track superposition, so as to realize the auxiliary review of a large number of short videos. Furthermore, based on the detection results of the video targets, it is determined whether the video to be detected is compliant or meets the copyright information requirements, and prevent the videos uploaded by users from being pirated.
[0168] As the number of multimedia information in the multimedia information server continues to increase, the copyright information of the multimedia information can be stored in a blockchain network or a cloud server to implement the judgment of the similarity of the multimedia information. Among them, the embodiments of the present invention can be implemented in combination with cloud technology or blockchain network technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing. It can also be understood as the general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. The background services of the technical network system require a large amount of computing and storage resources, such as multimedia information websites, picture websites, and more portal websites. Therefore, cloud technology needs to be supported by cloud computing.
[0169] It should be noted that cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool platform, abbreviated as a cloud platform, is generally referred to as Infrastructure as a Service (IaaS). Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (which can be virtual machines, including operating systems), storage devices, and network devices.
[0170] In some embodiments of the present invention, the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information can also be sent to the blockchain network so that
[0171] the nodes of the blockchain network fill the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information into a new block, and when consensus is reached on the new block, append the new block to the tail of the blockchain.
[0172] In the above solution, the method further includes:
[0173] Receive a data synchronization request from other nodes in the blockchain network; in response to the data synchronization request, verify the permissions of the other nodes; when the permissions of the other nodes pass the verification, control data synchronization between the current node and the other nodes, so as to enable the other nodes to obtain the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information.
[0174] In the above solution, the method further includes: in response to a query request, parse the query request to obtain the corresponding user identifier; according to the user identifier, obtain the permission information in the target block in the blockchain network; verify the matching degree between the permission information and the user identifier; when the permission information matches the user identifier, obtain the corresponding target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information in the blockchain network; in response to the query request, push the obtained corresponding target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information to the corresponding client, so as to enable the client to obtain the corresponding target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information saved in the blockchain network.
[0175] Continue to refer to Figure 9 , Figure 9 FIG. is a schematic diagram of the architecture of the blockchain network provided by an embodiment of the present invention, including a blockchain network 200 (exemplarily showing consensus nodes 210-1 to 210-3), an authentication center 300, a business entity 400, and a business entity 500, which will be described separately below.
[0176] The type of the blockchain network 200 is flexible and diverse. For example, it can be any one of a public chain, a private chain, or a consortium chain. Taking the public chain as an example, the electronic devices of any business entity, such as user terminals and servers, can access the blockchain network 200 without authorization; taking the consortium chain as an example, the electronic devices (such as terminals / servers) under the jurisdiction of a business entity can access the blockchain network 200 after obtaining authorization. At this time, they become client nodes in the blockchain network 200.
[0177] In some embodiments, the client node can only be an observer of the blockchain network 200, that is, it provides functions to support business entities to initiate transactions (for example, for storing data on the chain or querying data on the chain). For the functions of the consensus nodes 210 in the blockchain network 200, such as sorting functions, consensus services, and ledger functions, etc., the client node can implement them by default or selectively (for example, depending on the specific business needs of the business entity). Thus, the data and business processing logic of the business entity can be migrated to the blockchain network 200 to the greatest extent, and the trustworthiness and traceability of the data and business processing process can be achieved through the blockchain network 200.
[0178] The consensus nodes in the blockchain network 200 receive transactions submitted by client nodes of different business entities (such as the client node 410 belonging to the business entity 400 and the client node 510 belonging to the database operator system shown in the previous embodiments), execute the transactions to update the ledger or query the ledger, and various intermediate results or final results of the executed transactions can be returned to the client nodes of the business entity for display.
[0179] For example, the client nodes 410 / 510 can subscribe to events of interest in the blockchain network 200, such as transactions occurring in a specific organization / channel in the blockchain network 200. The consensus node 210 pushes corresponding transaction notifications to the client nodes 410 / 510, thereby triggering the corresponding business logic in the client nodes 410 / 510.
[0180] Taking the access of multiple business entities to the blockchain network to implement the management of instruction information and the business process matching the instruction information as an example, the exemplary application of the blockchain network is described below.
[0181] See Figure 9 , for the multiple business entities involved in the management link, such as the business entity 400 can be a multimedia information processing device, and the business entity 500 can be a display system with multimedia information processing function. They obtain their respective digital certificates by registering with the certification center 300. The digital certificate includes the public key of the business entity and the digital signature signed by the certification center 300 for the public key and identity information of the business entity. It is used to be attached to the transaction together with the digital signature of the business entity for the transaction and sent to the blockchain network, so that the blockchain network can take out the digital certificate and signature from the transaction, verify the reliability of the message (that is, whether it has been tampered with) and the identity information of the business entity sending the message. The blockchain network will verify according to the identity, such as whether it has the permission to initiate a transaction. The client running on the electronic device (such as a terminal or a server) under the business entity can request to access the blockchain network 200 and become a client node.
[0182] The client node 410 of the service entity 400 is used to send the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information to the blockchain network, so that the nodes of the blockchain network fill the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information into a new block, and when reaching a consensus on the new block, append the new block to the tail of the blockchain.
[0183] Among them, to send the corresponding target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information to the blockchain network 200, the service logic can be pre-set in the client node 410. When it is determined that the target multimedia information is not similar to the source multimedia information, the client node 410 automatically sends the target multimedia information identifier to be processed, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information to the blockchain network 200. It can also be that the business personnel of the service entity 400 log in to the client node 410, manually package the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, the type of the target audio of the target multimedia information, and the corresponding conversion process information, and send them to the blockchain network 200. When sending, the client node 410 generates a transaction for the corresponding update operation according to the target multimedia information identifier, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information. The smart contract to be called to implement the update operation and the parameters to be passed to the smart contract are specified in the transaction. The transaction also carries the digital certificate of the client node 410 and the signed digital signature (for example, encrypted using the private key in the digital certificate of the client node 410 for the digest of the transaction), and broadcasts the transaction to the consensus node 210 in the blockchain network 200.
[0184] When the consensus node 210 in the blockchain network 200 receives the transaction, it verifies the digital certificate and digital signature carried in the transaction. After successful verification, it confirms whether the service entity 400 has the transaction permission according to the identity of the service entity 400 carried in the transaction. Any verification judgment in the digital signature and permission verification will cause the transaction to fail. After successful verification, it signs its own digital signature (for example, encrypted using the private key of the consensus node 210-1 for the digest of the transaction), and continues to broadcast in the blockchain network 200.
[0185] After receiving a successfully verified transaction, the consensus node 210 in the blockchain network 200 fills the transaction into a new block and broadcasts it. When the consensus node 210 in the blockchain network 200 broadcasts a new block, a consensus process is performed on the new block. If the consensus is successful, the new block is appended to the tail of the blockchain stored by itself, and the state database is updated according to the result of the transaction, and the transaction in the new block is executed: for a transaction that submits an identification of a target multimedia information to be updated, a first audio feature vector of the target multimedia information, the first audio feature vector, the type of the target audio of the target multimedia information, and the corresponding process trigger information, a key-value pair including the identification of the target multimedia information, the first audio feature vector of the target multimedia information, the first audio feature vector, the type of the target audio of the target multimedia information, and the corresponding process trigger information is added to the state database.
[0186] The business personnel of the business entity 500 log in to the client node 510 and input a query request for the identification of the target multimedia information, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information. The client node 510 generates a transaction corresponding to an update operation / query operation according to the query request for the identification of the target multimedia information, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information. The smart contract to be called for implementing the update operation / query operation and the parameters to be passed to the smart contract are specified in the transaction. The transaction also carries the digital certificate of the client node 510 and the signed digital signature (for example, encrypted using the private key in the digital certificate of the client node 510 for the digest of the transaction), and broadcasts the transaction to the consensus node 210 in the blockchain network 200.
[0187] After receiving the transaction, the consensus node 210 in the blockchain network 200 verifies the transaction, fills the block, and reaches a consensus. Then, the filled new block is appended to the tail of the blockchain stored by itself, and the state database is updated according to the result of the transaction, and the transaction in the new block is executed: for a transaction that submits an updated manual recognition result corresponding to the copyright information data of a certain multimedia information, the key-value pair corresponding to the copyright information data of the multimedia information in the state database is updated according to the manual recognition result; for a transaction that submits a query for the copyright information data of a certain multimedia information, the key-value pair corresponding to the identification of the target multimedia information, the first audio feature vector of the target multimedia information, the first audio feature vector, and the type of the target audio of the target multimedia information is queried from the state database, and the transaction result is returned.
[0188] It should be noted that in Figure 9Exemplarily shown is the process of directly uploading the target multimedia information identifier, the first audio feature vector of the target multimedia information, the type of the first audio feature vector and the target audio of the target multimedia information, and the corresponding process trigger information to the chain. However, in some other embodiments, for the case where the data volume of the target multimedia information identifier, the first audio feature vector of the target multimedia information, the type of the first audio feature vector and the target audio of the target multimedia information is large, the client node 410 can upload the hash of the target multimedia information identifier, the first audio feature vector of the target multimedia information, the type of the first audio feature vector and the target audio of the target multimedia information and the corresponding hash in pairs, and store the target multimedia information identifier, the first audio feature vector of the target multimedia information, the type of the first audio feature vector and the target audio of the target multimedia information, and the corresponding process trigger information in a distributed file system or a database. After obtaining the target multimedia information identifier, the first audio feature vector of the target multimedia information, the type of the first audio feature vector and the target audio of the target multimedia information, and the corresponding process trigger information from the distributed file system or the database, the client node 510 can perform verification in combination with the corresponding hash in the blockchain network 200, thereby reducing the workload of the uploading operation.
[0189] As an example of the blockchain, refer to Figure 10 , Figure 10 is a schematic structural diagram of the blockchain in the blockchain network 200 provided by the embodiments of the present invention. The header of each block can include both the hash value of all transactions in the block and the hash value of all transactions in the previous block. After the record of the newly generated transaction is filled into the block and consensus is reached by the nodes in the blockchain network, it will be appended to the tail of the blockchain to form a chain-like growth. The chain-like structure based on the hash value between blocks ensures the anti-tampering and anti-counterfeiting of the transactions in the block.
[0190] The following describes the exemplary functional architecture of the blockchain network provided by the embodiments of the present invention. Refer to Figure 11 , Figure 11 is a schematic diagram of the functional architecture of the blockchain network 200 provided by the embodiments of the present invention, including an application layer 201, a consensus layer 202, a network layer 203, a data layer 204, and a resource layer 205, which will be described separately below.
[0191] The resource layer 205 encapsulates the computing resources, storage resources, and communication resources of each consensus node 210 in the blockchain network 200.
[0192] The data layer 204 encapsulates various data structures for implementing the ledger, including a blockchain implemented as files in a file system, a key-value state database, and proofs of existence (such as the hash tree of transactions in a block).
[0193] The network layer 203 encapsulates the functions of the peer-to-peer (P2P) network protocol, data dissemination mechanism, data verification mechanism, access authentication mechanism, and business entity identity management.
[0194] Among them, the P2P network protocol realizes the communication between consensus nodes 210 in the blockchain network 200. The data dissemination mechanism ensures the dissemination of transactions in the blockchain network 200. The data verification mechanism is used to achieve the reliability of data transmission between consensus nodes 210 based on cryptographic methods (such as digital certificates, digital signatures, public / private key pairs). The access authentication mechanism is used to authenticate the identity of business entities joining the blockchain network 200 according to the actual business scenario and grant the business entity the permission to access the blockchain network 200 when the authentication is passed. The business entity identity management is used to store the identities of business entities allowed to access the blockchain network 200 and their permissions (such as the types of transactions that can be initiated).
[0195] The consensus layer 202 encapsulates the mechanism for consensus nodes 210 in the blockchain network 200 to reach consensus on blocks (i.e., the consensus mechanism), as well as the functions of transaction management and ledger management. The consensus mechanism includes consensus algorithms such as POS, POW, and DPOS, and supports the pluggability of consensus algorithms.
[0196] Transaction management is used to verify the digital signatures carried in the transactions received by consensus nodes 210, verify the identity information of business entities, and determine whether they have the permission to conduct transactions based on the identity information (reading relevant information from the business entity identity management). For business entities that have obtained authorization to access the blockchain network 200, they all have digital certificates issued by the certification authority. The business entity uses the private key in its digital certificate to sign the submitted transaction, thereby declaring its legal identity.
[0197] Ledger management is used to maintain the blockchain and the state database. For the blocks that have reached consensus, append them to the end of the blockchain. Execute the transactions in the blocks that have reached consensus. When the transaction includes an update operation, update the key-value pairs in the state database. When the transaction includes a query operation, query the key-value pairs in the state database and return the query results to the client node of the business entity. Support various-dimensional query operations on the state database, including: querying blocks according to the block vector number (such as the hash value of a transaction); querying blocks according to the block hash value; querying blocks according to the transaction vector number; querying transactions according to the transaction vector number; querying the account data of a business entity according to the account (vector number) of the business entity; querying the blockchain in a channel according to the channel name.
[0198] The application layer 201 encapsulates various services that can be implemented by the blockchain network, including the traceability, storage, and verification of transactions, etc.
[0199] Thus, the copyright information of the target multimedia information identified through similarity recognition can be stored in the blockchain network. When a new user uploads multimedia information to the multimedia information server, the multimedia information server can call the copyright information in the blockchain network (at this time, the target multimedia information uploaded by the user can be used as the source multimedia information) to verify the copyright compliance of the multimedia information.
[0200] Figure 12 FIG. is a schematic diagram of the usage scenario of the multimedia information processing method provided by the embodiment of the present invention. Among them, the multimedia information is a short video. On the terminal (including terminal 10-1 and terminal 10-2), there is a client of software capable of displaying the corresponding short video, such as a client or plug-in for short video playback. Through the corresponding client, the user can obtain the target video and display it; the terminal is connected to the short video server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and uses a wireless link to implement data transmission. Of course, the user can also upload a video through the WeChat mini-program in the terminal for other users in the network to watch. In this process, the video server of the operator needs to detect the video uploaded by the user, compare and analyze different video information, determine whether the copyright of the video uploaded by the user is compliant, and recommend the compliant video to different users to avoid the unauthorized broadcast of the user's short video.
[0201] The present invention provides an information processing method. The following describes the usage process of the multimedia information processing method provided by the present invention. Among them, refer to Figure 13 , Figure 13 FIG. is a schematic diagram of an optional usage process of the multimedia information processing method in the embodiment of the present invention, which specifically includes the following steps:
[0202] Step 1301: Obtain the audio information corresponding to the target short video, and preprocess the audio information through a preprocessing process.
[0203] Step 1302: Obtain the training sample set of the multimedia information processing model.
[0204] Step 1303: Train the multimedia information processing model to determine the corresponding model parameters.
[0205] Step 1304: Deploy the trained multimedia information processing model in the corresponding video detection server.
[0206] Step 1305: Detect the audio in different video information through the multimedia information processing model to determine whether the audio of the target short video is compliant.
[0207] When it is determined that the audio of the target short video is compliant, obtain the copyright information of the target short video; for example, the video upload user uploads the corresponding copyright information through the WeChat mini-program run by the terminal 10-1, or the storage location of the video copyright information in the cloud server network. Determine the legality of the target short video through the copyright information of the target short video and the copyright information of the source video; when the copyright information of the target short video is inconsistent with the copyright information of the source video, send a warning message. At the same time, when it is determined that the target short video is not similar to the source video, add the target short video to the video source; sort the recall order of all the videos to be recommended in the video source; based on the sorting result of the recall order of the videos to be recommended, perform video recommendation to the target user, which is more conducive to the push of original videos.
[0208] Beneficial technical effects:
[0209] In each embodiment of the present invention, by obtaining target multimedia information and parsing the target multimedia information to separate the target audio included in the multimedia information; performing conversion processing on the target audio to form a Mel spectrogram that matches the time-domain characteristics and frequency-domain characteristics of the target audio; through the first sub-model network in the multimedia information processing model, based on the Mel spectrogram that matches the time-domain characteristics and frequency-domain characteristics of the target audio, determine the first audio feature vector corresponding to the target audio; through the second sub-model network in the multimedia information processing model, based on the Mel spectrogram that matches the time-domain characteristics and frequency-domain characteristics of the target audio, determine the second audio feature vector corresponding to the target audio; based on the first audio feature vector and the second audio feature vector, determine the type of the target audio in the target multimedia information. Thus, the type of the target audio in the target multimedia information can be determined, reducing the workload of manual review, improving the speed and accuracy of multimedia information review, and enhancing the user experience.
[0210] The above is only the embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multimedia information processing method, characterized in that, The method includes: Obtain target multimedia information, and parse the target multimedia information to separate the target audio included in the target multimedia information; Perform conversion processing on the target audio to form a Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio; Convert the Mel spectrogram into a corresponding grayscale image; According to the grayscale image, extract the feature vector of the Mel spectrogram through the convolutional neural network of the first sub-model network in the multimedia information processing model; Based on the number of the Mel spectrograms, determine the number of channels of the gated recurrent unit in the first sub-model network; According to the time domain characteristics and frequency domain characteristics of the target audio, determine the time series parameters; Based on the number of channels of the gated recurrent unit in the first sub-model network and the time series parameters, determine the recurrent neural network in the first sub-model network; Determine the first audio feature vector corresponding to the target audio through the recurrent neural network in the first sub-model network; Based on the Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio, determine the output information of the average pooling layer network through the residual network in the second sub-model network in the multimedia information processing model; According to the output information of the average pooling layer network, adjust the parameters of the image classification network in the second sub-model network; Through the image classification network in the second sub-model network, based on the Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio, determine the second audio feature vector corresponding to the target audio; Based on the first audio feature vector and the second audio feature vector, determine the type of the target audio in the target multimedia information, and the type of the target audio includes at least one of the following: compliant audio, accelerated and distorted audio, and multi-layer audio track superimposed audio.
2. The method according to claim 1, wherein The obtaining the target multimedia information and parsing the target multimedia information to separate the target audio included in the target multimedia information includes: Parse the target multimedia information to obtain the timing information of the target multimedia information; According to the timing information of the target multimedia information, parse the video parameters corresponding to the target multimedia information to obtain the playback duration parameter and the audio track information parameter corresponding to the target multimedia information; Based on the playback duration parameter and the audio track information parameter corresponding to the target multimedia information, extract the target multimedia information to obtain the target audio corresponding to the target multimedia information.
3. The method according to claim 1, wherein The performing conversion processing on the target audio to form a Mel spectrogram that matches the time domain characteristics and frequency domain characteristics of the target audio includes: Perform channel conversion processing on the target audio to form mono audio data; Based on the window function corresponding to the multimedia information processing model, perform short-time Fourier transform on the mono audio data to form a corresponding spectrogram; Determine the duration parameter corresponding to the multimedia information processing model; Process the spectrogram according to the duration parameter to form a Mel spectrogram that matches the time-domain and frequency-domain characteristics of the target audio.
4. The method according to claim 1, wherein The method further includes: Establish a data storage mapping according to the information source of the target multimedia information; In response to the established data storage mapping, adjust the file format of the target audio to match the information source.
5. The method according to claim 1, wherein The method further includes: Obtain a first training sample set, where the first training sample set is an audio sample in video information collected by a terminal; Add noise to the first training sample set to form a corresponding second training sample set; Process the second training sample set through a multimedia information processing model to determine the initial parameters of the multimedia information processing model; In response to the initial parameters of the multimedia information processing model, process the second training sample set through the multimedia information processing model to determine the updated parameters of the multimedia information processing model; According to the updated parameters of the multimedia information processing model, iteratively update the network parameters of the multimedia information processing model through the second training sample set.
6. The method according to claim 5, characterized in that, The adding noise to the first training sample set to form a corresponding second training sample set includes: Determine the dynamic noise type that matches the usage environment of the multimedia information processing model; Add noise to the first training sample set according to the dynamic noise type to change the background noise, volume, or sampling rate of the audio samples in the first training sample set, and form a corresponding second training sample set.
7. The method according to claim 5, characterized in that The responding to the initial parameters of the multimedia information processing model and processing the second training sample set through the multimedia information processing model to determine the updated parameters of the multimedia information processing model includes: Substitute different audio samples in the second training sample set into the loss functions corresponding to the first sub-model network and the second sub-model network of the multimedia information processing model respectively; Determine the parameters corresponding to the first sub-model network and the second sub-model network in the multimedia information processing model when the loss function meets the corresponding convergence condition; Use the parameters corresponding to the first sub-model network and the second sub-model network respectively as the updated parameters of the multimedia information processing model.
8. The method according to claim 5, wherein The iteratively updating the network parameters of the multimedia information processing model through the second training sample set according to the updated parameters of the multimedia information processing model includes: Determine the convergence conditions respectively matching the first sub-model network and the second sub-model network in the multimedia information processing model; Iteratively update the parameters corresponding to the first sub-model network and the second sub-model network respectively until the loss functions corresponding to the first sub-model network and the second sub-model network meet the corresponding convergence conditions.
9. The method according to claim 1, wherein The determining the type of the target audio in the target multimedia information based on the first audio feature vector and the second audio feature vector includes: Perform vector fusion processing on the first audio feature vector and the second audio feature vector; Based on the result of the vector fusion processing, determine the type of the target audio in the target multimedia information.
10. A multimedia information processing device, characterized in that, The device includes: An information transmission module, configured to obtain target multimedia information and parse the target multimedia information to separate the target audio included in the target multimedia information; An information processing module, configured to perform conversion processing on the target audio to form a Mel spectrogram that matches the time domain features and frequency domain features of the target audio; The information processing module is configured to convert the Mel spectrogram into a corresponding grayscale image; according to the grayscale image, extract the feature vector of the Mel spectrogram through the convolutional neural network of the first sub-model network in the multimedia information processing model; Based on the number of the Mel spectrograms, determine the number of channels of the gated recurrent unit in the first sub-model network; according to the time domain features and frequency domain features of the target audio, determine the time series parameters; based on the number of channels of the gated recurrent unit in the first sub-model network and the time series parameters, determine the recurrent neural network in the first sub-model network; determine the first audio feature vector corresponding to the target audio through the recurrent neural network in the first sub-model network; The information processing module is configured to, based on the Mel spectrogram that matches the time domain features and frequency domain features of the target audio, determine the output information of the average pooling layer network through the residual network in the second sub-model network in the multimedia information processing model; according to the output information of the average pooling layer network, adjust the parameters of the image classification network in the second sub-model network; through the image classification network in the second sub-model network, based on the Mel spectrogram that matches the time domain features and frequency domain features of the target audio, determine the second audio feature vector corresponding to the target audio; The information processing module is configured to, based on the first audio feature vector and the second audio feature vector, determine the type of the target audio in the target multimedia information, and the type of the target audio includes at least one of the following: compliant audio, accelerated and distorted audio, and multi-layer audio track superimposed audio.
11. An electronic device, characterized in that, The electronic device includes: A memory, configured to store executable instructions; A processor, configured to implement the multimedia information processing method according to any one of claims 1 to 9 when running the executable instructions stored in the memory.
12. A computer-readable storage medium storing executable instructions, characterized in that, When the executable instructions are executed by the processor, the multimedia information processing method according to any one of claims 1 to 9 is implemented.
13. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, the multimedia information processing method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Recording attack prevention voiceprint recognition method and device and access control system
CN108039176A
Video processing method and device, electronic equipment and storage medium
CN110213670A
System and method for tone recognition in spoken languages
CN112074903A
Multimedia information processing method and device, electronic equipment and storage medium
CN112104892A