Multimedia information processing method and device, electronic equipment, computer readable storage medium and computer program product

By introducing a pre-approval stage in multimedia information processing, using violation information labels and review notifications, the problems of long liberation cycle for user account complaints and high cost of manual review are solved, achieving more efficient review and better user experience.

CN120568097APending Publication Date: 2025-08-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410224998.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the prior art, the complaint relief period is long after the user account is suppressed, and the manual review cost is high, which affects the user experience.

Method used

The pre-appeal stage is introduced, and preset violation information tags are retrieved through multimedia information, and the review notification is sent to the target object, the review information is received and the operation is performed based on the results to avoid direct suppression of processing.

Benefits of technology

Shorten the appeal relief cycle, reduce manual review costs, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120568097A_ABST
    Figure CN120568097A_ABST
Patent Text Reader

Abstract

The invention provides a multimedia information processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps that multimedia information associated with a target object is acquired, and the multimedia information is used for being published to a social network; a plurality of preset violation information tags are retrieved based on the multimedia information, the violation information tags are configured for a pre-appeal stage, and the pre-appeal stage is a time period when it is detected that the multimedia information is violation information and the target object is not subjected to suppression processing; in response to the retrieved violation information tag matched with the multimedia information, sending an auditing notification to the target object; receiving auditing data of the target object; obtaining an audit result of the target object according to the audit data; and executing an operation for the target object according to an audit result. Through the method and the device, the auditing efficiency of the multimedia information can be improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to artificial intelligence technology and Internet technology, and in particular to a method, device, electronic device, computer-readable storage medium, and computer program product for processing multimedia information. Background Art

[0002] With the development of Internet technology, various instant messaging applications or servers have emerged, allowing users to log in to communicate. During this process, if the management system of the instant messaging application or social network server detects an abnormality in the multimedia messages posted by a user, it will suppress the user's account on the instant messaging application or social network server, preventing the user from using the instant messaging application or social network server to communicate.

[0003] In related technologies, after a user's account is suppressed, the user can submit review materials according to the review notice and make an appeal to release the account. However, the appeal release cycle is long and the efficiency is low, which affects the user experience. Summary of the Invention

[0004] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium, and computer program product for processing multimedia information, which can improve the review efficiency of multimedia information and enhance user experience.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] The present invention provides a method for processing multimedia information, the method comprising:

[0007] Acquiring multimedia information associated with a target object, wherein the multimedia information is used to be published to a social network;

[0008] Retrieving a plurality of preset violation information labels based on the multimedia information, wherein the violation information labels are configured for a pre-appeal phase, the pre-appeal phase being a period of time when the multimedia information is detected as violation information and suppression processing has not yet been performed on the target object;

[0009] In response to retrieving the violation information tag that matches the multimedia information, sending a review notification to the target object;

[0010] receiving audit information of the target object, wherein the audit information is sent by the target object in response to the audit notification;

[0011] Obtaining an audit result of the target object according to the audit data, wherein the audit result indicates whether the target object has passed the audit;

[0012] An operation is performed on the target object according to the audit result.

[0013] The present invention provides a multimedia information processing device, comprising:

[0014] an acquisition module, configured to acquire multimedia information associated with a target object, wherein the multimedia information is used to be published to a social network;

[0015] a retrieval module, configured to retrieve a plurality of preset violation information tags based on the multimedia information, wherein the violation information tags are configured for a pre-appeal phase, which is a period of time when the multimedia information is detected as violation information and the target object has not yet been suppressed;

[0016] a sending module, configured to send a review notification to the target object in response to retrieving the violation information tag matching the multimedia information;

[0017] a receiving module, configured to receive audit information of the target object, wherein the audit information is sent by the target object in response to the audit notification;

[0018] A processing module is configured to obtain an audit result of the target object based on the audit data, wherein the audit result indicates whether the target object has passed the audit; and perform an operation on the target object based on the audit result.

[0019] An embodiment of the present application provides an electronic device, comprising:

[0020] a memory for storing computer-executable instructions;

[0021] The processor is used to implement the multimedia information processing method provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.

[0022] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing a method for processing multimedia information provided in an embodiment of the present application when executed by a processor.

[0023] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the multimedia information processing method provided in the embodiment of the present application is implemented.

[0024] The embodiments of the present application have the following beneficial effects:

[0025] Before suppressing the target object, a pre-appeal stage is introduced. By retrieving multiple preset violation information tags based on the multimedia information sent by the target object, it is detected whether the multimedia information is in violation. If it is in violation, the target object is reminded to appeal by submitting review materials, and corresponding processing is performed according to the review results of the review materials. Since the target object has not been suppressed during this period, compared with the related technology of first suppressing the target object and then letting the target object appeal for relief, the user waiting time is reduced, the review cost is saved, and the user experience is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 1 is a schematic structural diagram of a multimedia information processing system 100 provided in an embodiment of the present application;

[0027] Figure 2 2 is a schematic diagram of the structure of the server 200 provided in an embodiment of the present application;

[0028] Figure 3A 1 is a flowchart of a method for processing multimedia information provided by an embodiment of the present application;

[0029] Figure 3B This is a flowchart of retrieving violation information provided by an embodiment of the present application;

[0030] Figure 3C This is a flowchart of the pre-appeal stage provided in an embodiment of the present application;

[0031] Figure 3D This is a flow chart of obtaining audit results based on audit materials provided in an embodiment of the present application;

[0032] Figure 3E This is a schematic diagram of a first process for obtaining an audit result by comparing information provided in an embodiment of the present application;

[0033] Figure 3F This is a second flow chart of obtaining an audit result by comparing information provided in an embodiment of the present application;

[0034] Figure 3G This is a flowchart of the processing related to the time limit for submitting review materials provided in the embodiment of the present application;

[0035] Figure 3H This is a flowchart of reviewing an appeal request provided by an embodiment of the present application;

[0036] Figure 3I This is a first flow chart of performing operations according to the audit results provided by an embodiment of the present application;

[0037] Figure 3J This is a second flow diagram of performing operations according to the audit results provided in an embodiment of the present application;

[0038] Figure 4A is a schematic diagram of extracting multimodal features of multimedia information provided by an embodiment of the present application;

[0039] Figure 4B This is a schematic diagram of extracting text information provided by an embodiment of the present application;

[0040] Figure 4C This is a schematic diagram of a social network structure provided by an embodiment of the present application;

[0041] Figure 5A This is a schematic diagram of the complaint relief process provided in an embodiment of the present application;

[0042] Figure 5B This is a schematic diagram of the pre-appeal process provided by the embodiment of the present application;

[0043] Figure 6 This is a schematic diagram of the face image alignment process provided by an embodiment of the present application;

[0044] Figure 7 This is a flow chart of the qualification review method provided in the embodiment of the present application;

[0045] Figure 8A This is a schematic diagram of the first process of self-help release provided by an embodiment of the present application;

[0046] Figure 8B This is a second flow diagram of the self-help release process provided by the embodiment of the present application;

[0047] Figure 8C This is a schematic diagram of the third process of self-help release provided by the embodiment of the present application;

[0048] Figure 8D This is a fourth process diagram of self-help release provided in an embodiment of the present application.

[0049] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0051] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0052] In the following description, the terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0053] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0054] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0055] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0056] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0057] 1) Creator: The publisher of multimedia information. Multimedia information can be any of video, audio, and images. For example, a video can be a short video. If a creator's multimedia information contains illegal content, it will be suppressed and a notification will be issued.

[0058] 2) Anchor: A natural person who broadcasts live. If the video stream contains any illegal content during the live broadcast, it will be suppressed and a notification will be issued.

[0059] 3) Release method: When the target object (such as the creator or anchor mentioned above) is suppressed, the method of requesting the server to release the suppression may include: submitting materials to the server for review and facial verification.

[0060] 4) Euclidean distance: Euclidean distance, also known as Euclidean distance, measures the absolute distance between two eigenvectors in a vector space.

[0061] 5) Cosine Distance: Also known as cosine similarity, this measure uses the cosine of the angle between two vectors in vector space as a measure of the difference between two individuals. If two vectors have the same direction and the angle is close to zero, then the two vectors are close. In machine learning, features are typically represented as vectors, so cosine similarity is often used when analyzing the similarity between two feature vectors.

[0062] 6) Feature extraction: Mapping data of a specific modality (such as images, audio, or text) to a vector space through a series of processes to obtain a vector representation of the data, which is convenient for computer processing.

[0063] 7) Fusion processing: The method of fusing multiple vectors may be, for example, direct concatenation, product (including inner product and outer product), averaging, weighted processing (such as weighted averaging and weighted summation), and the method to be used may be determined based on the actual application scenario.

[0064] 8) Violation information label: a label that identifies the violation of the video, which can be in the form of numbers or text. Different labels represent different types of violations, such as deception, insult, etc.

[0065] 9) Pre-appeal phase: After detecting that the target subject's multimedia information violates a policy and before suppressing the target subject, the target subject is allowed to appeal the violation by submitting review materials to remove the subsequent suppression operation. For example, after a social network server receives a video file uploaded by a user to be shared, if it detects that the video file violates a policy, the period from the moment the policy is detected to the moment the user is suppressed is called the pre-appeal phase. During this phase, the user can appeal the violation. If the appeal is successful, the server removes the subsequent suppression operation. If the appeal fails, the server will continue to suppress the user.

[0066] In the prior art, the process for appealing a user's account after it has been suppressed typically involves submitting review materials to the server in response to a server notification. Once the server verifies that no issues exist, the suppression is lifted. However, this approach has the following issues: 1) the appeals process is long, resulting in a poor user experience; 2) the manual review process in the prior art is costly and inefficient, impacting the user experience.

[0067] Based on the above analysis, the applicant found that the method of the related technology of first suppressing the user account and then notifying the user to submit review materials for appeal and relief has a high manual review cost and low efficiency, which easily affects the user experience. To address the above problems, the embodiment of the present application provides a method for processing multimedia information, which can reduce the manual review cost, improve the efficiency of appeal and relief processing, and enhance the user experience.

[0068] The embodiments of the present application provide a method, device, electronic device, computer-readable storage medium, and computer program product for processing multimedia information, which can improve the efficiency of reviewing multimedia information and enhance the user experience. The following describes exemplary applications of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and car terminals, and can also be implemented as servers. The following describes exemplary applications when the electronic device is implemented as a server.

[0069] See also Figure 1 , Figure 1 This is an architectural diagram of a multimedia information processing system 100 provided in an embodiment of the present application. To support a multimedia information processing application, a terminal 400 (a human-computer interaction interface 410 is shown as an example) is connected to a server 200 via a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0070] In some embodiments, the terminal 400 is used to obtain multimedia information, such as reading a pre-shot video file selected by the user at the terminal 400, or capturing the user's live video stream, and displaying the review notification and review results sent by the server 200 on the human-computer interaction interface 410; the server 200 is used to receive multimedia information sent by the user through the terminal 400; retrieve multiple violation information tags preset in the pre-appeal stage based on the multimedia information; send a review notification to the user in response to retrieving a violation information tag that matches the multimedia information; receive the review materials submitted by the user, and obtain the user's review results based on the review materials; and perform corresponding processing according to the review results.

[0071] In some embodiments, server 200 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence servers. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not prohibited in the embodiments of the present application.

[0072] See also Figure 2 , Figure 2 is a structural diagram of the server 200 provided in an embodiment of the present application, Figure 2 The server 200 shown includes: at least one processor 210, a memory 230 and at least one network interface 220. The various components in the server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 240 is not described in detail. Figure 2 Various buses are labeled as bus system 240 .

[0073] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0074] The memory 230 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 230 may optionally include one or more storage devices that are physically remote from the processor 210.

[0075] The memory 230 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 230 described in the embodiments of the present application is intended to include any suitable type of memory.

[0076] In some embodiments, memory 230 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0077] Operating system 231, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0078] A network communication module 232 for reaching other electronic devices via one or more (wired or wireless) network interfaces 220 , exemplary network interfaces 220 including Bluetooth, Wi-Fi, and Universal Serial Bus (USB);

[0079] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 The multimedia information processing device 233 stored in the memory 230 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: an acquisition module 2331, a retrieval module 2332, a sending module 2333, a receiving module 2334, and a processing module 2335. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.

[0080] In some embodiments, the server can implement the multimedia information processing method provided in the embodiments of the present application by running various computer-executable instructions or computer programs. For example, the computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. In general, the above-mentioned computer-executable instructions can be instructions in any form, and the above-mentioned computer program can be a program, module, or plug-in in any form.

[0081] The multimedia information processing method provided in the embodiments of the present application can be implemented using artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a manner similar to human intelligence. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, giving machines the functions of perception, reasoning, and decision-making.

[0082] The multimedia information processing method provided in the embodiment of the present application will be described in conjunction with the exemplary application and implementation of the electronic device provided in the embodiment of the present application.

[0083] The following describes the multimedia information processing method provided in the embodiment of the present application. As mentioned above, the electronic device that implements the multimedia information processing method in the embodiment of the present application can be a server, so the execution entity of each step will not be repeated below.

[0084] See also Figure 3A , Figure 3A This is a flowchart of a method for processing multimedia information provided by an embodiment of the present application, with the server as the execution subject, combining Figure 3A The steps shown are explained.

[0085] In step 101, multimedia information associated with a target object is obtained, wherein the multimedia information is used to be published to a social network.

[0086] In some embodiments, multimedia information may be in different data formats in different scenarios.

[0087] For example, in a video sharing scenario, multimedia information could be a pre-recorded video file, and the target audience is any user who posted the video file on a social network, such as the creator of the video file. In a live broadcast scenario, multimedia information could be a data stream of an unpublished live broadcast, intended for posting to a social network as a live broadcast, and the target audience is the live broadcaster. On the social network or live broadcast server, the target audience can be represented by a pre-registered user account.

[0088] In step 102, a plurality of preset violation information tags are retrieved based on the multimedia information, wherein the violation information tags are configured for the pre-appeal stage, which is a time period when the multimedia information is detected to be violation information and the target object has not yet been suppressed.

[0089] By adding a pre-appeal stage before suppressing the target object, the embodiment of the present application can shorten the appeal and release cycle compared to the related art of directly suppressing the target object corresponding to the violation information, and to a certain extent avoid the poor user experience caused by misjudgment and suppression of the target object.

[0090] As an example, in the case where the multimedia information is a video file, audio file or image posted to a social network, the timing for detecting the multimedia information may be after the target object posts the multimedia information to the social network, or after the target object submits the multimedia information and before the server posts it to the social network.

[0091] As an example, in a sharing scenario based on a social network, the multimedia information is a video file, audio file or image to be published on the social network. The timing for detecting the video file can be after receiving the multimedia information sent by the target object and publishing the multimedia information on the social network, or after receiving the multimedia information submitted by the target object and before publishing it on the social network.

[0092] As an example, in a live broadcast scenario, the multimedia information is a live stream. The timing for detecting the live stream can be periodic, and the detection can be performed at fixed intervals (for example, five minutes). The live stream can be detected for violations before the anchor uploads the live stream to the server and synchronizes it to the client of the live audience.

[0093] In some embodiments, the violation information tag represents the violation type, such as deception and verbal sarcasm. Figure 3B , Figure 3B This is a flowchart of retrieving violation information provided by an embodiment of the present application. Figure 3A Step 102 can be achieved by Figure 3B Steps 1021 to 1023 are implemented as described below.

[0094] In step 1021 , features of multiple modalities are extracted from the multimedia information, and the features of the multiple modalities are combined into multimodal features.

[0095] In some embodiments, multimodal features can be extracted using artificial intelligence technology. Artificial intelligence technology is an interdisciplinary discipline covering a wide range of fields, encompassing both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0096] In some embodiments, a feature extraction network can be used to extract features of multiple modalities from multimedia information, where the features of multiple modalities include features of text modality, image modality, and audio modality. Figure 4A , Figure 4A This is a schematic diagram of extracting multimodal features of multimedia information provided by an embodiment of the present application. Figure 4A The feature extraction of three modes of multimedia information, text mode, image mode and audio mode, is shown in the figure. Figure 4A Specific instructions.

[0097] like Figure 4A As shown, the first mode is the text mode. Figure 4A The text in the multimedia information shown includes the title text and the text in the video frame image. For the text in the multimedia information, feature extraction processing can be performed using a Bidirectional Encoder Representations from Transformer (BERT) model to obtain a text vector. The training process of the BERT model can include the pre-training stage and the fine-tuning stage. In the pre-training stage, the training tasks may include the Masked Language Model (MLM) task and the Next Sentence Prediction (NSP) task. The MLM task is used to mask part of the input text and then predict the masked text through the BERT model. The NSP task is used to determine whether sentence B is the next sentence of sentence A for the input sentences A and B. In the fine-tuning stage, it can include the Recognizing Textual Entailment (RTE) task, the Named Entity Recognition (NER) task and the Stanford Question Answering Dataset (SQuAD) task.

[0098] The second modality is the image modality, which involves extracting features from the video images in three dimensions: target (object), scene, and face. For example, the image can be subjected to feature extraction using a target detection model to obtain a target feature vector, where the target detection model can be a Single Shot MultiBox Detector (SSD) model; the image can be subjected to feature extraction using a scene recognition model to obtain a scene feature vector, where the scene recognition model can be a Vector of Locally Aggregated Descriptors (NetVLAD); and the image can be subjected to feature extraction using a face detection model to obtain a face feature vector, where the face detection model can be a Multi-task Convolutional Neural Network (MTCNN) model. Since the contribution of these three dimensions of information varies across different videos, after feature extraction, the target feature vector, scene feature vector, and face feature vector can be fused using a collaborative gating attention mechanism, i.e., weighted summation is performed to obtain an image vector. The mechanism of collaborative gate attention is actually equivalent to a multi-channel switch, which is used to adjust the weight of information in each dimension of the image. For example, if the video subject is a face, the weight of the face feature vector is increased accordingly, making it greater than the weights of the target feature vector and the scene feature vector; for example, if the video subject is a landscape, the weight of the scene feature vector is increased accordingly, making it greater than the weights of the target feature vector and the face feature vector.

[0099] In some embodiments, since there is a temporal correlation between the image vectors of different video frame images in the same video, the NetVLAD model can be used as an aggregation network for image vectors, wherein the NetVLAD model is used to perform weighted summation of the image vectors of multiple video frame images through learnable weights to obtain a global image vector.

[0100] The third modality is the audio modality. First, the audio signal is separated from the video and converted by calculating the Mel Frequency Cepstrum Coefficient (MFCC) features. Then, the VGGish model is used to extract features from the converted audio signal to obtain audio vectors corresponding to different video frame images. Similar to image vectors, the NetVLAD model can be used as an aggregation network for audio vectors. The NetVLAD model is used to perform weighted summation of audio vectors corresponding to multiple video frame images using learnable weights to obtain a global audio vector.

[0101] like Figure 4A As shown, after obtaining the image vector in the visual modality, the audio vector in the audio modality, and the text vector in the text modality, multimodal fusion processing is performed to obtain multimodal features corresponding to the multimedia information. Among them, the multimodal fusion processing methods include but are not limited to direct splicing, multiplication (including inner product and outer product), and weighted averaging.

[0102] For example, if the accuracy requirement for similar video recognition is high, that is, the requirement for accuracy is greater than the requirement for efficiency, then a direct splicing method can be used to obtain a longer video representation vector; if the efficiency requirement for similar video recognition is high, that is, the requirement for efficiency is greater than the requirement for accuracy, then a weighted average method can be used to obtain a shorter video representation vector at the cost of sacrificing a small amount of information. This can reduce the amount of computation for subsequent calculations and improve computational efficiency.

[0103] Continue to see Figure 3B , continue with step 1021 above for explanation.

[0104] In step 1022 , the similarities between the multimodal features and a plurality of preset violation information labels are determined.

[0105] In some embodiments, the similarities between the multimodal features and a plurality of preset violation information labels can be determined by calculating a similarity network.

[0106] In some embodiments, the multimodal features extracted from the multimedia information and the preset multiple violation information labels can be used as inputs to calculate the similarity network to obtain the feature vectors corresponding to the multimodal features and the feature vectors corresponding to the multiple violation information labels; by calculating the Euclidean distance between the feature vectors corresponding to the multimodal features and the feature vectors corresponding to the multiple violation information labels, the similarity between the multimodal features and the multiple violation information labels is obtained.

[0107] As an example, the preset multiple violation information labels include violation information label 1, violation information label 2, and violation information label 3. By calculating the similarity network, it is obtained that the multimodal features corresponding to the multimedia information have a similarity of 30% with violation information label 1, 50% with violation information label 2, and 90% with violation information label 3.

[0108] In step 1023 , in response to the maximum similarity being greater than the similarity threshold, the violation information label corresponding to the maximum similarity is determined as the violation information label that matches the multimedia information.

[0109] As an example, when the similarity threshold is 80%, the multimodal feature with the greatest similarity to the multimedia information in the above example is the violation information label 3, which is greater than the similarity threshold. Therefore, the violation information label 3 is determined as the violation information label that matches the multimedia information.

[0110] In some embodiments, a violation information tag includes at least one keyword of a violation type, that is, different violation information tags include keywords of different violation types. Figure 3A Step 102 can also be implemented in the following manner: comparing the illegal information tag with the multimedia information to obtain a comparison result, wherein the comparison result indicates whether the multimedia information matches the illegal information tag, that is, whether the multimedia information contains keywords corresponding to the illegal information tag; in response to the comparison result being that the multimedia information contains keywords corresponding to the illegal information tag, determining that the illegal information tag matches the multimedia information.

[0111] For example, when the multimedia information is a video file, the video frame image of the video file is first obtained. Text recognition is performed on the video frame image to obtain the text information corresponding to the video frame image. Then, the audio information of the video file is obtained and speech recognition is performed to obtain the text information corresponding to the audio information. The text information corresponding to the video frame image and the text information corresponding to the audio information are used as the text information of the multimedia information. Finally, the text information of the multimedia information is matched with the keywords corresponding to the violation information label. If there is a match, the violation information label corresponding to the matching keyword is determined as the violation information label that matches the multimedia information.

[0112] The embodiment of the present application determines the illegal information label that matches the multimedia information by performing multimodal feature extraction or keyword extraction on the multimedia information. Compared with the method of manually reviewing and determining the illegal information label in the related art, it can improve the accuracy of determining the illegal information label of the multimedia information.

[0113] Continue to see Figure 3A , continue with step 102 above for explanation.

[0114] In step 103, in response to retrieving the illegal information tag matching the multimedia information, a review notification is sent to the target object.

[0115] In some embodiments, different violation information labels represent different violation types. The audit notification is used to prompt the target object to submit audit materials corresponding to the violation type. The audit materials include the target object's information to be audited. Different violation types correspond to different audit materials.

[0116] As an example, when the server detects that multimedia information matches a violation information tag with an identity anomaly, it can prompt the user to provide review materials related to the user's identity as required, such as photos of specified actions; when the server detects that multimedia information matches a violation information tag with a qualification anomaly, it can prompt the user to submit review materials related to the qualification certificate as required.

[0117] In some embodiments, see Figure 3C , Figure 3C This is a flow chart of the pre-appeal stage provided by the embodiment of this application. Figure 3A Before step 103, you can Figure 3C Steps 201 to 202 implement the pre-appeal, which is described in detail below.

[0118] In step 201 , in response to multimedia information meeting conditions for immediate suppression processing, a first suppression processing is performed on a target object based on the violation information tag, and notification information corresponding to the first suppression processing is sent to the target object.

[0119] In some embodiments, the multimedia information meeting the condition of immediate suppression processing includes that the cumulative number of multimedia information matching the violation information tag posted by the target object reaches a violation number threshold.

[0120] As an example, when the preset violation number threshold is 5, and the cumulative number of multimedia information matching the violation information tag published by the target object is 6, the condition for immediate suppression processing is met.

[0121] In some embodiments, the multimedia information meeting the conditions for immediate suppression includes the multimedia information matching a violation information tag with a weight greater than a violation weight threshold, i.e., a serious violation. The weight of the violation information tag can be preset or based on the proportion of violation information tags detected within a current period.

[0122] As an example, when the preset violation weight threshold is 80% and the weight of the violation information label matched to the multimedia information is 85%, the condition for immediate suppression processing is met.

[0123] In some embodiments, the multimedia information satisfies the condition of immediate suppression processing, including that the influence of the target object in the social network is greater than a threshold. Figure 4C , Figure 4C This is a schematic diagram of a social network structure provided in an embodiment of the present application. Figure 4CThe figure shows a knowledge graph built based on a social network. Each node in the knowledge graph corresponds to a user. The edges connecting the nodes (i.e., each link) represent the frequency of contact between users. The influence link of the target object in the knowledge graph is found, that is, starting from the node corresponding to user 1, the number of nodes within a preset number of hops (for example, the number of hops from the node of the target object to the directly connected node is called 1 hop) is found; a weighted sum is performed based on the number of hops to the neighboring nodes, and the weight is the contact frequency recorded on the edge from the target node to the neighboring node.

[0124] like Figure 4C As shown in the figure, taking user 1 as the target object, when the preset number of hops is 2, the number of nodes reachable in the first hop is 3, corresponding to users 4, 5, and 7, with a contact frequency of 3; the number of nodes reachable in the second hop is 4, corresponding to users 3, 4, 5, and 6, with a contact frequency of 4. Therefore, when the preset number of hops is 2, the influence of the target object in the knowledge graph is the cumulative value of the product of each hop and the corresponding contact frequency, that is, 1 (representing the first hop) * 3 plus 2 (representing the second hop) * 4, that is, the influence of the target object (user 1) is 11. If the degree threshold is 10, the conditions for immediate suppression are met.

[0125] In some embodiments, the first suppression process includes at least one of the following:

[0126] The first is to issue a violation warning to the target object.

[0127] The second method is to set a login ban period for the target object. The period can be positively correlated with the number of illegal multimedia information released, the degree of influence of the target object, and the weight of the matched illegal information label.

[0128] For example, if the target object posts 5 illegal multimedia messages, the target object is prohibited from logging in for 1 hour; if the target object posts 10 illegal multimedia messages, the target object is prohibited from logging in for 2 hours.

[0129] The third method is to set a comment ban period for the target object. The period can be positively correlated with the number of illegal multimedia information released, the degree of influence of the target object, and the weight of the matched illegal information label.

[0130] For example, if the target object posts 5 illegal multimedia messages, the target object is prohibited from logging in and commenting for 1 hour; if the target object posts 10 illegal multimedia messages, the target object is prohibited from commenting for 2 hours.

[0131] The fourth method is to deduct points from the target object. The number of points deducted may be positively correlated with the number of illegal multimedia information released, the degree of influence of the target object, and the weight of the matched illegal information label.

[0132] For example, if the target object posts 5 illegal multimedia messages, 10 points will be deducted from the target object; if the target object posts 10 illegal multimedia messages, 20 points will be deducted from the target object.

[0133] As an example, when multimedia information meets the conditions for immediate suppression processing, based on the illegal information label, the target object can be set to be prohibited from logging in for 1 hour, prohibited from commenting for 1 hour, or deducted 10 points, etc.

[0134] In step 202, in response to the multimedia information not meeting the conditions for immediate suppression processing, the process proceeds to sending a review notification to the target object.

[0135] Continue to see Figure 3A , continue with step 103 above for explanation.

[0136] In step 104 , audit information of the target object is received, wherein the audit information is sent by the target object in response to an audit notification.

[0137] In some embodiments, the audit data sent by the target object in response to the audit notification is received to audit the audit data.

[0138] As an example, the display interface of the review notification will display the options of "Submit" and "Do not submit for now" for uploading review materials. In response to the user's trigger operation of "Submit", the terminal displays the review material template for the user to fill in. In response to the user's filling operation on the review material template, the content of the review material filled in by the user is obtained. In response to the user's submission operation, the review material is sent to the server via the terminal.

[0139] In step 105 , the audit result of the target object is obtained according to the audit data, wherein the audit result indicates whether the target object has passed the audit.

[0140] In some embodiments, the audit information includes the target object's pending audit information, and the audit notification is used to prompt the target object to submit the audit information corresponding to the violation type. Figure 3D , Figure 3D This is a flow chart of obtaining audit results based on audit materials provided in an embodiment of the present application. Figure 3A Step 105 can be achieved by Figure 3D Steps 1051 to 1052 are implemented as described below.

[0141] In step 1051, authentication reference information of the target object is obtained.

[0142] In some embodiments, the authentication reference information of the target object can be a real-time biometric image of the target object, such as a real-time face or palm print image of the target object; it can also be pre-authentication text information of the target object, such as a qualification certificate, which can be the target object's identity document or industry qualification certificate.

[0143] In some embodiments, the information to be reviewed includes a biometric image of the target object to be reviewed, and the authentication reference information includes a real-time biometric image of the target object. Acquiring the authentication reference information of the target object can be achieved by acquiring a real-time biometric image, wherein the real-time biometric image is obtained by collecting the biometric characteristics of the target object in real time using a terminal device.

[0144] As an example, the target object's biometric image to be reviewed includes an image containing biometric features such as the target object's face or palm print. The target object's authentication reference information includes a real-time biometric image of the target object, such as a real-time image of the target object's face or palm print.

[0145] In some embodiments, the information to be reviewed includes an image of the target object's certificate to be authenticated. Obtaining the authentication reference information of the target object can be achieved by querying the database for pre-authentication text information of the target object as the authentication reference information.

[0146] As an example, the target object's certificate to be authenticated in the information to be reviewed may be the target object's identity document or industry qualification certificate; the target object's pre-authentication text information may be the target object's qualification certificate.

[0147] In some embodiments, the pre-authentication certificate information can be uploaded to the database in advance by the target object through the terminal device. Before uploading, the target object can be authenticated by collecting the real-time biometric image of the target object and comparing it with the image in the certificate of the pre-authentication certificate information. If the comparison is consistent, the authentication is passed.

[0148] In step 1052, the target object's information to be reviewed is compared with the authentication reference information to obtain the review result of the target object.

[0149] In some embodiments, see Figure 3E , Figure 3E This is a schematic diagram of the first process of obtaining the audit result by comparing the information provided in the embodiment of the present application. When the information to be audited includes the biometric image to be audited of the target object, Figure 3D Step 1052 can be achieved by Figure 3ESteps 10521A to 10522A are implemented as described below.

[0150] In step 10521A, a first feature vector corresponding to the biometric image to be reviewed and a second feature vector corresponding to the real-time biometric image are obtained.

[0151] As an example, when the biometric image is a facial image of the target object, the first feature vector corresponding to the biometric image to be reviewed is the first feature vector corresponding to the facial image to be reviewed of the target object, and the second feature vector corresponding to the real-time biometric image is the second feature vector corresponding to the real-time facial image of the target object.

[0152] In step 10522A, the similarity between the first feature vector and the second feature vector is determined. If the similarity is greater than the similarity threshold, an audit result indicating compliance of the target object is generated. If the similarity is less than or equal to the similarity threshold, an audit result indicating violation of the target object is generated.

[0153] As an example, when the similarity threshold is 90%, if the similarity between the first feature vector and the second feature vector is 92%, an audit result indicating compliance of the target object is generated; if the similarity between the first feature vector and the second feature vector is 80%, an audit result indicating violation of the target object is generated.

[0154] In some embodiments, see Figure 3F , Figure 3F This is a second flow chart of comparing information to obtain audit results provided by the embodiment of the present application. When the information to be audited includes the image of the target object's ID to be authenticated, Figure 3D Step 1052 can be achieved by Figure 3F Steps 10521B to 10522B are implemented as described below.

[0155] In step 10521B, text recognition is performed on the image of the document to be authenticated to obtain text information.

[0156] In some embodiments, the image of the document to be authenticated can be pre-processed by an optical character recognition (OCR) model, and text detection can be performed to obtain a text area in the image of the document to be authenticated; text recognition can be performed on the text area to obtain text information. Figure 4B , Figure 4B is a schematic diagram of extracting text information provided by an embodiment of the present application, Figure 3F Step 10521B can be achieved by Figure 4B The method shown is implemented.

[0157] First, the document image to be authenticated undergoes image preprocessing. Image preprocessing typically addresses image imaging issues and includes geometric transformations (perspective, distortion, rotation, etc.), distortion correction, blur removal, image enhancement, and lighting correction. Next, text detection is performed on the preprocessed image. Text detection involves detecting the location, scope, and layout of text. This typically includes layout analysis and text line detection. Text detection primarily addresses the location and scope of text. Finally, text recognition is performed. Building on text detection, text recognition identifies text content and converts the text information in the image into text.

[0158] Continue to see Figure 3F , continue with step 10521B above for explanation.

[0159] In step 10522B, the text information is compared with the pre-authentication text information. If the comparison is consistent, an audit result indicating compliance of the target object is generated. If the comparison is inconsistent, an audit result indicating violation of the target object is generated.

[0160] In some embodiments, the comparison of text information with pre-authentication text information can be performed through keyword matching. First, keywords, such as name and ID serial number, are extracted from the text information of the document image to be authenticated. The extracted keywords are then compared with the pre-authentication text information. If the pre-authentication text information contains the keywords, the comparison is consistent, and a verification result indicating compliance with the target object is generated. If the comparison is inconsistent, a verification result indicating non-compliance with the target object is generated.

[0161] The embodiment of the present application uses a method of reviewing the review materials through a server, which reduces the waiting time for user complaints to be resolved, improves the review efficiency of the review materials, and reduces the manual review cost compared to the manual review method.

[0162] In some embodiments, Figure 3A Step 105 can also be implemented by: obtaining authentication reference information of the target object; sending audit materials and a violation information tag matching the multimedia information to a human agent device, so that the human agent device performs the following processing: comparing the audit materials with the authentication reference information of the target object; if the comparison is consistent, generating an audit result indicating compliance of the target object; if the comparison is inconsistent, generating an audit result indicating violation of the target object based on the violation information tag.

[0163] In some embodiments, the content involved in the authentication reference information of the audit material or template target object may be a password, name, photo, answers to preset questions, etc., or face, qualification certificate information, etc.

[0164] In some embodiments, the audit notice includes a deadline for submitting audit materials. Figure 3G , Figure 3G This is a flow chart of the relevant processing of the time limit for submitting audit materials provided by the embodiment of this application. Figure 3A Before step 103, different processing for different receiving times of audit materials can be achieved by Figure 3G Steps 301 to 302 are implemented as described below.

[0165] In step 301 , in response to the fact that the time of receiving the audit document exceeds the deadline, a second suppression process is performed on the target object, and notification information corresponding to the second suppression process is sent to the target object.

[0166] In some embodiments, the suppression level of the second suppression process is greater than the suppression level of the first suppression process that is triggered when the review materials are submitted within the deadline.

[0167] In some embodiments, the second suppression process includes setting a login prohibition time, setting a comment prohibition time, deducting points, and other aspects.

[0168] For example, if the deadline is 24:00 on March 1, 2024, and the review materials were received at 01:00 on March 2, 2024, the second suppression process will be performed on the target object. For example, the target object may be prohibited from logging in for 1 day, prohibited from commenting for 1 day, or deducted 100 points.

[0169] In some embodiments, see Figure 3H , Figure 3H This is a flowchart of the review of the appeal request provided by the embodiment of this application. Figure 3G After step 301, you can Figure 3H Steps 401 to 404 implement the review and processing of the appeal request, which is described in detail below.

[0170] In step 401, a complaint request from a target object is received.

[0171] In some embodiments, for the second suppression process, the target object may send a complaint request to the server via a terminal device.

[0172] As an example, the display interface of the notification information corresponding to the second suppression process will display two options: "Appeal" and "Do not appeal for now". In response to the user's trigger operation on "Appeal", the terminal displays an appeal request template for the user to fill in. In response to the user's filling operation on the appeal request template, the appeal request content filled in by the user is obtained. In response to the user's submission operation, the appeal request is sent to the server via the terminal.

[0173] In step 402, the appeal request is reviewed.

[0174] In some embodiments, the appeal request may be sent to a human agent device for review.

[0175] As an example, the appeal request includes supplementary review materials. The manual agent device reviews the supplementary review materials in the appeal request and compares the supplementary review materials with the authentication reference information of the target object. If the comparison is consistent, the suppression state of the target object after the second suppression processing is executed is released.

[0176] In some other embodiments, the appeal request may be reviewed by a machine review method.

[0177] As an example, the appeal request includes supplementary review materials. The supplementary review materials in the appeal request are reviewed by the server through a machine review method, and the supplementary review materials are compared with the authentication reference information of the target object. If the comparison is consistent, the suppression state of the target object after the second suppression processing is executed is released.

[0178] In step 403 , in response to the approval of the appeal request, the state of the target object after being subjected to the second suppression process is released.

[0179] In some embodiments, if the appeal request is approved, the suppression state of the target object after the second suppression process is executed is released.

[0180] In step 404 , in response to the appeal request failing to pass the review, the target object is maintained in the state after the second suppression process is executed.

[0181] In some embodiments, if the appeal request is not reviewed and approved, the target object continues to be in the suppressed state after the second suppression process is executed.

[0182] Continue to see Figure 3G , continue with step 301 above for explanation.

[0183] In step 302 , in response to the fact that the time of receiving the audit document does not exceed the deadline, the process proceeds to executing a process of obtaining the audit result of the target object based on the audit document.

[0184] In some embodiments, when the time of receiving the audit material does not exceed the deadline, the process proceeds to executing a process of obtaining the audit result of the target object based on the audit material.

[0185] As an example, if the deadline is 24:00 on March 1, 2024, and the audit materials are received at 12:00 on February 28, 2024, the process will proceed to obtain the audit results of the target object based on the audit materials.

[0186] Continue to see Figure 3A, continue with step 105 above for explanation.

[0187] In step 106, an operation is performed on the target object according to the audit result.

[0188] In some embodiments, see Figure 3I , Figure 3I This is a first flow chart of performing operations based on the audit results provided in an embodiment of the present application. Figure 3A Step 106 can be achieved by Figure 3I Steps 1061A to 1062A are implemented as described below.

[0189] In step 1061A, in response to the audit result indicating that the target object has not violated any regulations, the unexecuted suppression processing for the target object is released.

[0190] In some embodiments, if the audit result of the received audit data indicates that the target object has not violated any regulations, the suppression process that has not yet been executed on the target object is released.

[0191] In step 1062A, in response to the audit result indicating that the audit data violates the rules, a second suppression process is performed on the target object, and notification information corresponding to the second suppression process is sent to the target object.

[0192] As an example, if the audit result indicates that the audit data is in violation of regulations, the target object may be prohibited from logging in for 1 day, or prohibited from commenting for 1 day, or 100 points may be deducted, and corresponding notification information may be sent to the target object.

[0193] In some embodiments, see Figure 3J , Figure 3J This is a second flow chart of performing operations based on the audit results provided in an embodiment of the present application. Figure 3A Step 106 can be achieved by Figure 3J Steps 1061B to 1062B are implemented as described below.

[0194] In step 1061B, in response to the multimedia information having been published to the social network, if the audit result indicates that the target object is compliant, the multimedia information is kept accessible in the social network; if the audit result indicates that the target object is in violation, the multimedia information in the social network is set to an inaccessible state.

[0195] For example, if the multimedia information is a video file posted by the user in the example above, the video file will be reviewed after it is posted to the social network. If the review result indicates that the target object is compliant, the video file will remain accessible on the social network. If the review result indicates that the target object is not compliant, the video file will be made inaccessible or deleted from the social network.

[0196] In step 1062B, in response to the multimedia information not being published to the social network, if the audit result indicates that the target object is compliant, the multimedia information is kept published to the social network; if the audit result indicates that the target object is in violation, the multimedia information is blocked from being published to the social network.

[0197] For example, if the multimedia information is a live stream used by the host in the example above, it will be reviewed before being uploaded to the server and synchronized to the audience. If the review result indicates that the target audience complies with the requirements, the live stream will be published to the social network. If the review result indicates that the target audience violates the requirements, the live stream will be blocked from being published to the social network and synchronized to the audience.

[0198] Below, exemplary applications of the embodiments of the present application in actual application scenarios will be described.

[0199] The multimedia information processing method provided in the embodiment of the present application can be applied in the scenario where a creator in an instant messaging application or server publishes a short video, or in the scenario where a host in an instant messaging application or live broadcast server performs a live broadcast.

[0200] In the multimedia information processing method provided in the embodiment of the present application, a pre-appeal process is proposed, and machine review (referred to as machine review) can be used instead of manual review. At the same time, a lightweight means of relief is added to improve the efficiency of user complaint relief and enhance the user experience.

[0201] See also Figure 5A , Figure 5A This is a flowchart of the complaint relief process provided by the embodiment of this application. Figure 5A As shown, when a user posts a video that violates the rules, he or she will receive a review notification and be asked to submit review materials; when the review result of the review materials is passed, the penalty on the user will be lifted (suppression processing); when the review fails, the user will be subject to an upgraded penalty operation (escalation suppression processing); at this time, if the user submits an appeal and the review of the appeal is passed, the penalty on the user will be lifted; if the user does not submit an appeal, or the review of the appeal fails, the status after the upgraded penalty will be maintained.

[0202] Among them, escalated penalties are reflected in the length of login bans, comment bans, and point deductions. For example, if a video file posted by a user is detected to be in violation, the user will be subject to at least one of the following penalties: a one-hour login ban, a one-hour comment ban, or a deduction of 10 points. If the user submits review materials that fail the review, the user will be subject to at least one of the following escalated penalties: a one-day login ban, a one-day comment ban, or a deduction of 100 points.

[0203] To prevent users' accounts from being suppressed during the appeal release phase, a pre-appeal process has been introduced. This pre-appeal process means that before suppressing a user's account, the server will issue a warning and notify the user to submit review materials. If the user submits the review materials within the deadline and the review proves the compliance of their actions, the server will not continue to suppress the user's account. If the user does not submit the review materials corresponding to the review notice within the deadline, or the submitted review materials fail the review, the user's account will be suppressed again.

[0204] The server will require different review documents for different violation types. For example, if the server detects an identity anomaly, it will ask the user to provide a photo of a specific action as review document. The review document will be compared with the person in the video, and the review result will be determined based on the comparison. For another example, if the server detects an abnormality in the user's professional qualifications, it will ask the user to submit proof of qualifications as review document.

[0205] See also Figure 5B , Figure 5B This is a schematic diagram of the pre-appeal process provided by the embodiment of this application. Figure 5B As shown, the pre-appeal process involves an interactive process between a user terminal 501 and a backend server 502. When content or other actions posted by a user on the user terminal 501 match a pre-appeal tag configured by the backend server 502, the backend server 502 first determines whether to immediately penalize the user. If the backend server 502 immediately penalizes the user, a minor penalty is imposed. Otherwise, the backend server 502 issues a review notice, requiring the user to submit review materials. The user terminal 501 then receives the review notice, informing the user to submit the review materials. The review notice includes an X-day deadline for submitting the review materials. If the terminal submits the review materials within the deadline and passes manual review, the penalty is lifted. If the terminal submits the review materials within the deadline but fails manual review, the penalty is escalated, and a notification of the escalated penalty is displayed on the terminal and sent to the user, informing them of the escalated penalty. If the terminal fails to submit the review materials within the deadline, the penalty is escalated, and a notification of the escalated penalty is displayed on the terminal and sent to the user, informing them of the escalated penalty. The user can then submit an appeal through the terminal. If the appeal submitted by the terminal is approved, the penalty will be lifted. If the terminal does not submit an appeal or the appeal submitted by the terminal is not approved, the user's status after the upgraded penalty will be maintained.

[0206] In order to save the cost of manual review during the document review process and reduce the waiting time for users to obtain relief from their complaints, the embodiment of the present application adopts a solution of replacing manual review with machine review to improve the efficiency of complaint relief.

[0207] For example, when the backend server detects a user identity anomaly, it will issue a notification asking the user to verify that they are the same person as the person in the video. At this point, the backend system will compare the user's uploaded photos with the face of the person in the video to ensure consistency. To reduce labor costs during review, a machine review process will first compare the user's photos and the person in the video before submitting the video to a human reviewer. This is explained below.

[0208] First, face alignment.

[0209] Different images of the same person may show different postures and expressions, which is not conducive to facial feature extraction. Therefore, it is necessary to transform all facial images to a unified angle or posture, that is, face alignment. Figure 6 , Figure 6 This is a schematic diagram of the face image alignment process provided by the embodiment of the present application, and the following is a specific description using the server as the execution subject. Figure 6 As shown in the figure, first, face detection is performed on the face image; secondly, the key points in the face image are detected; then, based on the detected key points, the key points are rotated, scaled, and translated through similarity transformation technology to transform the face into a standard template face; finally, the aligned face image after transformation is obtained.

[0210] Then, facial features are extracted.

[0211] Based on the aligned facial images, deep learning methods are used to extract facial features. Facial feature extraction refers to the process of extracting compact and discriminative feature vectors, or feature templates, from facial images through calculation.

[0212] Finally, feature comparison.

[0213] The resulting feature vectors are then compared to determine whether the person in the user-submitted photo and video is the same person. Faces are extracted from video frames and compared with the faces in the user-submitted photos to determine their similarity. Appeals with similarity below a first similarity threshold are rejected directly by the machine review; those with similarity between the first and second similarity thresholds are sent to a manual review agent for re-examination. Appeals with similarity greater than or equal to the second similarity threshold are considered approved.

[0214] From the perspective of accurate recognition, feature vectors extracted from different photos of the same person are relatively close in feature space, while feature vectors extracted from photos of different people's faces are farther apart. First, the distance between the two feature vectors is calculated; then, the obtained distance between the feature vectors is compared with a preset threshold to determine the feature comparison result. The distance between the feature vectors can be calculated by calculating the Euclidean distance or cosine distance between the two feature vectors. The smaller the Euclidean distance, the higher the vector similarity; the smaller the cosine angle, that is, the larger the cosine distance, the higher the vector similarity.

[0215] For example, when the server determines that a user's published content requires qualification review, it will ask the user to upload proof of qualifications. However, if there are many proofs of qualifications, leaving all of them to the reviewer to judge will increase the review cost. Therefore, the embodiment of this application improves review efficiency and reduces the human cost of review by using machine review.

[0216] See also Figure 7 , Figure 7 This is a flow chart of the qualification review method provided by the embodiment of this application. Figure 7 The steps shown are explained.

[0217] In step 701, obtain qualification certificate.

[0218] Obtain the qualification certification materials uploaded by the user in response to the review notification, where the qualification certification materials include text information, image information, etc.

[0219] In step 702, the image is pre-processed.

[0220] Image preprocessing usually corrects imaging problems of the image. The image preprocessing process includes geometric transformation (perspective, distortion, rotation, etc.), distortion correction, blur removal, image enhancement and light correction.

[0221] In step 703, text detection is performed.

[0222] Text detection involves detecting the location, scope, and layout of text. This typically includes layout analysis and text line detection. The primary issues addressed by text detection are the location of text and its extent.

[0223] In step 704, text recognition is performed.

[0224] Text recognition is based on text detection and is the process of identifying text content, converting text information in an image into text. When the recognized content consists of words from a lexicon, it is called lexicon-based recognition; otherwise, it is called lexicon-free recognition.

[0225] In step 705, the audit result is determined.

[0226] The system compares the text extracted from the user's uploaded qualification certificate with the required qualification certificate text. The system determines the audit result based on whether the comparison results are consistent. Inconsistent results will be considered as failed audits. Qualification audits with doubtful machine-assessed results will be sent to a manual audit seat for further review. Inconsistent results will be considered as passed audits.

[0227] At the same time, in order to improve the efficiency of appeals, the embodiment of the present application provides a lightweight self-service appeal process. Users only need to complete the task as required, without going through the manual review process, to complete the appeal.

[0228] For example, see Figure 8A , Figure 8A This is a first flow chart of self-help relief provided by the embodiment of the present application. When a user submits a complaint relief application, Figure 8A As shown, the terminal interface will display a verification prompt area 801A and a "Learn More" function key 802A, wherein the verification prompt area 801A displays relevant information such as "Video number identity verification needs to be completed".

[0229] In response to the user's triggering operation of the "Learn Details" function key 802A, the terminal interface will jump to the identity verification interface for the video account application, see Figure 8B , Figure 8B This is a second flow diagram of the self-help release provided by the embodiment of the present application. Figure 8B A personal identity confirmation area 801B, a service authorization area 802B, and a "next" function key 803B are shown.

[0230] In response to the user's triggering operation of the "Next" function key 803B, the terminal interface jumps to the video authentication interface, see Figure 8C , Figure 8C This is a third flow diagram of the self-help release method provided in the embodiment of the present application. Figure 8C The review content prompt area 801C and the video extraction area 802C are shown, wherein the review content prompt area 801C prompts the user to perform a specific action, such as "please open your mouth", "please blink", etc.

[0231] When passing Figure 8C The video frame extracted from the video extraction area 802C in the server is reviewed and approved by the server. The terminal interface will jump to the appeal approval interface. Figure 8D , Figure 8D This is a fourth process diagram of self-help relief provided by an embodiment of the present application. Figure 8DThe prompt message "Appeal has been approved" is displayed, and the self-service appeal relief process is completed.

[0232] The embodiment of the present application adds a pre-appeal stage to the appeal relief process, retrieves multiple preset violation information tags based on the multimedia information sent by the target object, and sends an audit notification to the target object in response to the retrieval of the violation information tag matching the multimedia information, wherein the audit notification is used to prompt the target object to submit audit materials; obtain the audit results of the target object based on the audit materials; and perform corresponding processing based on the audit results, which can reduce the cost of manual review, improve the efficiency of appeal relief processing, and enhance the user experience.

[0233] The following continues to describe the exemplary structure of the multimedia information processing device 233 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the multimedia information processing device 233 of the memory 230 may include:

[0234] The acquisition module 2331 is configured to acquire multimedia information associated with the target object, wherein the multimedia information is used to be published to a social network.

[0235] The retrieval module 2332 is used to retrieve multiple preset violation information tags based on multimedia information, wherein the violation information tags are configured for the pre-appeal stage, which is the time period when the multimedia information is detected to be violation information and the target object has not yet been suppressed.

[0236] The sending module 2333 is configured to send a review notification to a target object in response to retrieving a violation information tag that matches the multimedia information.

[0237] The receiving module 2334 is configured to receive the audit information of the target object, wherein the audit information is sent by the target object in response to the audit notification.

[0238] The processing module 2335 is used to obtain the audit result of the target object based on the audit data, wherein the audit result indicates whether the target object has passed the audit; and perform corresponding processing based on the audit result.

[0239] In some embodiments, the audit materials include pending audit information of the target subject, and the audit notification prompts the target subject to submit audit materials corresponding to the violation type. Processing module 2335 is further configured to obtain authentication reference information of the target subject, compare the pending audit information of the target subject with the authentication reference information, and obtain an audit result for the target subject.

[0240] In some embodiments, the information to be reviewed includes a biometric image of the target object to be reviewed, and the authentication reference information includes a real-time biometric image of the target object. Processing module 2335 is further configured to obtain a real-time biometric image, where the real-time biometric image is obtained by real-time biometric collection of the target object by a terminal device; obtain a first feature vector corresponding to the biometric image to be reviewed and a second feature vector corresponding to the real-time biometric image; determine the similarity between the first feature vector and the second feature vector; if the similarity is greater than a similarity threshold, generate an audit result indicating compliance of the target object; if the similarity is less than or equal to the similarity threshold, generate an audit result indicating non-compliance of the target object.

[0241] In some embodiments, the information to be reviewed includes an image of the target subject's ID to be authenticated. Processing module 2335 is further configured to query a database for pre-authenticated text information of the target subject as authentication reference information; perform text recognition on the image of the ID to be authenticated to obtain text information; and compare the text information with the pre-authenticated text information. If the comparison is consistent, an audit result indicating compliance with the target subject is generated; if the comparison is inconsistent, an audit result indicating non-compliance with the target subject is generated.

[0242] In some embodiments, the processing module 2335 is further configured to obtain authentication reference information of the target object; and to send audit data and a violation information tag matching the multimedia information to a human agent device, so that the human agent device performs the following processing: comparing the audit data with the authentication reference information of the target object; if the comparison is consistent, generating an audit result indicating compliance of the target object; if the comparison is inconsistent, generating an audit result indicating violation of the target object based on the violation information tag.

[0243] In some embodiments, the processing module 2335 is also used to, in response to the multimedia information meeting the conditions for immediate suppression processing, perform a first suppression processing on the target object based on the violation information label, and send notification information corresponding to the first suppression processing to the target object; in response to the multimedia information not meeting the conditions for immediate suppression processing, proceed to the processing of sending an audit notification to the target object.

[0244] In some embodiments, the processing module 2335 is also used to release the suppression processing that has not yet been executed on the target object in response to the audit result indicating that the target object has not violated the rules; in response to the audit result indicating that the audit data has violated the rules, perform a second suppression processing on the target object, and send notification information corresponding to the second suppression processing to the target object.

[0245] In some embodiments, the audit notification includes a deadline for submitting audit materials. Processing module 2335 is further configured to, in response to the audit materials being received beyond the deadline, perform a second suppression process on the target object and send a notification message corresponding to the second suppression process to the target object; and, in response to the audit materials being received within the deadline, proceed to the process of obtaining the audit result of the target object based on the audit materials.

[0246] In some embodiments, the receiving module 2334 is also used to receive a complaint request from the target object; perform an audit operation on the complaint request; in response to the approval of the complaint request, release the target object from the state after the second suppression process is executed; in response to the failure of the audit of the complaint request, maintain the target object from the state after the second suppression process is executed.

[0247] In some embodiments, the violation information tag represents the violation type. Retrieval module 2332 is further configured to extract features of multiple modalities from the multimedia information, combine the features of the multiple modalities into a multimodal feature, determine the similarity between the multimodal features and a plurality of preset violation information tags, and, in response to the maximum similarity being greater than a similarity threshold, determine the violation information tag corresponding to the maximum similarity as the violation information tag that matches the multimedia information.

[0248] In some embodiments, the processing module 2335 is also used to, in response to the multimedia information having been published to the social network, if the audit result indicates that the target object is compliant, keep the multimedia information accessible in the social network; if the audit result indicates that the target object is in violation, set the multimedia information in the social network to an inaccessible state; in response to the multimedia information not having been published to the social network, if the audit result indicates that the target object is compliant, keep the multimedia information published to the social network; if the audit result indicates that the target object is in violation, block the multimedia information from being published to the social network.

[0249] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the multimedia information processing method described in the embodiment of the present application.

[0250] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the multimedia information processing method provided by the embodiment of the present application, for example, Figure 3A The method for processing multimedia information is shown.

[0251] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0252] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0253] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0254] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0255] In summary, the embodiment of the present application introduces a pre-appeal stage before suppressing the target object, and detects whether the multimedia information violates the rules by retrieving multiple preset violation information tags based on the multimedia information sent by the target object. If it violates the rules, the target object is reminded to submit an appeal by submitting review materials, so that corresponding processing can be performed according to the review results of the review materials. Since the target object has not been suppressed during this period, compared with the related technology of suppressing the target object first and then having the target object appeal for relief, the user waiting time is reduced, the review cost is saved, and the user experience is improved. The method of reviewing the review materials by the server reduces the waiting time for the user to appeal and release the review, improves the review efficiency of the review materials, and reduces the manual review cost.

[0256] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A method for processing multimedia information, characterized in that: The method comprises: Acquiring multimedia information associated with a target object, wherein the multimedia information is used to be published to a social network; Retrieving a plurality of preset violation information labels based on the multimedia information, wherein the violation information labels are configured for a pre-appeal phase, the pre-appeal phase being a period of time when the multimedia information is detected as violation information and suppression processing has not yet been performed on the target object; In response to retrieving the violation information tag that matches the multimedia information, sending a review notification to the target object; receiving audit information of the target object, wherein the audit information is sent by the target object in response to the audit notification; Obtaining an audit result of the target object according to the audit data, wherein the audit result indicates whether the target object has passed the audit; An operation is performed on the target object according to the audit result.

2. The method according to claim 1, characterized in that The audit information includes the target object's pending audit information, and the audit notification is used to prompt the target object to submit the audit information corresponding to the violation type; The obtaining of the audit result of the target object according to the audit data includes: Obtaining authentication reference information of the target object; The to-be-audited information of the target object is compared with the authentication reference information to obtain an audit result of the target object.

3. The method according to claim 2, characterized in that The information to be reviewed includes a biometric image of the target object to be reviewed, and the authentication reference information includes a real-time biometric image of the target object; The obtaining of the authentication reference information of the target object includes: Acquiring the real-time biometric image, wherein the real-time biometric image is obtained by collecting the real-time biometric characteristics of the target object through a terminal device; The step of comparing the information to be reviewed of the target object with the authentication reference information to obtain the review result of the target object includes: Obtaining a first feature vector corresponding to the to-be-verified biometric image and a second feature vector corresponding to the real-time biometric image; Determine the similarity between the first feature vector and the second feature vector. If the similarity is greater than a similarity threshold, generate an audit result indicating compliance of the target object. If the similarity is less than or equal to the similarity threshold, generate an audit result indicating violation of the target object.

4. The method according to claim 2, characterized in that The information to be reviewed includes the image of the target object's certificate to be authenticated; The obtaining of the authentication reference information of the target object includes: Querying the pre-authentication text information of the target object from a database as authentication reference information; The step of comparing the information to be reviewed of the target object with the authentication reference information to obtain the review result of the target object includes: Performing text recognition on the image of the document to be authenticated to obtain text information; The text information is compared with the pre-authentication text information. If the comparison is consistent, an audit result indicating that the target object is compliant is generated. If the comparison is inconsistent, an audit result indicating that the target object is in violation of the regulations is generated.

5. The method according to claim 1, wherein The obtaining of the audit result of the target object according to the audit data includes: Obtaining authentication reference information of the target object; Sending the audit information and the violation information tag that matches the multimedia information to a human agent device, so that the human agent device performs the following processing: The audit data is compared with the authentication reference information of the target object. If the comparison is consistent, an audit result indicating compliance of the target object is generated. If the comparison is inconsistent, an audit result indicating violation of the target object is generated based on the violation information tag.

6. The method according to any one of claims 1 to 5, characterized in that Before sending the audit notification to the target object, the method further includes: In response to the multimedia information meeting the conditions for immediate suppression processing, performing a first suppression processing on the target object based on the violation information tag, and sending notification information corresponding to the first suppression processing to the target object; In response to the multimedia information not meeting the conditions for immediate suppression processing, the process proceeds to sending an audit notification to the target object.

7. The method according to any one of claims 1 to 5, characterized in that The performing of an operation on the target object according to the audit result includes: In response to the audit result indicating that the target object has not violated any regulations, releasing the unexecuted suppression process for the target object; In response to the audit result indicating that the audit data violates the rules, a second suppression process is performed on the target object, and notification information corresponding to the second suppression process is sent to the target object.

8. The method according to any one of claims 1 to 5, characterized in that The review notice includes the deadline for submitting the review materials; Before obtaining the audit result of the target object according to the audit data, the method further includes: In response to the receipt time of the audit document exceeding the deadline, performing a second suppression process on the target object, and sending notification information corresponding to the second suppression process to the target object; In response to the fact that the time of receiving the audit material does not exceed the deadline, the process proceeds to executing the process of obtaining the audit result of the target object based on the audit material.

9. The method according to claim 8, characterized in that After performing the second suppression process on the target object and sending notification information corresponding to the second suppression process to the target object, the method further includes: receiving a complaint request from the target object; Reviewing the appeal request; In response to a review of the appeal request being passed, releasing the target object from the state in which the second suppression process was executed; In response to a failure in reviewing the appeal request, the target object is kept in a state after being subjected to the second suppression process.

10. The method according to any one of claims 1 to 5, characterized in that The violation information label represents the violation type; The plurality of violation information tags preset based on the multimedia information retrieval include: Extracting features of multiple modalities from the multimedia information, and combining the features of the multiple modalities into a multimodal feature; Determining similarities between the multimodal features and a plurality of preset violation information labels; In response to the maximum similarity being greater than a similarity threshold, the violation information label corresponding to the maximum similarity is determined as the violation information label that matches the multimedia information.

11. The method according to any one of claims 1 to 5, characterized in that The performing of an operation on the target object according to the audit result includes: In response to the multimedia information having been published on the social network, if the audit result indicates that the target object is compliant, maintaining the multimedia information in an accessible state in the social network; if the audit result indicates that the target object is in violation, setting the multimedia information in the social network to an inaccessible state; In response to the multimedia information not having been published to the social network, if the audit result indicates that the target object is compliant, the multimedia information is kept published to the social network; if the audit result indicates that the target object is in violation, the multimedia information is blocked from being published to the social network.

12. A multimedia information processing device, characterized in that: The device comprises: an acquisition module, configured to acquire multimedia information associated with a target object, wherein the multimedia information is used to be published to a social network; a retrieval module, configured to retrieve a plurality of preset violation information tags based on the multimedia information, wherein the violation information tags are configured for a pre-appeal phase, which is a period of time when the multimedia information is detected as violation information and the target object has not yet been suppressed; a sending module, configured to send a review notification to the target object in response to retrieving the violation information tag matching the multimedia information; a receiving module, configured to receive audit information of the target object, wherein the audit information is sent by the target object in response to the audit notification; A processing module is configured to obtain an audit result of the target object based on the audit data, wherein the audit result indicates whether the target object has passed the audit; and perform an operation on the target object based on the audit result.

13. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions; The processor is configured to implement the multimedia information processing method according to any one of claims 1 to 11 when executing the computer-executable instructions stored in the memory.

14. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the method for processing multimedia information according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the method for processing multimedia information according to any one of claims 1 to 11 is implemented.