Multimedia data processing method and device

By performing feature extraction and encoder processing on multimedia data, the confounding factor features are removed, and the problem of confounding variables in multimedia data affecting the accuracy of causal inference is solved, and more accurate multimedia data processing results are achieved.

CN119940520AActive Publication Date: 2025-05-06ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510421898.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

In multimedia data, there are mutually indistinguishable confounding variables that affect the accuracy of causal inference, resulting in false correlations between variables.

Method used

By extracting the multimedia data feature, multimedia features are obtained, and these features are processed by an encoder to obtain confounding factor features. Based on these features, the characteristics of noise factors are removed from the multimedia features to obtain standard features, thereby determining the target processing results of the multimedia data.

Benefits of technology

By eliminating the influence of confounding variables, the accuracy of causal inference is improved and the accuracy of multimedia data processing results is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940520A_ABST
    Figure CN119940520A_ABST
Patent Text Reader

Abstract

The invention provides a multimedia data processing method and device which can be applied to the technical field of computers. The method comprises the steps that feature extraction is carried out on multimedia data to obtain multimedia features of the multimedia data, the multimedia features comprise features of a plurality of different factors of the multimedia data, and the factors comprise noise factors; using an encoder to process the multimedia features to obtain confounding factor features, the confounding factor features including features of noise factors, and the noise factors including factors affecting causal relationship measurement between multimedia data and results; based on the hybrid factor features, removing features of noise factors from the multimedia features to obtain standard features of the multimedia data; and determining a target processing result of the multimedia data based on the standard features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for processing multimedia data. Background Art

[0002] In the field of data analysis, causal inference of multimedia data can determine the direct impact of one variable on another variable from a large number of complex variable relationships. Therefore, the results of causal inference can guide decision-making, such as providing optimization solutions for resource allocation and revealing potential problems. Therefore, causal inference is an important research direction in the field of data analysis. With the development of modern data analysis and computer technology, the objects of data analysis are no longer limited to numbers and texts. Artificial intelligence models can be used to analyze multimedia data such as video and voice that are difficult for traditional computers to process and analyze.

[0003] A common method of causal inference is to describe the causal relationship between variables through a causal graph. A causal graph is a directed acyclic graph in which nodes represent variables and edges represent causal relationships. In real data, the causes of a result are usually complex, and there may be confounding variables that affect both the cause and the result. In the presence of confounding variables in multimedia data, traditional correlation analysis methods such as causal graphs can only reveal the statistical relationship between variables, and it is difficult to accurately determine the causal relationship. Summary of the invention

[0004] In view of the above problems, the present invention provides a method and device for processing multimedia data.

[0005] According to a first aspect of the present invention, a method for processing multimedia data is provided, comprising: extracting features from multimedia data to obtain multimedia features of the multimedia data, wherein the multimedia features include features of respective factors of a plurality of different factors of the multimedia data, the factors including noise factors; processing the multimedia features using an encoder to obtain confounding factor features, wherein the confounding factor features include features of noise factors, and the noise factors include factors that affect the measurement of the causal relationship between the multimedia data and a result; based on the confounding factor features, removing features of the noise factors from the multimedia features to obtain standard features of the multimedia data; and determining a target processing result of the multimedia data based on the standard features.

[0006] According to an embodiment of the present invention, based on the confounding factor features, the features of the noise factors are removed from the multimedia features to obtain the standard features of the multimedia data, including: extracting features from the confounding factor features and the multimedia features to determine the joint features that characterize the correlation influence between the multimedia features and the confounding factor features; based on the joint features and the confounding factor features, determining the conditional features for removing the features of the noise factors; and determining the standard features using the conditional features and the multimedia features.

[0007] According to an embodiment of the present invention, feature extraction is performed on multimedia data to obtain multimedia features of the multimedia data, including: using a dimensionality reduction matrix to perform dimensionality reduction processing on the multimedia data to obtain the multimedia features.

[0008] According to an embodiment of the present invention, a target processing result of multimedia data is determined based on standard features, including: using an inverse matrix of a dimensionality reduction matrix to process the standard features to obtain standard multimedia data corresponding to the standard features; and determining the target processing result based on the standard multimedia data.

[0009] According to an embodiment of the present invention, the encoder is trained in the following manner: feature extraction is performed on sample multimedia data to obtain sample multimedia features of the sample multimedia data, wherein the sample multimedia features include features of multiple different sample factors of the sample multimedia data; the sample multimedia features are processed using the original encoder to obtain a sample multimedia encoding result; based on the discriminator's recognition result of the sample multimedia encoding result and the disordered sample multimedia encoding result, the parameters of the original encoder are adjusted to obtain a first encoder, wherein the disordered sample multimedia encoding result is obtained by disordering the sample multimedia encoding result; based on the sample multimedia decoding result and the sample multimedia features, the parameters of the first encoder are adjusted to obtain an encoder, wherein the sample multimedia decoding result is obtained by decoding the sample multimedia encoding result.

[0010] According to an embodiment of the present invention, the method for processing multimedia data also includes: performing random processing on sample multimedia coding to obtain a random sample multimedia coding result; using a discriminator to identify the sample multimedia coding result and the random sample multimedia coding result to obtain a recognition result, wherein the recognition result is used to represent the distance between the sample multimedia coding result and the random sample multimedia coding result; and using a decoder to process the coding result to obtain a sample multimedia decoding result.

[0011] According to an embodiment of the present invention, a discriminator is used to identify sample multimedia coding results and out-of-order sample multimedia coding results to obtain a recognition result, including: using a discriminator to process the sample multimedia coding results and the out-of-order sample multimedia coding results to obtain a sample multimedia judgment result and an out-of-order sample multimedia judgment result, respectively; and determining the recognition result based on a cross entropy result determined according to the sample multimedia judgment result and the out-of-order sample multimedia judgment result.

[0012] According to an embodiment of the present invention, based on the recognition result of the discriminator on the sample multimedia encoding result and the disordered sample multimedia encoding result, the parameters of the original encoder are adjusted to obtain a first encoder, including: determining the gradient direction and gradient value of the original encoder based on the recognition result; reversing the gradient direction to obtain a reversed gradient direction; and adjusting the parameters of the original encoder based on the reversed gradient direction and the gradient value until the recognition result indicates that the distance between the sample multimedia encoding result and the disordered sample multimedia encoding result is less than a preset value.

[0013] According to an embodiment of the present invention, based on the sample multimedia decoding results and the sample multimedia features, the parameters of the first encoder are adjusted to obtain the encoder, including: determining the loss value based on the sample multimedia decoding results and the sample multimedia features; and adjusting the parameters of the first encoder based on the loss value until the loss value converges.

[0014] The second aspect of the present invention provides a multimedia data processing device, including: a feature extraction module, used to extract features from multimedia data to obtain multimedia features of the multimedia data, wherein the multimedia features include features of multiple different factors of the multimedia data, and the factors include noise factors; a feature processing module, used to use an encoder to process the multimedia features to obtain confounding factor features, wherein the confounding factor features include features of noise factors, and the noise factors include features of noise factors that affect the causal relationship measurement between multimedia data and results; a feature determination module, used to remove the features of noise factors from the multimedia features based on the confounding factor features to obtain standard features of the multimedia data; and a result determination module, used to determine the target processing result of the multimedia data based on the standard features.

[0015] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0016] The fourth aspect of the present invention further provides a computer-readable storage medium on which a computer program or instruction is stored, and the steps of the above method are implemented when the above computer program or instruction is executed by a processor.

[0017] The fifth aspect of the present invention also provides a computer program product, including a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0018] According to an embodiment of the present invention, since there are confounding variables that influence each other and are difficult to distinguish in multimedia data such as video, voice and other streaming media data, these confounding variables will affect the causal relationship between the variables that need to be studied, thereby affecting the processing results when the multimedia data is processed. By using an encoder with adjusted parameters to process the multimedia features of the multimedia data, accurate confounding factor features can be determined. Since the confounding factor features are obtained based on the multimedia features, the confounding factor features are known quantities that can be determined. Using confounding factor features to perform causal intervention on multimedia features can eliminate the false correlation between the variables that need to be studied in the multimedia data due to the existence of confounding variables, thereby more effectively improving the accuracy of causal inference. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:

[0020] Figure 1 A diagram showing an application scenario of a method and device for processing multimedia data according to an embodiment of the present invention is shown.

[0021] Figure 2 A flowchart of a method for processing multimedia data according to an embodiment of the present invention is shown.

[0022] Figure 3A A schematic diagram of confounding variables according to an embodiment of the present invention is shown.

[0023] Figure 3B A schematic diagram of causal intervention according to an embodiment of the present invention is shown.

[0024] Figure 4 A structural diagram of a causal inference model of confounding factor features and adversarial learning according to an embodiment of the present invention is shown.

[0025] Figure 5 A structural block diagram of a multimedia data processing device according to an embodiment of the present invention is shown.

[0026] Figure 6 A block diagram of an electronic device suitable for implementing a method for processing multimedia data according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0027] Below, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of concepts of the present invention.

[0028] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the existence of features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0029] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0030] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0031] In the technical solution of the present invention, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, invention and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] In the scenario of using personal information for automated decision-making, the methods, devices, and systems provided by the embodiments of the present invention provide users with corresponding operation portals for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs, and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge and skills, and have reached a certain level of professionalism.

[0033] An embodiment of the present invention provides a method for processing multimedia data, comprising: extracting features from multimedia data to obtain multimedia features of the multimedia data, wherein the multimedia features include features of multiple different factors of the multimedia data, and the factors include noise factors; using an encoder to process the multimedia features to obtain confounding factor features, wherein the confounding factor features include features of noise factors, and the noise factors include factors that affect the measurement of the causal relationship between the multimedia data and the results; based on the confounding factor features, removing the features of the noise factors from the multimedia features to obtain standard features of the multimedia data; and determining the target processing result of the multimedia data based on the standard features.

[0034] Figure 1 A diagram showing an application scenario of a method and device for processing multimedia data according to an embodiment of the present invention is shown.

[0035] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0036] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0037] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0038] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0039] It should be noted that the multimedia data processing method provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the multimedia data processing device provided in the embodiment of the present invention can generally be set in the server 105. The multimedia data processing method provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Correspondingly, the multimedia data processing device provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0040] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0041] The following will be based on Figure 1 The scene described by Figure 2 , Figure 3A , Figure 3B and Figure 4 The method for processing multimedia data according to the embodiment of the present invention is described in detail.

[0042] Figure 2 A flowchart of a method for processing multimedia data according to an embodiment of the present invention is shown.

[0043] like Figure 2 As shown, the multimedia data processing method of this embodiment includes operations S210 to S240.

[0044] In operation S210, feature extraction is performed on multimedia data to obtain multimedia features of the multimedia data.

[0045] Multimedia data may include voice data, video data, etc. In addition to standard features that contribute to video semantics, multimedia data may also include confounding variables. For example, in the case of abnormal action detection of people, it is necessary to analyze the standard features of video data to determine whether the people in the video are performing abnormal actions. In this case, the standard features of video data are independent variables, the detection results of abnormal action detection are dependent variables, and factors such as video lighting conditions, resolution, and crowd density in the video can be used as confounding variables.

[0046] In the process of processing multimedia data, confounding variables may affect the standard characteristics of the multimedia data, thereby affecting the initial processing results. The processing of multimedia data may include detecting the correctness of the multimedia data, detecting the behavior of objects appearing in the multimedia data, etc.

[0047] For example, in dim light at night, the person in the video may move slowly, which is judged as abnormal movement. That is, the confounding variable can affect the dependent variable by affecting the independent variable. In dim light at night, it is difficult to judge the person's movements in the video, which affects the detection results. That is, the confounding variable can also directly affect the dependent variable.

[0048] There is a correlation between the confounding variable and the multimedia data, and there is also a correlation between the confounding variable and the output result determined by the multimedia data. The above correlation can be expressed as the confounding variable can affect both the multimedia data and the output result.

[0049] Since confounding variables are not variables that need to be directly concerned in the causal inference process, the existence of confounding variables will affect the mapping relationship between multimedia data and the causal inference results determined based on the multimedia data.

[0050] Since multimedia data has high data dimension and complexity, feature extraction methods can be used to extract multimedia features with dimensions smaller than the multimedia data to reduce the amount of data for subsequent processing, wherein the multimedia features include the characteristics of multiple different factors of the multimedia data, the multimedia features include standard features and confounding variables, and the factors include noise factors.

[0051] In operation S220, the multimedia features are processed using an encoder to obtain confounding factor features.

[0052] The multimedia features are encoded using an encoder, and the obtained encoding results can be used to represent the confounding factor features of the multimedia features, wherein the encoding methods may include Huffman coding, arithmetic coding, context-adaptive binary arithmetic coding, etc. The confounding factor features include features of noise factors that affect the causal relationship measurement between multimedia data and results, and the multiple factors include noise factors.

[0053] Similar to confounding variables, noise factors are associated with multimedia data and output results. Therefore, the characteristics of noise factors, namely, confounding factor characteristics, include characteristics of data that affect the causal relationship measurement between the semantics of multimedia data and the initial processing results of multimedia data. In addition, since confounding variables are difficult to predict or control directly, and confounding factor characteristics are obtained by encoding multimedia features using encoders, confounding factor characteristics can be used to replace confounding variables, so that confounding variables that are difficult to predict or control directly can be replaced with known confounding factor characteristics, thereby improving the accuracy of the statistical process.

[0054] In operation S230, based on the confounding factor feature, the feature of the noise factor is removed from the multimedia feature to obtain a standard feature of the multimedia data.

[0055] According to an embodiment of the present invention, a backdoor criterion can be used to remove confounding factor features from multimedia features. The backdoor criterion can identify and control confounding variables in causal inference by analyzing a directed acyclic graph of causal inference. In a directed acyclic graph, a node represents a variable, and a directed edge between nodes represents an influence relationship between two nodes. For example, if there is an edge from node A to node B in a directed acyclic graph, it means that the variable represented by node A has an influence on the variable represented by node B. For the causal relationship between the independent variable and the dependent variable, if there is a backdoor path from the independent variable to the dependent variable, the variable on the backdoor path can be determined as a confounding variable. In the directed acyclic graph, if there is a specific path, and there is at least one edge pointing to the independent variable in the specific path, it can be determined that the path will cause a non-causal association between the independent variable and the dependent variable, interfering with the judgment of the true causal relationship between the independent variable and the dependent variable.

[0056] For example, when studying the effect of a certain rehabilitation training method on the recovery status of athletes with fractures, in addition to "training method → ​​recovery status", the directed acyclic graph may also have directed edges such as "age → training method" and "age → recovery status". That is, athletes of different ages undergo rehabilitation training with different intensities, and athletes of different ages recover at different speeds and effects. Therefore, age will interfere with the judgment of the true causal relationship between rehabilitation training methods and recovery status. It can be determined that "age → training method" and "age → recovery status" are backdoor paths.

[0057] After determining the backdoor path, a set of control variables can be found. When the set of control variables is controlled, all backdoor paths from the independent variable to the dependent variable can be blocked, so that the dependent variable is only affected by the dependent variable, so as to judge the true causal relationship between the independent variable and the dependent variable. Among them, the control variable can be a node on the backdoor path or other variables. The control variable is not affected by the independent variable and can also block the influence of the variables on the backdoor path on the independent variable and the dependent variable.

[0058] For example, when studying the relationship between rehabilitation training methods and recovery status, age can be used as a control variable, and athletes of the same age can be selected for comparison. Gender, weight, competitive status and other variables can also be used as control variables. By controlling the values ​​of the above set of control variables, the influence of different ages on the causal relationship can be eliminated.

[0059] Therefore, in the embodiment of the present invention, the confounding factor feature can be used as a control variable. The confounding factor feature can be used to block the influence of the confounding variable on the multimedia feature, so that the multimedia feature is no longer affected by other variables.

[0060] Since the encoding result is obtained by encoding the confounding variables, the encoding result includes the characteristics of the confounding variables. By adjusting the encoding result, the influence of the confounding variables on the multimedia feature body can be blocked, so the encoding result can represent the confounding factor characteristics of the multimedia feature. In addition, the encoding result is a known result. By using the backdoor criterion, the standard feature can be obtained according to the encoding result and the multimedia feature, wherein the standard feature is a feature representation obtained after eliminating the influence of other variables on the multimedia feature. In the directed acyclic graph, there is only one edge from the standard feature to the causal inference result of the multimedia data, and there is no edge pointing to the standard feature.

[0061] In operation S240, a target processing result of the multimedia data is determined based on the standard feature.

[0062] Since the standard features are not affected by other variables, accurate target processing results can be obtained by using the standard features.

[0063] For example, when studying the relationship between rehabilitation training methods and recovery status, it is necessary to continuously adjust factors such as gender to block the influence of factors such as age on rehabilitation training methods. In this process, the recovery status will also change with changes in factors such as gender. Therefore, after completely blocking the influence of factors such as age on rehabilitation training methods, the relationship between the two can be accurately determined based on the current rehabilitation training methods and recovery status.

[0064] According to an embodiment of the present invention, since there are confounding variables that influence each other and are difficult to distinguish in multimedia data such as video, voice and other streaming media data, these confounding variables will affect the causal relationship between the variables that need to be studied, thereby affecting the processing results when the multimedia data is processed. By using an encoder with adjusted parameters to process the multimedia features of the multimedia data, accurate confounding factor features can be determined. Since the confounding factor features are obtained based on the multimedia features, the confounding factor features are known quantities that can be determined. Through backdoor adjustment, the confounding factor features are used to perform causal intervention on the multimedia features, which can eliminate the false correlation between the variables that need to be studied in the multimedia data due to the existence of confounding variables, thereby effectively improving the accuracy of causal inference.

[0065] According to an embodiment of the present invention, based on the confounding factor features, the features of the noise factors are removed from the multimedia features to obtain the standard features of the multimedia data, including: extracting features from the confounding factor features and the multimedia features to determine the joint features that characterize the correlation influence between the multimedia features and the confounding factor features; based on the joint features and the confounding factor features, determining the conditional features for removing the features of the noise factors; and determining the standard features using the conditional features and the multimedia features.

[0066] The mutual influence between multimedia features and confounding factor features can be represented by joint features. j It can be determined by calculating the result of element-wise multiplication between the multimedia feature X and the confounding factor feature C, as shown in formula (1):

[0067] (1)

[0068] in, Represents element-wise multiplication between matrices.

[0069] Among them, the confounding factor feature C is obtained by formula (2):

[0070] (2)

[0071] Among them, L encoder (·) indicates an encoder.

[0072] The influence of noise factors on multimedia features can be expressed by conditional features. C It can be determined by formula (3):

[0073] (3)

[0074] in, Can be used to express the use of joint features Pj The confounding factor feature C is weighted and then Perform skip connection with the confounding factor feature C. The above conditional feature P C The confounding factor feature can be used to indicate that the multimedia feature is affected by the confounding variable. Therefore, the conditional feature can be used to remove the noise factor feature from the multimedia feature, thereby obtaining the standard feature that is not affected by other variables. , can be determined by formula (4):

[0075] (4)

[0076] Among them, by training the model, the model can adaptively learn the conditional feature P C , therefore, the conditional feature P C The negative feature that can be used to represent the feature of the noise factor is calculated by summing it with the multimedia feature X through formula (4), so that the feature of the noise factor can be removed from the multimedia feature, thereby eliminating the influence of the noise factor on the multimedia feature.

[0077] Through the confounding factor feature, multiple confounding variables existing in the multimedia feature can be weighted averaged to obtain the true distribution of the multimedia feature after the intervention of the confounding factor feature, which is the standard feature.

[0078] According to an embodiment of the present invention, based on the backdoor criterion, the confounding factor features are used to perform causal intervention on the multimedia features, ensuring that the standard features obtained by the causal intervention can characterize the real distribution of the multimedia features and ensure the accuracy of causal inference.

[0079] Figure 3A A schematic diagram of confounding variables according to an embodiment of the present invention is shown.

[0080] like Figure 3A As shown, X, Y, and C are multimedia data, results determined based on the multimedia data, and confounding variables, respectively. C affects both the multimedia data and the results. Therefore, it is difficult to accurately determine the impact of X on Y based on the relationship between X, Y, and C.

[0081] Figure 3B A schematic diagram of causal intervention according to an embodiment of the present invention is shown.

[0082] like Figure 3B As shown, by using the above-mentioned operation of causal intervention using the confounding factor characteristics, the standard characteristics can be obtained , the influence of C on X is eliminated. Therefore, there is no predecessor node of X in the directed graph. By fixing C, the influence of X on Y can be accurately determined.

[0083] According to an embodiment of the present invention, feature extraction is performed on multimedia data to obtain multimedia features of the multimedia data, including: using a dimensionality reduction matrix to perform dimensionality reduction processing on the multimedia data to obtain the multimedia features.

[0084] Since multimedia data is usually multi-channel and multi-dimensional data, a dimensionality reduction matrix can be set to convert the multimedia data into a matrix form, and then matrix multiplication can be performed with the dimensionality reduction matrix to obtain multimedia features with a dimension lower than that of the multimedia data. The method of converting multimedia data into a matrix form may include converting multimedia data in a picture form into a pixel value matrix of the picture, etc.

[0085] Since the processing result is performed on multimedia data, after processing the multimedia features to obtain standard features, the inverse process opposite to the dimensionality reduction process can be used to restore the standard features to obtain multimedia data corresponding to the standard features, and then process them to obtain standard results for the multimedia data.

[0086] According to an embodiment of the present invention, a target processing result of multimedia data is determined based on standard features, including: using an inverse matrix of a dimensionality reduction matrix to process the standard features to obtain standard multimedia data corresponding to the standard features; and determining the target processing result based on the standard multimedia data.

[0087] According to an embodiment of the present invention, when the dimensionality reduction process is to right-multiply the dimensionality reduction matrix by the multimedia data, the inverse process of the dimensionality reduction process may be to determine the inverse matrix of the dimensionality reduction matrix, calculate the inverse matrix of the standard feature right-multiplied dimensionality reduction matrix, and determine the calculation result as the standard multimedia data.

[0088] After determining the standard multimedia data, the processing method for multimedia data can be used to determine the processing result of the standard multimedia data. Since the standard multimedia data is data obtained after causal intervention on the multimedia data, the processing result is the target processing result of the multimedia data.

[0089] According to an embodiment of the present invention, the encoder is trained in the following manner: feature extraction is performed on sample multimedia data to obtain sample multimedia features of the sample multimedia data; the sample multimedia features are processed using an untrained original encoder to obtain a sample multimedia encoding result; based on the discriminator's recognition result of the sample multimedia encoding result and the disordered sample multimedia encoding result, the parameters of the original encoder are adjusted to obtain a first encoder, wherein the disordered sample multimedia encoding result is obtained by disordering the sample multimedia encoding result; based on the sample multimedia decoding result and the sample multimedia features, the parameters of the first encoder are adjusted to obtain an encoder, wherein the sample multimedia decoding result is obtained by decoding the sample multimedia encoding result.

[0090] After the shuffle processing, the numerical values ​​in the shuffled sample multimedia encoding result are the same as the numerical values ​​in the sample multimedia encoding, but the arrangement order is different. During the training process, the encoder parameters can be adjusted according to the properties of the sample multimedia encoding results obtained by the encoder and the properties of the confounding factor characteristics of the sample multimedia data, so that the encoding results obtained by the encoder after parameter adjustment can be used as the confounding factor characteristics of the multimedia data after processing the multimedia data.

[0091] Confounding factor features need to be independent. After the confounding factor features are represented as a sequence, if multiple elements in the sequence can independently express their respective semantics, it can be determined that the confounding factor features are independent. In this case, the sequence of the confounding factor features is shuffled, and the semantics of the obtained shuffled sequence remains unchanged, so the shuffled sequence and the sequence before shuffling can express the same semantics. Therefore, the confounding factor features are shuffled, and if the distance between the result obtained after shuffling and the confounding factor features is small enough, it can be determined that the alternative confounding variables are independent. Based on the distance between the sample multimedia encoding result and the shuffled sample multimedia encoding result, the original encoder can be adjusted in parameters to obtain a first encoder that can output a sample multimedia encoding result with independence, wherein the distance between the shuffled sample multimedia encoding result and the sample multimedia encoding result feature can be represented by the Euclidean distance between the above two variables.

[0092] The confounding factor feature also needs to be reconfigurable. Reconfigurability means that after the multimedia feature is reconstructed using the confounding factor feature, the distance between the reconstructed variable and the multimedia feature is small enough, that is, the multimedia feature is encoded to obtain the confounding factor feature, and after the confounding factor feature is decoded, the reconstructed variable can express the same semantics as the variable before reconstruction. Therefore, based on the sample multimedia decoding results and the sample multimedia features, the parameters of the first encoder can be adjusted to obtain an encoder that can output a sample multimedia encoding result with reconfigurability.

[0093] According to an embodiment of the present invention, by checking the data output by the encoder, and when it is determined that the data output by the encoder does not have the properties required by the confounding factor feature, the encoder is parameter adjusted based on the results determined by other modules, so that the encoder obtained after parameter adjustment can output the properties required by the confounding factor feature. This ensures that in the application stage, the encoding result output by the encoder can be directly used as the confounding factor feature for subsequent causal intervention.

[0094] According to an embodiment of the present invention, a discriminator is used to identify sample multimedia coding results and out-of-order sample multimedia coding results to obtain a recognition result, including: using a discriminator to process the sample multimedia coding results and the out-of-order sample multimedia coding results to obtain a sample multimedia judgment result and an out-of-order sample multimedia judgment result, respectively; and determining the recognition result based on a cross entropy result determined according to the sample multimedia judgment result and the out-of-order sample multimedia judgment result.

[0095] Use the discriminator to encode the sample multimedia results C s and out-of-order sample multimedia encoding results perm Cs Processing is performed to obtain the sample multimedia encoding result C s The judgment result D (C s ) and the out-of-order sample multimedia encoding results perm Cs The judgment result D (perm Cs ), wherein the processing of the discriminator may include prediction and classification. For example, the discriminator is used to respectively perform the sample multimedia encoding results C s and out-of-order sample multimedia encoding results perm Cs By classifying, the category D (C s ) and the category D corresponding to the multimedia coding result of the out-of-order samples (perm Cs ).

[0096] Among them, perm Cs It can be obtained by formula (5):

[0097] (5)

[0098] Permutation (·) means randomly permuting matrices and vectors. Random permutation means that the values ​​of elements in matrices and vectors remain unchanged but their positional relationships are changed.

[0099] In order to ensure the training effect and generalization ability of the discriminator, the discriminator can be trained using streaming data including multiple sample multimedia data. Based on the cross entropy result BCE determined according to the multiple sample multimedia judgment results and the disordered sample multimedia judgment results, the recognition result of the discriminator can be determined, as shown in formula (6):

[0100] (6)

[0101] Wherein, n represents the sequence of the current sample multimedia data in the stream data.

[0102] According to an embodiment of the present invention, based on the recognition result of the discriminator on the sample multimedia encoding result and the disordered sample multimedia encoding result, the parameters of the original encoder are adjusted to obtain a first encoder, including: determining the recognition gradient of the original encoder based on the recognition result; processing the recognition gradient using a gradient reversal layer to obtain a reverse gradient; and adjusting the parameters of the original encoder based on the reverse gradient until the recognition result indicates that the distance between the sample multimedia encoding result and the disordered sample multimedia encoding result is less than a preset value.

[0103] Based on the above recognition results, the recognition gradient of the original encoder can be determined, wherein the recognition gradient is used to indicate the parameter adjustment direction of the original encoder. When the parameters are adjusted in this direction, the discriminator has a stronger recognition ability for the sample multimedia encoding results and the out-of-order sample multimedia encoding results. Since the parameters of the discriminator are not adjusted, it can be determined that by adjusting the parameters of the original encoder according to the recognition gradient, the distance between the output sample multimedia encoding results and the out-of-order sample multimedia encoding results obtained after the out-of-order processing will be greater, thereby ensuring that the discriminator has a stronger recognition ability.

[0104] Since the purpose of adjusting the parameters of the original encoder is to make the model unable to distinguish the sample multimedia encoding results from the disordered sample multimedia encoding results obtained after the disordered processing in the feature space, it is necessary to adjust the parameters in the opposite direction of the recognition gradient. The gradient reversal layer can be used to multiply the recognition gradient with a fixed value less than 0, and the product is determined as the reverse gradient. The parameter adjustment direction is determined based on the reverse gradient, and the parameters of the original encoder are adjusted according to the parameter adjustment direction and the preset parameter adjustment step size and learning rate, so that the distance between the sample multimedia encoding results output by the adjusted encoder and the disordered sample multimedia encoding results obtained after the disordered processing can be reduced.

[0105] The model parameters are continuously adjusted until the recognition result indicates that the distance between the sample multimedia encoding result and the shuffled sample multimedia encoding result is less than a preset value, and the first encoder can be determined based on the current model parameters, wherein the distance between the sample multimedia encoding result output by the first encoder before and after the shuffle processing is less than the preset value, that is, the shuffle processing does not affect the semantics of each element in the sample multimedia encoding result, and therefore the first encoder can output independent sample multimedia encoding results.

[0106] According to an embodiment of the present invention, by fixing the model parameters of the discriminator and training the encoder, the distance between the encoding result generated by the encoder and the result after the out-of-order processing can be reduced when the recognition ability of the discriminator is fixed, thereby enabling the encoder to have the ability to generate independent encoding results.

[0107] The discriminator is trained in the following manner: determining an out-of-order loss value between a sample multimedia encoding result and an out-of-order sample multimedia encoding result; and adjusting the model parameters of the discriminator based on the out-of-order loss until the out-of-order loss value converges.

[0108] The out-of-order loss value can be used to represent the loss caused by out-of-order processing of the sample multimedia coding result. Based on the sample multimedia coding result and the out-of-order sample multimedia coding, the out-of-order loss threshold between the two is determined, wherein when the sample multimedia coding results are independent, the out-of-order loss threshold between the two is determined to be a first threshold, and when the sample multimedia coding results are not independent, the out-of-order loss threshold between the two is determined to be a second threshold, wherein the first threshold and the second threshold can be preset, and the first threshold is less than the second threshold.

[0109] Based on the above determined disorder loss threshold and the disorder loss value between the sample multimedia encoding result determined by the discriminator and the disordered sample multimedia encoding result, the model parameters of the discriminator are adjusted until the difference between the disorder loss value output by the discriminator and the disorder loss threshold is less than the preset value, and the model parameters of the discriminator are determined. This ensures that the discriminator has good recognition ability and ensures the training effect of the encoder trained according to the results of the discriminator.

[0110] According to an embodiment of the present invention, based on the sample multimedia decoding results and the sample multimedia features, the parameters of the first encoder are adjusted to obtain the encoder, including: determining the loss value based on the sample multimedia decoding results and the sample multimedia features; and adjusting the parameters of the first encoder based on the loss value until the loss value converges.

[0111] According to an embodiment of the present invention, the loss value between the sample multimedia decoding result and the sample multimedia feature may include a mean square error, a root mean square error, etc. Taking the mean square error as an example, the loss value MSE between the sample multimedia decoding result and the sample multimedia feature may be determined by formula (7):

[0112] (7)

[0113] in, represents the multimedia features of the ith sample in the stream data of the multimedia data, Represents the multimedia decoding result of the i-th sample in the stream data of the multimedia data.

[0114] in, It can be obtained by formula (8):

[0115] (8)

[0116] Among them, Ci Represents the multimedia encoding result of the i-th sample in the stream data of the multimedia data.

[0117] The gradient descent method is used to adjust the parameters of the first encoder based on the loss value, determine the model parameters when the loss value converges, and determine the encoder based on the model parameters.

[0118] According to an embodiment of the present invention, nonlinear independent component analysis is used to determine the difference between the sample multimedia features before and after reconstruction based on the loss value between the encoded and decoded sample multimedia decoding results and the sample multimedia features, thereby determining the reconstructibility of the sample multimedia features, to ensure that after the model parameters are adjusted based on the loss value, the encoder has the ability to generate reconstructible encoding results.

[0119] Figure 4 A structural diagram of a causal inference model of confounding factor features and adversarial learning according to an embodiment of the present invention is shown.

[0120] like Figure 4 As shown, in the training phase, the encoder is used to process the sample multimedia features X' of the sample multimedia data to obtain the sample multimedia encoding result C'.

[0121] The model training part can be divided into two stages, namely the adversarial training stage and the autoencoder training stage.

[0122] Arrange the sample multimedia encoding results C' in random order to obtain the random sample multimedia encoding results perm C , use the discriminator to process the sample multimedia coding result C' and the disordered sample multimedia coding result perm C , and determine the recognition gradient based on the processing results. The gradient reversal layer can be used to achieve gradient reversal in order to perform subsequent training.

[0123] Input the sample multimedia encoding result C' into the decoder to obtain the reconstructed sample multimedia decoding result , and determine the loss value according to the sample multimedia features X' and the multimedia decoding results, to ensure that the sample multimedia features X' processed by the encoder can be restored to obtain the sample multimedia features after decoding, and complete the parameter adjustment of the autoencoder training stage.

[0124] After training, the encoder has the ability to generate independent and reconstructible coding results. Therefore, after the encoder is used to process the multimedia feature X' to obtain C', C' can be determined as the confounding factor feature of the sample multimedia feature X'. Therefore, based on the confounding factor feature C', causal intervention is performed on the sample multimedia feature X' to obtain the standard features of the multimedia data.

[0125] In one embodiment, when performing a video anomaly detection task, the multimedia data is video stream data, and the feature extraction model is used to process the video frames to obtain a feature sequence of the video, wherein the feature sequence includes a plurality of multimedia features arranged in the order of the plurality of video frames in the video stream data. The feature sequence is processed by an encoder to obtain a confounding factor feature, and the confounding factor feature can be used to perform causal intervention on the feature sequence to eliminate the influence of the confounding variables on the feature sequence and the detection result of the video.

[0126] Based on the above-mentioned multimedia data processing method, the present invention also provides a multimedia data processing device. Figure 5 The device is described in detail.

[0127] Figure 5 A structural block diagram of a multimedia data processing device according to an embodiment of the present invention is shown.

[0128] like Figure 5 As shown, the multimedia data processing device 500 of this embodiment includes a feature extraction module 510 , a feature processing module 520 , a feature determination module 530 and a result determination module 540 .

[0129] The feature extraction module 510 is used to extract features from multimedia data to obtain multimedia features of the multimedia data, wherein the multimedia features include features of multiple different factors of the multimedia data, and the factors include noise factors. In one embodiment, the feature extraction module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0130] The feature processing module 520 is used to process the multimedia features using an encoder to obtain a confounding factor feature, wherein the confounding factor feature includes a feature of a noise factor, and the noise factor includes a factor that affects the causal relationship measurement between the multimedia data and the result. In one embodiment, the feature processing module 520 can be used to perform the operation S220 described above, which will not be repeated here.

[0131] The feature determination module 530 is used to remove the features of noise factors from the multimedia features based on the confounding factor features to obtain standard features of the multimedia data. In one embodiment, the feature determination module 530 can be used to perform the operation S230 described above, which will not be described in detail here.

[0132] The result determination module 540 is used to determine the target processing result of the multimedia data based on the standard features. In one embodiment, the result determination module 540 can be used to perform the operation S240 described above, which will not be described in detail here.

[0133] According to an embodiment of the present invention, the feature determination module 530 includes a first feature determination submodule, a second feature determination submodule and a feature determination submodule.

[0134] The first feature determination submodule is used to extract features from the confounding factor features and the multimedia features, and determine a joint feature that represents the correlation effect between the multimedia features and the confounding factor features.

[0135] The second feature determination submodule is used to determine conditional features for removing features of noise factors based on the joint features and the confounding factor features.

[0136] The feature determination submodule is used to determine standard features using conditional features and multimedia features.

[0137] According to an embodiment of the present invention, the feature extraction module 510 includes a data dimension reduction submodule.

[0138] The data dimension reduction submodule is used to use the dimension reduction matrix to reduce the dimension of multimedia data and obtain multimedia features.

[0139] According to an embodiment of the present invention, the result determination module 540 includes a feature restoration submodule and a result determination submodule.

[0140] The feature restoration submodule is used to process the standard features using the inverse matrix of the dimension reduction matrix to obtain standard multimedia data corresponding to the standard features.

[0141] The result determination submodule is used to determine the target processing result based on the standard multimedia data.

[0142] According to an embodiment of the present invention, the multimedia data processing device 500 further includes a sample feature extraction module, a sample feature processing module, a first parameter adjustment module and a second parameter adjustment module.

[0143] The sample feature extraction module is used to extract features from the sample multimedia data to obtain sample multimedia features of the sample multimedia data, wherein the sample multimedia features include features of multiple different sample factors of the sample multimedia data.

[0144] The sample feature processing module is used to process the sample multimedia features using the original encoder to obtain the sample multimedia encoding result.

[0145] The first parameter adjustment module is used to adjust the parameters of the original encoder based on the recognition result of the discriminator on the sample multimedia encoding result and the disordered sample multimedia encoding result to obtain the first encoder, wherein the disordered sample multimedia encoding result is obtained by disordering the sample multimedia encoding result.

[0146] The second parameter adjustment module is used to adjust the parameters of the first encoder based on the sample multimedia decoding result and the sample multimedia feature to obtain the encoder, wherein the sample multimedia decoding result is obtained by decoding the sample multimedia encoding result.

[0147] According to an embodiment of the present invention, the multimedia data processing device 500 further includes a coding disorder module, a result identification module and a result processing module.

[0148] The coding disorder module is used to perform disorder processing on the sample multimedia coding to obtain disordered sample multimedia coding results.

[0149] The result recognition module is used to use a discriminator to recognize the sample multimedia coding result and the disordered sample multimedia coding result to obtain a recognition result, wherein the recognition result is used to represent the distance between the sample multimedia coding result and the disordered sample multimedia coding result.

[0150] The result processing module is used to process the encoding result by using a decoder to obtain a sample multimedia decoding result.

[0151] According to an embodiment of the present invention, the result identification module includes a first result identification submodule and a second result identification submodule.

[0152] The first result identification submodule is used to process the sample multimedia encoding result and the disordered sample multimedia encoding result by using the discriminator to obtain the sample multimedia judgment result and the disordered sample multimedia judgment result respectively.

[0153] The second result identification submodule is used to determine the identification result based on the cross entropy result determined according to the sample multimedia judgment result and the disordered sample multimedia judgment result.

[0154] According to an embodiment of the present invention, the first parameter adjustment module includes a gradient determination submodule, a gradient reversal submodule and a first parameter adjustment submodule.

[0155] The gradient determination submodule is used to determine the gradient direction and gradient value of the original encoder based on the recognition result.

[0156] The gradient reversal submodule is used to reverse the gradient direction to obtain a reversed gradient direction.

[0157] The first parameter adjustment submodule is used to adjust the parameters of the original encoder based on the inverted gradient direction and gradient value until the recognition result indicates that the distance between the sample multimedia encoding result and the disordered sample multimedia encoding result is less than a preset value.

[0158] According to an embodiment of the present invention, the second parameter adjustment module includes a loss value determination submodule and a second parameter adjustment submodule.

[0159] The loss value determination submodule is used to determine the loss value based on the sample multimedia decoding result and the sample multimedia features.

[0160] The second parameter adjustment submodule is used to adjust the parameters of the first encoder based on the loss value until the loss value converges.

[0161] According to an embodiment of the present invention, any multiple modules of the feature extraction module 510, the feature processing module 520, the feature determination module 530 and the result determination module 540 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the feature extraction module 510, the feature processing module 520, the feature determination module 530 and the result determination module 540 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in a suitable combination of any of them. Alternatively, at least one of the feature extraction module 510, the feature processing module 520, the feature determination module 530 and the result determination module 540 may be at least partially implemented as a computer program module, which may perform a corresponding function when executed.

[0162] Figure 6 A block diagram of an electronic device suitable for implementing a method for processing multimedia data according to an embodiment of the present invention is shown.

[0163] like Figure 6 As shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 to a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include an onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0164] In RAM 603, various programs and data required for the operation of electronic device 600 are stored. Processor 601, ROM 602 and RAM 603 are connected to each other via bus 604. Processor 601 performs various operations of the method flow according to the embodiment of the present invention by executing the programs in ROM 602 and / or RAM 603. It should be noted that the program can also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 can also perform various operations of the method flow according to the embodiment of the present invention by executing the programs stored in one or more memories.

[0165] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage portion 608 as needed.

[0166] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.

[0167] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than ROM 602 and RAM 603.

[0168] The embodiment of the present invention also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present invention.

[0169] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when it is executed by the processor 601. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0170] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0171] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.

[0172] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0173] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0174] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention may be combined and / or combined in various ways. All of these combinations and / or combinations fall within the scope of the present invention.

[0175] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination advantageously. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A method for processing multimedia data, characterized in that: The method comprises: Extracting features from the multimedia data to obtain multimedia features of the multimedia data, wherein the multimedia features include features of multiple different factors of the multimedia data, and the factors include noise factors; Using an encoder, processing the multimedia features to obtain confounding factor features, wherein the confounding factor features include features of the noise factors, and the noise factors include factors that affect the causal relationship measurement between the multimedia data and the results; Based on the confounding factor feature, removing the feature of the noise factor from the multimedia feature to obtain a standard feature of the multimedia data; and Based on the standard features, a target processing result of the multimedia data is determined.

2. The method according to claim 1, characterized in that The removing the feature of the noise factor from the multimedia feature based on the confounding factor feature to obtain the standard feature of the multimedia data includes: Performing feature extraction on the confounding factor feature and the multimedia feature to determine a joint feature representing a correlation influence between the multimedia feature and the confounding factor feature; Determining conditional features for removing features of the noise factor based on the joint features and the confounding factor features; and The standard feature is determined using the condition feature and the multimedia feature.

3. The method according to claim 1, characterized in that The extracting features of the multimedia data to obtain multimedia features of the multimedia data includes: The multimedia data is subjected to dimensionality reduction processing by using a dimensionality reduction matrix to obtain the multimedia features.

4. The method according to claim 3, characterized in that The determining, based on the standard feature, a target processing result of the multimedia data includes: Processing the standard features using the inverse matrix of the dimension reduction matrix to obtain standard multimedia data corresponding to the standard features; and Based on the standard multimedia data, the target processing result is determined.

5. The method according to claim 1, characterized in that The encoder is trained in the following way: Extracting features from the sample multimedia data to obtain sample multimedia features of the sample multimedia data, wherein the sample multimedia features include features of a plurality of different sample factors of the sample multimedia data; Using an original encoder, processing the sample multimedia features to obtain a sample multimedia encoding result; Adjusting the parameters of the original encoder based on the recognition result of the sample multimedia encoding result and the disordered sample multimedia encoding result by the discriminator to obtain a first encoder, wherein the disordered sample multimedia encoding result is obtained by disordering the sample multimedia encoding result; Based on the sample multimedia decoding result and the sample multimedia feature, the parameters of the first encoder are adjusted to obtain the encoder, wherein the sample multimedia decoding result is obtained by decoding the sample multimedia encoding result.

6. The method according to claim 5, characterized in that The method further comprises: Performing random processing on the sample multimedia coding to obtain the random sample multimedia coding result; Using the discriminator to identify the sample multimedia coding result and the disordered sample multimedia coding result, to obtain the identification result, wherein the identification result is used to represent the distance between the sample multimedia coding result and the disordered sample multimedia coding result; and The encoding result is processed by a decoder to obtain the sample multimedia decoding result.

7. The method according to claim 6, characterized in that The step of using the discriminator to identify the sample multimedia encoding result and the disordered sample multimedia encoding result to obtain the identification result includes: Processing the sample multimedia coding result and the disordered sample multimedia coding result by the discriminator to obtain a sample multimedia judgment result and a disordered sample multimedia judgment result respectively; and The recognition result is determined based on a cross entropy result determined according to the sample multimedia judgment result and the disordered sample multimedia judgment result.

8. The method according to claim 5, characterized in that The method of adjusting the parameters of the original encoder based on the recognition result of the sample multimedia encoding result and the disordered sample multimedia encoding result by the discriminator to obtain the first encoder includes: Based on the recognition result, determining the gradient direction and gradient value of the original encoder; Reversing the gradient direction to obtain a reversed gradient direction; and Based on the reversed gradient direction and the gradient value, the parameters of the original encoder are adjusted until the recognition result indicates that the distance between the sample multimedia encoding result and the out-of-order sample multimedia encoding result is less than a preset value.

9. The method according to claim 6, characterized in that The adjusting the parameters of the first encoder based on the sample multimedia decoding result and the sample multimedia feature to obtain the encoder includes: Determining a loss value based on the sample multimedia decoding result and the sample multimedia feature; and Parameters of the first encoder are adjusted based on the loss value until the loss value converges.

10. A multimedia data processing device, characterized in that: The device comprises: A feature extraction module, configured to extract features from the multimedia data to obtain multimedia features of the multimedia data, wherein the multimedia features include features of respective factors of the multimedia data, and the factors include noise factors; A feature processing module, configured to process the multimedia features using an encoder to obtain a confounding factor feature, wherein the confounding factor feature includes a feature of the noise factor, and the noise factor includes a factor that affects the causal relationship measurement between the multimedia data and the result; a feature determination module, configured to remove the feature of the noise factor from the multimedia feature based on the confounding factor feature, so as to obtain a standard feature of the multimedia data; and The result determination module is used to determine the target processing result of the multimedia data based on the standard features.

Citation Information

Patent Citations

  • Multimedia file recommendation method, device, parameter adjustment method, device, medium and electronic equipment

    CN110990600A

  • Competitive risk survival analysis method based on causal inference

    CN114418420A

  • System and method for generating unbiased scene graph based on causal reasoning

    CN118015117A

  • Road crack detection method based on causal decoupling, electronic equipment and storage medium

    CN119360085A

  • Back door defense method based on causal diffusion model

    CN119416212A