Method and apparatus for processing multimedia data
Through feature extraction and encoder processing, the noise factors in multimedia data are removed using backdoor criterion, and the problem of the influence of confounding variables in causal inference is solved, and the accuracy of causal relationship and the reliability of processing results are achieved.
Patent Information
- Application Number
- CN202510421898.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In multimedia data, confounding variables are difficult to distinguish, affecting the accuracy of causal relationships, and existing methods are difficult to effectively remove noise factors, resulting in inaccurate causal inference results.
Through feature extraction and encoder processing, the confounding factor characteristics are determined, the noise factors are removed using backdoor criterion, standard features are obtained, and data processing is performed using dimensionality reduction matrix to ensure the accuracy of causal inference.
The accuracy of causal inference is improved, the influence of confounding variables on causal relationship is eliminated, and the accuracy and reliability of the processing results are ensured.
Smart Images

Figure CN119940520B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and apparatus for processing multimedia data. Background Art
[0002] In the field of data analysis, causal inference of multimedia data can determine the direct impact of one variable on another from a large number of complex variable relationships. Therefore, based on the results of causal inference, decision-making can be guided, such as providing an optimized resource allocation plan and revealing potential problems. Therefore, causal inference is an important research direction in the field of data analysis. With the development of modern data analysis and computer technology, the objects of data analysis are no longer limited to numbers and texts, and artificial intelligence models can be used to analyze multimedia data such as videos and voices that are difficult to process and analyze by traditional computers.
[0003] A commonly used method for causal inference is to describe the causal relationship between variables through a causal graph. A causal graph is a directed acyclic graph, where nodes represent variables and edges represent causal relationships. In real data, the causes leading to a result are usually complex, and there may be confounding variables that affect both the cause and the result. In the case of the existence of confounding variables in multimedia data, using traditional correlation analysis methods such as causal graphs can only reveal the statistical relationship between variables, and it is difficult to accurately determine the causal relationship. Summary of the Invention
[0004] In view of the above problems, the present invention provides a method and apparatus for processing multimedia data.
[0005] According to a first aspect of the present invention, there is provided a method for processing multimedia data, including: extracting features from the multimedia data to obtain multimedia features of the multimedia data, where the multimedia features include features of multiple different factors of the multimedia data, and the factors include noise factors; using an encoder to process the multimedia features to obtain confounding factor features, where the confounding factor features include features of the noise factors, and the noise factors include factors that affect the causal relationship measurement between the multimedia data and the result; based on the confounding factor features, removing the features of the noise factors from the multimedia features to obtain standard features of the multimedia data; and based on the standard features, determining a target processing result of the multimedia data.
[0006] According to an embodiment of the present invention, based on the confounding factor features, removing the features of the noise factors from the multimedia features to obtain the standard features of the multimedia data, including: extracting features from the confounding factor features and the multimedia features to determine the joint features representing the associated influence between the multimedia features and the confounding factor features; determining the conditional features for removing the features of the noise factors based on the joint features and the confounding factor features; and determining the standard features by using the conditional features and the multimedia features.
[0007] According to an embodiment of the present invention, extracting features from the multimedia data to obtain the multimedia features of the multimedia data, including: using a dimensionality reduction matrix to perform dimensionality reduction processing on the multimedia data to obtain the multimedia features.
[0008] According to an embodiment of the present invention, based on the standard features, determining the target processing result of the multimedia data, including: using the inverse matrix of the dimensionality reduction matrix to process the standard features to obtain the standard multimedia data corresponding to the standard features; and determining the target processing result based on the standard multimedia data.
[0009] According to an embodiment of the present invention, the encoder is trained in the following manner: extracting features from the sample multimedia data to obtain the sample multimedia features of the sample multimedia data, where the sample multimedia features include the features of multiple different sample factors of the sample multimedia data; using the original encoder to process the sample multimedia features to obtain the sample multimedia coding result; adjusting the parameters of the original encoder based on the recognition result of the discriminator for the sample multimedia coding result and the scrambled sample multimedia coding result, where the scrambled sample multimedia coding result is obtained by scrambling the sample multimedia coding result, to obtain the first encoder; and adjusting the parameters of the first encoder based on the sample multimedia decoding result and the sample multimedia features to obtain the encoder, where the sample multimedia decoding result is obtained by decoding the sample multimedia coding result.
[0010] According to an embodiment of the present invention, the processing method of the multimedia data further includes: scrambling the sample multimedia coding to obtain the scrambled sample multimedia coding result; using the discriminator to recognize the sample multimedia coding result and the scrambled sample multimedia coding result to obtain the recognition result, where the recognition result is used to represent the distance between the sample multimedia coding result and the scrambled sample multimedia coding result; and using the decoder to process the coding result to obtain the sample multimedia decoding result.
[0011] According to an embodiment of the present invention, a discriminator is used to identify the sample multimedia encoding result and the scrambled sample multimedia encoding result, and an identification result is obtained, including: processing the sample multimedia encoding result and the scrambled sample multimedia encoding result by the discriminator to obtain a sample multimedia judgment result and a scrambled sample multimedia judgment result respectively; and determining the identification result based on the cross-entropy result determined according to the sample multimedia judgment result and the scrambled sample multimedia judgment result.
[0012] According to an embodiment of the present invention, based on the identification result of the discriminator for the sample multimedia encoding result and the scrambled sample multimedia encoding result, the parameters of the original encoder are adjusted to obtain a first encoder, including: determining the gradient direction and gradient value of the original encoder based on the identification result; performing an inversion process on the gradient direction to obtain an inverted gradient direction; and adjusting the parameters of the original encoder based on the inverted gradient direction and the gradient value until the identification result indicates that the distance between the sample multimedia encoding result and the scrambled sample multimedia encoding result is less than a preset value.
[0013] According to an embodiment of the present invention, based on the sample multimedia decoding result and the sample multimedia features, the parameters of the first encoder are adjusted to obtain an encoder, including: determining a loss value based on the sample multimedia decoding result and the sample multimedia features; and adjusting the parameters of the first encoder based on the loss value until the loss value converges.
[0014] A second aspect of the present invention provides a multimedia data processing device, including: a feature extraction module for extracting features from multimedia data to obtain multimedia features of the multimedia data, where the multimedia features include the features of multiple different factors of the multimedia data, and the factors include noise factors; a feature processing module for using an encoder to process the multimedia features to obtain a confounding factor feature, where the confounding factor feature includes the features of noise factors, and the noise factors include the features of noise factors that affect the causal relationship measurement between the multimedia data and the result; a feature determination module for removing the features of noise factors from the multimedia features based on the confounding factor feature to obtain standard features of the multimedia data; and a result determination module for determining a target processing result of the multimedia data based on the standard features.
[0015] A third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, where the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0016] A fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0017] The fifth aspect of the present invention further provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps of the above method.
[0018] According to an embodiment of the present invention, in multimedia data such as video, voice and other streaming media data, there are confounding variables that affect each other and are difficult to distinguish. These confounding variables will affect the causal relationship between the variables to be studied, thus affecting the processing result when processing multimedia data. By using an encoder with adjusted parameters to process the multimedia features of multimedia data, accurate confounding factor features can be determined. Since the confounding factor features are obtained based on the multimedia features, the confounding factor features are known quantities that can be determined. Using the confounding factor features to perform causal intervention on the multimedia features can eliminate the spurious correlation between the variables to be studied due to the existence of confounding variables in the multimedia data, thereby effectively improving the accuracy of causal inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features and advantages of the present invention will become clearer. In the drawings:
[0020] Figure 1 The application scenario diagram of the method and device for processing multimedia data according to an embodiment of the present invention is shown.
[0021] Figure 2 The flowchart of the method for processing multimedia data according to an embodiment of the present invention is shown.
[0022] Figure 3A The schematic diagram of the confounding variables according to an embodiment of the present invention is shown.
[0023] Figure 3B The schematic diagram of the causal intervention according to an embodiment of the present invention is shown.
[0024] Figure 4 The structural diagram of the confounding factor features and the causal inference model of adversarial learning according to an embodiment of the present invention is shown.
[0025] Figure 5 The structural block diagram of the device for processing multimedia data according to an embodiment of the present invention is shown.
[0026] Figure 6 The block diagram of the electronic device suitable for implementing the method for processing multimedia data according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0028] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0030] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0031] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, invention, and application, etc., all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or reject.
[0032] In the scenario of making automated decisions using personal information, the method, device, and system provided by the embodiments of the present invention all provide corresponding operation entrances for users to choose to agree or reject the results of automated decisions; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies, or economic, health, credit status, etc. through a computer program and making decisions. The expression "expert decision" here refers to the activity of making decisions by personnel who are engaged in work in a specific field, have specialized experience, knowledge, and skills, and have reached a certain professional level.
[0033] Embodiments of the present invention provide a method for processing multimedia data, including: extracting features from the multimedia data to obtain multimedia features of the multimedia data, where the multimedia features include the features of multiple different factors of the multimedia data, and the factors include noise factors; using an encoder to process the multimedia features to obtain confounding factor features, where the confounding factor features include the features of the noise factors, and the noise factors include factors that affect the measurement of the causal relationship between the multimedia data and the result; based on the confounding factor features, removing the features of the noise factors from the multimedia features to obtain standard features of the multimedia data; and based on the standard features, determining the target processing result of the multimedia data.
[0034] Figure 1 The application scenario diagram of the method and device for processing multimedia data according to an embodiment of the present invention is shown.
[0035] As Figure 1 shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0036] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).
[0037] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and the like.
[0038] The server 105 may be a server that provides various services, such as a background management server (only an example) that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0039] It should be noted that the multimedia data processing method provided in the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the multimedia data processing device provided in the embodiments of the present invention can generally be set in the server 105. The multimedia data processing method provided in the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the multimedia data processing device provided in the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0040] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in
[0041] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the scenario described below Figure 2 , Figure 3A , Figure 3B and Figure 4 the multimedia data processing method of the embodiments of the present invention will be described in detail.
[0042] Figure 2 shows a flowchart of the multimedia data processing method according to the embodiments of the present invention.
[0043] As Figure 2 shown, the multimedia data processing method of this embodiment includes operation S210 to operation S240.
[0044] In operation S210, feature extraction is performed on the multimedia data to obtain the multimedia features of the multimedia data.
[0045] Multimedia data may include voice data, video data, etc. In multimedia data, in addition to including standard features that contribute to video semantics, it may also include confounding variables. For example, in the case of detecting abnormal actions of a person, it is necessary to analyze the standard features of the video data to determine whether the person in the video is performing an abnormal action. In this case, the standard features of the video data are independent variables, and the detection result of the abnormal action detection is the dependent variable, while factors such as the lighting condition of the video, the resolution, and the crowd density in the video can be used as confounding variables.
[0046] During the process of processing multimedia data, confounding variables will affect the standard features of the multimedia data, thereby affecting the initial processing result. Among them, the processing performed on the multimedia data may include detecting the correctness of the multimedia data, detecting the behavior of the objects appearing in the multimedia data, etc.
[0047] For example, in the case of dim light at night, the people in the video may move slowly, thus being judged as having abnormal actions, that is, this confounding variable can affect the dependent variable by affecting the independent variable. In the case of dim light at night, it is difficult to judge the actions of the people in the video, thus affecting the detection result, that is, this confounding variable can also directly affect the dependent variable.
[0048] There is an association between the confounding variable and the multimedia data, and there is also an association between the confounding variable and the output result determined by the multimedia data. The above association can be expressed as that the confounding variable can affect both the multimedia data and the output result.
[0049] Since the confounding variable is not a variable that needs to be directly concerned about in the causal inference process, the existence of the confounding variable will affect the mapping relationship between the multimedia data and the causal inference result determined according to the multimedia data.
[0050] Since the data dimension and complexity of the multimedia data are relatively high, feature extraction methods can be used to extract multimedia features with dimensions smaller than those of the multimedia data to reduce the amount of data for subsequent processing. Among them, the multimedia features include the features of multiple different factors of the multimedia data. The multimedia features include standard features and confounding variables, and the factors include noise factors.
[0051] In operation S220, an encoder is used to process the multimedia features to obtain confounding factor features.
[0052] The encoder is used to encode the multimedia features, and the obtained encoding result can be used to represent the confounding factor features of the multimedia features. Among them, the encoding methods can include Huffman coding, arithmetic coding, context - adaptive binary arithmetic coding, etc. The confounding factor features include the features of the noise factors that affect the causal relationship measurement between the multimedia data and the result, and the multiple factors include the noise factors.
[0053] Similar to the confounding variables, there are associations between the noise factors and both the multimedia data and the output result. Therefore, the features of the noise factors, that is, the confounding factor features, include the features of the data that affect the causal relationship measurement between the semantics of the multimedia data and the initial processing result of the multimedia data. In addition, since the confounding variables are difficult to directly predict or control, and the confounding factor features are obtained by encoding the multimedia features using an encoder, the confounding variables can be replaced with the confounding factor features, so that the confounding variables that are difficult to directly predict or control can be replaced with the known confounding factor features, improving the accuracy of the statistical process.
[0054] In operation S230, based on the confounding factor features, the features of the noise factors are removed from the multimedia features to obtain the standard features of the multimedia data.
[0055] According to the embodiments of the present invention, the backdoor criterion can be used to remove the confounding factor features from the multimedia features. The backdoor criterion can analyze the directed acyclic graph of causal inference to identify and control the confounding variables in causal inference. In a directed acyclic graph, nodes represent variables, and the directed edges between nodes represent the influence relationship between two nodes. For example, if there is an edge from node A to node B in the directed acyclic graph, it means that the variable represented by node A has an influence on the variable represented by node B. For the causal relationship between the independent variable and the dependent variable, if there is a backdoor path from the independent variable to the dependent variable, the variables on the backdoor path can be determined as confounding variables. Among them, in a directed acyclic graph, if there is a specific path and there is at least one edge pointing to the independent variable in the specific path, it can be determined that this path will cause a non - causal association between the independent variable and the dependent variable, interfering with the judgment of the true causal relationship between the independent variable and the dependent variable.
[0056] For example, in the case of studying the influence of a certain rehabilitation training method on the recovery status of fractured athletes, in addition to "training method → recovery status" in the directed acyclic graph, there may also be directed edges such as "age → training method" and "age → recovery status", that is, athletes of different ages have different intensities of this rehabilitation training, and athletes of different ages also have different speeds and effects of self - recovery. Therefore, age will interfere with the judgment of the true causal relationship between the rehabilitation training method and the recovery status, and "age → training method" and "age → recovery status" can be determined as the backdoor paths.
[0057] After determining the backdoor path, a set of control variables can be found. When these control variables are controlled, all backdoor paths from the independent variable to the dependent variable can be blocked, so that the dependent variable is only affected by the independent variable, in order to judge the true causal relationship between the independent variable and the dependent variable. Among them, the control variables can be the nodes on the backdoor path or other variables. The control variables are not affected by the independent variable and can also block the influence of the variables on the backdoor path on the independent variable and the dependent variable.
[0058] For example, in the case of studying the relationship between the rehabilitation training method and the recovery status, age can be used as a control variable, and athletes of the same age can be selected for comparison. A set of variables such as gender, weight, and competitive status can also be used as control variables. By controlling the values of the above set of control variables, the influence of different ages on the causal relationship can be eliminated.
[0059] Therefore, in the embodiments of the present invention, the confounding factor feature can be used as a control variable, and the influence of the confounding variable on the multimedia feature can be blocked by using the confounding factor feature, so that the multimedia feature is no longer affected by other variables.
[0060] Since the encoding result is obtained by encoding the confounding variable, the encoding result includes the features of the confounding variable. By adjusting the encoding result, the influence of the confounding variable on the multimedia feature can be blocked. Therefore, the encoding result can represent the confounding factor feature of the multimedia feature. And the encoding result is a known result. Using the backdoor criterion, the standard feature can be obtained according to the encoding result and the multimedia feature. Among them, the standard feature is the feature representation obtained after eliminating the influence of other variables on the multimedia feature. In the directed acyclic graph, there is only one edge pointing from the standard feature to the causal inference result of the multimedia data, and there is no edge pointing to the standard feature.
[0061] In operation S240, based on the standard feature, the target processing result of the multimedia data is determined.
[0062] Since the standard feature is not affected by other variables, the accurate target processing result can be obtained by using the standard feature.
[0063] For example, in the case of studying the relationship between the rehabilitation training method and the recovery status, factors such as gender need to be continuously adjusted to block the influence of factors such as age on the rehabilitation training method. During this process, the recovery status will also change with the change of factors such as gender. Therefore, after completely blocking the influence of factors such as age on the rehabilitation training method, the relationship between the current rehabilitation training method and the recovery status can be accurately determined.
[0064] According to an embodiment of the present invention, in multimedia data such as video, voice and other streaming media data, there are confounding variables that interact with each other and are difficult to distinguish. These confounding variables will affect the causal relationship between the variables to be studied, thus affecting the processing result when processing multimedia data. By using an encoder with adjusted parameters to process the multimedia features of multimedia data, accurate confounding factor features can be determined. Since the confounding factor features are obtained by processing based on the multimedia features, the confounding factor features are known quantities that can be determined. Through backdoor adjustment, using the confounding factor features to perform causal intervention on the multimedia features can eliminate the spurious correlation between the variables to be studied due to the existence of confounding variables in the multimedia data, thereby effectively improving the accuracy of causal inference.
[0065] According to an embodiment of the present invention, based on the confounding factor features, the features of noise factors are removed from the multimedia features to obtain the standard features of the multimedia data, including: extracting features from the confounding factor features and the multimedia features to determine the joint features representing the associated influence between the multimedia features and the confounding factor features; determining the conditional features for removing the features of noise factors based on the joint features and the confounding factor features; and determining the standard features by using the conditional features and the multimedia features.
[0066] The mutual influence between the multimedia features and the confounding factor features can be represented by the joint features. The joint feature P j can be determined by the calculation result of the element-wise multiplication between the multimedia feature X and the confounding factor feature C, as shown in formula (1):
[0067] (1)
[0068] where, represents the element-wise multiplication between matrices.
[0069] Among them, the confounding factor feature C is obtained by formula (2):
[0070] (2)
[0071] where, L encoder (·) represents the encoder.
[0072] The influence of the noise factors on the multimedia features can be represented by the conditional features. The conditional feature P C can be determined by formula (3):
[0073] (3)
[0074] where, can be used to represent using the joint feature Pj Weight the confounding factor feature C, and then perform a skip connection with the confounding factor feature C. The above conditional feature P C can utilize the confounding factor feature, indicating that the multimedia feature is affected by the confounding variable. Therefore, the feature of the noise factor can be removed from the multimedia feature by using the conditional feature, so as to obtain the standard feature that is not affected by other variables , which can be determined by formula (4):
[0075] (4)
[0076] Among them, by training the model, the model can adaptively learn the conditional feature P C . Therefore, the conditional feature P C can be used to represent the negative feature of the feature of the noise factor. By calculating the sum of it and the multimedia feature X through formula (4), the feature of the noise factor can be removed from the multimedia feature, so as to eliminate the influence of the noise factor on the multimedia feature.
[0077] Through the confounding factor feature, the multiple confounding variables existing in the multimedia feature can be weighted and averaged, so as to obtain the true distribution of the multimedia feature after being intervened by the confounding factor feature, which is the standard feature.
[0078] According to the embodiments of the present invention, based on the backdoor criterion, the multimedia feature is causally intervened by using the confounding factor feature to ensure that the standard feature obtained by the causal intervention can represent the true distribution of the multimedia feature and ensure the accuracy of causal inference.
[0079] Figure 3A shows a schematic diagram of the confounding variable according to the embodiments of the present invention.
[0080] As Figure 3A shown, X, Y, and C are multimedia data, the result determined according to the multimedia data, and the confounding variable respectively. C affects both the multimedia data and the result at the same time. Therefore, it is difficult to accurately determine the influence of X on Y according to the relationship between X, Y, and C.
[0081] Figure 3B shows a schematic diagram of the causal intervention according to the embodiments of the present invention.
[0082] As Figure 3B shown, through the above operation of causally intervening by using the confounding factor feature, the standard feature can be obtained, and the influence of C on X is eliminated. Therefore, there are no precursor nodes of X in the directed graph. By fixing C, the influence situation of X on Y can be accurately determined.
[0083] According to an embodiment of the present invention, feature extraction is performed on multimedia data to obtain multimedia features of the multimedia data, including: using a dimensionality reduction matrix to perform dimensionality reduction processing on the multimedia data to obtain multimedia features.
[0084] Since multimedia data is usually multi-channel and multi-dimensional data, a dimensionality reduction matrix can be set. After converting the multimedia data into matrix form, matrix multiplication is performed with the dimensionality reduction matrix to obtain multimedia features with dimensions lower than those of the multimedia data. Among them, the method of converting multimedia data into matrix form may include converting multimedia data in picture form into a pixel value matrix of the picture, etc.
[0085] Since the processing result is for multimedia data, after processing the multimedia features to obtain standard features, the inverse process opposite to the dimensionality reduction process can be used to restore the standard features to obtain the multimedia data corresponding to the standard features, and then process it to obtain the standard result of the multimedia data.
[0086] According to an embodiment of the present invention, determining the target processing result of multimedia data based on standard features includes: using the inverse matrix of the dimensionality reduction matrix to process the standard features to obtain standard multimedia data corresponding to the standard features; and determining the target processing result based on the standard multimedia data.
[0087] According to an embodiment of the present invention, in the case where the dimensionality reduction process is to right-multiply the multimedia data by the dimensionality reduction matrix, the inverse process of this dimensionality reduction process can be to determine the inverse matrix of the dimensionality reduction matrix, calculate the right multiplication of the standard features by the inverse matrix of the dimensionality reduction matrix, and determine the calculation result as the standard multimedia data.
[0088] After determining the standard multimedia data, the processing method for multimedia data can be used to determine the processing result of the standard multimedia data. Since the standard multimedia data is the data obtained after causal intervention on the multimedia data, this processing result is the target processing result of the multimedia data.
[0089] According to an embodiment of the present invention, the encoder is trained in the following manner: feature extraction is performed on sample multimedia data to obtain sample multimedia features of the sample multimedia data; using an untrained original encoder to process the sample multimedia features to obtain a sample multimedia coding result; based on the recognition result of the discriminator for the sample multimedia coding result and the scrambled sample multimedia coding result, adjusting the parameters of the original encoder to obtain a first encoder, where the scrambled sample multimedia coding result is obtained by scrambling the sample multimedia coding result; based on the sample multimedia decoding result and the sample multimedia features, adjusting the parameters of the first encoder to obtain an encoder, where the sample multimedia decoding result is obtained by decoding the sample multimedia coding result.
[0090] After scrambling, the values in the multimedia coding result of the scrambled sample are the same as those in the multimedia coding of the sample, but the arrangement order is different. During the training process, the parameters of the encoder can be adjusted according to the properties of the multimedia coding result of the sample obtained by the encoder processing and the properties of the confounding factor characteristics of the sample multimedia data, so that after the adjusted encoder processes the multimedia data, the obtained coding result can be used as the confounding factor characteristics of the multimedia data.
[0091] The confounding factor characteristics need to be independent. After representing the confounding factor characteristics as a sequence, when multiple elements in the sequence can independently represent their respective semantics, it can be determined that the confounding factor characteristics are independent. In this case, after scrambling the sequence of the confounding factor characteristics, the semantics in the scrambled sequence remain unchanged. Therefore, the scrambled sequence and the sequence before scrambling can express the same semantics. Therefore, if the distance between the result obtained after scrambling the confounding factor characteristics and the confounding factor characteristics is small enough, it can be determined that the alternative confounding variable is independent. The parameters of the original encoder can be adjusted based on the distance between the multimedia coding result of the sample and the multimedia coding result of the scrambled sample, to obtain a first encoder that can output a multimedia coding result of the sample with independence, where the distance between the multimedia coding result of the scrambled sample and the characteristics of the multimedia coding result of the sample can be represented by the Euclidean distance between the above two variables.
[0092] The confounding factor characteristics also need to be reconstructable. Reconstructability means that after reconstructing the multimedia characteristics using the confounding factor characteristics, the distance between the obtained reconstruction variable and the multimedia characteristics is small enough, that is, the multimedia characteristics are encoded to obtain the confounding factor characteristics, and after decoding the confounding factor characteristics, the reconstructed variable and the variable before reconstruction can express the same semantics. Therefore, the parameters of the first encoder can be adjusted based on the multimedia decoding result of the sample and the multimedia characteristics of the sample, to obtain an encoder that can output a multimedia coding result of the sample with reconstructability.
[0093] According to the embodiments of the present invention, by checking the data output by the encoder, and when it is determined that the data output by the encoder does not have the properties that the confounding factor characteristics need to have, based on the results determined by other modules, the parameters of the encoder are adjusted so that the adjusted encoder can output the properties that the confounding factor characteristics need to have. Thus, it is ensured that in the application stage, the coding result output by the encoder can be directly used as the confounding factor characteristics for subsequent causal intervention.
[0094] According to an embodiment of the present invention, a discriminator is used to identify the sample multimedia coding result and the shuffled sample multimedia coding result to obtain an identification result, including: processing the sample multimedia coding result and the shuffled sample multimedia coding result by the discriminator to respectively obtain a sample multimedia judgment result and a shuffled sample multimedia judgment result; and determining the identification result based on the cross-entropy result determined according to the sample multimedia judgment result and the shuffled sample multimedia judgment result.
[0095] Use the discriminator to process the sample multimedia coding result C s and the shuffled sample multimedia coding result perm Cs respectively, to obtain the judgment result D(C s ) of the sample multimedia coding result C s and the judgment result D(perm Cs ) of the shuffled sample multimedia coding result perm Cs . Among them, the processing of the discriminator may include prediction and classification. For example, using the discriminator to classify the sample multimedia coding result C s and the shuffled sample multimedia coding result perm Cs respectively, the category D(C s ) corresponding to the sample multimedia coding result and the category D(perm Cs ) corresponding to the shuffled sample multimedia coding result can be determined respectively.
[0096] Among them, perm Cs can be obtained by formula (5):
[0097] (5)
[0098] Among them, permutation(·) represents random permutation of a matrix or vector. Random permutation means that the element values in the matrix or vector remain unchanged while the positional relationship changes.
[0099] To ensure the training effect and generalization ability of the discriminator, a stream of data including multiple sample multimedia data can be used to train the discriminator. Based on the cross-entropy result BCE determined according to multiple sample multimedia judgment results and shuffled sample multimedia judgment results, the identification result of the discriminator can be determined, as shown in formula (6):
[0100] (6)
[0101] Among them, n represents the sequence of the current sample multimedia data in the stream of data.
[0102] According to an embodiment of the present invention, based on the recognition result of the discriminator on the sample multimedia encoding result and the scrambled sample multimedia encoding result, the parameters of the original encoder are adjusted to obtain a first encoder, including: determining the recognition gradient of the original encoder based on the recognition result; processing the recognition gradient using a gradient reversal layer to obtain a reversed gradient; and adjusting the parameters of the original encoder based on the reversed gradient until the recognition result indicates that the distance between the sample multimedia encoding result and the scrambled sample multimedia encoding result is less than a preset value.
[0103] Based on the above recognition result, the recognition gradient of the original encoder can be determined, where the recognition gradient is used to represent the parameter adjustment direction of the original encoder. When adjusting the parameters in this direction, the discriminator has a stronger recognition ability for the sample multimedia encoding result and the scrambled sample multimedia encoding result. Since the parameters of the discriminator are not adjusted, it can be determined that by adjusting the parameters of the original encoder according to the recognition gradient, the distance between the output sample multimedia encoding result and the scrambled sample multimedia encoding result obtained after scrambling it will be greater to ensure that the discriminator has a stronger recognition ability.
[0104] Since the purpose of adjusting the parameters of the original encoder is to make the model unable to distinguish between the sample multimedia encoding result and the scrambled sample multimedia encoding result obtained after scrambling it in the feature space, it is necessary to adjust the parameters in the opposite direction of the recognition gradient. A gradient reversal layer can be used to multiply the recognition gradient by a fixed value less than 0, and this product is determined as the reversed gradient. Based on the reversed gradient to determine the parameter adjustment direction, and adjusting the parameters of the original encoder according to the parameter adjustment direction and the preset parameter adjustment step and learning rate can reduce the distance between the sample multimedia encoding result output by the adjusted encoder and the scrambled sample multimedia encoding result obtained after scrambling it.
[0105] Continuously adjust the model parameters until the recognition result indicates that the distance between the sample multimedia encoding result and the scrambled sample multimedia encoding result is less than a preset value. The first encoder can be determined based on the current model parameters, where the distance between the sample multimedia encoding result output by the first encoder before and after scrambling is less than the preset value, that is, scrambling does not affect the semantics of each element in the sample multimedia encoding result. Therefore, the first encoder can output a sample multimedia encoding result with independence.
[0106] According to an embodiment of the present invention, by fixing the model parameters of the discriminator and training the encoder, it is possible to reduce the distance between the encoding result generated by the encoder and the result after scrambling it while the recognition ability of the discriminator is fixed, so that the encoder has the ability to generate an encoding result with independence.
[0107] Among them, the discriminator is trained as follows: determining the scrambling loss value between the sample multimedia encoding result and the scrambled sample multimedia encoding result; and adjusting the model parameters of the discriminator based on the scrambling loss until the scrambling loss value converges.
[0108] The scrambling loss value can be used to represent the loss caused by scrambling the sample multimedia encoding result. Based on the sample multimedia encoding result and the scrambled sample multimedia encoding, the scrambling loss threshold between the two is determined. Among them, when the sample multimedia encoding result has independence, the scrambling loss threshold between the two is determined as the first threshold, and when the sample multimedia encoding result does not have independence, the scrambling loss threshold between the two is determined as the second threshold. Among them, the first threshold and the second threshold can be set in advance, and the first threshold is less than the second threshold.
[0109] Based on the determined scrambling loss threshold and the scrambling loss value between the sample multimedia encoding result determined by the discriminator and the scrambled sample multimedia encoding result, the model parameters of the discriminator are adjusted until the difference between the scrambling loss value output by the discriminator and the scrambling loss threshold is less than the preset value, and the model parameters of the discriminator are determined. Thus, it is ensured that the discriminator has good recognition ability and the training effect of training the encoder according to the result of the discriminator is ensured.
[0110] According to an embodiment of the present invention, based on the sample multimedia decoding result and the sample multimedia feature, the parameters of the first encoder are adjusted to obtain an encoder, including: determining a loss value based on the sample multimedia decoding result and the sample multimedia feature; and adjusting the parameters of the first encoder based on the loss value until the loss value converges.
[0111] According to an embodiment of the present invention, the loss value between the sample multimedia decoding result and the sample multimedia feature may include the mean square error, root mean square error, etc. between the two. Taking the mean square error as an example, the loss value MSE between the sample multimedia decoding result and the sample multimedia feature can be determined by formula (7):
[0112] (7)
[0113] Among them, represents the i-th sample multimedia feature in the stream data of the multimedia data, represents the i-th sample multimedia decoding result in the stream data of the multimedia data.
[0114] Among them, can be obtained by formula (8):
[0115] (8)
[0116] Among them, Ci Denote the multimedia encoding result of the i-th sample in the streaming data representing multimedia data.
[0117] Using the gradient descent method, adjust the parameters of the first encoder based on the loss value, determine the model parameters when the loss value converges, and determine the encoder based on the model parameters.
[0118] According to an embodiment of the present invention, nonlinear independent component analysis is adopted. According to the loss value between the decoded result of the sample multimedia after encoding and decoding and the sample multimedia feature, determine the difference between the sample multimedia features before and after reconstruction, so as to determine the reconstructibility of the sample multimedia feature, so as to ensure that after adjusting the model parameters based on the loss value, the encoder has the ability to generate a reconstructible encoding result.
[0119] Figure 4 Shows the structural diagram of the causal inference model of the confounding factor feature and adversarial learning according to an embodiment of the present invention.
[0120] As Figure 4 shown, in the training stage, use the encoder to process the sample multimedia feature X' of the sample multimedia data to obtain the sample multimedia encoding result C'.
[0121] The model training part can be divided into two stages, namely the adversarial training stage and the autoencoder training stage.
[0122] Randomly arrange the sample multimedia encoding result C' to obtain the scrambled sample multimedia encoding result perm C , and use the discriminator to process the sample multimedia encoding result C' and the scrambled sample multimedia encoding result perm C , and determine the recognition gradient according to the processing result. The gradient reversal layer can be used to reverse the gradient for subsequent training.
[0123] Input the sample multimedia encoding result C' into the decoder to obtain the reconstructed sample multimedia decoding result , and determine the loss value according to the sample multimedia feature X' and the multimedia decoding result, ensuring that the sample multimedia feature X' processed by the encoder can be restored to the sample multimedia feature after decoding, and complete the parameter adjustment in the autoencoder training stage.
[0124] After training is completed, the encoder has the ability to generate an independent and reconstructible encoding result. Therefore, after using the encoder to process the multimedia feature X' to obtain C', it can be determined that C' is the confounding factor feature of the sample multimedia feature X'. Therefore, based on the confounding factor feature C', causal intervention on the sample multimedia feature X' can obtain the standard feature of the multimedia data.
[0125] In one embodiment, in the case of performing a video anomaly detection task, the multimedia data is video stream data. A feature extraction model is used to process video frames to obtain a feature sequence of the video, where the feature sequence includes a plurality of multimedia features arranged in the order of a plurality of video frames in the video stream data. An encoder is used to process the feature sequence to obtain a confounding factor feature, and the confounding factor feature can be used to perform causal intervention on the feature sequence to eliminate the influence of confounding variables on the detection result between the feature sequence and the video.
[0126] Based on the above method for processing multimedia data, the present invention also provides a device for processing multimedia data. The following will be combined with Figure 5 to describe the device in detail.
[0127] Figure 5 The structural block diagram of the device for processing multimedia data according to an embodiment of the present invention is shown.
[0128] As Figure 5 shown, the device 500 for processing multimedia data in this embodiment includes a feature extraction module 510, a feature processing module 520, a feature determination module 530, and a result determination module 540.
[0129] The feature extraction module 510 is used to extract features from the multimedia data to obtain the multimedia features of the multimedia data, where the multimedia features include the features of multiple different factors of the multimedia data, and the factors include noise factors. In one embodiment, the feature extraction module 510 can be used to perform the operation S210 described above, which will not be elaborated here.
[0130] The feature processing module 520 is used to use an encoder to process the multimedia features to obtain a confounding factor feature, where the confounding factor feature includes the features of noise factors, and the noise factors include the factors that affect the causal relationship measurement between the multimedia data and the result. In one embodiment, the feature processing module 520 can be used to perform the operation S220 described above, which will not be elaborated here.
[0131] The feature determination module 530 is used to remove the features of noise factors from the multimedia features based on the confounding factor feature to obtain the standard features of the multimedia data. In one embodiment, the feature determination module 530 can be used to perform the operation S230 described above, which will not be elaborated here.
[0132] The result determination module 540 is used to determine the target processing result of the multimedia data based on the standard features. In one embodiment, the result determination module 540 can be used to perform the operation S240 described above, which will not be elaborated here.
[0133] According to an embodiment of the present invention, the feature determination module 530 includes a first feature determination sub-module, a second feature determination sub-module, and a feature determination sub-module.
[0134] The first feature determination sub-module is configured to perform feature extraction on the confounding factor feature and the multimedia feature, and determine a joint feature characterizing the associated influence between the multimedia feature and the confounding factor feature.
[0135] The second feature determination sub-module is configured to determine a conditional feature for removing the feature of the noise factor based on the joint feature and the confounding factor feature.
[0136] The feature determination sub-module is configured to determine a standard feature by using the conditional feature and the multimedia feature.
[0137] According to an embodiment of the present invention, the feature extraction module 510 includes a data dimensionality reduction sub-module.
[0138] The data dimensionality reduction sub-module is configured to perform dimensionality reduction processing on the multimedia data by using a dimensionality reduction matrix to obtain multimedia features.
[0139] According to an embodiment of the present invention, the result determination module 540 includes a feature restoration sub-module and a result determination sub-module.
[0140] The feature restoration sub-module is configured to process the standard feature by using the inverse matrix of the dimensionality reduction matrix to obtain standard multimedia data corresponding to the standard feature.
[0141] The result determination sub-module is configured to determine a target processing result based on the standard multimedia data.
[0142] According to an embodiment of the present invention, the multimedia data processing device 500 further includes a sample feature extraction module, a sample feature processing module, a first parameter adjustment module, and a second parameter adjustment module.
[0143] The sample feature extraction module is configured to perform feature extraction on the sample multimedia data to obtain sample multimedia features of the sample multimedia data, where the sample multimedia features include features of multiple different sample factors of the sample multimedia data.
[0144] The sample feature processing module is configured to process the sample multimedia features by using an original encoder to obtain a sample multimedia coding result.
[0145] The first parameter adjustment module is configured to adjust the parameters of the original encoder based on the recognition result of the discriminator for the sample multimedia coding result and the scrambled sample multimedia coding result to obtain a first encoder, where the scrambled sample multimedia coding result is obtained by scrambling the sample multimedia coding result.
[0146] The second parameter adjustment module is used to adjust the parameters of the first encoder based on the sample multimedia decoding result and the sample multimedia feature to obtain an encoder, where the sample multimedia decoding result is obtained by decoding the sample multimedia encoding result.
[0147] According to an embodiment of the present invention, the multimedia data processing device 500 further includes an encoding scrambling module, a result recognition module, and a result processing module.
[0148] The encoding scrambling module is used to scramble the sample multimedia encoding to obtain a scrambled sample multimedia encoding result.
[0149] The result recognition module is used to use a discriminator to recognize the sample multimedia encoding result and the scrambled sample multimedia encoding result to obtain a recognition result, where the recognition result is used to represent the distance between the sample multimedia encoding result and the scrambled sample multimedia encoding result.
[0150] The result processing module is used to use a decoder to process the encoding result to obtain a sample multimedia decoding result.
[0151] According to an embodiment of the present invention, the result recognition module includes a first result recognition sub-module and a second result recognition sub-module.
[0152] The first result recognition sub-module is used to use a discriminator to process the sample multimedia encoding result and the scrambled sample multimedia encoding result to obtain a sample multimedia judgment result and a scrambled sample multimedia judgment result respectively.
[0153] The second result recognition sub-module is used to determine the recognition result based on the cross-entropy result determined according to the sample multimedia judgment result and the scrambled sample multimedia judgment result.
[0154] According to an embodiment of the present invention, the first parameter adjustment module includes a gradient determination sub-module, a gradient inversion sub-module, and a first parameter adjustment sub-module.
[0155] The gradient determination sub-module is used to determine the gradient direction and gradient value of the original encoder based on the recognition result.
[0156] The gradient inversion sub-module is used to invert the gradient direction to obtain an inverted gradient direction.
[0157] The first parameter adjustment sub-module is used to adjust the parameters of the original encoder based on the inverted gradient direction and the gradient value until the recognition result represents that the distance between the sample multimedia encoding result and the scrambled sample multimedia encoding result is less than a preset value.
[0158] According to an embodiment of the present invention, the second parameter adjustment module includes a loss value determination sub-module and a second parameter adjustment sub-module.
[0159] A loss value determination sub-module, configured to determine a loss value based on a sample multimedia decoding result and sample multimedia features.
[0160] A second parameter adjustment sub-module, configured to adjust parameters of the first encoder based on the loss value until the loss value converges.
[0161] According to an embodiment of the present invention, any plurality of modules among the feature extraction module 510, the feature processing module 520, the feature determination module 530, and the result determination module 540 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the feature extraction module 510, the feature processing module 520, the feature determination module 530, and the result determination module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the feature extraction module 510, the feature processing module 520, the feature determination module 530, and the result determination module 540 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute corresponding functions.
[0162] Figure 6 A block diagram of an electronic device suitable for implementing a method for processing multimedia data according to an embodiment of the present invention is shown.
[0163] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0164] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to an embodiment of the present invention by executing programs in the ROM 602 and / or the RAM 603. It should be noted that the programs can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to an embodiment of the present invention by executing programs stored in one or more memories.
[0165] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.
[0166] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to an embodiment of the present invention is implemented.
[0167] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than the ROM 602 and RAM 603.
[0168] An embodiment of the present invention also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present invention.
[0169] When the computer program is executed by the processor 601, it executes the above functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0170] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0171] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or be installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0172] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, programming languages such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0174] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0175] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A method for processing multimedia data, characterized in that: The method comprises: Extracting features from the multimedia data to obtain multimedia features of the multimedia data, wherein the multimedia features include features of multiple different factors of the multimedia data, and the factors include noise factors; Using an encoder, processing the multimedia features to obtain confounding factor features, wherein the confounding factor features include features of the noise factors, and the noise factors include factors that affect the causal relationship measurement between the multimedia data and the results; Based on the confounding factor feature, removing the feature of the noise factor from the multimedia feature to obtain a standard feature of the multimedia data; and Based on the standard features, determining a target processing result of the multimedia data; The encoder is trained in the following manner: Extracting features from the sample multimedia data to obtain sample multimedia features of the sample multimedia data, wherein the sample multimedia features include features of a plurality of different sample factors of the sample multimedia data; Using an original encoder, processing the sample multimedia features to obtain a sample multimedia encoding result; Based on the recognition result of the discriminator on the sample multimedia encoding result and the disordered sample multimedia encoding result, the parameters of the original encoder are adjusted to obtain a first encoder, wherein the disordered sample multimedia encoding result is obtained by disordering the sample multimedia encoding result; Based on the sample multimedia decoding result and the sample multimedia feature, the parameters of the first encoder are adjusted to obtain the encoder, wherein the sample multimedia decoding result is obtained by decoding the sample multimedia encoding result.
2. The method according to claim 1, characterized in that The removing the feature of the noise factor from the multimedia feature based on the confounding factor feature to obtain the standard feature of the multimedia data includes: Performing feature extraction on the confounding factor feature and the multimedia feature to determine a joint feature representing a correlation influence between the multimedia feature and the confounding factor feature; Determining conditional features for removing features of the noise factor based on the joint features and the confounding factor features; and The standard feature is determined using the condition feature and the multimedia feature.
3. The method according to claim 1, characterized in that The extracting features of the multimedia data to obtain multimedia features of the multimedia data includes: The multimedia data is subjected to dimensionality reduction processing by using a dimensionality reduction matrix to obtain the multimedia features.
4. The method according to claim 3, characterized in that The determining, based on the standard feature, a target processing result of the multimedia data includes: Processing the standard features using the inverse matrix of the dimension reduction matrix to obtain standard multimedia data corresponding to the standard features; and Based on the standard multimedia data, the target processing result is determined.
5. The method according to claim 1, characterized in that The method further comprises: Performing random processing on the sample multimedia coding to obtain the random sample multimedia coding result; Using the discriminator to identify the sample multimedia coding result and the disordered sample multimedia coding result to obtain the identification result, wherein the identification result is used to represent the distance between the sample multimedia coding result and the disordered sample multimedia coding result; and The encoding result is processed by a decoder to obtain the sample multimedia decoding result.
6. The method according to claim 5, characterized in that The step of using the discriminator to identify the sample multimedia encoding result and the disordered sample multimedia encoding result to obtain the identification result includes: Processing the sample multimedia coding result and the disordered sample multimedia coding result by the discriminator to obtain a sample multimedia judgment result and a disordered sample multimedia judgment result respectively; and The recognition result is determined based on a cross entropy result determined according to the sample multimedia judgment result and the disordered sample multimedia judgment result.
7. The method according to claim 1, characterized in that The method of adjusting the parameters of the original encoder based on the recognition result of the sample multimedia encoding result and the disordered sample multimedia encoding result by the discriminator to obtain the first encoder includes: Based on the recognition result, determining the gradient direction and gradient value of the original encoder; Reversing the gradient direction to obtain a reversed gradient direction; and Based on the reversed gradient direction and the gradient value, the parameters of the original encoder are adjusted until the recognition result indicates that the distance between the sample multimedia encoding result and the out-of-order sample multimedia encoding result is less than a preset value.
8. The method according to claim 5, characterized in that The adjusting the parameters of the first encoder based on the sample multimedia decoding result and the sample multimedia feature to obtain the encoder includes: Determining a loss value based on the sample multimedia decoding result and the sample multimedia feature; and Parameters of the first encoder are adjusted based on the loss value until the loss value converges.
9. A multimedia data processing device, characterized in that: The device comprises: A feature extraction module, configured to extract features from the multimedia data to obtain multimedia features of the multimedia data, wherein the multimedia features include features of respective factors of the multimedia data, and the factors include noise factors; A feature processing module, configured to process the multimedia features using an encoder to obtain a confounding factor feature, wherein the confounding factor feature includes a feature of the noise factor, and the noise factor includes a factor that affects the causal relationship measurement between the multimedia data and the result; a feature determination module, configured to remove the feature of the noise factor from the multimedia feature based on the confounding factor feature, so as to obtain a standard feature of the multimedia data; and A result determination module, used for determining a target processing result of the multimedia data based on the standard feature; The device also includes: A sample feature extraction module, used to extract features from sample multimedia data to obtain sample multimedia features of the sample multimedia data, wherein the sample multimedia features include features of multiple different sample factors of the sample multimedia data; A sample feature processing module, used to process the sample multimedia features using the original encoder to obtain a sample multimedia encoding result; A first parameter adjustment module is used to adjust the parameters of the original encoder based on the recognition result of the sample multimedia encoding result and the disordered sample multimedia encoding result by the discriminator to obtain a first encoder, wherein the disordered sample multimedia encoding result is obtained by disordering the sample multimedia encoding result; The second parameter adjustment module is used to adjust the parameters of the first encoder based on the sample multimedia decoding result and the sample multimedia feature to obtain the encoder, wherein the sample multimedia decoding result is obtained by decoding the sample multimedia encoding result.
Citation Information
Patent Citations
Road crack detection method based on causal decoupling, electronic equipment and storage medium
CN119360085A