A Multimodal Cross-Language Detection Method and Device

Through multimodal cross-language detection method, combined with deep learning and knowledge distillation technology, the neglect of multimodal information on papers and the shortcomings of multiple cross-language detection in the existing technology are solved, and the accurate similarity detection of paper data in different languages ​​is achieved, and academic integrity is effectively maintained.

CN118395205BActive Publication Date: 2025-05-30BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410597296.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-14
Publication Date
2025-05-30
Estimated Expiration
2044-05-14

AI Technical Summary

Technical Problem

The existing paper detection methods mainly focus on a single text modality, ignore other modal information such as tables and formulas in the paper, and cannot effectively detect the behavior of multiple copies of one draft across languages, resulting in the erosion of academic integrity.

Method used

A multimodal cross-language detection method is proposed. By obtaining the original data of the paper and decomposing it into ordinary text modes, structural text modes and image modes, the features are extracted using deep learning technology, and combined with knowledge distillation and fine-tuning technology to improve the cross-language detection ability. Finally, the Gaussian naive Bayes model is used for multimodal fusion detection.

Benefits of technology

It realizes accurate similarity detection of paper data in different languages, effectively reduces the behavior of multiple copies of one draft, and maintains academic integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118395205B_ABST
    Figure CN118395205B_ABST
Patent Text Reader

Abstract

This application proposes a multi-modal cross-language detection method and device, which relates to the field of computer technology. Among them, the method includes: obtaining the original paper data, and decomposing the original paper data into ordinary text modality, structured text modality, and image modality through a parsing tool; inputting the ordinary text modality into a cross-language similarity model to output first vector data, and inputting the structured text modality into the cross-language similarity model to output second vector data; inputting the image modality into an image similarity model to output third vector data; inputting the first vector data, second vector data, and third vector data into a multi-modal fusion learning model to output a detection result. The application adopting the above solution can effectively realize the similarity detection of paper data in different languages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a multi-modal cross-language detection method and device. Background Art

[0002] The increasing number of academic misconduct cases year by year is eroding the integrity foundation of the scientific research environment. And duplicate publication refers to the act of an author publishing two or more papers using the same data, core ideas or conclusions without mutual citation. Duplicate publication will double count research results, thus enabling the author to obtain unfair benefits, and this behavior seriously undermines the atmosphere of fairness and justice in scientific research.

[0003] Currently, most research methods for paper detection focus on single text modality, ignoring other modality information such as tables and formulas in papers. At the same time, for text modality, illegal authors often take advantage of the differences in syntactic structures and word usage between different languages to conduct cross-language duplicate publication in order to evade paper detection systems.

[0004] Therefore, based on multi-modal fusion technology, developing a multi-modal cross-language detection method can effectively reduce duplicate publication behavior and maintain academic integrity. Summary of the Invention

[0005] This application aims to solve at least one of the technical problems in the related art to some extent.

[0006] To this end, the first object of this application is to propose a multi-modal cross-language detection method, which realizes the accurate detection of the similarity of paper data in different languages.

[0007] The second object of this application is to propose a multi-modal cross-language detection device.

[0008] To achieve the above object, the first aspect embodiment of this application proposes a multi-modal cross-language detection method, including: obtaining the original paper data, and decomposing the original paper data into ordinary text modality, structured text modality and image modality through a parsing tool; inputting the ordinary text modality into a cross-language similarity model to output first vector data, and inputting the structured text modality into the cross-language similarity model to output second vector data; inputting the image modality into an image similarity model to output third vector data; inputting the first vector data, the second vector data and the third vector data into a multi-modal fusion learning model to output a detection result.

[0009] The multi-modal cross-language detection method of the embodiments of the present application uses the object detection technology in deep learning to extract ordinary images, tables, and formulas, extracts the main text, major headings, and chapter headings in the paper by recognizing the fonts and font sizes of the PDF, and extracts the keywords and abstracts in the PDF by using regular expressions, making the parsed data more comprehensive; uses the knowledge distillation technology to enable the single-language semantic similarity model SBERT to obtain cross-language capabilities, and at the same time uses the fine-tuning technology to enable this cross-language semantic similarity model to have the ability to understand academic proper nouns, effectively improving the cross-language detection ability of the model; uses the late fusion method, adopts the Gaussian Naive Bayes model as the multi-modal fusion learning model, combines the text similarity, image similarity, and structural similarity to complete the detection of multiple submissions of the same manuscript, and realizes the fusion detection of multiple modalities.

[0010] Optionally, in an embodiment of the present application, the original paper data is decomposed into ordinary text modality, structural text modality, and image modality through a parsing tool, including:

[0011] Extract the main text and headings in the original paper data based on the fonts and font sizes of the text, and use regular expressions to extract the keywords and abstracts in the original paper data. The extracted main text is used as the ordinary text modality, and the extracted headings, keywords, and abstracts are used as the structural text modality, where the headings include the major headings and chapter headings of the paper;

[0012] Use an object detection model to extract the images, tables, and formulas in the original paper data to obtain the initial image modality, and correct the initial image modality to obtain the image modality.

[0013] Optionally, in an embodiment of the present application, correcting the initial image modality to obtain the image modality includes:

[0014] Determine the image center coordinates and recognition types of each image in the initial image modality, where the recognition types include images, tables, and formulas;

[0015] Through the text detection method, obtain the text blocks related to the recognition type in the original paper data, and determine the center coordinates and keyword types of the text blocks, where the keyword types include table types, image types, and ordinary text types, and the formula recognition type corresponds to the ordinary text type;

[0016] Calculate the Euclidean distance based on the image center coordinates of each image and the center coordinates of the text blocks, and determine several text blocks with the smallest distance from each image;

[0017] Judge whether the recognition type of each image is in the keyword types of the corresponding several text blocks. If it is, keep the recognition type of the image. If not, modify the recognition type of the image to the keyword type of the text block with the closest distance.

[0018] Optionally, in an embodiment of the present application, constructing a cross-lingual similarity model includes:

[0019] Using the monolingual semantic similarity model SBERT as the teacher model and the multilingual model XLM-R as the student model, and performing knowledge distillation through multilingual parallel corpora to obtain the model after knowledge distillation;

[0020] Fine-tuning the model after knowledge distillation in an incremental training manner, and using multilingual academic parallel corpora as training data during fine-tuning to obtain a cross-lingual similarity model.

[0021] Optionally, in an embodiment of the present application, during knowledge distillation, setting the teacher model h T , the student model h S , given the multilingual parallel corpus as {(s 1 , t 1 ), (s 2 , t 2 ), …, (s n , t n )}, the corresponding training loss function is expressed as:

[0022]

[0023] where s n and t n represent corpora in different languages.

[0024] Optionally, in an embodiment of the present application, before inputting the ordinary text modality into the cross-lingual similarity model, it further includes:

[0025] Segmenting the ordinary text modality by sentence to obtain the processed ordinary text modality;

[0026] Inputting the processed ordinary text modality into the cross-lingual similarity model, and outputting the first vector data, including:

[0027] Vectorizing the processed ordinary text modality to generate a semantic matrix;

[0028] Using cosine similarity to recall and match each semantic vector in the semantic matrix in the semantic matrix of the article to be matched, and generating a similar semantic matrix for each article to be matched;

[0029] Determining the semantic similarity between the ordinary text modality and each article to be matched based on the semantic matrix and the similar semantic matrix of each article to be matched.

[0030] Optionally, in an embodiment of the present application, for the semantic matrix S 1Each semantic vector a in has a corresponding semantic matrix S of the article to be matched 2 with the similar semantic vector b in

[0031]

[0032] where x represents the semantic matrix S 2 and is a semantic vector in

[0033] Based on the semantic matrix S 1 each semantic vector a in and the corresponding semantic matrix S of the article to be matched 2 with the similar semantic vector b in form a similar semantic matrix S 3 , S 1 ={a 1 , a 2 , …, a n}, S 3 ={b 1 , b 2 , …, b n}, and the semantic similarity between the ordinary text modality and the article to be matched is expressed as:

[0034]

[0035] where a i is the i-th element in the semantic matrix S 1 , and b i is the i-th element in the similar semantic matrix S i corresponding to a 3 .

[0036] Optionally, in an embodiment of the present application, the first vector data, the second vector data, and the third vector data are input into a multi-modal fusion learning model, and a detection result is output, including:

[0037] Using the late fusion method, a Gaussian Naive Bayes model is used as the multi-modal fusion learning model to fuse the text similarity, the image similarity, and the structure similarity to obtain the detection result.

[0038] Optionally, in an embodiment of the present application, training the multi-modal fusion learning model includes:

[0039] Obtaining a training data set;

[0040] Using the training data set to train the multi-modal fusion learning model, and during training, weight update parameters are adopted to increase the weight of positive samples in the data set for model update and reduce the weight of negative samples in the data set for model update, where the weight formula is expressed as:

[0041]

[0042] Among them, α is the weight when updating positive samples, β is the weight when updating negative samples, N is the total number of samples, n 1 is the number of positive samples, and n 2 is the number of negative samples.

[0043] To achieve the above object, an embodiment of the second aspect of the present application proposes a multi-modal cross-language detection device, including a paper parsing module, a cross-language text similarity calculation module, an image similarity calculation module, and a similarity fusion module. Among them,

[0044] The paper parsing module is used to obtain the original paper data and decompose the original paper data into ordinary text modality, structured text modality, and image modality through a parsing tool;

[0045] The cross-language text similarity calculation module is used to input the ordinary text modality into a cross-language similarity model, output the first vector data, and input the structured text modality into the cross-language similarity model, output the second vector data;

[0046] The image similarity calculation module is used to input the image modality into an image similarity model and output the third vector data;

[0047] The similarity fusion module is used to input the first vector data, the second vector data, and the third vector data into a multi-modal fusion learning model and output the detection result.

[0048] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present application. Description of the Drawings

[0049] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0050] Figure 1 is a schematic flowchart of a multi-modal cross-language detection method provided by Embodiment 1 of the present application;

[0051] Figure 2 is a framework diagram of a multi-modal cross-language duplicate submission detection system according to an embodiment of the present application;

[0052] Figure 3 is an example diagram of the recognition effect according to an embodiment of the present application;

[0053] Figure 4 is a schematic diagram of a multi-modal fusion method according to an embodiment of the present application;

[0054] Figure 5 is a schematic structural diagram of a multi-modal cross-language detection device provided by an embodiment of the present application. Detailed Implementation Modes

[0055] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.

[0056] The multi-modal cross-language detection method and device according to the embodiments of the present application will be described below with reference to the accompanying drawings.

[0057] Figure 1 It is a schematic flowchart of a multi-modal cross-language detection method provided by Embodiment 1 of the present application.

[0058] As Figure 1 shown, the multi-modal cross-language detection method includes the following steps:

[0059] Step 101, obtain the original paper data, and decompose the original paper data into a plain text modality, a structured text modality, and an image modality through a parsing tool;

[0060] Step 102, input the plain text modality into a cross-language similarity model, output first vector data, and input the structured text modality into the cross-language similarity model, output second vector data;

[0061] Step 103, input the image modality into an image similarity model, output third vector data;

[0062] Step 104, input the first vector data, the second vector data, and the third vector data into a multi-modal fusion learning model, and output a detection result.

[0063] The multi-modal cross-language detection method according to the embodiments of the present application uses the object detection technology in deep learning to extract ordinary images, tables, and formulas, extracts the main text, big titles, and chapter titles in the paper by identifying the PDF font and font size, and extracts keywords and abstracts in the PDF by using regular expressions, making the parsed data more comprehensive; uses the knowledge distillation technology to endow the single-language semantic similarity model SBERT with cross-language capabilities, and at the same time uses the fine-tuning technology to enable the cross-language semantic similarity model to have the ability to understand academic proper nouns, effectively improving the cross-language detection ability of the model; uses the late fusion method, adopts the Gaussian Naive Bayes model as the multi-modal fusion learning model, combines text similarity, image similarity, and structural similarity to complete the detection of duplicate submissions, and realizes the fusion detection of multiple modalities.

[0064] Optionally, in an embodiment of the present application, decomposing the original paper data into a plain text modality, a structured text modality, and an image modality through a parsing tool includes:

[0065] Extract the main text and titles in the original data of the paper based on the font and font size of the text, and use regular expressions to extract the keywords and abstracts in the original data of the paper. Take the extracted main text as the ordinary text modality, and take the extracted titles, keywords and abstracts as the structured text modality, where the titles include the main title of the paper and the chapter titles;

[0066] Use the object detection model to extract images, tables and formulas in the original data of the paper to obtain the initial image modality, and correct the initial image modality to obtain the image modality.

[0067] Optionally, in an embodiment of the present application, correcting the initial image modality to obtain the image modality includes:

[0068] Determine the image center coordinates and recognition types of each image in the initial image modality, where the recognition types include images, tables and formulas;

[0069] Through the text detection method, obtain the text blocks related to the recognition type in the original data of the paper, and determine the center coordinates and keyword types of the text blocks, where the keyword types include table types, image types and ordinary text types, and the formula recognition type corresponds to the ordinary text type;

[0070] Calculate the Euclidean distance based on the image center coordinates of each image and the center coordinates of the text blocks, and determine several text blocks with the smallest distance from each image;

[0071] Judge whether the recognition type of each image is in the keyword types of the corresponding several text blocks. If so, keep the recognition type of the image. If not, modify the recognition type of the image to the keyword type of the text block with the closest distance.

[0072] Optionally, in an embodiment of the present application, constructing a cross-language similarity model includes:

[0073] Use the single-language semantic similarity model SBERT as the teacher model and the multilingual model XLM-R as the student model, and perform knowledge distillation through multilingual parallel corpora to obtain the model after knowledge distillation;

[0074] Fine-tune the model after knowledge distillation in an incremental training manner, and use multilingual academic parallel corpora as training data during fine-tuning to obtain a cross-language similarity model.

[0075] Optionally, in an embodiment of the present application, when performing knowledge distillation, set the teacher model h T , student model h S , given the multilingual parallel corpus as {(s 1 ,t 1),(s 2 ,t 2 ),…,(s n ,t n )}, and the corresponding training loss function is expressed as:

[0076]

[0077] Among them, s n and t n represent corpora in different languages.

[0078] Optionally, in an embodiment of the present application, before inputting the ordinary text modality into the cross - language similarity model, it further includes:

[0079] Segment the ordinary text modality into sentences to obtain the processed ordinary text modality;

[0080] Input the processed ordinary text modality into the cross - language similarity model and output the first vector data, including:

[0081] Vectorize the processed ordinary text modality to generate a semantic matrix;

[0082] Use cosine similarity to recall and match each semantic vector in the semantic matrix in the semantic matrix of the article to be matched, and generate a similar semantic matrix for each article to be matched;

[0083] Determine the semantic similarity between the ordinary text modality and each article to be matched based on the semantic matrix and the similar semantic matrix of each article to be matched.

[0084] Optionally, in an embodiment of the present application, for each semantic vector a in the semantic matrix S 1 there is a corresponding similar semantic vector b in the semantic matrix S 2 of the article to be matched,

[0085]

[0086] Among them, x represents a semantic vector in the semantic matrix S 2 ;

[0087] Based on each semantic vector a in the semantic matrix S 1 and the corresponding similar semantic vector b in the semantic matrix S 2 of the article to be matched, a similar semantic matrix S 3 is formed, S 1 = {a 1 , a 2 , …, a n}, S 3 = {b 1 , b 2,…,b n}, the semantic similarity between the ordinary text modality and the article to be matched is expressed as:

[0088]

[0089] where a i is the i-th element in the semantic matrix S 1 , and b i is the i-th element in the similar semantic matrix S i corresponding to a. 3

[0090] Optionally, in an embodiment of the present application, the first vector data, the second vector data, and the third vector data are input into a multi-modal fusion learning model, and a detection result is output, including:

[0091] Using the late fusion method, a Gaussian Naive Bayes model is used as the multi-modal fusion learning model to fuse the text similarity, the image similarity, and the structural similarity to obtain the detection result.

[0092] Optionally, in an embodiment of the present application, training the multi-modal fusion learning model includes:

[0093] Obtain a training data set;

[0094] Use the training data set to train the multi-modal fusion learning model. When training, weight update parameters are adopted to increase the weight of positive samples in the data set for model update and reduce the weight of negative samples in the data set for model update. The weight formula is expressed as:

[0095]

[0096] where α is the weight for positive sample update, β is the weight for negative sample update, N is the total number of samples, n 1 is the number of positive samples, and n 2 is the number of negative samples.

[0097] To describe the solution of the present application in detail, this embodiment also proposes a multi-modal cross-language one-submission-multiple-publication detection system. Figure 2 is the overall framework diagram of the multi-modal cross-language one-submission-multiple-publication detection system of this embodiment, as shown in Figure 2 , and the specific process is as follows:

[0098] (1) The original data of the paper is decomposed into ordinary text modality, structural modality, and image modality through a PDF parsing tool.

[0099] ​Specifically, the text type is determined by identifying the font and font size in the PDF. For example, the font of the body text is the most frequently occurring font. According to the corresponding font, the corresponding text can be extracted. The font of the chapter title should be among the top three in terms of font size, and the number of sentences is below 50. Based on the font and font size, this embodiment can extract the body text, main title, and chapter title. By using regular expressions, the "keywords" and "abstract" in the Chinese-English bilingual version of the paper are identified, and the keyword content and abstract content are extracted. The extracted body text and abstract text are segmented into sentence-level text fragments, and the separators between the extracted keywords are removed using regular matching.

[0100] When extracting images, the object detection model YOLO v7 in deep learning is used, and three recognition categories are defined: images, tables, and formulas. Based on these three categories, a PDF standard dataset is constructed, and the object detection model is trained on this dataset. The recognition effect is as Figure 3 shown.

[0101] However, there are still accuracy problems when only using ordinary object detection, and there may be incorrect recognition types. Therefore, this embodiment needs to combine text parsing + the nearest matching algorithm to correct the recognized categories. The specific process is as follows:

[0102] 1) After preliminary recognition, the central coordinates (x, y) of the image and the recognition type type are obtained;

[0103] 2) Using the above text detection method, keywords such as "table, Table, figure, Fig, Image" in the PDF of this page are obtained, and the central coordinates of the text block are obtained. At the same time, the keyword types are marked: "table, Table" is marked as the table type, "figure, Fig, Image" is marked as the image type, and the rest of the text blocks are marked as the ordinary text type (generally, one line of text is recognized as a text block);

[0104] 3) Combining the central coordinates of the recognized image with the central coordinates of the keywords, the corresponding Euclidean distance is calculated, and the 3 keyword text blocks with the smallest distance from the image are obtained.

[0105] 4) When the image type is among the 3 keyword text block types, the image type is maintained; otherwise, the image type is corrected to the nearest keyword text block type (the formula type corresponds to the ordinary text type).

[0106] (2) The modality of the text type After text preprocessing, the text is segmented at the sentence level, and useless symbols and stop words are removed.

[0107] (3) The ordinary text modality and the structure modality are input into the multi-draft cross-language similarity model to obtain the corresponding modality vector data; the image modality is input into the multi-draft image similarity model to obtain the corresponding modality vector data;

[0108] Specifically, SBERT is a powerful monolingual semantic similarity model with good semantic understanding ability, but it cannot handle cross-lingual scenarios. XLM-R is a multilingual model, although it has multilingual understanding ability, its effect in semantic recognition and semantic similarity calculation is not ideal. Therefore, to combine the advantages of the two models, this embodiment uses knowledge distillation technology to not only retain the powerful semantic similarity calculation ability of SBERT but also obtain the multilingual ability of XLM-R.

[0109] Using knowledge distillation technology, the monolingual semantic similarity model SBERT is used as the teacher model, and the multilingual model XLM-R is used as the student model, so as to not only retain the powerful semantic understanding ability of SBERT but also obtain the multilingual ability of XLM-R. For the teacher model h T , the student model h S , given the Chinese-English parallel corpus as {(s 1 , t 1 ), (s 2 , t 2 ), …, (s n , t n ), the corresponding training loss function is:

[0110]

[0111] After this knowledge distillation, only the general language semantic understanding ability is obtained, and the understanding effect of the proper nouns in the academic field is not good. Therefore, the increment training method is used to fine-tune the model after knowledge distillation, and the Chinese-English academic parallel corpus is used to fine-tune the model so that it can obtain the ability to understand academic proper nouns, and thus a semantic similarity model for academic terms in cross-lingual scenarios can be obtained.

[0112] (4) Perform the most similar recall in the vector database according to the vector, and use the cosine similarity formula to obtain the similarity with the closest vector.

[0113] Specifically, for the calculation of the similarity of multiple submissions of the same manuscript, the process is as follows: The input text is segmented into sentences, and the segmented sentences are input into the text similarity model for vectorization, so as to generate the semantic matrix of a paper. When comparing the similarity of the main texts of two articles, the cosine similarity is used to recall and match each semantic vector in the article to be detected, so as to generate a similar semantic matrix. The average value of the cosine similarities of the corresponding semantic vectors of the two semantic matrices is calculated as the main text semantic similarity of the two articles. In mathematical expression form, it is as follows:

[0114] Define the semantic matrix of the article to be detected as S 1 , and the semantic matrix of the matching article as S 2, then for each semantic vector a of the semantic matrix of the article to be detected, there is a corresponding similar semantic vector b, where a ∈ S 1 , b ∈ S 2 And there is:

[0115]

[0116] The matrix composed of the similar semantic vectors b of all a is the similar semantic matrix S 3 , that is, S 1 ={a 1 , a 2 , …, a n}, S 3 ={b 1 , b 2 , …, b n}, the semantic similarity between the article to be detected and the matching article is:

[0117]

[0118] The evaluation models are divided into cross - language semantic similarity models and machine translation + same - language semantic similarity models. The cross - language semantic similarity models are as follows: the multilingual semantic similarity model MUSE that supports 16 languages; the large - language model Chinese - LLaMA - Alpaca - 2 - 7B; mSimCSE that realizes multilingual semantic similarity by contrastive learning. For the machine translation + same - language semantic similarity model route, in this embodiment, the T5 model with better translation performance and SBERT that performs very well in the same - language semantics are selected.

[0119] The index uses the Pearson correlation coefficient between the semantic similarity calculated by the model and the label value. The specific formula is as follows, where X is the semantic similarity calculated by the model and Y is the label value.

[0120]

[0121] The experimental results are shown in Table 1.

[0122]

[0123] Table 1

[0124] (5) Input the vector data of the ordinary text modality, structural modality and image modality into the multi - modal fusion learning model to complete the result prediction.

[0125] Specifically, the late fusion method in multimodal fusion is used to comprehensively analyze the calculation results of each modality to obtain the prediction result. Among them, the similarity model of the text modality and the structural modality is the SBERT model after knowledge distillation, and the similarity model of the image modality is the open-source Vision Transformer. The calculation formula for the similarity of the image modality is the same as that of the text modality, so it will not be elaborated here. The multimodal fusion method is as Figure 4 shown.

[0126] Since the multiple submissions dataset is an imbalanced dataset, common methods to solve the sample imbalance include oversampling or undersampling. This method constructs a dataset with a balanced sample distribution by increasing the number of repeated positive samples or reducing the number of negative samples. However, this method will directly change the sample distribution of the multiple submissions dataset and destroy the authenticity. Therefore, this paper uses the method of updating the weights of parameters to handle the problem of sample imbalance, that is, increasing the weight of positive samples for model update and reducing the weight of negative samples for model update. The specific weight calculation formula is as follows:

[0127]

[0128] where α is the weight for positive sample update, β is the weight for negative sample update, N is the total number of samples, n 1 is the number of positive samples, and n 2 is the number of negative samples.

[0129] To ensure the unity of the model during multimodal fusion, a decision tree is used as the multimodal fusion learning model in this part. As shown in Table 2, since some papers do not have images themselves, the effect of image + structure fusion is not good, and even the performance is not as excellent as that of a single text modality. On this dataset, the performance of the text modality combined with the image modality is better than that combined with the structural modality. Compared with the single text modality, the AUC is improved by about 6.1% and 4.0% respectively, and the F1 is improved by about 18.5% and 8.0% respectively. The prediction performance of the three modalities combined is the best, with an AUC of 0.899 and an F1 of 0.843. Compared with the single text modality, the AUC is improved by 8.3% and the F1 is improved by 34.4%. This also shows that using multimodal information does help to improve the model effect in the detection of multiple submissions.

[0130]

[0131] Table 2

[0132] By comparing multiple machine learning models, the optimal Gaussian Naive Bayes was finally selected as the multi-modal fusion learning model. Gaussian Naive Bayes demonstrated the best performance in terms of AUC and F1. Although Logistic Regression had a high AUC value, its F1 was only 0.545 and Precision was only 0.391, indicating a relatively high misjudgment rate of this model for this task. The experimental results are shown in Table 3.

[0133]

[0134] Table 3

[0135] To implement the above embodiments, the present application also proposes a multi-modal cross-language detection device.

[0136] Figure 5 It is a schematic structural diagram of a multi-modal cross-language detection device provided for the embodiments of the present application.

[0137] As Figure 5 shown, the multi-modal cross-language detection device includes a paper parsing module, a cross-language text similarity calculation module, an image similarity calculation module, and a similarity fusion module. Among them,

[0138] The paper parsing module is used to obtain the original paper data and decompose the original paper data into ordinary text modality, structural text modality, and image modality through a parsing tool;

[0139] The cross-language text similarity calculation module is used to input the ordinary text modality into the cross-language similarity model, output the first vector data, and input the structural text modality into the cross-language similarity model, output the second vector data;

[0140] The image similarity calculation module is used to input the image modality into the image similarity model and output the third vector data;

[0141] The similarity fusion module is used to input the first vector data, the second vector data, and the third vector data into the multi-modal fusion learning model and output the detection result.

[0142] It should be noted that the foregoing explanations of the embodiments of the multi-modal cross-language detection method also apply to the multi-modal cross-language detection device of this embodiment, and will not be elaborated here.

[0143] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0144] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0145] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of this application pertain.

[0146] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0147] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0148] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0149] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0150] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A multimodal cross-language detection method, characterized in that: The following steps are involved: The obtained raw data of the paper is decomposed into ordinary text mode, structured text mode and image mode through parsing tools; The ordinary text modality is segmented into sentences and then input into the cross-language similarity model, and the first vector data is output, including: Vectorize the common text modality to generate a semantic matrix; Using cosine similarity, recall matching is performed on each semantic vector in the semantic matrix in the semantic matrix of the article to be matched, and a similar semantic matrix of each article to be matched is generated; Determine the semantic similarity between the common text modality and each article to be matched based on the semantic matrix and the similar semantic matrix; For the semantic matrix Each semantic vector There is a corresponding semantic matrix of articles to be matched Similar semantic vectors in , Representation semantic matrix A semantic vector in Based on similar semantic vector Constructing a similar semantic matrix , , , the semantic similarity between the common text modality and the article to be matched is: The semantic matrix The elements, For The corresponding similarity matrix The elements; Inputting the structured text modality and the image modality into the cross-language similarity model and the image similarity model respectively, and outputting the second vector data and the third vector data; The first vector data, the second vector data and the third vector data are input into a multimodal fusion learning model, and a detection result is output.

2. The multimodal cross-language detection method according to claim 1, characterized in that: The original data of the paper is decomposed into ordinary text mode, structured text mode and image mode through parsing tools, including: Extracting the text and title from the original data of the paper based on the font and size of the text, and extracting the keywords and abstract from the original data of the paper using regular expressions, taking the extracted text as the ordinary text mode, and taking the extracted title, keywords and abstract as the structured text mode, wherein the title includes the main title of the paper and the chapter title; The target detection model is used to extract images, tables and formulas in the original data of the paper to obtain an initial image modality, and the initial image modality is corrected to obtain an image modality.

3. The multimodal cross-language detection method according to claim 2, characterized in that: The initial image modality is modified to obtain an image modality, including: Determining image center coordinates and a recognition type of each image in the initial image modality, wherein the recognition type includes an image, a table, and a formula; By using a text detection method, a text block related to the recognition type in the original data of the paper is obtained, and the center coordinates and keyword type of the text block are determined, wherein the keyword type includes a table type, an image type and an ordinary text type, and the formula recognition type corresponds to the ordinary text type; The Euclidean distance is calculated based on the image center coordinates of each image and the center coordinates of the text block, and a number of text blocks with the smallest distance to each image are determined; It is determined whether the recognition type of each image is in the keyword types of the corresponding text blocks. If so, the recognition type of the image is maintained; if not, the recognition type of the image is modified to the keyword type of the nearest text block.

4. The multimodal cross-language detection method according to claim 1, characterized in that: Constructing the cross-language similarity model includes: The single-language semantic similarity model SBERT is used as the teacher model, and the multi-language model XLM-R is used as the student model. Knowledge distillation is performed through multi-language parallel corpora to obtain the model after knowledge distillation. The model after knowledge distillation is fine-tuned by incremental training, and multilingual academic parallel corpora are used as training data during fine-tuning to obtain a cross-language similarity model.

5. The multimodal cross-language detection method according to claim 1, characterized in that: When distilling knowledge, set the teacher model , student model , given a multilingual parallel corpus , the corresponding training loss function is expressed as: in, , Represents corpora in different languages.

6. The multimodal cross-language detection method according to claim 1, characterized in that: Inputting the first vector data, the second vector data, and the third vector data into a multimodal fusion learning model, and outputting a detection result, including: By using the late fusion method, the Gaussian Naive Bayes model is used as the multimodal fusion learning model to fuse the text similarity, image similarity and structure similarity to obtain the detection result.

7. The multimodal cross-language detection method according to claim 1, characterized in that: Training the multimodal fusion learning model includes: Get the training dataset; The multimodal fusion learning model is trained using the training data set. During the training, a weight update parameter is used to increase the weight of the positive samples in the data set for the model update and reduce the weight of the negative samples in the data set for the model update. The weight formula is expressed as: in, is the weight of the positive sample update, is the weight when negative samples are updated, is the total number of samples, is the number of positive samples, is the number of negative samples.

8. A multimodal cross-language detection device, characterized in that: It includes paper parsing module, cross-language text similarity calculation module, image similarity calculation module and similarity fusion module, among which: The paper parsing module is used to decompose the acquired original paper data into ordinary text mode, structured text mode and image mode through a parsing tool; The cross-language text similarity calculation module is used to divide the common text modality into sentences and input them into the cross-language similarity model, and output the first vector data, including: Vectorize the common text modality to generate a semantic matrix; Using cosine similarity, recall matching is performed on each semantic vector in the semantic matrix in the semantic matrix of the article to be matched, and a similar semantic matrix of each article to be matched is generated; Determine the semantic similarity between the common text modality and each article to be matched based on the semantic matrix and the similar semantic matrix; For the semantic matrix Each semantic vector There is a corresponding semantic matrix of articles to be matched Similar semantic vectors in , Representation semantic matrix A semantic vector in Based on similar semantic vector Constructing a similar semantic matrix , , , the semantic similarity between the common text modality and the article to be matched is: The semantic matrix The elements, For The corresponding similarity matrix The elements; The cross-language text similarity calculation module is further used to input the structure text modality into the cross-language similarity model and output second vector data; The image similarity calculation module is used to input the image modality into the image similarity model and output the third vector data; The similarity fusion module is used to input the first vector data, the second vector data and the third vector data into the multimodal fusion learning model and output the detection result.

Citation Information

Patent Citations

  • Electric power defect image detection method based on image-text question-answer multi-modal model

    CN117763107A

  • Discovery of semantic similarities between images and text

    US20170061250A1