A method, apparatus, computer device and readable storage medium for screening works based on email

By obtaining emails from the mailbox server, filtering and compliance testing, combining multi-dimensional scoring and user review, the problem of low manual screening efficiency is solved, and efficient and accurate work screening is achieved.

CN119537571BActive Publication Date: 2025-07-11DARK MATTER ARTIFICIAL INTELLIGENT (BEIJING) TECHNOLOGY CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411469971.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-07-11
Estimated Expiration
2044-10-21

AI Technical Summary

Technical Problem

In the prior art, when receiving submissions from works through mailboxes, manual screening efficiency is low and subjective errors are prone to occur, making it difficult to achieve efficient and accurate work screening.

Method used

By obtaining emails from the target mailbox server, filtering and compliance detection, and using multi-dimensional scoring and user review, automated filtering is achieved.

Benefits of technology

Efficient and accurate work screening is achieved, manual intervention is reduced, and screening efficiency and accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537571B_ABST
    Figure CN119537571B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device, computer device and readable storage medium for screening works based on an email box, including: first obtaining original emails from a target email server and filtering them to obtain pending email works, then performing compliance detection based on the work type, obtaining a scoring result through multi-dimensional scoring after passing the compliance, and finally determining the final screening result according to the size of the result after the user reviews and approves the scoring result, so as to achieve efficient and accurate work screening.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method, device, computer device, and readable storage medium for screening works based on an email box. Background Art

[0002] In today's digital age, it has become increasingly common to receive submissions of various works (such as literary creations, art works, etc.) through email boxes. However, in the face of a large number of submission emails, how to efficiently and accurately screen out the works that meet the requirements has become a challenge. The traditional manual screening method is inefficient and prone to subjective errors. With the development of information technology, there is a need for a screening method with a high degree of automation and the ability to comprehensively evaluate works according to the type of works. Summary of the Invention

[0003] The purpose of the present invention is to provide a method, device, computer device, and readable storage medium for screening works based on an email box.

[0004] In a first aspect, an embodiment of the present invention provides a method for screening works based on an email box, including:

[0005] Obtain original emails from a target email server and filter them to obtain pending email works;

[0006] Perform work compliance detection based on the type of the pending email works;

[0007] In the case of passing the work compliance detection, perform multi-dimensional scoring on the pending email works to obtain a multi-dimensional scoring result of the pending email works;

[0008] In response to a user review passing instruction for the multi-dimensional scoring result, obtain a final screening result of the pending email works according to the size of the multi-dimensional scoring result.

[0009] In a possible implementation manner, the obtaining original emails from a target email server and filtering them to obtain pending email works includes:

[0010] Obtain the IMAP server address and port number of the target email server;

[0011] Establish a connection with the target email server according to the IMAP server address and the port number, in combination with login credentials;

[0012] Download the original emails in the target email server according to a preset submission time range, where the original emails include an original email body and an original email attachment;

[0013] Filter the original email body and the original email attachment according to the preset filtering rules. When the filtering result indicates compliance with the filtering rules, use the original email attachment as the pending email work.

[0014] In a possible implementation, the work compliance detection based on the work type of the pending email work includes:

[0015] Determine the detection standard for the pending email work based on the work type of the pending email work;

[0016] Perform sensitive information detection, duplicate submission detection, and theme compliance detection on the pending email work based on the detection standard.

[0017] In a possible implementation, the sensitive information detection for the pending email work includes:

[0018] When the pending email work is a text work, preprocess the text work and perform sentiment analysis, and combine with a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection;

[0019] When the pending email work is an image work, segment the image work, and combine with a preset scene recognition model to perform sensitive information detection; when the image work includes text content, extract the text from the image work, preprocess the text content and perform sentiment analysis, and combine with a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection;

[0020] When the pending email work is an audio work, use ASR to convert the video sound into audio text, preprocess the audio text and perform sentiment analysis, and combine with a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection;

[0021] When the pending email work is a video work, use ASR to convert the video sound into audio text, preprocess the audio text and perform sentiment analysis, and combine with a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection; extract multiple video images from the video work at preset time intervals, segment the video images, and combine with a preset scene recognition model to perform sensitive information detection; when the video images include text content, extract the text from the video images, preprocess the text content and perform sentiment analysis, and combine with a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection.

[0022] In a possible implementation, the duplicate submission detection for the pending email work includes:

[0023] When the to-be-determined email work is a text work, perform word segmentation on the text work to obtain multiple lexical units; remove the stop words from the multiple lexical units, and convert the multiple lexical units into multiple text vectors; perform text similarity comparison and semantic similarity comparison between the multiple text vectors and the archived text vectors in the work database, and perform the duplicate submission detection.

[0024] When the to-be-determined email work is an image work, use the SURF algorithm to extract the key points and descriptors of the image work; perform deep feature extraction based on the key points and descriptors to obtain the deep features of the image work; use the FLANN and Brute-Force algorithms, combined with the archived image works in the work database, to perform feature matching similarity comparison, and perform the duplicate submission detection.

[0025] In a possible implementation manner, the to-be-determined email work is subjected to theme conformity detection, including:

[0026] Extract keywords and phrases from the to-be-determined email work;

[0027] Use the pre-trained LDA topic model to determine the current topic distribution;

[0028] Perform syntactic analysis on the keywords and phrases to obtain the deep features of the to-be-determined email work;

[0029] Perform similarity comparison based on the current topic distribution and the deep features to achieve the theme conformity detection.

[0030] In a possible implementation manner, the multi-dimensional scoring of the to-be-determined email work to obtain the multi-dimensional scoring result of the to-be-determined email work includes:

[0031] When the to-be-determined email work is a writing work, perform multi-dimensional scoring based on the BERT model with part-of-speech tagging, syntactic analysis, and sentiment analysis as scoring indicators to obtain the multi-dimensional scoring result of the writing work;

[0032] When the to-be-determined email work is an image work or a video work, perform multi-dimensional scoring based on the LSTM and Transformer models with painting style, painting technique, and artistic elements as scoring indicators to obtain the multi-dimensional scoring result of the image work or the video work;

[0033] When the to-be-determined email work is a calligraphy work, perform multi-dimensional scoring according to the area difference, aspect ratio difference, distance difference, and orientation difference between the calligraphy work and the standard character to obtain the multi-dimensional scoring result of the calligraphy work.

[0034] In a second aspect, an embodiment of the present invention provides a work screening device based on an email box, including:

[0035] An acquisition module, configured to acquire original emails from a target email server and perform filtering to obtain pending email works;

[0036] A screening module, configured to perform work compliance detection based on the work type of the pending email works; in the case of passing the work compliance detection, perform multi-dimensional scoring on the pending email works to obtain a multi-dimensional scoring result of the pending email works; in response to an indication that the user review of the multi-dimensional scoring result passes, obtain a final screening result of the pending email works according to the magnitude of the multi-dimensional scoring result.

[0037] In a third aspect, an embodiment of the present invention provides a computer device, where the computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method described in the first aspect.

[0038] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, where the readable storage medium includes a computer program, and when the computer program runs, it controls a computer device where the readable storage medium is located to execute the method described in the first aspect.

[0039] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a work screening method, device, computer device and readable storage medium based on an email box disclosed in the present invention, obtaining original emails from a target email server and filtering to obtain pending email works, then performing compliance detection based on the work type, performing multi-dimensional scoring after passing the compliance to obtain a scoring result, and finally determining the final screening result according to the magnitude of the result after the user reviews and passes the scoring result, thereby realizing efficient and accurate work screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is a schematic flowchart of the steps of the work screening method based on an email box provided by an embodiment of the present invention;

[0042] Figure 2 It is a schematic block diagram of the structure of the work screening device based on an email box provided by an embodiment of the present invention;

[0043] Figure 3 It is a structural schematic block diagram of the computer device provided by the embodiment of the present invention. Specific embodiments

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0045] The following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.

[0046] To solve the technical problems in the foregoing background art, Figure 1 It is a flowchart of the work screening method based on email provided by the embodiment of the present disclosure. The work screening method based on email will be introduced in detail below.

[0047] Step S201: Obtain the original emails from the target email server and filter them to obtain the pending email works.

[0048] Step S202: Perform work compliance detection based on the work type of the pending email works.

[0049] Step S203: When the work compliance detection is passed, perform multi-dimensional scoring on the pending email works to obtain the multi-dimensional scoring results of the pending email works.

[0050] Step S204: In response to the user's review pass indication for the multi-dimensional scoring results, obtain the final screening results of the pending email works according to the size of the multi-dimensional scoring results.

[0051] In the embodiment of the present invention, by way of example, it is assumed that the server needs to screen works for a specific submission email box, which is set up by an art competition organization. This email box uses the IMAP (Internet Message Access Protocol) protocol. The server knows in advance that the IMAP server address corresponding to this submission email box is "imap.example.com" and the port number is 143 (this is a common non-encrypted IMAP port number. If it is an encrypted connection, the port number may be 993). This address and port number are like the gate address and entrance number to the submission email box. The server can find the location of the target email server and attempt to establish a connection through this information.

[0052] After the server obtains the IMAP server address and port number, it also needs a login credential to establish a connection. Suppose the login account for this art competition submission email is "contest@example.com" and the password is "Contest1234". The server uses these login credentials and sends a connection request to the target email server according to the provisions of the IMAP protocol. It's like two people verifying each other's identities between the server and the target email server. The target email server will check whether the received account and password are correct. If correct, it will allow the server to establish a connection, just like opening the door to the submission email box, and the server can enter to view the email content inside.

[0053] The art competition stipulates that the submission period is from January 1, 2023, to March 31, 2023. After successfully connecting to the target email server, the server will filter the emails to be downloaded according to this preset submission time range. The server will traverse each email in the mailbox and check the received time of the email. For example, when it encounters an email with a received time of February 15, 2023, which is within the preset submission time range, the server will download the original email body and the original email attachment of this email. The original email body may contain some text information such as the author's description of the work and creative ideas, while the original email attachment may be the actual entry work, such as an image file of a painting or a document file of a literary work.

[0054] Suppose the preset filtering rules are: the email body cannot contain malicious advertising links, and the attachment size cannot exceed 10MB. After downloading the original email, the server starts to filter it. For the email body, the server will use a special link detection algorithm to check whether there are malicious advertising links. If a suspicious link is found in the body of a certain email and it is found to be a malicious advertising link by comparing it with the malicious link database, then this email does not meet the filtering rules. For the attachment, the server will check its file size. If the size of an attachment is 12MB, exceeding the 10MB limit, this email also does not meet the filtering rules. Only when the email body has no malicious advertising links and the attachment size is within 10MB, is this email considered to meet the filtering rules, and at this time, the original email attachment is regarded as a pending email work. For example, there is an email with only normal work descriptions in the body, no malicious links, and the attachment is an 8MB image file of a painting, then this image file of the painting will be regarded as a pending email work.

[0055] After the server obtains the pending email work, it first needs to determine its work type. Assuming the pending email work is an image file, the detection criteria for image works may include: not containing sensitive scenes such as porn, violence, horror, etc.; not being highly similar to existing participating works (duplicate submission detection); the work content should conform to the competition theme, etc. If the pending email work is a text file, the detection criteria may include: not containing sensitive words such as abusive or discriminatory words; not being a plagiarized existing work (duplicate submission detection); the article theme should be relevant to the competition theme, etc.

[0056] Suppose this text work is a short story. The server first preprocesses this short story, such as converting all letters to lowercase and removing punctuation marks. Then it performs sentiment analysis using a pre-trained sentiment analysis model, such as the TextBlob sentiment analysis model. If the sentiment analysis result shows that the overall sentiment of this short story is positive. Next, it combines a preset sensitive word library and a preset sensitive word detection model to detect sensitive information. The preset sensitive word library contains words such as "discriminatory words" and "abusive words". Suppose there is a sentence in the short story "He is a fool". Although this is not a very serious abusive word, within the scope of civilized language required by the competition, this may be regarded as a non-compliant expression. If such words are found after detection, it is determined that this text work contains sensitive information.

[0057] Suppose this image work is a landscape painting. The server segments this landscape painting using an image segmentation algorithm, such as the deep learning-based U-Net algorithm, to segment the image into different regions such as the sky, mountains, and rivers. Then it combines a preset scene recognition model to detect sensitive information. This preset scene recognition model is trained with a large amount of data and can identify porn, violence and other scenes. If these bad scenes are not recognized in the image, but the image contains some text content, such as there is a sign with some text in the picture. The server will extract the text from the image work using OCR (Optical Character Recognition) technology, such as Tesseract-OCR. After extracting the text, preprocess these text contents, such as removing special characters, and then perform sentiment analysis and combine with the sensitive word library and sensitive word detection model for detection, just like the detection of text works.

[0058] Suppose this audio work is an audio of a poem recitation. The server uses ASR (Automatic Speech Recognition) technology, such as Baidu's speech recognition API, to convert the video sound into audio text. Suppose the converted audio text is "In that faraway place, there are beautiful landscapes...". Then preprocess this audio text, such as removing some meaningless modal particles and other operations, and then conduct sentiment analysis using appropriate sentiment analysis algorithms. At the same time, combine the preset sensitive word library and the preset sensitive word detection model to detect sensitive information. If no sensitive information is found in the audio text, it means that this audio work has passed the sensitive information detection link.

[0059] If the pending email work is a video work

[0060] Suppose this video work is a creative short video. The server first uses ASR to convert the video sound into audio text, then preprocesses and conducts sentiment analysis on the audio text, combines the preset sensitive word library and the preset sensitive word detection model to detect sensitive information, and the operation process is similar to that of the audio work. At the same time, extract multiple video images from this video work at preset time intervals (such as every 5 seconds). For example, for a 60-second video, extracting at 5-second intervals will result in 12 video images. Segment these video images using an image segmentation algorithm, such as the U-Net algorithm mentioned earlier, and then combine the preset scene recognition model to detect sensitive information. If there is text content in the video image, such as subtitles in the video, extract the text from the video image using OCR technology, and then preprocess, conduct sentiment analysis on the text content, and combine the sensitive word library and the sensitive word detection model to detect.

[0061] Suppose this text work is a scientific and technological paper. The server performs word segmentation on this paper using tools such as Jieba Segmentation to obtain multiple lexical units, such as "artificial intelligence", "algorithm optimization", etc. Then remove the stop words in these lexical units, such as "de", "shi", etc., and convert the remaining lexical units into multiple text vectors. Models such as Word2Vec can be used to convert words into vectors. Then compare the text similarity and semantic similarity of these text vectors with the archived text vectors in the work database. Suppose there is an archived paper in the work database. After calculation, the text similarity of these two papers reaches 80%, and the semantic similarity is also very high, then it may be determined that this text work has the suspicion of duplicate submission.

[0062] If the pending email work is an image work

[0063] Suppose the image work is a portrait painting. The server uses the SURF (Speeded Up Robust Features) algorithm to extract the key points and descriptors of this portrait painting, which can characterize the features of the image. Then, based on these key points and descriptors, deep feature extraction is performed to obtain the deep features of the image work. Next, using the FLANN (Fast Library for Approximate Nearest Neighbors) and Brute-Force algorithms, in combination with the archived image works in the work database, a comparison of the feature matching similarity is carried out. If a work is found in the archived image works whose feature matching similarity with the current portrait painting reaches more than 90%, it may be determined that there is a situation of duplicate submission of this portrait painting.

[0064] Suppose the work to be determined in the email is an essay on the theme of environmental protection. First, the server extracts keywords and phrases from this essay. Using natural language processing techniques, keywords and phrases such as "environmental protection" and "energy conservation and emission reduction" may be extracted. Then, the pre-trained LDA (Latent Dirichlet Allocation) topic model is used to determine the current topic distribution. This topic model can analyze the distribution of these keywords and phrases in the entire article to determine the topic structure of the article. Next, grammatical analysis is performed on these keywords and phrases to obtain the deep features of this essay, such as sentence structure, part of speech, and other information. Finally, based on the current topic distribution and deep features, a similarity comparison is made with the set environmental protection theme of the competition. If it is found that the topic structure, keywords, and phrases of the article are highly relevant to the environmental protection theme, it is determined that the theme of this work is in line.

[0065] Suppose the work to be determined in the email is a poem. Based on the BERT (Bidirectional Encoder Representations from Transformers) model, part-of-speech tagging, syntactic analysis, and sentiment analysis are used as scoring metrics for multi-dimensional scoring.

[0066] In terms of part-of-speech tagging: If various parts of speech, such as nouns, verbs, adjectives, etc., are correctly used in the poem and are used creatively and accurately, for example, some uncommon but appropriate adjectives are used to describe the scenery, the score for this part will be relatively high.

[0067] In terms of syntactic analysis: If the sentence structure of the poem is reasonable and the rhythm is harmonious, for example, the antithesis between verses is neat and the combination of long and short sentences is reasonable, a better score will be obtained in this scoring metric of syntactic analysis.

[0068] In terms of sentiment analysis: If the sentiment expressed in the poem is sincere and profound and can resonate with readers, such as expressing a strong yearning for love or a deep longing for hometown, a higher score will be obtained in this sentiment analysis metric. Combining the scores of these three aspects, the multi-dimensional scoring result of this poem is obtained.

[0069] Suppose the work of the email to be determined is an oil painting. Based on the LSTM (Long Short-Term Memory) and Transformer models, the painting style, painting techniques, and artistic elements are used as scoring indicators for multi-dimensional scoring.

[0070] In terms of painting style: If this oil painting is in the Impressionist style and well reflects the blurry and colorful characteristics of Impressionism, such as bold use of colors and loose brushstrokes, it will receive a higher score in the painting style indicator.

[0071] In terms of painting techniques: If the color transitions in the oil painting are natural and the depiction of people or scenery is delicate and vivid, for example, the facial expressions of people are lifelike, it indicates proficient painting techniques and will receive a good score in this indicator.

[0072] In terms of artistic elements: If the oil painting contains rich artistic elements, such as unique composition, ingenious use of light and shadow, etc., it will receive a higher score in the artistic elements indicator. By synthesizing the scores of these three aspects, the multi-dimensional scoring result of this oil painting is obtained.

[0073] Suppose the work of the email to be determined is a running script calligraphy work. Multi-dimensional scoring is carried out according to the area difference, width-to-height ratio difference, distance difference, and orientation difference between the calligraphy work and the standard characters.

[0074] In terms of area difference: If the size of the characters in the calligraphy work is not much different from that of the standard characters, for example, the area of the standard characters is 10 square centimeters and the area of the characters in the calligraphy work is between 8 - 12 square centimeters, it will receive a good score in the area difference indicator.

[0075] In terms of width-to-height ratio difference: If the width-to-height ratio of the characters in the calligraphy work is close to that of the standard characters, such as the width-to-height ratio of the standard characters is 1:1.2 and the width-to-height ratio of the characters in the calligraphy work is between 1:1 - 1:1.4, it will receive a good score in this indicator.

[0076] In terms of distance difference: If the spacing between characters and the spacing between lines in the calligraphy work are similar to the standard calligraphy layout spacing, it will receive a good score in the distance difference indicator.

[0077] In terms of orientation difference: If there is no obvious skew in the writing orientation of the characters in the calligraphy work, it will receive a good score in the orientation difference indicator. By synthesizing the scores of these four aspects, the multi-dimensional scoring result of this running script calligraphy work is obtained.

[0078] Suppose the server has obtained the multi-dimensional scoring results of multiple submitted works (pending email works). For example, the multi-dimensional scoring results of 10 prose works have been calculated. These results are displayed on a management interface, and the administrator (user) reviews these scoring results. The administrator will carefully check the basis for scoring each work. For example, for a prose work, the specific scores in parts-of-speech tagging, syntactic analysis, and sentiment analysis are checked. If the administrator has no objections to these scoring results, a user review passed indication will be given.

[0079] After the server receives this indication, it sorts these prose works according to the size of the multi-dimensional scoring results. Suppose one of the prose works has the highest comprehensive score in the three indicators of parts-of-speech tagging, syntactic analysis, and sentiment analysis and ranks first. Then this prose work will become the first work in the final screening result, and the other works are arranged in order of score. In this way, the final screening result of the pending email works is obtained, and this result can be used to determine the winning works, excellent works, etc. of the competition.

[0080] In the embodiment of the present invention, the obtaining of the original emails from the target mailbox server and filtering to obtain the pending email works can be implemented through the following examples.

[0081] Obtain the IMAP server address and port number of the target mailbox server;

[0082] Establish a connection with the target mailbox server according to the IMAP server address and the port number in combination with the login credentials;

[0083] Download the original emails in the target mailbox server according to the preset submission time range, where the original emails include the original email body and the original email attachments;

[0084] Filter the original email body and the original email attachments according to the preset filtering rules. When the filtering result indicates compliance with the filtering rules, use the original email attachment as the pending email work.

[0085] In the embodiment of the present invention, for example, in a large-scale art work solicitation activity scenario, the server is responsible for screening the works in the submission mailbox. This submission mailbox is set up by the activity organizer on a specific mail server and uses the IMAP protocol so that the server can obtain the email content.

[0086] Assume that this art work solicitation activity is global. To facilitate submissions from contestants in different regions, the organizer has selected a well-known email service provider. This email service provider has assigned a specific IMAP server address "imap.artcontestmail.com" to this submission email box, with a port number of 143 (a commonly used port number for non-encrypted connections). At the initial stage of the entire work screening process, the server needs to accurately obtain this address and port number information. This information is like a part of a key and is an important basis for establishing a connection with the target email server.

[0087] After the server obtains the IMAP server address "imap.artcontestmail.com" and the port number 143, it also needs login credentials to establish a connection. The activity organizer has created a dedicated login account "contestserver@artcontestmail.com" and the corresponding password "ServerContestPass123" for the server.

[0088] The server starts the operation of establishing a connection in accordance with the specifications of the IMAP protocol. It first sends a connection request to the target email server, and this request contains information such as the IMAP server address, port number, login account, and password. After receiving this request, the target email server will verify the account and password. This process is similar to a security access control system checking the identity of a visitor. If the account and password match correctly, the target email server will allow the server to establish a connection. At this time, a data transmission channel is established between the server and the target email server, just like opening the door to the submission email box, preparing for obtaining the email content subsequently.

[0089] In this art work solicitation activity, the organizer has stipulated that the submission time range is from 00:00:00 on May 1, 2023 to 23:59:59 on August 31, 2023. After the server successfully establishes a connection with the target email server, it starts to download the original emails according to this preset submission time range.

[0090] The server traverses each email in the target mailbox. For each email, it first checks the receiving time of the email. For example, there is an email with a receiving timestamp of 2023-06-15-10:30:00, and this time is within the preset submission time range, so this email meets the download conditions. The original email contains two important parts, namely the original email body and the original email attachment. The original email body may contain textual information such as the contestant's elaboration on the creative concept of the work and the work introduction. For example, in the original email body of a submission, the contestant writes: "The inspiration for my painting comes from the tranquility and harmony of nature, and I hope to convey this beauty through colors and lines." The original email attachment is the actual entry work, which may be various types of files, such as image files (.jpg,.png, etc.), document files (.docx,.pdf, etc.). If it is a submission of a painting work, the attachment may be a beautiful painting stored in.jpg format; if it is a submission of a literary work, the attachment may be a.docx document containing the content of a novel or poem.

[0091] In this art work solicitation activity, a series of strict filtering rules are preset.

[0092] Filtering of the original email body:

[0093] Rule 1: Prohibit the inclusion of malicious links:

[0094] The server scans the body of each email that meets the submission time range. It uses a special network link analysis tool to detect whether there are malicious links. This tool parses each link in the email body and then compares it with the global malicious link database. For example, in an email body, it is mentioned that "Click here to view more work examples: http: / / xxxxx.xxxxx.com". The server discovers through comparison that this link has a record in the malicious link database, and this email does not meet the requirements.

[0095] Rule 2: Restrict the use of specific keywords:

[0096] The event organizer stipulates that the email body cannot contain some specific keywords related to inappropriate content, such as discriminatory words, words related to prohibited drugs, etc. The server performs a full-text search on the email body. Suppose a discriminatory word "inferior race" appears in an email body. After the server detects this keyword, it determines that this email does not meet the filtering rules.

[0097] Filtering of the original email attachment:

[0098] Rule 1: Restrict the attachment type and size:

[0099] The activity regulations only accept specific types of attachments. For example, for painting works, only image files in.jpg and.png formats are accepted, and for literary works, only document files in.docx and.pdf formats are accepted. The server will check the file format of the attachments. If the format of one of the attachments is.exe (executable file format), this obviously does not meet the requirements. At the same time, there are also restrictions on the attachment size. For example, the maximum size cannot exceed 20MB. If there is an attachment that is a 30MB image file, the server will determine that this email does not meet the filtering rules.

[0100] Rule Two: Preliminary Detection of Attachment Content (for Specific Types of Attachments):

[0101] For image attachments, if it is a participating painting work, the server will conduct a preliminary image integrity check. It will check whether the image is damaged and whether it can be opened and displayed normally. If an image attachment is damaged due to problems during the transmission process and the image content cannot be displayed normally, then this email also does not meet the requirements.

[0102] Only when both the original email body and the original email attachments pass all the preset filtering rules will the server regard the original email attachments as pending email works. For example, there is an email whose body has no malicious links and prohibited keywords, and the attachment is a 15MB.jpg format image file of a painting work and the image can be displayed normally. Then this painting work image file will be regarded as a pending email work by the server and wait for subsequent further work compliance detection and other operations.

[0103] In the embodiment of the present invention, the work compliance detection based on the work type of the pending email work can be implemented through the following examples.

[0104] Determine the detection criteria for the pending email work based on the work type of the pending email work;

[0105] Conduct sensitive information detection, duplicate submission detection, and theme compliance detection on the pending email work based on the detection criteria.

[0106] In the embodiment of the present invention, exemplarily, determine the detection criteria for the pending email work based on the work type of the pending email work:

[0107] Painting works:

[0108] In a painting work solicitation activity, the server receives multiple painting image attachments as pending email works. For painting works, the detection criteria mainly focus on the following aspects:

[0109] In terms of content:

[0110] The picture must not contain any inappropriate content such as violence, porn, horror, etc. For example, there should be no depiction of bloody violence scenes, and no figure images or picture elements with pornographic implications, etc.

[0111] The content of the painting should conform to public order and good customs, and there should be no content that violates social morality and ethics, such as malicious slander or spoof images of specific groups.

[0112] Technically:

[0113] The painting should have a certain degree of integrity. For example, if it is an oil painting, there should be no large areas of unfinished areas or chaotic paint smears that make the picture unrecognizable.

[0114] The use of colors should be reasonable. For example, there should be no serious disharmony in color combinations that affect the overall visual effect, unless it is a deliberate artistic expression style and can be reasonably judged as an artistic innovation.

[0115] Literary works:

[0116] In the scenario of soliciting literary works, such as poems, novels, essays, etc., after being received by the server as works in the pending email, the detection criteria are as follows:

[0117] In terms of content:

[0118] The written content must be positive and healthy, and cannot contain discriminatory language, whether based on race, gender, religion or other factors. For example, there should be no derogatory descriptions of a certain race.

[0119] It cannot contain content that promotes violence, terrorism or bad values, such as glorifying criminal acts, etc.

[0120] In terms of creation norms:

[0121] If it is a poem, it should follow certain rhyme rules or have a unique poem structure and expression form. For example, classical poems should meet the metrical requirements (if it is a solicitation of ancient-style poems), and modern poems should also have a unique rhythm and sense of rhyme.

[0122] For novels and essays, they should have a clear logical structure, distinct character images, and reasonable and coherent plot developments.

[0123] Music works:

[0124] When a music work (such as a work in the pending email in the form of an audio file) is processed by the server, the detection criteria include:

[0125] In terms of content:

[0126] The lyrics (if any) cannot contain bad information, such as discriminatory, violent or pornographically suggestive words.

[0127] Music melodies should not plagiarize existing well-known works and should have a certain degree of originality.

[0128] Technical aspects:

[0129] The audio quality should meet certain standards and there should be no serious noise, distortion, or chaotic audio clips.

[0130] For works performed by multiple instruments, the cooperation between the instruments should be harmonious and the rhythm should be stable.

[0131] Based on the detection standards, perform sensitive information detection, duplicate submission detection, and topic consistency detection on the pending email works:

[0132] Sensitive information detection:

[0133] Painting works:

[0134] The server will use image recognition technology for sensitive information detection. For painting works, the image will first be converted into a digital matrix form for processing.

[0135] When detecting whether the picture contains pornographic content, the server uses an image classification model based on deep learning, which has been trained with a large amount of image data containing and not containing pornographic content. For example, after the server inputs the image of the painting work into the model, the model will analyze the human figures, poses, and picture elements in the image. If there are overly exposed or suggestive human poses in the picture, the model will recognize them and determine that it contains sensitive information.

[0136] In detecting horror and violent content, the server will identify the scene elements in the picture, such as whether there are bloody scenes, terrifying creature images, or improper depictions of weapons. For example, if there is an image of a person holding a sharp knife and covered in blood, the server may determine that this painting work contains sensitive information.

[0137] Literary works:

[0138] For literary works, the server will detect them word by word.

[0139] In detecting discriminatory language, the server will use a predefined discriminatory vocabulary library. When detecting content that promotes violence and terrorism, the server will analyze the semantics of the sentences. If there is a passage of text that describes a detailed criminal process with a glorifying or encouraging attitude, such as "He brutally killed that person with a knife, and this is really a cool behavior", the server will determine that this literary work contains sensitive information.

[0140] Music works:

[0141] When a music work has lyrics, the server processes the lyrics.

[0142] The server will perform word segmentation on the lyrics and then compare them with a predefined sensitive vocabulary library. For example, if the word "drug" appears in the lyrics and is in the context of promoting drug use, such as "Drugs make me feel really high", the server will determine that sensitive information exists.

[0143] If the music work has no lyrics, the server will analyze the melody style of the music. Although it is difficult for the music melody itself to directly contain sensitive information, if the melody style is highly similar to certain specific bad music (such as music promoting violence or horror), the server will also mark it for further manual review.

[0144] Duplicate submission detection:

[0145] Painting works:

[0146] The server will extract the feature vectors of the painting works. First, an image feature extraction algorithm such as the SIFT (Scale-Invariant Feature Transform) algorithm is used.

[0147] The server processes the painting works and extracts features such as color distribution features, line direction features, and object shape features. For example, for a painting work mainly in blue tones with many curved lines depicting waves, these features will be extracted and converted into feature vectors.

[0148] Then, the server compares this feature vector with the feature vectors of the painting works stored in the work database. The similarity is judged by calculating the distance between the vectors (such as the Euclidean distance). If this distance is less than a preset threshold (for example, after testing, a value of 0.3 indicates a high similarity), it is determined that the work may be a duplicate submission.

[0149] Literary works:

[0150] The server performs text processing on the literary works. First, lexical analysis is carried out to break down the article into words and phrases.

[0151] For example, for an essay, the server breaks it down into words and phrases such as "morning", "sunshine", "woods", etc. Then, some common stop words such as "de", "shi", "zai" (the, is, in in Chinese) are removed.

[0152] Next, the server uses text fingerprint technology to generate a unique fingerprint (a digital representation form) for the processed text. This fingerprint is compared with the text fingerprints of the archived literary works. If the two fingerprints are highly similar, for example, the similarity reaches more than 90%, it is determined that there may be a case of duplicate submission.

[0153] Music works:

[0154] For music works, the server extracts the features of the audio. First, the audio is converted into a spectrogram, and then features such as rhythm features, pitch features, and timbre features are extracted from the spectrogram.

[0155] For example, if a pop music work has a rhythm of 4 / 4 time, the main pitch is around C major, and has a bright timbre (obtained through spectral analysis), these features will be combined into a feature vector.

[0156] The server compares this feature vector with the feature vectors of existing music works. If the similarity between the two feature vectors exceeds a set threshold (such as 0.8), it is determined that the work may be a duplicate submission.

[0157] Detection of theme compliance:

[0158] Painting works:

[0159] The server first conducts content analysis on the painting works. For example, in a painting solicitation activity with the theme of "the beauty of nature".

[0160] The server will identify the elements in the picture and determine whether they are related to nature. If the picture mainly depicts urban architecture and has only a few natural elements (such as a small patch of sky), then the compliance with the theme of "the beauty of nature" is relatively low.

[0161] The server will also analyze whether the painting style can reflect the theme. If the theme is "peaceful nature" and the painting style uses strong contrasting colors and chaotic lines, showing a restless feeling, it may not meet the theme requirements.

[0162] Literary works:

[0163] The server conducts keyword and semantic analysis on literary works. In a essay solicitation activity with the theme of "friendship".

[0164] The server will extract keywords in the work, such as "friend", "company", "trust", etc. Then it analyzes the distribution and weight of these keywords in the article. If an article rarely mentions keywords related to friendship but more describes the individual's lonely experiences, then the theme compliance is poor.

[0165] The server will also analyze the semantics of the article. If the article mentions the word "friend", but the overall semantics is about betrayal and hatred between friends, this also does not meet the theme requirements of "friendship".

[0166] Music works:

[0167] In a music work solicitation activity with the theme of "happy festival", the server analyzes the music.

[0168] From the perspective of the rhythm of music, if the rhythm of the music is slow and heavy, such as a slow rhythm of only 40 beats per minute, while the rhythm of the usually happy festival music is relatively brisk (such as about 120 beats per minute), then it does not meet the theme requirements in terms of rhythm.

[0169] Analyzing from the aspect of melody, if the melody is low and sad instead of lively and bright, it also does not meet the theme requirements of "happy festival". If there are lyrics in the music, the server will also analyze whether the content of the lyrics is related to the happy festival, such as whether the lyrics mention festival elements, a happy atmosphere, etc.

[0170] In the embodiment of the present invention, for the to-be-determined email work, sensitive information detection can be implemented through the following examples.

[0171] When the to-be-determined email work is a text work, preprocess the text work and perform sentiment analysis, and combine a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection;

[0172] When the to-be-determined email work is an image work, segment the image work, and combine a preset scene recognition model to perform sensitive information detection; when the image work includes text content, extract the text from the image work, preprocess the text content and perform sentiment analysis, and combine a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection;

[0173] When the to-be-determined email work is an audio work, use ASR to convert the video sound into audio text, preprocess the audio text and perform sentiment analysis, and combine a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection;

[0174] When the to-be-determined email work is a video work, use ASR to convert the video sound into audio text, preprocess the audio text and perform sentiment analysis, and combine a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection; extract multiple video images from the video work at preset time intervals, segment the video images, and combine a preset scene recognition model to perform sensitive information detection; when the video images include text content, extract the text from the video images, preprocess the text content and perform sentiment analysis, and combine a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection.

[0175] In the embodiment of the present invention, exemplarily, when the to-be-determined email work is a text work:

[0176] In a literary work submission activity, the server receives numerous text works as pending email works.

[0177] Preprocessing:

[0178] The server first performs preprocessing operations on the text works. This includes uniformly converting the text to a standard format, such as converting all uppercase letters to lowercase letters, for the consistency of subsequent processing. For example, for an article that originally contains sentences starting with uppercase letters and a mixture of some uppercase proper nouns, after conversion, all letters become lowercase.

[0179] Then, the server removes the punctuation marks in the text because punctuation marks may interfere with subsequent analysis in some cases. For example, "Hello, world!" becomes "hello world". At the same time, for some special characters, such as tab characters, line break characters, etc., appropriate processing will also be carried out to organize the text into a continuous string form.

[0180] Next, the server also performs lexical analysis on the text, splitting the text into individual words, which is the basis for subsequent analysis. For example, "I love reading books" will be split into words such as "I", "love", "reading", "books", etc.

[0181] Sentiment analysis:

[0182] The server uses a pre-trained sentiment analysis model for sentiment analysis. This model is trained based on a large amount of text data with sentiment annotations (such as positive, negative, or neutral). For example, for an article describing natural scenery, which contains words such as "beautiful scenery", "fresh air", "pleasant sunshine", etc., the model judges the overall sentiment of this article as positive through the analysis of these words and the overall structure of the sentence.

[0183] If it is an article describing the cruelty of war, which contains words such as "bloody battlefield", "sorrow of people", etc., the model will determine it as a negative sentiment. Although this sentiment analysis cannot directly detect sensitive information, it can be used as auxiliary information because some texts with extremely negative or bad sentiment may be more likely to contain sensitive information.

[0184] Combined with a preset sensitive word library and a preset sensitive word detection model for sensitive information detection:

[0185] The server has a large preset sensitive word library, which contains various types of sensitive words, such as discriminatory words (such as racial discrimination words, gender discrimination words, etc.), abusive words, words related to porn and violence, etc.

[0186] At the same time, the server also uses a preset sensitive word detection model, which can identify content that is essentially sensitive information even though it has been distorted or euphemistically expressed. For example, some cases where homophonic words, abbreviations, or special symbols are used to replace sensitive words. For an article, the server will compare each word with the sensitive word library and conduct in-depth detection through the sensitive word detection model.

[0187] For example, in an article, there is the word "f*ck" (a veiled abusive word with some letters replaced by asterisks). Although this word is not in its complete form, it can be identified through the sensitive word detection model. If the article contains multiple such sensitive words or suspected sensitive words, the server will determine that this text work contains sensitive information.

[0188] In the case where the pending email work is an image work:

[0189] Segment the image work:

[0190] In a scenario where a painting work is submitted, after the server receives the image work, it starts the image segmentation operation. The server uses a deep learning-based image segmentation algorithm, such as the Mask R-CNN algorithm.

[0191] Suppose the received is a landscape painting image work, which contains elements such as sky, mountains, rivers, and trees. Through the image segmentation algorithm, the server can accurately segment this image into different regions, with each region corresponding to an element in the image. For example, the sky part is segmented into one region, and the mountain part is segmented into another region, etc. Such segmentation helps in subsequent detailed analysis of each region to detect whether there is sensitive information.

[0192] Combine with a preset scene recognition model for sensitive information detection:

[0193] The server has a preset scene recognition model, which is trained with a large number of annotated scene images (such as normal scenes, violent scenes, pornographic scenes, etc.). For the segmented image regions, the model will conduct recognition.

[0194] For example, if in the segmented image region, there is a scene part where human organs are exposed and the pose implies pornographic content, the model will recognize this as pornographic scene content, thus determining that this image work contains sensitive information. Or if there is a region that looks like a bloody battle scene, containing many violent elements such as weapons and injured people, the model will also determine that there is sensitive information.

[0195] In the case where the image work includes text content:

[0196] If in this landscape image, there is a sign with some text, such as "No Trespassing" or something like that.

[0197] The server will first perform text extraction on the image. It will use OCR (Optical Character Recognition) technology, such as Tesseract-OCR tool. This tool can recognize the text in the image and convert it into a text format that can be processed by the computer.

[0198] Then, the extracted text content is preprocessed, and the preprocessing process is similar to that of the text work, including operations such as unifying the uppercase and lowercase letters and removing punctuation marks. For example, "NO ENTRY" is converted to "no entry".

[0199] Then, we conduct sentiment analysis to determine the sentiment of these words. If the text content is relatively neutral, such as "no entry", the sentiment is neutral. But if it is "get out you idiot" (including the insulting word "idiot"), the sentiment is negative and contains sensitive information. Finally, we combine the preset sensitive word library and the preset sensitive word detection model to perform detection, just like treating pure text works. If sensitive words are present, it is determined that this image work contains sensitive information.

[0200] In the case where the pending email work is an audio work:

[0201] In a music work submission activity, there may be some audio works with singing content or narration.

[0202] Use ASR to convert video sound into audio text:

[0203] The server uses ASR (automatic speech recognition) technology, such as Baidu's speech recognition service. After receiving the audio work, the sound in the audio is converted into audio text. For example, the lyrics of the singing part of a piece of music are "We run on sunny days", which are converted into corresponding text content through ASR technology.

[0204] Preprocessing and sentiment analysis of audio text:

[0205] Preprocess the converted audio text. Similar to the preprocessing of the text work, convert all letters to lowercase letters, remove punctuation marks, etc. For example, change "We run on sunny days." to "We run on sunny days."

[0206] Then, perform sentiment analysis using a pre-trained sentiment analysis model. If the lyrics are positive, such as describing a wonderful life, happy times, etc., the sentiment is positive. If it describes sadness, pain, or contains negative emotions, the sentiment is negative.

[0207] Perform sensitive information detection by combining a preset sensitive word library and a preset sensitive word detection model

[0208] The server compares the preprocessed audio text with the preset sensitive word library and simultaneously uses the sensitive word detection model for detection. For example, if the audio text contains a swear word like "damn", through comparison with the sensitive word library and model detection, it will be determined that this audio work contains sensitive information.

[0209] In the case where the pending email work is a video work:

[0210] Use ASR to convert the video sound into audio text, preprocess the audio text and perform sentiment analysis, and combine a preset sensitive word library and a preset sensitive word detection model to perform sensitive information detection

[0211] In a video work submission activity, the server receives a video work. First, use ASR technology to convert the sound in the video into audio text. For example, in a video work, there is a narrator saying "This is a place full of hope", which is converted into corresponding text content through ASR.

[0212] Then preprocess this audio text, such as converting uppercase letters to lowercase letters, removing punctuation marks, etc., to become "This is a place full of hope". Then perform sentiment analysis and judge the sentiment according to the text content. If it is a positive expression, the sentiment is positive.

[0213] Finally, combine a preset sensitive word library and a preset sensitive word detection model for detection. If sensitive words, such as "fool" (a swear word), appear in the audio text, it is determined that this video work contains sensitive information.

[0214] Extract multiple video images from the video work at preset time intervals, segment the video images, and combine a preset scene recognition model to perform sensitive information detection:

[0215] The server extracts one video image from the video work at preset time intervals, for example, every 5 seconds. For a 60 - second video work, 12 video images will be extracted.

[0216] Perform a segmentation operation on each extracted video image using an image segmentation algorithm such as the Mask R-CNN algorithm mentioned above. Assume that there are elements such as people and background buildings in the video image, and use the segmentation algorithm to divide it into different regions.

[0217] Then, combine it with a preset scene recognition model for detection. If a violent scene such as fighting or weapon display, or a pornographic scene such as inappropriate body exposure is found in the segmented area of a certain video image, it is determined that this video work contains sensitive information.

[0218] In the case where the video image includes text content:

[0219] If there is text content such as subtitles or signs in the picture in a certain video image, for example, the subtitle shows "Dangerous Area".

[0220] The server first extracts the text from the video image, uses OCR technology to extract the text and convert it into a text format.

[0221] Then, preprocess the extracted text content, such as unifying the case, removing punctuation marks, etc., to become "Dangerous Area". Then perform sentiment analysis. The sentiment of the text in this example is neutral.

[0222] Finally, combine it with a preset sensitive word library and a preset sensitive word detection model for detection. If there are sensitive words, it is determined that this video work contains sensitive information.

[0223] In the embodiment of the present invention, for the to-be-determined email work, duplicate submission detection can be performed through the following examples.

[0224] In the case where the to-be-determined email work is a text work, perform word segmentation on the text work to obtain multiple lexical units; remove the stop words from the multiple lexical units, and convert the multiple lexical units into multiple text vectors; perform text similarity comparison and semantic similarity comparison based on the multiple text vectors and the archived text vectors in the work database to perform the duplicate submission detection;

[0225] In the case where the to-be-determined email work is an image work, use the SURF algorithm to extract the key points and descriptors of the image work; perform deep feature extraction based on the key points and descriptors to obtain the deep features of the image work; use the FLANN and Brute-Force algorithms, combined with the archived image works in the work database, to perform feature matching similarity comparison to perform the duplicate submission detection.

[0226] In the embodiment of the present invention, for example, in the case where the to-be-determined email work is a text work:

[0227] Perform word segmentation on the text work to obtain multiple vocabulary units:

[0228] In a scenario of a large-scale literary work submission platform, the server receives many text works as pending email works. For each text work, the server will perform word segmentation. For example, for a prose work "In the early morning, the sun shines through the gaps in the leaves and sprinkles on the ground, forming patches of light. The birds sing happily on the branches, as if telling the beauty of a new day." The server uses professional word segmentation tools, such as Jieba word segmentation, to segment this article and obtain vocabulary units: "morning", "sunlight", "through", "leaves", "gap", "sprinkled on", "ground", "formed", "pieces", "light spots", "birds", "branches", "happy", "singing", "as if", "telling", "a new day", "beautiful", etc. These vocabulary units are the basis for subsequent processing.

[0229] Remove stop words from multiple vocabulary units and convert the multiple vocabulary units into multiple text vectors:

[0230] Remove stop words:

[0231] After obtaining the vocabulary units, the server will remove the stop words. Stop words are words that appear frequently in the text but do not contribute substantially to the semantic expression, such as "的", "地", "得", "是", "在", etc. After removing the stop words, the vocabulary units in the above prose may be "早上", "阳光", "叶", "裂", "光斑", "鸟儿", "枝头", "快乐", "唱", "新一天", "美好", etc.

[0232] Convert to text vector:

[0233] The server then converts these vocabulary units into text vectors. This process is usually done with the help of a pre-trained word vector model, such as the Word2Vec model. The Word2Vec model is trained based on a large amount of text data, and it can map each word into a vector space with a fixed dimension. For example, the word "morning" may be mapped to a 100-dimensional vector [0.1, 0.2, -0.3, ...], and "sunshine" may also be mapped to a 100-dimensional vector [0.3, -0.1, 0.2, ...], etc. In this way, the entire text work is converted into a series of text vectors.

[0234] Perform duplicate submission detection by comparing text similarity and semantic similarity between multiple text vectors and archived text vectors in the work database:

[0235] Text similarity comparison:

[0236] The server has a large database of works, which stores the text vectors corresponding to the text documents received and processed previously. For the text vector of the current pending email work, the server will calculate its text similarity with the archived text vectors in the database. There are various methods for calculating text similarity, such as cosine similarity. Suppose the text vector of the current pending email work after processing is vector set A, and an archived text vector in the database is vector set B. By calculating the cosine similarity formula: \(cosine\_similarity=\frac{A\cdot B}{\vert A\vert\times\vert B\vert}\), a similarity value is obtained. If this similarity value is very high, such as exceeding 0.8 (this threshold is set based on experience or platform regulations), it indicates that the two articles have a high similarity at the lexical level.

[0237] Semantic similarity comparison:

[0238] In addition to text similarity comparison, the server will also perform semantic similarity comparison. Semantic similarity comparison takes into account not only the surface matching of words but also the semantic meanings of sentences and articles. For example, the server may use a deep learning-based semantic analysis model, such as the BERT (Bidirectional Encoder Representations from Transformers) model. The current pending email work and the archived text in the database are respectively input into the BERT model to obtain their semantic representation vectors. Then, the similarity between these two semantic representation vectors is calculated, also using methods such as cosine similarity. If the semantic similarity is also very high, such as exceeding 0.7, then it is more likely to be a duplicate submission of the work. Combining the results of text similarity and semantic similarity, if both exceed the corresponding thresholds, the server determines that this text work is suspected of being a duplicate submission.

[0239] In the case where the pending email work is an image work:

[0240] Use the SURF algorithm to extract the key points and descriptors of the image work:

[0241] In a painting submission system, the server receives an image work as a pending email work. For each image work, the server processes it using the SURF (Speeded-Up Robust Features) algorithm. For example, for an oil painting depicting a landscape, the SURF algorithm will detect some representative key points in the image. These key points are usually places where the gray level changes significantly in the image, such as the junction of the sky and the mountains, the contour edges of the trees, etc. At the same time, the SURF algorithm will generate a descriptor for each key point. This descriptor is a vector that describes the pixel information around the key point, such as the gradient direction and magnitude of the pixels, etc. Through the SURF algorithm, this oil painting may obtain hundreds of key points and corresponding descriptors, and these key points and descriptors can characterize the features of this image to a certain extent.

[0242] Based on the key points and descriptors, deep feature extraction is performed to obtain the deep features of the image work:

[0243] After obtaining the key points and descriptors, the server will perform deep feature extraction. This process may involve more complex neural network models or feature combination algorithms. For example, the server can use a convolutional neural network (CNN) to further process these key points and descriptors. The CNN can automatically learn the high-level semantic relationships between the key points and descriptors, so as to extract more representative deep features. These deep features can better reflect the overall features of the image, such as the style of the image, the shape combination of objects, etc. For the previously mentioned landscape oil painting, after deep feature extraction, the obtained deep features may contain information about the layout of the scenery and color matching in the picture.

[0244] Using the FLANN and Brute-Force algorithms, combined with the archived image works in the work database, feature matching similarity comparison is carried out to perform duplicate submission detection:

[0245] FLANN (Fast Library for Approximate Nearest Neighbors) algorithm:

[0246] The server compares the deep features of the current pending email work with the deep features of the archived image works in the work database. First, the FLANN algorithm is used. The FLANN algorithm is an algorithm for quickly searching for approximate nearest neighbors in a high-dimensional space. It can quickly find the archived image features that are most similar to the current image features in a large amount of image feature data. For example, for the deep features of the current landscape oil painting, the FLANN algorithm will search in the database for the deep features of the archived image work that is closest to it.

[0247] Brute-Force Algorithm (Brute-Force Matching Algorithm):

[0248] In addition to the FLANN algorithm, the server also uses the Brute-Force algorithm to compare the similarity of feature matches. The Brute-Force algorithm is a simple and direct method that compares the depth features of the current image work one by one with the depth features of all archived image works in the database. Although this method has a large computational amount, it can obtain more accurate matching results. The similarity is measured by calculating the distance between depth features (such as Euclidean distance). For example, if the Euclidean distance between the depth features of the current landscape oil painting and the depth features of an archived landscape oil painting in the database is very small, such as less than a set threshold (based on experience or platform regulations, such as 0.1), it indicates that these two images are very similar in features.

[0249] Comprehensive Judgment:

[0250] Based on the results of the FLANN and Brute-Force algorithms, if an archived image work with a very high similarity (judged according to the set threshold) to the depth features of the current pending email work is found, the server determines that there is a possibility of duplicate submission for this image work.

[0251] In the embodiment of the present invention, for the subject conformity detection of the pending email work, the following examples can be used for implementation.

[0252] Extract keywords and phrases from the pending email work;

[0253] Use the pre-trained LDA topic model to determine the current topic distribution;

[0254] Perform syntactic analysis on the keywords and phrases to obtain the deep features of the pending email work;

[0255] Perform similarity comparison according to the current topic distribution and the deep features to achieve the subject conformity detection.

[0256] In the embodiment of the present invention, by way of example, extract keywords and phrases from the pending email work:

[0257] In the scenario of a news submission system, the server receives a large number of pending email works, which contain various types of news reports.

[0258] For text-based news manuscripts:

[0259] The server first performs word segmentation on the text using professional word segmentation tools such as Jieba (for Chinese text) or NLTK (for English text). For example, for a news article about new breakthroughs in the technology field: "Recently, scientists have made a major breakthrough in the field of artificial intelligence. They have developed a new algorithm that can significantly improve the accuracy of image recognition." After word segmentation, the following lexical units are obtained: "Recently", "scientists", "artificial intelligence", "field", "made", "a", "major", "breakthrough", "developed", "a", "new", "algorithm", "can", "significantly", "improve", "image recognition", "accuracy", etc.

[0260] Then, the server will identify the keywords and phrases among them according to certain rules. These rules may include word frequency statistics, part-of-speech tagging, etc. In this example, words such as "artificial intelligence", "algorithm", "image recognition" are recognized as keywords because they have relatively high importance and word frequency in the field of technology news; "major breakthrough" is recognized as a phrase.

[0261] For pending email works of the image type (such as images with descriptive text or labels):

[0262] If the image is accompanied by some descriptive text, for example, a picture showing a new solar panel with the descriptive text "New solar panel, efficiently converting solar energy, contributing to the development of environmental protection energy." The server also performs word segmentation on this descriptive text, obtaining lexical units such as "new", "solar panel", "efficiently", "converting", "solar energy", "environmental protection energy", "development", "contributing". Then, keywords such as "solar panel", "environmental protection energy" and phrases such as "efficiently converting" are extracted from it.

[0263] For pending email works of the audio or video type (if they contain text information such as audio transcriptions or video subtitles):

[0264] Taking a video work about history and culture as an example, the subtitles in the video contain "In the ancient Egyptian civilization, the pyramid is its most representative building, which contains rich historical and cultural connotations." After the server performs word segmentation on the subtitle text, the following lexical units are obtained: "ancient", "Egyptian civilization", "pyramid", "most", "representative", "building", "contains", "rich", "historical and cultural", "connotations", etc. Furthermore, keywords such as "Egyptian civilization", "pyramid", "historical and cultural" and phrases such as "most representative" are extracted.

[0265] Use the pre-trained LDA topic model to determine the current topic distribution:

[0266] The server has pre-trained the LDA (Latent Dirichlet Allocation) topic model using a large amount of text data. These text data cover various topic fields, such as technology, culture, entertainment, sports, etc.

[0267] For the previously extracted keywords and phrases, the server inputs them into the LDA topic model. Taking the just-mentioned technology news manuscript as an example, keywords such as "artificial intelligence", "algorithm", and "image recognition" are input into the LDA model.

[0268] The LDA model will determine the topic distribution of the current email work to be determined according to the distribution probability of these keywords under different topics. Suppose in the training data, the topic highly related to words such as "artificial intelligence", "algorithm", and "image recognition" is the "technology - computer technology" topic. The LDA model will calculate the probability of this topic in the current work. For example, it may be 0.6, indicating that this work has a 60% possibility of belonging to the "technology - computer technology" topic. At the same time, the model will also calculate the probabilities of other related topics. For example, the probability of the "technology - data processing" topic may be 0.2, and the probability of the "technology - automation" topic may be 0.1, etc.

[0269] Perform syntactic analysis on the keywords and phrases to obtain the deep features of the email work to be determined:

[0270] The server performs syntactic analysis on the previously extracted keywords and phrases. Taking the keyword "new algorithm" in the previous technology news manuscript as an example.

[0271] Part-of-speech tagging:

[0272] The server first performs part-of-speech tagging to determine that "new" is an adjective and "algorithm" is a noun. This part-of-speech tagging can help the server understand the role of the keywords in the sentence. For example, the combination of an adjective + a noun is used to describe a specific thing or concept in many cases.

[0273] Dependency analysis:

[0274] The server will also perform dependency analysis. In the sentence "They developed a new algorithm", "developed" is the predicate verb, "algorithm" is the object, and "new" is the attributive of "algorithm". This dependency relationship indicates that "new" is used to modify "algorithm", and the action of "developed" is targeted at the object "algorithm". Through this dependency analysis, the server can further understand the semantic relationship of the keywords and phrases in the sentence.

[0275] Named entity recognition (if applicable):

[0276] If the keywords contain named entities, such as "Egyptian civilization" in the video subtitles about history and culture, the server will recognize this as a named entity, which is a specific name in the cultural category. This helps to more accurately grasp the theme content of the work. Through these syntactic analysis operations, the server obtains the deep features of the pending email work, which are not only the keywords and phrases themselves, but also the syntactic relationships and semantic roles between them.

[0277] Perform a similarity comparison based on the current theme distribution and deep features to achieve theme compliance detection:

[0278] The server compares the theme distribution and deep features of the current pending email work with the predefined theme requirements. In a technology news submission platform, assume that the platform stipulates some specific theme classifications, such as "Computer Technology Innovation", "Biotechnology Progress", "New Energy Development", etc.

[0279] For previous technology news manuscripts, their theme distribution shows a high probability of belonging to the "Technology - Computer Technology" theme. At the same time, the deep features obtained through syntactic analysis indicate that "algorithm" is the core concept and is related to actions such as "research and development". If the theme requirement of the platform's "Computer Technology Innovation" is about new achievements, new methods, etc. in the field of computer technology, then the server will determine that this work is in line with the "Computer Technology Innovation" theme.

[0280] However, if a news manuscript has keywords and phrases such as "star" and "concert" after extraction, and its theme distribution shows that it belongs to the "Entertainment - Performing Arts" theme, but the platform requires technology-themed manuscripts, then the server will determine that this work does not match the theme required by the platform.

[0281] In the case of image works, assume it is an art work submission platform with the theme requirement of "Natural Landscape Painting". For an image work, the keywords extracted from its descriptive text are "high-rise buildings" and "city streets", and its theme distribution shows a correlation with the "Urban Architecture" theme. The deep features obtained through syntactic analysis also indicate that these keywords describe the urban environment. The server will determine that this image work does not match the "Natural Landscape Painting" theme.

[0282] The same is true for audio or video works. For example, in a historical and cultural documentary submission platform, the required theme is "Exploration of Ancient Civilizations". If the keywords and phrases extracted from the subtitles of a video work are "modern urban life" and "fashion trends", and its theme distribution is related to the "Modern Lifestyle" theme, and the deep features obtained through syntactic analysis also indicate that it is about modern life, then the server will determine that this video work does not match the theme required by the platform. Through such a comparison process, the server achieves the theme compliance detection of the pending email work.

[0283] In an embodiment of the present invention, for multi-dimensionally scoring the to-be-determined email work to obtain the multi-dimension scoring result of the to-be-determined email work, the following examples may be executed for implementation.

[0284] When the to-be-determined email work is a writing work, multi-dimensionally score based on the BERT model with part-of-speech tagging, syntactic analysis, and sentiment analysis as scoring metrics to obtain the multi-dimension scoring result of the writing work;

[0285] When the to-be-determined email work is an image work or a video work, multi-dimensionally score based on the LSTM and Transformer models with painting style, painting skills, and artistic elements as scoring metrics to obtain the multi-dimension scoring result of the image work or the video work;

[0286] When the to-be-determined email work is a calligraphy work, multi-dimensionally score according to the area difference, aspect ratio difference, distance difference, and orientation difference between the calligraphy work and the standard characters to obtain the multi-dimension scoring result of the calligraphy work.

[0287] In an embodiment of the present invention, by way of example, when the to-be-determined email work is a writing work, multi-dimensionally score based on the BERT model with part-of-speech tagging, syntactic analysis, and sentiment analysis as scoring metrics to obtain the multi-dimension scoring result of the writing work.

[0288] Part-of-speech tagging scoring:

[0289] In a literary work submission platform, the server receives numerous writing works as to-be-determined email works. For each writing work, the server first performs part-of-speech tagging. For example, for an essay: "In the early morning, the sun shines gently on the earth, and the birds sing merrily." The server uses the BERT model to perform part-of-speech tagging.

[0290] The BERT model will identify that "early morning" is a noun, "sunshine" is a noun, "gently" is an adverb, "shines" is a verb, "on" is a preposition, "on the earth" is a noun phrase, "birds" is a noun, "merrily" is an adverb, and "sing" is a verb.

[0291] Then, a score is given based on the accuracy and diversity of part-of-speech tagging. If the part-of-speech tagging in the article is accurate and uses a rich variety of parts of speech, such as more adjectives and adverbs to vividly describe scenes and actions, then a higher score will be obtained in the part-of-speech tagging indicator. For example, if an article, in addition to common nouns and verbs, also skillfully uses many expressive adjectives, such as "gorgeous sunset glow" and "quiet forest", this indicates that the author is proficient in using parts of speech and may get 8 - 10 points (assuming a full score of 10 points) in this indicator. On the contrary, if there are part-of-speech tagging errors, such as mislabeling a verb as a noun, or the use of parts of speech is very single, only 3 - 5 points may be obtained.

[0292] Syntactic analysis score:

[0293] Next, the server performs syntactic analysis on the writing work. Still taking the above-mentioned prose as an example, the BERT model will analyze the sentence structure. For example, "In the early morning, the sun shines gently on the earth" is a sentence with a subject-predicate structure, where "the sun" is the subject, "shines" is the predicate, "on the earth" is the complement, and "gently" is the adverbial; "The birds are singing cheerfully" is also a subject-predicate structure, where "the birds" is the subject, "are singing" is the predicate, and "cheerfully" is the adverbial.

[0294] A score is given according to the reasonableness and complexity of syntactic analysis. If the sentence structure is complete, reasonable, and has a certain degree of complexity, such as using multiple clause and compound sentence structures to express complex ideas, a higher score will be obtained. For example, "When night falls, those elves hidden in the darkness, they shuttle among the trees, as if telling ancient stories, and I, standing there quietly, am attracted by this mysterious atmosphere." This sentence contains multiple structures such as a time adverbial clause and a coordinate sentence, showing the author's good syntactic application ability and may get 7 - 9 points in the syntactic analysis indicator. If the sentence structure is simple and there are many grammar errors, such as incomplete sentence components and chaotic word order, only 2 - 4 points may be obtained.

[0295] Sentiment analysis score:

[0296] The server will also perform sentiment analysis on the writing work. For the description in the prose, such as "The sun shines gently on the earth, and the birds are singing cheerfully", the BERT model can judge that the overall sentiment tendency is positive.

[0297] Score according to the depth, authenticity, and coherence of emotional expression. If an article can express a positive or negative emotion deeply, authentically, and coherently, such as an article describing the grief of losing a loved one, through delicate descriptions and progressive emotional layers, enabling readers to deeply feel the author's pain, then it will receive a relatively high score in the emotional analysis index, perhaps 8 - 10 points. If the emotional expression is rather vague, incoherent, or inconsistent with the content of the article, for example, the article describes beautiful scenery at the beginning and then suddenly expresses sadness illogically, it may only get 3 - 5 points.

[0298] Finally, by integrating the scores of part-of-speech tagging, syntactic analysis, and emotional analysis, a multi-dimensional scoring result of the writing work is obtained. For example, if an article gets 7 points in part-of-speech tagging, 6 points in syntactic analysis, and 8 points in emotional analysis, then its multi-dimensional scoring result may be (7 + 6 + 8) / 3 = 7 points (assuming the weights of each index are the same).

[0299] In the case where the work to be determined is an image work or a video work, based on the LSTM and Transformer models, painting style, painting techniques, and artistic elements are used as scoring indicators for multi-dimensional scoring to obtain the multi-dimensional scoring result of the image work or video work.

[0300] Painting style scoring (taking an image work as an example):

[0301] In an art work solicitation activity, the server receives an image work. For an oil painting, the server first analyzes its painting style. If this oil painting is in the Impressionist style, the LSTM and Transformer models will identify its typical features, such as loose brushstrokes, emphasizing the light and shadow effects, the mixing and interweaving of colors to create a hazy visual effect, etc.

[0302] Score according to the typicality, uniqueness, and innovativeness of the painting style. If this oil painting very typically embodies the Impressionist style, such as the feeling of capturing the instantaneous light and shadow through colors and brushstrokes like Monet's "Water Lilies", and has some unique innovations on this basis, such as being more daring in the use of colors, then it may get 8 - 10 points in the painting style index. If the painting style is not clear, seemingly being both Impressionist and having a shadow of Realism, but neither well reflects their respective characteristics, it may only get 3 - 5 points.

[0303] Painting technique scoring (taking an image work as an example):

[0304] Next, we analyze the painting techniques. For oil paintings, painting techniques include color matching, brushstroke application, object shape depiction, etc. For example, in terms of color matching, if the color transition in the picture is natural and harmonious, and there is no obvious color conflict (unless it is a deliberate artistic effect), such as the natural transition from light blue to dark blue sky color, then it will get a good score in this aspect.

[0305] In terms of brushstroke application, if the brushstrokes are delicate and expressive, and can vividly depict the texture of the object, such as soft and delicate brushstrokes when depicting human skin, and rough and powerful brushstrokes when depicting rocks, this will also increase the score of painting skills. Scoring is based on the comprehensive performance of these painting skills. If the painting skills are skillful and exquisite, it may get 7-9 points on this indicator; if the color matching is not coordinated, the brushstrokes are stiff and lack of expression, it may only get 2-4 points.

[0306] Artistic element scoring (taking image works as an example):

[0307] Then analyze the artistic elements. Artistic elements include composition, lines, light and shadow, etc. For example, in terms of composition, if the picture adopts a unique composition method, such as golden ratio composition, places the subject in a key position in the picture, and makes the whole picture balanced and attractive, then it will get a good score in the artistic element of composition.

[0308] In terms of the use of lines, if the lines are smooth and rhythmic, and can well outline the contours and shapes of objects, such as using smooth curves to depict the contours of a woman's body, the score will also increase. Scoring is based on the comprehensive use of artistic elements. If the artistic elements are used properly and richly, this indicator may get 7-9 points; if the composition is messy and the lines lack beauty, it may only get 2-4 points.

[0309] For video works, the analysis of painting style, painting techniques and artistic elements is similar, but the dynamic elements of the video, such as the switching of shots and the rhythm of the picture, need to be considered. For example, in terms of shot switching, if the switching is natural and smooth and can guide the audience's sight well, it will get a good score in the artistic element of shot switching. The scores of painting style, painting techniques and artistic elements are combined to obtain the multi-dimensional scoring results of the image work or video work. For example, if an image work gets 8 points in painting style, 7 points in painting techniques, and 8 points in artistic elements, then its multi-dimensional scoring result may be (8+7+8) / 3=7.67 points (assuming that the weight of each indicator is the same).

[0310] When the pending email work is a calligraphy work, a multi-dimensional score is performed based on the area difference, aspect ratio difference, distance difference and orientation difference between the calligraphy work and the standard characters to obtain a multi-dimensional score result of the calligraphy work.

[0311] Area Difference Rating:

[0312] In a calligraphy work submission platform, the server receives calligraphy works. For each calligraphy work, the server first compares it with the standard characters to analyze the area difference. For example, the area of ​​the standard characters is set to 10 square centimeters (this is an assumed standard, which can be set according to the font type and requirements in practice).

[0313] If the average area of ​​the characters in the calligraphy work is 9-11 square centimeters, it means that the area difference with the standard characters is small, and it may get 8-10 points on this indicator. Because the small area difference means that the font size is more in line with the standard and the writing is more standardized. If the average area of ​​the characters is 6-8 square centimeters or 12-14 square centimeters, that is, the area difference with the standard characters is large, it may only get 3-5 points, because this may affect the overall visual effect and standardization.

[0314] Aspect Ratio Difference Rating:

[0315] Next, we analyze the difference in aspect ratio. Assuming that the aspect ratio of a standard character is 1:1.2, for characters in calligraphy works, if the aspect ratio of most characters is between 0.9:1.1-1.1:1.3, it means that the aspect ratio difference is small, and the character may get 7-9 points on this indicator. This shows that the shape ratio of the character is close to the standard and meets the aesthetic standards of calligraphy.

[0316] If the aspect ratio is 0.7:1.5 or 1.3:1.0, which is significantly different from the standard aspect ratio, you may only get 2-4 points, because such characters may appear uncoordinated and not conform to the proportional beauty of traditional calligraphy.

[0317] Distance Difference Rating:

[0318] Then analyze the distance difference, where the distance refers to the spacing between characters and lines. If the spacing between characters and lines in the calligraphy work is close to the standard calligraphy layout spacing (assuming the standard spacing is 1 / 2 of the character height), for example, the character spacing is between 0.4-0.6 times the character height, and the line spacing is between 1.4-1.6 times the character height, this indicator may get 8-10 points. This shows that the layout is reasonable and the density is appropriate.

[0319] If the character spacing is too large or too small, and the line spacing does not meet the standard, such as the character spacing is 0.2 times or 0.8 times the character height, and the line spacing is 1.0 times or 2.0 times the character height, you may only get 3-5 points, because such a layout will make the work look crowded or loose, affecting the overall beauty.

[0320] Azimuth difference score:

[0321] Finally, analyze the azimuth difference, that is, whether the writing azimuth of the characters is regular. If the characters in the calligraphy work are basically horizontal and vertical without obvious skewing, a score of 8-10 may be obtained in this index. This reflects the writer's control ability over the structure and layout of the characters.

[0322] If the characters have obvious skewing, such as some characters being tilted by more than 5 degrees (this is a hypothetical threshold), a score of only 2-4 may be obtained because this will affect the overall neatness and beauty of the calligraphy work.

[0323] Combine the scores of the area difference, width-to-height ratio difference, distance difference, and azimuth difference to obtain the multi-dimensional scoring result of the calligraphy work. For example, if a calligraphy work gets 8 points in the area difference, 7 points in the width-to-height ratio difference, 9 points in the distance difference, and 8 points in the azimuth difference, then its multi-dimensional scoring result may be (8 + 7 + 9 + 8) / 4 = 8 points (assuming the weights of each index are the same).

[0324] In order to more clearly describe the solution provided by the embodiments of the present invention, a more complete implementation manner is provided below.

[0325] 1. A method for collecting works from an email box:

[0326] The organizing committee generally has the following ways to organize the competition to collect submission works:

[0327] a) Set up a dedicated email box: Contestants submit their works via email, and the organizing committee can batch download or manage these works through the email box.

[0328] b) Third-party competition management platform: Use a third-party competition management platform, which generally provides an integrated service for work submission, review, and display.

[0329] c) Official website: Integrate the work submission function in the official website or dedicated application of the organizing committee, so that contestants can directly upload works through the website or application. Among them, collecting contestants' works by email is the most extensive way.

[0330] The present invention provides an efficient method for collecting works from an email box to efficiently download and manage a large number of participating email works, and automatically screen out works that obviously do not meet the participation conditions, saving time and reducing human errors. The following are some key steps and methods:

[0331] S1 Configure the email server parameters: Obtain the IMAP server address and port number of the target email server. Configure the SSL / TLS secure connection to ensure the security of data transmission.

[0332] S2 Connect to the mail server: Establish a connection through the IMAP address and port of the mail server configured in step S1, as well as the login credentials.

[0333] S3 Download emails:

[0334] S3-1 Define the submission start time T_start and submission end time T_end according to the submission time requirements of the current competition event. Filter out emails with submission times greater than T_start and less than T_end, and use the FETCH command to obtain the email body and attachments.

[0335] S3-2 Store the email content in a relational database and a file server in the required format.

[0336] S4 Filter out email works that do not meet the conditions:

[0337] S4-1 Configure filtering and screening rules according to the work submission requirements of the current event theme, such as the submission email title S_title, attachment name F_name, attachment file format F_type, attachment file size F_size, attachment size F_spec, etc.

[0338] S4-2 Perform filtering and screening on the downloaded email works according to the filtering rules configured in step S4-1. Determine whether the submission meets the conditions by defining rules:

[0339] a) Whether the title meets the specifications

[0340] The system presets the email delivery title rules for this event, such as: "school_class_name_contact phone number", which can be verified through regular expressions: ^(\w+)_(\w+)_([\u4e00-\u9fa5]+)_1[3-9]\d{9}$. The following is the explanation of each part of the regular expression: (\w+): Matches the school name. \w matches letters, numbers, and underscores, and + means one or more. _: The literal underscore character. (\w+): Matches the class name, with the same rule. _: The literal underscore character. ([\u4e00-\u9fa5]+): Matches one or more Chinese characters. \u4e00-\u9fa5 is the Chinese character range in the Unicode encoding. _: The literal underscore character. 1[3-9]\d{9}: Matches the mobile phone number in the preset area, starting with 1, the second digit is a number between 3 and 9, and \d{9} means followed by 9 numbers.

[0341] By matching regular expressions, we can quickly filter out submissions whose S_titles meet the standards. For some submissions that are slightly non-compliant with the standards, the system can also make some compatible processing. For example, if the separator is mistakenly written as a hyphen "-", in order to be compatible with such email titles, we can adjust the regular expression appropriately and change it to: ^(\w+)[_-](\w+)[_-]([\u4e00-\u9fa5\w]+)[_-]1[3-9]\d{9}$.

[0342] b) Whether the attachments meet the specifications

[0343] The system presets the attachment naming standard for this event, such as: "School_Class_Name_Work Name.docx". The attachment name F_name can be verified by regular expression, such as: ^(\w+)[_-](\w+)[_-]([\u4e00-\u9fa5\w]+)[_-][\u4e00-\u9fa5\w]+)\.docx$.

[0344] The system automatically downloads the attachments and parses them accordingly to obtain the file type and file size of the attachments; based on the file type, it obtains additional parameter information. For example, for document files such as .doc, .docx, .txt, etc., the system can automatically obtain the number of words in the file; for image files such as .jpg, .jpeg, .png, .bmp, .gif, .svg, etc., it can obtain the corresponding image size, aspect ratio, resolution and other parameter information; for audio and video files such as .mp3, .mp4 file types, it can obtain the corresponding duration and resolution information;

[0345] Combined with the competition rules, we will filter the attachments. For example, if the video length must not exceed 3 minutes, the works that are longer than 3 minutes will be automatically marked as "Exceeded" and archived. If the file size cannot exceed 50MB, the works with attachment size F_size>50Mb will be automatically marked as "Oversized Attachment" and archived.

[0346] S5 further analyzes and processes:

[0347] S5-1 After step S4, the works that obviously do not meet the requirements are filtered out, and the works that meet the requirements are further processed, the content is analyzed, and useful information is extracted, such as title, attachment, email address, mobile phone number, school, grade, class, work name, etc.

[0348] S5-2 classifies the works according to title rules, such as painting, music, writing, calligraphy, etc. Each category is configured with different review and scoring standards and review and scoring models.

[0349] 2. A method for reviewing the compliance of works:

[0350] The compliance check of works mainly examines whether the works comply with laws, regulations, moral standards and social norms. It checks whether the content of the works includes risk content or elements such as porn, violence, political sensitivity, abuse, etc.; whether there is duplicate submission of works; whether the content of the works conforms to the theme, etc.

[0351] The present invention provides a method for reviewing the compliance of works, which helps the competition organizers to reduce the manual review cost and improve the efficiency and accuracy of preliminary screening in academic and creative competitions (such as calligraphy competitions, composition competitions, creative design competitions, etc.). The following are some key steps and methods:

[0352] S1 Preset rules:

[0353] According to laws, regulations and social guidelines, set predefined rules and standards for preliminary screening of content. Configure a sensitive word library dictionary to detect whether the works hit sensitive and illegal words.

[0354] S2 Work classification:

[0355] The forms of works participating in academic and creative competitions are diverse, but generally fall into the following types of work forms:

[0356] Document type: such as text documents (txt, docx), PDF documents;

[0357] Presentation type: presentations in PPT, PPTX format;

[0358] Image type: bitmap images (JPEG, PNG, GIF, etc.), vector graphics (SVG, AI, EPS, etc.);

[0359] Audio type: lossless audio files (WAV), compressed audio files (MP3, AAC, etc.);

[0360] Video type: standard video files (MP4, MOV, AVI), high-definition video files (HD, 4K);

[0361] Classify the works according to the type by parsing the title and the attachment suffix, and configure different detection methods and detection models for different classifications.

[0362] S3 Content compliance detection:

[0363] The compliance detection includes a) sensitive information detection; b) duplicate submission detection; c) work integrity detection; c) theme conformity detection;

[0364] Based on the classification results obtained in step S2, different models are used for detection according to different classifications. a) Sensitive information detection: S3-1, for text types, use NLP technology to perform natural language processing on the works, and conduct inspections in aspects such as language style, sensitive vocabulary, and political sensitivity. Split continuous text strings into meaningful units, assign grammatical categories to each word in the text. Identify specific entities in the text, analyze the dependency relationships between words in the text, and construct the grammatical tree structure of the sentence. Utilize machine learning models to quantitatively analyze the sentiment tendency of the text based on the vocabulary and context in the text. Identify whether the sensitive vocabulary or inappropriate vocabulary in step S1 is hit in the text through pattern matching and machine learning models, and these words may include inappropriate content such as political topics or remarks or pornographic violence. If the text contains an image link or description, use computer vision technology to analyze the image content. Map the vocabulary to vectors in a high-dimensional space to capture the semantic relationships between words. Apply algorithms such as logistic regression, support vector machine (SVM), and random forest to train the model to identify text features. Judge the sentiment tendency of the text through sentiment analysis and identify the text that may contain negative emotions or offensive language.

[0365] S3-2, for image types, use computer vision technology to detect inappropriate content. Use deep learning models, convolutional neural networks (CNNs) to identify specific objects or scenes in the pictures. Train deep learning models using a large number of labeled picture datasets so that they can identify and distinguish different types of content. Segment the pictures to identify different regions or objects in the images for more accurate analysis and identification. Use face recognition technology to detect faces in the pictures and identify the content that may involve personal privacy. Identify scenes in the pictures, such as violent scenes, pornographic scenes, etc., through training specific scene recognition models. If the picture contains text, OCR (Optical Character Recognition) technology can be used to extract the text and further analyze the text content using step S3-1 to detect the compliance of the text.

[0366] S3-3, for audio types, use ASR technology to convert the video sound into text, and then perform the detection in step S3-1 on the text; S3-4, for video types, use ASR technology to convert the video sound into text, and then perform the detection in step S3-1 on the text; use video decoding technology to extract frames at a certain time interval interval and convert them into images, and then perform the detection in step S3-2 on the extracted frames respectively;

[0367] b) Duplicate submission / plagiarism detection:

[0368] S3-5, utilize data fingerprint technology and hash algorithms to identify and prevent the same work from being submitted repeatedly. Detect plagiarism problems by comparing existing works in the database and using text matching and image comparison algorithms.

[0369] The steps for document class detection are as follows:

[0370] S3 - 5.1 Word Segmentation and Stop Word Filtering:

[0371] Perform word segmentation on the text, splitting the continuous text into discrete lexical units. Remove stop words, such as common words like "de", "he", "shi", etc., which contribute less to the analysis.

[0372] S3 - 5.2 Feature Extraction:

[0373] Use the TF - IDF method to convert the text into a vectorized representation for easy calculation of similarity.

[0374] S3 - 5.3 Similarity Calculation:

[0375] Apply cosine similarity and Jaccard similarity algorithms to calculate the similarity between texts.

[0376] S3 - 5.4 Machine Learning Model Training:

[0377] Use the SVM machine learning algorithm to train a model to identify plagiarized texts.

[0378] Use the BERT deep learning model for deeper semantic comparison.

[0379] S3 - 5.5 Threshold Setting and Plagiarism Judgment:

[0380] Set a similarity threshold. Texts exceeding this threshold are considered potential plagiarisms. Mark the texts with similarity exceeding the threshold.

[0381] The steps for image class detection are as follows:

[0382] S3 - 5.6 Feature Extraction:

[0383] Use the SURF algorithm to extract the key points and descriptors of the image. Use the CNN deep learning model to extract the deep features of the image.

[0384] S3 - 5.7 Feature Matching:

[0385] Use the FLANN and Brute - Force algorithms for feature matching.

[0386] S3 - 5.8 Similarity Calculation:

[0387] Calculate the similarity of matching features, mainly by measuring the similarity between data points through the Euclidean distance calculation formula. For two points x and y in the dataset, their Euclidean distance can be calculated by the following formula: d(x,y)=sqrt((x1 - y1)^2+(x2 - y2)^2+...+(xn - yn)^2) where x1, x2,..., xn and y1, y2,..., yn are the coordinate values of points x and y in each dimension respectively, and n is the dimension of the point data.

[0388] S3-5.9 Similar Image Comparison:

[0389] Compare the image to be detected with the images in the database, and find similar images according to the similarity results calculated in step S3-5.8.

[0390] S3-5.10 Threshold Setting and Decision Making:

[0391] Set a similarity threshold T to determine whether an image is plagiarized. Mark the works whose similarity exceeds the threshold.

[0392] c) Theme Conformity Detection:

[0393] For composition and keynote speech contests, it is possible to detect whether the content conforms to the contest theme. The specific steps are as follows:

[0394] S3-5.11 Keyword and Phrase Extraction

[0395] Use TF-IDF and TextRank techniques to extract keywords and phrases in the text. Determine the named entities in the text as indicators of theme relevance.

[0396] S3-5.12 Build a Theme Model

[0397] Use the LDA theme model algorithm to extract the theme distribution from the reference data.

[0398] According to the vocabulary and theme assignment in the document, calculate the occurrence probability of each vocabulary on each theme.

[0399] Update the theme assignment: According to the occurrence probability of each vocabulary on each theme, update the theme assignment of the document.

[0400] Repeat the above steps until the theme assignment and the occurrence probability of vocabulary on the theme reach stability.

[0401] The core formula of the mathematical model of the LDA theme model is as follows:

[0402] p(w,z∣α,β)=p(w∣z,β)∗p(z∣α). Here, w represents the words in the document; z represents the topic to which the word belongs; α and β are the hyperparameters of the model, which respectively control the prior distributions of the document topic distribution and the topic-word distribution; p(w∣z,β) represents the probability that word w appears under the condition of a given topic z and parameter β; p(z∣α) represents the topic distribution of the document under the condition of a given parameter α; compare the extracted topics with the keywords of the work.

[0403] S3-5.13 Semantic Understanding and Dependency Parsing

[0404] Use the dependency parsing tool (spaCy library) to understand the semantic structure of the text. The dependency parsing in spaCy depends on the weights and activation functions of the neural network model. The key steps are as follows:

[0405] Feature extraction: xt=f(w,ct), where xt is the feature representation of the t-th word, w is the model parameter, and ct is the context information of the word. Neural network forward propagation: ht=σ(W⋅xt+b), where ht is the hidden state of the t-th word, σ is the activation function, such as ReLU. Dependency relationship and label prediction: yt=g(V⋅ht) where yt is the predicted probability distribution of the dependency relationship and label of the t-th word, V is the weight matrix of the output layer, and g is the softmax function.

[0406] S3-5.14 Similarity Calculation

[0407] Use measurement methods such as cosine similarity to calculate the similarity Similar between text vectors. The cosine value between two vectors can be obtained by using the Euclidean dot product formula:. Given two attribute vectors, A and B, their cosine similarity θ is given by the dot product and vector length, as follows:

[0408]

[0409] Here, Ai and Bi represent the respective components of vectors A and B.

[0410] The similarity given ranges from -1 to 1: -1 means that the directions pointed by the two vectors are exactly opposite, 1 means that their directions are exactly the same, 0 usually means that they are independent of each other, and values in between represent intermediate similarity or dissimilarity.

[0411] S3-5.15 Threshold Setting:

[0412] Set the similarity threshold T as the criterion for judging whether the text conforms to the topic.

[0413] S3-5.16 Automated Detection and Scoring:

[0414] Automatically detect works using the trained model and the set threshold.

[0415] The score Score = the similarity Similar calculated in step S3 - 5.14 * 100. Compare the score Score and the threshold T algorithmically. When Score >= T, it indicates compliance with the theme.

[0416] Generate a score or classification result that complies with the theme for each work.

[0417] Through this multi - level and multi - dimensional review method, effectively ensure the compliance of works, while improving the efficiency and fairness of the review process.

[0418] 3. A method for preliminary scoring of works:

[0419] The present invention provides a method for preliminary scoring of works, which can conduct preliminary scoring on participating works (calligraphy, literary works, creative writing, painting, etc.), provide data support and decision - making suggestions for the review team, and improve the efficiency of the review process. The following are some key steps and methods:

[0420] The work scoring module conducts multi - dimensional evaluations on different types of participating works according to the defined scoring criteria and uses an AI technology model for training to obtain a comprehensive reference score.

[0421] S1 Set the scoring criteria and weights:

[0422] Clarify the quantitative dimensions of work scoring, including but not limited to key indicators such as creativity, technicality, and originality.

[0423] Define the dimensions DIM and weights W of the score, such as creativity, technicality, expressiveness, etc. Set quantitative indicators and scoring rules for each dimension. For example, for writing works, evaluate grammar, spelling, content, logic, and creativity; for calligraphy works, evaluate whether there are writing errors, writing neatness, pen tips, strength, etc.

[0424] S2 Feature standardization:

[0425] Standardize the feature values to the unified range of 0 to 1 to eliminate the influence of different dimensions and magnitudes.

[0426] S3 Feature extraction:

[0427] Digitally process the works, clean and standardize the data. Extract key features according to different work types.

[0428] a) For writing: Use NLP techniques to analyze the content of the article, including part-of-speech tagging, syntactic analysis, sentiment analysis, etc. Evaluate grammar errors in the article, such as subject-verb agreement, preposition collocations, etc. Analyze the use of vocabulary in the article, including vocabulary richness and appropriateness. Evaluate the complexity and diversity of sentences, such as the use of simple sentences, compound sentences, and complex sentences. Check whether the content of the article is relevant to the topic and whether it is developed around the given theme. Analyze the logical relationships between paragraphs and the overall coherence of the article. Evaluate the writing style and innovation of the article, such as the use of rhetorical devices like metaphors and similes. Construct scoring dimensions from multiple aspects such as text smoothness, literary grace, idea analysis, and text structure.

[0429] b) For image and video:

[0430] Use image processing techniques to perform preprocessing operations on the input painting images, such as denoising and normalization. To make the data more uniformly distributed in each dimension and accelerate the faster convergence of the optimization algorithm. Avoid the model from overfitting in one feature dimension and improve the generalization ability of the model. Use a convolutional neural network (CNN) to extract key visual features from the painting works.

[0431] c) For calligraphy:

[0432] Use pattern recognition techniques to extract data on strokes, structures, layouts, sizes, and components from the works, and perform multi-dimensional analysis on calligraphy characters, including multiple dimensions such as component area, component width-to-height ratio, distance between components, and orientation between components

[0433] S4 dimension score calculation:

[0434] According to the scoring dimensions and weights defined for different work types, calculate the dimension scores based on the extracted key features.

[0435] For writing:

[0436] Based on the pre-trained language model BERT, use the output of the last layer corresponding to [CLS] to represent the article and input it into the classifier to achieve scoring: Score = Max(L, MIN(B + F, U)).

[0437] In the above formula, Score represents the final score, the MAX function represents taking the larger value of the two, the MIN function represents taking the smaller value of the two, L represents the lowest score of a certain grade, U represents the highest score of a certain grade, B represents the basic score of a certain grade, and F represents the floating score.

[0438] For image and video:

[0439] Based on the extracted features, evaluate the style, techniques, and artistic elements of the paintings. Select 1 million image works from the historical works library. To ensure the generality of the training model for various categories and styles of pictures, these pictures need to be evenly distributed among various styles and themes, such as: figures, still lifes, color, black and white, cities, scientific creativity, humanistic photography, etc. And select the most complete score distribution map as the data label (to prevent the influence of extreme scores and unobjective scores, the score of each picture is taken as the median score of all scores).

[0440] Train LSTM and Transformer deep learning models to identify and evaluate the creativity and uniqueness of painting works. Through a multi-task learning framework, simultaneously evaluate the technical and artistic aspects of the paintings. Utilize the model pre-trained on a large-scale dataset and transfer it to the painting scoring task to improve the learning efficiency.

[0441] Based on Early Stopping, minimize the training time as much as possible to avoid overfitting, and stop increasing the training cycle in a timely manner when there is no obvious improvement in accuracy. Test different optimizers and learning rates to ensure that the error between the final model score and the manual score is within an acceptable range.

[0442] Select the deep residual network ResNet50 for training, modify the number of output units to 1, and ensure that the predicted score S of the output is between 0 and 1 through the activation function limit. The final score calculation formula is as follows: Score = S * 100. Through the above model training method, the score Score of the image work is finally obtained.

[0443] For calligraphy works:

[0444] Calculate the difference between the area ratio of two components of the calligraphy character and the area ratio of two components of the standard character: D_area = ((area_hw^(i + 1) / area_hw^i)) ⁄ ((area_gt^(i + 1) / area_gt^i) - 1). Among them, area represents the area dimension, the subscript hw represents the handwritten calligraphy character, gt represents the corresponding standard character, and i represents the component serial number.

[0445] Calculate the score coefficient of the current component area ratio difference of the handwritten character: Score_area = 1 - ((|D_area| - |α_area|)) ⁄ β_area. Among them, α represents the area threshold, which is used to judge whether the area ratio between components is reasonable and is obtained by kmeans clustering of the pre-classified data. β represents the maximum tolerance threshold, which is set manually.

[0446] The total score of calligraphy character evaluation is the average score of multiple component dimensions:

[0447] Score_total=(Score_area+Score_aspect+Score_distance+Score_angle) / 4。

[0448] Among them, Score_area is the score for the area dimension, Score_aspect is the score for the aspect ratio dimension, Score_distance is the score for the distance dimension, and Score_angle is the score for the orientation dimension. Through the above method, the calligraphy works can be automatically evaluated and the initial score data can be output.

[0449] S6 Final score calculation:

[0450] Design an automated scoring process to ensure the consistency and repeatability of scoring.

[0451] Integrate the evaluation and score data of multiple dimensions into a single score, set thresholds according to the score result distribution, and distinguish works of different levels.

[0452] Multiply the scoring result of each dimension by its corresponding weight, and then add the weighted results to obtain the preliminary score of the work.

[0453] Final_Score=w1×s1+w2×s2+…+wn×sn+Post-Processing.

[0454] Where: wi is the weight of the i-th feature, si is the score of the i-th feature, n is the total number of features, and Post-Processing represents any necessary post-processing steps, such as normalization or adjustment.

[0455] Through this method, multiple dimensions can be comprehensively considered to give the final score data of a work.

[0456] S7 Manual review feedback:

[0457] Introduce a manual review mechanism to review the preliminary AI scoring results to ensure the fairness and accuracy of scoring.

[0458] Collect review feedback for further training and optimizing the model. S8 Model iteration and performance optimization:

[0459] According to the feedback and review results, continuously iterate and optimize the model to improve the scoring accuracy. Regularly update the training set to incorporate new scoring data and expert review opinions. Through the above multi-dimensional and multi-index scoring method and the ability of deep learning and self-optimization, provide an efficient and fair scoring method for the review work of participating works.

[0460] Please refer to Figure 2 , Figure 2A work screening device 110 provided by an embodiment of the present invention based on an email box, comprising:

[0461] An acquisition module 1101, configured to acquire original emails from a target email server and perform filtering to obtain pending email works;

[0462] A screening module 1102, configured to perform work compliance detection based on the work type of the pending email works; in the case of passing the work compliance detection, perform multi-dimensional scoring on the pending email works to obtain a multi-dimensional scoring result of the pending email works; in response to a user review pass indication for the multi-dimensional scoring result, obtain a final screening result of the pending email works according to the magnitude of the multi-dimensional scoring result.

[0463] It should be noted that the implementation principle of the foregoing work screening device 110 based on an email box can refer to the implementation principle of the foregoing work screening method based on an email box, which will not be elaborated here. It should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the work screening device 110 based on an email box can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the work screening device 110 based on an email box. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together or independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the hardware integrated logic circuit or software-form instructions in the processor element.

[0464] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more microprocessors (digital signal processors, DSP), or one or more field programmable gate arrays (FPGA), etc. For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0465] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned mailbox-based work screening device 110. Figure 3 As shown, Figure 3 The computer device 100 is a block diagram of a structure of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a mailbox-based work screening device 110 , a memory 111 , a processor 112 and a communication unit 113 .

[0466] To achieve data transmission or interaction, the memory 111, processor 112 and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The mailbox-based work screening device 110 includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the mailbox-based work screening device 110 stored in the memory 111, such as the software function modules and computer programs included in the mailbox-based work screening device 110.

[0467] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the aforementioned mailbox-based work screening device 110.

[0468] For purposes of illustration, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the disclosure and its practical application, to thereby enable others skilled in the art to best utilize the disclosure and various embodiments with various modifications as are suited to the particular application contemplated.

Claims

1. A method for screening works based on email, characterized in that, including: Obtain the original emails from the target email server and filter them to obtain the pending email works; Conduct work compliance detection based on the work type of the pending email works; In the case of passing the work compliance detection, perform multi-dimensional scoring on the pending email works to obtain the multi-dimensional scoring result of the pending email works; In response to the user's review and approval instruction for the multi-dimensional scoring result, obtain the final screening result of the pending email works according to the size of the multi-dimensional scoring result; The obtaining the original emails from the target email server and filtering them to obtain the pending email works includes: Obtain the IMAP server address and port number of the target email server; Establish a connection with the target email server based on the IMAP server address and the port number in combination with the login credentials; Download the original emails in the target email server according to the preset submission time range, where the original emails include the original email text and the original email attachments; Filter the original email text and the original email attachments according to the preset filtering rules. In the case where the filtering result indicates compliance with the filtering rules, use the original email attachment as the pending email work; The conducting work compliance detection based on the work type of the pending email works includes: Determine the detection criteria for the pending email works based on the work type of the pending email works; Conduct sensitive information detection, duplicate submission detection, and theme consistency detection on the pending email works based on the detection criteria; The performing multi-dimensional scoring on the pending email works to obtain the multi-dimensional scoring result of the pending email works includes: In the case where the pending email work is a writing work, perform multi-dimensional scoring based on the BERT model using part-of-speech tagging, syntactic analysis, and sentiment analysis as scoring indicators to obtain the multi-dimensional scoring result of the writing work; In the case where the pending email work is an image work or a video work, perform multi-dimensional scoring based on the LSTM and Transformer models using painting style, painting skills, and artistic elements as scoring indicators to obtain the multi-dimensional scoring result of the image work or the video work; In the case where the pending email work is a calligraphy work, perform multi-dimensional scoring according to the area difference, aspect ratio difference, distance difference, and orientation difference between the calligraphy work and the standard characters to obtain the multi-dimensional scoring result of the calligraphy work.

2. The method according to claim 1, characterized in that The conducting sensitive information detection on the pending email works includes: In the case where the pending email work is a text work, preprocess the text work and perform sentiment tendency analysis, and combine with a preset sensitive word library and a preset sensitive word detection model to conduct sensitive information detection; In the case where the pending email work is an image work, segment the image work and combine with a preset scene recognition model to conduct sensitive information detection; in the case where the image work includes text content, extract the text from the image work, preprocess the text content and perform sentiment tendency analysis, and combine with a preset sensitive word library and a preset sensitive word detection model to conduct sensitive information detection; When the to-be-determined email work is an audio work, use ASR to convert the video sound into audio text, preprocess and perform sentiment analysis on the audio text, and combine a preset sensitive word library and a preset sensitive word detection model to detect sensitive information; When the to-be-determined email work is a video work, use ASR to convert the video sound into audio text, preprocess and perform sentiment analysis on the audio text, and combine a preset sensitive word library and a preset sensitive word detection model to detect sensitive information; Extract the video work at preset time intervals to obtain multiple video images, segment the video images, and combine a preset scene recognition model to detect sensitive information; When the video image includes text content, extract the text from the video image, preprocess and perform sentiment analysis on the text content, and combine a preset sensitive word library and a preset sensitive word detection model to detect sensitive information.

3. The method according to claim 1, wherein Perform duplicate submission detection on the to-be-determined email work, including: When the to-be-determined email work is a text work, perform word segmentation on the text work to obtain multiple lexical units; Remove the stop words from the multiple lexical units, and convert the multiple lexical units into multiple text vectors; Perform text similarity comparison and semantic similarity comparison according to the multiple text vectors and the archived text vectors in the work database to perform the duplicate submission detection; When the to-be-determined email work is an image work, use the SURF algorithm to extract the key points and descriptors of the image work; Extract the deep features of the image work according to the key points and descriptors to obtain the deep features of the image work; Use the FLANN and Brute-Force algorithms, combined with the archived image works in the work database, to perform feature matching similarity comparison to perform the duplicate submission detection.

4. The method according to claim 1, characterized in that, Perform topic compliance detection on the to-be-determined email work, including: Extract keywords and phrases from the to-be-determined email work; Use a pre-trained LDA topic model to determine the current topic distribution; Perform syntactic analysis on the keywords and phrases to obtain the deep features of the to-be-determined email work; Perform similarity comparison according to the current topic distribution and the deep features to achieve the topic compliance detection.

5. A work screening device based on an email box, characterized in that, Including: An acquisition module for acquiring the original email from the target email server and filtering it to obtain the to-be-determined email work; A screening module for performing work compliance detection based on the work type of the to-be-determined email work; When the work compliance detection is passed, perform multi-dimensional scoring on the to-be-determined email work to obtain the multi-dimensional scoring result of the to-be-determined email work; In response to the user review pass instruction for the multi-dimensional scoring result, obtain the final screening result of the to-be-determined email work according to the size of the multi-dimensional scoring result; The acquisition module is specifically used for: Obtain the IMAP server address and port number of the target mailbox server; establish a connection with the target mailbox server based on the IMAP server address and the port number, in combination with the login credentials; download the original emails in the target mailbox server according to the preset submission time range, where the original emails include the original email body and the original email attachments; Filter the original email body and the original email attachments according to the preset filtering rules. When the filtering result indicates compliance with the filtering rules, use the original email attachments as the pending email works; The screening module is specifically used for: Determine the detection criteria for the pending email works based on the work type of the pending email works; perform sensitive information detection, duplicate submission detection, and theme compliance detection on the pending email works based on the detection criteria; When the pending email work is a writing work, perform multi-dimensional scoring using part-of-speech tagging, syntactic analysis, and sentiment analysis as scoring indicators based on the BERT model to obtain the multi-dimensional scoring result of the writing work; When the pending email work is an image work or a video work, perform multi-dimensional scoring using painting style, painting technique, and artistic elements as scoring indicators based on the LSTM and Transformer models to obtain the multi-dimensional scoring result of the image work or the video work; When the pending email work is a calligraphy work, perform multi-dimensional scoring according to the area difference, aspect ratio difference, distance difference, and orientation difference between the calligraphy work and the standard characters to obtain the multi-dimensional scoring result of the calligraphy work.

6. A computer device, characterized in that, The computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1-4.

7. A readable storage medium, characterized in that, The readable storage medium includes a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Automated mail piece quality analysis tool

    US20080033738A1