Data processing method and system for food detection platform
By constructing a keyword list and a collection of synonyms, using named entity recognition and word vector analysis, replacing synonyms in food testing reports, the inaccuracy and abuse of information in the report text when citing is solved, and the safe and reliable use of the text is achieved.
Patent Information
- Application Number
- CN202510464436.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Due to the customization of the testing items and the customization of experimental operations in food testing reports, the information may be inaccurate and abused when citing the report text, especially because the synonymous expression and format differences are difficult to effectively prevent counterfeiting through existing methods.
By constructing keyword lists and synonyms collections, using named entity recognition and word vector analysis, replacing synonyms in the detection report, and controlling text overwriting through discreteness and semantic approximation, a reliable second text is generated to ensure information security.
It realizes that modifications and misuses can be identified when citing text, ensures the accuracy and security of the report content, and prevents text from being incorrectly cited.
Smart Images

Figure CN120337920A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and particularly relates to a data processing method and system for a food detection platform. Background Art
[0002] Food detection reports involving specifications can be anti-counterfeited in various ways.
[0003] Especially for reports with detection standards, by attaching a QR code and scanning the QR code with a mobile phone, and by checking whether the information in the detection report is consistent with the information displayed by mobile phone query, the authenticity of the food detection report can be determined; it is also possible to query by entering the institution website in a browser and entering the report number, anti-counterfeiting code, and verification code. For reports with clear detection content, since their formats are fixed, when performing data verification, only the items and conclusions in the table need to be noted. However, for customized reports, since their detection items include items that need to be detected in the laboratory according to the requirements of the sample provider, and may not include a formatted template, and such reports need to record the operations and experimental records of the experimenters, their anti-counterfeiting is different from that of standardized reports.
[0004] Since the conclusion part and the experimental operation part contain self-entered content, and this part may be rewritten when being cited, resulting in a difference between the presented meaning and the actual expressed meaning. Therefore, when the text in the food inspection report is cited, it is necessary to ensure that the information is not misused through feasible methods.
[0005] It is found in implementation that there are multiple synonymous expressions for the detection components involved in food detection reports. For example, a food additive may have a trade name, a chemical name, and a common name, and some compounds have some approximate expressions. These cause inaccurate information to be presented when the text of the report is cited, and it is impossible to obtain a synonymous text by making large-scale changes to the text of the report, thus posing a challenge to the security of information. Summary of the Invention
[0006] The purpose of the present invention is to provide a data processing method and system for a food detection platform to solve one or more of the above-mentioned technical defects and to achieve the safe and reliable use of information.
[0007] According to an embodiment of the present invention, a data processing method for a food detection platform includes: Determining a first keyword set used in a first text according to the input detection data; Determining a first vector corresponding to the first text; Rewriting the first text to obtain a second text; The similarity between the second vector corresponding to the second text and the first vector is greater than a first threshold; The rewriting includes replacing the corresponding synonyms in the first text with the synonyms in the keyword table, and the distribution of the replaced words in the second text satisfies a preset degree of dispersion. The keyword table includes a key name and a set of synonyms corresponding to the key name; The word vector corresponding to the key name in the keyword table and the word vector similarity of the synonyms in the corresponding synonym set are lower than a second threshold and higher than a third threshold.
[0008] According to an embodiment of the present invention, the first keyword set is obtained in the following manner: Perform named entity recognition on the first text to obtain a first named entity list; Sort the first named entities in descending order of the frequency of occurrence of the named entities to obtain a first named entity list; Determine the named entities in the first named entity list that are the same as the key names in the keyword table as the first keyword set.
[0009] According to an embodiment of the present invention, in response to the first keyword set being empty, determine a second named entity list with no more than m named entities based on the frequency of occurrence of the named entities, where 2 ≤ m ≤ 10; Determine the second named entity list based on the similarity between the named entities in the first named entity list. The semantic similarity between any two named entities in the second named entity list is not higher than the second threshold; Obtain words similar to the named entities in the second named entity list from a database or a user-defined dictionary as alternative words. In response to the user's input, determine the associated words of the named entity selected or input by the user; Save the word input by the user to the keyword table and update it to the first keyword set.
[0010] According to an embodiment of the present invention, in response to the user's input on the user interface, obtain one or more alternative words selected by the user; Filter the alternative words according to the similarity between the alternative words selected by the user and the named entities prompted by the user interface; Perform anomaly recognition based on the word vectors corresponding to the alternative words and remove the abnormal words; Use the alternative words after removing the abnormal words as the synonym set, use the named entities prompted by the user interface as the key names, and update the keyword table and the first keyword set.
[0011] According to an embodiment of the present invention, in response to the first text containing synonyms in the synonym set of the keyword table, replace the synonyms in the first text with the key names corresponding to the elements in the synonym set.
[0012] According to an embodiment of the present invention, the process of obtaining the second text includes: Replacing the first keywords included in the first text with synonyms in the corresponding synonym table in the keyword table to obtain one or more candidate texts; Determining a plurality of different candidate texts based on the semantic similarity between the candidate texts and the first text; Traversing the candidate text set, and determining the candidate text that meets the preset dispersion degree and the similarity degree with the first vector greater than the first threshold as the second text.
[0013] According to an embodiment of the present invention, outlier analysis is performed on the sentence vectors corresponding to a plurality of candidate texts, and the candidate texts with an outlier index greater than 1 are removed.
[0014] According to an embodiment of the present invention, the data processing method further includes: Splitting the rewritten first text according to the first step length and the first window size to obtain a first string array; Determining the elements in the first string array that do not meet the preset rules; Adjusting the elements that do not meet the preset rules according to the preset rules; Generating a second text based on the adjusted first string array.
[0015] According to an embodiment of the present invention, the adjustment includes splitting of strings, adding explanatory text, or merging based on text.
[0016] According to an embodiment of the present invention, a data processing system of a food detection platform includes: A text rewriting unit, configured to determine a set of first keywords used in the first text according to the input detection data, and replace and rewrite the first keywords included in the first text with a keyword table to obtain a second text; A vector obtaining unit, configured to determine the vector corresponding to the text; The keyword table includes a key name and a set of synonyms corresponding to the key name; The similarity degree between the second vector corresponding to the second text and the first vector corresponding to the first text is greater than the first threshold; And the distribution of the replaced words in the second text satisfies a preset discrete value; The similarity degree between the word vector corresponding to the key name in the keyword table and the word vectors of the words in the set of synonyms corresponding to the key name is lower than the second threshold and higher than the third threshold.
[0017] The present invention can achieve the following beneficial effects: 1. When the text is cited, it is possible to determine whether the text has been modified and expresses inconsistent meanings through the meaning expressed by the text; 2. It is possible to determine whether the content of the report has been misused or abused by analyzing the text; 3. Prevent the text from being incorrectly cited. Description of the Drawings
[0018] Figure 1 is a flowchart of the data processing method of the food detection platform provided by the embodiment of the present invention; Figure 2 is a structural diagram of the data processing system of the food detection platform provided by the embodiment of the present invention. Detailed Embodiments
[0019] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0020] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices. According to an embodiment of the present invention, a data processing method for a food detection platform, referring to Figure 1 the flowchart shown, the data processing method includes: Determine the first keyword set used in the first text according to the input detection data; Determine the first vector corresponding to the first text; Rewrite the first text to obtain the second text; The similarity between the second vector corresponding to the second text and the first vector is greater than the first threshold; The rewriting includes replacing the synonyms corresponding to the first text with synonyms in the keyword table, and the distribution of the replaced words in the second text satisfies a preset degree of dispersion. The keyword table includes a key name and a set of synonyms corresponding to the key name; The word vectors corresponding to the key names in the keyword table and the word vectors of the synonyms in the corresponding synonym set have an approximation lower than the second threshold and higher than the third threshold.
[0021] Generally, a report includes fixed-format content and non-fixed-format content. The first text is the text input by the laboratory staff in the non-standard text area of the report.
[0022] For example, content such as "This test report is only responsible for the submitted samples" in the remarks section of the report is in a fixed format and cannot be modified. When it comes to the test method, if it is the content stipulated by the standard, it is in a fixed format. If this part of the content is executed according to the customer's requirements, then it is non-fixed-format content.
[0023] Here, the first text referred to in the present invention is all non-fixed-format text. If a report does not contain non-fixed-format text, then obviously there is no possibility for the user to input the first text, and naturally it is not applicable to the method of the present invention for processing.
[0024] Determine the first keyword set used in the first text according to the input test data. This part can be implemented in various forms. For example, it can be implemented through traditional natural language processing methods or through language large models. It should be noted that since there are many specialized terms in the testing industry, constructing a term dictionary is more important than which tool to use. The construction of the term dictionary can be achieved by collecting common terms and providing them as the dictionary of the selected natural language processing tool.
[0025] When collecting the dictionary, a list of synonyms and near-synonyms can be further sorted out, and the language model can be optimized by segmenting the existing documents. For example, through the training of the Bert model or the conventional word2vec model, the acquisition of word vectors and the retrieval of near-synonyms are realized. It should be noted that the dimensions of the word vectors and the distances or approximations between near-synonyms obtained by different methods are inconsistent. Based on this, the present invention constructs texts with different characteristics under the condition of consistent meanings and processes the food test report data based on this.
[0026] The vectors of the first text or the second text can be obtained in various ways. However, when comparing the two, the vectors obtained by the text embedded expression processed by the same model should be compared.
[0027] When rewriting the first text, it includes replacing the text with keyword pairs. Specifically, it traverses the keywords in the first keyword set, determines the corresponding word sets for these keywords, and replaces them with a word from the word set to obtain the second text after replacement. An implementable operation is to sort the first keyword set in descending order according to the length of the keywords, and then traverse the sorted first keyword set. When the keyword table contains the corresponding first keyword, a keyword is taken from the word set corresponding to the first keyword in the keyword table and replaced.
[0028] The cosine similarity between the word vectors corresponding to the key names in the keyword table and the word vectors of the words in the corresponding word sets of the key names is lower than the second threshold and higher than the third threshold. For example, "appearance" corresponds to a set of words including ["packaging", "outer packaging", "packaging surface", etc.]. Using the dictionary constructed above and the existing text containing these words for training with word2vec, among the obtained word vectors, the cosine similarities between "appearance" and the three corresponding words are 0.94, 0.97, and 0.95 respectively, while the similarity between its word vector and the word vector corresponding to "packaging" is 0.87.
[0029] Simply using this method to replace the text cannot obtain a verifiable text. To ensure that the text can be analyzed, the replaced words should be analyzable, and their distribution, that is, the degree of dispersion, in the second text should have a certain pattern. And semantically, their similarity with the first text should still be within a suitable range. That is, when using a specific model to analyze the two texts, there are differences between them, but the differences are within the controlled range. In this way, when used without authorization, although the literal meanings seem the same, when analyzed using the model, there are recognizable differences between the two. Typically, such as "abnormal smell, damage on the appearance", "appearance" is the identified first keyword, which corresponds to a set of words including ["packaging", "outer packaging", "packaging surface", etc.]. After replacement, it becomes "abnormal smell, damage on the outer packaging". Although the meanings expressed by the two are the same, the cosine similarity of the corresponding sentence vectors is 0.93, and the preset value is between 0.93 - 0.95, which is within the expected range. Based on this, the authenticity of the text can be considered. However, if the text mentioned in a text is "abnormal smell, damage on the outer packaging outside", the cosine similarity of the corresponding sentence vectors is 0.88, indicating that the reported text has been modified. Based on this, it can be considered that the report has been misquoted.
[0030] It should be understood that when the model parameters used are not leaked, the difference in the text is only known to the provider of the report. By providing the corresponding interface and the key text, it can be determined whether the content of the report has been misused or abused.
[0031] In addition, the distribution of the replaced words in the second text is restricted, that is, a rule is preset. For example, when the length of the second text is m, the preset ratio is set to r, such as 0.2 < r < 0.25. When m is 100, the sum of the string lengths corresponding to the replaced words is between 20 and 25. In another case, such as setting a sliding window, if the window size is set to 100 and the range of r is the same as above, then in any sliding window corresponding to the second text, the sum of the string lengths corresponding to the replaced words is between 20 and 25. Obviously, considering the length of the terms and the scenario, it is more appropriate to set the latter to a larger range, such as 0.15 - 0.35.
[0032] Obviously, after constructing the second text in the above manner, the detection of the second text can be carried out in the following manner: (1)Detect the named entities contained in the second text to obtain the second keyword set; (2)Detect whether the distribution of the second keywords in the second text meets the preset rules; (3)Detect the keywords corresponding to the second keywords, replace the keywords in the second text to obtain the third text, calculate the cosine similarity between the third text and the second text, and determine whether the text has been substantially tampered with based on the obtained cosine similarity.
[0033] According to an embodiment of the present invention, the first keyword set is obtained in the following manner: Perform named entity recognition on the first text to obtain the first named entity list; Sort the first named entities from high to low according to the frequency of occurrence of the named entities to obtain the first named entity list; Determine the named entities in the first named entity list that are the same as the key names in the keyword table as the first keyword set.
[0034] In this way, a series of words can be obtained as the basis for reconstructing the text.
[0035] In an embodiment of the present invention, the first keyword set is obtained in the following manner: Perform named entity recognition on the first text to obtain the first named entity list; Sort the first named entities from high to low according to the frequency of occurrence of the named entities to obtain the first named entity list; Define the first numerical value as 0, construct an empty set of strings, traverse the first list of named entities. When the keyword table contains the current element, obtain the ratio of the element in the first list of named entities to the first text, and use the sum of it and the first numerical value as the new value of the first numerical value. Add the current element to the set of strings. When the ratio of the first numerical value to the length of the first text is within the preset range of discreteness, stop traversing; when the keyword table does not contain the current element, traverse the next element; Use the set of strings as the first set of keywords.
[0036] In this way, keywords that meet the design requirements can be determined and replaced to obtain the second text.
[0037] According to an embodiment of the present invention, in response to the first set of keywords being empty, determine a second list of named entities with a number not exceeding m based on the frequency of occurrence of named entities, where 2 ≤ m ≤ 10; Determine the second list of named entities based on the approximation between the named entities in the first list of named entities, and the semantic approximation between any two named entities in the second list of named entities is not higher than the second threshold; Obtain words similar to the named entities in the second list of named entities from the database or user-defined dictionary as alternative words, and in response to the user's input, determine the associated words of the named entity selected or input by the user; Save the word input by the user to the keyword table and update it to the first set of keywords.
[0038] In this way, the expansion of undefined keywords can be achieved, avoiding the defect that data processing cannot be performed when there are no keywords.
[0039] By limiting the semantic approximation between any two named entities to not be higher than the second threshold, it is avoided that multiple named entities correspond to the same word, thereby avoiding using one technical term corresponding to multiple technical terms in the obtained second text.
[0040] By setting an appropriate value of m, the problem of insufficient keywords in the first text can be avoided. However, when setting a larger value of m, if the user needs to input too many options, it may cause a decrease in accuracy or the possibility of operation errors.
[0041] Regarding the construction of the user interface, it can be achieved by providing a web interface and setting a search window and an interactive interface to associate words with each other. However, it should be realized that the user may select words 2, 3, 4, and 5 as synonyms for word 1. However, when the number of model training times is different, only some synonyms may be selected as objects.
[0042] According to an embodiment of the present invention, in response to an input by a user at a user interface, one or more alternative words selected by the user are obtained; The alternative words are filtered according to the proximity between the alternative words selected by the user and the named entities prompted by the user interface; Anomaly recognition is performed according to the word vectors corresponding to the alternative words, and abnormal words are removed; The alternative words with abnormal words removed are used as a synonym set, and the named entity prompted by the user interface is used as a key name to update the keyword table and the first keyword set.
[0043] In this way, the keyword table can be updated, which is applicable to a pre-trained model or an online training model. When using the latter, based on the newly submitted keywords as a dictionary, the documents in the system are trained to obtain an update of the model weights. This process may cause changes in the weight values in some existing keyword tables. However, in actual tests, it is found that when the meaning of the words is clear and polysemous words are not used, the word vectors of the corresponding words may change, but the change in the proximity between the words is basically unchanged.
[0044] According to an embodiment of the present invention, when a synonym in the synonym set of the keyword table is included in the first text, the synonym in the first text is replaced with the key name corresponding to the element in the synonym set.
[0045] In this way, it is possible to avoid directly introducing the elements in the synonym set included in the keyword table into the first text, and avoid the problem of inaccurate data caused when calculating the dispersion degree and proximity.
[0046] For example, "After the sample is placed in an environment with a temperature of 35 degrees Celsius for 14 days, the reduction ratio of the carbon dioxide content is...", where the keyword table contains the following mapping "reduction rate" and ["reduction ratio", "reduction proportion"], which includes the elements in the synonym set and should be replaced with "After the sample is placed in an environment with a temperature of 35 degrees Celsius for 14 days, the reduction rate of the carbon dioxide content is...", and then other text processing is performed. According to an embodiment of the present invention, the process of obtaining the second text includes: The first keyword included in the first text is replaced with the synonym in the corresponding synonym table in the keyword table to obtain one or more candidate texts; Based on the semantic proximity between the candidate text and the first text, multiple different candidate texts are determined; Traverse the candidate text set, and determine the candidate text that meets the preset dispersion degree and the proximity to the first vector greater than the first threshold as the second text.
[0047] In a more specific implementation, it is described that the first keyword set contains m key names. Among them, the first key name corresponds to a synonym set with a length of N1, the second key name corresponds to a synonym set with a length of N2, and so on. The mth key name corresponds to a synonym set with a length of Nm. In the case of replacing all synonyms, a total of N1 * N2 *... * Nm combinations can be formed. Obviously, in a long paragraph, such a computational workload is huge and unnecessary. To avoid this situation, the length of the first text can be set to a preset value, for example, less than 500 words, so that the formed combinations are controllable.
[0048] In some cases, the order of the text can also be sorted to provide text that meets the preset degree of dispersion without changing the semantics.
[0049] Whether to sort the text can be determined in the following way: The text is segmented according to punctuation marks (such as "."), and then the correlation between adjacent sentences is compared. If the correlation between adjacent sentences is greater than the fourth threshold (such as 0.85), it is considered that the two are related and the order cannot be reversed, and both sentences are marked as non-exchangeable; otherwise, the order of the two can be exchanged. When marking, if a sentence has been marked as non-exchangeable, even if it is marked as exchangeable in the next round of matching, it is still non-exchangeable; and after a sentence is exchanged to be exchangeable, it can still be marked as non-exchangeable.
[0050] After that, the order of the sentences marked as exchange is adjusted to make it meet the preset degree of dispersion. According to an embodiment of the present invention, outlier analysis is performed on the sentence vectors corresponding to multiple candidate texts, and candidate texts with an outlier index greater than 1 are removed.
[0051] By performing outlier analysis on the text sentence vectors, some candidate texts with significantly different semantics from other sentences can be removed. To achieve screening faster, traditional methods such as bert or word2vec can be used here. According to an embodiment of the present invention, the data processing method further includes: The rewritten first text is split according to the first step length and the first window size to obtain a first string array; Determine the elements in the first string array that do not meet the preset rules; Adjust the elements that do not meet the preset rules according to the preset rules; Generate a second text based on the adjusted first string array.
[0052] In a few cases, such as when the text is relatively long, the text generated based on the previous rules may not meet the requirement of the degree of discreteness. In this case, this defect can be overcome by the above method.
[0053] Taking an embodiment as an example, the text length is 300, the step size is set to 50, and the size of the first window is 100. First, set the offset to 0, add the characters from 0 to 100 in the string to an array, then set the offset to the current offset plus the step size, and add the characters from 50 to 150 in the string to an array; and so on, to obtain an array. Determine whether the proportion of the replaced words in this array meets the requirements. For the strings that do not meet the requirements, rewrite the strings to meet the preset rules. For example, when the first window is 50, in the text "Measure 25.0 mL of the prepared solution A and 25.0 mL of solution B, add solid C after dissolution, and make the volume up to the mark after uniform dissolution, and then at room temperature", the words "Measure", "uniform dissolution", and "prepared" are the replaced words, and their proportion is 20%, which is higher than the preset value of 0.10 - 0.18. Then rewrite it, such as adding a space after "solution A" and "solution B" to meet the requirements. The methods that can be adopted include segmenting, inserting spaces, inserting parentheses, and inserting explanatory text, such as modifying "25 mL" to "(25 mL)". One or more regular expressions are set here to achieve matching and replacement. According to an embodiment of the present invention, the adjustment includes splitting the string, adding explanatory text, or merging based on the text.
[0054] In this way, the adjustment of the degree of discreteness of the replaced words in the obtained second text can be achieved.
[0055] According to an embodiment of the present invention, a data processing system for a food detection platform, referring to Figure 2 the shown structure, the data processing system includes: A text rewriting unit, configured to determine a first keyword set used in the first text according to the input detection data, replace and rewrite the first keywords included in the first text using a keyword table to obtain a second text; A vector obtaining unit, configured to determine the vector corresponding to the text; The keyword table includes a key name and a set of synonyms corresponding to the key name; The similarity between the second vector corresponding to the second text and the first vector corresponding to the first text is greater than a first threshold; and the distribution of the replaced words in the second text meets the preset discrete value; The similarity between the word vector corresponding to the key name in the keyword table and the word vectors of the words in the set of synonyms corresponding to the key name is lower than a second threshold and higher than a third threshold.
[0056] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program runs on an electronic device, the electronic device is caused to execute the foregoing method.
[0057] The embodiments of the present application also provide a computer program product, including: computer program code. When the computer program code runs on an electronic device, the electronic device is caused to execute the foregoing method.
[0058] The embodiments of the present application also provide a chip, including: a processor, configured to call and run a computer program from a memory, so that an electronic device installed with the chip executes the foregoing method.
[0059] Through the description of the above embodiments, those skilled in the art can understand that, for the convenience and conciseness of description, only the above division of each functional module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0060] It should be understood that the devices and processes disclosed in several embodiments of the present application can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device. In addition, some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of devices or units may be in electrical, mechanical or other forms.
[0061] The units described as separate components may or may not be physically separated. The components displayed as units may be a physical unit or multiple physical units. That is, they can be located in one place, or can be distributed to multiple different places. The parts or all of the units can be selected according to actual needs to achieve the purpose of this solution.
[0062] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit; can also exist physically alone; can also be that some units are integrated in one unit and some units exist physically alone. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0063] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, all or part of the technical solutions of the embodiments of this application can be embodied in the form of a software product. This software product is stored in a storage medium. The software product includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0064] It should be noted that all or part of the above-mentioned various embodiments provided in this application (for example, part or all of any feature) can be arbitrarily combined or used in combination with each other.
[0065] As described above, the foregoing are only specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A data processing method for a food detection platform, characterized in that, Including: Determine a first keyword set used in the first text according to the input detection data; Determine a first vector corresponding to the first text; Rewrite the first text to obtain a second text; The similarity between the second vector corresponding to the second text and the first vector is greater than a first threshold; The rewriting includes replacing the corresponding synonyms in the first text with synonyms in a keyword table, and the distribution of the replaced words in the second text satisfies a preset degree of dispersion. The keyword table includes a key name and a set of synonyms corresponding to the key name; The similarity between the word vector corresponding to the key name in the keyword table and the word vectors of the synonyms in the corresponding synonym set is lower than a second threshold and higher than a third threshold.
2. The data processing method of a food detection platform according to claim 1, characterized in that, The first keyword set is obtained in the following manner: Perform named entity recognition on the first text to obtain a first named entity list; Sort the first named entities in descending order according to the frequency of occurrence of the named entities to obtain a first named entity list; Determine the named entities in the first named entity list that are the same as the key name in the keyword table as the first keyword set.
3. The data processing method of a food detection platform according to claim 2, wherein, In response to the first keyword set being empty, determine a second named entity list with no more than m named entities based on the frequency of occurrence of the named entities, where 2 ≤ m ≤ 10; Determine the second named entity list based on the similarity between the named entities in the first named entity list. The semantic similarity between any two named entities in the second named entity list is not higher than the second threshold; Obtain words similar to the named entities in the second named entity list from a database or a user-defined dictionary as alternative words. In response to the user's input, determine the associated words of the named entity selected or input by the user; Save the word input by the user to the keyword table and update it to the first keyword set.
4. The data processing method of a food detection platform according to claim 3, characterized in that In response to the user's input on the user interface, obtain one or more alternative words selected by the user; Filter the alternative words according to the similarity between the alternative words selected by the user and the named entity prompted by the user interface; Perform anomaly recognition based on the word vectors corresponding to the alternative words and remove the abnormal words; Use the alternative words after removing the abnormal words as the synonym set, use the named entity prompted by the user interface as the key name, and update the keyword table and the first keyword set.
5. The data processing method of a food detection platform according to claim 1, characterized in that, In response to the first text containing a synonym in the synonym set of the keyword table, replace the synonym in the first text with the key name corresponding to the element in the synonym set.
6. The data processing method of a food detection platform according to claim 1, characterized in that, The process of obtaining the second text includes: Replace the first keywords included in the first text with synonyms in the corresponding synonym table in the keyword table to obtain one or more candidate texts; Determine multiple different candidate texts based on the semantic similarity between the candidate texts and the first text; Traverse the candidate text set, and determine the candidate text that meets the preset degree of dispersion and has a similarity greater than the first threshold with the first vector as the second text.
7. The data processing method of a food detection platform according to claim 6, characterized in that Perform outlier analysis on the sentence vectors corresponding to multiple candidate texts and remove the candidate texts with an outlier index greater than 1.
8. The data processing method of a food detection platform according to claim 1, wherein, The data processing method further includes: Split the rewritten first text according to a first step length and a first window size to obtain a first string array; Determine the elements in the first string array that do not meet the preset rules; Adjust the elements that do not meet the preset rules according to the preset rules; Generate a second text based on the adjusted first string array.
9. The data processing method of a food detection platform according to claim 8, characterized in that, The adjustment includes splitting of strings, adding explanatory text, or merging based on text.
10. A data processing system for a food detection platform, characterized in that, Comprising: A text rewriting unit, configured to determine a first keyword set used in a first text according to input detection data, replace and rewrite the first keywords included in the first text using a keyword table to obtain a second text; A vector obtaining unit, configured to determine a vector corresponding to the text; The keyword table contains key names and a set of synonyms corresponding to the key names; The similarity between the second vector corresponding to the second text and the first vector corresponding to the first text is greater than a first threshold; And the distribution of the replaced words in the second text satisfies a preset discrete value; The similarity between the word vector corresponding to the key name in the keyword table and the word vectors of the words in the set of synonyms corresponding to the key name is lower than a second threshold and higher than a third threshold.