Text detection method and device, equipment and medium
By performing word segmentation, clustering and semantic matching on text, the problem of insufficient semantic understanding in text difference detection in the prior art is solved, and the accuracy and efficiency of the detection results are improved.
Patent Information
- Application Number
- CN202311394869.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2025-05-02
AI Technical Summary
The prior art cannot accurately understand the semantic differences between two texts in text difference detection, resulting in inaccurate detection results and increasing unnecessary inspection workload.
By performing word segmentation, clustering and element type annotation on the detected text and the target text, semantic matching is used for clustered word segmentation to obtain the semantic detection results of the text.
It improves the accuracy of text semantic detection results, reduces the difficulty of determining whether there are semantic differences in text, and reduces unnecessary inspection work.
Smart Images

Figure CN119918525A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a text detection method, device, equipment and medium. Background Art
[0002] When a user uses a service (such as a navigation service) provided by an application software (hereinafter referred to as an application), the application can deliver information that the user needs to pay attention to to the user through voice broadcast or text display. Usually, to realize the function of voice broadcast or text display of information, the application developer needs to pre-prepare the text content. When the application is running, the voice broadcast engine or text rendering engine can broadcast the text content to the user by voice, or display the text content to the user in the form of text.
[0003] In order to continuously improve the product experience, application developers will iterate the application. The inventors of the present disclosure have found that after the application iteration, the text content formulated by the application developer may change compared to before the iteration. In order to ensure the product experience of the application after the iteration, it is necessary to perform text difference detection on the text content before and after the iteration to determine the impact of the difference on the product experience. In the related art, the difference comparison is performed through a text difference comparison algorithm, such as the Mayers algorithm, or a text similarity or distance algorithm, such as the Euclidean distance, cosine similarity, minimum edit distance, etc. However, the related art has the problem of being unable to accurately understand whether there is a semantic difference between the two texts, thereby reducing the accuracy of the detection results. Summary of the invention
[0004] In order to solve the problems in the related art, the embodiments of the present disclosure provide a text detection method, apparatus, device and medium.
[0005] In a first aspect, an embodiment of the present disclosure provides a text detection method, comprising:
[0006] Get the text to be detected and the target text;
[0007] Perform word segmentation on the text to be detected and the target text respectively to obtain the word segmentation of the text to be detected and the word segmentation of the target text;
[0008] Cluster the words of the text to be detected and label the element types of the clustered words;
[0009] Cluster the target text's segmented words and label the element types of the clustered segmented words;
[0010] Based on the element type of the cluster segmentation annotation, the cluster segmentation of the to-be-detected text and the cluster segmentation of the target text are semantically matched to obtain the cluster segmentation matching result of the to-be-detected text and the target text;
[0011] Based on the clustering and word segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained.
[0012] In one embodiment of the present disclosure, the text to be detected and the target text are navigation guidance speech texts, and the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text, including:
[0013] Based on the pre-defined navigation guidance speech template, the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text, and the segmentation is a word or a single character that cannot be divided any further.
[0014] In one embodiment of the present disclosure, the navigation guidance speech template includes at least one of a main action speech template, an auxiliary action speech template, a road direction speech template, and a traffic information speech template.
[0015] In one embodiment of the present disclosure, clustering the segmented words of the text to be detected and labeling the element type of the clustered segmented words includes:
[0016] Determine the element type of the word segmentation of the text to be detected through the pre-defined element dictionary and element rules;
[0017] According to the word order of the segmented words in the text to be detected, adjacent segmented words with the same element type are clustered into a clustered segmented word and the element type is marked.
[0018] In one embodiment of the present disclosure, the element type includes at least one of driving action, driving distance, driving speed, driving time and electronic eye.
[0019] In one embodiment of the present disclosure, based on the element type of the cluster segmentation annotation, the cluster segmentation of the to-be-detected text and the cluster segmentation of the target text are semantically matched to obtain the cluster segmentation matching result of the to-be-detected text and the target text, including:
[0020] The element types and word orders in the clustered word segments of the to-be-detected text and the clustered word segments of the target text meet the pre-established semantic matching rules, and are determined as the clustered word segments that match the to-be-detected text and the target text;
[0021] For the clustered word segments of the text to be detected that cannot be matched with the clustered word segments of the target text through the semantic matching rules, semantic matching is performed through a preset text matching algorithm to obtain the corresponding clustered word segmentation matching results.
[0022] In one embodiment of the present disclosure, based on the cluster segmentation matching results of the text to be detected and the target text, obtaining the semantic detection results of the text to be detected and the target text includes:
[0023] Based on the cluster segmentation matching result, if it is determined that the cluster segmentation of the text to be detected and the cluster segmentation of the target text do not match, then the semantic detection results of the text to be detected and the target text are determined to be semantically different;
[0024] If it is determined that the clustered word segmentation of the text to be detected matches the clustered word segmentation of the target text, or the proportion of the number of mutually matching clustered word segmentations exceeds a preset threshold, then the semantic detection results of the text to be detected and the target text are determined to be semantically indistinguishable;
[0025] If it is determined that the proportion of the number of mutually matching clustered word segments is lower than a preset threshold, the semantic difference of the mutually non-matching clustered word segments is determined according to a predefined semantic difference rule.
[0026] In a second aspect, an embodiment of the present disclosure provides a text detection device, which includes:
[0027] A text acquisition module is configured to acquire the text to be detected and the target text;
[0028] The word segmentation module is configured to segment the text to be detected and the target text respectively to obtain the word segmentation of the text to be detected and the word segmentation of the target text;
[0029] The to-be-detected clustering module is configured to cluster the word segments of the to-be-detected text and label the element types of the clustered word segments;
[0030] A target clustering module is configured to cluster the word segments of the target text and label the element types of the clustered word segments;
[0031] The semantic matching module is configured to perform semantic matching on the clustered word segmentation of the to-be-detected text and the clustered word segmentation of the target text based on the element type of the clustered word segmentation annotation, and obtain the clustered word segmentation matching result of the to-be-detected text and the target text;
[0032] The semantic detection module is configured to obtain semantic detection results of the text to be detected and the target text based on the clustering word segmentation matching results of the text to be detected and the target text.
[0033] In a third aspect, an embodiment of the present disclosure provides an electronic device, comprising a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement a method as described in any one of the first aspect or any one of the embodiments of the first aspect.
[0034] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, a method as described in any one of the first aspect or any one of the embodiments of the first aspect is implemented.
[0035] According to the technical solution provided by the embodiments of the present disclosure, a text to be detected and a target text are obtained; the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text; the segmentation of the text to be detected and the target text are clustered, and the element types of the clustered segmentations are respectively marked; based on the element types marked by the clustered segmentations, the clustered segmentations of the text to be detected and the clustered segmentations of the target text are semantically matched to obtain the clustered segmentation matching results of the text to be detected and the target text; based on the clustered segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained. Among them, since the element type of the cluster segmentation of the corresponding text can be understood as indicating which type of information the cluster segmentation is used to express, the cluster segmentation of the text to be detected and the cluster segmentation of the target text are semantically matched based on the element type marked by the cluster segmentation. The obtained cluster segmentation matching result can be used to determine the cluster segmentations expressing the same type of information (i.e., the semantics are relatively close) in the text to be detected and the target text. When the semantic detection results of the text to be detected and the target text are obtained based on the cluster segmentation matching result, it can be determined whether there is a semantic difference between the two texts based on whether there is a difference in the cluster segmentations used to express the same type of information in the two texts. This reduces the difficulty of determining whether there is a semantic difference between the text to be detected and the target text, and improves the accuracy of the semantic detection results.
[0036] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Other features, objectives and advantages of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments in conjunction with the accompanying drawings. In the accompanying drawings:
[0038] Figure 1 A flowchart of a text detection method according to an embodiment of the present disclosure is shown.
[0039] Figure 2 A flowchart of a text detection method according to an embodiment of the present disclosure is shown.
[0040] Figure 3 A structural block diagram of a text detection device according to an embodiment of the present disclosure is shown.
[0041] Figure 4 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0042] Figure 5 A schematic diagram showing the structure of a computer system suitable for implementing the method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0043] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts not related to the description of the exemplary embodiments are omitted in the accompanying drawings.
[0044] In the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate the presence of features, numbers, steps, behaviors, components, parts, or a combination thereof disclosed in the present specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, behaviors, components, parts, or a combination thereof exist or are added.
[0045] It should also be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0046] In the present disclosure, if it involves operations of obtaining user information or user data or displaying user information or user data to others, the operations are all authorized and confirmed by the user, or actively selected by the user.
[0047] In practice, in order to continuously improve the product experience, application developers will iterate (upgrade) the application. If the application requires the application developer to formulate text content, the text content may change after the application iteration compared with before the iteration. In order to ensure the product experience of the application after the iteration, it is necessary to perform text difference detection on the text content before and after the iteration to determine the impact of the difference on the product experience. Taking the application with map navigation function as an example, the developer needs to formulate the navigation guidance text. When the user uses the navigation function of the application, the application can broadcast the navigation guidance text to the user in the form of voice broadcast through the speaker of the terminal device. After the application is iterated, the description of the navigation guidance text may be different from that before the iteration. If the difference is too large, it will cause the user to be unable to quickly understand the navigation guidance text after the iteration, affecting driving safety. Therefore, the developer needs to check the navigation text with differences. In the related art, the difference comparison is performed by a text difference comparison algorithm, such as the Mayers algorithm, or a text similarity or distance algorithm, such as Euclidean distance, cosine similarity, minimum edit distance, etc. However, the inventors have discovered that these solutions can be understood as: physically calculating the difference or similarity between the two detected texts, without understanding the difference between the two texts semantically. That is, if the semantics of the two detected texts are actually the same, but the description methods are different, then based on the existing solutions, a detection result indicating that the two detected texts are significantly different may be obtained, and the detection result is inaccurate. When the detection result is inaccurate, for the aforementioned map navigation scenario, it will cause developers to increase unnecessary inspection workload.
[0048] In order to solve the above problems, the embodiments of the present disclosure provide a text detection method, apparatus, device and medium.
[0049] According to the technical solution provided by the embodiments of the present disclosure, a text to be detected and a target text are obtained; the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text; the segmentation of the text to be detected and the target text are clustered, and the element types of the clustered segmentations are respectively marked; based on the element types marked by the clustered segmentations, the clustered segmentations of the text to be detected and the clustered segmentations of the target text are semantically matched to obtain the clustered segmentation matching results of the text to be detected and the target text; based on the clustered segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained. Among them, since the element type of the cluster segmentation of the corresponding text can be understood as indicating which type of information the cluster segmentation is used to express, the cluster segmentation of the text to be detected and the cluster segmentation of the target text are semantically matched based on the element type marked by the cluster segmentation. The obtained cluster segmentation matching result can be used to determine the cluster segmentations expressing the same type of information (i.e., the semantics are relatively close) in the text to be detected and the target text. When the semantic detection results of the text to be detected and the target text are obtained based on the cluster segmentation matching result, it can be determined whether there is a semantic difference between the two texts based on whether there is a difference in the cluster segmentations used to express the same type of information in the two texts. This reduces the difficulty of determining whether there is a semantic difference between the text to be detected and the target text, and improves the accuracy of the semantic detection results.
[0050] Figure 1 FIG. 1 is a flowchart of a text detection method according to an embodiment of the present disclosure. Figure 1 As shown, the text detection method includes the following steps S101-S106:
[0051] In step S101, the text to be detected and the target text are obtained;
[0052] In step S102, the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text;
[0053] In step S103, the word segments of the text to be detected are clustered, and the element types are marked on the clustered word segments;
[0054] In step S104, the word segments of the target text are clustered, and the element types are labeled for the clustered word segments;
[0055] In step S105, based on the element type marked by the cluster segmentation, semantic matching is performed on the cluster segmentation of the text to be detected and the cluster segmentation of the target text to obtain the cluster segmentation matching result of the text to be detected and the target text;
[0056] In step S106, based on the clustering word segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained.
[0057] In one implementation of the present disclosure, the text to be detected can be understood as the text after iteration, and the target text can be understood as the text before iteration. The text to be detected and the target text can be used for prompting. Exemplarily, an application with a map navigation function is used as an example for explanation, in which the text to be detected is the navigation guidance text after iteration, and the target text is the navigation guidance text before iteration.
[0058] In an implementation of the present disclosure, performing word segmentation on the text to be detected and the target text respectively can be understood as performing word segmentation on the text to be detected and the target text respectively according to a pre-acquired word segmentation algorithm or word segmentation model.
[0059] Exemplarily, taking the example that both the text to be detected and the target text are navigation guidance texts, if the text to be detected or the target text is "drive right front after nine hundred meters", after word segmentation, the text will be divided into "five hundred", "meters", "after", "keep left", "drive", "ahead", "traffic light intersection", and "go straight". Since some of the words in the word segmentation results are meaningless, the present disclosure needs to cluster the word segmentations. Continuing with the previous example, the clustered word segmentations obtained by clustering include: "five hundred meters", "after", "drive left", "ahead", "traffic light intersection", and "go straight".
[0060] In one implementation of the present disclosure, the element type can be understood as the type of information expressed by the corresponding cluster segmentation. When the to-be-detected text and the target text are both navigation guidance texts, the element type can include at least one of driving action, driving distance, driving speed, driving time, or electronic eyes. Exemplarily, the driving action can also be subdivided into primary driving action and secondary driving action.
[0061] If the cluster segmentation words are turn left, turn right, U-turn, etc., then the element type of the cluster segmentation words can be marked as the main driving action in the driving action; if the cluster segmentation words are go up the elevated road, go down the elevated road, enter the main road, etc., then the element type of the cluster segmentation words can be marked as the auxiliary driving action in the driving action. If the cluster segmentation is a phrase consisting of distance numerical words and distance unit words (such as meters, kilometers, etc.), or the cluster segmentation is a phrase consisting of any one of ahead, immediately, about to, in advance and distance numerical words and distance unit words, then the element type of the cluster segmentation can be marked as driving distance; if the cluster segmentation is a phrase consisting of speed numerical words and speed unit words (such as kilometers per hour, kilometers per hour, meters per minute, etc.), or if the cluster segmentation is a phrase consisting of any one of vehicle speed, speed limit, speed and speed numerical words and speed unit words, then the element type of the cluster segmentation can be marked as driving speed; if the cluster segmentation is a phrase consisting of time numerical words and time unit words (such as minutes, hours, seconds, etc.), or if any cluster segmentation is a phrase consisting of any one of morning, afternoon, morning, evening and time numerical words and time unit words, then the element type of the cluster segmentation can be marked as driving time; if the cluster segmentation is any one of red light running photo taking and illegal parking photo taking, then the element type of the cluster segmentation can be marked as electronic eye.
[0062] In one implementation of the present disclosure, labeling the element type for cluster segmentation can be implemented by substituting the corresponding cluster segmentation into the element recognition algorithm for calculation based on a pre-acquired element recognition algorithm, and labeling the element type for the cluster segmentation according to the calculation result; it can also be implemented by using the corresponding cluster segmentation as the input of the element recognition model based on a pre-acquired element recognition model (such as a Markov model, a Bayesian network model, a conditional random field model, etc.), outputting the element recognition result through the element recognition model, and labeling the element type for the cluster segmentation according to the element recognition result.
[0063] In one implementation of the present disclosure, semantic matching is performed on the clustered segmentation of the text to be detected and the clustered segmentation of the target text based on the element type annotated by the clustered segmentation. This can be implemented as an algorithm obtained according to a budget, such as the Text-Match algorithm, to semantically match the clustered segmentation of the text to be detected and the clustered segmentation of the target text based on the element type annotated by the clustered segmentation to obtain the clustered segmentation matching result. It can also be implemented as obtaining a pre-trained semantic matching model, taking the clustered segmentations in the text to be detected and the target text and the element type annotated by each clustered segmentation as input, and inputting the semantic matching model to obtain the clustered segmentation matching result output by the semantic matching model.
[0064] In one implementation of the present disclosure, based on the clustering word segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained. It can be implemented as based on the clustering word segmentation matching results. If the semantically mismatched parts in the clustering word segmentation of the text to be detected and the clustering word segmentation of the target text are determined, and the semantic detection results of the text to be detected and the target text are obtained based on the semantically mismatched parts, the semantic detection results can be understood as indicating whether there are differences in content between the text to be detected and the target text. Furthermore, when there are differences in content between the text to be detected and the target text, the semantic detection results can also be understood as indicating the parts where there are differences in content between the text to be detected and the target text.
[0065] According to the technical solution provided by the embodiments of the present disclosure, a text to be detected and a target text are obtained; the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text; the segmentation of the text to be detected and the target text are clustered, and the element types of the clustered segmentations are respectively marked; based on the element types marked by the clustered segmentations, the clustered segmentations of the text to be detected and the clustered segmentations of the target text are semantically matched to obtain the clustered segmentation matching results of the text to be detected and the target text; based on the clustered segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained. Among them, since the element type of the cluster segmentation of the corresponding text can be understood as indicating which type of information the cluster segmentation is used to express, the cluster segmentation of the text to be detected and the cluster segmentation of the target text are semantically matched based on the element type marked by the cluster segmentation. The obtained cluster segmentation matching result can be used to determine the cluster segmentations expressing the same type of information (i.e., the semantics are relatively close) in the text to be detected and the target text. When the semantic detection results of the text to be detected and the target text are obtained based on the cluster segmentation matching result, it can be determined whether there is a semantic difference between the two texts based on whether there is a difference in the cluster segmentations used to express the same type of information in the two texts. This reduces the difficulty of determining whether there is a semantic difference between the text to be detected and the target text, and improves the accuracy of the semantic detection results.
[0066] In one embodiment of the present disclosure, the text to be detected and the target text are navigation guidance speech texts, and the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text, including:
[0067] Based on the pre-defined navigation guidance speech template, the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text, and the segmentation is a word or a single character that cannot be divided any further.
[0068] In one implementation of the present disclosure, a navigation guidance speech template is pre-formulated, which can be implemented by pre-segmenting the collected text and counting the word frequency of each word after the segmentation; selecting a plurality of representative words (such as start, whole process, purpose, after, prepare, subsequently, please, pay attention, keep, enter, etc.) according to the word frequency of each word combined with prior knowledge, and based on the plurality of representative words, using a One-Hot method to vectorize each text; after obtaining the vectorized representation of each text, using a clustering algorithm to perform unsupervised clustering on the collected text, and corresponding each clustered category to a navigation guidance speech template.
[0069] In one implementation of the present disclosure, the navigation guidance speech template can be understood as indicating what content should be included in the navigation guidance speech text and the position of the corresponding content in the navigation guidance speech text.
[0070] In one embodiment of the present disclosure, the navigation guidance speech template may include a main action speech template, an auxiliary action speech template, a road direction speech template, etc. The navigation guidance speech template may be set by those skilled in the art as required, and the present disclosure cannot be exhaustive, and the above examples shall not be regarded as limiting the present disclosure.
[0071] The composition structure of the words indicated by the main action speech template may include: basic words of orientation + distance information words + connecting words + ordinal words + traffic facility words + main action words. For the main action speech template, the basic words of orientation can be understood as words used to describe directions, such as front, left front, right front, etc.; the distance information words can include distance quantity words (such as five hundred, one and a half, etc.) and distance unit words (such as meters, kilometers, etc.); the connecting words can be understood as words that do not represent any actual content and are used to connect distance information words and ordinal words, such as in, for, etc.; the ordinal words can be understood as words used to describe order, such as first, second, third, and each, etc.; the traffic facility words can be understood as words used to describe traffic facilities, traffic signs, etc. on the road, such as traffic lights, speed limit signs, crosswalks, intersections, etc.; the main action words can be understood as words used to describe driving actions that users may make, such as turn left, turn right, drive on the right, drive on the left, turn around, etc.
[0072] For example, if the navigation guidance words are: Turn right at the second traffic light intersection five hundred meters ahead, then the word segmentation of the navigation guidance words based on the main action words template may be: (ahead|basic direction word)(five hundred|distance quantifier)(meter|distance unit word)(in|connection word)(second|ordinal number)(piece|ordinal number)(traffic light|traffic facility phrase)(intersection|traffic facility phrase)(turn right|main action word).
[0073] The composition structure of the words indicated by the auxiliary action speech template may include: reminder words + prepositions + ordinal words + traffic facility words + driving action words. For the auxiliary action speech template, the reminder words can be understood as words that can attract the user's attention and thus remind the user, such as please, don't, etc.; the prepositions can include in, from, to, along, through, through, etc.; the ordinal words can be understood as words used to describe the order, such as the first, second, third, and so on; the traffic facility words can be understood as words used to describe traffic facilities, traffic signs, etc. on the road, such as traffic lights, speed limit signs, crosswalks, intersections, etc.; the driving action words can be understood as words used to describe the driving actions that the user may make, such as turn left, turn right, drive on the right, drive on the left, turn around, etc. It should be noted that compared with the main action speech template, the importance of reminding the user to make or prohibit making the corresponding driving action in the navigation guidance speech matching the auxiliary action speech template is generally lower than the importance of reminding the user to make or prohibit making the corresponding driving action in the navigation guidance speech matching the main action speech template.
[0074] Exemplarily, if the navigation guidance speech is: Please exit from the second exit, then the segmentation results obtained by segmenting the navigation guidance speech based on the auxiliary action speech template may be: ((please | reminder word) + (from | preposition) + (second | ordinal number) + (exit | traffic facility word) + (exit | driving action word); if the navigation guidance speech is: Do not drive in the second lane, then the segmentation results obtained by segmenting the navigation guidance speech based on the auxiliary action speech template may be: (do not | reminder word) + (in | preposition) + (second | ordinal number) + (lane | traffic facility word) + (drive | driving action word).
[0075] The word structure indicated by the road direction speech template includes: "towards" + road information words + "direction". For the road direction speech template, the road information words can be understood as words used to express the traffic conditions on the road, which can include at least one of the road nouns, road location words, and traffic condition words.
[0076] For example, if the navigation guidance words are: to Yuetan South Street, Guangningbo Street exit direction, then the word segmentation obtained by segmenting the navigation guidance words based on the road direction words template can be: (to) + (Yuetan South Street | road noun) + (Guangningbo Street exit | road location word) + (direction); if the navigation guidance words are: to Huanzhou 3rd Road direction, then the word segmentation obtained by segmenting the navigation guidance words based on the road direction words template can be: (to) + (Huanzhou 3rd Road | road noun) + (direction).
[0077] The composition structure of the words indicated by the traffic information speech template includes: road noun + connecting word + distance information word + traffic information word. For the traffic information speech template, the road noun can be understood as a word that can identify the corresponding road; the connecting word can be understood as a word that is used to connect the distance information word and the ordinal word and does not represent any actual content, such as "have", "appear", etc.; the distance information word can include distance quantity words and distance unit words; the traffic information word can be understood as a word used to indicate specific road conditions, such as congestion, traffic accident, smooth, etc.
[0078] For example, if the navigation guidance words are: There is a 500-meter congestion on Zhichunli Road, then the word segmentation obtained by segmenting the navigation guidance words based on the road condition information words template can be (Zhichunli Road|road name)+(there|connecting word)+(five hundred|distance quantifier)+(meter|distance unit word)+(congestion|road condition information).
[0079] In one implementation of the present disclosure, if the text to be detected and the target text are separately segmented based on a pre-established navigation guidance speech template, and the segmentation of the text to be detected and the target text cannot be obtained, a pre-trained segmentation model can be obtained, such as a hidden Markov model (HMM), a conditional random field (CRF), a long short-term memory network (LSTM), a bidirectional encoder representation from transformers (BERT) model, etc., and the text to be detected and the target text are separately segmented based on the segmentation model.
[0080] According to the technical solution provided by the embodiment of the present disclosure, by limiting the text to be detected and the target text to be navigation guidance speech text, and by segmenting the text to be detected and the target text respectively based on a pre-established navigation guidance speech template, the segmentation of the text to be detected and the segmentation of the target text are obtained, which can improve the accuracy of the segmentation of the text to be detected and the target text, help cluster the segmentations of the text to be detected, and improve the accuracy of the element type of the clustered segmentations.
[0081] In one embodiment of the present disclosure, clustering the segmented words of the text to be detected and labeling the element type of the clustered segmented words includes:
[0082] Determine the element type of the word segmentation of the text to be detected through the pre-defined element dictionary and element rules;
[0083] According to the word order of the segmented words in the text to be detected, adjacent segmented words with the same element type are clustered into a clustered segmented word and the element type is marked.
[0084] In one implementation of the present disclosure, an element dictionary can be understood as a correspondence between fixed or rarely changing words and element types. Exemplarily, the element dictionary can be used to indicate that the element types corresponding to words such as turn left, turn right, drive on the right, keep left, to the right front, to the left rear, turn around, enter the roundabout, etc. are the main driving actions in the driving action; the element dictionary can also be used to indicate that the element types corresponding to words such as turn right around the roundabout, go up the elevated road, enter the small road, go down the elevated road, go up the bridge, enter the main road, and enter the community road are the auxiliary driving actions in the driving action.
[0085] In one implementation of the present disclosure, the element rule can be understood as the correspondence between the word order of the segmentation, at least one of the word classes and the element type. For example, the element type corresponding to the segmentation consisting of a distance numerical word and a distance unit word can be the driving distance; the element type corresponding to the segmentation consisting of a speed numerical word and a speed unit word can be the driving speed.
[0086] For example, if the text to be detected is: There is a 500-meter congestion on Zhichunli Road, the word segmentation of the text to be detected is: Zhichunli Road + There is + 500 + Meters + Congestion, where the element types of "500" and "Meters" in the word segmentation of the text to be detected are determined through the pre-established element dictionary and element rules. Both belong to driving distance. Then, "500" and "Meters" can be aggregated into a clustered word segmentation, and the element type of "500 meters" can be marked as driving distance.
[0087] According to the technical solution provided by the embodiment of the present disclosure, the element type of the word segmentation of the text to be detected is determined by means of a pre-established element dictionary and element rules; according to the word order of the word segmentation of the text to be detected in the text to be detected, adjacent word segmentations with the same element type are clustered into a clustered word segmentation and the element type is marked, which can improve the accuracy of marking the element type of the clustered word segmentation and help improve the accuracy of the semantic detection results.
[0088] In one embodiment of the present disclosure, clustering the word segments of the target text and labeling the element types of the clustered word segments include:
[0089] Determine the element type of the target text segmentation through the pre-established element dictionary and element rules;
[0090] According to the word order of the target text's segmented words in the text to be detected, adjacent segmented words with the same element type are clustered into a clustered segmented word and the element type is marked.
[0091] According to the technical solution provided by the embodiment of the present disclosure, the element type of the target text's word segmentation is determined by a pre-established element dictionary and element rules; according to the word order of the target text's word segmentation in the target text, adjacent word segmentations with the same element type are clustered into a clustered word segmentation and the element type is labeled, which can improve the accuracy of labeling the element type for the clustered word segmentation and help improve the accuracy of the semantic detection results.
[0092] In one embodiment of the present disclosure, based on the element type of the cluster segmentation annotation, the cluster segmentation of the to-be-detected text and the cluster segmentation of the target text are semantically matched to obtain the cluster segmentation matching result of the to-be-detected text and the target text, including:
[0093] The element types and word orders in the clustered word segments of the to-be-detected text and the clustered word segments of the target text meet the pre-established semantic matching rules, and are determined as the clustered word segments that match the to-be-detected text and the target text;
[0094] For the clustered word segments of the text to be detected that cannot be matched with the clustered word segments of the target text through the semantic matching rules, semantic matching is performed through a preset text matching algorithm to obtain the corresponding clustered word segmentation matching results.
[0095] In one implementation of the present disclosure, the word order of the cluster segmentation conforms to the pre-established semantic matching rules, which can be understood as the relative position of other cluster segmentations whose element types meet the corresponding requirements in the corresponding text conforms to the pre-established semantic matching rules.
[0096] For example, the text to be detected is: driving to the right front, entering the ramp, uphill, there is a camera on the ramp, the speed limit is 40 kilometers per hour, 500 meters ahead, driving towards Yan'an Elevated Road, Hongqiao Hub, and the clustered word segmentation of the target text is: driving into the ramp on the right front, uphill, 500 meters later along Yan'an Elevated Road, driving towards Hongqiao Hub. " Among them, the clustered word segmentation of the text to be detected and the corresponding element type can be: (driving to the right front|driving action) + (entering the ramp|driving action) + (uphill|driving action) + (there is a camera on the ramp Head, speed limit 40 kilometers per hour | electronic eye) + (five hundred meters ahead | driving distance) + (driving to Yan'an Elevated Road, Hongqiao Hub direction | driving action), the clustered word segmentation of the target text and the corresponding element type can be: (enter the ramp in the right front | driving action) + (uphill | driving action) + (five hundred meters later | driving distance) + (driving along Yan'an Elevated Road, Hongqiao Hub direction | driving action). By referring to the element types and word order in the clustered word segmentation of the text to be detected and the clustered word segmentation of the target text, it can be determined that the " It is determined that "drive to the right front and enter the ramp" matches with "drive to the right front into the ramp" in the target text, and it is determined that "uphill" in the text to be detected matches with "uphill" in the target text. It is determined that "five hundred meters ahead, drive toward Yan'an Elevated Road, Hongqiao Hub direction" in the text to be detected matches with "five hundred meters later along Yan'an Elevated Road, Hongqiao Hub direction" in the target text. "There is a camera on the ramp, the speed limit is 40 kilometers per hour" in the text to be detected does not match any clustered word segmentation in the target text. At the same time, the text to be detected includes a clustered word segmentation "There is a camera on the ramp, the speed limit is 40 kilometers per hour|electronic eye" that does not match any clustered word segmentation in the target text. The element type corresponding to the unmatched clustered word segmentation is electronic eye. Since the clustered word segmentations that affect the user's driving choice are all matched, the electronic eye is the broadcast content that will not affect the driving choice, and the proportion of matched clustered word segmentations is also very high. Therefore, it can be determined that the semantic detection result of the text to be detected and the target text does not have a semantic difference. For texts without semantic differences, no further review is required, which reduces the review workload.
[0097] According to the technical solution provided by the embodiments of the present disclosure, the text to be detected and the target text are determined to be clustered word segmentations that match each other if the element types and word orders in the clustered word segmentations of the text to be detected and the target text meet the pre-established semantic matching rules; for the clustered word segmentations of the text to be detected and the target text that cannot be matched through the semantic matching rules, semantic matching is performed through a preset text matching algorithm to obtain corresponding clustered word segmentation matching results, which can improve the accuracy of the obtained clustered word segmentation matching results and help improve the accuracy of the semantic detection results obtained based on the clustered word segmentation matching results.
[0098] In one embodiment of the present disclosure, based on the cluster segmentation matching results of the text to be detected and the target text, obtaining the semantic detection results of the text to be detected and the target text includes:
[0099] Based on the cluster segmentation matching result, if it is determined that the cluster segmentation of the text to be detected and the cluster segmentation of the target text do not match, then the semantic detection results of the text to be detected and the target text are determined to be semantically different;
[0100] If it is determined that the clustered word segmentation of the text to be detected matches the clustered word segmentation of the target text, or the proportion of the number of mutually matching clustered word segmentations exceeds a preset threshold, then the semantic detection results of the text to be detected and the target text are determined to be semantically indistinguishable;
[0101] If it is determined that the proportion of the number of mutually matching clustered word segments is lower than a preset threshold, the semantic difference of the mutually non-matching clustered word segments is determined according to a predefined semantic difference rule.
[0102] Exemplarily, based on the cluster segmentation matching result, if it is determined that the proportion of cluster segmentations of the to-be-detected text and the target text that match each other is lower than a preset threshold, then the cluster segmentations of the to-be-detected text and the target text that do not match can be queried in the previously acquired whitelist. When it is determined according to the query result that the whitelist includes any of the unmatched cluster segmentations, it can be understood that there is a large difference that cannot be ignored between the content in the to-be-detected text and the content in the target text, and the two texts are different semantically, so it can be determined that there is a semantic difference between the to-be-detected text and the target text. Among them, the cluster segmentations in the above whitelist can be understood as cluster segmentations that have a strong semantic impact on the text, and the cluster segmentations not saved in the whitelist can be understood as cluster segmentations that have a weaker semantic impact on the text.
[0103] According to the technical solution provided by the embodiments of the present disclosure, based on the cluster segmentation matching results, if it is determined that the cluster segmentation of the text to be detected and the cluster segmentation of the target text do not match, then the semantic detection results of the text to be detected and the target text are determined to have semantic differences; if it is determined that the cluster segmentation of the text to be detected and the cluster segmentation of the target text are matched, or the proportion of the number of mutually matching cluster segmentations exceeds a preset threshold, then the semantic detection results of the text to be detected and the target text are determined to have no semantic differences; if it is determined that the proportion of the number of mutually matching cluster segmentations is lower than a preset threshold, then the semantic differences of the mutually non-matching cluster segmentations are determined according to a pre-defined semantic difference rule, which can effectively improve the accuracy of determining whether there are semantic differences in the semantic detection results of the text to be detected and the target text.
[0104] Figure 2FIG. 1 is a flowchart of a text detection method according to an embodiment of the present disclosure. Figure 1 As shown, the text detection method includes the following steps S201-S210:
[0105] In step S201, a text to be detected and a target text are obtained, where the text to be detected and the target text are navigation guidance speech texts;
[0106] In step S202, based on the pre-defined navigation guidance speech template, the to-be-detected text and the target text are segmented to obtain the segmentation of the to-be-detected text and the segmentation of the target text;
[0107] The navigation guidance speech template includes at least one of a main action speech template, an auxiliary action speech template, a road direction speech template, and a road condition information speech template, and the segmented words are words or single characters that cannot be further divided;
[0108] In step S203, the element types of the segmented words in the to-be-detected text and the target text are determined respectively by using the pre-defined element dictionary and element rules; the element types include at least one of driving action, driving distance, driving speed, driving time and electronic eye;
[0109] In step S204, according to the word order of the segmented words in the text to be detected, adjacent segmented words with the same element type are clustered into a clustered segmented word and the element type is marked;
[0110] In step S205, according to the word order of the target text segmentation words in the target text, the adjacent segmentation words with the same element type are clustered into a clustered segmentation word and the element type is marked;
[0111] In step S206, the clustered word segments of the to-be-detected text and the target text whose element types and word orders meet the pre-established semantic matching rules are determined as the clustered word segments that match the to-be-detected text and the target text;
[0112] In step S207, for the clustered word segments of the to-be-detected text that cannot be matched with the clustered word segments of the target text by the semantic matching rule, semantic matching is performed by a preset text matching algorithm to obtain a corresponding clustered word segmentation matching result;
[0113] In step S208, based on the cluster segmentation matching result, if it is determined that the cluster segmentation of the text to be detected and the cluster segmentation of the target text do not match, then it is determined that the semantic detection result of the text to be detected and the target text is that there is a semantic difference;
[0114] In step S209, if it is determined that the clustered word segmentation of the to-be-detected text matches the clustered word segmentation of the target text, or the proportion of the number of mutually matching clustered word segmentations exceeds a preset threshold, it is determined that the semantic detection result of the to-be-detected text and the target text is semantically indistinguishable;
[0115] In step S210, if it is determined that the proportion of the number of mutually matching cluster segmentations is lower than a preset threshold, the semantic difference of mutually non-matching cluster segmentations is determined according to a predefined semantic difference rule.
[0116] Figure 3 The structural block diagram of the text detection device according to the embodiment of the present disclosure is shown. The device can be implemented as part or all of the electronic device through software, hardware or a combination of both.
[0117] like Figure 3 As shown, the text detection device 300 includes:
[0118] A text acquisition module is configured to acquire the text to be detected and the target text;
[0119] The word segmentation module is configured to segment the text to be detected and the target text respectively to obtain the word segmentation of the text to be detected and the word segmentation of the target text;
[0120] The to-be-detected clustering module is configured to cluster the word segments of the to-be-detected text and label the element types of the clustered word segments;
[0121] A target clustering module is configured to cluster the word segments of the target text and label the element types of the clustered word segments;
[0122] The semantic matching module is configured to perform semantic matching on the clustered word segmentation of the to-be-detected text and the clustered word segmentation of the target text based on the element type of the clustered word segmentation annotation, and obtain the clustered word segmentation matching result of the to-be-detected text and the target text;
[0123] The semantic detection module is configured to obtain semantic detection results of the text to be detected and the target text based on the clustering word segmentation matching results of the text to be detected and the target text.
[0124] According to the technical solution provided by the embodiments of the present disclosure, a text to be detected and a target text are obtained; the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text; the segmentation of the text to be detected and the target text are clustered, and the element types of the clustered segmentations are respectively marked; based on the element types marked by the clustered segmentations, the clustered segmentations of the text to be detected and the clustered segmentations of the target text are semantically matched to obtain the clustered segmentation matching results of the text to be detected and the target text; based on the clustered segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained. Among them, since the element type of the cluster segmentation of the corresponding text can be understood as indicating which type of information the cluster segmentation is used to express, the cluster segmentation of the text to be detected and the cluster segmentation of the target text are semantically matched based on the element type marked by the cluster segmentation. The obtained cluster segmentation matching result can be used to determine the cluster segmentations expressing the same type of information (i.e., the semantics are relatively close) in the text to be detected and the target text. When the semantic detection results of the text to be detected and the target text are obtained based on the cluster segmentation matching result, it can be determined whether there is a semantic difference between the two texts based on whether there is a difference in the cluster segmentations used to express the same type of information in the two texts. This reduces the difficulty of determining whether there is a semantic difference between the text to be detected and the target text, and improves the accuracy of the semantic detection results.
[0125] The present disclosure also discloses an electronic device, Figure 4 A structural block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0126] like Figure 4 As shown, the electronic device includes a memory and a processor, wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method according to an embodiment of the present disclosure.
[0127] The present disclosure provides a text detection method, including:
[0128] Get the text to be detected and the target text;
[0129] Perform word segmentation on the text to be detected and the target text respectively to obtain the word segmentation of the text to be detected and the word segmentation of the target text;
[0130] Cluster the words of the text to be detected and label the element types of the clustered words;
[0131] Cluster the target text's segmented words and label the element types of the clustered segmented words;
[0132] Based on the element type of the cluster segmentation annotation, the cluster segmentation of the to-be-detected text and the cluster segmentation of the target text are semantically matched to obtain the cluster segmentation matching result of the to-be-detected text and the target text;
[0133] Based on the clustering and word segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained.
[0134] In one embodiment of the present disclosure, the text to be detected and the target text are navigation guidance speech texts, and the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text, including:
[0135] Based on the pre-defined navigation guidance speech template, the text to be detected and the target text are segmented respectively to obtain the segmentation of the text to be detected and the segmentation of the target text, and the segmentation is a word or a single character that cannot be divided any further.
[0136] In one embodiment of the present disclosure, the navigation speech template includes at least one of a main action speech template, an auxiliary action speech template, a road direction speech template, and a traffic information speech template.
[0137] In one embodiment of the present disclosure, clustering the segmented words of the text to be detected and labeling the element type of the clustered segmented words includes:
[0138] Determine the element type of the word segmentation of the text to be detected through the pre-defined element dictionary and element rules;
[0139] According to the word order of the segmented words in the text to be detected, adjacent segmented words with the same element type are clustered into a clustered segmented word and the element type is marked.
[0140] In one embodiment of the present disclosure, the element type includes at least one of driving action, driving distance, driving speed, driving time and electronic eye.
[0141] In one embodiment of the present disclosure, based on the element type of the cluster segmentation annotation, the cluster segmentation of the to-be-detected text and the cluster segmentation of the target text are semantically matched to obtain the cluster segmentation matching result of the to-be-detected text and the target text, including:
[0142] The element types and word orders in the clustered word segments of the to-be-detected text and the clustered word segments of the target text meet the pre-established semantic matching rules, and are determined as the clustered word segments that match the to-be-detected text and the target text;
[0143] For the clustered word segments of the text to be detected that cannot be matched with the clustered word segments of the target text through the semantic matching rules, semantic matching is performed through a preset text matching algorithm to obtain the corresponding clustered word segmentation matching results.
[0144] In one embodiment of the present disclosure, based on the cluster segmentation matching results of the text to be detected and the target text, obtaining the semantic detection results of the text to be detected and the target text includes:
[0145] Based on the cluster segmentation matching result, if it is determined that the cluster segmentation of the text to be detected and the cluster segmentation of the target text do not match, then the semantic detection results of the text to be detected and the target text are determined to be semantically different;
[0146] If it is determined that the clustered word segmentation of the text to be detected matches the clustered word segmentation of the target text, or the proportion of the number of mutually matching clustered word segmentations exceeds a preset threshold, then the semantic detection results of the text to be detected and the target text are determined to be semantically indistinguishable;
[0147] If it is determined that the proportion of the number of mutually matching clustered word segments is lower than a preset threshold, the semantic difference of the mutually non-matching clustered word segments is determined according to a predefined semantic difference rule.
[0148] Figure 5 A schematic diagram showing the structure of a computer system suitable for implementing the method according to an embodiment of the present disclosure is shown.
[0149] like Figure 5 As shown, the computer system includes a processing unit, which can perform the various methods in the above-mentioned embodiments according to the program stored in the read-only memory (ROM) or the program loaded from the storage part into the random access memory (RAM). In the RAM, various programs and data required for the operation of the computer system are also stored. The processing unit, ROM and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0150] The following components are connected to the I / O interface: an input part including a keyboard, a mouse, etc.; an output part including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage part including a hard disk, etc.; and a communication part including a network interface card such as a LAN card, a modem, etc. The communication part performs a communication process via a network such as the Internet. The drive is also connected to the I / O interface as needed. Removable media, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on the drive as needed so that the computer program read therefrom is installed into the storage part as needed. Among them, the processing unit can be implemented as a processing unit such as a CPU, a GPU, a TPU, an FPGA, an NPU, etc.
[0151] In particular, according to an embodiment of the present disclosure, the method described above can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly contained on a machine-readable medium, and the computer program includes a program code for executing the above method. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium.
[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0153] The units or modules involved in the embodiments described in the present disclosure may be implemented by software or programmable hardware. The units or modules described may also be set in a processor, and the names of these units or modules do not constitute limitations on the units or modules themselves in some cases.
[0154] As another aspect, the present disclosure further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the electronic device or computer system in the above embodiment; or a computer-readable storage medium that exists independently and is not assembled into a device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the method described in the present disclosure.
[0155] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other.
Claims
1. A text detection method, wherein: include: Get the text to be detected and the target text; Performing word segmentation on the text to be detected and the target text respectively to obtain word segmentation of the text to be detected and word segmentation of the target text; Clustering the word segments of the text to be detected, and marking element types for the clustered word segments; Clustering the word segments of the target text and labeling the element types of the clustered word segments; Based on the element type marked by the cluster segmentation, semantic matching is performed on the cluster segmentation of the text to be detected and the cluster segmentation of the target text to obtain a cluster segmentation matching result between the text to be detected and the target text; Based on the clustering word segmentation matching results of the text to be detected and the target text, the semantic detection results of the text to be detected and the target text are obtained.
2. The text detection method according to claim 1, wherein: The text to be detected and the target text are navigation guidance speech texts, and the word segmentation of the text to be detected and the target text is performed respectively to obtain the word segmentation of the text to be detected and the word segmentation of the target text, including: Based on a pre-defined navigation guidance speech template, the text to be detected and the target text are segmented respectively to obtain the segmented words of the text to be detected and the segmented words of the target text, wherein the segmented words are indivisible words or single characters.
3. The text detection method according to claim 2, wherein: The navigation guidance speech template includes at least one of a main action speech template, an auxiliary action speech template, a road direction speech template, and a traffic information speech template.
4. The text detection method according to any one of claims 1 to 3, wherein: Clustering the word segments of the text to be detected and labeling the element types of the clustered word segments include: Determining the element type of the word segmentation of the text to be detected by using a pre-defined element dictionary and element rules; According to the word order of the segmented words of the to-be-detected text in the to-be-detected text, adjacent segmented words with the same element type are clustered into a clustered segmented word and the element type is marked.
5. The text detection method according to claim 4, wherein: The element type includes at least one of driving action, driving distance, driving speed, driving time and electronic eyes.
6. The text detection method according to claim 4, wherein: Based on the element type marked by the cluster segmentation, semantic matching is performed on the cluster segmentation of the text to be detected and the cluster segmentation of the target text to obtain the cluster segmentation matching result of the text to be detected and the target text, including: The element types and word orders in the clustered word segments of the to-be-detected text and the clustered word segments of the target text meet the pre-established semantic matching rules, and are determined as the clustered word segments in which the to-be-detected text and the target text match each other; For the clustered word segments of the text to be detected and the clustered word segments of the target text that cannot be matched through the semantic matching rule, semantic matching is performed through a preset text matching algorithm to obtain corresponding clustered word segmentation matching results.
7. The text detection method according to claim 6, wherein: The step of obtaining semantic detection results of the text to be detected and the target text based on the cluster segmentation matching results of the text to be detected and the target text includes: Based on the cluster segmentation matching result, if it is determined that the cluster segmentation of the text to be detected and the cluster segmentation of the target text do not match, then it is determined that the semantic detection results of the text to be detected and the target text are that there is a semantic difference; If it is determined that the clustered word segmentation of the to-be-detected text matches the clustered word segmentation of the target text, or the proportion of the number of mutually matching clustered word segmentations exceeds a preset threshold, it is determined that the semantic detection results of the to-be-detected text and the target text are semantically indistinguishable; If it is determined that the proportion of the number of mutually matching clustered word segments is lower than a preset threshold, the semantic difference of the mutually non-matching clustered word segments is determined according to a predefined semantic difference rule.
8. A text detection device, wherein: include: A text acquisition module is configured to acquire the text to be detected and the target text; A word segmentation module is configured to perform word segmentation on the text to be detected and the target text respectively to obtain word segmentations of the text to be detected and the target text; A clustering module to be detected is configured to cluster the word segments of the text to be detected and mark element types for the clustered word segments; A target clustering module is configured to cluster the word segments of the target text and label the element types of the clustered word segments; A semantic matching module is configured to perform semantic matching on the clustered word segmentation of the to-be-detected text and the clustered word segmentation of the target text based on the element type marked by the clustered word segmentation, so as to obtain a clustered word segmentation matching result between the to-be-detected text and the target text; The semantic detection module is configured to obtain semantic detection results of the text to be detected and the target text based on the clustering word segmentation matching results of the text to be detected and the target text.
9. An electronic device, wherein: The method comprises a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions are executed by the processor to implement the method steps described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer instructions stored thereon, wherein: When the computer instructions are executed by a processor, the method steps of any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Sensitive information detection method and device, equipment, storage medium and product
CN121077729A