Text similarity recognition method, device, equipment and storage medium
Converting log text data into image format for similarity analysis addresses inefficiencies in log text recognition, enhancing accuracy and reducing manual effort in system startup assessments.
Patent Information
- Application Number
- CN202210858888.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-21
AI Technical Summary
In the prior art, the similarity recognition efficiency of log text data is low, and it is difficult to deal with the system's sudden garbled code and abnormal foreign language situations, resulting in high false alarm rates and missed alarm rates, and it is impossible to effectively judge the startup status of the application system.
After picturing the text-format startup log, the startup process is determined by calculating the similarity between pictures. The specific steps include obtaining text characterization data, converting it into an encoded string, generating a target log image, performing preprocessing and calculating image similarity.
It improves the similarity recognition efficiency of log text data, can more accurately judge the startup status of the application system, reduce false alarm rates and missed alarm rates, and adapt to the system's sudden garbled code and abnormal foreign language situations.
Smart Images

Figure CN115424284B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method, device, equipment and storage medium for text similarity recognition. Background Art
[0002] During the process of updating the application system version, it is necessary to replace the program code package, restart the application system, and then obtain the startup log of the application system immediately; judge whether the application system starts normally according to the log. The traditional log judgment method requires setting a keyword list for normal startup of the application system and a keyword list for abnormal error reporting, etc. Then, according to the characteristics of the application system, personalized black and white list libraries are set; the rules are recorded in the blacklist library. During the startup process of a certain system, it is necessary to determine that a certain keyword must not appear, or a certain process must start before another process, otherwise it is regarded as a startup failure.
[0003] For the normal grammar keyword recognition rules and the additional black and white list rules related to the specific resource environment, manpower needs to be invested in maintenance daily, and the recognition efficiency of the program will drop significantly, the false alarm rate and missed alarm rate will increase, and it cannot handle the situation of sudden garbled characters or abnormal foreign languages in the system. At the same time, there is a method in the industry to separate and semantically extract log texts, and then recognize semantics and judge startup. Because it involves different selections such as language, common words, word segmentation granularity, interception step length, vector conversion, comparison algorithm, sample library, etc., and is closely combined with the industry attributes, it is difficult to continuously improve the accuracy and generality. Therefore, how to improve the similarity recognition efficiency of log text data has become a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] The main purpose of the present invention is to solve the technical problem of low similarity recognition efficiency of log text data in the prior art by calculating the similarity between pictures after picturing the startup log, the last log and the historical log in text format, and judging whether the current startup process is normal according to the similarity.
[0005] The first aspect of the present invention provides a method for text similarity recognition, including: obtaining real-time log text data to be processed, and determining multi-level text representation data corresponding to the real-time log text data; converting the real-time log text data into a coded string according to the text representation data and a preset coding specification; determining a picture specification according to the coded string and a preset rule, and generating a target log image based on the picture specification; obtaining a historical log image corresponding to historical log text data, respectively preprocessing the target log image and the historical log image to obtain image features and identification text feature data of each image; calculating an image similarity between the target log image and the historical log image according to the image features and the identification text feature data.
[0006] Optionally, in the first implementation manner of the first aspect of the present invention, the determining the multi-level text representation data corresponding to the real-time log text data includes: obtaining the real-time log text data to be processed, extracting features from the real-time log text data based on a preset text encoder to obtain the sentence-level features and word-level features of the real-time log text data; labeling each word in the real-time log text data according to the word-level features; extracting the level information corresponding to the real-time log text data based on a preset regular expression, and determining the multi-level text representation data corresponding to the real-time log text data according to the level information.
[0007] Optionally, in the second implementation manner of the first aspect of the present invention, the converting the real-time log text data into a coded string according to the text representation data and a preset coding specification includes: performing conversion processing on the text representation data based on a preset Unicode character set to obtain an initial code point; determining the number of bytes of the initial code point; and converting the real-time log text data into a coded string according to the number of bytes and the preset coding specification.
[0008] Optionally, in the third implementation manner of the first aspect of the present invention, the determining the picture specification according to the coded string and a preset rule and generating a target log image based on the picture specification includes: converting the coded string into a plurality of RGB color values; determining the picture specification according to the preset rule and the RGB color values; determining the text order in the real-time log text data, arranging the plurality of RGB color values according to the text order and the picture specification to obtain image parameters; and generating the target log image corresponding to the real-time log text data according to the image parameters.
[0009] Optionally, in the fourth implementation manner of the first aspect of the present invention, the obtaining the historical log image corresponding to the historical log text data, and respectively preprocessing the target log image and the historical log image to obtain the image features and the identification text feature data of each image includes: obtaining the images to be compared, where the images to be compared are the historical log image and the target log image, and the historical log image is the corresponding historical log image converted from the historical log text data; respectively performing rotation correction detection on the images to be compared to obtain the images to be compared after angle correction; performing feature extraction on the images to be compared after angle correction to obtain a feature extraction map corresponding to the images to be compared after angle correction; performing target detection on the images to be compared after angle correction according to the feature extraction map to obtain the identification position data corresponding to the images to be compared, and performing text feature extraction on the images to be compared after angle correction to obtain the identification text feature data corresponding to the images to be compared.
[0010] Optionally, in the fifth implementation manner of the first aspect of the present invention, calculating the image similarity between the target log image and the historical log image according to the image feature and the identification text feature data includes: cropping the image to be compared according to the identification position data to obtain a target identification image corresponding to the image to be compared; extracting features from the target identification image to obtain a feature vector of the target identification image; calculating the similarity between the historical log image and the target log image according to the feature vector and the identification text feature data to obtain a similarity comparison result between the historical log image and the target log image.
[0011] Optionally, in the sixth implementation manner of the first aspect of the present invention, calculating the similarity between the historical log image and the target log image according to the feature vector and the identification text feature data to obtain a similarity comparison result between the historical log image and the target log image includes: calculating the feature distance between the target identification images according to the feature vector; determining whether the feature distance is greater than a preset threshold; if so, determining the similarity between the images to be compared according to the identification text feature data, and obtaining the similarity according to the value of the similarity.
[0012] The second aspect of the present invention provides a text similarity recognition device, including: a determination module, configured to obtain real-time log text data to be processed and determine multi-level text representation data corresponding to the real-time log text data; a conversion module, configured to convert the real-time log text data into a coded string according to the text representation data and a preset coding specification; a generation module, configured to determine a picture specification according to the coded string and a preset rule, and generate a target log image based on the picture specification; a preprocessing module, configured to obtain a historical log image corresponding to historical log text data, and respectively preprocess the target log image and the historical log image to obtain image features and identification text feature data of each image; a calculation module, configured to calculate the image similarity between the target log image and the historical log image according to the image features and the identification text feature data.
[0013] Optionally, in the first implementation manner of the second aspect of the present invention, the determination module is specifically configured to: obtain real-time log text data to be processed, extract features from the real-time log text data based on a preset text encoder to obtain sentence-level features and word-level features of the real-time log text data; label each word in the real-time log text data according to the word-level features; extract level information corresponding to the real-time log text data based on a preset regular expression, and determine multi-level text representation data corresponding to the real-time log text data according to the level information.
[0014] Optionally, in the second implementation manner of the second aspect of the present invention, the conversion module includes: a first conversion unit configured to perform conversion processing on the text representation data based on a preset unified code character set to obtain an initial code point; a determination unit configured to determine the number of bytes of the initial code point; and a second conversion unit configured to convert the real-time log text data into a coded string according to the number of bytes and a preset coding specification.
[0015] Optionally, in the third implementation manner of the second aspect of the present invention, the generation module is specifically configured to: convert the coded string into a plurality of RGB color values; determine a picture specification according to a preset rule and the RGB color values; determine the text order in the real-time log text data, and arrange the plurality of RGB color values according to the text order and the picture specification to obtain image parameters; and generate a target log image corresponding to the real-time log text data according to the image parameters.
[0016] Optionally, in the fourth implementation manner of the second aspect of the present invention, the preprocessing module is specifically configured to: obtain an image to be compared, where the image to be compared is a historical log image and a target log image, and the historical log image is a corresponding historical log image converted from historical log text data; respectively perform rotation correction detection on the images to be compared to obtain the images to be compared after angle correction; perform feature extraction on the images to be compared after angle correction to obtain a feature extraction map corresponding to the images to be compared after angle correction; perform target detection on the images to be compared after angle correction according to the feature extraction map to obtain identification position data corresponding to the images to be compared, and perform text feature extraction on the images to be compared after angle correction to obtain identification text feature data corresponding to the images to be compared.
[0017] Optionally, in the fifth implementation manner of the second aspect of the present invention, the calculation module is specifically configured to: crop the image to be compared according to the identification position data to obtain a target identification image corresponding to the image to be compared; perform feature extraction on the target identification image to obtain a feature vector of the target identification image; and calculate the similarity between the historical log image and the target log image according to the feature vector and the identification text feature data to obtain a similarity comparison result between the historical log image and the target log image.
[0018] Optionally, in the sixth implementation manner of the second aspect of the present invention, the calculation module is further specifically configured to: calculate a feature distance between the target identification images according to the feature vector; determine whether the feature distance is greater than a preset threshold; if so, determine the similarity between the images to be compared according to the identification text feature data, and obtain a similarity according to the value of the similarity.
[0019] In a third aspect of the present invention, there is provided a text similarity recognition device, including: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line;
[0020] The at least one processor invokes the instructions in the memory to cause the text similarity recognition device to execute each step of the above-mentioned text similarity recognition method.
[0021] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, in which instructions are stored. When the instructions are run on a computer, the computer is caused to execute each step of the above-mentioned text similarity recognition method.
[0022] In the technical solution provided by the present invention, by determining the text characterization data corresponding to the acquired real-time log data; converting the real-time log data into a coded string according to the text characterization data and a preset coding specification; determining a picture specification according to the coded string and a preset rule, and generating a target log image based on the picture specification; acquiring a historical log image corresponding to historical log data, preprocessing the target log image and the historical log image respectively, and calculating the similarity between the target log image and the historical log image according to the obtained image features and identification character feature data. This solution processes text data into image data in a picture format, calculates the similarity between images, so as to determine whether the current startup process is normal, and solves the technical problem of low efficiency in recognizing the similarity of log data in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Schematic diagram of the first embodiment of the text similarity recognition method provided by the present invention;
[0024] Figure 2 Schematic diagram of the second embodiment of the text similarity recognition method provided by the present invention;
[0025] Figure 3 Schematic diagram of the third embodiment of the text similarity recognition method provided by the present invention;
[0026] Figure 4 Schematic diagram of the fourth embodiment of the text similarity recognition method provided by the present invention;
[0027] Figure 5 Schematic diagram of the fifth embodiment of the text similarity recognition method provided by the present invention;
[0028] Figure 6 Schematic diagram of the first embodiment of the text similarity recognition device provided by the present invention;
[0029] Figure 7 Schematic diagram of the second embodiment of the text similarity recognition device provided by the present invention;
[0030] Figure 8 Schematic diagram of an embodiment of the text similarity recognition device provided by the present invention. Detailed implementation manners
[0031] The embodiments of the present invention provide a text similarity recognition method, device, equipment and storage medium. In the technical solution of the present invention, first, text representation data corresponding to the obtained real-time log data is determined; the real-time log data is converted into a coded string according to the text representation data and a preset coding specification; the picture specification is determined according to the coded string and a preset rule, and a target log image is generated based on the picture specification; a historical log image corresponding to historical log data is obtained, the target log image and the historical log image are respectively preprocessed, and the similarity between the target log image and the historical log image is calculated according to the obtained image features and identification text feature data. This solution performs image processing on text data to obtain image data in a picture format, calculates the similarity between the images, and determines whether the current startup process is normal, thereby solving the technical problem of low efficiency in recognizing the similarity of log data in the prior art.
[0032] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0033] For easy understanding, the specific processes of the embodiments of the present invention are described below. Please refer to Figure 1 , the first embodiment of the text similarity recognition method in the embodiments of the present invention includes:
[0034] 101. Obtain the real-time log text data to be processed, and determine the multi-level text representation data corresponding to the real-time log text data;
[0035] In this embodiment, real-time log text data to be processed is obtained, and multi-level text representation data corresponding to the real-time log text data is determined. Specifically, the server obtains the real-time log text data to be processed, and the real-time log text data to be processed is the text content of the attachment. Among them, the real-time log text data to be processed is the text content in the attachment that needs to be uploaded on the web page.
[0036] It should be noted that before obtaining the real-time log text data to be processed, the server receives a parsing signal, which is used to instruct the server to parse the uploaded attachment, parse the text content of the attachment through the FILE API of JavaScript, and use the parsed result as the real-time log text data to be processed. After receiving the parsing signal, the server also receives an encryption signal, which is used to instruct the server to encrypt the real-time log text data parsed from the front end into a picture and then transmit it to the background.
[0037] In this embodiment, comprehensive text semantic representation plays a crucial role in the text-to-image task. In the embodiment of the present application, the text semantics are represented at multiple levels, including sentence-level, aspect-level, and word-level, and the sentence-level features, aspect-level features, and word-level features of the text sentence are extracted accordingly.
[0038] 102. Convert the real-time log text data into a coded string according to the text representation data and the preset coding specification;
[0039] In this embodiment, the real-time log text data is converted into a coded string according to the text representation data and the preset coding specification.
[0040] Specifically, the server converts the real-time log text data to be processed into an initial coded string according to the preset coding specification. Specifically, the server converts the real-time log text data to be processed into an initial code point according to the preset Unicode character set, and the initial code point is in hexadecimal; the server determines the number of bytes of the initial code point; the server converts the initial code point into an initial coded string according to the number of bytes and the preset coding specification, and the initial coded string is binary.
[0041] In this embodiment, Unicode is an industry standard in the field of computer science, including a character set, an encoding scheme, etc. The encoding range is integers between 0 and 65535. It was developed to address the limitations of traditional character encoding schemes. It sets a unified and unique binary encoding for each character in every language to meet the requirements of text conversion and processing across languages and platforms. Since a computer can only process numbers, if it is to process real-time log text data to be processed, the real-time log text data to be processed must first be converted into numbers before it can be processed. Among them, each character in the real-time log text data to be processed corresponds to an encoded character.
[0042] Currently, characters are arranged in 17 groups, from 0x0000 to 0x10FFFF. Each group is called a Plane, and each Plane has 65536 code positions, for a total of 1114112. However, only a few planes are currently in use. UTF-8, UTF-16, and UTF-32 are all encoding schemes for converting numbers into program data.
[0043] 103. Determine the picture specification according to the encoded string and the preset rules, and generate a target log image based on the picture specification;
[0044] In this embodiment, the picture specification is determined according to the encoded string and the preset rules, and a target log image is generated based on the picture specification. Specifically, the server converts the initial encoded string into multiple RGB color values according to the first preset rule. Specifically, the server converts the initial encoded string into an RGB digital string through JavaScript, and the RGB digital string is in decimal; the server determines three consecutive values in the RGB digital string as the pixel value of an RGB color value; the server determines multiple RGB color values according to the pixel values.
[0045] Further, determine the picture specification according to the second preset rule and multiple RGB color values, and generate a target log image based on the picture specification. The picture specification is used to indicate the number of rows and columns of the RGB color values. The second preset rule stipulates the number of RGB color values included in the same row. Specifically, the server determines the picture specification according to the second preset rule, and the second preset rule stipulates the number of RGB color values included in the same row; the server determines the text order in the target text to obtain a first order; the server arranges the multiple RGB color values in order according to the first order and the picture specification to obtain picture parameters, and the picture parameters include the number of rows and columns of the RGB color values; the server generates a target log image according to the picture parameters.
[0046] 104. Obtain the historical log image corresponding to the historical log text data, and preprocess the target log image and the historical log image respectively to obtain the image features and the identification text feature data of each image;
[0047] In this embodiment, a historical log image corresponding to the historical log text data is obtained, and the target log image and the historical log image are preprocessed respectively to obtain the image features and the identification text feature data of each image. Among them, the preprocessing includes object detection and text feature extraction of the picture.
[0048] Specifically, the object detection refers to identifying the object identifier and the position of the object identifier in the picture. The object identifier can be set in advance according to the business scenario, and specifically can refer to the object to be compared for similarity. For example, when performing the similarity verification of the storefront, the object detection can specifically refer to identifying the storefront and the position of the storefront in the picture. The storefront refers to the plaque and related facilities set at the entrance of an enterprise, institution, or individual business, and is a decorative form outside the door of a store. Another example is that when performing the similarity verification of a person, the object detection can specifically refer to identifying the person and the position of the person in the picture. For illustration, the commonly used object detection frameworks include the object detection framework that extracts classification by region proposal + CNN (Convolutional Neural Network) and the end-to-end object detection framework.
[0049] Among them, the text feature extraction refers to identifying the text feature information in the picture. For example, when performing the similarity verification of the storefront, the text feature extraction can specifically refer to identifying the storefront text in the picture. For illustration, the commonly used text feature extraction method can be to perform text feature extraction using OCR (Optical Character Recognition). The identification text feature data refers to the text information extracted from the picture to be compared by using text feature extraction. The identification position data refers to the position of the object identifier identified from the picture to be compared through object detection. For example, the identification position data can specifically refer to the coordinate information of the object identifier in the picture to be compared.
[0050] 105. Calculate the image similarity between the target log image and the historical log image according to the image features and the identification text feature data.
[0051] In this embodiment, the image similarity between the target log image and the historical log image is calculated according to the image features and the identification text feature data. Specifically, the similarity is used to characterize whether the pictures to be compared are of the same scene. For example, the similarity can specifically indicate that the pictures to be compared represent the same scene. Another example is that the similarity can specifically indicate that the pictures to be compared represent different scenes.
[0052] In this embodiment, the server first calculates the feature distance between the target identification pictures according to the feature vectors, and determines whether the similarity can be obtained by comparing the feature distance with a preset first distance threshold and a preset second distance threshold. When the feature distance is less than the preset first distance threshold, the similarity can be directly obtained. When the feature distance is greater than the preset second distance threshold, the server further obtains the similarity of the picture to be compared according to the identification text feature data. When the feature distance is greater than the preset first distance threshold and less than the preset second distance threshold, the server pushes the picture to be compared to the artificial terminal, and the staff of the artificial terminal compares the picture to be compared.
[0053] In the embodiment of the present invention, the text representation data corresponding to the acquired real-time log data is determined; the real-time log data is converted into a coded string according to the text representation data and a preset coding specification; the picture specification is determined according to the coded string and a preset rule, and a target log image is generated based on the picture specification; the historical log image corresponding to the historical log data is acquired, the target log image and the historical log image are respectively preprocessed, and the similarity between the target log image and the historical log image is calculated according to the obtained image features and identification text feature data. This solution processes the text data into image data in picture format by performing image processing on the text data, calculates the similarity between the images, and determines whether the current startup process is normal, thereby solving the technical problem of low efficiency in identifying the similarity of log data in the prior art.
[0054] Please refer to Figure 2 , the second embodiment of the text similarity recognition method in the embodiment of the present invention includes:
[0055] 201. Acquire the real-time log text data to be processed, and perform feature extraction on the real-time log text data based on a preset text encoder to obtain the sentence-level features and word-level features of the real-time log text data;
[0056] In this embodiment, the real-time log text data to be processed is acquired, and feature extraction is performed on the real-time log text data based on a preset text encoder to obtain the sentence-level features and word-level features of the real-time log text data. Specifically, comprehensive text semantic representation plays a crucial role in the text-to-image task. In the embodiment of the present application, the text semantics are represented at multiple levels, including sentence-level, aspect-level, and word-level, and the sentence-level features, aspect-level features, and word-level features of the text sentence are correspondingly extracted.
[0057] In a specific application scenario, the extracted original sentence-level features can be directly adopted to participate in the subsequent text-to-image processing process. Alternatively, optionally, the Conditioning Augmentation (CA) method can be further used to enhance the extracted sentence-level features (to make the sentence-level representation more accurate), and the enhanced sentence-level features are used to participate in the subsequent text-to-image processing process.
[0058] 202. Annotate each word in the real-time log text data according to the word-level features;
[0059] In this embodiment, each word in the real-time log text data is annotated according to the word-level features. Among them, for the aspect-level features, the aspect-level information of the text sentence can be extracted according to the syntactic structure of the text sentence, and the aspect-level features corresponding to the aspect-level information are further extracted, so as to realize the extraction of the aspect-level features of the text sentence.
[0060] Specifically, tools such as NLTK (natural language toolkit) can be first used to perform part-of-speech tagging on each word in the text sentence, and then regular expressions are used to extract the aspect information contained therein.
[0061] 203. Extract the level information corresponding to the real-time log text data based on a preset regular expression, and determine the multi-level text representation data corresponding to the real-time log text data according to the level information;
[0062] In this embodiment, the level information corresponding to the real-time log text data is extracted based on a preset regular expression, and the multi-level text representation data corresponding to the real-time log text data is determined according to the level information. Specifically, a comprehensive text semantic representation plays a crucial role in the text-to-image task. In the embodiments of the present application, the text semantics are represented at multiple levels, including the sentence-level, aspect-level, and word-level, and the sentence-level features, aspect-level features, and word-level features of the text sentence are correspondingly extracted.
[0063] 204. Perform conversion processing on the text representation data based on a preset unified character set to obtain an initial code point;
[0064] In this embodiment, conversion processing is performed on the text representation data based on a preset unified character set to obtain an initial code point. Specifically, the server converts the target text into an initial code point according to the preset unified character set, and the initial code point is in hexadecimal; the server determines the number of bytes of the initial code point; the server converts the initial code point into an initial encoded string according to the number of bytes and the preset encoding specification, and the initial encoded string is in binary.
[0065] 205. Determine the number of bytes of the initial code point;
[0066] In this embodiment, the number of bytes of the initial code point is determined. Specifically, the Unicode is an industry standard in the field of computer science, including character sets, encoding schemes, etc. The encoding range is an integer between 0 and 65535. It is created to solve the limitations of traditional character encoding schemes. It sets a unified and unique binary encoding for each character in each language to meet the requirements of text conversion and processing across languages and platforms. Because computers can only process numbers, if you want to process the target text, you must first convert the target text into numbers before processing. Among them, each character in the target text corresponds to a coded character.
[0067] 206. Convert the real-time log text data into a coded string according to the number of bytes and the preset coding specification;
[0068] In this embodiment, the real-time log text data is converted into a coded string according to the number of bytes and the preset coding specification. Among them, the current characters are divided into 17 groups, 0x0000 to 0x10FFFF, each group is called a plane, and each plane has 65536 code positions, a total of 1114112. However, only a few planes are currently used. UTF-8, UTF-16, and UTF-32 are all encoding schemes for converting numbers to program data.
[0069] Among them, the characteristic of UTF-8 is that different lengths of encoding are used for characters in different ranges. For characters between 0x00-0x7F, UTF-8 encoding is exactly the same as ASCII encoding. The maximum length of UTF-8 encoding is 4 bytes. The 4-byte template has 21 xs, which means it can accommodate 21 binary digits. The maximum code position 0x10FFFF also has only 21 bits. Take UTF-8 as an example to illustrate. For example, the initial code point of the word "Han" is 0x6C49; 0x6C49 is between 0x0800-0xFFFF, so a 3-byte template is used: 1110xxxx 10xxxxxx 10xxxxxx, with a byte count of 3; convert the initial code point 0x6C49 into a binary target code point: 0110 1100 0100 1001, and use this bit stream to replace the x in the 3-byte template from back to front, and get: 11100110 1011000110001001.
[0070] 207. Determine the image specifications according to the encoding string and the preset rules, and generate a target log image based on the image specifications;
[0071] 208. Obtain the historical log image corresponding to the historical log text data, and preprocess the target log image and the historical log image respectively to obtain the image features and the identification text feature data of each image;
[0072] 209. Calculate the image similarity between the target log image and the historical log image according to the image features and the identification text feature data.
[0073] In this embodiment, steps 207-209 are similar to steps 103-105 in the first embodiment, and will not be elaborated here.
[0074] In the embodiment of the present invention, by determining the text representation data corresponding to the obtained real-time log data; converting the real-time log data into a coded string according to the text representation data and the preset coding specification; determining the picture specification according to the coded string and the preset rule, and generating a target log image based on the picture specification; obtaining the historical log image corresponding to the historical log data, preprocessing the target log image and the historical log image respectively, and calculating the similarity between the target log image and the historical log image according to the obtained image features and the identification text feature data. This solution processes the text data into image data in picture format and calculates the similarity between the images to determine whether the current startup process is normal, solving the technical problem of low similarity recognition efficiency of log data in the prior art.
[0075] Please refer to Figure 3 , the third embodiment of the text similarity recognition method in the embodiment of the present invention includes:
[0076] 301. Obtain the real-time log text data to be processed, and determine the multi-level text representation data corresponding to the real-time log text data;
[0077] 302. Convert the real-time log text data into a coded string according to the text representation data and the preset coding specification;
[0078] 303. Convert the coded string into multiple RGB color values;
[0079] In this embodiment, the coded string is converted into multiple RGB color values. Specifically, the initial coded string is converted into multiple RGB color values according to the first preset rule. Specifically, the server converts the initial coded string into an RGB digital string through JavaScript, and the RGB digital string is in decimal; the server determines three consecutive values in the RGB digital string as the pixel value of an RGB color value according to the first preset rule; the server determines multiple RGB color values according to the pixel value. Among them, the value range of the pixel value is 0-255.
[0080] For example, if the initial encoded string is 11111111 11111010 11111010, and the corresponding RGB digital string is 255 250 250, these three values are determined as the pixel values of an RGB color value. The corresponding RGB color value is R255G250B250, and the corresponding hexadecimal color is #FFFAFA. Another example, if the initial encoded string is 11111000 11111000 11111111, and the corresponding RGB digital string is 248 248 255, these three values are determined as the pixel values of an RGB color value. The corresponding RGB color value is R248G248B255, and the corresponding hexadecimal color value is #F8F8FF. Details are not elaborated here.
[0081] 304. Determine the picture specification according to the preset rules and the RGB color value;
[0082] In this embodiment, the picture specification is determined according to the preset rules and the RGB color value. Specifically, the picture specification is determined according to the second preset rule and multiple RGB color values, and a target picture is generated based on the picture specification. The picture specification is used to indicate the number of rows and columns of the RGB color values. The second preset rule stipulates the number of RGB color values included in the same row. Specifically, the server determines the picture specification according to the second preset rule, and the second preset rule stipulates the number of RGB color values included in the same row.
[0083] 305. Determine the text order in the real-time log text data, and arrange multiple RGB color values according to the text order and the picture specification to obtain image parameters;
[0084] In this embodiment, the text order in the real-time log text data is determined, and multiple RGB color values are arranged according to the text order and the picture specification to obtain image parameters.
[0085] Specifically, the server determines the text order in the target text to obtain the first order; the server arranges multiple RGB color values in sequence according to the first order and the picture specification to obtain picture parameters, and the picture parameters include the number of rows and columns of the RGB color values; the server generates a target picture according to the picture parameters. Among them, the picture parameters include the number of RGB color values in each row and each column of the picture.
[0086] 306. Generate a target log image corresponding to the real-time log text data according to the image parameters;
[0087] In this embodiment, a target log image corresponding to the real-time log text data is generated according to the image parameters. Specifically, in the original order (the first order) of the text in the target text, each text is replaced with the corresponding RGB color value, and a target picture is generated according to the picture parameters after replacement.
[0088] Among them, the server can set the conversion rule according to the preset rule, that is, set the parameter threshold of the image, and adjust the size of the generated target image. For example, it can be stipulated how many RGB color values are included in the same row to generate a regular target image, and the specific details are not elaborated here.
[0089] 307. Obtain the historical log image corresponding to the historical log text data, and preprocess the target log image and the historical log image respectively to obtain the image features and the identification text feature data of each image.
[0090] 308. Calculate the image similarity between the target log image and the historical log image according to the image features and the identification text feature data.
[0091] In this embodiment, steps 301-302, 307-308 are similar to steps 101-102, 104-105 in the first embodiment, and the details are not elaborated here.
[0092] In the embodiment of the present invention, by determining the text representation data corresponding to the obtained real-time log data; converting the real-time log data into a coded string according to the text representation data and the preset coding specification; determining the picture specification according to the coded string and the preset rule, and generating a target log image based on the picture specification; obtaining the historical log image corresponding to the historical log data, preprocessing the target log image and the historical log image respectively, and calculating the similarity between the target log image and the historical log image according to the obtained image features and the identification text feature data. This solution processes the text data into image data in picture format and calculates the similarity between the images to determine whether the current startup process is normal, solving the technical problem of low efficiency in identifying the similarity of log data in the prior art.
[0093] Please refer to Figure 4 , the fourth embodiment of the text similarity recognition method in the embodiment of the present invention includes:
[0094] 401. Obtain the real-time log text data to be processed, and determine the multi-level text representation data corresponding to the real-time log text data.
[0095] 402. Convert the real-time log text data into a coded string according to the text representation data and the preset coding specification.
[0096] 403. Determine the picture specification according to the coded string and the preset rule, and generate a target log image based on the picture specification.
[0097] 404. Obtain the images to be compared, where the images to be compared are the historical log image and the target log image, and the historical log image is the corresponding historical log image converted from the historical log text data.
[0098] In this embodiment, a to-be-compared image is obtained, where the to-be-compared image is a historical log image and a target log image, and the historical log image is the corresponding historical log image converted from historical log text data. Specifically, the historical log text data is obtained. The historical log text data is the text content of the attachment. Among them, the historical log text data is the text content of the attachment that needs to be uploaded on the web page.
[0099] It should be noted that before obtaining the historical log text data, the server receives a parsing signal, which is used to instruct the server to parse the uploaded attachment, parse the text content of the attachment through the FILE API of JavaScript, and use the parsed result as the historical log text data. After receiving the parsing signal, the server also receives an encryption signal, which is used to instruct the server to encrypt the historical log text data parsed from the front end into a picture and then transmit it to the background.
[0100] 405. Rotation correction detection is respectively performed on the to-be-compared images to obtain the to-be-compared images after angle correction;
[0101] In this embodiment, rotation correction detection is respectively performed on the to-be-compared images to obtain the to-be-compared images after angle correction. Specifically, the rotation correction detection refers to detecting whether the to-be-compared picture is a preset normal angle. When the to-be-compared picture is not a preset normal angle, the to-be-compared picture needs to be angle-corrected. For example, rotation correction detection can detect the situations where the to-be-compared picture is a preset normal angle, rotated 90 degrees, rotated 180 degrees, and rotated 270 degrees. When it is detected that the to-be-compared picture is rotated 90 degrees, rotated 180 degrees, or rotated 270 degrees, the to-be-compared picture needs to be angle-corrected to the preset normal angle. The feature extraction map refers to the feature map obtained by performing feature extraction on the to-be-compared picture after angle correction, and is used to characterize the picture features of the to-be-compared picture after angle correction.
[0102] 406. Feature extraction is performed on the to-be-compared image after angle correction to obtain a feature extraction map corresponding to the to-be-compared image after angle correction;
[0103] In this embodiment, feature extraction is performed on the to-be-compared image after angle correction to obtain a feature extraction map corresponding to the to-be-compared image after angle correction. Specifically, the server can use the mobilenetv2 network as a feature extractor to perform feature extraction to obtain a feature extraction map corresponding to the to-be-compared picture after angle correction. After obtaining the feature extraction map, the server will perform target detection through a preset target detection network to obtain the identification position data corresponding to the to-be-compared picture.
[0104] 407. Perform object detection on the image to be compared after angle correction according to the feature extraction map, obtain the identification position data corresponding to the image to be compared, and perform text feature extraction on the image to be compared after angle correction to obtain the identification text feature data corresponding to the image to be compared.
[0105] In this embodiment, object detection is performed on the image to be compared after angle correction according to the feature extraction map to obtain the identification position data corresponding to the image to be compared, and text feature extraction is performed on the image to be compared after angle correction to obtain the identification text feature data corresponding to the image to be compared. Among them, the preset object detection network can specifically be a network composed of multiple convolutional layers and average pooling layers. Each convolutional layer can perform feature extraction on the feature extraction map and output feature maps with different receptive field sizes. By predicting the object position and category on these feature maps with different receptive field sizes, the identification position data corresponding to the image to be compared can be obtained. When performing text feature extraction on the image to be compared after angle correction, the server first inputs the image to be compared after angle correction into the ResNet network for convolution to extract the feature map of the image to be compared after angle correction, and then inputs the feature map into the preset text line detection module and the text line recognition bidirectional LSTM (Long Short-Term Memory) network to obtain the identification text feature data corresponding to the image to be compared.
[0106] 408. Calculate the image similarity between the target log image and the historical log image according to the image features and the identification text feature data.
[0107] Steps 401-403 and 408 in this embodiment are similar to steps 101-103 and 105 in the first embodiment, and will not be elaborated here.
[0108] In the embodiment of the present invention, by determining the text representation data corresponding to the acquired real-time log data; converting the real-time log data into a coded string according to the text representation data and the preset coding specification; determining the picture specification according to the coded string and the preset rule, and generating a target log image based on the picture specification; obtaining the historical log image corresponding to the historical log data, preprocessing the target log image and the historical log image respectively, and calculating the similarity between the target log image and the historical log image according to the obtained image features and identification text feature data. This solution performs image processing on text data to obtain image data in picture format, calculates the similarity between images to determine whether the current startup process is normal, and solves the technical problem of low efficiency in identifying the similarity of log data in the prior art.
[0109] Please refer to Figure 5 , the fifth embodiment of the text similarity recognition method in the embodiment of the present invention includes:
[0110] 501. Obtain real-time log text data to be processed, and determine multi-level text characterization data corresponding to the real-time log text data;
[0111] 502. Convert the real-time log text data into a coded string according to the text characterization data and a preset coding specification;
[0112] 503. Determine a picture specification according to the coded string and a preset rule, and generate a target log image based on the picture specification;
[0113] 504. Obtain a historical log image corresponding to historical log text data, and perform preprocessing on the target log image and the historical log image respectively to obtain image features and identification character feature data of each image;
[0114] 505. Crop the image to be compared according to the identification position data to obtain a target identification image corresponding to the image to be compared;
[0115] In this embodiment, the image to be compared is cropped according to the identification position data to obtain a target identification image corresponding to the image to be compared. Specifically, after obtaining the identification position data, the server will mark the target identification picture in the image to be compared according to the identification position data, crop the image to be compared, and crop the accurate target identification picture from the image to be compared. For example, when performing the similarity audit of the storefront, after obtaining the storefront position information, the server will crop the image to be compared according to the storefront position information and crop the storefront photo from the image to be compared. In this way, an accurate target identification picture can be cropped from the image to be compared, so that the similarity comparison can be realized according to the accurate target identification picture, reducing the interference of other picture features unrelated to the target identification picture in the image to be compared on the similarity comparison, and facilitating obtaining an accurate similarity comparison result.
[0116] 506. Extract features from the target identification image to obtain a feature vector of the target identification image;
[0117] In this embodiment, features are extracted from the target identification image to obtain a feature vector of the target identification image. Among them, features are extracted from the target identification picture to obtain a feature vector, and the target identification picture is classified according to the feature vector. In the process of classifying the target identification picture, the trained classification model will first extract features from the target identification picture multiple times through a multi-layer network to obtain a feature vector of the target identification picture, and then classify the target identification picture according to the feature vector. The feature vector of the target identification picture refers to the vector used to represent the picture features of the target identification picture.
[0118] 507. Calculate the feature distance between the target identification images according to the feature vector;
[0119] In this embodiment, the feature distance between target identification images is calculated according to the feature vectors. Specifically, the server first calculates the feature distance between target identification pictures according to the feature vectors, and determines whether a similarity comparison result can be obtained by comparing the feature distance with a preset first distance threshold and a preset second distance threshold. When the feature distance is less than the preset first distance threshold, the similarity comparison result can be directly obtained. When the feature distance is greater than the preset second distance threshold, the server further obtains the similarity comparison result of the picture to be compared according to the target identification text feature information.
[0120] 508. Determine whether the feature distance is greater than a preset threshold;
[0121] In this embodiment, it is determined whether the feature distance is greater than a preset threshold. The server first calculates the feature distance between target identification pictures according to the feature vectors, and determines whether similarity can be obtained by comparing the feature distance with a preset first distance threshold and a preset second distance threshold.
[0122] 509. If so, determine the similarity between the pictures to be compared according to the identification text feature data, and obtain the similarity according to the value of the similarity.
[0123] In this embodiment, if the feature distance is greater than the preset threshold, the similarity between the pictures to be compared is determined according to the identification text feature data, and the similarity comparison result is obtained according to the value of the similarity.
[0124] Steps 501 - 504 in this embodiment are similar to 101 - 105 in the first embodiment, and will not be elaborated here.
[0125] In the embodiment of the present invention, by determining the text representation data corresponding to the acquired real - time log data; converting the real - time log data into a coded string according to the text representation data and a preset coding specification; determining the picture specification according to the coded string and a preset rule, and generating a target log image based on the picture specification; acquiring the historical log image corresponding to the historical log data, pre - processing the target log image and the historical log image respectively, and calculating the similarity between the target log image and the historical log image according to the obtained image features and identification text feature data. This solution processes text data into image data in picture format, calculates the similarity between images to determine whether the current startup process is normal, and solves the technical problem of low efficiency in identifying the similarity of log data in the prior art.
[0126] The text similarity recognition method in the embodiment of the present invention is described above. Next, the text similarity recognition device in the embodiment of the present invention will be described. Please refer to Figure 6 , the first embodiment of the text similarity recognition device in the embodiment of the present invention includes:
[0127] A determination module 601, configured to obtain real-time log text data to be processed and determine multi-level text representation data corresponding to the real-time log text data;
[0128] A conversion module 602, configured to convert the real-time log text data into a coded string according to the text representation data and a preset coding specification;
[0129] A generation module 603, configured to determine a picture specification according to the coded string and a preset rule, and generate a target log image based on the picture specification;
[0130] A preprocessing module 604, configured to obtain a historical log image corresponding to historical log text data, preprocess the target log image and the historical log image respectively, and obtain image features and identification text feature data of each image;
[0131] A calculation module 605, configured to calculate an image similarity between the target log image and the historical log image according to the image features and the identification text feature data.
[0132] In an embodiment of the present invention, by determining text representation data corresponding to the obtained real-time log data; converting the real-time log data into a coded string according to the text representation data and a preset coding specification; determining a picture specification according to the coded string and a preset rule, and generating a target log image based on the picture specification; obtaining a historical log image corresponding to historical log data, preprocessing the target log image and the historical log image respectively, and calculating a similarity between the target log image and the historical log image according to the obtained image features and identification text feature data. This solution performs image processing on text data to obtain image data in a picture format, calculates the similarity between images, so as to determine whether the current startup process is normal, and solves the technical problem of low efficiency in recognizing the similarity of log data in the prior art.
[0133] Please refer to Figure 7 , a second embodiment of the text similarity recognition device in the embodiment of the present invention. The text similarity recognition device specifically includes:
[0134] A determination module 601, configured to obtain real-time log text data to be processed and determine multi-level text representation data corresponding to the real-time log text data;
[0135] A conversion module 602, configured to convert the real-time log text data into a coded string according to the text representation data and a preset coding specification;
[0136] A generation module 603, configured to determine a picture specification according to the encoded string and a preset rule, and generate a target log image based on the picture specification;
[0137] A preprocessing module 604, configured to obtain historical log image corresponding to historical log text data, and perform preprocessing on the target log image and the historical log image respectively to obtain image features and identification text feature data of each image;
[0138] A calculation module 605, configured to calculate an image similarity between the target log image and the historical log image according to the image features and the identification text feature data.
[0139] In this embodiment, the determination module 601 is specifically configured to:
[0140] Obtain real-time log text data to be processed, perform feature extraction on the real-time log text data based on a preset text encoder to obtain sentence-level features and word-level features of the real-time log text data;
[0141] Label each word in the real-time log text data according to the word-level features;
[0142] Extract level information corresponding to the real-time log text data based on a preset regular expression, and determine multi-level text representation data corresponding to the real-time log text data according to the level information.
[0143] In this embodiment, the conversion module 602 includes:
[0144] A first conversion unit 6021, configured to perform conversion processing on the text representation data based on a preset unified character set to obtain an initial code point;
[0145] A determination unit 6022, configured to determine the number of bytes of the initial code point;
[0146] A second conversion unit 6023, configured to convert the real-time log text data into an encoded string according to the number of bytes and a preset encoding specification.
[0147] In this embodiment, the generation module 603 is specifically configured to:
[0148] Convert the encoded string into a plurality of RGB color values;
[0149] Determine a picture specification according to a preset rule and the RGB color values;
[0150] Determine the text order in the real-time log text data, and arrange the plurality of RGB color values according to the text order and the picture specification to obtain image parameters;
[0151] Generate a target log image corresponding to the real-time log text data according to the image parameters.
[0152] In this embodiment, the preprocessing module 604 is specifically configured to:
[0153] Obtain the images to be compared, where the images to be compared are historical log images and target log images, and the historical log images are corresponding historical log images converted from historical log text data;
[0154] Perform rotation correction detection on the images to be compared respectively to obtain the images to be compared after angle correction;
[0155] Extract features from the images to be compared after angle correction to obtain a feature extraction map corresponding to the images to be compared after angle correction;
[0156] Perform target detection on the images to be compared after angle correction according to the feature extraction map to obtain identification position data corresponding to the images to be compared, and perform text feature extraction on the images to be compared after angle correction to obtain identification text feature data corresponding to the images to be compared.
[0157] In this embodiment, the calculation module 605 is specifically configured to:
[0158] Crop the images to be compared according to the identification position data to obtain target identification images corresponding to the images to be compared;
[0159] Extract features from the target identification images to obtain feature vectors of the target identification images;
[0160] Calculate the similarity between the historical log images and the target log images according to the feature vectors and the identification text feature data to obtain a similarity comparison result between the historical log images and the target log images.
[0161] In this embodiment, the calculation module 605 is specifically further configured to:
[0162] Calculate the feature distance between the target identification images according to the feature vectors;
[0163] Determine whether the feature distance is greater than a preset threshold;
[0164] If so, determine the similarity between the images to be compared according to the identification text feature data, and obtain the similarity according to the value of the similarity.
[0165] In an embodiment of the present invention, text characterization data corresponding to the acquired real-time log data is determined; the real-time log data is converted into a coded string according to the text characterization data and a preset coding specification; a picture specification is determined according to the coded string and a preset rule, and a target log image is generated based on the picture specification; a historical log image corresponding to historical log data is acquired, the target log image and the historical log image are respectively preprocessed, and the similarity between the target log image and the historical log image is calculated according to the obtained image features and identification character feature data. This solution performs image processing on text data to obtain image data in a picture format, and calculates the similarity between the images to determine whether the current startup process is normal, solving the technical problem of low efficiency in identifying the similarity of log data in the prior art.
[0166] Above Figure 6 And Figure 7 The text similarity recognition device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. Below, the text similarity recognition device in the embodiment of the present invention is described in detail from the perspective of hardware processing.
[0167] Figure 8 FIG. is a schematic structural diagram of a text similarity recognition device provided by an embodiment of the present invention. The text similarity recognition device 800 may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) 810 (for example, one or more processors) and a memory 820, and one or more storage media 830 (for example, one or more mass storage devices) storing application programs 833 or data 832. Among them, the memory 820 and the storage media 830 may be transient storage or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the text similarity recognition device 800. Further, the processor 810 may be configured to communicate with the storage media 830 and execute a series of instruction operations in the storage media 830 on the text similarity recognition device 800 to implement the steps of the text similarity recognition method provided in the above method embodiments.
[0168] The text similarity recognition device 800 may further include one or more power supplies 840, one or more wired or wireless network interfaces 850, one or more input / output interfaces 860, and / or one or more operating systems 831, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand, Figure 8The structure of the text similarity recognition device shown does not limit the text similarity recognition device provided by this application. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0169] The present invention also provides a computer-readable storage medium. This computer-readable storage medium can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the steps of the above-mentioned text similarity recognition method.
[0170] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0172] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A text similarity recognition method, characterized in that, The described text similarity recognition method includes: Obtain real-time log text data to be processed, and determine multi-level text representation data corresponding to the real-time log text data; Convert the real-time log text data into a coded string according to the text representation data and a preset coding specification; Determine a picture specification according to the coded string and a preset rule, and generate a target log image based on the picture specification; Obtain a historical log image corresponding to historical log text data, and perform preprocessing on the target log image and the historical log image respectively to obtain image features and identification text feature data of each image; Calculate the image similarity between the target log image and the historical log image according to the image features and the identification text feature data; The determining of the multi-level text representation data corresponding to the real-time log text data includes: Obtain real-time log text data to be processed, and perform feature extraction on the real-time log text data based on a preset text encoder to obtain sentence-level features and word-level features of the real-time log text data; Annotate each word in the real-time log text data according to the word-level features; Extract level information corresponding to the real-time log text data based on a preset regular expression, and determine multi-level text representation data corresponding to the real-time log text data according to the level information; The converting of the real-time log text data into a coded string according to the text representation data and a preset coding specification includes: Perform conversion processing on the text representation data based on a preset Unicode character set to obtain an initial code point; Determine the number of bytes of the initial code point; Convert the real-time log text data into a coded string according to the number of bytes and a preset coding specification; The determining of the picture specification according to the coded string and a preset rule, and the generating of a target log image based on the picture specification includes: Convert the coded string into multiple RGB color values; Determine a picture specification according to a preset rule and the RGB color values; Determine the text order in the real-time log text data, and arrange the multiple RGB color values according to the text order and the picture specification to obtain image parameters; Generate a target log image corresponding to the real-time log text data according to the image parameters.
2. The text similarity recognition method according to claim 1, characterized in that The obtaining of a historical log image corresponding to historical log text data, and the performing of preprocessing on the target log image and the historical log image respectively to obtain image features and identification text feature data of each image includes: Obtain an image to be compared, where the image to be compared is a historical log image and a target log image, and the historical log image is a corresponding historical log image converted from historical log text data; Perform rotation correction detection on the image to be compared respectively to obtain the image to be compared after angle correction; Perform feature extraction on the image to be compared after angle correction to obtain a feature extraction map corresponding to the image to be compared after angle correction; Perform object detection on the angle-corrected image to be compared according to the feature extraction diagram, obtain the identification position data corresponding to the image to be compared, and perform text feature extraction on the angle-corrected image to be compared to obtain the identification text feature data corresponding to the image to be compared.
3. The text similarity recognition method according to claim 2, wherein, Calculating the image similarity between the target log image and the historical log image according to the image feature and the identification text feature data includes: Crop the image to be compared according to the identification position data to obtain a target identification image corresponding to the image to be compared; Extract features from the target identification image to obtain a feature vector of the target identification image; Calculate the similarity between the historical log image and the target log image according to the feature vector and the identification text feature data to obtain the similarity comparison result between the historical log image and the target log image.
4. The text similarity recognition method according to claim 3, wherein Calculating the similarity between the historical log image and the target log image according to the feature vector and the identification text feature data to obtain the similarity comparison result between the historical log image and the target log image includes: Calculate the feature distance between the target identification images according to the feature vector; Determine whether the feature distance is greater than a preset threshold; If so, determine the similarity between the images to be compared according to the identification text feature data, and obtain the similarity according to the value of the similarity.
5. A text similarity recognition device, characterized in that, The text similarity recognition device includes: A determination module for obtaining real-time log text data to be processed and determining multi-level text representation data corresponding to the real-time log text data; A conversion module for converting the real-time log text data into a coded string according to the text representation data and a preset coding specification; A generation module for determining a picture specification according to the coded string and a preset rule, and generating a target log image based on the picture specification; A preprocessing module for obtaining a historical log image corresponding to historical log text data, and preprocessing the target log image and the historical log image respectively to obtain the image feature and the identification text feature data of each image; A calculation module for calculating the image similarity between the target log image and the historical log image according to the image feature and the identification text feature data; The determination module is specifically configured to: obtain real-time log text data to be processed, perform feature extraction on the real-time log text data based on a preset text encoder to obtain the sentence-level feature and the word-level feature of the real-time log text data; label each word in the real-time log text data according to the word-level feature; extract the level information corresponding to the real-time log text data based on a preset regular expression, and determine the multi-level text representation data corresponding to the real-time log text data according to the level information. The conversion module includes: a first conversion unit for performing conversion processing on the text representation data based on a preset unified character set to obtain an initial code point; a determination unit for determining the number of bytes of the initial code point; a second conversion unit for converting the real-time log text data into a coded string according to the number of bytes and a preset coding specification; The generation module is specifically configured to: convert the coded string into a plurality of RGB color values; determine a picture specification according to a preset rule and the RGB color values; determine the text order in the real-time log text data, and arrange the plurality of RGB color values according to the text order and the picture specification to obtain image parameters; generate a target log image corresponding to the real-time log text data according to the image parameters.
6. The text similarity recognition device according to claim 5, characterized in that, The preprocessing module is specifically configured to: obtain an image to be compared, where the image to be compared is a historical log image and a target log image, and the historical log image is a corresponding historical log image converted from historical log text data; respectively perform rotation correction detection on the image to be compared to obtain the image to be compared after angle correction; perform feature extraction on the image to be compared after angle correction to obtain a feature extraction map corresponding to the image to be compared after angle correction; perform target detection on the image to be compared after angle correction according to the feature extraction map to obtain identification position data corresponding to the image to be compared, and perform text feature extraction on the image to be compared after angle correction to obtain identification text feature data corresponding to the image to be compared.
7. The text similarity recognition device according to claim 6, wherein The calculation module is specifically configured to: crop the image to be compared according to the identification position data to obtain a target identification image corresponding to the image to be compared picture; perform feature extraction on the target identification image to obtain a feature vector of the target identification image; calculate the similarity between the historical log image and the target log image according to the feature vector and the identification text feature data to obtain a similarity comparison result between the historical log image and the target log image.
8. The text similarity recognition device according to claim 7, wherein The calculation module is specifically further configured to: calculate the feature distance between the target identification images according to the feature vector; determine whether the feature distance is greater than a preset threshold; if so, determine the similarity between the images to be compared according to the identification text feature data, and obtain the similarity according to the value of the similarity.
9. A text similarity recognition device, characterized in that, The text similarity recognition device includes: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; The at least one processor calls the instructions in the memory so that the text similarity recognition device executes each step of the text similarity recognition method according to any one of claims 1-4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it realizes each step of the text similarity recognition method according to any one of claims 1-4.
Citation Information
Patent Citations
Artificial intelligence-based text conversion method, device, equipment and storage medium
CN110515892A
Discovery of semantic similarities between images and text
US20170061250A1